DNA modifying enzymes and active fragments and variants thereof and methods of use thereof

Novel adenine deaminases and fusion proteins with RNA-guided nucleases enable efficient targeted editing of DNA, addressing inefficiencies in existing methods by introducing specific mutations, enhancing genetic correction and trait introduction in cells.

JP7719172B2Active Publication Date: 2025-08-05LIFEEDIT THERAPEUTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023516171
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-08
Filing Date
2021-09-10
Publication Date
2025-08-05
Estimated Expiration
2041-09-10

AI Technical Summary

Technical Problem

Existing genome editing methods, such as meganucleases, zinc finger fusion proteins, and TALENs, require the creation of chimeric nucleases for each target sequence, which is costly and inefficient, while RNA-guided nucleases (RGNs) face limitations in efficiently introducing targeted mutations in DNA.

Method used

Development of novel adenine deaminases and fusion proteins comprising RNA-guided nucleases and DNA-binding polypeptides, which can introduce targeted mutations by deaminating adenine to inosine, allowing polymerases to incorporate cytosine, thereby changing A:T to G:C base pairs.

Benefits of technology

The novel adenine deaminases enable efficient and targeted editing of DNA in vitro, ex vivo, and in vivo, facilitating genetic corrections and introducing beneficial traits in mammalian and plant cells, with high specificity and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007719172000001
    Figure 0007719172000001
  • Figure 0007719172000002
    Figure 0007719172000002
  • Figure 0007719172000003
    Figure 0007719172000003
Patent Text Reader

Abstract

Compositions and methods are provided that include novel deaminase polypeptides for targeted editing of nucleic acids. The compositions include the deaminase polypeptides. Fusion proteins are also provided that include the DNA-binding polypeptides of the present invention and the deaminase. The fusion proteins include an RNA-guided nuclease fused to the deaminase, optionally complexed with a guide RNA. The compositions also include nucleic acid molecules encoding the deaminase or the fusion protein. Vectors and host cells are also provided that include nucleic acid molecules encoding the deaminase or the fusion protein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 077,089, filed September 11, 2020, and U.S. Provisional Patent Application No. 63 / 146,840, filed February 8, 2021. Each provisional patent application is incorporated herein by reference in its entirety.

[0002] Sequence Listing Description The Sequence Listing associated with this application is provided in ASCII format in lieu of a paper copy and is incorporated herein by reference. The ASCII copy is named L103438_1230WO_0108_1_SL.txt, is 1,071,246 bytes in size, was created on September 9, 2021, and has been submitted electronically via EFS-Web.

[0003] The present invention relates to the fields of molecular biology and gene editing. [Background technology]

[0004] Targeted genome editing or modification is rapidly becoming an important tool in basic and applied research. Early methods involved engineering nucleases, such as meganucleases, zinc finger fusion proteins, or TALENs, which required the creation of chimeric nucleases with engineered, programmable, sequence-specific DNA-binding domains specific to each specific target sequence. RNA-guided nucleases (RGNs), such as the CRISPR-associated (Cas) proteins of the clustered regularly interspaced short palindromic repeats (CRISPR)-Cas bacterial system, can target specific sequences by complexing the nuclease with a guide RNA that specifically hybridizes to the specific target sequence. Creating target-specific guide RNAs is more cost-effective and efficient than creating chimeric nucleases for each target sequence. Such RNA-guided nucleases can be used to edit genomes through the introduction of sequence-specific double-strand breaks that are repaired by error-prone non-homologous end joining (NHEJ), introducing mutations at specific genomic locations.

[0005] In addition, RGNs are useful for targeted DNA editing approaches. Targeted editing of nucleic acid sequences (e.g., targeted cleavage that allows specific modifications to be introduced into genomic DNA) enables highly nuanced approaches to studying gene function and gene expression. RGNs can also be used to create chimeric proteins that use the RNA-guided activity of RGNs in combination with DNA-modifying enzymes, such as deaminases, for targeted base editing. Targeted editing can also be used to target genetic diseases in humans or to introduce agriculturally beneficial mutations into crop genomes. The development of genome editing tools provides new approaches to mammalian therapeutics and agricultural biotechnology based on gene editing. Summary of the Invention

[0006] Compositions and methods for modifying a target DNA molecule are provided. The compositions are used to modify a target DNA molecule of interest. The provided compositions comprise a deaminase polypeptide. Fusion proteins comprising a nucleic acid molecule-binding polypeptide (e.g., a DNA-binding polypeptide) and a deaminase polypeptide, as well as ribonucleoprotein complexes comprising a fusion protein comprising an RNA-guided nuclease and a deaminase polypeptide and a ribonucleic acid, are also provided. The provided compositions also include nucleic acid molecules encoding the deaminase polypeptide or fusion protein, as well as vectors and host cells comprising the nucleic acid molecules. The methods disclosed herein relate to binding of a target sequence of interest within a target DNA molecule of interest and modification of the target DNA molecule of interest. DETAILED DESCRIPTION OF THE INVENTION

[0007] Many modifications and other embodiments of the inventions described herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions. It is to be understood, therefore, that the invention is not to be limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0008] I. Overview The present disclosure provides novel adenine deaminases and fusion proteins comprising a nucleic acid molecule-binding polypeptide, such as a DNA-binding polypeptide, and a novel deaminase polypeptide. In certain embodiments, the DNA-binding polypeptide is a sequence-specific DNA-binding polypeptide in that it binds to a target sequence more frequently than to randomized background sequences. In some embodiments, the DNA-binding polypeptide is or is derived from a meganuclease, a zinc finger fusion protein, or a TALEN. In some embodiments, the fusion protein comprises an RNA-guided DNA-binding polypeptide and a deaminase polypeptide. In some embodiments, the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease, such as a Cas9 polypeptide domain, that binds to a guide RNA (also referred to as gRNA) and then binds to the target nucleic acid sequence via strand hybridization.

[0009] The deaminase polypeptides disclosed herein can deaminate nucleic acid bases, such as adenine. Deamination of nucleic acid bases by deaminases can generate point mutations at the respective residues, which is referred to herein as "nucleic acid editing" or "base editing." Thus, fusion proteins comprising an RNA-guided nuclease (RGN) polypeptide and a deaminase can be used for targeted editing of nucleic acid sequences.

[0010] Such fusion proteins are useful for targeted editing of DNA in vitro, e.g., for generating genetically modified cells. These genetically modified cells can be plant or animal cells. Such fusion proteins can also be useful for introducing targeted mutations, e.g., correcting genetic defects in mammalian cells ex vivo (e.g., in cells obtained from a subject and then reintroduced into the same or another subject), and for introducing targeted mutations (e.g., correcting genetic defects in disease-associated genes in mammalian subjects or introducing inactivating mutations in those genes). Such fusion proteins can also be useful for introducing targeted mutations in plant cells, e.g., for introducing beneficial or agronomically valuable traits or alleles.

[0011] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of multiple amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids in length. A protein, peptide, or polypeptide can refer to an individual protein or a population of proteins. One or more amino acids within a protein, peptide, or polypeptide can be modified, for example, by adding a chemical feature (e.g., a hydrocarbon group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc.). A protein, peptide, or polypeptide can be simply a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, synthetic, or any combination thereof.

[0012] Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced through recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0013] II. Deaminases The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. The deaminase of the present invention is a nucleobase deaminase, and the terms "deaminase" and "nucleobase deaminase" are used interchangeably herein. The deaminase can be a naturally occurring deaminase enzyme or an active fragment or variant thereof. The deaminase can be active on single-stranded nucleic acids, such as ssDNA or ssRNA, or double-stranded nucleic acids, such as dsDNA or dsRNA. In some embodiments, the deaminase can only deaminate ssDNA and does not act on dsDNA.

[0014] The disclosed methods and compositions include adenine deaminase. In some embodiments, the deaminase is an ADAT family deaminase or a variant thereof. Deamination of adenine, adenosine, or deoxyadenosine produces inosine, which is processed by polymerases to guanine. To date, no naturally occurring adenine deaminase is known to deaminate adenine in DNA. Several methods have been employed to evolve and optimize adenine deaminase (ADAT) proteins acting on tRNA in mammalian cells to act on DNA molecules (Gaudelli et al., 2017; Koblan, L. Wet et al., 2018, Nat Biotechnol 36, 843-846; Richter, M. F. et al., 2020, Nat Biotechnol, doi:10.1038 / s41587-020-0562-8, each of which is incorporated herein by reference in its entirety). One such method uses a bacterial selection assay in which only cells capable of activating antibiotic resistance through an A:T>G:C transition can survive.

[0015] The present invention relates to novel adenine deaminase polypeptides generated through the evolution and optimization of bacterial deaminases. The novel adenine deaminases are disclosed herein and set forth as SEQ ID NOS: 1-10 and 399-441. The deaminases of the present invention can be used for editing DNA or RNA molecules. In some embodiments, the deaminases of the present invention can be used for editing ssDNA or ssRNA molecules. The adenine deaminases described herein are useful as deaminases alone or as components in fusion proteins. Fusion proteins comprising a DNA-targeting polypeptide and an adenine deaminase polypeptide, referred to herein as "A base editors," "adenine base editors," or "ABEs," can be used for targeted editing of nucleic acid sequences.

[0016] A "base editor" is a fusion protein comprising a DNA-targeting polypeptide, such as RGN, and a deaminase. An adenine base editor (ABE) comprises a DNA-targeting protein, such as RGN, and an adenine deaminase. ABEs function by deaminating adenine in DNA target molecules, converting it to inosine (Gaudelli, NM et al. 2017). Inosine is recognized as guanine by polymerases, allowing for the incorporation of cytosine on the complementary DNA strand opposite the inosine. After a round of replication following deamination, an A:T to G:C base pair change occurs in the genome. In some embodiments, the disclosed adenine deaminases or active variants or fragments thereof introduce an A>N mutation into a DNA molecule, where N is C, G, or T. In further embodiments, they introduce an A>G mutation into a DNA molecule.

[0017] In embodiments in which the deaminase is targeted to a specific region of a nucleic acid molecule via fusion with a DNA-binding polypeptide, the mutation rate of adenines within or adjacent to the target sequence bound by the DNA-binding polypeptide can be measured using any method known in the art, including polymerase chain reaction (PCR), restriction fragment length polymorphism (RFLP), or DNA sequencing.

[0018] To increase the efficiency of introducing the desired A>G mutation into a target DNA molecule, the novel deaminases of the present disclosure, or active variants or fragments thereof that retain deaminase activity, may be introduced into a cell as part of a deaminase-DNA binding polypeptide fusion and / or co-expressed with a DNA binding polypeptide-deaminase fusion. The deaminases of the present disclosure have the amino acid sequence of any of SEQ ID NOS: 1-10 and 399-441, or variants or fragments thereof that retain deaminase activity. In some embodiments, the deaminase has an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of any one of SEQ ID NOs: 1-10 and 399-441. In certain embodiments, the deaminase comprises an amino acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441. In some embodiments, the deaminase comprises an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 407. For example, the deaminase comprises an amino acid sequence having at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to SEQ ID NO: 407.In some embodiments, the deaminase comprises an amino acid sequence having at least 80% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.5% identity, or at least 99.9% identity to SEQ ID NO: 407. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 407. In some embodiments, the deaminase comprises an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 399. For example, the deaminase comprises an amino acid sequence having at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to SEQ ID NO: 399. In some embodiments, the deaminase comprises an amino acid sequence having at least 80% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.5% identity, or at least 99.9% identity to SEQ ID NO: 399. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 399. In some embodiments, the deaminase comprises an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 405. For example, the deaminase comprises an amino acid sequence having at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to SEQ ID NO: 405.In some embodiments, the deaminase comprises an amino acid sequence having at least 80% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.5% identity, or at least 99.9% identity to SEQ ID NO: 405. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:405.

[0019] III. Nucleic Acid Molecule-Binding Polypeptides Some aspects of the present disclosure provide fusion proteins comprising a nucleic acid molecule-binding polypeptide and a deaminase polypeptide. While binding to and targeted editing of RNA molecules is contemplated by the present invention, in some embodiments, the nucleic acid molecule-binding polypeptide of the fusion protein is a DNA-binding polypeptide. Such fusion proteins are useful for targeted editing of DNA in vitro, ex vivo, or in vivo. These novel fusion proteins are active in mammalian cells and are useful for targeted editing of DNA molecules.

[0020] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains derived from at least two different proteins. A fusion protein may comprise two or more different domains, for example, a DNA-binding domain and a deaminase. In some embodiments, the fusion protein is complexed with or associated with a nucleic acid, for example, RNA.

[0021] In some embodiments, a fusion protein of the present disclosure comprises a DNA-binding polypeptide. As used herein, the term "DNA-binding polypeptide" refers to any polypeptide capable of binding to DNA. In certain embodiments, the DNA-binding polypeptide portion of a fusion protein of the present disclosure binds to double-stranded DNA. In certain embodiments, the DNA-binding polypeptide binds to DNA in a sequence-specific manner. As used herein, the terms "sequence-specific" or "sequence-specifically" refer to selective interaction with a specific nucleotide sequence.

[0022] Two polynucleotide sequences can be considered substantially complementary if the two sequences hybridize to each other under stringent conditions. Similarly, a DNA-binding polypeptide is considered to bind to a particular target sequence in a sequence-specific manner if it binds to that sequence under stringent conditions. By "stringent conditions" or "stringent hybridization conditions" is intended conditions under which two polynucleotide sequences (or polypeptides binding to their specific target sequences) bind to each other at a detectably higher degree (e.g., at least twice background) than to other sequences. Stringent conditions are sequence-dependent and will vary in various circumstances. Typically, stringent conditions are conditions in which the salt concentration is less than about 1.5 M Na ion, typically about 0.01 to 1.0 M Na ion (or other salt) at pH 7.0 to 8.3, and the temperature is at least about 30°C for short sequences (e.g., 10 to 50 nucleotides) and at least about 60°C for long sequences (e.g., more than 50 nucleotides). Stringent conditions can also be achieved by adding destabilizing agents such as formamide. Exemplary low stringency conditions include hybridization using a buffer of 30-35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate) at 37°C, followed by washing with 1x to 2x SSC (20x SSC = 3.0 M NaCl / 0.3 M trisodium citrate) at 50-55°C. Exemplary medium stringency conditions include hybridization using 40-45% formamide, 1.0 M NaCl, and 1% SDS at 37°C, followed by washing with 0.5x to 1x SSC at 55-60°C. Exemplary high stringency conditions include hybridization using 50% formamide, 1 M NaCl, and 1% SDS at 37°C, followed by washing with 0.1x SSC at 60-65°C. Optionally, the wash buffer may contain about 0.1% to about 1% SDS. The duration of hybridization is generally less than about 24 hours, usually about 4 to about 12 hours. The duration of the wash period is at least long enough to reach equilibrium.

[0023] Tm is the temperature (under defined ionic strength and pH) at which 50% of a complementary target sequence hybridizes to a perfectly matched sequence. For DNA-DNA hybrids, Tm can be estimated using the formula of Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5°C + 16.6 (log M) + 0.41 (% GC) - 0.61 (% form) - 500 / L, where M is the molar concentration of monovalent cations, % GC is the percentage of guanosine and cytosine nucleotides in the DNA, % form is the percentage of formamide in the hybridization solution, and L is the hybrid length in base pairs. Generally, stringent conditions are selected to be approximately 5°C lower than the melting point (Tm) of a specific sequence and its complement at a defined ionic strength and pH. However, very stringent conditions can utilize hybridization and / or wash temperatures that are 1, 2, 3, or 4° C. below the melting temperature (Tm), moderately stringent conditions can utilize hybridization and / or wash temperatures that are 6, 7, 8, 9, or 10° C. below the melting temperature (Tm), and low stringency conditions can utilize hybridization and / or wash temperatures that are 11, 12, 13, 14, 15, or 20° C. below the melting temperature (Tm). Those of skill in the art will understand that variations in stringency of hybridization and / or wash solutions are inherently described using the above formula, hybridization and wash compositions, and desired Tm. Extensive guides to nucleic acid hybridization can be found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York); and Ausubel et al., eds. (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York).See Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2nd ed., Cold Spring Harbor Laboratory Press, Plainview, New York).

[0024] In certain embodiments, the sequence-specific DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide (RGDBP). As used herein, "RNA-guided DNA-binding polypeptide" and "RGDBP" refer to a polypeptide that can bind to DNA through hybridization of an associated RNA molecule with a target DNA sequence.

[0025] In some embodiments, the DNA-binding polypeptide of the fusion protein is a nuclease, such as a sequence-specific nuclease. As used herein, the term "nuclease" refers to an enzyme that catalyzes the cleavage of phosphodiester bonds between nucleotides in a nucleic acid molecule. In some embodiments, the DNA-binding polypeptide is an endonuclease capable of cleaving phosphodiester bonds between nucleotides in a nucleic acid molecule, while in certain embodiments, the DNA-binding polypeptide is an exonuclease capable of cleaving nucleotides at either end (5' or 3') of a nucleic acid molecule. In some embodiments, the sequence-specific nuclease is selected from the group consisting of meganucleases, zinc finger nucleases, TAL-effector DNA-binding domain-nuclease fusion proteins (TALENs), and RNA-guided nucleases (RGNs) or variants thereof, and the nuclease activity is reduced or inhibited.

[0026] As used herein, the term "meganuclease" or "homing endonuclease" refers to an endonuclease that binds to a recognition site within double-stranded DNA that is 12 to 40 bp in length. Non-limiting examples of meganucleases include those belonging to the LAGLIDADG family, which contains the conserved amino acid motif LAGLIDADG (SEQ ID NO: 49). The term "meganuclease" can refer to dimeric or single-chain meganucleases.

[0027] As used herein, the term "zinc finger nuclease" or "ZFN" refers to a chimeric protein comprising a zinc finger DNA-binding domain and a nuclease domain.

[0028] As used herein, the term "TAL-effector DNA binding domain-nuclease fusion protein" or "TALEN" refers to a chimeric protein comprising a TAL effector DNA binding domain and a nuclease domain.

[0029] As used herein, the term "RNA-guided nuclease" or "RGN" refers to an RNA-guided DNA-binding polypeptide with nuclease activity. RGN is considered "RNA-guided" because a guide RNA forms a complex with the RNA-guided nuclease, directing and binding the RNA-guided nuclease to a target sequence, and in some embodiments, introducing a single- or double-strand break in the target sequence. RGNs are also known as CasX, CasY, C2c1, C2c2, C2c3, GeoCas9, SpCas9, SaCas9, Nme2Cas9, CjCas9, Cas12a (formerly known as Cpf1), Cas12b, Cas12g, Cas12h, Cas12i, LbCas12a, AsCas12a, CasMINI, Cas13b, Cas13c, Cas13d, The RGN may be Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, or Spy-macCas9, a Spy-macCas9 domain, or an RGN having the amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368. In some embodiments, the RGN provided herein is an RGN nickase, as described below.

[0030] According to the present invention, RGN proteins, such as dCas9, that have been mutated to be nuclease-inactive or "deactivated" can be referred to as RNA-guided DNA-binding polypeptides or nuclease-inactive RGNs or nuclease-deactivated RGNs. Additionally, suitable nuclease-inactive Cas9 domains can be determined for other known RNA-guided nucleases (RGNs) (e.g., nuclease-inactive variants of RGN APG08290.1 disclosed in U.S. Patent Application Publication No. 2019 / 0367949, the entire contents of which are incorporated herein by reference).

[0031] In some embodiments, the fusion protein comprises RGN fused to a deaminase described herein. In the above fusion protein embodiments, the deaminase is selected from deaminases comprising an amino acid sequence having at least 80% sequence identity to any one of SEQ ID NOS: 1-10 and 399-441. In some embodiments, the deaminase comprises an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 407. In some embodiments, the deaminase comprises an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 399. In some embodiments, the deaminase comprises an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 405. In embodiments of the above fusion proteins, the RGN is selected from CasX, CasY, C2c1, C2c2, C2c3, GeoCas9, SpCas9, SaCas9, Nme2Cas9, CjCas9, Cas12a (formerly known as Cpf1), Cas12b, Cas12g, Cas12h, Cas12i, LbCas12a, AsCas12a, CasMINI, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368. In certain embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 407. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 399. In certain embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 405. The Cas9 nickase may be any Cas9 nickase disclosed in WO2020181195, the entire contents of which are incorporated herein by reference.

[0032] The term "RGN polypeptide" encompasses RGN polypeptides, referred to herein as nickases, that cleave only one strand of a target nucleotide sequence. Such RGNs have a single functional nuclease domain. An RGN nickase may be a naturally occurring nickase or an RGN protein that naturally cleaves both strands of a double-stranded nucleic acid molecule mutated in one or more nuclease domains such that the nuclease activity of these mutated domains is reduced or eliminated. In some embodiments, the nickase RGN of the fusion protein contains a mutation (e.g., a D10A mutation) that enables RGN to cleave only the non-base-edited target strand of a nucleic acid duplex (the strand that contains the PAM and base-pairs with the gRNA). This D10A mutation mutates the first aspartic acid residue of the split RuvC nuclease domain of RGN. The present application discloses several D10A nickase variants or homologous nickase variants of RGN described (see Example 4). nAPG07433.1 and nAPG08290.1 (represented as SEQ ID NOs: 42 and 61, respectively) are nickase variants of APG07433.1 and APG08290.1, represented as SEQ ID NOs: 41 and 60, respectively, and are described in WO 2019 / 236566 (incorporated herein by reference in their entireties). nAPG00969 (represented as SEQ ID NO: 52) and nAPG09748 (represented as SEQ ID NO: 54) are nickase variants of APG00969 and APG09748, respectively, and are described in WO 2020 / 139783 (incorporated herein by reference in their entireties). nAPG06646 (represented as SEQ ID NO: 53) and nAPG09882 (represented as SEQ ID NO: 55) are nickase variants of APG06646 and APG09882, respectively, and are described in WO 2021 / 030344 (incorporated herein by reference in its entirety).nAPG03850, nAPG07553, nAPG055886, and nAPG01604 are set forth as SEQ ID NOS: 56-59, respectively, and are nickase variants of APG03850, APG07553, APG055886, and APG01604, which are described in pending International Application PCT / US2021 / 028843, which is incorporated herein by reference in its entirety. Various RGN nickases, their variants, and their sequences are disclosed in International Publication No. WO2020181195, the entire contents of which are incorporated herein by reference. One exemplary suitable nuclease-inactive Cas9 is the D10A / H840A Cas9 mutant (see, e.g., Qi et al., Cell. 2013;152(5):1173-83, the entire contents of which are incorporated herein by reference).

[0033] In some embodiments, the nickase RGN of the fusion protein contains a mutation (e.g., an H840A mutation) that allows RGN to cleave only the base-editing non-target strand of a nucleic acid duplex (the strand that does not contain a PAM and is not base-paired with the gRNA). The H840A mutation mutates the first histidine of the HNH nuclease domain. A nickase RGN containing the H840A mutation or an equivalent mutation has an inactivated HNH domain. A nickase RGN containing the H840A mutation cleaves the non-target strand. A nickase containing the D10A mutation or an equivalent mutation has an inactivated RuvC nuclease domain and cleaves the target strand. A D10A nickase cannot cleave the non-target strand of DNA, i.e., the strand where base editing is desired.

[0034] Other additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, D10A / D839A / H840A and D10A / D839A / H840A / N863A mutant domains (see, e.g., Mali et al., Nature Biotechnology. 2013;31(9):833-838, the entire contents of which are incorporated herein by reference). Additional suitable RGN proteins mutated into nickases will be apparent to those skilled in the art based on this disclosure and knowledge in the art (e.g., RGNs disclosed in WO 2019 / 236566 and WO 2020181195, the entire contents of which are incorporated herein by reference) and are within the scope of the present disclosure. In a preferred embodiment, an RGN having nickase activity against a target strand nicks the target strand, while the complementary non-target strand is modified by a deaminase. The cell's DNA repair machinery can use the modified non-target strand as a template to repair the nicked target strand, thereby introducing a mutation into the DNA.

[0035] In some embodiments, an RGN nickase that retains nickase activity comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to SEQ ID NO: 42, or any one of SEQ ID NOs: 52-59, 61, 397, and 398.

[0036] Any method known in the art for introducing mutations into amino acid sequences, such as PCR-mediated mutagenesis and site-directed mutagenesis, can be used to generate nickase or nuclease-inactive RGNs. See, for example, U.S. Patent Application Publication Nos. 2014 / 0068797 and 9,790,490, each of which is incorporated herein by reference in its entirety. RNA-guided nucleases (RGNs) enable targeted manipulation of single sites within the genome and are useful in the context of gene targeting for therapeutic and research applications. In various organisms, including mammals, RNA-guided nucleases have been used for genome manipulation by stimulating either non-homologous end joining or homologous recombination. RGNs include CRISPR-Cas proteins, or active variants or fragments thereof, which are RNA-guided nucleases directed to target sequences by guide RNAs (gRNAs) as part of the clustered regularly interspaced short palindromic repeats (CRISPR) RNA-guided nuclease system.

[0037] Further provided herein are RGN polypeptides (and nucleic acid molecules encoding RGN polypeptides) comprising the amino acid sequence set forth as SEQ ID NO: 41 or 60, but lacking amino acid residues 590-597 of SEQ ID NO: 41 or 60, or active variants or fragments thereof. In certain embodiments, the RGN polypeptide comprises the amino acid sequence set forth as SEQ ID NO: 366, 368, 397, or 398, or an active variant or fragment thereof.

[0038] Some aspects of the present disclosure provide a fusion protein comprising an RNA-guided DNA-binding polypeptide and a deaminase polypeptide, specifically an adenine deaminase polypeptide. In some embodiments, the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease. In further embodiments, the RNA-guided nuclease is a naturally occurring CRISPR-Cas protein or an active variant or fragment thereof. CRISPR-Cas systems are classified as class 1 or class 2 systems. Class 2 systems are composed of a single effector nuclease and include types II, V, and VI. Class 1 and 2 systems are subdivided into multiple types (I, II, III, IV, V, VI), and some types are further subdivided into multiple subtypes (e.g., type II-A, type II-B, type II-C, type VA, type VB).

[0039] In certain embodiments, the CRISPR-Cas protein is a naturally occurring type II CRISPR-Cas protein or an active variant or fragment thereof. As used herein, the terms "type II CRISPR-Cas protein," "type II CRISPR-Cas effector protein," or "Cas9" refer to a CRISPR-Cas effector protein that requires a trans-activating RNA (tracrRNA) and contains two nuclease domains (i.e., RuvC and HNH), each of which is involved in cleaving a single strand of a double-stranded DNA molecule. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to Streptococcus pyogenes Cas9 (SpCas9) or SpCas9 nickase, the sequences of which are set forth as SEQ ID NOs: 555 and 556, respectively, and which are described in U.S. Patent Nos. 10,000,772 and 8,697,359, each of which is incorporated herein by reference in its entirety. In some embodiments, the present invention provides fusion proteins comprising a deaminase of the present disclosure fused to Streptococcus thermophilus Cas9 (StCas9) or StCas9 nickase, the sequences of which are set forth as SEQ ID NOs: 557 and 558, respectively, and are described in U.S. Patent No. 10,113,167, which is incorporated herein by reference in its entirety. In some embodiments, the present invention provides fusion proteins comprising a deaminase of the present disclosure fused to Streptococcus aureus Cas9 (SaCas9) or SaCas9 nickase, the sequences of which are set forth as SEQ ID NOs: 559 and 560, respectively, and are disclosed in U.S. Patent No. 9,752,132, which is incorporated herein by reference in its entirety.

[0040] In some embodiments, the CRISPR-Cas protein is a naturally occurring type V CRISPR-Cas protein or an active variant or fragment thereof. As used herein, "Type V CRISPR-Cas protein," "Type V CRISPR-Cas effector protein," or "Cas12" refers to a CRISPR-Cas effector protein that cleaves dsDNA and contains a single RuvC nuclease domain or a split RuvC nuclease domain and lacks an HNH domain (Zetsche et al. 2015, Cell doi:10.1016 / j.cell.2015.09.038; Shmakov et al. 2017, Nat Rev Microbiol doi:10.1038 / nrmicro.2016.184; Yan et al. 2018, Science doi:10.1126 / science.aav7271; Harrington et al. 2018, Science doi:10.1126 / science.aav4294). Note that Cas12a, also known as Cpf1, does not require tracrRNA, while other type V CRISPR-Cas proteins, such as Cas12b, do. Most type V effectors can also target ssDNA (single-stranded DNA), often without the need for a PAM (Zetsche et al., 2015; Yan et al., 2018; Harrington et al., 2018). The term "type V CRISPR-Cas protein" includes unique RGNs containing a split RuvC nuclease domain, such as those disclosed in U.S. Provisional Application Nos. 62 / 955,014, filed December 30, 2019, and 63 / 058,169, filed July 29, 2020, and International Application No. PCT / US2020 / 067138, filed December 28, 2020 (the contents of each application are incorporated by reference in their entirety into this specification).In some embodiments, the present invention provides fusion proteins comprising a deaminase of the present disclosure fused to either Francisella novicida Cas12a (FnCas12a) (the sequence of which is set forth as SEQ ID NO:561 and is disclosed in U.S. Patent No. 9,790,490, which is incorporated by reference herein in its entirety), or a nuclease-inactive mutant of FnCas12a disclosed in U.S. Patent No. 9,790,490.

[0041] In some embodiments, the CRISPR-Cas protein is a naturally occurring Type VI CRISPR-Cas protein or an active variant or fragment thereof. As used herein, "Type VI CRISPR-Cas protein," "Type VI CRISPR-Cas effector protein," or "Cas13" refers to a CRISPR-Cas effector protein that contains two HEPN domains that cleave RNA without requiring tracrRNA.

[0042] The term "guide RNA" refers to a nucleotide sequence that has sufficient complementarity with a target nucleotide sequence to hybridize with the target sequence and direct and bind an associated RGN to the target nucleotide sequence in a sequence-specific manner. In the case of a CRISPR-Cas RGN, each guide RNA is one or more RNA molecules (generally one or two) that can bind to the RGN and guide the RGN to bind to a specific target nucleotide sequence, and also cleave the target nucleotide sequence if the RGN has nickase or nuclease activity. Guide RNAs include CRISPR RNAs (crRNAs), and in some embodiments, trans-activating CRISPR RNAs (tracrRNAs).

[0043] CRISPR RNA comprises a spacer sequence and a CRISPR repeat sequence. A "spacer sequence" is a nucleotide sequence that directly hybridizes to a target nucleotide sequence of interest. The spacer sequence is engineered to be fully or partially complementary to the target sequence of interest. In various embodiments, the spacer sequence comprises from about 8 nucleotides to about 30 nucleotides or more. For example, the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the spacer sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the spacer sequence is about 10 to about 26 nucleotides in length or about 12 to about 30 nucleotides in length. In some embodiments, the spacer sequence is 10 to 26 nucleotides in length or 12 to 30 nucleotides in length. In certain embodiments, the spacer sequence is about 30 nucleotides in length. In certain embodiments, the spacer sequence is 30 nucleotides in length. In some embodiments, the degree of complementarity between a spacer sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is between 50% and 99% or more, including, but not limited to, about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more. In certain embodiments, the degree of complementarity between a spacer sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more.In certain embodiments, the spacer sequence is free of secondary structure, which can be predicted using any suitable polynucleotide folding algorithm known in the art, such as, but not limited to, mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148)) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).

[0044] CRISPR RNA repeats comprise a nucleotide sequence that, by itself or in concert with a hybridized tracrRNA, forms a structure recognized by an RGN molecule. In various embodiments, the CRISPR RNA repeat comprises from about 8 nucleotides to about 30 nucleotides or more. In certain embodiments, the CRISPR RNA repeat comprises from 8 nucleotides to 30 nucleotides or more. For example, the CRISPR repeat can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In certain embodiments, the CRISPR repeat sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA sequence, when optimally aligned using a suitable alignment algorithm, is between 50% and 99% or more, including, but not limited to, about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more. In certain embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA sequence, when optimally aligned using a suitable alignment algorithm, is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more.

[0045] In some embodiments, the guide RNA further comprises a tracrRNA molecule. A transactivating CRISPR RNA or tracrRNA molecule comprises a nucleotide sequence comprising a region sufficiently complementary to the CRISPR repeat sequence of the crRNA, referred to herein as the anti-repeat region. In some embodiments, the tracrRNA molecule further comprises a region with a secondary structure (e.g., a stem-loop) or forms a secondary structure upon hybridization with its corresponding crRNA. In certain embodiments, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is at the 5' end of the molecule, and the 3' end of the tracrRNA comprises a secondary structure. This region of secondary structure generally comprises several hairpin structures, including a nexus hairpin found adjacent to the anti-repeat sequence. The 3' end of the tracrRNA often contains a terminal hairpin, which can vary in structure and number, but often includes a GC-rich Rho-independent transcription terminator hairpin followed by a string of Us at the 3' end. See, e.g., Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Protoc; doi:10.1101 / pdb.top090902, and U.S. Patent Application Publication No. 2017 / 0275648, each of which is incorporated herein by reference.

[0046] In various embodiments, the anti-repeat region of the tracrRNA, which is fully or partially complementary to the CRISPR repeat sequence, comprises from about 6 nucleotides to about 30 nucleotides or more. For example, the base-paired region between the tracrRNA anti-repeat sequence and the CRISPR repeat sequence can be about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In certain embodiments, the base-paired region between the tracrRNA anti-repeat sequence and the CRISPR repeat sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In certain embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is about 10 nucleotides in length. In certain embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is 10 nucleotides in length. In some embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA anti-repeat sequence, when optimally aligned using a suitable alignment algorithm, is between 50% and 99% or more, including, but not limited to, about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more. In certain embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA anti-repeat sequence is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more when optimally aligned using a suitable alignment algorithm.

[0047] In various embodiments, the full-length tracrRNA comprises from about 60 nucleotides to more than about 210 nucleotides. In certain embodiments, the full-length tracrRNA comprises from about 60 nucleotides to more than 210 nucleotides. For example, the tracrRNA can be about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, or more nucleotides in length. In certain embodiments, the tracrRNA is 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210, or more nucleotides in length. In certain embodiments, the tracrRNA is about 100 to about 210 nucleotides in length, including about 95, about 96, about 97, about 98, about 99, about 100, about 105, about 106, about 107, about 108, about 109, and about 100 nucleotides in length. In certain embodiments, the tracrRNA is 100-110 nucleotides in length, including 95, 96, 97, 98, 99, 100, 105, 106, 107, 108, 109, and 110 nucleotides in length.

[0048] The guide RNA forms a complex with an RNA-guided DNA-binding polypeptide or an RNA-guided nuclease, directing and binding the RNA-guided nuclease to the target sequence. When the guide RNA forms a complex with an RGN, the bound RGN introduces a single- or double-strand break in the target sequence. After cleavage of the target sequence, the break can be repaired, resulting in a modification of the DNA sequence of the target sequence during the repair process. Provided herein are methods for modifying a target sequence in the DNA of a host cell using a mutant variant of an RNA-guided nuclease that is either a nuclease-inactive form or a nickase linked to a deaminase. A mutant variant of an RNA-guided nuclease with inactivated or significantly reduced nuclease activity can bind to a target sequence but does not necessarily cleave the target sequence, and therefore can be referred to as an RNA-guided DNA-binding polypeptide. An RNA-guided nuclease that can cleave only one strand of a double-stranded nucleic acid molecule is referred to as a nickase herein.

[0049] The target nucleotide sequence is bound by the RNA-guided DNA-binding polypeptide and hybridizes with the guide RNA associated with RGDBP, and the target sequence can then be subsequently cleaved if RGDBP has nuclease activity, including nickase activity (i.e., is an RGN).

[0050] The guide RNA may be a single guide RNA or a dual guide RNA system. A single guide RNA comprises a crRNA and optionally a tracrRNA on a single RNA molecule, while a dual guide RNA system comprises a crRNA and a tracrRNA on two different RNA molecules, hybridized to each other via at least a portion of the CRISPR repeat sequence of the crRNA and at least a portion of the tracrRNA, which may be fully or partially complementary to the CRISPR repeat sequence of the crRNA. In some embodiments where the guide RNA is a single guide RNA, the crRNA and optionally the tracrRNA are separated by a linker nucleotide sequence.

[0051] Generally, the linker nucleotide sequence does not contain complementary bases to avoid the formation of secondary structures within or involving the nucleotides of the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between the crRNA and the tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or more nucleotides in length. In certain embodiments, the linker nucleotide sequence between the crRNA and the tracrRNA is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more nucleotides in length. In certain embodiments, the linker nucleotide sequence of a single guide RNA is at least 4 nucleotides in length. In certain embodiments, the linker nucleotide sequence of a single guide RNA is 4 nucleotides in length.

[0052] In certain embodiments, the guide RNA may be introduced into a target cell, organelle, or embryo as an RNA molecule. The guide RNA may be in vitro transcribed or chemically synthesized. In some embodiments, a nucleotide sequence encoding the guide RNA is introduced into a cell, organelle, or embryo. In some embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., an RNA polymerase III promoter). The promoter may be a native promoter or may be heterologous to the nucleotide sequence encoding the guide RNA.

[0053] In various embodiments, the guide RNA may be introduced into a target cell, organelle, or embryo as a ribonucleoprotein complex, as described herein, wherein the guide RNA binds to an RNA-guided nuclease polypeptide.

[0054] The guide RNA directs the associated RNA-guided nuclease to a specific target nucleotide sequence of interest upon hybridization of the guide RNA to the target nucleotide sequence. The target nucleotide sequence may comprise DNA, RNA, or a combination of both, and may be single-stranded or double-stranded. The target nucleotide sequence may be genomic DNA (i.e., chromosomal DNA), plasmid DNA, or an RNA molecule (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). The target nucleotide sequence may be bound (and, in some embodiments, cleaved) by an RNA-guided DNA-binding polypeptide in vitro or in a cell. The chromosomal sequence targeted by the RGDBP may be a nuclear, plastid, or mitochondrial chromosomal sequence. In some embodiments, the target nucleotide sequence is unique to the target genome.

[0055] In some embodiments, the target nucleotide sequence is adjacent to a protospacer adjacent motif (PAM). The PAM is generally within about 1 to about 10 nucleotides of the target nucleotide sequence (including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides of the target nucleotide sequence). In certain embodiments, the PAM is within 1 to 10 nucleotides of the target nucleotide sequence (including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the target nucleotide sequence). The PAM can be 5' or 3' to the target sequence. In some embodiments, the PAM is 3' to the target sequence. Generally, the PAM is a consensus sequence of about 2 to 6 nucleotides, but in certain embodiments, it is 1, 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length.

[0056] Because the PAM must be close to the target nucleotide sequence, the PAM that can target a given RGDBP or RGN sequence is limited. Upon recognizing its corresponding PAM sequence, the RGN can cleave the target nucleotide sequence at a specific cleavage site. As used herein, a cleavage site is composed of two specific nucleotides within the target nucleotide sequence, between which the nucleotide sequence is cleaved by the RGN. The cleavage site can include the first and second, second and third, third and fourth, fourth and fifth, fifth and sixth, seventh and eighth, or eighth and ninth nucleotides in either the 5' or 3' direction from the PAM. Because the RGN can cleave the target nucleotide sequence to create a cohesive end, in some embodiments, the cleavage site is defined based on a distance of two nucleotides from the PAM on the plus (+) strand of the polynucleotide and a distance of two nucleotides from the PAM on the minus (-) strand of the polynucleotide.

[0057] RGDBP and RGN can be used to deliver fusion polypeptides, polynucleotides, or small molecule payloads to specific genomic locations.

[0058] In embodiments in which the DNA-binding polypeptide comprises a meganuclease, the target sequence comprises a pair of inverted 9-base pair "half-sites" separated by 4 base pairs. In the case of a single-chain meganuclease, the N-terminal domain of the protein contacts the first half-site, and the C-terminal domain of the protein contacts the second half-site. Cleavage by the meganuclease generates a 4-base pair 3' overhang. In embodiments in which the DNA-binding polypeptide comprises a compact TALEN, the recognition sequence comprises a first CNNNGN sequence recognized by the I-TevI domain, followed by a 4-16 base pair long nonspecific spacer, followed by a 16-22 bp long second sequence (this sequence typically has a 5' T base) recognized by the TAL effector domain. In embodiments in which the DNA-binding polypeptide comprises a zinc finger, the DNA-binding domain typically recognizes an 18-bp recognition sequence containing a pair of 9-base pair "half-sites" separated by 2-10 base pairs, and cleavage by the nuclease generates either a blunt end or a 5' overhang of variable length (often 4 base pairs).

[0059] IV. Fusion Proteins In some embodiments, a DNA-binding polypeptide (e.g., a nuclease-inactive or nickase RGN) is operably linked to a deaminase of the invention. In some embodiments, a DNA-binding polypeptide (e.g., a nuclease-inactive RGN or nickase RGN) fused to a deaminase of the invention can target a specific location of a nucleic acid molecule (i.e., a target nucleic acid molecule), in some embodiments, a specific genomic locus, to alter expression of a desired sequence. In some embodiments, binding of the fusion protein to the target sequence deaminates a nucleobase, converting one nucleobase to another. In some embodiments, binding of the fusion protein to the target sequence deaminates a nucleobase adjacent to the target sequence. The nucleobases adjacent to the target sequence to be deaminated and mutated using the compositions and methods of the present disclosure can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 base pairs from the 5' or 3' end of the target sequence (bound by the gRNA) within the target nucleic acid molecule. Some aspects of the present disclosure provide fusion proteins comprising (i) a DNA-binding polypeptide (e.g., a nuclease-inactive or nickase RGN polypeptide), (ii) a deaminase polypeptide, and optionally (iii) a second deaminase. The second deaminase can be the same deaminase as the first deaminase or a different deaminase. In some embodiments, both the first and second deaminases are adenine deaminases of the invention.

[0060] The present disclosure provides fusion proteins of various configurations. In some embodiments, the deaminase polypeptide is fused to the N-terminus of the DNA-binding polypeptide (e.g., RGN polypeptide). In some embodiments, the deaminase polypeptide is fused to the C-terminus of the DNA-binding polypeptide (e.g., RGN polypeptide).

[0061] In some embodiments, the deaminase and the DNA-binding polypeptide (e.g., an RNA-guided DNA-binding polypeptide) are fused to each other via a peptide linker. The linker between the deaminase and the DNA-binding polypeptide (e.g., an RNA-guided DNA-binding polypeptide) determines the editing window of the fusion protein, thereby increasing deaminase specificity and reducing off-target mutations. To achieve optimal length and rigidity for deaminase activity for a particular application, a peptide linker such as (GGGGS) n and (G) n From a very flexible linker in the form of (EAAAK) n and (XP) n A variety of linker lengths and flexibility can be employed, ranging from more rigid linkers of the form: As used herein, the term "linker" refers to a chemical group or molecule that links two molecules or moieties, such as the binding domain and cleavage domain of a nuclease. In some embodiments, a linker links an RNA-guided nuclease and a deaminase. In some embodiments, a linker links a dead or inactive RGN and a deaminase. In further embodiments, a linker links two deaminases. Typically, a linker is positioned between or adjacent to two groups, molecules, or other moieties, covalently linking each and thus linking the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 3 to 100 amino acids in length, e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, shorter linkers are preferred to reduce the overall size or length of the fusion protein or its coding sequence.

[0062] In some embodiments, the linker is (GGGGS) n , (G) n , (EAAAK) n , or (XP) n motif, or any combination thereof, where n is independently an integer from 1 to 30. In some embodiments, n is independently 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30, or any combination thereof when more than one linker or more than one linker motif is present. Further suitable linker motifs and linker configurations will be apparent to those of skill in the art. In some embodiments, suitable linker motifs and configurations include those described in Chen et al., 2013 (Adv Drug Deliv Rev. 65(10):1357-69, the entire contents of which are incorporated herein by reference). Further suitable linker sequences will be apparent to those of skill in the art. In some embodiments, the linker sequence comprises the amino acid sequence set forth as SEQ ID NO: 45 or 442.

[0063] In some embodiments, the general architecture of exemplary fusion proteins provided herein comprises the structure: [NH]-[deaminase]-[DBP]-[COOH]; [NH]-[DBP]-[deaminase]-[COOH]; [NH]-[DBP]-[deaminase]-[deaminase]-[COOH]; [NH]-[deaminase]-[DBP]-[deaminase]-[COOH]; or [NH]-[deaminase]-[deaminase]-[DBP]-[COOH], where DBP is the DNA-binding polypeptide, NH is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein. In some embodiments, the fusion protein comprises three or more deaminase polypeptides.

[0064] In certain embodiments, the general architecture of exemplary fusion proteins provided herein comprises the structure: [NH]-[deaminase]-[RGN]-[COOH]; [NH]-[RGN]-[deaminase]-[COOH]; [NH]-[RGN]-[deaminase]-[deaminase]-[COOH]; [NH]-[deaminase]-[RGN]-[deaminase]-[COOH]; or [NH]-[deaminase]-[deaminase]-[RGN]-[COOH], where NH is the N-terminus of the fusion protein and COOH is the C-terminus of the fusion protein. In some embodiments, the fusion protein comprises three or more deaminase polypeptides.

[0065] In some embodiments, the fusion protein comprises the structure: [NH2]-[deaminase]-[nuclease-inactive RGN]-[COOH]; [NH2]-[deaminase]-[deaminase]-[nuclease-inactive RGN]-[COOH]; [NH2]-[nuclease-inactive RGN]-[deaminase]-[COOH]; [NH2]-[deaminase]-[nuclease-inactive RGN]-[deaminase]-[COOH]; or [NH2]-[nuclease-inactive RGN]-[deaminase]-[deaminase]-[COOH]. It should be understood that "nuclease-inactive RGN" refers to any RGN, including any CRISPR-Cas protein that has been mutated to be nuclease-inactive. In some embodiments, the fusion protein comprises three or more deaminase polypeptides.

[0066] In some embodiments, the fusion protein comprises the structure: [NH2]-[deaminase]-[RGN nickase]-[COOH]; [NH2]-[deaminase]-[deaminase]-[RGN nickase]-[COOH]; [NH2]-[RGN nickase]-[deaminase]-[COOH]; [NH2]-[deaminase]-[RGN nickase]-[deaminase]-[COOH]; or [NH2]-[RGN nickase]-[deaminase]-[deaminase]-[COOH]. It should be understood that "RGN nickase" refers to any RGN, including any CRISPR-Cas protein that has been mutated to be active as a nickase.

[0067] In some embodiments, the "-" used in the above general architecture indicates the presence of an optional linker sequence. In some embodiments, the fusion proteins provided herein do not include a linker sequence. In some embodiments, at least one of the optional linker sequences is present.

[0068] Other exemplary features that may be present include localization sequences, such as nuclear localization sequences, cytoplasmic localization sequences, transport sequences, such as nuclear export sequences, or other localization sequences, as well as sequence tags useful for solubilizing, purifying, or detecting the fusion protein. Suitable localization signal sequences and protein tag sequences provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, polyhistidine tags (also known as histidine tags or His tags), maltose-binding protein (MBP) tags, nus tags, glutathione-S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softags (e.g., Softag1, Softag3), strep tags, biotin ligase tags, FlAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art.

[0069] In certain embodiments, the fusion proteins of the present disclosure comprise at least one cell-penetrating domain that facilitates cellular uptake of the fusion protein. Cell-penetrating domains are known in the art and generally comprise a series of positively charged amino acid residues (i.e., polycationic cell-penetrating domains), alternating polar and non-polar amino acid residues (i.e., amphipathic cell-penetrating domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domains) (see, e.g., Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell-penetrating domain is the trans-activating transcriptional activator (TAT) from human immunodeficiency virus 1.

[0070] In some embodiments, the deaminase or fusion protein provided herein further comprises a nuclear localization sequence (NLS). The nuclear localization signal, plastid localization signal, mitochondrial localization signal, dual-targeting localization signal, and / or cell-penetrating domain may be located at the amino-terminus (N-terminus), carboxyl-terminus (C-terminus), or internal position of the fusion protein.

[0071] In some embodiments, the NLS is fused to the N-terminus of the fusion protein or deaminase. In some embodiments, the NLS is fused to the C-terminus of the fusion protein or deaminase. In some embodiments, the NLS is fused to the N-terminus of the deaminase of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the deaminase of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the DNA-binding polypeptide (e.g., RGN polypeptide) of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the DNA-binding polypeptide (e.g., RGN polypeptide) of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the deaminase polypeptide of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the deaminase polypeptide of the fusion protein. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises any one of the amino acid sequences of the NLS sequences provided or referenced herein. In some embodiments, the NLS comprises the amino acid sequence set forth in SEQ ID NO: 43 or SEQ ID NO: 46. In some embodiments, the fusion protein or deaminase comprises SEQ ID NO: 43 at its N-terminus and SEQ ID NO: 46 at its C-terminus.

[0072] In some embodiments, the fusion proteins provided herein comprise the full-length sequence of a deaminase, e.g., any one of SEQ ID NOs: 1-10 and 399-441. However, in some embodiments, the fusion proteins provided herein do not comprise the full-length sequence of a deaminase, but only a fragment thereof. For example, in some embodiments, the fusion proteins provided herein further comprise a DNA-binding polypeptide (e.g., RNA-guided DNA-binding) domain and a deaminase domain.

[0073] In some embodiments, a fusion protein of the invention comprises a DNA-binding polypeptide (e.g., RGN) and a deaminase, wherein the deaminase has an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any of SEQ ID NOs: 1-10 and 399-441. Examples of such fusion proteins are described in the Examples section herein.

[0074] In some embodiments, the fusion protein comprises one deaminase polypeptide. In some embodiments, the fusion protein comprises at least two deaminase polypeptides operably linked directly or via a peptide linker. In some embodiments, the fusion protein comprises one deaminase polypeptide, and a second deaminase polypeptide is co-expressed with the fusion protein.

[0075] Also provided herein is a ribonucleoprotein complex comprising a fusion protein comprising a deaminase and RGDBP, and a guide RNA, which is either a single-guide or dual-guide RNA (collectively also referred to as gRNA).

[0076] V. Nucleotides encoding deaminases, fusion proteins, and / or gRNAs The present disclosure provides polynucleotides (SEQ ID NOS: 11-20 and 443-485) encoding the deaminase polypeptides of the present disclosure. The present disclosure further provides polynucleotides encoding fusion proteins comprising a deaminase and a DNA-binding polypeptide (e.g., a meganuclease, a zinc finger fusion protein, or a TALEN). The present disclosure further provides polynucleotides encoding fusion proteins comprising a deaminase domain and an RNA-guided DNA-binding polypeptide. Such RNA-guided DNA-binding polypeptides can be RGN or RGN variants. The protein variants can be nuclease-inactive or nickase. RGN can be a CRISPR-Cas protein or an active variant or fragment thereof. SEQ ID NOS: 41 and 42 are non-limiting examples of RGN and nickase RGN variants, respectively. Examples of CRISPR-Cas nucleases are well known in the art, and similar corresponding mutations can be made to generate mutant variants that are also nickase or nuclease-inactive.

[0077] Embodiments of the present invention provide polynucleotides encoding a fusion protein comprising an RGDBP described herein and a deaminase (SEQ ID NOS: 1-10 and 399-441, or variants thereof). In some embodiments, a second polynucleotide encodes a guide RNA required by the RGDBP for targeting to a nucleotide sequence of interest. In some embodiments, the guide RNA and the fusion protein are encoded by the same polynucleotide.

[0078] The use of the term "polynucleotide" is not intended to limit the present disclosure to polynucleotides comprising DNA, although such DNA polynucleotides are contemplated. As will be appreciated by those of skill in the art, polynucleotides can include ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogues. Polynucleotides disclosed herein also encompass sequences in all forms, including, but not limited to, single-stranded forms, double-stranded forms, stem-loop structures, circular forms (including, for example, circular RNA), and the like.

[0079] One embodiment of the present invention is a nucleic acid molecule comprising a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOS: 11-20 and 443-485, wherein the nucleic acid molecule encodes a deaminase having adenine deaminase activity. The nucleic acid molecule may further comprise a heterologous promoter or terminator. The nucleic acid molecule may encode a fusion protein, wherein the encoded deaminase is operably linked to a DNA-binding polypeptide and, optionally, a second deaminase. In some embodiments, the nucleic acid molecule encodes a fusion protein, wherein the encoded deaminase is operably linked to an RGN and, optionally, a second deaminase.

[0080] In some embodiments, nucleic acid molecules comprising polynucleotides encoding deaminases of the invention are codon-optimized for expression in a target organism. A "codon-optimized" coding sequence is a polynucleotide coding sequence whose codon usage is designed to mimic the preferred codon usage or transcription conditions of a particular host cell. Expression in a particular host cell or organism is enhanced as a result of altering one or more codons at the nucleic acid level such that the translated amino acid sequence is unchanged. Nucleic acid molecules may be fully or partially codon-optimized. Codon tables and other references providing preference information for a wide range of organisms are available in the art (e.g., see Campbell and Gowri (1990) Plant Physiol. 92:1-11 for a discussion of preferred codon usage in plants). Methods for synthesizing plant-preferred genes are available in the art. See, e.g., U.S. Patent Nos. 5,380,831 and 5,436,391, and Murray et al. (1989) Nucleic Acids Res. 17:477-498, which are incorporated herein by reference.

[0081] In some embodiments, polynucleotides encoding the deaminases, fusion proteins, and / or gRNAs described herein are provided within expression cassettes for in vitro expression or for expression in a cell, organelle, embryo, or organism of interest. The cassettes may include 5' and 3' regulatory sequences operably linked to the polynucleotides encoding the deaminases provided herein and / or fusion proteins comprising the deaminase, RNA-guided DNA-binding polypeptide, and optionally a second deaminase, and / or gRNA, allowing for expression of the polynucleotides. The cassettes may further include at least one additional gene or genetic element cotransformed into the organism. When additional genes or elements are included, those components are operably linked. The term "operably linked" is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a promoter and a coding region of interest (e.g., a region encoding a deaminase, RNA-guided DNA-binding polypeptide, and / or gRNA) is a functional linkage that allows for expression of the coding region of interest. Operably linked elements may or may not be contiguous. When used to refer to the joining of two protein-coding regions, operably linked intends that the coding regions are in the same reading frame. In some embodiments, additional genes or elements can be provided in multiple expression cassettes. For example, a nucleotide sequence encoding a deaminase of the present disclosure can be present on one expression cassette, either alone or as a component of a fusion protein, and a nucleotide sequence encoding a gRNA can be present on a separate expression cassette. Another example can have a nucleotide sequence encoding a deaminase of the present disclosure alone in a first expression cassette, a second expression cassette encoding a fusion protein containing the deaminase, and a third expression cassette encoding a gRNA. Such expression cassettes include multiple restriction and / or recombination sites for inserting a polynucleotide under the transcriptional control of the regulatory region. An expression cassette containing a selectable marker gene may also be present.

[0082] An expression cassette may include, in the 5' to 3' direction of transcription, a transcriptional (and in some embodiments, translational) initiation region (i.e., promoter), a polynucleotide encoding a deaminase of the present invention, and a transcriptional (and in some embodiments, translational) termination region (i.e., termination region) functional in the organism of interest. A promoter of the present invention is capable of directing or driving expression of a coding sequence in a host cell. Regulatory regions (e.g., promoter, transcriptional regulatory region, and translational termination region) may be endogenous or heterologous to the host cell or to each other. As used herein, "heterologous" with respect to a sequence refers to a sequence that is derived from a foreign species, or, if derived from the same species, has been substantially modified by human intervention from its natural form in composition and / or genomic locus. As used herein, a chimeric gene comprises a coding sequence operably linked to a transcriptional initiation region that is heterologous to the coding sequence.

[0083] Convenient termination regions are available from the Ti plasmid of A. tumefaciens, such as the octopine synthase and nopaline synthase termination regions. Guerineau et al.(1991)Mol.Gen.Genet.262:141-144;Proudfoot(1991)Cell 64:671-674;Sanfacon et al.(1991)Genes Dev.5:141-149;Mogen et al.(1990)Plant Cell 2:1261-1272;Munroe et al. See also Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903; and Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.

[0084] Additional control signals include, but are not limited to, transcription initiation sites, operators, activators, enhancers, other control elements, ribosome binding sites, start codons, termination signals, etc. See, e.g., U.S. Patent Nos. 5,039,523 and 4,853,331; European Patent No. 0480762(A2); Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY) (hereinafter "Sambrook 11"); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and references cited therein.

[0085] In preparing expression cassettes, various DNA fragments may be manipulated to provide DNA sequences in the proper orientation and, if necessary, the proper reading frame. To this end, adapters or linkers may be used to join DNA fragments, or other manipulations may be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. To this end, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, e.g., transitions and transversions, may be involved.

[0086] Many promoters can be used in the practice of the present invention. The promoter can be selected based on the desired results. The nucleic acid can be combined with a constitutive promoter, an inducible promoter, a developmental stage-specific promoter, a cell type-specific promoter, a tissue-preferred promoter, a tissue-specific promoter, or other promoter for expression in the target organism. See, for example, WO 99 / 43838 and U.S. Patent Nos. 8,575,425; 7,790,846; 8,147,856; 8,586,832; 7,772,369; 7,534,939; 6,072,050; 5,659,026; 5,608,149; See the promoters set forth in US Pat. Nos. 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; 5,608,142; and 6,177,611, which are incorporated herein by reference.

[0087] For expression in plants, constitutive promoters also include the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812); rice actin (McElroy et al. (1990) Plant Cell 2:163-171); ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et al. (1992) Plant Mol. Biol. 18:675-689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588); and MAS (Velten et al. (1984) EMBO J. 3:2723-2730).

[0088] Examples of inducible promoters are the Adh1 promoter, which is induced by hypoxia or cold stress, the Hsp70 promoter, which is induced by heat stress, and the PPDK promoter and PEP carboxylase promoter, both of which are induced by light. Chemically inducible promoters, such as the more safely induced In2-2 promoter (U.S. Pat. No. 5,364,780), the auxin-inducible, tapetum-specific, and callus-active Axig1 promoter (International Application US 01 / 22169), steroid-responsive promoters (e.g., the ERE promoter, which is estrogen-inducible, and see the glucocorticoid-inducible promoters of Schena et al. (1991) Proc. Natl. Acad. Sci. USA 88:10421-10425 and McNellis et al. (1998) Plant J. 14(2):247-257), and tetracycline-inducible and tetracycline-repressible promoters (e.g., Gatz et al. (1991) Mol. Gen. Genet. 227:229-237 and U.S. Patent Nos. 5,814,618 and 5,789,156, which are incorporated herein by reference, are also useful.

[0089] In some embodiments, tissue-specific or tissue-preferential promoters are utilized to target expression of an expression construct in a particular tissue. In certain embodiments, tissue-specific or tissue-preferential promoters are active in plant tissues. Examples of promoters under plant developmental control include promoters that preferentially initiate transcription in specific tissues, such as leaves, roots, fruits, seeds, or flowers. A "tissue-specific" promoter is a promoter that initiates transcription only in a specific tissue. Unlike constitutive expression of a gene, tissue-specific expression is the result of several interacting levels of gene regulation. Therefore, it may be preferable to use a promoter from a homologous or closely related plant species to achieve efficient and reliable expression of a transgene in a specific tissue. In some embodiments, expression involves a tissue-preferential promoter. A "tissue-preferential" promoter is a promoter that preferentially initiates transcription in a specific tissue, but not necessarily exclusively or exclusively in a specific tissue.

[0090] In some embodiments, the nucleic acid molecules encoding the deaminases described herein comprise a cell-type-specific promoter. A "cell-type-specific" promoter is a promoter that primarily drives expression in a particular cell type within one or more organs. Some examples of plant cells in which a cell-type-specific promoter functional in a plant may be primarily active include, for example, BETL cells, roots, leaves, vascular cells of stem cells, and stem cells. The nucleic acid molecule may also comprise a cell-type-preferential promoter. A "cell-type-preferential" promoter is a promoter that primarily drives expression in a particular cell type within one or more organs, but not necessarily exclusively or only in a particular cell type. Some examples of plant cells in which a cell-type-preferential promoter functional in a plant may be primarily active include, for example, BETL cells, roots, leaves, vascular cells of stem cells, and stem cells.

[0091] In some embodiments, the nucleic acid sequence encoding the deaminase, fusion protein, and / or gRNA is operably linked to a promoter sequence recognized by a phage RNA polymerase, e.g., for in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA may be purified for use in the methods described herein. For example, the promoter sequence may be a T7, T3, or SP6 promoter sequence, or a variant of a T7, T3, or SP6 promoter sequence. In such embodiments, the expressed protein and / or RNA may be purified for use in the methods of genome modification described herein.

[0092] In certain embodiments, the polynucleotide encoding the deaminase, fusion protein, and / or gRNA is linked to a polyadenylation signal (e.g., the SV40 polyA signal and other signals that function in plants) and / or at least one transcription termination sequence. In some embodiments, as described elsewhere herein, the sequence encoding the deaminase or fusion protein is linked to a sequence encoding at least one nuclear localization signal, at least one cell-penetrating domain, and / or at least one signal peptide capable of transporting the protein to a specific intracellular location.

[0093] In some embodiments, the polynucleotides encoding the deaminase, fusion protein, and / or gRNA are present in one or more vectors. A "vector" refers to a polynucleotide composition for transferring, delivering, or introducing a nucleic acid into a host cell. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, baculoviral vectors). In some embodiments, the vector contains additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Further information can be found in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.

[0094] In some embodiments, the vector contains a selectable marker gene for selecting transformed cells. The selectable marker gene is used to select transformed cells or tissues. Marker genes include genes encoding antibiotic resistance, such as those encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), and genes conferring resistance to herbicide compounds, such as glufosinate ammonium, bromoxynil, imidazolinone, and 2,4-dichlorophenoxyacetic acid (2,4-D).

[0095] In some embodiments, an expression cassette or vector containing a sequence encoding a fusion protein containing an RNA-guided DNA-binding polypeptide, such as RGN, further comprises a sequence encoding a gRNA. In some embodiments, the sequence encoding the gRNA is operably linked to at least one transcriptional control sequence to express the gRNA in an organism or host cell of interest. For example, a polynucleotide encoding the gRNA may be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters, and rice U6 and U3 promoters.

[0096] As described, an organism of interest can be transformed using an expression construct containing a nucleotide sequence encoding a deaminase, a fusion protein, and / or a gRNA. Methods of transformation involve introducing a nucleotide construct into the organism of interest. By "introducing" is meant introducing the nucleotide construct into a host cell so that it can access the interior of the host cell. The methods of the present invention do not require a particular method for introducing the nucleotide construct into the host organism, only that the nucleotide construct be able to access the interior of at least one cell of the host organism. The host cell can be a eukaryotic or prokaryotic cell. In certain embodiments, the eukaryotic host cell is a plant cell, a mammalian cell, or an insect cell. Methods for introducing nucleotide constructs into plant and other host cells are known in the art and include, but are not limited to, stable transformation, transient transformation, and virus-mediated methods.

[0097] This method results in transformed organisms, such as plants, e.g., whole plants, as well as plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos, and their progeny. Plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, pollen).

[0098] A "transgenic organism" or "transformed organism" or "stably transformed" organism, cell, or tissue refers to an organism into which a polynucleotide encoding a deaminase of the present invention has been incorporated or integrated. It is recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments can also be incorporated into a host cell. Agrobacterium- and biolistic-mediated transformation remain the two predominant approaches for transforming plant cells. However, host cell transformation can also be carried out by infection, transfection, microinjection, electroporation, microprojection, biolistic or particle gun methods, electroporation, silica / carbon fiber, ultrasound-mediated, PEG-mediated, calcium phosphate co-precipitation, polycation DMSO technology, DEAE-dextran methods, and viral, liposome-mediated, etc. Viral-mediated transfer of polynucleotides encoding deaminases, fusion proteins, and / or gRNAs includes retrovirus-, lentivirus-, adenovirus-, and adeno-associated virus-mediated transfer and expression, as well as the use of caulimoviruses (e.g., cauliflower mosaic virus), geminiviruses (e.g., bean golden yellow mosaic virus or corn stripe virus), and RNA plant viruses (e.g., tobacco mosaic virus).

[0099] Transformation protocols, as well as protocols for introducing polypeptide or polynucleotide sequences into plants, may vary depending on the type of host cell (e.g., monocotyledonous or cotyledonous plant cell) to be transformed. Transformation methods are known in the art and include those described in U.S. Patent Nos. 8,575,425; 7,692,068; 8,802,934; and 7,541,517, each of which is incorporated herein by reference. Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett.7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4:1-12;Bates, GW (1999) Methods in Molecular Biology 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606; Christou, P. (1992) The Plant Journal 2:275-281; Christou, P. (1995) Euphytica 85:13-27;Tzfira et See also: Yao et al. (2004) TRENDS in Genetics 20:375-383; Yao et al. (2006) Journal of Experimental Botany 57:3737-3746; Zupan and Zambryski (1995) Plant Physiology 107:1041-1047; Jones et al. (2005) Plant Methods 1:5. Transformation can result in stable or transient integration of a nucleic acid into a cell. "Stable transformation" is intended to mean that a nucleotide construct introduced into a host cell is integrated into the genome of the host cell and can be inherited by its progeny. "Transient transformation" is intended to mean that a polynucleotide is introduced into a host cell without being integrated into the genome of the host cell.

[0100] Methods for chloroplast transformation are known in the art. See, for example, Svab et al. (1990) Proc. Natl. Acad. Sci. USA 87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J. 12:601-606. This method relies on particle gun delivery of DNA containing a selectable marker and targeting the DNA to the plastid genome by homologous recombination. Additionally, plastid transformation can be achieved by transactivating a silent transgene carrying the plastid through tissue-preferential expression of a nuclear-encoded, plastid-directed RNA polymerase. Such a system is reported in McBride et al. (1994) Proc. Natl. Acad. Sci. USA 91:7301-7305.

[0101] The transformed cells may be grown into transgenic organisms, e.g., plants, according to conventional methods. See, e.g., McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants may then be grown and pollinated with either the same transformed line or a different line, and hybrids carrying the deaminase or fusion protein polynucleotide may be identified. They may be grown for two or more generations to ensure that the deaminase or fusion protein polynucleotide is stably maintained and inherited, and seeds may be harvested to ensure the presence of the deaminase or fusion protein polynucleotide. In this manner, the present invention provides transformed seeds (also referred to as "transgenic seeds") having stably integrated into their genomes a nucleotide construct of the present invention, e.g., an expression cassette of the present invention.

[0102] In some embodiments, the transformed cells are introduced into an organism, which may be derived from the organism and transformed by an ex vivo approach.

[0103] The sequences provided herein may be used to transform any plant species, including, but not limited to, monocotyledons and dicotyledons. Examples of plants of interest include, but are not limited to, corn, sorghum, wheat, sunflower, tomato, cruciferous, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley, and rapeseed, Brassica, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus fruits, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, oat, vegetables, ornamental plants, and conifers.

[0104] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and plants belonging to the genus Cucumis, such as cucumbers, cantaloupes, and muskmelons. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. Preferably, the plants of the present invention are crop plants (e.g., corn, sorghum, wheat, sunflowers, tomatoes, crucifers, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).

[0105] As used herein, the term plant includes plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, plant mass, and intact plant cells from plants or plant parts (e.g., embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, grains, ears, cobs, husks, stems, roots, root tips, anthers, etc.). Grain is intended to mean mature seeds produced by commercial growers for purposes other than the cultivation or propagation of the species. Progeny, variants, and mutants of regenerated plants are also included within the scope of the present invention, provided that these parts comprise the introduced polynucleotide. Additionally, processed plant products or by-products (including, for example, soybean meal) carrying the sequences disclosed herein are also provided.

[0106] In some embodiments, polynucleotides encoding deaminases, fusion proteins, and / or gRNAs are used to transform any eukaryotic species, including, but not limited to, animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoebas, algae, and yeast. In some embodiments, polynucleotides encoding deaminases, fusion proteins, and / or gRNAs are used to transform any prokaryotic species, including, but not limited to, archaea and bacteria (e.g., Bacillus, Klebsiella, Streptomyces, Rhizobium, Escherichia, Pseudomonas, Salmonella, Shigella, Vibrio, Yersinia, Mycoplasma, Agrobacterium, and Lactobacillus).

[0107] In some embodiments, conventional viral and non-viral gene transfer methods are used to introduce nucleic acids into mammalian cells or target tissues. Using such methods, nucleic acids encoding the deaminase or fusion proteins of the present invention, and optionally gRNAs, can be administered to cells in culture or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with a delivery vehicle (e.g., liposomes). Viral vector delivery systems include DNA and RNA viruses that have either episomal or integrated genomes after delivery to cells. Non-limiting examples include vectors utilizing caulimoviruses (e.g., cauliflower mosaic virus), geminiviruses (e.g., bean golden yellow mosaic virus or corn stripe virus), and RNA plant viruses (e.g., tobacco mosaic virus). For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and See Immunology, Doerfler and Bohm (eds) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).

[0108] Non-viral methods for nucleic acid delivery include lipofection, Agrobacterium-mediated transformation, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid:nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those described in Felgner's International Publication Nos. WO 91 / 17424 and WO 91 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or to target tissues (e.g., in vivo administration). The preparation of lipid:nucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known to those of skill in the art (e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); see U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0109] The use of RNA or DNA virus-based systems for nucleic acid delivery takes advantage of highly evolved processes for targeting viruses to specific cells in the body and transporting the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or they can be used to engineer cells in vitro, with the modified cells then optionally administered to patients (ex vivo). Traditional virus-based systems include retroviral, lentiviral, adenoviral, adeno-associated viral, and herpes simplex viral vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated viral gene transfer methods, often resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiencies have been observed in many different cell types and target tissues.

[0110] The tropism of retroviruses can be altered by incorporating foreign envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system depends on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats that have the capacity to package foreign sequences up to 6-10 kb. The minimal cis-acting LTRs are sufficient for vector replication and packaging, which are then used to integrate therapeutic genes into target cells and permanently express the transgene. Widely used retroviral vectors include vectors based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731-2739 (1992); Johann et al., J. Virol. 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., J. Virol. 63:2374-2378 (1989); Miller et al., J. Virol. 65:2220-2224 (1991); International Application US 94 / 05700).

[0111] For applications where transient expression is preferred, adenovirus-based systems may be used. Adenovirus-based vectors are capable of very high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained using such vectors. This vector can be produced in large quantities using a relatively simple system. Adeno-associated virus ("AAV") vectors may be used to transduce target nucleic acids into cells, for example, in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994)). The construction of recombinant AAV vectors has been described in numerous publications, including U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989). Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells, which package adenovirus, and ΨJ2 or PA317 cells, which package retrovirus.

[0112] Viral vectors used in gene therapy are usually produced by engineering cell lines that package nucleic acid vectors into viral particles. The vectors typically contain minimal viral sequences necessary for packaging and subsequent integration into the host, with other viral sequences replaced with an expression cassette for polynucleotide expression. The missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome necessary for packaging and integration into the host genome. Viral DNA is packaged into a cell line containing a helper plasmid encoding other AAV genes, namely rep and cap, but lacking the ITR sequences.

[0113] Alternatively, the cell line may be infected with adenovirus as a helper. The helper virus promoter promotes the replication of the AAV vector and the expression of AAV genes from the helper plasmid. The helper plasmid is not packaged in large quantities due to the absence of ITR sequences. Contamination by adenovirus can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV. Additional methods for delivering nucleic acids into cells are known to those skilled in the art. See, for example, U.S. Patent No. 20030087817, incorporated herein by reference.

[0114] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, the cells are transfected in a manner similar to that which occurs naturally in a subject. In some embodiments, the transfected cells are obtained from a subject.

[0115] In some embodiments, the transfected cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal cell (e.g., mammalian, insect, fish, bird, and reptile). In some embodiments, the transfected cell is a human cell. In some embodiments, the transfected cell is a cell from the hematopoietic system, such as an immune cell (i.e., a cell of the innate or adaptive immune system) (including, but not limited to, B cells, T cells, natural killer (NK) cells), pluripotent stem cells and induced pluripotent stem cells, chimeric antigen receptor T (CAR-T) cells, monocytes, macrophages, and dendritic cells.

[0116] In some embodiments, the cells are derived from cells, e.g., cell lines, taken from a subject. In some embodiments, the cells or cell lines are prokaryotic cells. In some embodiments, the cells or cell lines are eukaryotic cells. In further embodiments, the cells or cell lines are derived from insects, birds, plants, or fungi. In some embodiments, the cells or cell lines may be mammalian, e.g., human, monkey, mouse, cow, pig, goat, hamster, rat, cat, or dog. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFL, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182 , A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI-231, HB56, T IB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2 780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC Examples of suitable cell lines include NIH-69, MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic variants thereof. Cell lines are available from a variety of sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC), Manassas, Va.).

[0117] In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines containing one or more vector-derived sequences. In some embodiments, cells transiently transfected with a fusion protein of the invention and optionally a gRNA, or a ribonucleoprotein complex of the invention, and modified by the activity of the fusion protein or ribonucleoprotein complex, are used to establish new cell lines comprising cells containing the modification but lacking other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used to evaluate one or more test compounds.

[0118] In some embodiments, one or more vectors described herein are used to generate non-human transgenic animals or transgenic plants. In some embodiments, the transgenic animal is an insect. In further embodiments, the insect is a pest such as a mosquito or a tick. In some embodiments, the insect is a plant pest such as a corn rootworm or fall armyworm. In some embodiments, the transgenic animal is a bird, such as a chicken, turkey, goose, or duck. In some embodiments, the transgenic animal is a mammal, such as a human, mouse, rat, hamster, monkey, ape, rabbit, pig, cow, horse, goat, sheep, cat, or dog.

[0119] VI. Polypeptide and Polynucleotide Variants and Fragments The present disclosure provides novel adenine deaminases active against DNA molecules (the amino acid sequences of which are set forth in SEQ ID NOS: 1-10 and 399-441), active variants or fragments thereof, and polynucleotides encoding the same.

[0120] The activity of the variants or fragments may be altered compared to the polynucleotide or polypeptide of interest, but the variants and fragments should retain the functionality of the polynucleotide or polypeptide of interest. For example, the variants or fragments may have increased activity, decreased activity, a different range of activity, or any other altered activity compared to the polynucleotide or polypeptide of interest.

[0121] Fragments and variants of the deaminases of the present invention that have adenine deaminase activity retain said activity when they are part of a fusion protein that further comprises a DNA-binding polypeptide or a fragment thereof.

[0122] The term "fragment" refers to a portion of a polynucleotide or polypeptide sequence of the invention. A "fragment" or "biologically active portion" includes a polynucleotide comprising a sufficient number of contiguous nucleotides to retain biological activity (i.e., deaminase activity on nucleic acids). A "fragment" or "biologically active portion" includes a polypeptide comprising a sufficient number of contiguous amino acid residues to retain biological activity. Fragments of the deaminases disclosed herein include those that are shorter than the full-length sequence, for example, by using alternative downstream start sites. In some embodiments, a biologically active portion of a deaminase is a polypeptide comprising 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, or more contiguous amino acid residues of any of SEQ ID NOS: 1-10 and 399-441. Such biologically active portions may be prepared by recombinant techniques and evaluated for activity.

[0123] In general, "variant" is intended to mean a substantially similar sequence. In the case of polynucleotides, variants include deletions and / or additions of one or more nucleotides at one or more internal sites within the naturally occurring polynucleotide and / or substitutions of one or more nucleotides at one or more sites within the naturally occurring polynucleotide. As used herein, a "naturally occurring" or "wild-type" polynucleotide or polypeptide includes a naturally occurring nucleotide sequence or amino acid sequence, respectively. In the case of polynucleotides, conservative variants include sequences that, due to the degeneracy of the genetic code, encode the naturally occurring amino acid sequence of a gene of interest. Naturally occurring allelic variants such as these can be identified using well-known techniques of molecular biology, such as polymerase chain reaction (PCR) and hybridization techniques, as outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those generated using site-directed mutagenesis, but which still encode a polypeptide or polynucleotide of interest. Generally, variants of a particular polynucleotide disclosed herein will have at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity to that particular polynucleotide as determined by sequence alignment programs and parameters described elsewhere herein.

[0124] Variants of a particular polynucleotide (i.e., a reference polynucleotide) disclosed herein can also be evaluated by comparing the percent sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The percent sequence identity between any two polypeptides can be calculated using sequence alignment programs and parameters described elsewhere herein. When any given pair of polynucleotides disclosed herein is evaluated by comparing the percent sequence identity shared by the two polypeptides they encode, the percent sequence identity between the two encoded polypeptides is at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity.

[0125] In certain embodiments, a polynucleotide of the present disclosure encodes an adenine deaminase comprising an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more identity to the amino acid sequence of any of SEQ ID NOs: 1-10 and 399-441.

[0126] Biologically active variants of adenine deaminase of the present invention may differ by as few as 1-15 amino acid residues, as few as 1-10, e.g., 6-10, as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 amino acid residue. In certain embodiments, the polypeptides include N- or C-terminal truncations and may include deletions of at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids from either the N- or C-terminus of the polypeptide. In some embodiments, the polypeptides include internal deletions, which may include deletions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids.

[0127] It is recognized that the deaminases provided herein can be modified to create variant proteins and polynucleotides. Artificially designed changes can be introduced by applying site-directed mutagenesis techniques. In some embodiments, unknown or unidentified naturally occurring polynucleotides and / or polypeptides structurally and / or functionally related to the sequences disclosed herein may also be identified as being within the scope of the present invention. Conservative amino acid substitutions that do not alter the function of the polypeptide as an adenine deaminase may be made in non-conserved regions. In some embodiments, modifications are made that improve the adenine deaminase activity of the deaminase.

[0128] Variant polynucleotides and proteins also encompass sequences and proteins resulting from mutagenesis and recombination techniques, e.g., DNA shuffling. Such procedures are used to engineer one or more of the different deaminases disclosed herein (e.g., SEQ ID NOS: 1-10 and 399-441) to generate new adenine deaminases with desired properties. In this manner, libraries of recombinant polynucleotides are generated from a collection of related sequence polynucleotides that share substantial sequence identity and contain sequence regions that can be homologously recombined in vitro or in vivo. For example, this technique can be used to shuffle sequence motifs encoding domains of interest between the deaminase sequences provided herein and other subsequently identified deaminase genes to generate sequences with improved properties of interest, e.g., in the case of the enzyme K. mNew genes encoding proteins with increased activity may be obtained. Strategies for such DNA shuffling are known in the art. See, for example, Stemmer (1994) Proc. Natl. Acad. Sci. USA 91:10747-10751; Stemmer (1994) Nature 370:389-391; Crameri et al. (1997) Nature Biotech. 15:436-438; Moore et al. (1997) J. Mol. Biol. 272:336-347; Zhang et al. (1997) Proc. Natl. Acad. Sci. USA 94:4504-4509; Crameri et al. (1998) Nature 391:288-291; and U.S. Patent Nos. 5,605,793 and 5,837,458. "Shuffled" nucleic acids are nucleic acids produced by a shuffling procedure, such as any of the shuffling procedures described herein. Shuffled nucleic acids are produced by recombining (physically or virtually) two or more nucleic acids (or character strings), e.g., artificially, optionally in a recursive manner. Generally, the shuffling process uses one or more screening steps to identify nucleic acids of interest. This screening step can occur before or after any recombination step. In some (but not all) shuffling embodiments, it is desirable to perform multiple rounds of recombination before selection to increase the diversity of the pool being screened. The entire recombination and selection process is repeated, optionally recursively. Depending on the context, shuffling can refer to the overall process of recombination and selection, or simply the recombination portion of the overall process.

[0129] As used herein, "sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences refers to the residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When percentage sequence identity is used with respect to proteins, it is recognized that non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue with similar chemical properties (e.g., charge or hydrophobicity), thus not altering the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity may be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Means for making this adjustment are well known to those of skill in the art. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, conservative substitutions are given a score of 0-1, with identical amino acids being given a score of 1 and non-conservative substitutions being given a score of 0. Scoring of conservative substitutions is calculated, for example, as implemented in the PC / GENE program (Intelligenetics, Mountain View, California).

[0130] As used herein, "percentage of sequence identity" refers to a value determined by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (which contains no additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where the same nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.

[0131] Unless otherwise specified, the sequence identity / similarity values provided herein refer to values obtained using GAP Version 10 (using the following parameters: for nucleotide sequence % identity and % similarity, GAP Weight 50, Length Weight 3, and the nwsgapdna.cmp scoring matrix; for amino acid sequence % identity and % similarity, GAP Weight 8, Length Weight 2, and the BLOSUM62 scoring matrix), or any equivalent program. By "equivalent program" is intended any sequence comparison program that produces alignments with identical nucleotide or amino acid residue matches and identical percent sequence identity for any two sequences in question when compared to corresponding alignments produced by GAP Version 10.

[0132] Two sequences are "optimally aligned" when they are aligned using a predetermined amino acid substitution matrix (e.g., BLOSUM62), gap existence penalties, and gap extension penalties to achieve the highest possible score for the pair of sequences in terms of similarity score. Amino acid substitution matrices and their use to quantify similarity between two sequences are well known in the art and are described, for example, in Dayhoff et al. (1978) "A model of evolutionary change in proteins." (In "Atlas of Protein Sequence and Structure," Vol. 5, Suppl. 3 (ed. MO Dayhoff)), pp. 345-352. Natl. Biomed. Res. Found., Washington, DC, and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA 89:10915-10919. The BLOSUM62 matrix is often used as the default scoring substitution matrix in sequence alignment protocols. A gap existence penalty is imposed for introducing a single amino acid gap into one of the aligned sequences, and a gap extension penalty is imposed for introducing each additional empty amino acid position into an existing gap. The alignment is defined by the amino acid positions in each sequence where the alignment begins and ends, and optionally by inserting one or more gaps in one or both sequences to achieve the highest possible score. While optimal alignment and scoring can be achieved manually, this process can be facilitated using computer-implemented alignment algorithms, such as Gapped BLAST 2.0, described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402 and made publicly available on the National Center for Biotechnology Information website (www.ncbi.nlm.nih.gov).Optimal alignments containing multiple alignments may be prepared using, for example, PSI-BLAST, available at www.ncbi.nlm.nih.gov and described by Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402.

[0133] With respect to an amino acid sequence that is optimally aligned with a reference sequence, an amino acid residue "corresponds" to the position in the reference sequence with which it is paired in the alignment. A "position" is represented by a number that sequentially identifies each amino acid in the reference sequence based on its position relative to the N-terminus. Generally, the number of an amino acid residue in a test sequence, determined by simple counting from the N-terminus, is not necessarily the same as the number of its corresponding position in the reference sequence, due to deletions, insertions, truncations, fusions, etc., which must be considered when determining optimal alignment. For example, if an aligned test sequence contains a deletion, there will be no amino acid at the site of the deletion that corresponds to a position in the reference sequence. If an aligned reference sequence contains an insertion, the insertion will not correspond to any amino acid position in the reference sequence. In the case of a truncation or fusion, there may be a stretch of amino acids in either the reference sequence or the aligned sequence that does not correspond to any amino acid in the corresponding sequence.

[0134] VII. Antibodies Antibodies against the deaminases, fusion proteins, or ribonucleoproteins comprising the deaminases (including those having the amino acid sequence set forth as any one of SEQ ID NOS: 1-10 and 399-441, or active variants or fragments thereof) of the present invention are also encompassed. Methods for producing antibodies are well known in the art (see, e.g., Harlow and Lane (1988) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY; and U.S. Pat. No. 4,196,265). These antibodies can be used in kits for detecting and isolating the deaminases, or fusion proteins or ribonucleoproteins comprising the deaminases, described herein. Accordingly, the present disclosure provides kits comprising antibodies that specifically bind to the polypeptides or ribonucleoproteins described herein (e.g., including polypeptides comprising a sequence at least 85% identical to any one of SEQ ID NOS: 1-10 and 399-441).

[0135] VIII. SYSTEMS AND RIBONUCLEOPROTEIN COMPLEXES FOR BINDING TO AND / OR MODIFYING TARGET SEQUENCES OF INTEREST AND METHODS FOR PRODUCING SAME The present disclosure provides a system for targeting a nucleic acid sequence and modifying the target nucleic acid sequence. In some embodiments, an RNA-guided DNA-binding polypeptide, such as an RGN, and a gRNA are involved in targeting a ribonucleoprotein complex to a nucleic acid sequence of interest. A deaminase polypeptide fused to RGDBP is involved in modifying the target nucleic acid sequence from A>N. In some embodiments, the deaminase converts A>G. A guide RNA hybridizes to the target sequence of interest and further forms a complex with the RNA-guided DNA-binding polypeptide, thereby directing the RNA-guided DNA-binding polypeptide to the target sequence. The RNA-guided DNA-binding polypeptide is one domain of a fusion protein, and the second domain is a deaminase described herein. In some embodiments, the RNA-guided DNA-binding polypeptide is an RGN, such as Cas9. Other examples of RNA-guided DNA-binding polypeptides include RGNs, such as those described in WO 2019 / 236566 and WO 2020 / 139783. In some embodiments, the RNA-guided DNA-binding polypeptide is a type II CRISPR-Cas polypeptide, or an active variant or fragment thereof. In some embodiments, the RNA-guided DNA-binding polypeptide is a type V CRISPR-Cas polypeptide, or an active variant or fragment thereof. In some embodiments, the RNA-guided DNA-binding polypeptide is a type VI CRISPR-Cas polypeptide. In some embodiments, the DNA-binding domain of the fusion protein, e.g., a zinc finger nuclease, TALEN, or meganuclease polypeptide, does not require an RNA guide. In some embodiments, the nuclease activity of the DNA-binding domain is partially or completely inactivated. In further embodiments, the RNA-guided DNA-binding polypeptide comprises the amino acid sequence of RGN, e.g., APG07433.1 (SEQ ID NO: 41), or an active variant or fragment thereof, such as nickase nAPG07433.1 (SEQ ID NO: 42) or other nickase RGN variants described in the Examples (SEQ ID NOs: 52-59, 61, 397, and 398).

[0136] In some embodiments, the systems provided herein for binding to and / or modifying a target sequence of interest are ribonucleoprotein complexes that comprise at least one molecule of RNA bound to at least one protein. The ribonucleoprotein complexes provided herein comprise at least one guide RNA as the RNA component and a fusion protein comprising a deaminase of the present invention and an RNA-guided DNA-binding polypeptide as the protein component. In some embodiments, the ribonucleoprotein complex is purified from cells or organisms transformed with a polynucleotide encoding the fusion protein and the guide RNA and cultured under conditions that allow expression of the fusion protein and the guide RNA.

[0137] In various embodiments, a ribonucleoprotein complex is provided, comprising any of the fusion proteins described herein and a guide RNA linked to the DNA-binding polypeptide of the fusion protein. For example, provided herein is a ribonucleoprotein complex comprising a fusion protein with a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 407. In another example, provided is a ribonucleoprotein complex comprising a fusion protein with a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 399. In yet another example, provided is a ribonucleoprotein complex comprising a fusion protein with a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 405. In some embodiments of the above-described ribonucleoprotein complexes, the fusion protein is selected from the group consisting of CasX, CasY, C2c1, C2c2, C2c3, GeoCas9, SpCas9, SaCas9, Nme2Cas9, CjCas9, Cas12a (formerly known as Cpf1), Cas12b, Cas12g, Cas12h, Cas12i, LbCas12a, AsCas12a, CasMINI, Cas In some embodiments, the ribonucleoprotein complex comprises an RGN selected from Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368. In some embodiments, the ribonucleoprotein complex comprises a nickase having an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398 fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 407. In some embodiments, the ribonucleoprotein complex comprises a nickase having an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398 fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 399.In some embodiments, the ribonucleoprotein complex comprises a nickase having an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398 fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 405. In some embodiments, the ribonucleoprotein complex comprises a Cas9 nickase fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 407. In some embodiments, the ribonucleoprotein complex comprises a Cas9 nickase fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 399. In some embodiments, the ribonucleoprotein complex comprises a Cas9 nickase fused to a deaminase comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 405. The Cas9 nickase may be any Cas9 nickase disclosed in WO2020181195, the entire contents of which are incorporated herein by reference. In various embodiments described herein, the ribonucleoprotein complex may also include a gRNA as described herein.

[0138] Methods for producing a deaminase, a fusion protein, or a fusion protein-ribonucleoprotein complex are provided. Such methods include culturing cells containing a nucleotide sequence encoding the deaminase, the fusion protein (and in some embodiments, a nucleotide sequence encoding the guide RNA) under conditions in which the deaminase or the fusion protein (and in some embodiments, the guide RNA) is expressed. The deaminase, the fusion protein, or the fusion ribonucleoprotein may then be purified from a lysate of the cultured cells.

[0139] Methods for purifying deaminases, fusion proteins, or fusion ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., size exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reverse-phase chromatography, immunoprecipitation). In certain methods, the deaminase or fusion protein is recombinantly produced and includes a purification tag to aid in its purification. These include, but are not limited to, glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tags, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, biotin carboxyl carrier protein (BCCP), and calmodulin. Generally, the tagged deaminase, fusion protein, or fusion ribonucleoprotein complex is purified using immunoprecipitation or other similar methods known in the art.

[0140] An "isolated" or "purified" polypeptide, or biologically active portion thereof, is substantially or essentially free from components that normally accompany or interact with the polypeptide as found in its naturally occurring environment. Thus, an isolated or purified polypeptide is substantially free of other cellular material or culture medium if produced by recombinant techniques, or substantially free of chemical precursors or other chemicals if chemically synthesized. A protein that is substantially free of cellular material includes preparations of the protein having less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% (by dry weight) of contaminating proteins. When the proteins of the invention, or biologically active portions thereof, are recombinantly produced, an optimal culture medium represents less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% (by dry weight) of chemical precursors or chemicals other than the protein of interest.

[0141] Certain methods provided herein for binding to and / or cleaving a target sequence of interest involve the use of a ribonucleoprotein complex. In some embodiments, the ribonucleoprotein complex is assembled in vitro. In vitro assembly of the ribonucleoprotein complex can be performed using any method known in the art for contacting an RGDBP polypeptide or a fusion protein comprising the same with a guide RNA under conditions that allow the RGDBP polypeptide or a fusion protein comprising the same to bind to the guide RNA. As used herein, "contact," "contacting," or "contacted" refers to placing components of a desired reaction together under conditions suitable for the desired reaction to occur. In some embodiments of the described methods for modifying a target DNA molecule, the contacting step is performed in vitro. In some embodiments, the contacting step is performed in vivo. In some embodiments, the contacting step is performed in a subject (e.g., a human subject or a non-human animal subject). In some embodiments, the contacting step is performed in a cell, e.g., a human or non-human animal cell. The RGDBP polypeptide or a fusion protein comprising the same may be purified from a biological sample, cell lysate, or culture medium, produced by in vitro translation, or chemically synthesized. The guide RNA may be purified from a biological sample, cell lysate, or culture medium, transcribed in vitro, or chemically synthesized. The RGDBP polypeptide or a fusion protein comprising the RGDBP polypeptide and the guide RNA may be contacted in a solution (e.g., a buffered saline solution) to assemble a ribonucleoprotein complex in vitro.

[0142] IX. Methods for Modifying Target Sequences The present disclosure provides methods for modifying a target nucleic acid molecule of interest (e.g., a target DNA molecule). The method comprises delivering a fusion protein comprising a DNA-binding polypeptide and at least one deaminase of the present invention or a polynucleotide encoding same to the target sequence, or to a cell, organelle, or embryo containing the target sequence. In certain embodiments, the method comprises delivering a system comprising at least one guide RNA or a polynucleotide encoding same and at least one fusion protein comprising at least one deaminase of the present invention and an RNA-guided DNA-binding polypeptide or a polynucleotide encoding same to the target sequence, or to a cell, organelle, or embryo containing the target sequence. In some embodiments, the fusion protein comprises any one of the amino acid sequences set forth in SEQ ID NOs: 1-10 and 399-441, or an active variant or fragment thereof.

[0143] In some embodiments, the method comprises contacting a DNA molecule with (a) a fusion protein comprising a deaminase and an RNA-guided DNA-binding polypeptide (e.g., a nuclease-inactive or nickase Cas9 domain), and (b) a gRNA that targets the fusion protein of (a) to a target nucleotide sequence in the DNA molecule, wherein the DNA molecule is contacted with the fusion protein and the gRNA in an amount and under suitable conditions effective for deamination of nucleobases. In some embodiments, the target DNA molecule comprises a sequence associated with a disease or disorder, and deamination of the nucleobases results in a sequence not associated with the disease or disorder. In some embodiments, the disease or disorder affects an animal. In further embodiments, the disease or disorder affects a mammal, such as a human, cow, horse, dog, cat, goat, sheep, pig, monkey, rat, mouse, or hamster. In some embodiments, the target DNA sequence is present in an allele of a crop plant, and a particular allele of a trait of interest results in a plant with reduced agricultural value. Deamination of nucleobases can result in alleles that improve the trait, increasing the agricultural value of the plant.

[0144] In embodiments where the method includes delivering a polynucleotide encoding a guide RNA and / or a fusion protein, the cell or embryo may then be cultured under conditions where the guide RNA and / or fusion protein is expressed. In various embodiments, the method includes contacting the target sequence with a ribonucleoprotein complex comprising a gRNA and a fusion protein (comprising a deaminase of the invention and an RNA-guided DNA-binding polypeptide). In certain embodiments, the method includes introducing a ribonucleoprotein complex of the invention into a cell, organelle, or embryo containing the target sequence. The ribonucleoprotein complex of the invention may be purified from a biological sample, recombinantly produced and then purified, or constructed in vitro, as described herein. In embodiments where the ribonucleoprotein complex contacting the target sequence, cell, organelle, or embryo is constructed in vitro, the method may further include assembling the complex in vitro prior to contacting the target sequence, cell, organelle, or embryo.

[0145] Purified or in vitro assembled ribonucleoprotein complexes of the invention can be introduced into cells, organelles, or embryos using any method known in the art, including, but not limited to, electroporation. In some embodiments, a fusion protein comprising a deaminase and an RNA-guided DNA-binding polypeptide of the invention and a polynucleotide encoding or comprising a guide RNA is introduced into a cell, organelle, or embryo using any method known in the art, such as, for example, electroporation.

[0146] Upon delivery to or contact with a target sequence, or a cell, organelle, or embryo containing the target sequence, the guide RNA directs and binds the fusion protein to the target sequence in a sequence-specific manner. The target sequence can then be modified via the deaminase domain of the fusion protein. In some embodiments, binding of the fusion protein to the target sequence results in modification of nucleotides adjacent to the target sequence. The nucleobases adjacent to the target sequence that are modified by the deaminase can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 base pairs from the 5' or 3' end of the target sequence. A fusion protein comprising a deaminase of the invention and an RNA-guided DNA-binding polypeptide is capable of introducing targeted A>N, preferably targeted A>G, mutations into a target DNA molecule.

[0147] In some embodiments of the described methods of modifying a target DNA molecule, the contacting step is performed in vitro. In certain embodiments, the contacting step is performed in vivo. In some embodiments, the contacting step is performed in a subject (e.g., a human subject or a non-human animal subject). In some embodiments, the contacting step is performed in a cell, e.g., a human or non-human animal cell.

[0148] Methods for measuring binding of fusion proteins to target sequences are known in the art and include chromatin immunoprecipitation assays, gel mobility shift assays, DNA pull-down assays, reporter assays, and microplate capture and detection assays. Similarly, methods for measuring cleavage or modification of target sequences are known in the art and include in vitro or in vivo cleavage assays that confirm cleavage using PCR, sequencing, or gel electrophoresis, with or without attaching appropriate labels (e.g., radioisotopes, fluorescent substances) to the target sequence to facilitate detection of degradation products. In some embodiments, a nicking-triggered exponential amplification reaction (NTEXPAR) assay is used (see, e.g., Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage may also be assessed using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).

[0149] In some embodiments, this method involves the use of an RNA-binding DNA-guided domain complexed with two or more guide RNAs as part of a fusion protein.Two or more guide RNAs can target different regions of a single gene, or can target multiple genes.This multiple targeting allows the deaminase domain of the fusion protein to modify nucleic acid, thereby introducing multiple mutations into the target nucleic acid molecule of interest (e.g., genome).

[0150] In embodiments where the method involves the use of an RNA-guided nuclease (RGN), such as a nickase RGN (i.e., capable of cleaving only one strand of a double-stranded polynucleotide, e.g., nAPG07433.1 (SEQ ID NO: 42 or SEQ ID NOs: 50-57), the method may include targeting the same or overlapping target sequence and introducing two different RGN or RGN variants that cleave different strands of the polynucleotide. For example, an RGN nickase that cleaves only the plus (+) strand of a double-stranded polynucleotide can be introduced along with a second RGN nickase that cleaves only the minus (-) strand of the double-stranded polynucleotide. In some embodiments, two different fusion proteins are provided, each containing a different RGN with a different PAM recognition sequence, such that a greater variety of nucleotide sequences can be targeted for mutation.

[0151] Those skilled in the art will understand that any of the methods disclosed herein can be used to target a single target sequence or multiple target sequences. Thus, the methods involve the use of a fusion protein comprising a single RNA-guided DNA-binding polypeptide in combination with multiple distinct guide RNAs that can target multiple distinct sequences within a single gene and / or multiple genes. The deaminase domain of the fusion protein then introduces mutations into each of the target sequences. Also encompassed herein are methods of introducing multiple distinct guide RNAs in combination with multiple distinct RNA-guided DNA-binding polypeptides. Such RNA-guided DNA-binding polypeptides can be multiple RGNs or RGN variants. These guide RNAs and guide RNA / fusion protein systems can target multiple distinct sequences within a single gene and / or multiple genes.

[0152] In some embodiments, fusion proteins comprising an RNA-guided DNA-binding polypeptide and a deaminase polypeptide of the present invention can be used to generate mutations in a target gene or target region of a gene of interest. In some embodiments, the fusion proteins of the present invention can be used for saturation mutagenesis of a target gene or region of a target gene of interest, followed by high-throughput forward genetic screening to identify novel mutations and / or phenotypes. In some embodiments, the fusion proteins described herein can be used to generate mutations at target genomic locations that may or may not contain coding DNA sequences. Libraries of cell lines generated by the above-described targeted mutagenesis can also be useful for studying gene function or gene expression.

[0153] X. Target Polynucleotide In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method comprises harvesting a cell or population of cells from a human or non-human animal or plant (including microalgae) and modifying one or more of the cells. Culturing can be performed ex vivo at any stage. The one or more cells may then be reintroduced into a human, non-human animal, or plant (including microalgae).

[0154] Plant breeders exploit natural diversity to combine genes that are most useful for desirable qualities, such as yield, quality, uniformity, durability, and pest resistance. These desirable qualities include growth, photoperiod preference, temperature requirements, the onset of flowering or reproductive development, fatty acid content, insect resistance, disease resistance, nematode resistance, fungus resistance, herbicide resistance, and tolerance to various environmental factors (including adverse soil conditions such as drought, heat, humidity, cold, wind, and high salinity). Sources of these useful genes include native or exotic varieties, heirloom species, wild plant relatives, and induced mutations, such as treating plant material with mutagens. The present invention provides plant breeders with a new tool for inducing mutagenesis. Thus, those skilled in the art can use the present invention to induce increases in useful genes more precisely than previous mutagenizing agents, thereby facilitating and improving plant breeding programs.

[0155] The target polynucleotide of a deaminase or fusion protein of the present invention can be any polynucleotide, endogenous or exogenous to a eukaryotic cell. For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. In some embodiments, the target polynucleotide is a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some embodiments, the target sequence of a fusion protein of the present invention is associated with a PAM (protospacer adjacent motif), i.e., a short sequence recognized by an RNA-guided DNA-binding polypeptide. While the exact sequence and length requirements of the PAM vary depending on the RNA-guided DNA-binding polypeptide used, the PAM is typically a 2-5 base pair sequence adjacent to the protospacer (i.e., the target sequence).

[0156] The target polynucleotide of the fusion protein of the present invention can include numerous disease-associated genes and polynucleotides, as well as signal transduction biochemical pathway-related genes and polynucleotides. Examples of target polynucleotides include sequences associated with signal transduction biochemical pathways, such as signal transduction biochemical pathway-related genes or polynucleotides. Examples of target polynucleotides include disease-associated genes or polynucleotides. A "disease-associated" gene or polynucleotide refers to any gene or polynucleotide whose transcription or translation product is produced at an abnormal level or in an abnormal form in cells derived from diseased tissue compared to non-diseased control tissue or cells. This may be a gene that becomes expressed at an abnormally high level or an abnormally low level, and the change in expression correlates with the onset and / or progression of the disease. A disease-associated gene also refers to a gene that is directly involved in the etiology of the disease (e.g., a causative mutation) or that has a mutation or genetic variation in linkage disequilibrium with a gene involved in the etiology of the disease (e.g., a causative mutation). The transcription or translation product may be known or unknown, and may be at a normal or abnormal level.

[0157] Non-limiting examples of disease-associated genes that can be targeted using the methods and compositions of the present disclosure are provided in Table 34. In some embodiments, the disease-associated gene targeted is a gene disclosed in Table 34 that has a G>A mutation. Further examples of disease-associated genes and polynucleotides are available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.).

[0158] In some embodiments, the target polynucleotide comprises the cystic fibrosis transmembrane conductance regulator (5) gene.

[0159] As used herein, the term "cystic fibrosis transmembrane conductance regulator" or "CFTR" refers to a cAMP-regulated chloride channel located in the apical membrane of epithelial cells that catalyzes the passage of small ions across the membrane. A non-limiting example of a CFTR gene is set forth as SEQ ID NO:51.

[0160] As used herein, the terms "target" or "targeting" with respect to a spacer sequence and a target sequence refer to the localization of an RNA-guided nuclease to a target sequence based on the ability of the spacer sequence in the associated guide RNA to sufficiently hybridize with the target sequence.

[0161] CRISPR RNA (crRNA) or a nucleic acid molecule encoding the same is provided, wherein the crRNA includes a spacer sequence that targets a CFTR target sequence. Also provided are guide RNAs comprising such crRNAs, one or more nucleic acid molecules encoding guide RNAs comprising such crRNAs, vectors comprising one or more nucleic acid molecules encoding guide RNAs comprising such crRNAs, and systems comprising such crRNAs. Also provided are methods for binding to, cleaving, and / or regulating a target sequence using such crRNAs or a nucleic acid molecule encoding the same, guide RNAs comprising such crRNAs, one or more nucleic acid molecules encoding guide RNAs comprising such crRNAs, vectors comprising one or more nucleic acid molecules encoding guide RNAs comprising such crRNAs, and systems comprising such crRNAs.

[0162] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, 562, and 563, or a complement thereof. In some embodiments, a single guide RNA (sgRNA) comprising a crRNA with a spacer sequence targeting a CFTR target sequence has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0163] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 62-68, 80-85, 116-119, 128-131, 163, 164, 180, 181, 203-209, 219-225, 256-258, 274-276, 310-313, and 330-333, or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 53. In some embodiments, the sgRNA comprising the crRNA with a spacer sequence targeting the CFTR target sequence has a sequence similar to any one of SEQ ID NOs: 98-104, 140-143, 197, 198, 235-241, 292-294, and 350-353 by at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, or 94% of any one of SEQ ID NOs: 98-104, 140-143, 197, 198, 235-241, 292-294, and 350-353. , 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:53, and related RGN polypeptides have an amino acid sequence that has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:53.

[0164] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 68-71, 86-89, 120-122, 132-134, 152-156, 169-173, 213-215, 229-231, 251-255, 269-273, 305-309, and 325-329, or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 55. In some embodiments, the sgRNA comprising the crRNA with a spacer sequence targeting the CFTR target sequence has a sequence similar to any one of SEQ ID NOs: 104-107, 144-146, 186-190, 245-247, 287-291, and 345-349 by at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, or 94% of any one of SEQ ID NOs: 104-107, 144-146, 186-190, 245-247, 287-291, and 345-349. , 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:55, and related RGN polypeptides have an amino acid sequence that has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:55.

[0165] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 72, 73, 90, 91, 161, 162, 178, 179, 265, 266, 283, and 284, or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 52. In some embodiments, an sgRNA comprising a crRNA with a spacer sequence targeting a CFTR target sequence has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 108, 109, 195, 196, 301, and 302, and an associated RGN polypeptide has an amino acid sequence with at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 52.

[0166] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 74, 75, 92, 93, 123, 124, 135, 136, 167, 184, 216-218, 232-234, 259-261, 277-279, 314-317, and 334-337, or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 56. In some embodiments, the sgRNA comprising the crRNA with a spacer sequence targeting a CFTR target sequence has a sequence identity that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110, 111, 147, 148, 201, 248-250, 295-297, and 354-357. and related RGN polypeptides have an amino acid sequence that has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:56.

[0167] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 76, 94, 210-212, 226-228, 322, 342, 562, and 563, or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 42. In some embodiments, an sgRNA comprising a crRNA with a spacer sequence targeting a CFTR target sequence has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 112, 242-244, 362, and 564, and an associated RGN polypeptide has an amino acid sequence with at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 42.

[0168] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 77, 95, 125, 137, 157-160, 174-177, 323, and 343, or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 54. In some embodiments, an sgRNA comprising a crRNA with a spacer sequence targeting a CFTR target sequence has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 113, 149, 191-194, and 363, and an associated RGN polypeptide has an amino acid sequence with at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 54.

[0169] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 78, 96, 126, 138, 168, 185, 267, 285, 318, 319, 338, and 339 or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 57. In some embodiments, an sgRNA comprising a crRNA with a spacer sequence that targets a CFTR target sequence has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 114, 150, 202, 303, 358, and 359, and an associated RGN polypeptide has an amino acid sequence with at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 57.

[0170] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 79, 97, 127, 139, 262-264, 280-282, 324, and 344, or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 58. In some embodiments, an sgRNA comprising a crRNA with a spacer sequence targeting a CFTR target sequence has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 115, 151, 298-300, and 364, and an associated RGN polypeptide has an amino acid sequence with at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 58.

[0171] In some embodiments, the CFTR target sequence of the crRNA or guide RNA has a sequence set forth in any one of SEQ ID NOs: 165, 166, 182, 183, 268, 286, 320, 321, 340, and 341 or a complement thereof, and the associated RGN polypeptide has an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 59. In some embodiments, an sgRNA comprising a crRNA with a spacer sequence that targets a CFTR target sequence has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 199, 200, 304, 360, and 361, and an associated RGN polypeptide has an amino acid sequence with at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 59.

[0172] In some embodiments, the method comprises contacting a DNA molecule comprising a target DNA sequence with a DNA-binding polypeptide-deaminase fusion protein of the invention, wherein the DNA molecule is contacted with the fusion protein in an amount and under suitable conditions effective for nucleobase deamination. In certain embodiments, the method comprises contacting a DNA molecule comprising the target DNA sequence with (a) an RGN-deaminase fusion protein of the invention and (b) a gRNA that targets the fusion protein of (a) to a target nucleotide sequence in a DNA strand, wherein the DNA molecule is contacted with the fusion protein and gRNA in an amount and under suitable conditions effective for nucleobase deamination. In some embodiments, the target DNA sequence comprises a sequence associated with a disease or disorder, and nucleobase deamination results in a sequence that is not associated with the disease or disorder. In some embodiments, the target DNA sequence is present in an allele of a crop plant, and a particular allele for a trait of interest results in a plant with reduced agronomic value. Nucleobase deamination results in an allele that improves the trait, thereby increasing the agricultural value of the plant.

[0173] In some embodiments, the target DNA sequence contains a G>A point mutation associated with a disease or disorder, and deamination of the mutant A base results in a sequence that is not associated with the disease or disorder. In some embodiments, deamination corrects the point mutation in the sequence associated with the disease or disorder.

[0174] In some embodiments, the disease or disorder-associated sequence encodes a protein, and deamination introduces a stop codon into the disease or disorder-associated sequence, truncating the encoded protein. In some embodiments, the contacting is performed in vivo in a subject susceptible to, suffering from, or diagnosed with a disease or disorder. In some embodiments, the disease or disorder is a disease associated with a point mutation or single base mutation in the genome. In some embodiments, the disease is a genetic disease, cancer, metabolic disease, or lysosomal storage disease.

[0175] XI. Pharmaceutical Compositions and Methods of Treatment Provided herein is a method of treating a disease in a subject in need thereof, comprising administering to the subject in need thereof an effective amount of a fusion protein of the present disclosure or a polynucleotide encoding same, a gRNA of the present disclosure or a polynucleotide encoding same, a fusion protein system of the present disclosure, a ribonucleoprotein complex of the present disclosure, or a cell modified by or comprising any one of these compositions.

[0176] In some embodiments, treatment involves in vivo gene editing by administering a fusion protein, gRNA, or fusion protein system of the present disclosure, or a polynucleotide encoding the same, to a subject in need thereof. In some embodiments, treatment involves ex vivo gene editing, in which cells are genetically modified ex vivo with a fusion protein, gRNA, or fusion protein system of the present disclosure, or a polynucleotide encoding the same, and then the modified cells are administered to the subject. In some embodiments, the genetically modified cells are derived from the subject to whom the modified cells will subsequently be administered, and the transplanted cells are referred to herein as autologous. In some embodiments, the genetically modified cells are derived from a different subject (i.e., donor) within the same species as the subject to whom the modified cells will be administered (i.e., recipient), and the transplanted cells are referred to herein as allogeneic. In some examples described herein, the cells may be expanded in culture and then administered to a subject in need thereof.

[0177] For example, in some embodiments, methods are provided that include administering to a subject having such a disease (e.g., a genetic defect associated with the CFTR gene) an effective amount of a ribonucleoprotein complex comprising a fusion protein with a deaminase having an amino acid sequence at least 80% identical to the sequence set forth in any one of SEQ ID NOS: 399 and 405-407. In embodiments described herein, administration of the ribonucleoprotein complex corrects a point mutation or introduces an inactivating mutation in the disease-associated CFTR gene. Other diseases that can be treated by correcting a point mutation or introducing an inactivating mutation into a disease-associated gene are known to those of skill in the art, and the disclosure is not limited in this respect.

[0178] In some embodiments, the disease treated with the compositions of the present disclosure is a disease treatable with immunotherapy, such as chimeric antigen receptor (CAR) T cells, including, but not limited to, cancer.

[0179] In some embodiments, deamination of the targeted nucleobase corrects a genetic defect (e.g., for correction of the CFTR gene) or a point mutation that results in a loss of function in the gene product. In some embodiments, the genetic defect is associated with a disease or disorder, such as a lysosomal storage disease or a metabolic disease, such as type I diabetes. Thus, in some embodiments, the disease treated with the compositions of the present disclosure is associated with a sequence that is mutated (i.e., the sequence is causative of the disease or disorder or is responsible for symptoms associated with the disease or disorder) to treat the disease or disorder or alleviate symptoms associated with the disease or disorder.

[0180] In some embodiments, the disease treated with the compositions of the present disclosure is associated with a causative mutation. As used herein, a "causative mutation" refers to a specific nucleotide or nucleotides or nucleotide sequence in a genome that is responsible for the severity or presence of a disease or disorder in a subject. Correcting the causative mutation ameliorates at least one symptom caused by the disease or disorder. In some embodiments, correcting the causative mutation ameliorates at least one symptom caused by the disease or disorder. In some embodiments, the causative mutation is adjacent to a PAM site recognized by an RGDBP (e.g., RGN) fused to a deaminase disclosed herein. The causative mutation can be corrected with a fusion polypeptide comprising an RGDBP (e.g., RGN) and a deaminase of the present disclosure. Non-limiting examples of diseases associated with causative mutations include cystic fibrosis, Hurler syndrome, Friedreich's ataxia, Huntington's disease, and sickle cell disease. Further non-limiting examples of disease-associated genes and mutations are available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.).

[0181] In some embodiments, the methods provided herein are used to introduce inactivating point mutations into genes or alleles encoding gene products associated with a disease or disorder. For example, in some embodiments, methods are provided herein for introducing inactivating point mutations into cancer genes (e.g., in the treatment of proliferative disorders) using fusion proteins. Inactivating mutations, in some embodiments, can generate premature stop codons in the coding sequence, resulting in the expression of truncated gene products, e.g., truncated proteins lacking the function of the full-length protein. In some embodiments, the goal of the methods provided herein is to restore the function of dysfunctional genes through genome editing. The fusion proteins provided herein can be validated for in vitro human therapy based on gene editing, e.g., by correcting disease-associated mutations in human cell culture. One of skill in the art will understand that the fusion proteins provided herein, e.g., fusion proteins comprising an RNA-guided DNA-binding polypeptide and a deaminase polypeptide, can be used to correct any single-point G>A mutation. The mutation is corrected by deaminating the mutant A to G.

[0182] As used herein, "treatment" or "treating" or "alleviating" or "ameliorating" are used interchangeably. These terms refer to an approach for obtaining beneficial or desired results, including, but not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit refers to any treatment-related improvement in or effect on one or more diseases, conditions, or symptoms being treated. For prophylactic benefit, compositions may be administered to subjects at risk of developing a particular disease, condition, or symptom, or to subjects reporting one or more physiological symptoms of a disease, even if the disease, condition, or symptom has not yet manifested. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In certain embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or to inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., taking into account the history of the condition and / or taking into account genetic or other susceptibility factors). Treatment may also be continued after symptoms have disappeared, for example to prevent or delay their recurrence.

[0183] The term "effective amount" or "therapeutically effective amount" refers to an amount of an agent sufficient to produce a beneficial or desired result. A therapeutically effective amount may vary depending on one or more of the subject and condition being treated, the subject's weight and age, the severity of the condition, the method of administration, etc., and can be readily determined by one of ordinary skill in the art. The specific dose may vary depending on one or more of the particular agent selected, the dosing regimen to be followed, whether it is administered in combination with other compounds, the timing of administration, and the delivery system in which it is delivered.

[0184] The term "administering" refers to placement of an active ingredient in a subject by a method or route that at least partially localizes the introduced active ingredient at a desired site, e.g., a site of injury or repair, thereby producing a desired effect. In some embodiments, the present disclosure provides methods comprising delivering any of the isolated polypeptides, nucleic acid molecule fusion proteins, ribonucleoprotein complexes, vectors, pharmaceutical compositions, and / or gRNAs described herein. In some embodiments, the present disclosure further provides cells produced by such methods, and organisms (e.g., animals or plants) comprising or produced from such cells. In some embodiments, the deaminases, fusion proteins, and / or nucleic acid molecules described herein are delivered to cells in combination with (and optionally complexed with) a guide sequence.

[0185] In some embodiments, administration includes administration via viral delivery. Viral vectors containing nucleic acids encoding the fusion proteins, ribonucleoprotein complexes, or vectors disclosed herein may be administered directly to a patient (i.e., in vivo) or used to treat cells in vitro, with the modified cells optionally administered to a patient (i.e., ex vivo). Traditional viral-based systems include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated viral, and herpes simplex viral vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated viral gene transfer methods, often resulting in long-term expression of the inserted transgene. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. For applications where transient expression is preferred, adenoviral-based systems may be used. Adenoviral-based vectors are capable of very high transduction efficiency in many cell types and do not require cell division.

[0186] In some embodiments, administering comprises administering by electroporation. In some embodiments, administering comprises administering by nanoparticle delivery. In some embodiments, administering comprises administering by liposome delivery. Any effective route of administration can be used to administer an effective amount of the pharmaceutical compositions described herein.

[0187] In some embodiments, administration includes administration via other non-viral delivery of nucleic acids. Exemplary non-viral delivery methods include, but are not limited to, RNP complexes, lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycations or lipid-nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those of Felgner, International Publication Nos. WO 1991 / 17424 and WO 1991 / 16024. Delivery can be to cells (eg, in vitro or ex vivo administration) or to target tissues (eg, in vivo administration).

[0188] As used herein, the term "subject" refers to any individual for whom diagnosis, treatment, or therapy is desired. In some embodiments, the subject is an animal. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0189] The effectiveness of treatment can be determined by a skilled clinician. However, treatment is considered "effective treatment" if any or all of the signs or symptoms of the disease or disorder are beneficially altered (e.g., reduced by at least 10%) or other clinically recognized symptoms or markers of the disease are improved or alleviated. Efficacy can also be measured by the failure of an individual to deteriorate (e.g., the progression of the disease is stopped or at least slowed) as assessed by the need for hospitalization or medical intervention. Methods for measuring these indicators are known to those skilled in the art. Treatment includes: (1) inhibiting the disease, e.g., halting or slowing the progression of symptoms; or (2) ameliorating the disease, e.g., causing regression of symptoms; and (3) preventing or reducing the likelihood of onset of symptoms.

[0190] Pharmaceutical compositions are provided comprising an RGN polypeptide of the present disclosure or a polynucleotide encoding it, a gRNA of the present disclosure or a polynucleotide encoding it, a deaminase of the present disclosure or a polynucleotide encoding it, a fusion protein of the present disclosure, a system of the present disclosure (such as one comprising a fusion protein), a ribonucleoprotein complex of the present disclosure, or a cell comprising an RGN polypeptide or a polynucleotide encoding RGN, a gRNA or a polynucleotide encoding a gRNA, a polynucleotide encoding a fusion protein, or the system, and a pharmaceutically acceptable carrier.

[0191] As used herein, "pharmaceutically acceptable carrier" refers to a substance that does not cause significant irritation to an organism and does not impair the activity and properties of the active ingredient (i.e., the deaminase or fusion protein or nucleic acid molecule encoding it). The carrier must be of sufficiently high purity and sufficiently low toxicity to make it suitable for administration to the subject being treated. The carrier may be inert or may have pharmaceutical benefits. In some embodiments, a pharmaceutically acceptable carrier comprises one or more compatible solid or liquid fillers, diluents, or encapsulating substances suitable for administration to humans or other vertebrates. In some embodiments, a pharmaceutical composition comprises a non-naturally occurring pharmaceutically acceptable carrier. In some embodiments, the pharmaceutically acceptable carrier and the active ingredient are not found together in nature and are therefore heterologous.

[0192] Pharmaceutical compositions used in the methods of the present disclosure may be formulated with suitable carriers, excipients, and other agents to provide suitable transport, delivery, tolerability, etc. Many suitable formulations are known to those skilled in the art. See, for example, Remington, *The Science and Practice of Pharmacy* (21st ed. 2005). Non-limiting examples include: sterile diluents such as water for injection, saline solution, fixed oils, polyethylene glycols, glycerin, propylene glycol, or other synthetic solvents; antibacterial agents such as benzyl alcohol or methylparabens; antioxidants such as ascorbic acid or sodium bisulfite; chelating agents such as ethylenediaminetetraacetic acid; buffers such as acetates, citrates, or phosphates, and isotonicity adjusting agents such as sodium chloride or dextrose. For intravenous administration, particular carriers are saline or phosphate-buffered saline (PBS). Pharmaceutical compositions for oral or parenteral use may be prepared in a unit dosage form appropriate for the dose of the active ingredient. Such dosage forms in unit doses include, for example, tablets, pills, capsules, injections (ampoules), suppositories, etc. These compositions may also contain adjuvants such as preservatives, wetting agents, emulsifying agents, and dispersing agents. The action of microorganisms can be prevented by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, sorbic acid, etc. It may also be desirable to include isotonic agents, for example, sugars, sodium chloride, etc. Sustained absorption of injectable pharmaceutical forms can be achieved by the use of agents delaying absorption, for example, aluminum monostearate and gelatin.

[0193] In some embodiments in which cells containing or modified by the disclosed RGN, gRNA, deaminase, fusion protein, system (including those containing fusion proteins), or polynucleotides encoding same are administered to a subject, the cells are administered as a suspension with a pharmaceutically acceptable carrier. Those skilled in the art will recognize that pharmaceutically acceptable carriers used in cell compositions do not contain buffers, compounds, cryopreservatives, preservatives, or other agents in amounts that substantially interfere with the viability of the cells delivered to a subject. Cell-containing formulations may also include, for example, an osmotic buffer that allows for maintenance of cell membrane integrity, and optionally, nutrients to maintain cell viability or enhance engraftment upon administration. Such formulations and suspensions are known to those skilled in the art and / or can be adapted for use with the cells described herein using routine experimentation.

[0194] The cell compositions may also be emulsified or presented as liposomal compositions, provided that the emulsification procedure does not adversely affect cell viability. The cells and any other active ingredients may be mixed with excipients that are pharmaceutically acceptable and compatible with the active ingredients, in amounts suitable for use in the therapeutic methods described herein.

[0195] Additional agents included in the cell compositions may include pharmaceutically acceptable salts of the components therein. Pharmaceutically acceptable salts include acid addition salts (formed with the free amino groups of the polypeptide) formed with inorganic acids (e.g., hydrochloric acid or phosphoric acid) or organic acids (e.g., acetic acid, tartaric acid, mandelic acid, etc.). Salts formed with the free carboxyl groups may also be derived from inorganic bases (e.g., sodium hydroxide, potassium hydroxide, ammonium hydroxide, calcium hydroxide, or ferric hydroxide) and organic bases (e.g., isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, procaine, etc.).

[0196] Suitable routes of administration of the pharmaceutical compositions described herein include, but are not limited to, topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, oral, intracochlear, intratympanic, intravisceral, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseous, periocular, intratumoral, intracerebral, and intraventricular administration.

[0197] In some embodiments, the pharmaceutical compositions described herein are administered locally to the affected site (e.g., lungs). In some embodiments, the pharmaceutical compositions described herein are administered to a subject by injection, inhalation (e.g., aerosol), catheter, suppository, or implant. Implants are porous, non-porous, or gelatinous materials, including membranes such as sialastic membranes or fibers. In some embodiments, the pharmaceutical compositions are formulated for delivery to a subject, for example, for gene editing.

[0198] In some embodiments, pharmaceutical compositions are formulated according to routine procedures as compositions adapted for intravenous or subcutaneous administration to a subject, e.g., a human. In some embodiments, pharmaceutical compositions for administration by injection are solutions in sterile isotonic aqueous buffer. If necessary, the medicament can also include a solubilizing agent and a local anesthetic, such as lignocaine, to relieve pain at the injection site. Generally, the ingredients are supplied, either separately or mixed together, as a dry lyophilized powder or water-free concentrate in a hermetically sealed container, such as an ampoule or sachet indicating the quantity of active agent. When the medicament is administered by infusion, it can be dispensed using an infusion bottle containing pharmaceutical-grade sterile water or saline. When the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration.

[0199] In some embodiments, the pharmaceutical composition may be contained within a lipid particle or pore, such as a liposome or microcrystal, which may also be suitable for parenteral administration.

[0200] Although the description of pharmaceutical compositions provided herein primarily relates to pharmaceutical compositions suitable for administration to humans, those skilled in the art will understand that such compositions are generally suitable for administration to all types of animals or organisms.

[0201] Modifying the causative mutation using base editing An example of a genetic disease that can be corrected using a method that relies on the RGN-deaminase fusion protein of the present invention is cystic fibrosis. Cystic fibrosis (CF) is an autosomal recessive disease caused by mutations in the cystic fibrosis transmembrane conductance regulator (CFTR) gene (shown as SEQ ID NO: 51). CFTR encodes a cAMP-regulated chloride channel located in the apical membrane of epithelial cells, catalyzing the passage of small ions across the membrane. Dysregulation of this mechanism disrupts salt and fluid homeostasis, leading to multiple organ failure and ultimately to death from respiratory failure.

[0202] Approximately 2,000 mutations in the CFTR gene have been identified as the cause of CF. CFTR mutations are classified into six classes based on their impairment of either CFTR protein synthesis, transport, function, or stability, although many CFTR variants have been recognized to present with multiple disorders. Class I mutations result in the production of severely defective proteins. These are primarily nonsense or frameshift mutations that introduce premature termination codons (PTCs), resulting in unstable mRNAs that are degraded by the messenger RNA (mRNA) decay pathway (NMD). Nonsense mutations due to single nucleotide changes comprise a major subset of Class I mutations (Marangi, M. and Pistritto, G, 2018, Front Pharmacol 9, 396, doi:10.3389 / fphar.2018.00396; Pranke, I., et al., 2019, Front Pharmacol 10, 121, doi:10.3389 / fphar.2019.00121, both incorporated herein by reference). Treating patients with Class I cystic fibrosis can be difficult because functional CFTR protein is not produced. Notably, a significant proportion of these nonsense mutations could potentially be addressed by A-to-G base editors (Geurts, MH et al, 2020, Cell Stem Cell 26, 503-510 e507, doi:10.1016 / j.stem.2020.01.019, incorporated herein by reference).

[0203] Geurts et al. were the first group to perform precise base editing in cultured lung epithelial cells carrying class I mutations from cystic fibrosis patients using a fusion protein containing adenine deaminase operably linked to RGN, either SpyCas9 or an xSpyCas9 variant. SpyCas9 recognizes the 5'-nGG-3' PAM, whereas the xSpyCas9 variant recognizes the shortened 5'-nG-3'. The authors state that the main limitation of base editing technology is the PAM requirement of the Cas protein used. They found that many of the nonsense mutations identified in the CFTR gene were not within the target window required for fusion proteins containing RGN SpyCas9. PAMs are short motifs (approximately 1-4 nucleotides) in the target DNA sequence recognized by RGN. PAM sequences are unique to each RGN protein, and as a result, RGN can only reach the genomic space surrounding a suitable PAM. Furthermore, the base editing window of base editors is limited, often confined to only a portion of the nucleotides in the target sequence. If the target nucleotide is too close to the PAM, RGN blocks access to the nucleotide. If the nucleotide is too far from the PAM, the deaminase linked to RGN cannot reach the nucleotide. The amount of ssDNA exposed by the R-loop also limits the accessibility of the deaminase. The present invention includes an RGN-deaminase fusion protein in which RGN recognizes the PAM adjacent to a Class I mutation in the CFTR gene, allowing the deaminase to successfully modify the targeted causative mutation.

[0204] Another limitation of RGN-deaminase fusion proteins known in the art is that the vector constructs encoding the fusion proteins are too large for in vivo delivery methods. AAV delivery of these fusion proteins is not an option for SpyCas9-based fusion proteins because their size exceeds the limit of efficient AAV packaging. The RGN component of the fusion proteins described herein is a viable candidate for AAV vector delivery strategies due to its smaller size. The present invention also discloses guide RNAs specific to the RGN described herein, which guide the fusion proteins of the present invention to previously inaccessible nonsense mutation target sites in the CFTR gene. The present invention also teaches methods for using the above fusion proteins for targeted base editing via in vivo AAV vector delivery.

[0205] Ideally, the coding sequences for the RGN-deaminase fusion protein of the present invention and the corresponding guide RNA for targeting the fusion protein to the CFTR gene can all be packaged into a single AAV vector. The general size limit for AAV vectors is 4.7 kb, but larger sizes can be contemplated by reducing packing efficiency. The RGN nickases in Table 28 have coding sequence lengths of approximately 3.15 to 3.45 kB. Novel activity-deleted variants of RGN are described herein to ensure that the expression cassettes for both the fusion protein and its corresponding guide RNA fit into an AAV vector. In addition to shortening the amino acid sequence and thus the RGN coding sequence of the fusion protein, the peptide linker connecting RGN and the deaminase can also be shortened. Finally, genetic elements such as promoters, enhancers, and / or terminators can also be engineered through deletion analysis to determine the minimum size required for each to be functional.

[0206] Some embodiments of the present disclosure provide methods for editing nucleic acids using a deaminase or RGN complex described herein to achieve a nucleobase change, e.g., an A:T base pair to a G:C base pair. In some embodiments, the method is a method for editing a nucleobase of a nucleic acid (e.g., a base pair of a double-stranded DNA sequence). In some embodiments, a deaminase or RGN complex described herein is used to introduce a point mutation into a nucleic acid by deaminating and excising a targeted "A" nucleobase. In some embodiments, deamination and excision of the targeted nucleobase corrects a genetic defect, e.g., a point mutation in the CFTR gene. In some embodiments, the genetic defect is associated with a disease, disorder, or condition, e.g., cystic fibrosis. For example, in some embodiments, provided herein are methods for correcting a gene associated with a genetic defect, e.g., correcting a point mutation in the CFTR gene (e.g., in the treatment of a proliferative disease), using a base-editing RGN complex comprising a fusion protein with a deaminase having an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 399 and 405-407. In certain embodiments, the target sequence in the CFTR gene is SEQ ID NO: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, or 563.

[0207] In some embodiments, the objective of the methods provided herein is to restore the function of a dysfunctional gene through genome editing. The base editor proteins provided herein can be validated for in vitro human therapy based on gene editing, for example, by correcting disease-associated mutations in human cell culture. Those skilled in the art will understand that the fusion proteins and / or RGN complexes provided herein comprising a nucleic acid-binding protein (e.g., nCas9) and a nucleobase-modifying domain (e.g., a deaminase having the amino acid sequence set forth in SEQ ID NO:407, 399, or 405) can be used to modify any single point from T to G or change the pairing from T:A to G:C.

[0208] In some embodiments, provided herein are methods for treating a subject diagnosed with a disease associated with or caused by a point mutation (e.g., a mutation in the CFTR gene) that can be corrected by a fusion protein or RGN complex described herein. For example, in some embodiments, methods are provided that include administering to a subject with such a disease, e.g., cystic fibrosis, an effective amount of a fusion protein or RGN complex disclosed herein that corrects the point mutation or introduces an inactivating mutation in the disease-associated gene. In some embodiments, methods are provided that include administering to a subject with such a disease, e.g., cancer associated with the point mutation, an effective amount of a fusion protein, RGN complex, or pharmaceutical composition disclosed herein that corrects the point mutation or introduces an inactivating mutation in the disease-associated gene. In certain embodiments, methods for treating cystic fibrosis are provided in conjunction with methods for alleviating at least one symptom of cystic fibrosis by administering an effective amount of a pharmaceutical composition disclosed herein. An effective amount of a pharmaceutical composition for treating or alleviating a symptom of cystic fibrosis can reduce (i.e., treat) the symptoms of cystic fibrosis by about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more, or about 10-20%, 15-25%, 20-40%, 30-50%, 40-60%, 50-70%, 60-80%, 70-90%, 80-95%, or 90-95% compared to a control patient. In certain embodiments, the control patient can be the same patient prior to administration of an effective amount of a pharmaceutical composition disclosed herein. Symptoms of cystic fibrosis include, but are not limited to, sneezing, persistent cough that produces mucus or phlegm, shortness of breath especially during exercise, recurrent lung infections, stuffy nose, blocked sinuses, foul-smelling fatty stools, constipation, nausea, abdominal bloating, and loss of appetite, among others. Methods of identifying and measuring symptoms of cystic fibrosis are known in the art.

[0209] In some embodiments of the described methods for modifying a target DNA molecule, the contacting step occurs in vitro. In certain embodiments, the contacting step occurs in vivo. In some embodiments, the contacting step occurs in a subject (e.g., a human subject or a non-human animal subject). In some embodiments, the contacting step occurs in a cell, e.g., a human or non-human animal cell.

[0210] XII. Cells Containing Polynucleotide Gene Modifications Provided herein are cells and organisms comprising a target nucleic acid molecule of interest that has been modified using a process mediated by the fusion proteins and, optionally, a gRNA described herein. In some embodiments, the fusion protein comprises a deaminase polypeptide comprising the amino acid sequence of any of SEQ ID NOS: 1-10 and 399-441, or an active variant or fragment thereof. In some embodiments, the fusion protein comprises an adenine deaminase comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any of SEQ ID NOS: 1-10 and 399-441. In some embodiments, the fusion protein comprises a deaminase and a DNA-binding polypeptide (e.g., an RNA-guided DNA-binding polypeptide). In further embodiments, the fusion protein comprises a deaminase and RGN or a variant thereof, such as APG07433.1 (SEQ ID NO: 41) or its nickase variant nAPG07433.1 (SEQ ID NO: 42). In some embodiments, the fusion protein comprises a deaminase and Cas9 or a variant thereof, such as dCas9 or nickase-Cas9. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a Type II CRISPR-Cas polypeptide. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a Type V CRISPR-Cas polypeptide. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a Type VI CRISPR-Cas polypeptide.

[0211] The modified cells may be eukaryotic (e.g., mammalian, plant, insect, or avian) or prokaryotic. Organelles and embryos containing at least one nucleotide sequence modified by a process utilizing the fusion proteins described herein are also provided. Genetically modified cells, organisms, organelles, and embryos may be heterozygous or homozygous for the modified nucleotide sequence. The mutation introduced by the deaminase domain of the fusion protein may result in altered (increased or decreased) expression, inactivation, or expression of an altered protein product or integrated sequence. If the mutation results in either gene inactivation or expression of a non-functional protein product, the genetically modified cell, organism, organelle, or embryo is referred to as a "knockout." The knockout phenotype may be the result of a deletion mutation (i.e., deletion of at least one nucleotide), an insertion mutation (i.e., insertion of at least one nucleotide), or a nonsense mutation (i.e., substitution of at least one nucleotide such that a stop codon is introduced).

[0212] In some embodiments, the mutations introduced by the deaminase domain of the fusion protein result in the production of a variant protein product. The expressed variant protein product may have at least one amino acid substitution and / or at least one amino acid addition or deletion. The variant protein product may exhibit modified properties or activities (including, but not limited to, altered enzymatic activity or substrate specificity) when compared to the wild-type protein.

[0213] In some embodiments, mutations introduced by the deaminase domain of the fusion protein result in altered expression patterns of the protein. As a non-limiting example, mutations in regulatory regions controlling expression of a protein product can result in overexpression or downregulation of the protein product, or altered tissue or temporal expression patterns. The modified cells may be grown into organisms such as plants according to conventional methods. See, e.g., McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants may then be grown and pollinated with the same modified strain or a different strain, and the resulting hybrids will possess the genetic modification. The present invention provides genetically modified seeds. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the present invention, provided that these portions contain the genetic modification. Additionally, processed plant products or by-products (e.g., soybean meal) that retain the genetic modification are also provided.

[0214] Examples of plants of interest for which the methods provided herein may be used to modify any plant species, including, but not limited to, monocotyledonous and dicotyledonous plants, include, but are not limited to, corn, sorghum, wheat, sunflower, tomato, Brassica, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley, and rapeseed, Brassica, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus fruits, cacao, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, oat, vegetables, ornamental plants, and conifers.

[0215] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and plants belonging to the genus Cucumis, such as cucumbers, cantaloupes, and muskmelons. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. Preferably, the plants of the present invention are crop plants (e.g., corn, sorghum, wheat, sunflowers, tomatoes, crucifers, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).

[0216] The methods provided herein can also be used to genetically modify any prokaryotic species, including, but not limited to, archaea and bacteria (e.g., Bacillus, Klebsiella, Streptomyces, Rhizobium, Escherichia, Pseudomonas, Salmonella, Shigella, Vibrio, Yersinia, Mycoplasma, Agrobacterium, Lactobacillus).

[0217] The methods provided herein can be used to genetically modify any eukaryotic species or cells derived therefrom, including, but not limited to, animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoebae, algae, and yeast. In some embodiments, cells modified by the methods of the present disclosure include cells from the hematopoietic system, such as immune cells (i.e., cells of the innate or adaptive immune system) (including, but not limited to, B cells, T cells, natural killer (NK) cells), pluripotent stem cells, induced pluripotent stem cells, chimeric antigen receptor T (CAR-T) cells, monocytes, macrophages, and dendritic cells.

[0218] The modified cells may be introduced into an organism. These cells may be derived from the same organism (e.g., a human) for autologous cell transplantation, in which the cells are modified by ex vivo techniques. In some embodiments, the cells are derived from another organism within the same species (e.g., another human) for allogeneic cell transplantation.

[0219] XIII. Kit Some aspects of the present disclosure provide kits comprising the deaminase of the present invention. In certain embodiments, the present disclosure provides kits comprising a fusion protein comprising the deaminase of the present invention and a DNA-binding polypeptide (e.g., an RNA-guided DNA-binding polypeptide such as an RGN polypeptide, e.g., a nuclease-inactive Cas9 domain), and optionally a linker located between the DNA-binding polypeptide domain and the deaminase. In addition, in some embodiments, the kit includes reagents, buffers, and / or instructions suitable for using the fusion protein, e.g., for in vitro or in vivo DNA or RNA editing. In some embodiments, the kit includes instructions for designing and using a gRNA suitable for targeted editing of nucleic acid sequences.

[0220] In some embodiments, the pharmaceutical composition may be provided as a pharmaceutical kit comprising (a) a container with a composition of the present disclosure in lyophilized form and (b) a second container with a pharmaceutically acceptable diluent for injection (e.g., sterile water). The pharmaceutically acceptable diluent may be used to reconstitute or dilute the lyophilized compound of the present disclosure. Optionally, associated with such a container may be a notice reflecting approval by a governmental agency of the manufacture, use, or sale for human administration, in a manner prescribed by the governmental agency regulating the manufacture, use, or sale of pharmaceutical or biological products.

[0221] The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. For example, "a polypeptide" means one or more polypeptides.

[0222] All publications and patent applications mentioned in this specification are indicative of the level of ordinary skill in the art to which this disclosure pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.

[0223] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims.

[0224] Non-limiting embodiments include the following: 1. An isolated polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441, wherein the isolated polypeptide has deaminase activity.

[0225] 2. The isolated polypeptide of embodiment 1, comprising an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0226] 3. The isolated polypeptide of embodiment 1, comprising an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0227] 4. A nucleic acid molecule comprising a polynucleotide encoding a deaminase polypeptide, wherein the deaminase is: a) has at least 80% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485; or b) encoding an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1 to 10, 400 to 404, 406, and 408 to 441; A nucleic acid molecule encoded by a nucleotide sequence.

[0228] 5. The nucleic acid molecule of embodiment 4, wherein the deaminase is encoded by a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485.

[0229] 6. The nucleic acid molecule of embodiment 4, wherein the deaminase is encoded by a nucleotide sequence having at least 95% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485.

[0230] 7. The nucleic acid molecule of embodiment 4, wherein the deaminase is encoded by a nucleotide sequence having 100% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485.

[0231] 8. The nucleic acid molecule of embodiment 4, wherein the deaminase polypeptide has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0232] 9. The nucleic acid molecule of embodiment 4, wherein the deaminase polypeptide has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0233] 10. The nucleic acid molecule of any one of embodiments 4 to 9, further comprising a heterologous promoter operably linked to said polynucleotide.

[0234] 11. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a polypeptide according to any one of embodiments 1 to 3 or a nucleic acid molecule according to any one of embodiments 4 to 10.

[0235] 12. The pharmaceutical composition according to embodiment 11, wherein the pharmaceutically acceptable carrier is heterologous to the polypeptide or nucleic acid molecule.

[0236] 13. The pharmaceutical composition of embodiment 11 or 12, wherein the pharmaceutically acceptable carrier is non-naturally occurring.

[0237] 14. A fusion protein comprising a DNA-binding polypeptide and a deaminase having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0238] 15. The fusion protein of embodiment 14, wherein the deaminase has at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0239] 16. The fusion protein of embodiment 14, wherein the deaminase has 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0240] 17. The fusion protein of any one of embodiments 14 to 16, wherein the deaminase is adenine deaminase.

[0241] 18. The fusion protein of any one of embodiments 14 to 17, wherein the DNA-binding polypeptide is a meganuclease, a zinc finger fusion protein, or a TALEN.

[0242] 19. The fusion protein of any one of embodiments 14 to 17, wherein the DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide.

[0243] 20. The fusion protein of embodiment 19, wherein the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0244] 21. The fusion protein of embodiment 20, wherein RGN is a type II CRISPR-Cas polypeptide.

[0245] 22. The fusion protein of embodiment 20, wherein RGN is a type V CRISPR-Cas polypeptide.

[0246] 23. The fusion protein of any one of embodiments 20 to 22, wherein RGN is RGN nickase.

[0247] 24. The fusion protein of embodiment 20, wherein RGN has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 41, 60, 366, and 368.

[0248] 25. The fusion protein of embodiment 20, wherein RGN has the amino acid sequence of any one of SEQ ID NOs: 41, 60, 366, and 368.

[0249] 26. The fusion protein of embodiment 23, wherein the RGN nickase is any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0250] 27. A fusion protein according to any of embodiments 14 to 26, further comprising at least one nuclear localization signal (NLS).

[0251] 28. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein comprising a DNA-binding polypeptide and a deaminase, wherein the deaminase is a) has at least 80% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485; or b) encoding an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1 to 10, 400 to 404, 406, and 408 to 441; A nucleic acid molecule encoded by a nucleotide sequence.

[0252] 29. The nucleic acid molecule of embodiment 28, wherein the nucleotide sequence has at least 90% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485.

[0253] 30. The nucleic acid molecule of embodiment 28, wherein the nucleotide sequence has at least 95% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485.

[0254] 31. The nucleic acid molecule of embodiment 28, wherein the nucleotide sequence has 100% sequence identity to any one of SEQ ID NOs: 451, 449, 443, 11-20, 444-448, 450, and 452-485.

[0255] 32. The nucleic acid molecule of embodiment 28, wherein the nucleotide sequence encodes an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0256] 33. The nucleic acid molecule of embodiment 28, wherein the nucleotide sequence encodes an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0257] 34. The nucleic acid molecule of any one of embodiments 28 to 33, wherein the deaminase is adenine deaminase.

[0258] 35. The nucleic acid molecule of any one of embodiments 28 to 34, wherein the DNA-binding polypeptide is a meganuclease, a zinc finger fusion protein, or a TALEN.

[0259] 36. The nucleic acid molecule of any one of embodiments 28 to 34, wherein the DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide.

[0260] 37. The nucleic acid molecule of embodiment 36, wherein the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0261] 38. The nucleic acid molecule of embodiment 37, wherein RGN is a type II CRISPR-Cas polypeptide.

[0262] 39. The nucleic acid molecule of embodiment 37, wherein RGN is a type V CRISPR-Cas polypeptide.

[0263] 40. The nucleic acid molecule of any one of embodiments 37 to 39, wherein RGN is RGN nickase.

[0264] 41. The nucleic acid molecule of embodiment 37, wherein RGN has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 41, 60, 366, and 368.

[0265] 42. The nucleic acid molecule of embodiment 37, wherein RGN is SEQ ID NO: 41, 60, 366, or 368.

[0266] 43. The nucleic acid molecule of embodiment 40, wherein the RGN nickase is any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0267] 44. The nucleic acid molecule of any of embodiments 28 to 43, wherein the polynucleotide encoding the fusion protein is operably linked at its 5' end to a heterologous promoter.

[0268] 45. The nucleic acid molecule of any of embodiments 28 to 44, wherein the polynucleotide encoding the fusion protein is operably linked at its 3' end to a heterologous terminator.

[0269] 46. The nucleic acid molecule of any of embodiments 28 to 45, wherein the fusion protein comprises one or more nuclear localization signals.

[0270] 47. The nucleic acid molecule of any of embodiments 28 to 46, wherein the fusion protein is codon-optimized for expression in a eukaryotic cell.

[0271] 48. The nucleic acid molecule of any of embodiments 28 to 46, wherein the fusion protein is codon-optimized for expression in a prokaryotic cell.

[0272] 49. A vector comprising a nucleic acid molecule according to any one of embodiments 28 to 48.

[0273] 50. A vector comprising the nucleic acid molecule of any one of embodiments 28 to 48, further comprising at least one nucleotide sequence encoding a guide RNA (gRNA) capable of hybridizing to a target sequence.

[0274] 51. The vector of embodiment 50, wherein the gRNA is a single guide RNA.

[0275] 52. The vector of embodiment 50, wherein the gRNA is a dual guide RNA.

[0276] 53. A cell comprising a fusion protein according to any one of embodiments 14 to 27.

[0277] 54. A cell comprising a fusion protein according to any one of embodiments 14 to 27, further comprising a guide RNA.

[0278] 55. A cell comprising a nucleic acid molecule according to any one of embodiments 28 to 48.

[0279] 56. A cell comprising a vector according to any one of embodiments 49 to 52.

[0280] 57. A cell according to any one of embodiments 53 to 56, wherein the cell is a prokaryotic cell.

[0281] 58. The cell according to any one of embodiments 53 to 56, wherein the cell is a eukaryotic cell.

[0282] 59. The cell of embodiment 58, wherein the eukaryotic cell is a mammalian cell.

[0283] 60. The cell of embodiment 59, wherein the mammalian cell is a human cell.

[0284] 61. The cell of embodiment 60, wherein the human cell is an immune cell.

[0285] 62. The cell of embodiment 61, wherein the immune cell is a stem cell.

[0286] 63. The cell of embodiment 62, wherein the stem cell is an induced pluripotent stem cell.

[0287] 64. The cell of embodiment 58, wherein the eukaryotic cell is an insect cell or an avian cell.

[0288] 65. The cell of embodiment 58, wherein the eukaryotic cell is a fungal cell.

[0289] 66. The cell of embodiment 58, wherein the eukaryotic cell is a plant cell.

[0290] 67. A plant comprising a cell according to embodiment 66.

[0291] 68. A seed comprising cells according to embodiment 66.

[0292] 69. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any one of embodiments 14 to 27, a nucleic acid molecule according to any one of embodiments 28 to 48, a vector according to any one of embodiments 49 to 52, or a cell according to any one of embodiments 59 to 63.

[0293] 70. A method for producing a fusion protein, comprising culturing a cell according to any one of embodiments 53 to 66 under conditions in which the fusion protein is expressed.

[0294] 71. A method for producing a fusion protein, comprising introducing a nucleic acid molecule according to any one of embodiments 28 to 48 or a vector according to any one of embodiments 49 to 52 into a cell and culturing the cell under conditions in which the fusion protein is expressed.

[0295] 72. The method of embodiment 70 or 71, further comprising purifying the fusion protein.

[0296] 73. A method for producing an RGN fusion ribonucleoprotein complex, comprising introducing into a cell a nucleic acid molecule described in any one of embodiments 37 to 43 and a nucleic acid molecule comprising an expression cassette encoding a guide RNA or a vector described in any one of embodiments 50 to 52, and culturing the cell under conditions in which the fusion protein and gRNA are expressed and an RGN fusion ribonucleoprotein complex is formed.

[0297] 74. The method of embodiment 73, further comprising purifying the RGN fusion ribonucleoprotein complex.

[0298] 75. A system for modifying a target DNA molecule comprising a target DNA sequence, comprising: a) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and a deaminase having an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441, or a nucleotide sequence encoding the fusion protein; b) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding one or more guide RNAs (gRNAs), The system, wherein one or more guide RNAs are capable of forming a complex with the fusion protein to direct and bind the fusion protein to the target DNA sequence and modify the target DNA molecule.

[0299] 76. The system of embodiment 75, wherein the deaminase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0300] 77. The system of embodiment 75, wherein the deaminase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0301] 78. The system of any one of embodiments 75 to 77, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to the nucleotide sequence.

[0302] 79. A system according to any one of embodiments 75 to 78, wherein the target DNA sequence is a eukaryotic target DNA sequence.

[0303] 80. A system described in any one of embodiments 75 to 79, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by RGN.

[0304] 81. A system described in any one of embodiments 75 to 80, wherein the target DNA molecule is intracellular.

[0305] 82. The system of embodiment 81, wherein the cell is a eukaryotic cell.

[0306] 83. The system of embodiment 82, wherein the eukaryotic cell is a plant cell.

[0307] 84. The system of embodiment 82, wherein the eukaryotic cell is a mammalian cell.

[0308] 85. The system of embodiment 84, wherein the mammalian cells are human cells.

[0309] 86. The system of embodiment 85, wherein the human cells are immune cells.

[0310] 87. The system of embodiment 86, wherein the immune cells are stem cells.

[0311] 88. The system of embodiment 87, wherein the stem cells are induced pluripotent stem cells.

[0312] 89. The system of embodiment 82, wherein the eukaryotic cell is an insect cell.

[0313] 90. The system of embodiment 81, wherein the cells are prokaryotic cells.

[0314] 91. A system described in any one of embodiments 75 to 90, wherein the RGN of the fusion protein is a type II CRISPR-Cas polypeptide.

[0315] 92. A system described in any one of embodiments 75 to 90, wherein the RGN of the fusion protein is a type V CRISPR-Cas polypeptide.

[0316] 93. The system of any one of embodiments 75 to 90, wherein the RGN of the fusion protein has an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 41, 60, 366, or 368.

[0317] 94. A system described in any one of embodiments 75 to 90, wherein the RGN of the fusion protein has the amino acid sequence of any one of SEQ ID NOs: 41, 60, 366, and 368.

[0318] 95. A system according to any one of embodiments 75 to 90, wherein the RGN of the fusion protein is RGN nickase.

[0319] 96. The system of embodiment 95, wherein the RGN nickase is any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0320] 97. A system described in any of embodiments 75 to 96, wherein the fusion protein comprises one or more nuclear localization signals.

[0321] 98. A system described in any of embodiments 75 to 97, wherein the fusion protein is codon-optimized for expression in eukaryotic cells.

[0322] 99. A system according to any of embodiments 75 to 98, wherein the nucleotide sequence encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein are located on one vector.

[0323] 100. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a system according to any one of embodiments 75 to 99.

[0324] 101. A method for modifying a target DNA molecule comprising a target DNA sequence, the method comprising delivering a system described in any one of embodiments 75 to 99 to the target DNA molecule or a cell comprising the target DNA molecule.

[0325] 102. The method of embodiment 101, wherein the modified target DNA molecule comprises an A>N mutation of at least one nucleotide in the target DNA molecule, wherein N is C, G, or T.

[0326] 103. The method of embodiment 102, wherein said modified target DNA molecule comprises an A>G mutation of at least one nucleotide in the target DNA molecule.

[0327] 104. A method for modifying a target DNA molecule containing a target sequence, comprising: a) In vitro i) one or more guide RNAs capable of hybridizing to a target DNA sequence; ii) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one deaminase having an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441; under conditions suitable for the formation of an RGN-deaminase ribonucleotide complex; b) contacting the target DNA molecule or a cell containing the target DNA molecule with an in vitro assembled RGN-deaminase ribonucleotide complex; The method, wherein one or more guide RNAs hybridize to a target DNA sequence to direct and bind the fusion protein to the target DNA sequence, resulting in modification of the target DNA molecule.

[0328] 105. The method of embodiment 104, wherein the deaminase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0329] 106. The method of embodiment 104, wherein the deaminase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0330] 107. The method according to any one of embodiments 104 to 106, wherein the modified target DNA molecule comprises an A>N mutation of at least one nucleotide in the target DNA molecule, wherein N is C, G, or T.

[0331] 108. The method of embodiment 107, wherein the modified target DNA molecule comprises an A>G mutation of at least one nucleotide in the target DNA molecule.

[0332] 109. The method of any one of embodiments 104 to 108, wherein the RGN of the fusion protein is a type II CRISPR-Cas polypeptide.

[0333] 110. The method of any of embodiments 104-108, wherein the RGN of the fusion protein is a type V CRISPR-Cas polypeptide.

[0334] 111. The method of any of embodiments 104-108, wherein the RGN of the fusion protein has an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 41, 60, 366, or 368.

[0335] 112. The method of any one of embodiments 104 to 108, wherein the RGN of the fusion protein has the amino acid sequence of any one of SEQ ID NOs: 41, 60, 366, and 368.

[0336] 113. The method of any of embodiments 104-108, wherein the RGN of the fusion protein is RGN nickase.

[0337] 114. The method of embodiment 113, wherein the RGN nickase is any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0338] 115. The method of any of embodiments 104-114, wherein the fusion protein comprises one or more nuclear localization signals.

[0339] 116. The method of any of embodiments 104-115, wherein the fusion protein is codon-optimized for expression in eukaryotic cells.

[0340] 117. The method according to any one of embodiments 104 to 116, wherein the target DNA sequence is a eukaryotic target DNA sequence.

[0341] 118. The method of any of embodiments 104 to 117, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM).

[0342] 119. The method of any of embodiments 104-118, wherein the target DNA molecule is intracellular.

[0343] 120. The method of embodiment 119, wherein the cell is a eukaryotic cell.

[0344] 121. The method of embodiment 120, wherein the eukaryotic cell is a plant cell.

[0345] 122. The method of embodiment 120, wherein the eukaryotic cell is a mammalian cell.

[0346] 123. The method of embodiment 122, wherein the mammalian cells are human cells.

[0347] 124. The method of embodiment 123, wherein the human cells are immune cells.

[0348] 125. The method of embodiment 124, wherein the immune cells are stem cells.

[0349] 126. The method of embodiment 125, wherein the stem cells are induced pluripotent stem cells.

[0350] 127. The method of embodiment 120, wherein the eukaryotic cell is an insect cell.

[0351] 128. The method of embodiment 119, wherein the cell is a prokaryotic cell.

[0352] 129. The method of any one of embodiments 119-128, further comprising selecting cells containing said modified DNA molecule.

[0353] 130. A cell comprising a target DNA sequence modified according to the method of embodiment 129.

[0354] 131. The cell according to embodiment 130, wherein the cell is a eukaryotic cell.

[0355] 132. The cell according to embodiment 131, wherein the eukaryotic cell is a plant cell.

[0356] 133. A plant comprising a cell according to embodiment 132.

[0357] 134. A seed comprising cells according to embodiment 132.

[0358] 135. The cell of embodiment 131, wherein the eukaryotic cell is a mammalian cell.

[0359] 136. The cell according to embodiment 135, wherein the mammalian cell is a human cell.

[0360] 137. The cell of embodiment 136, wherein the human cell is an immune cell.

[0361] 138. The cell of embodiment 137, wherein the immune cell is a stem cell.

[0362] 139. The cell of embodiment 138, wherein the stem cell is an induced pluripotent stem cell.

[0363] 140. The cell according to embodiment 131, wherein the eukaryotic cell is an insect cell.

[0364] 141. The cell according to embodiment 130, wherein the cell is a prokaryotic cell.

[0365] 142. A pharmaceutical composition comprising a cell according to any one of embodiments 135 to 139 and a pharmaceutically acceptable carrier.

[0366] 143. A method for producing a genetically modified cell that corrects a causative mutation of a genetic disease, comprising: a) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and a deaminase having an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441, or a polynucleotide encoding the fusion protein, operably linked to a promoter to enable expression of the fusion protein in a cell; b) one or more guide RNAs (gRNAs) capable of hybridizing to a target DNA sequence, or a polynucleotide encoding the gRNA, operably linked to a promoter to allow expression of the gRNA in the cell; including introducing This allows the fusion protein and gRNA to target the genomic location of the causative mutation and modify the genomic sequence to remove the causative mutation.

[0367] 144. The method of embodiment 143, wherein the deaminase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0368] 145. The method of embodiment 143, wherein the deaminase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0369] 146. The method of any one of embodiments 143 to 145, wherein said RGN of the fusion protein is RGN nickase.

[0370] 147. The method of embodiment 146, wherein the RGN nickase is any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0371] 148. The method of any one of embodiments 143 to 147, wherein the genomic modification comprises introducing an A>G mutation of at least one nucleotide in the target DNA sequence.

[0372] 149. The method according to any of embodiments 143 to 148, wherein the cells are animal cells.

[0373] 150. The method of embodiment 149, wherein the animal cell is a mammalian cell.

[0374] 151. The method of embodiment 150, wherein the cells are derived from a dog, cat, mouse, rat, rabbit, horse, sheep, goat, cow, pig, or human.

[0375] 152. The method of any one of embodiments 143 to 151, wherein correcting the causative mutation comprises correcting a nonsense mutation.

[0376] 153. The method of embodiment 149, wherein the genetic disease is a disease listed in Table 34.

[0377] 154. The method of embodiment 149, wherein the genetic disease is cystic fibrosis.

[0378] 155. The method of embodiment 154, wherein the gRNA further comprises a spacer sequence targeting any one of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, and 563, or a complement thereof.

[0379] 156. The method of embodiment 155, wherein the gRNA comprises any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0380] 157. CRISPR RNA (crRNA) or a nucleic acid molecule encoding same, wherein the CRISPR RNA comprises a spacer sequence that targets a target DNA sequence within the cystic fibrosis transmembrane conductance regulator (CFTR) gene, and the target sequence has a sequence set forth in any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, 562, and 563, or a complement thereof.

[0381] 158. A guide RNA comprising the crRNA described in embodiment 157.

[0382] 159. The guide RNA of embodiment 158, which is a dual guide RNA.

[0383] 160. The guide RNA of embodiment 158, which is a single guide RNA (sgRNA).

[0384] 161. The guide RNA of embodiment 160, wherein the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0385] 162. The guide RNA of embodiment 160, wherein the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0386] 163. The guide RNA of embodiment 160, wherein the sgRNA has a sequence set forth in any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0387] 164. A vector comprising one or more nucleic acid molecules encoding a guide RNA according to any one of embodiments 158 to 163.

[0388] 165. A system for binding a target DNA sequence of a DNA molecule, comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding one or more guide RNAs (gRNAs); b) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and adenine deaminase, or a polynucleotide comprising a nucleotide sequence encoding the fusion protein; one or more guide RNAs capable of hybridizing to a target DNA sequence; one or more guide RNAs are capable of forming a complex with an RGN polypeptide to direct and bind the RGN polypeptide to the target DNA sequence of the DNA molecule; 1. A system comprising: at least one guide RNA comprising a CRISPR RNA (crRNA) comprising a spacer sequence that targets a target DNA sequence within the cystic fibrosis transmembrane conductance regulator (CFTR) gene, wherein the target sequence has a sequence set forth in any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, 562, and 563, or a complement thereof.

[0389] 166. The system of embodiment 165, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to the nucleotide sequence.

[0390] 167. A system for binding a target DNA sequence of a DNA molecule, comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding one or more guide RNAs (gRNAs); b) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and an adenine deaminase; one or more guide RNAs capable of hybridizing to a target DNA sequence; one or more guide RNAs are capable of forming a complex with an RGN polypeptide to direct and bind the RGN polypeptide to the target DNA sequence of a DNA molecule; 1. A system comprising: at least one guide RNA comprising a CRISPR RNA (crRNA) comprising a spacer sequence that targets a target DNA sequence within the cystic fibrosis transmembrane conductance regulator (CFTR) gene, wherein the target sequence has a sequence set forth in any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, 562, and 563, or a complement thereof.

[0391] 168. The system of embodiment 167, wherein at least one of the nucleotide sequences encoding one or more guide RNAs is operably linked to a promoter heterologous to the nucleotide sequence.

[0392] 169. The system according to any one of embodiments 165 to 168, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1 to 10 and 399 to 441.

[0393] 170. The system of any one of embodiments 165-168, wherein the deaminase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-10 and 399-441.

[0394] 171. The system according to any one of embodiments 165 to 168, wherein the deaminase has an amino acid sequence having a sequence set forth in any one of SEQ ID NOs: 1 to 10 and 399 to 441.

[0395] 172. A system according to any one of embodiments 165 to 171, wherein the RGN polypeptide and the one or more guide RNAs are not found complexed to each other in nature.

[0396] 173. The system of any one of embodiments 165-172, a) the target DNA sequence has any one of SEQ ID NOs: 62-68, 80-85, 116-119, 128-131, 163, 164, 180, 181, 203-209, 219-225, 256-258, 274-276, 310-313, and 330-333, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 53; b) the target DNA sequence has any one of SEQ ID NOs: 68-71, 86-89, 120-122, 132-134, 152-156, 169-173, 213-215, 229-231, 251-255, 269-273, 305-309, and 325-329, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 55; c) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 72, 73, 90, 91, 161, 162, 178, 179, 265, 266, 283, and 284, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 52; d) the target DNA sequence has any one of SEQ ID NOs: 74, 75, 92, 93, 123, 124, 135, 136, 167, 184, 216-218, 232-234, 259-261, 277-279, 314-317, and 334-337, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 56; e) the target DNA sequence has a sequence set forth in any one of SEQ ID NOs: 76, 94, 210-212, 226-228, 322, 342, 562, and 563, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 42; f) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 77, 95, 125, 137, 157-160, 174-177, 323, and 343, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 54; g) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 78, 96, 126, 138, 168, 185, 267, 285, 318, 319, 338, and 339, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 57; h) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 79, 97, 127, 139, 262-264, 280-282, 324, and 344, or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 58; and i) the target DNA sequence has a sequence set forth in any one of SEQ ID NOs: 165, 166, 182, 183, 268, 286, 320, 321, 340, and 341 or a complement thereof, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 59.

[0397] 174. The system of any one of embodiments 165-172, a) the target DNA sequence has any one of SEQ ID NOs: 62-68, 80-85, 116-119, 128-131, 163, 164, 180, 181, 203-209, 219-225, 256-258, 274-276, 310-313, and 330-333, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 53; b) the target DNA sequence has any one of SEQ ID NOs: 68-71, 86-89, 120-122, 132-134, 152-156, 169-173, 213-215, 229-231, 251-255, 269-273, 305-309, and 325-329, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 55; c) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 72, 73, 90, 91, 161, 162, 178, 179, 265, 266, 283, and 284, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 52; d) the target DNA sequence has any one of SEQ ID NOs: 74, 75, 92, 93, 123, 124, 135, 136, 167, 184, 216-218, 232-234, 259-261, 277-279, 314-317, and 334-337, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 56; e) the target DNA sequence has a sequence set forth in any one of SEQ ID NOs: 76, 94, 210-212, 226-228, 322, 342, 562, and 563, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 42; f) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 77, 95, 125, 137, 157-160, 174-177, 323, and 343, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 54; g) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 78, 96, 126, 138, 168, 185, 267, 285, 318, 319, 338, and 339, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 57; h) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 79, 97, 127, 139, 262-264, 280-282, 324, and 344, or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 58; and i) the target DNA sequence has a sequence set forth in any one of SEQ ID NOs: 165, 166, 182, 183, 268, 286, 320, 321, 340, and 341 or a complement thereof, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 59.

[0398] 175. The system of any one of embodiments 165-172, a) the target DNA sequence has any one of SEQ ID NOs: 62 to 68, 80 to 85, 116 to 119, 128 to 131, 163, 164, 180, 181, 203 to 209, 219 to 225, 256 to 258, 274 to 276, 310 to 313, and 330 to 333, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 53; b) the target DNA sequence has any one of SEQ ID NOs: 68-71, 86-89, 120-122, 132-134, 152-156, 169-173, 213-215, 229-231, 251-255, 269-273, 305-309, and 325-329, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 55; c) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 72, 73, 90, 91, 161, 162, 178, 179, 265, 266, 283, and 284, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 52; d) the target DNA sequence has any one of SEQ ID NOs: 74, 75, 92, 93, 123, 124, 135, 136, 167, 184, 216-218, 232-234, 259-261, 277-279, 314-317, and 334-337, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 56; e) the target DNA sequence has a sequence set forth in any one of SEQ ID NOs: 76, 94, 210-212, 226-228, 322, 342, 562, and 563, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 42; f) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 77, 95, 125, 137, 157-160, 174-177, 323, and 343, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 54; g) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 78, 96, 126, 138, 168, 185, 267, 285, 318, 319, 338, and 339, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 57; h) the target DNA sequence has the sequence set forth in any one of SEQ ID NOs: 79, 97, 127, 139, 262-264, 280-282, 324, and 344, or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 58; and i) the target DNA sequence has a sequence set forth in any one of SEQ ID NOs: 165, 166, 182, 183, 268, 286, 320, 321, 340, and 341 or a complement thereof, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 59.

[0399] 176. A system according to any one of embodiments 165 to 175, wherein at least one guide RNA is a dual guide RNA.

[0400] 177. A system described in any one of embodiments 165 to 175, wherein at least one guide RNA is a single guide RNA (sgRNA).

[0401] 178. The system of embodiment 177, a) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 98-104, 140-143, 197, 198, 235-241, 292-294, and 350-353, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 53; b) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 104-107, 144-146, 186-190, 245-247, 287-291, and 345-349, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 55; c) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 108, 109, 195, 196, 301, and 302, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 52; d) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 110, 111, 147, 148, 201, 248-250, 295-297, and 354-357, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 56; e) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 112, 242-244, 362, and 564, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 42; f) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 113, 149, 191-194, and 363, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 54; g) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 114, 150, 202, 303, 358, and 359, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 57; h) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 115, 151, 298-300, and 364, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 58; and i) the sgRNA has at least 90% sequence identity to any one of SEQ ID NOs: 199, 200, 304, 360, and 361, and the RGN polypeptide has a sequence having at least 90% sequence identity to SEQ ID NO: 59.

[0402] 179. The system of embodiment 177, a) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 98-104, 140-143, 197, 198, 235-241, 292-294, and 350-353, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 53; b) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 104-107, 144-146, 186-190, 245-247, 287-291, and 345-349, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 55; c) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 108, 109, 195, 196, 301, and 302, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 52; d) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 110, 111, 147, 148, 201, 248-250, 295-297, and 354-357, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 56; e) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 112, 242-244, 362, and 564, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 42; f) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 113, 149, 191-194, and 363, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 54; g) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 114, 150, 202, 303, 358, and 359, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 57; h) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 115, 151, 298-300, and 364, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 58; and i) the sgRNA has at least 95% sequence identity to any one of SEQ ID NOs: 199, 200, 304, 360, and 361, and the RGN polypeptide has a sequence having at least 95% sequence identity to SEQ ID NO: 59.

[0403] 180. The system of embodiment 177, a) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 98 to 104, 140 to 143, 197, 198, 235 to 241, 292 to 294, and 350 to 353, and the RGN polypeptide has a sequence that has 100% sequence identity to SEQ ID NO: 53; b) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 104-107, 144-146, 186-190, 245-247, 287-291, and 345-349, and the RGN polypeptide has a sequence that has 100% sequence identity to SEQ ID NO: 55; c) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 108, 109, 195, 196, 301, and 302, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 52; d) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 110, 111, 147, 148, 201, 248-250, 295-297, and 354-357, and the RGN polypeptide has a sequence that has 100% sequence identity to SEQ ID NO: 56; e) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 112, 242 to 244, 362, and 564, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 42; f) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 113, 149, 191-194, and 363, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 54; g) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 114, 150, 202, 303, 358, and 359, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 57; h) the sgRNA has 100% sequence identity to any one of SEQ ID NOs: 115, 151, 298-300, and 364, and the RGN polypeptide has a sequence having 100% sequence identity to SEQ ID NO: 58; and i) The sgRNA has 100% sequence identity to any one of SEQ ID NOs: 199, 200, 304, 360, and 361, and the RGN polypeptide has a sequence that has 100% sequence identity to SEQ ID NO: 59.

[0404] 181. A cell comprising a crRNA or nucleic acid molecule according to embodiment 157, a guide RNA according to any one of embodiments 158 to 163, a vector according to embodiment 164, or a system according to any one of embodiments 165 to 180.

[0405] 182. A pharmaceutical composition comprising the crRNA or nucleic acid molecule of embodiment 157, the guide RNA of any one of embodiments 158 to 163, the vector of embodiment 164, the cell of embodiment 181, or the system of any one of embodiments 165 to 180, and a pharmaceutically acceptable carrier.

[0406] 183. a) a fusion protein comprising a DNA-binding polypeptide and adenine deaminase, or a nucleic acid molecule encoding the fusion protein; b) a second adenine deaminase having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441, or a nucleic acid molecule encoding the deaminase.

[0407] 184. The composition of embodiment 183, wherein the second adenine deaminase has at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0408] 185. The composition of embodiment 183, wherein the second adenine deaminase has 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0409] 186. The composition of any one of embodiments 183-185, wherein the first adenine deaminase has at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0410] 187. The composition of any one of embodiments 183-186, wherein the first adenine deaminase has at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0411] 188. The composition of any one of embodiments 183-186, wherein the first adenine deaminase has 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0412] 189. The composition of any one of embodiments 183-188, wherein the DNA-binding polypeptide is a meganuclease, a zinc finger fusion protein, or a TALEN.

[0413] 190. The composition of any one of embodiments 183-189, wherein the DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide.

[0414] 191. The composition of embodiment 190, wherein the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0415] 192. The composition of embodiment 191, wherein the RGN is RGN nickase.

[0416] 193. A vector comprising a nucleic acid molecule encoding a fusion protein and a nucleic acid molecule encoding a second deaminase, wherein the fusion protein comprises a DNA-binding polypeptide and a first adenine deaminase, and the second adenine deaminase has at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0417] 194. The vector of embodiment 193, wherein the second adenine deaminase has at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0418] 195. The vector of embodiment 193, wherein the second adenine deaminase has 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0419] 196. The vector of any one of embodiments 193-195, wherein the first adenine deaminase has at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0420] 197. The vector of any one of embodiments 193-195, wherein the first adenine deaminase has at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0421] 198. The vector of any one of embodiments 193-195, wherein the first adenine deaminase has 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0422] 199. The vector of any one of embodiments 193-198, wherein the DNA-binding polypeptide is a meganuclease, a zinc finger fusion protein, or a TALEN.

[0423] 200. The vector of any one of embodiments 193-198, wherein the DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide.

[0424] 201. The vector of embodiment 200, wherein the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0425] 202. The vector of embodiment 201, wherein RGN is RGN nickase.

[0426] 203. A cell comprising a vector according to any one of embodiments 193 to 202.

[0427] 204. a) a fusion protein comprising a DNA-binding polypeptide and a first adenine deaminase, or a nucleic acid molecule encoding the fusion protein; b) a second adenine deaminase having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441, or a nucleic acid molecule encoding the second adenine deaminase.

[0428] 205. The cell of embodiment 204, wherein the second adenine deaminase has at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0429] 206. The cell of embodiment 204, wherein the second adenine deaminase has 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0430] 207. The cell of any one of embodiments 204-206, wherein the first adenine deaminase has at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0431] 208. The cell of any one of embodiments 204-206, wherein the first adenine deaminase has at least 95% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0432] 209. The cell of any one of embodiments 204-206, wherein the first adenine deaminase has 100% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441.

[0433] 210. The cell of any one of embodiments 204-209, wherein the DNA-binding polypeptide is a meganuclease, a zinc finger fusion protein, or a TALEN.

[0434] 211. The cell of any one of embodiments 204-209, wherein the DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide.

[0435] 212. The cell of embodiment 211, wherein the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0436] 213. The cell of embodiment 212, wherein the RGN is RGN nickase.

[0437] 214. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a composition according to any one of embodiments 183 to 192, a vector according to any one of embodiments 193 to 202, or a cell according to any one of embodiments 203 to 213.

[0438] 215. A method for treating a disease, comprising administering to a subject in need thereof an effective amount of a pharmaceutical composition described in any one of embodiments 69, 100, 142, and 214.

[0439] 216. The method of embodiment 215, wherein the disease is associated with a causative mutation and the effective amount of the pharmaceutical composition corrects the causative mutation.

[0440] 217. Use of a fusion protein according to any one of embodiments 14 to 27, a nucleic acid molecule according to any one of embodiments 28 to 48, a vector according to any one of embodiments 49 to 52 and 193 to 202, a cell according to any one of embodiments 59 to 63, 135 to 139 and 203 to 213, a system according to any one of embodiments 75 to 99, or a composition according to any one of embodiments 183 to 192 for the treatment of a disease in a subject.

[0441] 218. The use according to embodiment 217, wherein the disease is associated with a causative mutation and the treatment comprises correcting the causative mutation.

[0442] 219. Use of a fusion protein according to any one of embodiments 14 to 27, a nucleic acid molecule according to any one of embodiments 28 to 48, a vector according to any one of embodiments 49 to 52 and 193 to 202, a cell according to any one of embodiments 59 to 63, 135 to 139 and 203 to 213, a system according to any one of embodiments 75 to 99, or a composition according to any one of embodiments 183 to 192 for the manufacture of a medicament useful for the treatment of a disease.

[0443] 220. The use according to embodiment 219, wherein the disease is associated with a causative mutation and an effective amount of the pharmaceutical agent corrects the causative mutation.

[0444] 221. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 41 or 60, but lacking amino acid residues 590-597 of SEQ ID NO: 41 or 60; The method, wherein the RGN polypeptide is capable of binding to a target DNA sequence in a sequence-specific manner guided by RNA when bound to a guide RNA (gRNA) that can hybridize to the target DNA sequence.

[0445] 222. The nucleic acid molecule of embodiment 221, wherein said polynucleotide encoding an RGN polypeptide is operably linked to a promoter heterologous to said polynucleotide.

[0446] 223. The nucleic acid molecule of embodiment 221 or 222, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 366 or 368.

[0447] 224. The nucleic acid molecule of embodiment 221 or 222, wherein the RGN polypeptide comprises the amino acid sequence of SEQ ID NO: 366 or 368.

[0448] 225. The nucleic acid molecule of any one of embodiments 221 to 223, wherein the RGN polypeptide is nuclease-inactivated or functions as a nickase.

[0449] 226. The nucleic acid molecule of embodiment 225, wherein the nickase has an amino acid sequence as set forth in SEQ ID NO: 397 or 398.

[0450] 227. The nucleic acid molecule of any one of embodiments 221 to 226, wherein the RGN polypeptide is operably fused to the base-editing polypeptide.

[0451] 228. A vector comprising a nucleic acid molecule according to any one of claims 221 to 227.

[0452] 229. An isolated polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 41 or 60, but lacking amino acid residues 590-597 of SEQ ID NO: 41 or 60, wherein the isolated polypeptide is an RNA-guided nuclease.

[0453] 230. The isolated polypeptide of embodiment 229, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 366 or 368.

[0454] 231. The isolated polypeptide of embodiment 230, wherein the RGN polypeptide comprises the amino acid sequence of SEQ ID NO: 366 or 368.

[0455] 232. The isolated polypeptide of embodiment 229 or 230, wherein the RGN polypeptide is nuclease-inactivated or functions as a nickase.

[0456] 233. The isolated polypeptide of embodiment 232, wherein the nickase has an amino acid sequence set forth in SEQ ID NO: 397 or 398.

[0457] 234. The isolated polypeptide of any one of embodiments 229-233, wherein the RGN polypeptide is operably fused to the base-editing polypeptide.

[0458] 235. A cell comprising a nucleic acid molecule according to any one of embodiments 221 to 227, a vector according to claim 228, or a polypeptide according to any one of claims 229 to 234.

[0459] 236. An isolated polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 407, wherein the isolated polypeptide has deaminase activity.

[0460] 237. The isolated polypeptide of embodiment 236, comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 407, and having deaminase activity.

[0461] 238. The isolated polypeptide of embodiment 236, comprising the amino acid sequence set forth in SEQ ID NO: 407.

[0462] 239. A nucleic acid molecule comprising a polynucleotide encoding a deaminase polypeptide, wherein the deaminase is: a) has at least 80% sequence identity to SEQ ID NO: 451, or b) encoding an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407; A nucleic acid molecule encoded by a nucleotide sequence.

[0463] 240. The nucleic acid molecule of embodiment 239, wherein the deaminase is encoded by a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 451.

[0464] 241. The nucleic acid molecule of embodiment 239, wherein the deaminase is encoded by a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 451.

[0465] 242. The nucleic acid molecule of embodiment 239, wherein the deaminase is encoded by a nucleotide sequence having at least 100% sequence identity to SEQ ID NO: 451.

[0466] 243. The nucleic acid molecule of any one of embodiments 239 to 242, further comprising a heterologous promoter operably linked to said polynucleotide.

[0467] 244. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a polypeptide according to any one of embodiments 236-238 or a nucleic acid molecule according to any one of embodiments 239-242.

[0468] 245. A fusion protein comprising a DNA-binding polypeptide and a deaminase having at least 90% sequence identity to SEQ ID NO: 407.

[0469] 246. The fusion protein of embodiment 245, comprising a DNA-binding polypeptide and a deaminase having at least 95% sequence identity to SEQ ID NO: 407.

[0470] 247. The fusion protein of embodiment 245, comprising a DNA-binding polypeptide and a deaminase having 100% sequence identity to SEQ ID NO: 407.

[0471] 248. The fusion protein of any one of embodiments 245-247, wherein the DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0472] 249. The fusion protein of embodiment 248, wherein the RGN polypeptide is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0473] 250. The fusion protein of embodiment 248 or 249, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN polypeptide having the amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0474] 251. The fusion protein of any one of embodiments 248-250, wherein the RGN polypeptide is a nickase.

[0475] 252. The fusion protein of embodiment 251, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0476] 253. The fusion protein of embodiment 251, wherein the nickase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0477] 254. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein comprising a DNA-binding polypeptide and a deaminase, wherein the deaminase is a) has at least 80% sequence identity to SEQ ID NO: 451, or b) encoding an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 407, A nucleic acid molecule encoded by a nucleotide sequence.

[0478] 255. The nucleic acid molecule of embodiment 254, wherein the deaminase is encoded by a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 451.

[0479] 256. The nucleic acid molecule of embodiment 254, wherein the deaminase is encoded by a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 451.

[0480] 257. The nucleic acid molecule of embodiment 254, wherein the deaminase is encoded by a nucleotide sequence having at least 100% sequence identity to SEQ ID NO: 451.

[0481] 258. The nucleic acid molecule according to any one of embodiments 254 to 257, wherein the DNA-binding polypeptide is an RGN polypeptide.

[0482] 259. The nucleic acid molecule of embodiment 258, wherein the RGN is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0483] 260. The nucleic acid molecule of embodiment 258 or 259, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN polypeptide having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0484] 261. The nucleic acid molecule of any one of embodiments 258-260, wherein the RGN polypeptide is a nickase.

[0485] 262. The nucleic acid molecule of embodiment 261, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0486] 263. The nucleic acid molecule of embodiment 262, wherein the nickase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0487] 264. A vector comprising a nucleic acid molecule according to any one of embodiments 254 to 263.

[0488] 265. The vector of embodiment 264, further comprising at least one nucleotide sequence encoding a guide RNA (gRNA) capable of hybridizing to a target sequence.

[0489] 266. A ribonucleoprotein (RNP) complex comprising a fusion protein according to any one of embodiments 245 to 253 and a guide RNA bound to the DNA-binding polypeptide of the fusion protein.

[0490] 267. A cell comprising a fusion protein according to any of embodiments 245 to 253, a nucleic acid molecule according to any of embodiments 254 to 263, a vector according to embodiment 264 or 265, or an RNP complex according to embodiment 266.

[0491] 268. A system for modifying a target DNA molecule comprising a target DNA sequence, comprising: a) a fusion protein comprising an RNA-guided nuclease (RGN) polypeptide and a deaminase having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 407, or a nucleotide sequence encoding the fusion protein; b) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding one or more guide RNAs (gRNAs), The system, wherein one or more guide RNAs are capable of forming a complex with the fusion protein to direct and bind the fusion protein to the target DNA sequence and modify the target DNA molecule.

[0492] 269. The system of embodiment 268, wherein the deaminase has an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 407.

[0493] 270. The system of embodiment 268, wherein the deaminase has an amino acid sequence having 100% sequence identity to SEQ ID NO: 407.

[0494] 271. The system of any one of embodiments 268 to 270, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to the nucleotide sequence.

[0495] 272. A system according to any one of embodiments 268 to 271, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by an RGN polypeptide.

[0496] 273. The system of any one of embodiments 268-272, wherein the target DNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562 and 563, or a complement thereof.

[0497] 274. The system of any one of embodiments 268-273, wherein the gRNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0498] 275. The system of any one of embodiments 268 to 274, wherein the RGN polypeptide of the fusion protein is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0499] 276. The system of any one of embodiments 272 to 275, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0500] 277. The system of embodiment 276, wherein the RGN polypeptide is a nickase.

[0501] 278. The system of embodiment 277, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0502] 279. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any of embodiments 245 to 253, a nucleic acid molecule according to any one of embodiments 254 to 263, a vector according to embodiment 264 or 265, an RNP complex according to embodiment 266, a cell according to embodiment 267, or a system according to any one of embodiments 268 to 269.

[0503] 280. A method for modifying a target DNA molecule containing a target sequence, comprising: a) i) one or more guide RNAs capable of hybridizing to a target DNA sequence; ii) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one deaminase having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 407; under conditions suitable for the formation of an RGN-deaminase ribonucleotide complex; b) contacting the target DNA molecule or a cell containing the target DNA molecule with the assembled RGN-deaminase ribonucleotide complex; The method, wherein one or more guide RNAs hybridize to a target DNA sequence to direct and bind the fusion protein to the target DNA sequence, resulting in modification of the target DNA molecule.

[0504] 281. The method of embodiment 280, wherein the target DNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, and 563, or a complement thereof.

[0505] 282. The method of embodiment 280 or 281, wherein the gRNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0506] 283. The method according to any one of embodiments 280 to 283, which is carried out in vitro, in vivo, or ex vivo.

[0507] 284. A method of treating a subject having or at risk of developing a disease, disorder, or condition, comprising: A method comprising administering to a subject a fusion protein according to any one of embodiments 245 to 253, a nucleic acid molecule according to any one of embodiments 254 to 263, a vector according to embodiment 264 or 265, an RNP complex according to embodiment 266, a cell according to embodiment 267, a system according to any one of embodiments 268 to 28, or a pharmaceutical composition according to embodiment 279.

[0508] 285. The method of embodiment 284, further comprising administering any one of gRNAs comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0509] 286. An isolated polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 405, wherein the isolated polypeptide has deaminase activity.

[0510] 287. The isolated polypeptide of embodiment 286, comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 405, wherein the isolated polypeptide has deaminase activity.

[0511] 288. The isolated polypeptide of embodiment 286, comprising the amino acid sequence set forth in SEQ ID NO: 407.

[0512] 289. A nucleic acid molecule comprising a polynucleotide encoding a deaminase polypeptide, wherein the deaminase is: a) has at least 80% sequence identity to SEQ ID NO: 449, or b) encoding an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 405; A nucleic acid molecule encoded by a nucleotide sequence.

[0513] 290. The nucleic acid molecule of embodiment 289, wherein the deaminase is encoded by a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 449.

[0514] 291. The nucleic acid molecule of embodiment 289, wherein the deaminase is encoded by a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 449.

[0515] 292. The nucleic acid molecule of embodiment 289, wherein the deaminase is encoded by a nucleotide sequence having at least 100% sequence identity to SEQ ID NO: 449.

[0516] 293. The nucleic acid molecule of any one of embodiments 289 to 292, further comprising a heterologous promoter operably linked to said polynucleotide.

[0517] 294. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a polypeptide according to any one of embodiments 286 to 288 or a nucleic acid molecule according to any one of embodiments 289 to 293.

[0518] 295. A fusion protein comprising a DNA-binding polypeptide and a deaminase having at least 90% sequence identity to SEQ ID NO: 405.

[0519] 296. The fusion protein of embodiment 295, comprising a DNA-binding polypeptide and a deaminase having at least 95% sequence identity to SEQ ID NO: 405.

[0520] 297. The fusion protein of embodiment 295, comprising a DNA-binding polypeptide and a deaminase having 100% sequence identity to SEQ ID NO: 405.

[0521] 298. The fusion protein of any one of embodiments 295-297, wherein the DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0522] 299. The fusion protein of embodiment 298, wherein the RGN polypeptide is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0523] 300. The fusion protein of embodiment 298 or 299, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN polypeptide having the amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0524] 301. The fusion protein of any one of embodiments 298-300, wherein the RGN polypeptide is a nickase.

[0525] 302. The fusion protein of embodiment 301, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0526] 303. The fusion protein of embodiment 301, wherein the nickase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0527] 304. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein comprising a DNA-binding polypeptide and a deaminase, wherein the deaminase is: a) has at least 80% sequence identity to SEQ ID NO: 449, or b) encoding an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 405, A nucleic acid molecule encoded by a nucleotide sequence.

[0528] 305. The nucleic acid molecule of embodiment 304, wherein the deaminase is encoded by a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 449.

[0529] 306. The nucleic acid molecule of embodiment 304, wherein the deaminase is encoded by a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 449.

[0530] 307. The nucleic acid molecule of embodiment 304, wherein the deaminase is encoded by a nucleotide sequence having at least 100% sequence identity to SEQ ID NO: 449.

[0531] 308. The nucleic acid molecule of any one of embodiments 304 to 307, wherein the DNA-binding polypeptide is an RGN polypeptide.

[0532] 309. The nucleic acid molecule of embodiment 308, wherein the RGN is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0533] 310. The nucleic acid molecule of embodiment 308 or 309, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN polypeptide having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0534] 311. The nucleic acid molecule of any one of embodiments 308 to 310, wherein the RGN polypeptide is a nickase.

[0535] 312. The nucleic acid molecule of embodiment 311, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0536] 313. The nucleic acid molecule of embodiment 312, wherein the nickase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0537] 314. A vector comprising a nucleic acid molecule according to any one of embodiments 304 to 313.

[0538] 315. The vector of embodiment 314, further comprising at least one nucleotide sequence encoding a guide RNA (gRNA) capable of hybridizing to a target sequence.

[0539] 316. A ribonucleoprotein (RNP) complex comprising a fusion protein according to any one of embodiments 295 to 303 and a guide RNA bound to the DNA-binding polypeptide of the fusion protein.

[0540] 317. A cell comprising a fusion protein according to any of embodiments 295 to 303, a nucleic acid molecule according to any of embodiments 304 to 313, a vector according to embodiment 314 or 315, or an RNP complex according to embodiment 316.

[0541] 318. A system for modifying a target DNA molecule comprising a target DNA sequence, comprising: a) a fusion protein comprising an RNA-guided nuclease (RGN) polypeptide and a deaminase having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 405, or a nucleotide sequence encoding the fusion protein; b) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding one or more guide RNAs (gRNAs), The system, wherein one or more guide RNAs are capable of forming a complex with the fusion protein to direct and bind the fusion protein to the target DNA sequence and modify the target DNA molecule.

[0542] 319. The system of embodiment 318, wherein the deaminase has an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 405.

[0543] 320. The system of embodiment 318, wherein the deaminase has an amino acid sequence having 100% sequence identity to SEQ ID NO: 405.

[0544] 321. The system of any one of embodiments 318 to 320, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to the nucleotide sequence.

[0545] 322. A system according to any one of embodiments 318 to 321, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by an RGN polypeptide.

[0546] 323. The system of any one of embodiments 318-322, wherein the target DNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, and 563, or a complement thereof.

[0547] 324. The system of any one of embodiments 318-323, wherein the gRNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0548] 325. The system of any one of embodiments 318 to 324, wherein the RGN polypeptide of the fusion protein is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0549] 326. The system of any one of embodiments 322 to 325, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0550] 327. The system of embodiment 326, wherein the RGN polypeptide is a nickase.

[0551] 328. The system of embodiment 327, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0552] 329. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any of embodiments 295 to 303, a nucleic acid molecule according to any one of embodiments 304 to 313, a vector according to embodiment 314 or 315, an RNP complex according to embodiment 316, a cell according to embodiment 317, or a system according to any one of embodiments 318 to 328.

[0553] 330. A method for modifying a target DNA molecule containing a target sequence, comprising: a) i) one or more guide RNAs capable of hybridizing to a target DNA sequence; ii) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one deaminase having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 405; under conditions suitable for the formation of an RGN-deaminase ribonucleotide complex; b) contacting the target DNA molecule or a cell containing the target DNA molecule with the assembled RGN-deaminase ribonucleotide complex; The method wherein one or more guide RNAs hybridize to a target DNA sequence to direct and bind the fusion protein to the target DNA sequence, resulting in modification of the target DNA molecule.

[0554] 331. The method of embodiment 330, wherein the target DNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, and 563, or a complement thereof.

[0555] 332. The method of embodiment 330 or 331, wherein the gRNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0556] 333. The method according to any one of embodiments 330-332, which is carried out in vitro, in vivo, or ex vivo.

[0557] 334. A method of treating a subject having or at risk of developing a disease, disorder, or condition, comprising: A method comprising administering to a subject a fusion protein according to any one of embodiments 295 to 303, a nucleic acid molecule according to any one of embodiments 304 to 313, a vector according to embodiment 314 or 315, an RNP complex according to embodiment 316, a cell according to embodiment 317, a system according to any one of embodiments 318 to 328, or a pharmaceutical composition according to embodiment 329.

[0558] 335. The method of embodiment 334, further comprising administering any one of gRNAs comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0559] 336. An isolated polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 399, wherein the isolated polypeptide has deaminase activity.

[0560] 337. The isolated polypeptide of embodiment 336, comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 399, wherein the isolated polypeptide has deaminase activity.

[0561] 338. The isolated polypeptide of embodiment 336, comprising the amino acid sequence set forth in SEQ ID NO: 399.

[0562] 339. A nucleic acid molecule comprising a polynucleotide encoding a deaminase polypeptide, wherein the deaminase is: a) has at least 80% sequence identity to SEQ ID NO: 443, or b) encoding an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 399; A nucleic acid molecule encoded by a nucleotide sequence.

[0563] 340. The nucleic acid molecule of embodiment 339, wherein the deaminase is encoded by a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 443.

[0564] 341. The nucleic acid molecule of embodiment 339, wherein the deaminase is encoded by a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 443.

[0565] 342. The nucleic acid molecule of embodiment 339, wherein the deaminase is encoded by a nucleotide sequence having at least 100% sequence identity to SEQ ID NO: 443.

[0566] 343. The nucleic acid molecule of any one of embodiments 339 to 342, further comprising a heterologous promoter operably linked to said polynucleotide.

[0567] 344. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a polypeptide according to any one of embodiments 336-338 or a nucleic acid molecule according to any one of embodiments 339-342.

[0568] 345. A fusion protein comprising a DNA-binding polypeptide and a deaminase having at least 90% sequence identity to SEQ ID NO: 399.

[0569] 346. The fusion protein of embodiment 345, comprising a DNA-binding polypeptide and a deaminase having at least 95% sequence identity to SEQ ID NO: 399.

[0570] 347. The fusion protein of embodiment 345, comprising a DNA-binding polypeptide and a deaminase having 100% sequence identity to SEQ ID NO: 399.

[0571] 348. The fusion protein of any one of embodiments 345-347, wherein the DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

[0572] 349. The fusion protein of embodiment 348, wherein the RGN polypeptide is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0573] 350. The fusion protein of embodiment 348 or 349, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN polypeptide having the amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0574] 351. The fusion protein of any one of embodiments 348-350, wherein the RGN polypeptide is a nickase.

[0575] 352. The fusion protein of embodiment 351, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0576] 353. The fusion protein of embodiment 351, wherein the nickase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0577] 354. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein comprising a DNA-binding polypeptide and a deaminase, wherein the deaminase is: a) has at least 80% sequence identity to SEQ ID NO: 443, or b) encoding an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 399; A nucleic acid molecule encoded by a nucleotide sequence.

[0578] 355. The nucleic acid molecule of embodiment 354, wherein the deaminase is encoded by a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 443.

[0579] 356. The nucleic acid molecule of embodiment 354, wherein the deaminase is encoded by a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 443.

[0580] 357. The nucleic acid molecule of embodiment 354, wherein the deaminase is encoded by a nucleotide sequence having at least 100% sequence identity to SEQ ID NO: 443.

[0581] 358. The nucleic acid molecule of any one of embodiments 354 to 357, wherein the DNA-binding polypeptide is an RGN polypeptide.

[0582] 359. The nucleic acid molecule of embodiment 358, wherein the RGN is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0583] 360. The nucleic acid molecule of embodiment 358 or 359, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN polypeptide having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0584] 361. The nucleic acid molecule of any one of embodiments 358 to 360, wherein the RGN polypeptide is a nickase.

[0585] 362. The nucleic acid molecule of embodiment 361, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0586] 363. The nucleic acid molecule of embodiment 362, wherein the nickase has an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0587] 364. A vector comprising a nucleic acid molecule according to any one of embodiments 354 to 363.

[0588] 365. The vector of embodiment 364, further comprising at least one nucleotide sequence encoding a guide RNA (gRNA) capable of hybridizing to a target sequence.

[0589] 366. A ribonucleoprotein (RNP) complex comprising a fusion protein according to any one of embodiments 345 to 353 and a guide RNA bound to the DNA-binding polypeptide of the fusion protein.

[0590] 367. A cell comprising a fusion protein according to any of embodiments 345 to 353, a nucleic acid molecule according to any of embodiments 354 to 363, a vector according to embodiment 364 or 365, or an RNP complex according to embodiment 366.

[0591] 368. A system for modifying a target DNA molecule comprising a target DNA sequence, comprising: a) a fusion protein comprising an RNA-guided nuclease (RGN) polypeptide and a deaminase having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 399, or a nucleotide sequence encoding the fusion protein; b) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding one or more guide RNAs (gRNAs), The system, wherein one or more guide RNAs are capable of forming a complex with the fusion protein to direct and bind the fusion protein to the target DNA sequence and modify the target DNA molecule.

[0592] 369. The system of embodiment 368, wherein the deaminase has an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 399.

[0593] 370. The system of embodiment 368, wherein the deaminase has an amino acid sequence having 100% sequence identity to SEQ ID NO: 399.

[0594] 371. The system of any one of embodiments 368 to 370, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to the nucleotide sequence.

[0595] 372. A system according to any one of embodiments 368 to 371, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by an RGN polypeptide.

[0596] 373. The system of any one of embodiments 368-372, wherein the target DNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, and 563, or a complement thereof.

[0597] 374. The system of any one of embodiments 368 to 373, wherein the gRNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0598] 375. The system of any one of embodiments 368 to 374, wherein the RGN polypeptide of the fusion protein is a type II CRISPR-Cas polypeptide or a type V CRISPR-Cas polypeptide.

[0599] 376. The system of any one of embodiments 372 to 375, wherein the RGN polypeptide is Cas9, CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, LbCas12a, AsCas12a, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9 domain, or an RGN having an amino acid sequence set forth in any one of SEQ ID NOs: 41, 60, 366, or 368.

[0600] 377. The system of embodiment 376, wherein the RGN polypeptide is a nickase.

[0601] 378. The system of embodiment 377, wherein the nickase has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0602] 379. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any one of embodiments 345 to 353, a nucleic acid molecule according to any one of embodiments 354 to 363, a vector according to any one of embodiments 364 to 365, an RNP complex according to embodiment 366, a cell according to embodiment 367, or a system according to any one of embodiments 368 to 378.

[0603] 380. A method for modifying a target DNA molecule containing a target sequence, comprising: a) i) one or more guide RNAs capable of hybridizing to a target DNA sequence; ii) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one deaminase having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 399; under conditions suitable for the formation of an RGN-deaminase ribonucleotide complex; b) contacting the target DNA molecule or a cell containing the target DNA molecule with the assembled RGN-deaminase ribonucleotide complex; The method, wherein one or more guide RNAs hybridize to a target DNA sequence to direct and bind the fusion protein to the target DNA sequence, resulting in modification of the target DNA molecule.

[0604] 381. The method of embodiment 380, wherein the target DNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, and 563, or a complement thereof.

[0605] 382. The method of any one of embodiments 380-381, wherein the gRNA sequence comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0606] 383. The method according to any one of embodiments 380-382, which is carried out in vitro, in vivo, or ex vivo.

[0607] 384. A method of treating a subject having or at risk of developing a disease, disorder, or condition, comprising: A method comprising administering to a subject a fusion protein according to any one of embodiments 345 to 353, a nucleic acid molecule according to any one of embodiments 354 to 363, a vector according to any one of embodiments 364 to 365, an RNP complex according to embodiment 366, a cell according to embodiment 367, a system according to any one of embodiments 368 to 378, or a pharmaceutical composition according to embodiment 379.

[0608] 385. The method of embodiment 384, further comprising administering any one of gRNAs comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0609] 386. A method for treating or alleviating at least one symptom of cystic fibrosis, comprising administering to a subject in need thereof an effective amount of: a) a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and a deaminase having an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 407, 405, 399, 1-10, 400-404, 406, and 408-441, or a polynucleotide encoding the fusion protein, operably linked to a promoter to enable expression of the fusion protein in a cell; b) one or more guide RNAs (gRNAs) capable of hybridizing to a target DNA sequence, or a polynucleotide encoding the gRNA, operably linked to a promoter to allow expression of the gRNA in the cell; administering This allows the fusion protein and gRNA to target the genomic location of the causative mutation and modify the genomic sequence to remove the causative mutation.

[0610] 387. The method of embodiment 386, wherein the gRNA comprises a spacer sequence that targets any one of SEQ ID NOs: 62-97, 116-139, 152-185, 203-234, 251-286, 305-344, 562, and 563, or a complement thereof.

[0611] 388. The method of embodiment 386 or 387, wherein the gRNA comprises any one of SEQ ID NOs: 98-115, 140-151, 186-202, 235-250, 287-304, 345-364, and 564.

[0612] 389. The method of any one of claims 386-388, wherein the RGN has an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 41, 60, 366, and 368.

[0613] 390. The method of any one of claims 386-389, wherein the RGN has an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 42, 52-59, 61, 397, and 398.

[0614] The following examples are offered by way of illustration and not by way of limitation.

[0615] experiment Example 1: Demonstration of base editing in mammalian cells The deaminases shown in Table 1 below were generated based on naturally occurring deaminases, then mutated and selected for adenine deaminase activity in prokaryotic cells.

[0616] [Table 1]

[0617] To determine whether the deaminases in Table 1 can perform adenine base editing in mammalian cells, each deaminase was operably fused to RGN nickase to generate a fusion protein. Residues predicted to inactivate the RuvC domain of RGN APG07433.1 (shown as SEQ ID NO: 41; described in WO 2019 / 236566 (incorporated herein by reference)) were identified, and this RGN was modified into a nickase variant (nAPG07433.1; SEQ ID NO: 42). Nickase variants of RGN are referred to herein as "nRGN." It should be understood that any nickase variant of RGN can be used to generate the fusion protein of the present invention.

[0618] The deaminase and nRGN nucleotide sequences, codon-optimized for mammalian expression, were synthesized as fusion proteins with an N-terminal nuclear localization tag and cloned into the pTwist CMV (Twist Biosciences) expression plasmid. Starting from the amino terminus, each fusion protein contains an SV40 NLS (SEQ ID NO: 43) operably linked to a 3X FLAG tag (SEQ ID NO: 44) at the C-terminus, a deaminase operably linked to a peptide linker (SEQ ID NO: 45) at the C-terminus, an nRGN (e.g., nAPG07433.1, SEQ ID NO: 42) operably linked to an SV40 NLS (SEQ ID NO: 43) at the C-terminus, and finally a nucleoplasmin NLS (SEQ ID NO: 45) at the C-terminus. All fusion proteins contain at least one NLS and a 3X FLAG tag, as described above.

[0619] An expression plasmid containing an expression cassette encoding an sgRNA driven by the human U6 promoter (SEQ ID NO: 50) was also generated. The human genomic target sequence and the sgRNA sequence for directing the fusion protein to the genomic target are shown in Table 2.

[0620] [Table 2]

[0621] 500 ng of plasmids containing expression cassettes containing the coding sequences for the fusion proteins of each deaminase listed in Table 1 and 500 ng of plasmids containing expression cassettes encoding the sgRNAs listed in Table 2 were co-transfected into 75-90% confluent HEK293FT cells in 24-well plates using Lipofectamine 2000 reagent (Life Technologies). The cells were then incubated at 37°C for 72 hours. After incubation, genomic DNA was extracted using a NucleoSpin 96 Tissue (Macherey-Nagel) according to the manufacturer's protocol. Genomic regions flanking the target genomic site were PCR amplified using the primers listed in Table 2, and the products were purified using a ZR-96 DNA Clean and Concentrator (Zymo Research) according to the manufacturer's protocol. The purified PCR products were subjected to next-generation sequencing on an Illumina MiSeq. Typically, 100,000 250-bp paired-end reads (2 × 100,000 reads) are generated per amplicon. Reads were analyzed using CRISPResso (Pinello, et al. 2016 Nature Biotech, 34:695-697) to calculate editing rates. Output alignments were analyzed for indel formation or specific adenine mutation introduction. Tables 3–7 show adenine base editing for each fusion protein containing nAPG07433.1 and the deaminases listed in Table 1 and the guide RNAs listed in Table 2. The deaminase components of each fusion protein are indicated. Editing rates for adenines within or adjacent to the target sequence are shown. For example, "A5" indicates the adenine at position 5 of the target sequence. The position of each nucleotide in the target sequence is determined by numbering the first nucleotide in the target sequence closest to the PAM as position 1, with position numbers increasing in the 3' direction away from the PAM sequence. The table also indicates which nucleotide the adenine was mutated to and at what rate. For example, Table 3 shows that for the APG09982-nAPG07433.1 fusion protein, the adenine at position 13 was mutated to guanine at a rate of 1.2%.

[0622] [Table 3]

[0623] All fusion proteins showed detectable A>G transversions at positions A12 and A13. APG09982 and APG06333 showed at least 1% editing at position A13.

[0624] [Table 4]

[0625] All fusion proteins showed A>G conversions at positions A11 and A14. APG09982 showed 4.5% A11 to G conversion rate and 1.7% A14 to G conversion rate.

[0626] [Table 5]

[0627] All fusion proteins showed greater than 1% base editing at multiple positions in the target SGN000186. APG09102 showed a 6.2% A>G conversion rate at position A16 and greater than 2% base editing at positions A9 and A18. Position A16 was the most edited position in all fusion proteins tested.

[0628] [Table 6]

[0629] With SGN00194, all fusion proteins showed 0.9% to 1.8% A>G editing at position A15, with no detectable editing at positions A21, A23, A26, and A27.

[0630] [Table 7]

[0631] A14 was the most edited position across all fusion proteins tested, including SGN000930. For A>G transitions, the editing rate ranged from 0.3% to 1.2%.

[0632] Example 2: Fluorescence assay for targeted adenine base editing A vector was constructed containing enhanced green fluorescent protein (EGFP) (GFP-STOP, SEQ ID NO: 47) containing the W58X mutation that causes a premature stop codon, allowing the W58 codon to be reverted from a stop codon (TGA) to a wild-type tryptophan (TGG) residue by changing the A to G at position 3 using adenine deaminase. Successful A to G conversion results in expression of EGFP, which can be quantified. A second vector was also generated capable of expressing a guide RNA that targets the deaminase-RGN fusion protein to the region surrounding the W58X mutation (SEQ ID NO: 48).

[0633] This GFP-STOP reporter vector was transfected into HEK293T cells together with a vector capable of expressing the deaminase-nRGN fusion protein and the corresponding guide RNA using either lipofection or electroporation. For lipofection, cells were transfected at 1x10 cells per well the day before transfection in growth medium (DMEM + 10% fetal bovine serum + 1% penicillin / streptomycin). 5 Cells were seeded into 24-well plates at 1000 cells / well. 500 ng each of the GFP-STOP reporter vector, deaminase-RGN expression vector, and guide RNA expression vector were transfected using Lipofectamine® 3000 Reagent (Thermo Fisher Scientific) according to the manufacturer's instructions. For electroporation, cells were electroporated using the Neon® Transfection System (Thermo Fisher Scientific) according to the manufacturer's instructions.

[0634] In addition to transient transfection of the fluorescent GFP-STOP reporter, stable cell lines carrying the chromosomally integrated GFP-STOP cassette were generated. Once stable lines were established, cells were transfected at 1x10 cells per well the day before transfection in growth medium (DMEM + 10% fetal bovine serum + 1% penicillin / streptomycin). 5 Cells were seeded at 1000 cells / well in a 24-well plate. 500 ng each of the deaminase-nRGN expression vector and guide RNA expression vector were transfected using Lipofectamine® 3000 Reagent (Thermo Fisher Scientific) according to the manufacturer's instructions. For electroporation, cells were electroporated using the Neon® Transfection System (Thermo Fisher Scientific) according to the manufacturer's instructions.

[0635] 24-48 hours after lipofection or electroporation, GFP expression was measured by examining cells microscopically for the presence of GFP+ cells. After visual inspection, the ratio of GFP+ cells to GFP- cells could be determined. Fluorescence was observed in mammalian cells expressing each deaminase-nRGN fusion protein, indicating that the fusion proteins successfully targeted the GFP-STOP mutation, edited the mutation, and restored GFP protein fluorescence.

[0636] After microscopic analysis, cells were lysed in RIPA buffer and the resulting lysates were analyzed in a fluorescent plate reader to determine the GFP fluorescence intensity (Table 8). One skilled in the art will appreciate that cells can be analyzed by flow cytometry or fluorescence-activated cell sorting to determine the exact percentage of GFP+ and GFP- cells.

[0637] [Table 8] ND = not detected; + = few GFP+ cells detected; ++ = some GFP+ cells detected; +++ = many GFP+ cells detected.

[0638] Example 3: Demonstration of A base editing in mammalian cells The deaminases shown in Table 9 below were generated based on naturally occurring deaminases, then mutated and selected for adenine deaminase activity in prokaryotic cells.

[0639] [Table 9]

[0640] To determine whether the deaminases in Table 9 can perform adenine base editing in mammalian cells, each deaminase was operably fused to RGN nickase to generate a fusion protein. Residues predicted to inactivate the RuvC domain of RGN APG07433.1 (shown as SEQ ID NO: 41; described in WO2019 / 236566 (incorporated herein by reference)) were identified, and this RGN was modified into a nickase variant (nAPG07433.1; SEQ ID NO: 42). The nickase variant of RGN is referred to herein as "nRGN." It should be understood that any nickase variant of RGN can be used to generate the fusion protein of the present invention.

[0641] The deaminase and nRGN nucleotide sequences, codon-optimized for mammalian expression, were synthesized as fusion proteins with an N-terminal nuclear localization tag and cloned into the pTwist CMV (Twist Biosciences) expression plasmid. Starting from the amino terminus, each fusion protein contains an SV40 NLS (SEQ ID NO: 43) operably linked to a 3X FLAG tag (SEQ ID NO: 44) at the C-terminus, a deaminase at the C-terminus, a peptide linker (SEQ ID NO: 442) at the C-terminus, an nRGN (e.g., nAPG07433.1, SEQ ID NO: 42) at the C-terminus, and finally a nucleoplasmin NLS (SEQ ID NO: 46) at the C-terminus. The nAPG07433.1 and peptide linker nucleotide sequences, codon-optimized for mammalian expression, are set forth as SEQ ID NOs: 486 and 487, respectively. Table 10 lists the fusion proteins that were generated and tested for activity. All fusion proteins cont...

Claims

1. A nucleic acid molecule comprising a polynucleotide encoding a deaminase polypeptide, the deaminase polypeptide has adenine deaminase activity, and a) has at least 90% sequence identity to SEQ ID NO: 451; and b) encoding an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 407; encoded by a nucleotide sequence, The nucleic acid molecule further comprising a heterologous promoter operably linked to said polynucleotide.

2. the nucleotide sequence encoding the deaminase polypeptide is a) has at least 95% sequence identity to SEQ ID NO: 451; and b) a nucleic acid molecule according to claim 1, encoding an amino acid sequence having at least 95% sequence identity to SEQ ID NO:

407.

3. A nucleic acid molecule comprising a polynucleotide encoding a deaminase polypeptide, the deaminase polypeptide has adenine deaminase activity, and a) having the sequence of SEQ ID NO: 451, or b) encoding the amino acid sequence of SEQ ID NO: 407, A nucleic acid molecule encoded by a nucleotide sequence.

4. A vector comprising the nucleic acid molecule of claim 1.

5. 5. The vector of claim 4, further comprising at least one nucleotide sequence encoding a guide RNA capable of hybridizing to the target nucleic acid.

6. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein, the fusion protein comprises a type II CRISPR-Cas protein nickase and a deaminase; the deaminase has adenine deaminase activity and comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO: 407; and The nickase a) a Cas9 nickase; or b) A nucleic acid molecule comprising an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52, 53, 55-59, 61, 397, or 398.

7. 7. The nucleic acid molecule of claim 6, wherein the deaminase comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:

407.

8. 7. The nucleic acid molecule of claim 6, wherein the deaminase comprises the amino acid sequence of SEQ ID NO:

407.

9. 7. The nucleic acid molecule of claim 6, wherein the nickase has the amino acid sequence of any one of SEQ ID NOs: 42, 52, 53, 55-59, 61, 397, or 398.

10. The nucleic acid molecule of claim 6 , wherein the polynucleotide encoding the fusion protein is operably linked at its 5′ end to a heterologous promoter.

11. The nucleic acid molecule of claim 6 , wherein the polynucleotide encoding the fusion protein is operably linked at its 3′ end to a heterologous terminator.

12. The nucleic acid molecule of claim 6 , wherein the fusion protein comprises one or more nuclear localization signals.

13. The nucleic acid molecule of claim 6 , wherein the polynucleotide encoding the fusion protein is codon-optimized for expression in a eukaryotic cell.

14. The nucleic acid molecule of claim 6 , wherein the polynucleotide encoding the fusion protein is mRNA.

15. The nucleic acid molecule of claim 6, wherein the fusion protein comprises the amino acid sequence of SEQ ID NO:

496.

16. 1. A system for modifying a target DNA molecule, comprising: The system comprises: a) a nucleic acid molecule according to claim 6; b) one or more guide RNAs (gRNAs) capable of hybridizing to the target DNA molecule or one or more nucleic acids encoding the one or more gRNAs; and wherein the one or more gRNAs are capable of forming a complex with the fusion protein such that the fusion protein targets and binds to the target DNA molecule to modify the target DNA molecule.

17. 17. The system of Claim 16, wherein at least one of the one or more nucleic acids encoding the one or more gRNAs is operably linked to a promoter.

18. 17. The system of claim 16, comprising a vector comprising the nucleic acid molecule of claim 6 and the one or more nucleic acids encoding the one or more gRNAs.

19. The system of claim 16 , wherein the target DNA molecule is intracellular.

20. The system of claim 19 , wherein the cell is a eukaryotic cell.

21. A vector comprising the nucleic acid molecule of claim 3.

22. 22. The vector of claim 21, further comprising at least one nucleotide sequence encoding a guide RNA capable of hybridizing to a target nucleic acid.

23. A fusion protein comprising a type II CRISPR-Cas protein nickase and a deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO: 407 and has adenine deaminase activity; The nickase a) a Cas9 nickase; or b) A fusion protein comprising an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52, 53, 55-59, 61, 397, or 398.

24. 24. The fusion protein of claim 23, wherein the nickase has the amino acid sequence of SEQ ID NO:

42.

25. 24. The fusion protein of claim 23, further comprising at least one nuclear localization signal (NLS).

26. 24. A cell comprising the fusion protein of claim 23, further comprising a guide RNA, wherein the cell is not a cell of a human embryo.

27. 1. A system for modifying a target DNA molecule comprising a target DNA sequence, comprising: a) a fusion protein according to claim 23; b) one or more guide RNAs (gRNAs) capable of hybridizing to the target DNA molecule or one or more nucleotide sequences encoding the one or more gRNAs; and wherein the one or more gRNAs are capable of forming a complex with the fusion protein such that the fusion protein is directed to, binds to, and modifies the target DNA molecule.

28. 28. The system of claim 27, wherein the nickase has the amino acid sequence of SEQ ID NO:

42.

29. A polypeptide having adenine deaminase activity and comprising an amino acid sequence having at least 90% sequence identity to the amino acid sequence of SEQ ID NO:

407.

30. 30. The polypeptide of claim 29, wherein the polypeptide comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO:

407.

31. 30. The polypeptide of claim 29, wherein the polypeptide has the amino acid sequence of SEQ ID NO:

407.

32. 30. The polypeptide of claim 29, wherein the polypeptide further comprises one or more nuclear localization signals.

33. 30. An adenine base editor comprising the polypeptide of claim 29 and a type II CRISPR-Cas nickase, The nickase a) a Cas9 nickase; or b) an adenine base editor comprising an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 42, 52, 53, 55-59, 61, 397, or 398.

34. 34. The adenine base editor of Claim 33, wherein the adenine base editor introduces an A>G mutation in a DNA molecule.

35. 24. The fusion protein of claim 23, wherein the deaminase comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence of SEQ ID NO: 407 and has adenine deaminase activity.

36. 24. The fusion protein of claim 23, wherein the deaminase comprises the amino acid sequence of SEQ ID NO:

407.

37. 25. The fusion protein of claim 24, wherein the fusion protein comprises the amino acid sequence of SEQ ID NO:

496.

38. The fusion protein of claim 23 , wherein the fusion protein further comprises a protein tag or a cell-penetrating domain.

39. 24. The fusion protein of claim 23, wherein the nickase has the amino acid sequence of SEQ ID NO: 42, 52, 53, 55-59, 61, 397, or 398.

40. 28. The system of claim 27, wherein the nickase has the amino acid sequence of SEQ ID NO: 42, 52, 53, 55-59, 61, 397, or 398.

Citation Information

Patent Citations

  • Polypeptides useful for gene editing and methods of use

    CA3125175A1