Nucleic acid base editing factor comprising nucleic acid-programmable DNA binding protein

JP2026009871A5Pending Publication Date: 2026-09-07PRESIDENT & FELLOWS OF HARVARD COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025135185
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-08-30
Filing Date
2025-08-14
Publication Date
2026-09-07

AI Technical Summary

Technical Problem

Current genome editing technologies, such as ZFNs, TALENs, and Cas9, suffer from stochastic processes like NHEJ and HDR, leading to modest gene editing efficiency and unwanted gene changes, making precise nucleotide changes at specific genomic locations challenging.

Method used

Development of nucleic acid-programmable DNA-binding proteins (napDNAbps) like CRISPR-Cas systems, particularly dCas9 fused with cytidine deaminase domains and uracil glycosylase inhibitors (UGIs), enabling precise and efficient deamination of target cytidine residues without significant indels.

Benefits of technology

The fusion proteins achieve high efficiency in creating desired mutations, such as C→T changes, with minimal off-target effects and reduced indel formation, paving the way for precise gene editing tools and therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000273_0000
    Figure 00000273_0000
  • Figure 00000273_0001
    Figure 00000273_0001
  • Figure 00000273_0002
    Figure 00000273_0002
Patent Text Reader

Abstract

Some aspects of this disclosure provide strategies, systems, reagents, methods, and kits that are useful for targeted editing of nucleic acids, including editing a single site in the genome of a cell or subject, e.g., in the human genome.SOLUTION: In some embodiments, fusion proteins between nucleic acid-programmable DNA-binding proteins (napDNAbp), such as Cpf1 or variants thereof, and nucleic acid-editing proteins or proteins domains, such as deaminase domains, are provided. In some embodiments, a method for targeted nucleic acid editing is provided. In some embodiments, reagents and kits for the creation of targeted nucleic acid editing proteins, e.g., fusion proteins of a nucleic acid editing protein or domain with a napDNAbp (e.g., CasX, CasY, Cpf1, C2c1, C2c2, C2c3, and Argonaute) are provided.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Background of the Invention Targeted editing of nucleic acid sequences, such as targeted cleavage or targeted introduction of specific modifications into genomic DNA, is a highly promising approach for studying gene function and also has the potential to provide new treatments for human genetic diseases. 1 An ideal nucleic acid editing technology would have three characteristics: (1) high efficiency in installing the desired modifications; (2) minimal off-target activity; and (3) the ability to be programmed to precisely edit any site in a given nucleic acid, e.g., any site in the human genome. 2 Artificial zinc finger nucleases (ZFNs) 3 , transcription activator-like effector nuclease (TALEN) 4 , and more recently the RNA-guided DNA endonuclease Cas9 5 Current genome engineering tools, including those described in the literature, achieve sequence-specific DNA cleavage of the genome. This programmable cleavage can result in mutation of the DNA at the cleavage site by non-homologous end joining (NHEJ) or replacement of the DNA surrounding the cleavage site by homology-directed repair (HDR). 6,7 .

[0002] One drawback of current technologies is that both NHEJ and HDR are stochastic processes, which typically result in modest gene editing efficiency and unwanted gene changes that can compete with the desired changes. 8 In principle, many genetic diseases can be treated by effecting specific nucleotide changes at specific locations in the genome (e.g., a C→T change in a specific codon of a disease-associated gene). 9 The development of programmable methods to achieve such precise gene editing would represent both a powerful new research tool and a promising new approach to gene-editing-based human therapeutics. Summary of the Invention

[0003] SUMMARY OF THE INVENTION Nucleic acid-programmable DNA-binding proteins (napDNAbp), such as the clustered regularly interspaced short palindromic repeat (CRISPR) system, are a recently discovered prokaryotic adaptive immune system. 10 , which has been modified to enable robust and general genome engineering in a variety of organisms and cell lines. 11 The CRISPR-Cas (CRISPR-associated) system is a protein-RNA complex that uses an RNA molecule (sgRNA) as a guide to localize the complex to a target DNA sequence through base pairing. 12 In natural systems, the Cas proteins then act as endonucleases to cleave the target DNA sequence. 13 For the system to function, the target DNA sequence must be complementary to the sgRNA and must also contain a "protospacer adjacent motif" (PAM) at the 3' end of the complementary region. 14 .

[0004] Among the known Cas proteins, S. pyogenes Cas9 is primarily used widely as a tool for genome engineering. 15 The Cas9 protein is a large multidomain protein containing two separate nuclease domains. Point mutations can be introduced into Cas9 to abolish nuclease activity, resulting in dead Cas9 (dCas9), which still retains its ability to bind DNA in a manner programmed by the sgRNA. 16 In principle, when fused to another protein or domain, dCas9 could target that protein or domain to virtually any DNA sequence simply by co-expression with the appropriate sgRNA.

[0005] The potential of the dCas9 complex for genome engineering purposes is immense. Its unique ability to target proteins to specific sites in the genome programmed by sgRNA could, in theory, be expanded into a variety of site-specific genome engineering tools beyond nucleases, including deaminases (e.g., cytidine deaminases), transcriptional activators, transcriptional repressors, histone-modifying proteins, integrases, and recombinases. 11 Some of these potential applications have recently been implemented by dCas9 fusions with transcriptional activators, enabling RNA-guided transcriptional activators. 17,18 , transcriptional repressor 16,19,20 , and chromatin-modifying enzymes 21 The simple co-expression of these fusions with various sgRNAs results in specific expression of the target gene. These groundbreaking studies paved the way for the design and construction of easily programmable sequence-specific effectors for precise manipulation of the genome.

[0006] Some aspects of the present disclosure are based on the recognition that certain configurations of a nucleic acid-programmable DNA-binding protein (napDNAbp), such as CasX, CasY, Cpf1, C2c1, C2c2, C2c3, or Argonaute protein, fused via a linker with a cytidine deaminase domain are useful for efficiently deaminating target cytidine residues. Another aspect of the disclosure relates to the recognition that a nucleic acid base editing fusion protein having a cytidine deaminase domain fused via a linker to the N-terminus of a napDNAbp was able to efficiently deaminate a target nucleic acid within a double-stranded DNA target molecule. See, e.g., Examples 3 and 4 below, which demonstrate that the fusion proteins, also referred to herein as base editors, create fewer indels and efficiently deaminate target nucleic acids than other base editors, such as base editors without a UGI domain. Another aspect of the present disclosure relates to the recognition that nucleic acid base editing fusion proteins having a cytidine deaminase domain fused to the N-terminus of a napDNAbp via a linker perform base editing with higher efficiency and greatly improved product purity when the fusion protein contains more than one UGI domain. See, e.g., Example 17, which demonstrates that fusion proteins (e.g., base editors) containing two UGI domains create fewer indels and efficiently deaminate target nucleic acids than other base editors, such as those containing one UGI domain.

[0007] In some embodiments, the fusion protein comprises: (i) a nucleic acid-programmable DNA-binding protein (napDNAbp); (ii) a cytidine deaminase domain; and (iii) a uracil glycosylase inhibitor (UGI) domain, wherein the napDNAbp is a CasX, CasY, Cpf1, C2c1, C2c2, C2c3, or Argonaute protein. In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a CasX protein. In some embodiments, the CasX protein comprises an amino acid sequence at least 90% identical to SEQ ID NO: 29 or 30. In some embodiments, the CasX protein comprises the amino acid sequence of SEQ ID NO: 29 or 30.

[0008] In some embodiments, the fusion protein comprises: (i) a nucleic acid-programmable DNA-binding protein (napDNAbp); (ii) a cytidine deaminase domain; (iii) a first uracil glycosylase inhibitor (UGI) domain; and (iv) a second uracil glycosylase inhibitor (UGI) domain, wherein the napDNAbp is Cas9, dCas9, or a Cas9 nickase protein. In some embodiments, the napDNAbp is a dCas9 protein. In some embodiments, the napDNAbp is a CasX, CasY, Cpf1, C2c1, C2c2, C2c3, or Argonaute protein. In some embodiments, the dCas9 protein is S. pyogenes dCas9 (SpCas9d). In some embodiments, the dCas9 protein is S. pyogenes dCas9 harboring a D10A mutation. In some embodiments, the dCas9 protein comprises an amino acid sequence at least 90% identical to SEQ ID NO: 6 or 7. In some embodiments, the dCas9 protein comprises the amino acid sequence of SEQ ID NO: 6 or 7. In some embodiments, the dCas9 protein is S. aureus dCas9 (SaCas9d). In some embodiments, the dCas9 protein is S. aureus dCas9 harboring a D10A mutation. In some embodiments, the dCas9 protein comprises an amino acid sequence at least 90% identical to SEQ ID NOs: 33-36. In some embodiments, the dCas9 protein comprises the amino acid sequence of SEQ ID NOs: 33-36.

[0009] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a CasY protein. In some embodiments, the CasY protein comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 31. In some embodiments, the CasY protein comprises the amino acid sequence of SEQ ID NO: 31.

[0010] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a Cpf1 or Cpf1 mutant protein. In some embodiments, the Cpf1 or Cpf1 mutant protein comprises an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 9-24. In some embodiments, the Cpf1 or Cpf1 mutant protein comprises the amino acid sequence of any one of SEQ ID NOs: 9-24.

[0011] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a C2c1 protein. In some embodiments, the C2c1 protein comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 26. In some embodiments, the C2c1 protein comprises the amino acid sequence of SEQ ID NO: 26.

[0012] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a C2c2 protein. In some embodiments, the C2c2 protein comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 27. In some embodiments, the C2c2 protein comprises the amino acid sequence of SEQ ID NO: 27.

[0013] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a C2c3 protein. In some embodiments, the C2c3 protein comprises an amino acid sequence at least 90% identical to SEQ ID NO: 28. In some embodiments, the C2c3 protein comprises the amino acid sequence of SEQ ID NO: 28.

[0014] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is an Argonaute protein. In some embodiments, the Argonaute protein comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 25. In some embodiments, the Argonaute protein comprises the amino acid sequence of SEQ ID NO: 25.

[0015] Some aspects of the present disclosure are based on the recognition that the fusion proteins provided herein can create one or more mutations (e.g., C→T mutations) without creating a large percentage of indels. In some embodiments, any of the fusion proteins (e.g., base-editing proteins) provided herein create fewer than 10% indels. In some embodiments, any of the fusion proteins (e.g., base-editing proteins) provided herein create fewer than 10%, 9%, 8%, 7%, 6%, 5.5%, 5%, 4.5%, 4%, 3.5%, 3%, 2.5%, 2%, 1.5%, 1%, 0.5%, or 0.1% indels.

[0016] In some embodiments, the fusion protein comprises a napDNAbp and an apolipoprotein B mRNA editing complex 1 (APOBEC1) deaminase domain, wherein the deaminase domain is fused to the N-terminus of the napDNAbp domain via a linker comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604). In some embodiments, the napDNAbp comprises any of the amino acid sequences of napDNAbp provided herein. In some embodiments, the deaminase is rat APOBEC1 (SEQ ID NO: 76). In some embodiments, the deaminase is human APOBEC1 (SEQ ID NO: 74). In some embodiments, the deaminase is pmCDA1 (SEQ ID NO: 81). In some embodiments, the deaminase is human APOBEC3G (SEQ ID NO: 60). In some embodiments, the deaminase is any one of the human APOBEC3G variants of (SEQ ID NOs: 82-84). In some embodiments, the fusion protein comprises a napDNAbp and an apolipoprotein B mRNA editing complex 1 catalytic polypeptide-like 3G (APOBEC3G) deaminase domain, wherein the deaminase domain is fused to the N-terminus of the napDNAbp domain via a linker (e.g., an amino acid sequence, peptide, polymer, or bond) of any length or composition. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605).

[0017] In some embodiments, the fusion protein comprises a napDNAbp and a cytidine deaminase 1 (CDA1) deaminase domain, wherein the deaminase domain is fused to the N-terminus of the napDNAbp domain via a linker comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605). In some embodiments, the napDNAbp comprises the amino acid sequence of any of the napDNAbp provided herein.

[0018] In some embodiments, the fusion protein comprises a napDNAbp and an activation-induced cytidine deaminase (AID) deaminase domain, wherein the deaminase domain is fused to the N-terminus of the napDNAbp domain via a linker comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605). In some embodiments, the napDNAbp comprises the amino acid sequence of any of the napDNAbp provided herein.

[0019] Some aspects of the present disclosure are based on the recognition that certain configurations of a napDNAbp and a cytidine deaminase domain fused via a linker are useful for efficiently deaminating target cytidine residues. Another aspect of the present disclosure relates to the recognition that a nucleobase editing fusion protein having an apolipoprotein B mRNA editing complex 1 (APOBEC1) deaminase domain fused to the N-terminus of a napDNAbp via a linker comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604) was able to efficiently deaminate a target nucleic acid in a double-stranded DNA target molecule. In some embodiments, the fusion protein comprises a napDNAbp domain and an apolipoprotein B mRNA editing complex 1 (APOBEC1) deaminase domain, wherein the deaminase domain is fused to the N-terminus of the napDNAbp via a linker comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604).

[0020] Some aspects of the present disclosure provide strategies, systems, reagents, methods, and kits that are useful for targeted editing of nucleic acids, including editing a single site in a subject's genome, e.g., a human genome. In some embodiments, fusion proteins of napDNAbp (e.g., CasX, CasY, Cpf1, C2c1, C2c2, C2c3, or Argonaute protein) and a deaminase or deaminase domain are provided. In some embodiments, methods for targeted nucleic acid editing are provided. In some embodiments, reagents and kits are provided for creating targeted nucleic acid editing proteins, e.g., fusion proteins of napDNAbp and a deaminase or deaminase domain.

[0021] Some aspects of the present disclosure provide fusion proteins, comprising a napDNAbp provided herein fused to a second protein (e.g., an enzymatic domain such as a cytidine deaminase domain), thereby forming the fusion protein. In some embodiments, the second protein comprises an enzymatic domain or a binding domain. In some embodiments, the enzymatic domain is a nuclease, nickase, recombinase, deaminase, methyltransferase, methylase, acetylase, acetyltransferase, transcriptional activator, or transcriptional repressor domain. In some embodiments, the enzymatic domain is a nucleic acid editing domain. In some embodiments, the nucleic acid editing domain is a deaminase domain. In some embodiments, the deaminase is a cytosine deaminase or cytidine deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is an APOBEC1 deaminase. In some embodiments, the deaminase is an APOBEC2 deaminase. In some embodiments, the deaminase is an APOBEC3 deaminase. In some embodiments, the deaminase is an APOBEC3A deaminase. In some embodiments, the deaminase is an APOBEC3B deaminase. In some embodiments, the deaminase is an APOBEC3C deaminase. In some embodiments, the deaminase is an APOBEC3D deaminase. In some embodiments, the deaminase is an APOBEC3E deaminase. In some embodiments, the deaminase is an APOBEC3F deaminase. In some embodiments, the deaminase is an APOBEC3G deaminase. In some embodiments, the deaminase is an APOBEC3H deaminase. In some embodiments, the deaminase is an APOBEC4 deaminase. In some embodiments, the deaminase is an activation-induced deaminase (AID). It should be understood that the deaminase can be from any suitable organism (e.g., human or rat).In some embodiments, the deaminase is from a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase is rat APOBEC1 (SEQ ID NO: 76). In some embodiments, the deaminase is human APOBEC1 (SEQ ID NO: 74). In some embodiments, the deaminase is pmCDA1.

[0022] Some aspects of the present disclosure provide fusion proteins comprising: (i) a CasX, CasY, Cpf1, C2c1, C2c2, C2c3, or Argonaute protein domain comprising the amino acid sequence of SEQ ID NO: 32; and (ii) an apolipoprotein B mRNA editing complex 1 (APOBEC1) deaminase domain, wherein the deaminase domain is fused to the N-terminus of napDNAbp via a linker comprising the amino acid sequence of SGSETPGTSESATPES (SEQ ID NO: 604). In some embodiments, the deaminase is rat APOBEC1 (SEQ ID NO: 76). In some embodiments, the deaminase is human APOBEC1 (SEQ ID NO: 74). In some embodiments, the fusion protein comprises the amino acid sequence of SEQ ID NO: 591. In some embodiments, the fusion protein comprises the amino acid sequence of SEQ ID NO: 5737. In some embodiments, the deaminase is pmCDA1 (SEQ ID NO: 81). In some embodiments, the deaminase is human APOBEC3G (SEQ ID NO: 60). In some embodiments, the deaminase is a human APOBEC3G variant of any one of SEQ ID NOs: 82-84.

[0023] Another aspect of the present disclosure relates to the recognition that a fusion protein comprising a deaminase domain, a napDNAbp domain, and a uracil glycosylase inhibitor (UGI) domain demonstrates improved efficiency in deaminating target nucleotides in nucleic acid molecules. Without wishing to be bound by any particular theory, the cellular DNA repair response to the presence of U:G heteroduplex DNA may be responsible for the decreased cellular nucleobase editing efficiency. Uracil DNA glycosylase (UDG) catalyzes the removal of U from cellular DNA, which can initiate base excision repair, the most common outcome of which is the reversion of a U:G pair to a C:G pair. As demonstrated herein, uracil DNA glycosylase inhibitor (UGI) can inhibit human UDG activity. Without wishing to be bound by any particular theory, base excision repair can be inhibited by molecules that bind to a single strand, block the edited base, inhibit UGI, inhibit base excision repair, protect the edited base, and / or promote "repair" of the unedited strand, etc. Therefore, the present disclosure contemplates a fusion protein comprising a napDNAbp-cytidine deaminase domain fused to a UGI domain.

[0024] A further aspect of the present disclosure relates to the recognition that fusion proteins comprising a deaminase domain, a napDNAbp domain, and more than one uracil glycosylase inhibitor (UGI) domain (e.g., one, two, three, four, five, or more UGI domains) demonstrate improved efficiency and / or improved nucleic acid product purity for deaminating target nucleotides in nucleic acid molecules. Without wishing to be bound by any particular theory, the addition of a second UGI domain can substantially reduce access of UDG to G:U base editing intermediates, thereby improving the efficiency of base editing.

[0025] Some aspects of the present disclosure are based on the recognition that any of the base editors provided herein can modify specific nucleotide bases without creating a significant proportion of indels. As used herein, "indel" refers to the insertion or deletion of a nucleotide base in a nucleic acid. Such insertions or deletions can lead to frameshift mutations within the coding region of a gene. In some embodiments, it is desirable to create base editors that efficiently modify (e.g., mutate or deaminate) specific nucleotides in a nucleic acid without creating insertions or deletions (i.e., indels) in the nucleic acid. In certain embodiments, any of the base editors provided herein can create a greater proportion of intended modifications (e.g., point mutations or deaminations) relative to indels.

[0026] In certain embodiments, any of the base editors provided herein can create a certain percentage of desired mutations. In some embodiments, the desired mutation is a C→T mutation. In some embodiments, the desired mutation is a C→A mutation. In some embodiments, the desired mutation is a C→G mutation. In some embodiments, any of the base editors provided herein can create at least 1% of desired mutations. In some embodiments, any of the base editors provided herein can create at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of desired mutations.

[0027] Some aspects of the present disclosure are based on the recognition that any of the base editors provided herein can efficiently create intended mutations, e.g., point mutations, in a nucleic acid (e.g., a nucleic acid in a subject's genome) without creating a significant number of unintended mutations, e.g., unintended point mutations.

[0028] In some embodiments, the deaminase domain of the fusion protein is fused to the N-terminus of the napDNAbp domain. In some embodiments, the UGI domain is fused to the C-terminus of the napDNAbp domain. In some embodiments, the napDNAbp and nucleic acid editing domains are fused via a linker. In some embodiments, the napDNAbp domain and the UGI domain are fused via a linker. In some embodiments, the second UGI domain is fused to the C-terminus of the first UGI domain. In some embodiments, the first UGI domain and the second UGI domain are fused via a linker.

[0029] In certain embodiments, a linker can be used to link any of the peptides or peptide domains of the present invention. The linker can be as simple as a covalent bond, or it can be a polymeric linker many atoms in length. In some embodiments, the linker is a polypeptide or based on amino acids. In other embodiments, the linker is not peptidic. In some embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, a disulfide bond, a carbon-heteroatom bond, etc.). In some embodiments, the linker is a carbon-nitrogen bond of an amide linkage. In some embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched, aliphatic or heteroaliphatic linker. In some embodiments, the linker is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In some embodiments, the linker comprises a monomer, dimer, or polymer of aminoalkanoic acid. In some embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In some embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In some embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In other embodiments, the linker comprises an amino acid. In some embodiments, the linker comprises a peptide. In some embodiments, the linker comprises an aryl or heteroaryl moiety. In some embodiments, the linker is based on a phenyl ring. The linker may include a functionalized moiety to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0030] In some embodiments, the linker comprises the amino acid sequence (GGGGS)n (SEQ ID NO: 607), (G) n (SEQ ID NO: 608), (EAAAK) n (SEQ ID NO: 609), (GGS) n (SEQ ID NO: 610), (SGGS) n (SEQ ID NO: 606), SGSETPGTSESATPES (SEQ ID NO: 604), (XP) n (SEQ ID NO: 611), SGGS (GGS) n (SEQ ID NO: 612), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605), or any combination thereof, wherein n is independently an integer between 1 and 30, and X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS) n (SEQ ID NO: 610), where n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGGS (GGS) n (SEQ ID NO: 612), wherein n is 2. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605).

[0031] In some embodiments, the fusion protein comprises the structure [nucleic acid editing domain]-[optional linker sequence]-[napDNAbp]-[optional linker sequence]-[UGI]. In some embodiments, the fusion protein comprises the structure [nucleic acid editing domain]-[optional linker sequence]-[UGI]-[optional linker sequence]-[napDNAbp]; [UGI]-[optional linker sequence]-[nucleic acid editing domain]-[optional linker sequence]-[napDNAbp]; [UGI]-[optional linker sequence]-[napDNAbp]-[optional linker sequence]-[nucleic acid editing domain]; [napDNAbp]-[optional linker sequence]-[nucleic acid editing domain]-[optional linker sequence]-[UGI]; or [nucleic acid editing domain]-[optional linker sequence]-[napDNAbp]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI].

[0032] In some embodiments, the nucleic acid editing domain comprises a deaminase. In some embodiments, the nucleic acid editing domain comprises a deaminase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is an APOBEC1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, an APOBEC3H deaminase, or an APOBEC4 deaminase. In some embodiments, the deaminase is an activation-induced deaminase (AID). In some embodiments, the deaminase is cytidine deaminase 1 (CDA1). In some embodiments, the deaminase is lamprey CDA1 (pmCDA1) deaminase.

[0033] In some embodiments, the deaminase is from a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase is from a human. In some embodiments, the deaminase is from a rat. In some embodiments, the deaminase is a rat APOBEC1 deaminase comprising the amino acid sequence defined by (SEQ ID NO: 76). In some embodiments, the deaminase is a human APOBEC1 deaminase comprising the amino acid sequence defined by (SEQ ID NO: 74). In some embodiments, the deaminase is pmCDA1 (SEQ ID NO: 81). In some embodiments, the deaminase is human APOBEC3G (SEQ ID NO: 60). In some embodiments, the deaminase is any one of the human APOBEC3G variants of (SEQ ID NOs: 82-84). In some embodiments, the deaminase is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences defined by SEQ ID NOs:49-84.

[0034] In some embodiments, the UGI domain comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to SEQ ID NO: 134. In some embodiments, the UGI domain comprises the amino acid sequence defined by SEQ ID NO: 134.

[0035] Some aspects of the present disclosure provide a complex comprising a napDNAbp fusion protein provided herein and a guide RNA bound to the napDNAbp.

[0036] Some aspects of the present disclosure provide methods using the napDNAbp, fusion protein, or complex provided herein. For example, some aspects of the present disclosure provide methods including contacting a DNA molecule with (a) the napDNAbp or fusion protein provided herein and a guide RNA (the guide RNA is about 15-100 nucleotides in length and includes a sequence of at least 10 consecutive nucleotides that is complementary to a target sequence); or (b) the napDNAbp, napDNAbp fusion protein, or napDNAbp or napDNAbp complex with gRNA provided herein.

[0037] Some aspects of the present disclosure provide kits that include a nucleic acid construct, the kit comprising: (a) a nucleotide sequence encoding a napDNAbp or napDNAbp fusion protein provided herein; and (b) a heterologous promoter that drives expression of the sequence of (a). In some embodiments, the kit further includes an expression construct encoding a guide RNA backbone, the construct including a cloning site positioned to allow cloning of a nucleic acid sequence identical or complementary to a target sequence into the guide RNA backbone.

[0038] Some aspects of the present disclosure provide polynucleotides encoding the napDNAbp of the fusion proteins provided herein. Some aspects of the present disclosure provide vectors comprising such polynucleotides. In some embodiments, the vector comprises a heterologous promoter driving expression of the polynucleotide.

[0039] Some aspects of the present disclosure provide cells comprising the napDNAbp proteins, fusion proteins, nucleic acid molecules, and / or vectors provided herein.

[0040] It should be understood that any of the fusion proteins provided herein that include a Cas9 domain (e.g., Cas9, nCas9, or dCas9) can be replaced with any of the napDNAbps provided herein, such as CasX, CasY, Cpf1, C2c1, C2c2, C2c3, or Argonaute proteins.

[0041] The above descriptions of exemplary embodiments of reporter systems are provided for illustrative purposes only and are not meant to be limiting. Additional reporter systems, such as variations of the exemplary systems described in detail above, are also encompassed by the present disclosure.

[0042] The above summary is meant to illustrate, in a non-limiting manner, some of the aspects, advantages, features, and uses of the technology disclosed herein. Other aspects, advantages, features, and uses of the technology disclosed herein will be apparent from the detailed description, drawings, examples, and claims. [Brief explanation of the drawings]

[0043] [Figure 1] Figure 1 shows the deaminase activity of deaminases on single-stranded DNA substrates. Single-stranded DNA substrates with randomized PAM sequences (NNN PAM) were used as negative controls. Canonical PAM sequences used include (NGG PAM).

[0044] [Figure 2] Figure 2 shows the activity of the Cas9:deaminase fusion protein on single-stranded DNA substrates.

[0045] [Figure 3] Figure 3 illustrates double-stranded DNA substrate binding by the Cas9:deaminase:sgRNA complex.

[0046] [Figure 4]FIG. 4 illustrates a double-stranded DNA deamination assay.

[0047] [Figure 5] Figure 5 demonstrates that Cas9 fusions can target positions 3-11 (numbered according to the schematic in Figure 5) of a double-stranded DNA target sequence. Upper gel: 1 μM rAPOBEC1-GGS-dCas9, 125 nM dsDNA, 1 equivalent of sgRNA. Middle gel: 1 μM rAPOBEC1-(GGS)3 (SEQ ID NO: 610)-dCas9, 125 nM dsDNA, 1 equivalent of sgRNA. Lower gel: 1.85 μM rAPOBEC1-XTEN-dCas9, 125 nM dsDNA, 1 equivalent of sgRNA.

[0048] [Figure 6] Figure 6 demonstrates that the correct guide RNA, e.g., the correct sgRNA, is required for deaminase activity.

[0049] [Figure 7] Figure 7 illustrates the mechanism of target DNA binding of in vivo target sequences by the deaminase-dCas9:sgRNA complex.

[0050] [Figure 8] FIG. 8 shows the successful deamination of an exemplary disease-associated target sequence.

[0051] [Figure 9] Figure 9 shows in vitro C→T editing efficiency using His6-rAPOBEC1-XTEN-dCas9.

[0052] [Figure 10] Figure 10 shows that fusion with UGI greatly improves C→T editing efficiency in HEK293T cells.

[0053] [Figure 11A]Figures 11A-11C show that NBE1 mediates C-to-U conversion in vitro, programmed by a specific guide RNA. Figure 11A: Nucleobase editing strategy. DNA with a target C at a locus specified by the guide RNA is bound by dCas9, which mediates local denaturation of the DNA substrate. Cytidine deamination by the tethered APOBEC1 enzyme converts the target C to a U. The resulting G:U heteroduplex can be permanently converted to an A:T base pair after DNA replication or repair. If a U is in the template DNA strand, it will also result in an RNA transcript containing a G-to-A mutation after transcription. [Figure 11B-C] Figure 11B: Deamination assay showing an approximately 5-nucleotide activity window. After incubation of the NBE1-sgRNA complex with the dsDNA substrate for 2 h at 37°C, the 5' fluorophore-labeled DNA was isolated and incubated with USER enzyme (uracil DNA glycosylase endonuclease VIII) for 1 h at 37°C to induce DNA cleavage at either uracil site. The resulting DNA was separated on a denaturing polyacrylamide gel to visualize the strands to which either fluorophore was attached. Each lane is labeled according to the position of the target C within the protospacer or by "-" if the target C is absent. The base distal to the PAM is counted as position 1. Figure 11C: Deaminase assay showing the sequence specificity and sgRNA dependence of NBE1. DNA substrates with a target C at position 7 were incubated with NBE1 with either the correct sgRNA, a mismatched sgRNA, or no sgRNA, as in Figure 11B. No C→U editing is observed with a mismatched sgRNA or no sgRNA. The positive control sample contains a DNA sequence with a synthetically incorporated U at position 7.

[0054] [Figure 12A]Figures 12A-12B show the effects of sequence context and target C position on in vitro nucleobase editing efficiency. Figure 12A: Effect of varying the sequence surrounding the target C on in vitro editing efficiency. For C7 of the protospacer sequence 5'-TTATTTCGTGGATTTATTTA-3' (SEQ ID NO: 591), the deamination yield of 80% of the target strand (40% of total sequencing reads from both strands) was defined as 1.0. Relative deamination efficiencies of substrates containing all possible single-base mutations at positions 1-6 and 8-13 are shown. Values ​​and error bars reflect the mean and standard deviation of two or more independent biological replicates performed on different days. [Figure 12B] Figure 12B: Effect of position of each NC motif on in vitro editing efficiency. Each NC target motif was varied from positions 1 to 8 within the protospacer as indicated in the sequence shown on the right (PAM is shown in red; the protospacer plus one base 5' of the protospacer is also shown). The percentage of total sequence reads containing a T at each of the numbered target C positions after incubation with NBE1 is shown graphically. Note that the maximum possible in vitro deamination yield is 50% of the total sequencing reads (100% of the target strand). Values ​​and error bars reflect the mean and standard deviation of two or three independent biological replicates performed on different days. Figure 12B illustrates SEQ ID NOs: 619-626, respectively, from top to bottom.

[0055] [Figure 13A] Figures 13A-13C show nucleobase editing in human cells. Figure 13A: Protospacer and PAM sequences of six mammalian cell genomic loci targeted by nucleobase editors. Target Cs are indicated by subscript numbers corresponding to their position within the protospacer. Figure 13A illustrates SEQ ID NOs: 127-132 from top to bottom, respectively. [Figure 13B]Figure 13B: HEK293T cells were transfected with plasmids expressing NBE1, NBE2, or NBE3 and the appropriate sgRNA. Three days after transfection, genomic DNA was extracted and analyzed by high-throughput DNA sequencing at six loci. Cellular C-to-T conversion percentages, determined as the percentage of total DNA sequencing reads with a T at the indicated target position, are shown for NBE1, NBE2, and NBE3 at all six genomic loci and for wtCas9 with the donor HDR template at three of the six sites (EMX1, HEK293 site 3, and HEK293 site 4). Values ​​and error bars reflect the mean and standard deviation of three independent biological replicates performed on different days. [Figure 13C] Figure 13C: The frequency of indel formation, calculated as described in the methods, is shown after treatment of HEK293T cells with NBE2 and NBE3 for all six genomic loci or with wtCas9 and single-stranded DNA template for HDR at three of the six sites (EMX1, HEK293 site 3, and HEK293 site 4). Values ​​reflect the average of at least three independent biological replicates performed on different days.

[0056] [Figure 14]Figures 14A-14C show NBE2- and NBE3-mediated correction of three disease-relevant mutations in mammalian cells. For each site, the sequence of the protospacer is indicated to the right of the mutation name, and the PAM and base carrying the mutation are indicated in bold with a subscripted number corresponding to its position within the protospacer. The amino acid sequence above each disease-associated allele is shown along with the corrected amino acid sequence after nucleobase editing in red. Directly below each sequence is the percentage of total sequencing reads with the corresponding base. Cells were nucleofected with plasmids encoding NBE2 or NBE3 and the appropriate sgRNA. Two days after nucleofection, genomic DNA was extracted and analyzed by HTS to assess pathogenic mutation correction. Figure 14A: The Alzheimer's disease-associated APOE4 allele is converted to APOE3' by NBE3 in mouse astrocytes in 11% of total reads (44% of nucleofected astrocytes). Two nearby Cs are also converted to Ts, but with no change in the predicted sequence of the resulting protein (SEQ ID NO: 627). Figure 14B: The cancer-associated p53 N239D mutation is corrected by NBE2 in 11% of treated human lymphoma cells (12% of nucleofected cells) that are heterozygous for the mutation (SEQ ID NO: 628). Figure 14C: The p53 Y163C mutation is corrected by NBE3 in 7.6% of nucleofected human breast cancer cells (SEQ ID NO: 629).

[0057] [Figure 15]Figures 15A-15D show the effect of deaminase-dCas9 linker length and composition on nucleobase editing. Gel-based deaminase assays showing the deamination window of nucleobase editors with the GGS (Figure 15A), (GGS)3 (SEQ ID NO: 610) (Figure 15B), XTEN (Figure 15C), or (GGS)7 (SEQ ID NO: 610) (Figure 15D) deaminase-Cas9 linkers. After incubation of 1.85 μM editor-sgRNA complexes with 125 μM dsDNA substrate at 37°C for 2 h, the dye-conjugated DNA was isolated and incubated with USER enzyme (uracil DNA glycosylase endonuclease VIII) for an additional hour at 37°C to cleave the DNA backbone at any uracil site. The resulting DNA was separated on a denaturing polyacrylamide gel, and the dye-conjugated strand was imaged. Each lane is numbered according to the position of the target C within the protospacer or - if the target C is absent. 8U is a positive control sequence with a synthetically incorporated U at position 8.

[0058] [Figure 16] Figures 16A-16B demonstrate that NBE1 can correct disease-relevant mutations in vitro. Figure 16A: Protospacer and PAM sequences for seven disease-relevant mutations. The disease-associated target C in each case is indicated by a subscripted number reflecting its position within the protospacer. For all mutations except for both APOE4 SNPs, the target C is located on the template (non-coding) strand. Figure 16A illustrates SEQ ID NOs: 631-636 from top to bottom, respectively. Figure 16B: Deaminase assay showing each dsDNA oligonucleotide before (-) and after (+) incubation with NBE1, DNA isolation, and incubation with USER enzyme to cleave DNA at U-containing positions. Positive control lanes from incubation of synthetic oligonucleotides containing U at various positions within the protospacer with USER enzyme are indicated by the corresponding numbers indicating the position of the U.

[0059] [Figure 17] Figure 17 shows the processivity of NBE1. The protospacer and PAM of a 60-mer DNA oligonucleotide containing eight consecutive Cs are shown at the top. The oligonucleotide (125 μM) was incubated with NBE1 (2 μM) for 2 h at 37°C. DNA was isolated and analyzed by high-throughput sequencing. The percent of total reads is shown for the nine most frequent sequences observed. The majority (>93%) of edited strands have more than one C converted to a T. This figure illustrates SEQ ID NO: 309.

[0060] [Figures 18A-D] Figures 18A-18H show the effect of fusing UGI to NBEl to create NBE2. Figure 18A: Protospacer and PAM sequences of six mammalian genomic loci targeted by nucleobase editors. Editable Cs are indicated by labels corresponding to their position within the protospacer. Figure 18A illustrates SEQ ID NOs: 127-132, respectively, from top to bottom. Figures 18B-18G: HEK293T cells were transfected with plasmids expressing NBEl, NBE2, or NBEl and UGI and the appropriate sgRNA. Three days after transfection, genomic DNA was extracted and analyzed by high-throughput DNA sequencing at the six loci. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with Ts at the indicated target positions, are shown for NBEl, NBEl and UGI, and NBE2 at all six genomic loci. [Figure 18E-H]Figures 18B-18G: HEK293T cells were transfected with plasmids expressing NBE1, NBE2, or NBE1 and UGI and the appropriate sgRNA. Three days after transfection, genomic DNA was extracted and analyzed by high-throughput DNA sequencing at six loci. Cellular C-to-T conversion percentages, determined as the percentage of total DNA sequencing reads with a T at the indicated target position, are shown for NBE1, NBE1 and UGI, and NBE2 at all six genomic loci. Figure 18H: C-to-T mutation rates at 510 Cs surrounding the protospacer are shown for NBE1, NBE1 plus UGI, NBE2 on separate plasmids, and untreated cells. Data represent the results of 3,000,000 DNA sequencing reads from 1.5 x 106 cells. Values ​​reflect the average of at least two biological experiments performed on different days.

[0061] [Figure 19] Figure 19 shows the nucleobase editing efficiency of NBE2 in U2OS and HEK293T cells. The percentage of cellular C-to-T conversions by NBE2 is shown for each of the six target genomic loci in HEK293T and U2OS cells. HEK293T cells were transfected with lipofectamine 2000, and U2OS cells were nucleofected. The U2OS nucleofection efficiency was 74%. Three days after plasmid delivery, genomic DNA was extracted and analyzed for nucleobase editing at the six genomic loci by HTS. Values ​​and error bars reflect the mean and standard deviation of at least two biological experiments performed on different days.

[0062] [Figure 20]Figure 20 shows that nucleobase editing persists across multiple cell divisions. The percentage of cellular C→T conversions by NBE2 is displayed at two genomic loci in HEK293T cells before and after passaging the cells. HEK293T cells were transfected using Lipofectamine 2000. Three days after transfection, the cells were harvested and split in half. One half was subjected to HTS analysis, and the other half was allowed to propagate for approximately five cell divisions, then harvested and subjected to HTS analysis.

[0063] [Figure 21] Figure 21 shows genetic variants from ClinVar that could, in principle, be corrected by nucleobase editing. The NCBI ClinVar database of human genetic variations and their corresponding phenotypes was searched for genetic diseases that could be corrected by current nucleobase editing technology. Results were filtered by imposing sequential constraints, listed on the left. The x-axis shows the number of events that satisfy that constraint and all the above constraints on a logarithmic scale.

[0064] [Figure 22-1] Figure 22 shows the in vitro identification of editable Cs at six genomic loci. Synthetic 80-mers with sequences matching six different genomic sites were incubated with NBE1 and then analyzed for nucleobase editing by HTS. For each site, the sequence of the protospacer is indicated to the right of the site name, and the PAM is highlighted in red. Directly below each sequence is the percentage of total DNA sequencing reads with the corresponding base. A targeted C was considered "editable" if the in vitro conversion efficiency was >10%. Note that the maximum yield is 50% of the total DNA sequencing reads, as the non-targeted strand is not a substrate for nucleobase editing. This figure illustrates SEQ ID NOs: 127-132, respectively, from top to bottom. [Figure 22-2]Figure 22 shows the in vitro identification of editable Cs at six genomic loci. Synthetic 80-mers with sequences matching six different genomic sites were incubated with NBE1 and then analyzed for nucleobase editing by HTS. For each site, the sequence of the protospacer is indicated to the right of the site name, and the PAM is highlighted in red. Directly below each sequence is the percentage of total DNA sequencing reads with the corresponding base. A targeted C was considered "editable" if the in vitro conversion efficiency was >10%. Note that the maximum yield is 50% of the total DNA sequencing reads, as the non-targeted strand is not a substrate for nucleobase editing. This figure illustrates SEQ ID NOs: 127-132, respectively, from top to bottom.

[0065] [Figure 23-1] Figure 23 shows the activity of NBE1, NBE2, and NBE3 at EMX1 off-targets. HEK293T cells were transfected with plasmids expressing NBE1, NBE2, or NBE3 and an sgRNA matching the EMX1 sequence using Lipofectamine 2000. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the on-target locus of the EMX1 sgRNA plus the top 10 known Cas9 off-target loci, as previously determined using the GUIDE-seq method. 5 EMX1 off-target loci were not amplified and are not shown. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. On the far right, the total number of sequencing reads reported for each sequence is displayed. This figure illustrates SEQ ID NOs: 127 and 637-645, respectively, from top to bottom. [Figure 23-2]Figure 23 shows the activity of NBE1, NBE2, and NBE3 at EMX1 off-targets. HEK293T cells were transfected with plasmids expressing NBE1, NBE2, or NBE3 and an sgRNA matching the EMX1 sequence using Lipofectamine 2000. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the on-target locus of the EMX1 sgRNA plus the top 10 known Cas9 off-target loci, as previously determined using the GUIDE-seq method. 5 EMX1 off-target loci were not amplified and are not shown. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. On the far right, the total number of sequencing reads reported for each sequence is displayed. This figure illustrates SEQ ID NOs: 127 and 637-645, respectively, from top to bottom. [Figure 23-3]Figure 23 shows the activity of NBE1, NBE2, and NBE3 at EMX1 off-targets. HEK293T cells were transfected with plasmids expressing NBE1, NBE2, or NBE3 and an sgRNA matching the EMX1 sequence using Lipofectamine 2000. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the on-target locus of the EMX1 sgRNA plus the top 10 known Cas9 off-target loci, as previously determined using the GUIDE-seq method. 5 EMX1 off-target loci were not amplified and are not shown. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. On the far right, the total number of sequencing reads reported for each sequence is displayed. This figure illustrates SEQ ID NOs: 127 and 637-645, respectively, from top to bottom.

[0066] [Figure 24-1]Figure 24 shows the activity of NBE1, NBE2, and NBE3 at FANCF off-targets. HEK293T cells were transfected with plasmids expressing NBE1, NBE2, or NBE3 and sgRNAs matching the FANCF sequence using Lipofectamine 2000. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined on-target locus of the FANCF sgRNA plus all known Cas9 off-target loci using the GUIDE-seq method. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. The total number of sequencing reads reported for each sequence is displayed on the far right. This figure illustrates SEQ ID NOs: 128 and 646-653 from top to bottom, respectively. [Figure 24-2]Figure 24 shows the activity of NBE1, NBE2, and NBE3 at FANCF off-targets. HEK293T cells were transfected with plasmids expressing NBE1, NBE2, or NBE3 and sgRNAs matching the FANCF sequence using Lipofectamine 2000. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined on-target locus of the FANCF sgRNA plus all known Cas9 off-target loci using the GUIDE-seq method. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. The total number of sequencing reads reported for each sequence is displayed on the far right. This figure illustrates SEQ ID NOs: 128 and 646-653 from top to bottom, respectively.

[0067] [Figure 25]Figure 25 shows the activity of NBE1, NBE2, and NBE3 at HEK293 site 2 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing NBE1, NBE2, or NBE3 and an sgRNA matching the HEK293 site 2 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined HEK293 site 2 sgRNA on-target locus plus all known Cas9 off-target loci using the GUIDE-seq method. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. On the far right, the total number of sequencing reads reported for each sequence is displayed. From top to bottom, this figure illustrates SEQ ID NOs: 129, 654, and 655, respectively.

[0068] [Figure 26]Figure 26 shows the activity of NBE1, NBE2, and NBE3 at HEK293 site 3 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing NBE1, NBE2, or NBE3 and an sgRNA matching the HEK293 site 3 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined HEK293 site 3 sgRNA on-target locus plus all known Cas9 off-target loci using the GUIDE-seq method. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. On the far right, the total number of sequencing reads reported for each sequence is displayed. This figure illustrates SEQ ID NOs: 130 and 656-660 from top to bottom, respectively.

[0069] [Figure 27-1]Figure 27 shows the activity of NBE1, NBE2, and NBE3 at HEK293 site 4 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing NBE1, NBE2, or NBE3 and an sgRNA matching the HEK293 site 4 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined HEK293 site 4 sgRNA on-target locus plus the top 10 known Cas9 off-target loci using the GUIDE-seq method. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. On the far right, the total number of sequencing reads reported for each sequence is displayed. This figure illustrates SEQ ID NOs: 131 and 661-670, respectively, from top to bottom. [Figure 27-2]Figure 27 shows the activity of NBE1, NBE2, and NBE3 at HEK293 site 4 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing NBE1, NBE2, or NBE3 and an sgRNA matching the HEK293 site 4 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined HEK293 site 4 sgRNA on-target locus plus the top 10 known Cas9 off-target loci using the GUIDE-seq method. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are displayed. Cellular C→T conversion percentages, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, are shown for NBE1, NBE2, and NBE3. On the far right, the total number of sequencing reads reported for each sequence is displayed. This figure illustrates SEQ ID NOs: 131 and 661-670, respectively, from top to bottom.

[0070] [Figure 28] Figure 28 shows the non-target C mutation rate, where C→T mutation rates at 2,500 distinct cytosines surrounding the six on-target and 34 off-target loci tested are shown, representing a total of 14,700,000 sequence reads derived from approximately 1.8 x 10 cells.

[0071] [Figure 29A]Figures 29A-29C show base editing in human cells. Figure 29A shows possible base editing outcomes in mammalian cells. The initial edit resulted in a U:G mismatch. Recognition and excision of the U by uracil DNA glycosylase (UDG) initiated base excision repair (BER), which led to a return to the C:G starting state. BER was prevented by BE2 and BE3, which inhibit UDG. The U:G mismatch was also processed by mismatch repair (MMR), which preferentially repaired the nicked strand of the mismatch. BE3 nicked the unedited strand containing the G, preferentially resolving the U:G mismatch to the desired U:A or T:A outcome. [Figure 29B] Figure 29B shows HEK293T cells treated as described in the Examples Materials and Methods below. The percentage of total DNA sequencing reads with a T at the indicated target position indicates treatment with BE1, BE2, or BE3, or with wtCas9 with the donor HDR template. [Figure 29C] Figure 29C shows the frequency of indel formation after the treatments in Figure 29B. Values ​​are listed in Figure 34. In Figures 29B and 29C, values ​​and error bars reflect the mean and sd of three independent biological replicates performed on different days.

[0072] [Figure 30A]Figures 30A-30B show BE3-mediated correction of two disease-relevant mutations in mammalian cells. The sequence of the protospacer is shown to the right of the mutation; the PAM and target base are in red with a subscripted number indicating their position within the protospacer. Directly below each sequence is the percentage of total sequencing reads with the corresponding base. Cells were treated as described in Materials and Methods. Figure 30A shows the Alzheimer's disease-associated APOE4 allele converted to APOE3r by BE3 in 74.9% of total reads in mouse astrocytes. Two nearby Cs were also converted to Ts, but without any predicted sequence changes in the resulting protein. Identical treatment of these cells with wtCas9 and donor ssDNA resulted in only 0.3% correction, with 26.1% indel formation. This figure illustrates SEQ ID NOs: 671 and 627. [Figure 30B] Figure 30B shows that the cancer-associated p53 Y163C mutation was corrected by BE3 in 7.6% of nucleofected human breast cancer cells, with 0.7% indel formation. Identical treatment of these cells with wtCas9 and donor ssDNA did not result in mutation correction, with 6.1% indel formation. This figure illustrates SEQ ID NOs: 672 and 629.

[0073] [Figure 31-1]Figure 31 shows the activity of BE1, BE2, and BE3 at HEK293 site 2 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing BE1, BE2, or BE3 and sgRNAs matching the HEK293 site 2 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the on-target locus of the HEK293 site 2 sgRNA, as previously determined by Joung et al. (63) using GUIDE-seq and Adli et al. (18) using chromatin immunoprecipitation high-throughput sequencing (ChIP-seq) experiments, plus all known Cas9 and dCas9 off-target loci. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are shown. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, is shown for BE1, BE2, and BE3. On the far right, the total number of reported sequencing reads and reported ChIP-seq signal intensity are displayed for each sequence. From top to bottom, this figure illustrates SEQ ID NOs: 129, 654, 655, and 673-677, respectively. [Figure 31-2]Figure 31 shows the activity of BE1, BE2, and BE3 at HEK293 site 2 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing BE1, BE2, or BE3 and sgRNAs matching the HEK293 site 2 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the on-target locus of the HEK293 site 2 sgRNA, as previously determined by Joung et al. (63) using GUIDE-seq and Adli et al. (18) using chromatin immunoprecipitation high-throughput sequencing (ChIP-seq) experiments, plus all known Cas9 and dCas9 off-target loci. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are shown. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, is shown for BE1, BE2, and BE3. On the far right, the total number of reported sequencing reads and reported ChIP-seq signal intensity are displayed for each sequence. From top to bottom, this figure illustrates SEQ ID NOs: 129, 654, 655, and 673-677, respectively.

[0074] [Figure 32-1]Figure 32 shows the activity of BE1, BE2, and BE3 at HEK293 site 3 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing BE1, BE2, or BE3 and sgRNAs matching the HEK293 site 3 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the on-target locus of the HEK293 site 3 sgRNA, as well as all known Cas9 off-target loci and the top five known dCas9 off-target loci previously determined by Joung et al. using the GUIDE-seq method and chromatin immunoprecipitation high-throughput sequencing (ChIP-seq) experiments, respectively. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are shown. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, is shown for BE1, BE2, and BE3. On the far right, the total number of reported sequencing reads and reported ChIP-seq signal intensity are displayed for each sequence. From top to bottom, this figure illustrates SEQ ID NOs: 130, 656-660, and 678-682, respectively. [Figure 32-2]Figure 32 shows the activity of BE1, BE2, and BE3 at HEK293 site 3 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing BE1, BE2, or BE3 and sgRNAs matching the HEK293 site 3 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the on-target locus of the HEK293 site 3 sgRNA, as well as all known Cas9 off-target loci and the top five known dCas9 off-target loci previously determined by Joung et al. using the GUIDE-seq method and chromatin immunoprecipitation high-throughput sequencing (ChIP-seq) experiments, respectively. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are shown. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at each position of the original C within the protospacer, is shown for BE1, BE2, and BE3. On the far right, the total number of reported sequencing reads and reported ChIP-seq signal intensity are displayed for each sequence. From top to bottom, this figure illustrates SEQ ID NOs: 130, 656-660, and 678-682, respectively.

[0075] [Figure 33-1]Figure 33 shows the activity of BE1, BE2, and BE3 at HEK293 site 4 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing BE1, BE2, or BE3 and sgRNAs matching the HEK293 site 4 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined HEK293 site 4 sgRNA on-target locus using GUIDE-seq and chromatin immunoprecipitation high-throughput sequencing (ChIP-seq) experiments, as well as the top 10 known Cas9 off-target loci and top 5 known dCas9 off-target loci. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are shown. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at each position of the original C in the protospacer, is shown for BE1, BE2, and BE3. On the far right, the total number of reported sequencing reads and reported ChIP-seq signal intensity are displayed for each sequence. From top to bottom, this figure illustrates SEQ ID NOs: 131, 661-670, 683, and 684, respectively. [Figure 33-2]Figure 33 shows the activity of BE1, BE2, and BE3 at HEK293 site 4 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing BE1, BE2, or BE3 and sgRNAs matching the HEK293 site 4 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined HEK293 site 4 sgRNA on-target locus using GUIDE-seq and chromatin immunoprecipitation high-throughput sequencing (ChIP-seq) experiments, as well as the top 10 known Cas9 off-target loci and top 5 known dCas9 off-target loci. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are shown. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at each position of the original C in the protospacer, is shown for BE1, BE2, and BE3. On the far right, the total number of reported sequencing reads and reported ChIP-seq signal intensity are displayed for each sequence. From top to bottom, this figure illustrates SEQ ID NOs: 131, 661-670, 683, and 684, respectively. [Figure 33-3]Figure 33 shows the activity of BE1, BE2, and BE3 at HEK293 site 4 off-targets. HEK293T cells were transfected using Lipofectamine 2000 with plasmids expressing BE1, BE2, or BE3 and sgRNAs matching the HEK293 site 4 sequence. Three days after transfection, genomic DNA was extracted, amplified by PCR, and analyzed by high-throughput DNA sequencing at the previously determined HEK293 site 4 sgRNA on-target locus using GUIDE-seq and chromatin immunoprecipitation high-throughput sequencing (ChIP-seq) experiments, as well as the top 10 known Cas9 off-target loci and top 5 known dCas9 off-target loci. The sequences of the on-target and off-target protospacers and protospacer adjacent motifs (PAMs) are shown. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at each position of the original C in the protospacer, is shown for BE1, BE2, and BE3. On the far right, the total number of reported sequencing reads and reported ChIP-seq signal intensity are displayed for each sequence. From top to bottom, this figure illustrates SEQ ID NOs: 131, 661-670, 683, and 684, respectively.

[0076] [Figure 34-1]Figure 34 shows the mutation rate of non-protospacer bases in mouse astrocytes after BE3-mediated correction of the Alzheimer's disease-associated APOE4 allele to APOE3r. The 50-base DNA sequence on either side of the protospacer from Figures 30A and 34B is shown, along with the position of each base relative to the protospacer. The side of the protospacer distal to the PAM is designated by a positive number, while the side encompassing the PAM, along with the PAM, is designated by a negative number. Directly below each sequence is the percentage of total DNA sequencing reads containing the corresponding base for untreated cells, cells treated with sgRNAs targeting BE3 and the APOE4 C158R mutation, or cells treated with sgRNAs targeting BE3 and the VEGFA locus. Neither BE3-treated sample resulted in a mutation rate above that of the untreated control. This figure illustrates SEQ ID NOs: 685-688, respectively, from top to bottom. [Figure 34-2] Figure 34 shows the mutation rate of non-protospacer bases in mouse astrocytes after BE3-mediated correction of the Alzheimer's disease-associated APOE4 allele to APOE3r. The 50-base DNA sequence on either side of the protospacer from Figures 30A and 34B is shown, along with the position of each base relative to the protospacer. The side of the protospacer distal to the PAM is designated by a positive number, while the side encompassing the PAM, along with the PAM, is designated by a negative number. Directly below each sequence is the percentage of total DNA sequencing reads containing the corresponding base for untreated cells, cells treated with sgRNAs targeting BE3 and the APOE4 C158R mutation, or cells treated with sgRNAs targeting BE3 and the VEGFA locus. Neither BE3-treated sample resulted in a mutation rate above that of the untreated control. This figure illustrates SEQ ID NOs: 685-688, respectively, from top to bottom.

[0077] [Figure 35-1]Figure 35 shows the mutation rate of non-protospacer bases after BE3-mediated correction of the cancer-associated p53 Y163C mutation in HCC1954 human cells. The 50-base DNA sequence on either side of the protospacer from Figures 30B and 39B is shown, along with the position of each base relative to the protospacer. The side of the protospacer distal to the PAM is designated by a positive number, while the side encompassing the PAM, along with the PAM, is designated by a negative number. Directly below each sequence is the percentage of total sequencing reads containing the corresponding base for untreated cells, cells treated with BE3 and an sgRNA targeting the TP53 Y163C mutation, or cells treated with BE3 and an sgRNA targeting the VEGFA locus. Neither BE3-treated sample resulted in a mutation rate above that of the untreated control. This figure illustrates SEQ ID NOs: 689-692, respectively, from top to bottom. [Figure 35-2] Figure 35 shows the mutation rate of non-protospacer bases after BE3-mediated correction of the cancer-associated p53 Y163C mutation in HCC1954 human cells. The 50-base DNA sequence on either side of the protospacer from Figures 30B and 39B is shown, along with the position of each base relative to the protospacer. The side of the protospacer distal to the PAM is designated by a positive number, while the side encompassing the PAM, along with the PAM, is designated by a negative number. Directly below each sequence is the percentage of total sequencing reads containing the corresponding base for untreated cells, cells treated with BE3 and an sgRNA targeting the TP53 Y163C mutation, or cells treated with BE3 and an sgRNA targeting the VEGFA locus. Neither BE3-treated sample resulted in a mutation rate above that of the untreated control. This figure illustrates SEQ ID NOs: 689-692, respectively, from top to bottom.

[0078] [Figure 36A-C]Figures 36A-36F show the effects of deaminase, linker length, and linker composition on base editing. Figure 36A shows a gel-based deaminase assay demonstrating the activity of rAPOBEC1, pmCDA1, hAID, hAPOBEC3G, rAPOBEC1-GGS-dCas9, rAPOBEC1-(GGS)3 (SEQ ID NO: 610)-dCas9, and dCas9-(GGS)3 (SEQ ID NO: 610)-rAPOBEC1 on ssDNA. The enzymes were expressed in an in vitro transcription-translation system from mammalian cell lysates and incubated with 1.8 μM dye-conjugated ssDNA and USER enzyme (uracil-DNA glycosylase endonuclease VIII) for 2 hours at 37°C. The resulting DNA was separated on a denaturing polyacrylamide gel and imaged. The positive control was a sequence with a synthetically incorporated U at the same position as the target C. Figure 36B shows a Coomassie-stained denaturing PAGE gel of the expressed and purified proteins used in Figures 36C-36F. Figures 36C-36F show gel-based deaminase assays, demonstrating the deamination window of base editors with deaminase-Cas9 linkers of GGS (Figure 36C), (GGS)3 (SEQ ID NO: 610) (Figure 36D), XTEN (Figure 36E), or (GGS)7 (SEQ ID NO: 610) (Figure 36F). After incubation of 1.85 μM deaminase-dCas9 fusion complexed with sgRNA with 125 nM dsDNA substrate at 37°C for 2 hours, the dye-conjugated DNA was isolated and incubated with USER enzyme for 1 hour at 37°C to cleave the DNA backbone at any uracil site. The resulting DNA was separated on a denaturing polyacrylamide gel, and the dye-conjugated strands were imaged. Each lane is numbered according to the position of the target C within the protospacer, or by - if the target C is absent. 8U is a positive control sequence with a synthetically incorporated U at position 8. [Fig. 36D-F]Figures 36C-36F show gel-based deaminase assays demonstrating the deamination window of base editors with deaminase-Cas9 linkers of GGS (Figure 36C), (GGS)3 (SEQ ID NO: 610) (Figure 36D), XTEN (Figure 36E), or (GGS)7 (SEQ ID NO: 610) (Figure 36F). Following a 2-hour incubation of 1.85 μM deaminase-dCas9 fusion complexed with sgRNA with 125 nM dsDNA substrate at 37°C, the dye-conjugated DNA was isolated and incubated with USER enzyme for 1 hour at 37°C to cleave the DNA backbone at any uracil site. The resulting DNA was separated on a denaturing polyacrylamide gel, and the dye-conjugated strand was imaged. Each lane is numbered according to the position of the target C within the protospacer, or - if no target C is present. 8U is a positive control sequence with a synthetically incorporated U at position 8.

[0079] [Figure 37A] Figures 37A-37C show that BE1 base editing efficiency is dramatically reduced in mammalian cells. Figure 37A: Protospacer and PAM sequences of six mammalian cell genomic loci targeted by base editors. Target Cs are indicated in red by subscript numbers corresponding to their position within the protospacer. [Figure 37B] Figure 37B shows that synthetic 80-mers with sequences matching six different genomic sites were incubated with BE1 and then analyzed by HTS for base editing. For each site, the sequence of the protospacer, along with the PAM, is indicated to the right of the site name. Directly below each sequence is the percentage of total DNA sequencing reads with the corresponding base. We considered the target C "editable" if the in vitro conversion efficiency was >10%. Note that the maximum yield is 50% of the total DNA sequencing reads, because the non-target strand is unaffected by BE1. Values ​​are shown from a single experiment. [Figure 37C]Figure 37C shows HEK293T cells were transfected with a plasmid expressing BE1 and the appropriate sgRNA. Three days after transfection, genomic DNA was extracted and analyzed by high-throughput DNA sequencing at six loci. The cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at the indicated target position, is shown for BE1 at all six genomic loci. Values ​​and error bars for all data from HEK293T cells reflect the mean and standard deviation of three independent biological replicates performed on different days. Figure 37A illustrates SEQ ID NOs: 127-132, respectively, from top to bottom. Figure 37B illustrates SEQ ID NOs: 127-132, respectively, from top to bottom.

[0080] [Figure 38] Figure 38 shows that base editing persists across multiple cell divisions. The percentage of cellular C-to-T conversions by BE2 and BE3 is shown for HEK293 sites 3 and 4 in HEK293T cells before and after passaging the cells. HEK293T cells were nucleofected with a plasmid expressing BE2 or BE3 and an sgRNA targeting HEK293 sites 3 or 4. Three days after nucleofection, the cells were harvested and split in half. One half was subjected to HTS analysis, and the other half was allowed to propagate for approximately five cell divisions, then harvested and subjected to HTS analysis. Values ​​and error bars reflect the mean and standard deviation of at least two biological experiments.

[0081] [Figure 39A]Figures 39A-39C show non-target C / G mutation rates. Shown here are C→T and G→A mutation rates at 2,500 distinct cytosines and guanines surrounding the six on-target and 34 off-target loci tested, representing a total of 14,700,000 sequence reads derived from approximately 1.8 × 10 cells. Figures 39A and 39B show the cellular non-target C→T and G→A conversion percentages by BE1, BE2, and BE3 plotted individually for all 2,500 cytosines / guanines against their position relative to the protospacer. The side of the protospacer distal to the PAM is designated by a positive number, and the side encompassing the PAM is designated by a negative number. [Figure 39B] Figures 39A and 39B show the cellular non-target C→T and G→A conversion percentages by BE1, BE2, and BE3 plotted individually for all 2,500 cytosines / guanines versus their position relative to the protospacer. The side of the protospacer distal to the PAM is designated by a positive number, and the side including the PAM is designated by a negative number. [Figure 39C] Figure 39C shows the average cellular non-targeted C→T and G→A conversion percentages by BE1, BE2, and BE3, as well as the highest and lowest individual conversion percentages are shown.

[0082] [Figure 40A]Figures 40A-40B show additional datasets of BE3-mediated correction of two disease-relevant mutations in mammalian cells. For each site, the sequence of the protospacer is indicated to the right of the mutation name, and the PAM and base carrying the mutation are indicated in bold red with a subscript number corresponding to its position within the protospacer. The amino acid sequence above each disease-associated allele is shown along with the corrected amino acid sequence after base editing. Directly below each sequence is the percentage of total sequencing reads with the corresponding base. Cells were nucleofected with a plasmid encoding BE3 and the appropriate sgRNA. Two days after nucleofection, genomic DNA was extracted from the nucleofected cells and analyzed by HTS to assess pathogenic mutation correction. Figure 40A shows that the Alzheimer's disease-associated APOE4 allele was converted to APOE3r by BE3 in 58.3% of total reads in mouse astrocytes only when treated with the correct sgRNA. Two nearby Cs are also converted to Ts, but with no predicted sequence change in the resulting protein. Identical treatment of these cells with wtCas9 and donor ssDNA results in 0.2% correction, with 26.7% indel formation. [Figure 40B] Figure 40B shows that the cancer-associated p53 Y163C mutation is corrected by BE3 in 3.3% of nucleofected human breast cancer cells only when treated with the correct sgRNA. Identical treatment of these cells with wtCas9 and donor ssDNA results in no detectable mutation correction, with 8.0% indel formation. Figures 40A to 40B depict SEQ ID NOs: 671, 627, 672, and 629.

[0083] [Figure 41] FIG. 41 shows a schematic of an exemplary USER (Uracil-Specific Excision Reagent) enzyme-based assay, which can be used to test the activity of various deaminases on single-stranded DNA (ssDNA) substrates.

[0084] [Figure 42] Figure 42 is a schematic diagram of the pmCDA-nCas9-UGI-NLS construct and its activity at the HeK-3 site relative to the base editor (rAPOBEC1) and negative control (untreated). This figure depicts SEQ ID NO: 693.

[0085] [Figure 43] Figure 43 is a schematic diagram of the pmCDA1-XTEN-nCas9-UGI-NLS construct and its activity at the HeK-3 site relative to the base editor (rAPOBEC1) and negative control (untreated). This figure depicts SEQ ID NO: 694.

[0086] [Figure 44] Figure 44 shows the percent of total sequencing reads with target C converted to T using cytidine deaminase (CDA) or APOBEC.

[0087] [Figure 45] Figure 45 shows the percent of total sequencing reads with target C converted to A using deaminase (CDA) or APOBEC.

[0088] [Figure 46] Figure 46 shows the percent of total sequencing reads with the target C converted to G using a deaminase (CDA) or APOBEC.

[0089] [Figure 47] Figure 47 is a schematic diagram of the huAPOBEC3G-XTEN-nCas9-UGI-NLS construct and its activity at the HeK-2 site relative to a mutant form (huAPOBEC3G*(D316R_D317R)-XTEN-nCas9-UGI-NLS), a base editor (rAPOBEC1), and a negative control (untreated). This figure depicts SEQ ID NO: 695.

[0090] [Figure 48]FIG. 48 shows a schematic diagram of the LacZ construct used in the selection assay of Example 7.

[0091] [Figure 49-1] Figure 49 shows the reversion data from different plasmids and constructs. [Figure 49-2] Figure 49 shows the reversion data from different plasmids and constructs.

[0092] [Figure 50] Figure 50 shows the verification of lacZ reversion and purification of reverted clones.

[0093] [Figure 51] FIG. 51 is a schematic diagram illustrating the deamination selection plasmid used in Example 7.

[0094] [Figure 52] FIG. 52 shows the results of a chloramphenicol reversion assay (pmCDA1 fusions).

[0095] [Figure 53A] Figures 53A-53B demonstrate DNA correction induction of the two constructs. [Figure 53B] Figures 53A-53B demonstrate DNA correction induction of the two constructs.

[0096] [Figure 54] Figure 54 shows the results of a chloramphenicol reversion assay (huAPOBEC3G fusions).

[0097] [Figure 55-1] Figure 55 shows the activity of BE3 and HF-BE3 at EMX1 off-targets. Sequences correspond, from top to bottom, to SEQ ID NOs: 127 and 637-645. [Figure 55-2] Figure 55 shows the activity of BE3 and HF-BE3 at EMX1 off-targets. Sequences correspond, from top to bottom, to SEQ ID NOs: 127 and 637-645.

[0098] [Figure 56] Figure 56 shows the on-target base editing efficiency of BE3 and HF-BE3.

[0099] [Figure 57] Figure 57 is a graph demonstrating that mutations affect cytidine deamination to varying degrees. Combinations of mutations, each of which mildly impairs catalysis, allow selective deamination at one position over another. The FANCF site was GGAATC6C7C8TTTC11TGCAGCACCTGG (SEQ ID NO: 128).

[0100] [Figure 58] Figure 58 is a schematic illustrating next generation base editors.

[0101] [Figure 59] Figure 59 is a schematic illustrating new base editors made from Cas9 variants.

[0102] [Figure 60] Figure 60 shows the base-edited percentage of different NGA PAM sites.

[0103] [Figure 61] Figure 61 shows the base-edited percentage of cytidines using the NGCG PAM EMX(VRER BE3) and C1TC3C4C5ATC8AC10ATCAACCGGT (SEQ ID NO: 696) spacer.

[0104] [Figure 62] Figure 62 shows the base-edited percentage resulting from different NNGRRT PAM sites.

[0105] [Figure 63] Figure 63 shows the base-edited percentage resulting from different NNHRRT PAM sites.

[0106] [Figure 64] Figures 64A-64C show the base-edited percentages resulting from different TTTN PAM sites using Cpfl BE2. The spacers used were: TTTCCTC3C4C5C6C7C8C9AC11AGGTAGAACAT (Figure 64A, SEQ ID NO: 697), TTTCC1C2TC4TGTC8C9AC11ACCCTCATCCTG (Figure 64B, SEQ ID NO: 698), and TTTCC1C2C3AGTC7C8TC10C11AC13AC15C16C17TGAAAC (Figure 64C, SEQ ID NO: 699).

[0107] [Figure 65] Figure 65 is a schematic diagram illustrating selective deamination achieved by kinetic modulation of cytidine deaminase point mutagenesis.

[0108] [Figure 66] Figure 66 is a graph showing the effect of various mutations on the deamination window probed by cell culture with multiple cytidines in the spacer. The spacer used was: TGC3C4C5C6TC8C9C10TC12C13C14TGGCCC (SEQ ID NO: 700).

[0109] [Figure 67] Figure 67 is a graph showing the effect of various mutations on the deamination window probed by cell culture with multiple cytidines in the spacer. The spacer used was: AGAGC5C6C7C8C9C10C11TC13AAAGAGA (SEQ ID NO: 701).

[0110] [Figure 68]Figure 68 is a graph showing the effect of various mutations on a FANCF site with a limited number of cytidines. The spacer used was: GGAATC6C7C8TTTC11TGCAGCACCTGG (SEQ ID NO: 128). Note that the triple mutant (W90Y, R126E, R132E) preferentially edits the cytidine at position 6.

[0111] [Figure 69] Figure 69 is a graph showing the effect of various mutations on a HEK3 site with a limited number of cytidines. The spacer used was: GGCC4C5AGACTGAGCACGTGATGG (SEQ ID NO: 702). Note that the double and triple mutants preferentially edit the cytidine at the fifth position compared to the cytidine at the fourth position.

[0112] [Figure 70] Figure 70 is a graph showing the effect of various mutations on an EMX1 site with a limited number of cytidines. The spacer used was: GAGTC5C6GAGCAGAAGAAGAAGGG (SEQ ID NO: 703). Note that the triple mutant only edits the cytidine at the fifth position, not the sixth.

[0113] [Figure 71] Figure 71 is a graph showing the effect of various mutations on a HEK2 site with a limited number of cytidines. The spacer used was: GAAC4AC6AAAGCATAGACTGCGGG (SEQ ID NO: 704).

[0114] [Figure 72] Figure 72 shows the on-target base editing efficiency of BE3 and BE3 containing the mutation W90Y R132E in immortalized astrocytes.

[0115] [Figure 73] Figure 73 illustrates a schematic of the three Cpf1 fusion constructs.

[0116] [Figure 74] Figure 74 shows a comparison of plasmid delivery of BE3 and HF-BE3 (EMX1, FANCF, and RNF2).

[0117] [Figure 75] Figure 75 shows a comparison of plasmid delivery of BE3 and HF-BE3 (HEK3 and HEK4).

[0118] [Figure 76-1] Figure 76 shows off-target editing of EMX-1 at all 10 sites. This figure illustrates SEQ ID NOs: 127 and 637-645. [Figure 76-2] Figure 76 shows off-target editing of EMX-1 at all 10 sites. This figure illustrates SEQ ID NOs: 127 and 637-645.

[0119] [Figure 77] Figure 77 shows deaminase protein lipofection into HEK cells using a GAGTCCGAGCAGAAGAAGAAG (SEQ ID NO: 705) spacer. EMX-1 on-target and EMX-1 off-target site 2 were investigated.

[0120] [Figure 78] Figure 78 shows deaminase protein lipofection into HEK cells using a GGAATCCCTTCTGCAGCACCTGG (SEQ ID NO: 706) spacer. FANCF on-target and FANCF off-target sites 1 were investigated.

[0121] [Figure 79] Figure 79 shows deaminase protein lipofection into HEK cells using the GGCCCAGACTGAGCACGTGA (SEQ ID NO: 707) spacer. HEK-3 on-target sites were investigated.

[0122] [Figure 80-1]Figure 80 shows deaminase protein lipofection into HEK cells using the GGCACTGCGGCTGGAGGTGGGGG (SEQ ID NO: 708) spacer. HEK-4 on-target, off-target site 1, site 3, and site 4. [Figure 80-2] Figure 80 shows deaminase protein lipofection into HEK cells using the GGCACTGCGGCTGGAGGTGGGGG (SEQ ID NO: 708) spacer. HEK-4 on-target, off-target site 1, site 3, and site 4.

[0123] [Figure 81] Figure 81 shows the results of in vitro assays of sgRNA activity for sgHR_13 (GTCAGGTCGAGGGTTCTGTC (SEQ ID NO: 709) spacer; C8 target: G51 → stop), sgHR_14 (GGGCCGCAGTATCCTCACTC (SEQ ID NO: 710) spacer; C7 target; C7 target: Q68 → stop), and sgHR_15 (CCGCCAGTCCCAGTACGGGA (SEQ ID NO: 711) spacer; C10 and C11 are targets: W239 or W237 → stop).

[0124] [Figure 82] Figure 82 shows the results of in vitro assays for sgHR_17 (CAACCACTGCTCAAAGATGC (SEQ ID NO: 712) spacer; C4 and C5 are targets: W410 → stop), and sgHR_16 (CTTCCAGGATGAGAACACAG (SEQ ID NO: 713) spacer; C4 and C5 are targets: W273 → stop).

[0125] [Figure 83] Figure 83 shows direct injection of BE3 protein complexed with sgHR_13 in zebrafish embryos.

[0126] [Figure 84]Figure 84 shows direct injection of BE3 protein complexed with sgHR_16 in zebrafish embryos.

[0127] [Figure 85] Figure 85 shows direct injection of BE3 protein complexed with sgHR_17 in zebrafish embryos.

[0128] [Figure 86] Figure 86 shows exemplary nucleic acid changes that can be made using base editors that can make cytosine to thymine changes.

[0129] [Figure 87] Figure 87 shows an example of an apolipoprotein E (APOE) isoform demonstrating how a base editor (e.g., BE3) can be used to edit one APOE isoform (e.g., APOE4) into another APOE isoform (e.g., APOE3r) that is associated with a reduced risk of Alzheimer's disease.

[0130] [Figure 88] Figure 88 shows APOE4 to APOE3r base editing in mouse astrocytes. This figure depicts SEQ ID NOs: 671 and 627.

[0131] [Figure 89] Figure 89 shows base editing of PRNP to cause early truncation of the protein at arginine residue 37. This figure depicts SEQ ID NOs: 577 and 714.

[0132] [Figure 90-1] Figure 90 shows that knocking out UDG (which UGI inhibits) dramatically improves the cleanliness of C→T base editing efficiency. [Figure 90-2] Figure 90 shows that knocking out UDG (which UGI inhibits) dramatically improves the cleanliness of C→T base editing efficiency. [Figure 90-3] Figure 90 shows that knocking out UDG (which UGI inhibits) dramatically improves the cleanliness of C→T base editing efficiency. [Figure 90-4] Figure 90 shows that knocking out UDG (which UGI inhibits) dramatically improves the cleanliness of C→T base editing efficiency.

[0133] [Figure 91-1] Figure 91 shows that use of a base editor with a nickase but without a UGI leads to a mixture of results, with a very high indel rate. [Figure 91-2] Figure 91 shows that use of a base editor with a nickase but without a UGI leads to a mixture of results, with a very high indel rate.

[0134] [Fig. 92A-B] Figures 92A-92G show that SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, and VRER-BE3 mediate efficient base editing at target sites containing non-NGG PAMs in human cells. Figure 92A shows the base editor architecture using S. pyogenes and S. aureus Cas9. Figure 92B shows recently characterized Cas9 variants with alternative or relaxed PAM requirements. [Figure 92C]

[00139] Figures 92C and 92D show HEK293T cells treated with the indicated base editor variants as described in Example 12. The percentage of total DNA sequencing reads (without enrichment for transfected cells) with C converted to T at the indicated target position is shown. The PAM sequence for each target tested is shown below the X-axis. The charts show results for SaBE3 and SaKKH-BE3 at a genomic locus with an NNGRRT PAM (Figure 92C), SaBE3 and SaKKH-BE3 at a genomic locus with an NNNRRT PAM (Figure 92D), VQR-BE3 and EQR-BE3 at a genomic locus with an NGAG PAM (Figure 92E) and an NGAH PAM (Figure 92F), and VRER-BE3 at a genomic locus with an NGCG PAM (Figure 92G). Values ​​and error bars reflect the mean and standard deviation of at least two biological replicates. [Figure 92D]

[00139] Figures 92C and 92D show HEK293T cells treated with the indicated base editor variants as described in Example 12. The percentage of total DNA sequencing reads (without enrichment for transfected cells) with C converted to T at the indicated target position is shown. The PAM sequence for each target tested is shown below the X-axis. The charts show results for SaBE3 and SaKKH-BE3 at a genomic locus with an NNGRRT PAM (Figure 92C), SaBE3 and SaKKH-BE3 at a genomic locus with an NNNRRT PAM (Figure 92D), VQR-BE3 and EQR-BE3 at a genomic locus with an NGAG PAM (Figure 92E) and an NGAH PAM (Figure 92F), and VRER-BE3 at a genomic locus with an NGCG PAM (Figure 92G). Values ​​and error bars reflect the mean and standard deviation of at least two biological replicates. [Fig. 92E-F]The chart shows the results for SaBE3 and SaKKH-BE3 at genomic loci with NNGRRT PAMs (Figure 92C), SaBE3 and SaKKH-BE3 at genomic loci with NNNRRT PAMs (Figure 92D), VQR-BE3 and EQR-BE3 at genomic loci with NGAG PAMs (Figure 92E) and NGAH PAMs (Figure 92F), and VRER-BE3 at genomic loci with NGCG PAMs (Figure 92G). [Figure 92G] The chart shows the results for SaBE3 and SaKKH-BE3 at genomic loci with NNGRRT PAMs (Figure 92C), SaBE3 and SaKKH-BE3 at genomic loci with NNNRRT PAMs (Figure 92D), VQR-BE3 and EQR-BE3 at genomic loci with NGAG PAMs (Figure 92E) and NGAH PAMs (Figure 92F), and VRER-BE3 at genomic loci with NGCG PAMs (Figure 92G).

[0135] [Figure 93A] Figures 93A-93C demonstrate that base editors with mutations in the cytidine deaminase domain exhibit narrowed editing windows. Figures 93A-93C show HEK293T cells transfected with plasmids expressing mutant base editors and appropriate sgRNAs. Three days after transfection, genomic DNA was extracted and analyzed by high-throughput DNA sequencing at the indicated loci. The percentage of total DNA sequencing reads (without enrichment for transfected cells) with a C changed to T at the indicated target positions is shown for the EMX1 site (SEQ ID NO: 721), HEK293 site 3 (SEQ ID NO: 719), FANCF site (SEQ ID NO: 722), HEK293 site 2 (SEQ ID NO: 720), site A (SEQ ID NO: 715), and site B (SEQ ID NO: 718) loci. Figure 93A illustrates certain cytidine deaminase mutations that narrow the base editing window. See Figure 98 for characterization of additional mutations. [Figure 93B]Figure 93B shows the effect of cytidine deaminase mutations on achieving editing window width at a genomic locus. Combining beneficial mutations has an additive effect on narrowing the editing window. [Figure 93C] Figure 93C shows that YE1-BE3, YE2-BE3, EE-BE3, and YEE-BE3 achieve a base-editing product distribution that predominantly produces single-modified products, in contrast to BE3. Values ​​and error bars reflect the mean and standard deviation of at least two biological replicates.

[0136] [Figure 94] Figures 94A and 94B show genetic variants from ClinVar that could, in principle, be corrected by the base editors developed in this study. The NCBI ClinVar database of human gene variants and their corresponding phenotypes was searched for genetic diseases that could theoretically be corrected by base editing. Figure 94A demonstrates the improved base editing targeting scope at all pathogenic T→C mutations in the ClinVar database by using base editors with altered PAM specificity. The white fraction indicates the proportion of accessible pathogenic T→C mutations based on the PAM requirements of either BE3 or BE3 with the five altered PAM base editors developed in this study. Figure 94B shows the improved base editing targeting scope at all pathogenic T→C mutations in the ClinVar database by using base editors with narrowed activity windows. As shown in Figures 93A-93C, BE3 was predicted to edit Cs at positions 4-8 with comparable efficiency. YEE-BE3 was expected to edit with a preference for C5 > C6 > C7 > others within its activity window. The white fraction indicates the proportion of pathogenic T→C mutations that can be edited by BE3 without equivalent editing of other Cs (left) or by BE3 or YEE-BE3 without equivalent editing of other Cs (right).

[0137] [Figure 95A]Figures 95A-95B show the effect of shortened guide RNAs on base editing window width. HEK293T cells were transfected with plasmids expressing BE3 and sgRNAs of different 5'-shortened lengths. Treated cells were analyzed as described in the Examples. Figure 95A shows the protospacer and PAM sequence (top, SEQ ID NO: 715) at a site within the EMX1 genomic locus, as well as the cellular C to T conversion percentage, determined as the percentage of total DNA sequencing reads with a T at the indicated target position. At this site, the base editing window was altered by using a 17-nt shortened gRNA. [Figure 95B] Figure 95B shows the protospacer and PAM sequences (top, SEQ ID NOs: 715 and 716) at sites within the HEK site 3 and site 4 genomic loci, as well as the cellular C to T conversion percentage, determined as the percentage of total DNA sequencing reads with a T at the indicated target position. At these sites, no change in the base editing window was observed, but a linear decrease in editing efficiency of all substrate bases was noted as the sgRNA was shortened.

[0138] [Figure 96] Figure 96 shows the effect of APOBEC1-Cas9 linker length on base editing window width. HEK293T cells were transfected with plasmids expressing base editors and sgRNAs with the rAPOBEC1-Cas9 linkers XTEN, GGS, (GGS)3 (SEQ ID NO: 610), (GGS)5 (SEQ ID NO: 610), or (GGS)7 (SEQ ID NO: 610). Treated cells were analyzed as described in the Examples. Cellular C to T conversion percentages, defined as the percentage of total DNA sequencing reads with T at the indicated target position, are shown for various base editors with different linkers.

[0139] [Fig. 97A-B] Figures 97A-97C show the effect of rAPOBEC mutations on base editing window width. [Figure 97C-1] Figure 97C shows HEK293T cells transfected with sgRNA targeting either site A or site B and a plasmid expressing the indicated BE3 point mutant. Treated cells were analyzed as described in the Examples. All Cs within the protospacer and within 3 base pairs of the protospacer are labeled, and the cellular C→T conversion percentage is shown. The "editing window width," defined as the calculated number of nucleotides above half-maximal editing efficiency, is displayed for all tested mutants. [Figure 97C-2] Figure 97C shows HEK293T cells transfected with sgRNA targeting either site A or site B and a plasmid expressing the indicated BE3 point mutant. Treated cells were analyzed as described in the Examples. All Cs within the protospacer and within 3 base pairs of the protospacer are labeled, and the cellular C→T conversion percentage is shown. The "editing window width," defined as the calculated number of nucleotides above half-maximal editing efficiency, is displayed for all tested mutants.

[0140] [Figure 98] Figure 98 shows the effect of APOBEC1 mutations on base editing product distribution in mammalian cells. HEK293T cells were transfected with a plasmid expressing BE3 or its mutants and the appropriate sgRNA. Treated cells were analyzed as described in the Examples. Cellular C→T conversion percentage, defined as the percentage of total DNA sequencing reads with a T at the indicated target position, is shown (left). The percent of total sequencing reads containing a C→T conversion is shown on the right. BE3 point mutations do not significantly affect base editing efficiency at HEK site 4, a site with only one target cytidine.

[0141] [Figure 99] Figure 99 shows a comparison of on-target editing plasmid delivery in BE3 and HF-BE3.

[0142] [Figure 100] Figure 100 shows a comparison of on-target editing in protein and plasmid delivery of BE3.

[0143] [Figure 101] Figure 101 shows a comparison of on-target editing in protein and plasmid delivery of HF-BE3.

[0144] [Figure 102] Figure 102 shows that both lipofection and installing the HF mutation reduce off-target deamination events. Diamonds indicate that no off-target events were detected and the specificity ratio was set to 100.

[0145] [Figure 103] Figure 103 shows in vitro C→T editing of a synthetic substrate with Cs placed at even positions within the protospacer (NNNNTC2TC4TC6TC8TC10TC12TC14TC16TC18TC20NGG, SEQ ID NO: 723).

[0146] [Figure 104] Figure 104 shows in vitro C→T editing of a synthetic substrate with Cs placed at odd positions within the protospacer (NNNNTC2TC4TC6TC8TC10TC12TC14TC16TC18TC20NGG, SEQ ID NO: 723).

[0147] [Figure 105] Figure 105 includes two graphs illustrating the specificity ratio of base editing by plasmid vs. protein delivery.

[0148] [Figure 106A]Figures 106A-106B show BE3 activity against non-NGG PAM sites. HEK293T cells were transfected with a plasmid expressing BE3 and the appropriate sgRNA. Treated cells were analyzed as described in the Examples. Figure 106A shows that BE3 activity against sites can be efficiently targeted by SaBE3 or SaKKH-BE3. BE3 shows low but significant activity against NAG PAM. This figure illustrates SEQ ID NOs: 728 and 729. [Figure 106B] Figure 106B shows that BE3 has significantly reduced editing at sites with NGA or NGCG PAMs, in contrast to VQR-BE3 or VRER-BE3. This figure depicts SEQ ID NOs: 730 and 731.

[0149] [Figure 107A] Figures 107A-107B show the effects of APOBEC1 mutations on VQR-BE3 and SaKKH-BE3. HEK293T cells were transfected with plasmids expressing VQR-BE3, SaKKH-BE3, or their mutants and the appropriate sgRNA. Treated cells were analyzed as described in the examples below. Cellular C→T conversion percentages, determined as the percentage of total DNA sequencing reads with T at the indicated target positions, are shown. Figure 107A shows that window modulation mutations can be applied to VQR-BE3 to enable selective base editing at sites targetable by NGA PAM. This figure illustrates SEQ ID NOs: 732 and 733. [Figure 107B] Figure 107B shows that when applied to SaKKH-BE3, the mutations cause an overall decrease in base editing efficiency without conferring base selectivity within the target window. This figure illustrates SEQ ID NOs: 728 and 734.

[0150] [Figure 108]Figure 108 shows a schematic diagram of nucleotide editing. The following abbreviations are used: (MMR) - mismatch repair, (BE3 nickase) - base editor 3, which contains the Cas9 nickase domain, (UGI) - uracil glycosylase inhibitor, (UDG) - uracil DNA glycosylase, (APOBEC) - APOBEC cytidine deaminase.

[0151] [Figure 109] Figure 109 shows a schematic diagram of exemplary base editing constructs. The structural arrangement of the base editing constructs is shown for BE3, BE4-pmCDA1, BE4-hAID, BE4-3G, BE4-N, BE4-SSB, BE4-(GGS)3, BE4-XTEN, BE4-32aa, BE4-2xUGI, and BE4. Linkers are shown in gray (XTEN, SGGS (SEQ ID NO: 606), (GGS)3 (SEQ ID NO: 610), and 32aa). Deaminases are shown (rAPOBEC1, pmCDA1, hAID, and hAPOBEC3G). Uracil-DNA glycosylase inhibitor (UGI) is shown. Single-stranded DNA-binding protein (SSB) is shown in purple. Cas9 nickase, dCas9(A840H), is shown in red. Figure 109 also shows the following target sequences: EMX1, FANCF, HEK2, HEK3, HEK4, and RNF2. The amino acid sequences are indicated from top to bottom by SEQ ID NOs: 127-132. The PAM sequence is the last three nucleotides. The target cytosine (C) is numbered and indicated in red.

[0152] [Figure 110]Figure 110 shows the base editing results of the indicated base editing constructs (BE3, pmCDA1, hAID, hAPOBEC3G, BE4-N, BE4-SSB, BE4-(GGS)3, BE-XTEN, BE4-32aa, and BE4-2xUGI) on the target cytosine (C5) of the EMX1 sequence GAGTC5CGAGCAGAAGAAGAAGGG (SEQ ID NO: 127). The total percentage of target cytosines (C5) that were mutated is indicated under "C5" for each base editing construct. The total percentage of indels is indicated under "Indels" for each base editing construct. The proportion of mutated cytosines that were mutated to adenine (A), guanine (G), or thymine (T) is indicated by a pie chart for each base editing construct.

[0153] [Figure 111] Figure 111 shows the base editing results of the indicated base editing constructs (BE3, pmCDA1, hAID, hAPOBEC3G, BE4-N, BE4-SSB, BE4-(GGS)3, BE-XTEN, BE4-32aa, and BE4-2xUGI) on the target cytosine (C8) of the FANCF sequence GGAATCCCC8TTCTGCAGCACCTGG (SEQ ID NO: 128). The total percentage of target cytosines (C8) that were mutated is indicated under "C8" for each base editing construct. The total percentage of indels is indicated under "Indels" for each base editing construct. The proportion of mutated cytosines that were mutated to adenine (A), guanine (G), or thymine (T) is indicated by a pie chart for each base editing construct.

[0154] [Figure 112]Figure 112 shows the base editing results of the indicated base editing constructs (BE3, pmCDA1, hAID, hAPOBEC3G, BE4-N, BE4-SSB, BE4-(GGS)3, BE-XTEN, BE4-32aa, and BE4-2xUGI) on the target cytosine (C6) of the HEK2 sequence GAACAC6AAAGCATAGACTGCGGG (SEQ ID NO: 129). The total percentage of target cytosines (C6) that were mutated is indicated under "C6" for each base editing construct. The total percentage of indels is indicated under "Indels" for each base editing construct. The proportion of mutated cytosines that were mutated to adenine (A), guanine (G), or thymine (T) is indicated by a pie chart for each base editing construct.

[0155] [Figure 113] Figure 113 shows the base editing results of the indicated base editing constructs (BE3, pmCDA1, hAID, hAPOBEC3G, BE4-N, BE4-SSB, BE4-(GGS)3, BE-XTEN, BE4-32aa, and BE4-2xUGI) on the target cytosine (C5) of the HEK3 sequence GGCCC5AGACTGAGCACGTGATGG (SEQ ID NO: 130). The total percentage of target cytosines (C5) that were mutated is indicated under "C5" for each base editing construct. The total percentage of indels is indicated under "Indels" for each base editing construct. The proportion of mutated cytosines that were mutated to adenine (A), guanine (G), or thymine (T) is indicated by a pie chart for each base editing construct.

[0156] [Figure 114]Figure 114 shows the base editing results of the indicated base editing constructs (BE3, pmCDA1, hAID, hAPOBEC3G, BE4-N, BE4-SSB, BE4-(GGS)3, BE-XTEN, BE4-32aa, and BE4-2xUGI) on the target cytosine (C5) of the HEK4 sequence GGCAC5TGCGGCTGGAGGTCCGGG (SEQ ID NO: 131). The total percentage of target cytosines (C5) that were mutated is indicated under "C5" for each base editing construct. The total percentage of indels is indicated under "Indels" for each base editing construct. The proportion of mutated cytosines that were mutated to adenine (A), guanine (G), or thymine (T) is indicated by a pie chart for each base editing construct.

[0157] [Figure 115] Figure 115 shows the base editing results of the indicated base editing constructs (BE3, pmCDA1, hAID, hAPOBEC3G, BE4-N, BE4-SSB, BE4-(GGS)3, BE-XTEN, BE4-32aa, and BE4-2xUGI) on the target cytosine (C6) of the RNF2 sequence GTCATC6TTAGTCATTACCTGAGG (SEQ ID NO: 132). The total percentage of target cytosines (C6) that were mutated is indicated under "C6" for each base editing construct. The total percentage of indels is indicated under "Indels" for each base editing construct. The proportion of mutated cytosines that were mutated to adenine (A), guanine (G), or thymine (T) is indicated by a pie chart for each base editing construct.

[0158] [Figure 116] Figure 116 shows exemplary fluorescently labeled (Cy3-labeled) DNA constructs used to test for Cpf1 mutants that nick the target strand. In DNA construct 1, both the non-target strand (top strand) and the target strand (bottom strand) are fluorescently labeled. In DNA construct 2, the non-target strand (top strand) is fluorescently labeled, and the target strand (bottom strand) is not. In DNA construct 3, the non-target strand (top strand) is not fluorescently labeled, and the target strand (bottom strand) is fluorescently labeled.

[0159] [Figure 117] Figure 117 shows data demonstrating the ability of various Cpf1 constructs (e.g., R836A, R1138A, wild type) to cleave the target and non-target strands of the DNA construct shown in Figure 116 over reaction times of either 30 minutes (30 min) or greater than 2 hours (2 h+).

[0160] [Figure 118] Figure 118 shows data demonstrating that base editors with the architecture APOBEC-AsCpf1(R912A)-UGI can edit C residues (e.g., in target sequences FANCF1, FANCF2, HEK3-3, and HEK3-4) with a window from base 7 to base 11 of the target sequence. BG indicates background mutation levels (untreated). AsCpf1 indicates treatment with AsCpf1 only (control), APOBEC-AsCpf1(R912A)-UGI indicates a base editor containing Cpf1 that preferentially cleaves the target strand, and APOBEC-AsCpf1(R1225A)-UGI indicates a suicidal base editor containing Cpf1 that cleaves the non-target strand. The target sequences for FANCF1, FANCF2, HEK3-3, and HEK3-4 are as follows: FANCF1 GCGGATGTTCCAATCAGTACGCA (SEQ ID NO: 724) FANCF2 CGAGCTTCTGGCGGTCTCAAGCA (SEQ ID NO: 725) HEK3-3 TGCTTCTCCAGCCCTGGCCTGG (SEQ ID NO: 726) HEK3-4 AGACTGAGCACGTGATGGCAGAG (SEQ ID NO: 727)

[0161] [Figures 119-121]Figure 119 shows a schematic diagram of a base editor comprising a Cpf1 protein (e.g., AsCpf1 or LbCpf1). Different linker sequences (e.g., XTEN, GGS, (GGS)3 (SEQ ID NO: 610), (GGS)5 (SEQ ID NO: 610), and (GGS)7 (SEQ ID NO: 610)) were tested as the moiety labeled "linker." The results are shown in Figure 120.

[0162] Figure 120 shows data demonstrating the ability of the construct shown in Figure 119 to edit the C8 residue of the HEK3 site TGCTTCTC8CAGCCCTGGCCTGG (SEQ ID NO: 592). Different linker sequences linking the APOBEC domain to the Cpf1 domain (e.g., LbCpf1(R836A) or AsCpf1(R912A)) were tested. Exemplary linkers tested include XTEN, GGS, (GGS)3 (SEQ ID NO: 610), (GGS)5 (SEQ ID NO: 610), and (GGS)7 (SEQ ID NO: 610).

[0163] Figure 121 shows data demonstrating the ability of the construct shown in Figure 119 with the LbCpf1 domain to edit the C8 and C9 residues of TGCTTCTC8C9AGCCCTGGCCTGG (SEQ ID NO: 592) in HEK3. Different linker sequences from the database maintained by the Centre of Integrative Bioinformatics VU that link the APOBEC domain to the LbCpf1 domain were tested. Exemplary linkers tested include 1au7, 1clk, 1c20, 1ee8, 1flz, 1ign, 1jmc, 1sfe, 2ezx, and 2reb.

[0164] [Figure 122-123] Figure 122 shows a schematic diagram of the structure of AsCpf1, where the N- and C-termini are indicated.

[0165] Figure 123 shows a schematic diagram of the structure of SpCas9, with the N- and C-termini indicated.

[0166] [Figure 124] Figure 124 shows a schematic diagram of AsCpf1, where the red circle indicates the predicted area where the editing window is located, and the box indicates the helical region where one might attempt to interfere with APOBEC activity.

[0167] [Figure 125] Figures 125A and 125B show the engineering and in vitro characterization of a high-fidelity base editor (HF-BE3). Figure 125A shows a schematic of HF-BE3. The point mutations introduced into BE3 to create HF-BE3 are shown. The figure uses PDB structures 4UN3 (Cas9), 4ROV (cytidine deaminase), and 1UGI (uracil DNA glycosylase inhibitor). Figure 125B shows the in vitro deamination of a synthetic substrate containing a "TC" repeat protospacer. Values ​​and error bars reflect the mean and range of two independent replicates performed on different days.

[0168] [Fig. 126A-B] Figures 126A-126C show the purification of base editor proteins. Figure 126A shows the selection of optimal E. coli strains for base editor expression. After 16 h of IPTG-induced protein expression at 18 °C, crude cell lysates were analyzed for protein content. BL21 Star (DE3) (Thermo Fisher) cells showed the most promising post-expression levels for both BE3 and HF-BE3 and were used for base editor expression. Figure 126B shows the purification of expressed base editor proteins. Placing a His6 tag at the C-terminus of the base editor leads to the generation of truncated products for both BE3 and HF-BE3 (lanes 1 and 2). Unexpectedly, this truncated product was removed by placing a His6 tag at the N-terminus of the protein (lanes 3-6). Inducing base editor expression at a cell density of OD600 = 0.7 (lanes 4-5), which is later than the optimal cell density for Cas9 expression (OD600 = 0.4), improves the yield of base editor protein. Purification was performed using a manual HisPur resin column followed by a cation-exchange FPLC (Akta). [Figure 126C] Figure 126C shows purified BE3 and HF-BE3. Different concentrations of purified BE3 and HF-BE3 were denatured using heat and LDS and loaded onto a polyacrylamide gel. Protein samples are representative of the proteins used in this study. The gels in Figures 126A-126C are BOLT Bis-Tris Plus 4-12% polyacrylamide (Thermo Fisher). Electrophoresis and staining were performed as described in the Methods.

[0169] [Figure 127A] Figures 127A-127D show the activity of a high-fidelity base editor (HF-BE3) in human cells. Figures 127A-127C show that on- and off-target editing associated with plasmid transfection of BE3 and HF-BE3 was assayed using high-throughput sequencing of genomic DNA from HEK293T cells treated with sgRNAs targeting the non-repetitive genomic loci EMX1 (Figure 127A), FANCF (Figure 127B), and HEK293 site 3 (Figure 127C). The on- and off-target loci associated with each sgRNA are separated by vertical lines. [Fig. 127B-C] Figures 127A-127C show that on- and off-target editing associated with plasmid transfection of BE3 and HF-BE3 was assayed using high-throughput sequencing of genomic DNA from HEK293T cells treated with sgRNAs targeting the non-repetitive genomic loci EMX1 (Figure 127A), FANCF (Figure 127B), and HEK293 site 3 (Figure 127C). The on- and off-target loci associated with each sgRNA are separated by vertical lines. [Figure 127D]Figure 127D shows on- and off-target editing associated with a highly repetitive sgRNA targeting VEGFA site 2. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days. In Figures 127A-127C, asterisks indicate significant editing based on a comparison between treated samples and untreated controls. *p≦0.05, **p≦0.01, and ***p≦0.001 (two-tailed Student's t-test). In Figure 127D, asterisks are not shown because all treated samples displayed significant editing relative to the control. Individual p-values ​​are listed in Table 16.

[0170] [Figure 128A] Figures 128A-128C show the effect of BE3 protein or plasmid dose on the efficiency of on-target and off-target base editing in cultured human cells. Figure 128A shows the on-target editing efficiency at each of four genomic loci, averaged for all edited cytosines within the activity window for each sgRNA. Values ​​and error bars reflect the mean ± SEM of three independent biological replicates performed on different days. [Fig. 128B-C] Figures 128B and 128C show on- and off-target editing at the EMX1 site resulting from BE3 plasmid titration (Figure 128B) or BE3 protein titration (Figure 128C) in HEK293T cells. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0171] [Figure 129]Figures 129A-129B show on-target:off-target base editing frequency ratios for BE3 and HF-BE3 plasmid and protein delivery. Base editing on-target:off-target specificity ratios were calculated by dividing the on-target editing percentage of a specific cytosine within the activity window by the off-target editing percentage of the corresponding cytosine at the indicated off-target locus (see Methods). When off-target editing was below the detection threshold (0.025% of sequencing reads), we set off-target editing to the limit of detection (0.025%) and divided the on-target editing percentage by this upper limit. In these cases, indicated by a ◆, the specificity ratio shown corresponds to the lower limit. Specificity ratios are shown for the non-repetitive sgRNAs FANCF, HEK293 site 3, and FANCF (Figure 129A) and the highly repetitive sgRNA VEGFA site 2 (Figure 129B). Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0172] [Fig. 130A-B] Figures 130A-130D show protein delivery of base editors into cultured human cells. Figures 130A-130D show on- and off-target editing associated with RNP delivery of base editors complexed with sgRNAs targeting EMX1 (Figure 130A), FANCF (Figure 130B), HEK293 site 3 (Figure 130C), and VEGFA site 2 (Figure 130D). Off-target base editing was undetectable at all sequenced loci with non-repetitive sgRNAs. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days. Asterisks indicate significant editing based on comparison between treated samples and untreated controls. *p≦0.05, **p≦0.01, and ***p≦0.001 (two-tailed Student's t-test). [Fig. 130C-D]Figures 130A-130D show on- and off-target editing associated with RNP delivery of base editors complexed with sgRNAs targeting EMX1 (Figure 130A), FANCF (Figure 130B), HEK293 site 3 (Figure 130C), and VEGFA site 2 (Figure 130D). Off-target base editing was undetectable with non-repetitive sgRNAs at all of the sequenced loci. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days. Asterisks indicate significant editing based on comparison between treated samples and untreated controls. *p≦0.05, **p≦0.01, and ***p≦0.001 (two-tailed Student's t-test).

[0173] [Fig. 131A-B] Figures 131A-131C show indel formation associated with base editing at genomic loci. Figure 131A shows indel frequencies at on-target loci for VEGFA site 2, EMX1, FANCF, and HEK293 site 3 sgRNAs. Figure 131B shows the ratio of base editing to indel formation. Diamonds (◆) indicate no indels were detected (no significant difference in indel frequency between treated samples and untreated controls). [Figure 131C] Figure 131C shows indels observed at off-target loci relative to the on-target sites interrogated in Figure 131A. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0174] [Figure 132A]Figures 132A-132D show DNA-free in vivo base editing in the inner ear of zebrafish embryos and adult mice using RNP delivery of BE3. Figure 132A shows on-target genome editing in zebrafish harvested 4 days after injection of BE3 complexed with the indicated sgRNA. Values ​​and error bars reflect the mean ± SD of three injections and three control zebrafish. Controls were injected with BE3 complexed with an unrelated sgRNA. [Fig. 132B-C] Figure 132B shows a schematic diagram illustrating in vivo injection of BE3:sgRNA complexes encapsulated in cationic lipid nanoparticles. Figure 132C shows base editing of cytosine residues within the base editor window of the VEGFA site 2 genomic locus. [Figure 132D] Figure 132D shows on-target edits at each cytosine within the base editing window of the VEGFA site 2 target locus. Figure 132D (Figures 132C and 132D) shows that values ​​and error bars reflect the mean ± SEM of three mice injected with an sgRNA targeting VEGFA site 2, three uninjected mice, and one mouse injected with an unrelated sgRNA.

[0175] [Figure 133A] Figures 133A-133E show on- and off-target base editing in mouse NIH / 3T3 cells. Figure 133A shows on-target base editing associated with the "VEGFA site 2" sgRNA (see Figure 132E for sequence). The negative control corresponds to cells treated with a plasmid encoding BE3 without sgRNA. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days. [Fig. 133B-C]Figures 133B to 133E show that off-target editing associated with this site was measured using high-throughput DNA sequencing at the top four predicted off-target loci for this sgRNA (sequences are shown in Figure 132E). Figure 133B shows off-target 2, Figure 133C shows off-target 1, Figure 133D shows off-target 3, and Figure 133E shows off-target 4. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days. [Fig. 133D-E] Figure 133D shows off-target 3, and Figure 133E shows off-target 4. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0176] [Figure 134] Figures 134A-134B show off-target base editing and on-target indel analysis from in vivo edited mouse tissues. Figure 134A shows edits plotted for each cytosine within the base editing window of an off-target locus relative to VEGFA site 2. Figure 134B shows the indel rate at the on-target base editor locus. Values ​​and error bars reflect the mean ± SEM of three injected and three control mice.

[0177] [Figure 135]Figures 135A-135C show the effect of knocking out UNG on base editing product purity. Figure 135A shows HAP1 (UNG+) and HAP1 UNG- cells treated with BE3 as described in Materials and Methods in Example 17. Product distribution of edited DNA sequencing reads (reads in which the target C is mutated) is shown. Figure 135B shows the protospacer and PAM sequences of the genomic loci tested, with the target C analyzed in Figure 135A shown in red. Figure 135C shows the frequency of indel formation after treatment with BE3 in HAP1 or HAP1 UNG- cells. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0178] [Figure 136A] Figures 136A-136D show the effect of multiple C base editing on product purity. Figure 136A shows representative high-throughput sequencing data for untreated, BE3-treated, and AID-BE3-treated human HEK293T cells. The sequence of the protospacer is shown above, with the PAM and target C in red and subscripted numbers indicating their position within the protospacer. Directly below each sequence is the percentage of total sequencing reads with the corresponding base. The relative percentage of target Cs cleanly edited to T rather than non-T bases is much higher in cells treated with AID-BE3, which edits three Cs at this locus, than in cells treated with BE3, which edits only one C. [Fig. 136B-D]Figure 136B shows HEK293T cells treated with BE3, CDA1-BE3, and AID-BE3 as described in Materials and Methods in Example 17. Product distribution of edited DNA sequencing reads (reads in which target C is mutated) is shown. Figure 136C shows the protospacer and PAM sequences of the studied genomic loci, with target C analyzed in Figure 136B shown in red. Figure 136D shows the frequency of indel formation after the treatments shown in Figure 136A. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0179] [Fig. 137A-B] Figures 137A-137C show the effect of varying the architecture of BE3 on C→T editing efficiency and product purity. Figure 137A shows the protospacer and PAM sequences of the studied genomic loci, with target Cs in Figure 137C shown in purple and red, and target Cs in Figure 137B shown in red. Figure 137B shows HEK293T cells treated with BE3, SSB-BE3, N-UGI-BE3, and BE3-2xUGI as described in Materials and Methods in Example 17. Product distributions of edited DNA sequencing reads (reads in which target Cs are mutated) are shown for BE3, N-UGI-BE3, and BE3-2xUGI. [Figure 137C] Figure 137C shows C→T base editing efficiency. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0180] [Fig. 138A-C]Figures 138A-138D show the effect of varying the linker length of BE3 on C→T editing efficiency and product purity. Figure 138A shows the architectures of BE3, BE3C, BE3D, and BE3E. Figure 138B shows the protospacer and PAM sequences of the studied genomic loci, with target C in Figure 138C shown in purple and red, and target C in Figure 138D shown in red. Figure 138C shows HEK293T cells treated with BE3, BE3C, BE3D, or BE3E as described in Materials and Methods in Example 17. C→T base editing efficiency is shown. [Figure 138D] Figure 138D shows the product distribution of edited DNA sequencing reads (reads in which target C is mutated) for BE3, BE3C, BE3D, and BE3E. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0181] [Fig. 139A-C] Figures 139A-139D show that BE4 increases base editing efficiency and product purity compared to BE3. Figure 139A shows the architectures of BE3, BE4, and Target-AID. Figure 139B shows the protospacer and PAM sequences of the studied genomic loci, with target C in Figure 139C shown in purple and red, and target C in Figure 139D shown in red. Figure 139C shows HEK293T cells treated with BE3, BE4, or Target-AID as described in Materials and Methods in Example 17. C→T base editing efficiency is shown. [Figure 139D] Figure 139D shows the product distribution of edited DNA sequencing reads (reads in which target C is mutated) for BE3 and BE4. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0182] [Figure 140]Figures 140A-140C show that CDA1-BE3 and AID-BE3 edit Cs following targeted Gs more efficiently than BE3. Figure 140A shows the protospacer and PAM sequences of the genomic loci studied, with targeted Cs edited by BE3, CDA1-BE3, and AID-BE3 shown in red, and targeted Cs (following Gs) edited only by CDA1-BE3 and AID-BE3 shown in purple. Figure 140B shows HEK293T cells treated with BE3, CDA1-BE3, AID-BE3, or APOBEC3G-BE3 as described in Materials and Methods in Example 17. C->T base editing efficiencies are shown. Figure 140C shows individual DNA sequencing reads binned and analyzed according to the sequence of the protospacer from HEK293T cells treated with BE3, CDA1-BE3, or AID-BE3 targeting the HEK2 locus, revealing that >85% of sequencing reads with clean C→T edits by CDA1-BE3 and AID-BE3 have both Cs edited to Ts (Figure 140C).

[0183] [Fig. 141A-B] Figures 141A-141C show that uneven editing at sites with multiple editable Cs results in lower product purity. Figure 141A shows the protospacer and PAM sequences of the studied genomic loci, with the target Cs in Figure 141C shown in purple and red, and the target C in Figure 141B shown in red. Figures 141B and 141C show HEK293T cells treated with BE3 as described in Materials and Methods in Example 17. Product distribution of edited DNA sequencing reads (reads in which the target C is mutated) is shown. When editing efficiencies are unequal for two Cs within the same locus, C-to-non-T editing is more frequent. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days. [Figure 141C]Figures 141B and 141C show HEK293T cells treated with BE3 as described in Materials and Methods in Example 17. Product distribution of edited DNA sequencing reads (reads in which the target C is mutated) is shown. When editing efficiencies are unequal for two Cs within the same locus, C-to-non-T editing is more frequent. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0184] [Fig. 142A-B] Figures 142A-142D show that base editing of multiple Cs results in higher base-edited product purity. Figure 142A shows the protospacer and PAM sequences of the studied genomic loci, and the investigated target Cs are shown in red in Figure 142B. Figure 142B shows HEK293T cells treated with BE3 or BE3B (which lacks UGI) as described in Materials and Methods in Example 17. Product distribution of edited DNA sequencing reads (reads in which the target Cs are mutated) is shown. [Fig. 142C-D] Figure 142C shows HTS reads from HEK293T cells targeting the HEK2 locus and treated with BE3 or BE3B (which lack UGI) were binned according to the identity of the primary target C at position 6. The resulting reads were then analyzed for the base identity of the secondary target C at position 4. When there is only one editing event in the read, C6 is more likely to be incorrectly edited to a non-T. Figure 142D shows the distribution of edited reads with A, G, and T at C5 in cells targeting the HEK4 locus (a site with only one editable C) and treated with BE3 or BE3B, illustrating that a single G:U mismatch is processed by UNG-initiated base excision repair to give a mixture of products. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0185] [Figure 143]Figure 143 shows that base editing of multiple Cs results in higher base-edited product purity at the HEK3 and RNF2 loci. DNA sequencing reads from HEK293T cells targeted to the HEK3 and RNF2 loci and treated with BE3 or BE3B (without UGI) were separated according to the base identity of the primary target C position (red). Four groups of sequencing reads were then interrogated for the base identity of the secondary target C position (purple). In BE3, when the primary target C (red) is incorrectly edited to G, the secondary target C is more likely to remain C. Conversely, when the primary target C (red) is converted to T, the secondary target C is also likely to be edited to T in the same sequencing read. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0186] [Figure 144] Figures 144A-144C show that BE4 induces a lower indel frequency than BE3, while Target-AID exhibits similar product purity to CDA1-BE3. Figure 144A shows HEK293T cells treated with BE3, BE4, or Target-AID as described in Materials and Methods in Example 17. The frequency of indel formation (see Materials and Methods in Example 17) is shown. Figure 144B shows HEK293T cells treated with CDA1-BE3 or Target-AID as described in Materials and Methods in Example 17. Product distribution of edited DNA sequencing reads (reads in which the target C is mutated) is shown. Figure 144C shows the protospacer and PAM sequences of the studied genomic loci; the target C examined in Figure 144B is shown in red. Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0187] [Figure 145]Figures 145A-145C show that SaBE4 exhibits increased base editing yield and product purity compared to SaBE3. Figure 145A shows HEK293T cells treated with SaBE3 and SaBE4 as described in Materials and Methods in Example 17. The percentage of total DNA sequencing reads with T at the indicated target positions is shown. Figure 145B shows the protospacer and PAM sequences of the studied genomic loci, with target C in Figure 145A shown in purple and red, and target C examined in Figure 145C shown in red. Figure 145C shows the product distribution of edited DNA sequencing reads (reads in which target C is mutated). Values ​​and error bars reflect the mean ± SD of three independent biological replicates performed on different days.

[0188] [Figure 146] Figure 146 shows base editing results from treatment with BE3, CDA1-BE3, AID-BE3, or APOBEC3G-BE3 at the EMX1 locus. The sequence of the protospacer is shown above, with the PAM and target bases in red and subscripted numbers indicating their position within the protospacer. Directly below the sequence is the percentage of total sequencing reads with the corresponding base. Cells were treated as described in Materials and Methods in Example 17. Values ​​shown are from one representative experiment.

[0189] [Figure 147] Figure 147 shows base editing results from treatment with BE3, CDA1-BE3, AID-BE3, or APOBEC3G-BE3 at the FANCF locus. The sequence of the protospacer is shown above, with the PAM and target bases in red and subscripted numbers indicating their position within the protospacer. Directly below the sequence is the percentage of total sequencing reads with the corresponding base. Cells were treated as described in Materials and Methods in Example 17. Values ​​shown are from one representative experiment.

[0190] [Figure 148]Figure 148 shows base editing results from treatment with BE3, CDA1-BE3, AID-BE3, or APOBEC3G-BE3 at the HEK2 locus. The sequence of the protospacer is shown above, with the PAM and target bases in red and subscripted numbers indicating their position within the protospacer. Immediately below the sequence is the percentage of total sequencing reads with the corresponding base. Cells were treated as described in Materials and Methods in Example 17. Values ​​shown are from one representative experiment.

[0191] [Figure 149] Figure 149 shows base editing results from treatment with BE3, CDA1-BE3, AID-BE3, or APOBEC3G-BE3 at the HEK3 locus. The sequence of the protospacer is shown above, with the PAM and target bases in red and subscripted numbers indicating their position within the protospacer. Immediately below the sequence is the percentage of total sequencing reads with the corresponding base. Cells were treated as described in Materials and Methods in Example 17. Values ​​shown are from one representative experiment.

[0192] [Figure 150] Figure 150 shows base editing results from treatment with BE3, CDA1-BE3, AID-BE3, or APOBEC3G-BE3 at the HEK4 locus. The sequence of the protospacer is shown above, with the PAM and target bases in red and subscripted numbers indicating their position within the protospacer. Immediately below the sequence is the percentage of total sequencing reads with the corresponding base. Cells were treated as described in Materials and Methods in Example 17. Values ​​shown are from one representative experiment.

[0193] [Figure 151]Figure 151 shows base editing results from treatment with BE3, CDA1-BE3, AID-BE3, or APOBEC3G-BE3 at the RNF2 locus. The sequence of the protospacer is shown above, with the PAM and target bases in red and subscripted numbers indicating their position within the protospacer. Immediately below the sequence is the percentage of total sequencing reads with the corresponding base. Cells were treated as described in Materials and Methods in Example 17. Values ​​shown are from one representative experiment.

[0194] [Figure 152] Figure 152 shows a schematic diagram of the LBCpf1 fusion constructs. Construct 10 has the domain arrangement of [Apobec]-[LbCpf1]-[UGI]-[UGI]; construct 11 has the domain arrangement of [Apobec]-[LbCpf1]-[UGI]; construct 12 has the domain arrangement of [UGI]-[Apobec]-[LbCpf1]; construct 13 has the domain arrangement of [Apobec]-[UGI]-[LbCpf1]; construct 14 has the domain arrangement of [LbCpf1]-[UGI]-[Apobec]; and construct 15 has the domain arrangement of [LbCpf1]-[Apobec]-[UGI]. For each construct, three different LbCpf1 proteins were used (D / N / A, which refers to nuclease-dead LbCpf1 (D); LbCpf1 nickase (N); and nuclease-active LbCpf1 (A)).

[0195] [Figure 153] Figure 153 shows the percentage of C to T editing of the six C residues of the EMX target TTTGTAC3TTTGTC9C10TC12C13GGTTC18TG (SEQ ID NO: 738) using a 19 nucleotide long guide, EMX19:TACTTTGTCCTCCGGTTCT (SEQ ID NO: 744). Editing was tested for several of the constructs shown in Figure 152.

[0196] [Fig. 154]Figure 154 shows the percentage of C to T editing of the six C residues of the EMX target TTTGTAC3TTTGTC9C10TC12C13GGTTC18TG (SEQ ID NO: 738) using the 18 nucleotide long guide, EMX18:TACTTTGTCCTCCGGTTC (SEQ ID NO: 745). Editing was tested for some of the constructs shown in Figure 152.

[0197] [Figure 155] Figure 155 shows the percentage of C to T editing of the six C residues of the EMX target TTTGTAC3TTTGTC9C10TC12C13GGTTC18TG (SEQ ID NO: 738) using the 17 nucleotide long guide, EMX17:TACTTTGTCCTCCGGTT (SEQ ID NO: 746). Editing was tested for some of the constructs shown in Figure 152.

[0198] [Figure 156] Figure 156 shows the percentage of C to T editing of the eight C residues of the HEK2 target TTTCC1AGC4C5C6GC8TGGC12C13C14TGTAAA (SEQ ID NO: 739) using the 23 nucleotide long guide, Hek2_23:CAGCCCGCTGGCCCTGTAAAGGA (SEQ ID NO: 747). Editing was tested for several of the constructs shown in Figure 152.

[0199] [Figure 157] Figure 157 shows the percentage of C to T editing of the eight C residues of the HEK2 target TTTCC1AGC4C5C6GC8TGGC12C13C14TGTAAA (SEQ ID NO: 739) using the 20 nucleotide long guide, Hek2_20:CAGCCCGCTGGCCCTGTAAA (SEQ ID NO: 748). Editing was tested for several of the constructs shown in Figure 152.

[0200] [Figure 158]Figure 158 shows the percentage of C to T editing of the eight C residues of the HEK2 target TTTCC1AGC4C5C6GC8TGGC12C13C14TGTAAA (SEQ ID NO: 739) using the 19 nucleotide long guide, Hek2_19:CAGCCCGCTGGCCCTGTAA (SEQ ID NO: 749). Editing was tested for several of the constructs shown in Figure 152.

[0201] [Figure 159] Figure 159 shows the percentage of C to T editing of the eight C residues of the HEK2 target TTTCC1AGC4C5C6GC8TGGC12C13C14TGTAAA (SEQ ID NO: 739) using the 18 nucleotide long guide, Hek2_18:CAGCCCGCTGGCCCTGTA (SEQ ID NO: 750). Editing was tested for several of the constructs shown in Figure 152.

[0202] [Figure 160] Figure 160 shows the editing percentage values ​​(adjusted based on indel counts) and percentage of indels for the experiment illustrated in Figure 153.

[0203] [Figure 161] Figure 161 shows the edit percentage values ​​(adjusted based on indel counts) and percentage of indels for the experiment illustrated in Figure 154.

[0204] [Figure 162] Figure 162 shows the edit percentage values ​​(adjusted based on indel counts) and percentage of indels for the experiment illustrated in Figure 155.

[0205] [Figure 163] Figure 163 shows the edit percentage values ​​(adjusted based on indel counts) and percentage of indels for the experiment illustrated in Figure 156.

[0206] [Fig. 164]Figure 164 shows the edit percentage values ​​(adjusted based on indel counts) and percentage of indels for the experiment illustrated in Figure 157.

[0207] [Figure 165] Figure 165 shows the edit percentage values ​​(adjusted based on indel counts) and percentage of indels for the experiment illustrated in Figure 158.

[0208] [Figure 166] Figure 166 shows the edit percentage values ​​(adjusted based on indel counts) and percentage of indels for the experiment illustrated in Figure 159.

[0209] definition As used in the specification and claims, the singular forms "a," "an," and "the" include singular and plural references unless the context clearly dictates otherwise. Thus, for example, reference to an "agent" includes a single agent as well as a plurality of such agents.

[0210] The term "nucleic acid-programmable DNA-binding protein" or "napDNAbp" refers to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid (e.g., gRNA), which guides the napDNAbp to a specific nucleic acid sequence, e.g., by hybridizing to a target nucleic acid sequence. For example, a Cas9 protein can associate with a guide RNA, which guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the napDNAbp is a class 2 bacterial CRISPR-Cas effector. In some embodiments, the napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of nucleic acid-programmable DNA-binding proteins include, without limitation, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2c3, and Argonaute. However, it should be understood that nucleic acid-programmable DNA binding proteins also encompass nucleic acid-programmable proteins that bind to RNA. For example, napDNAbp can be linked to a nucleic acid that guides the napDNAbp to RNA. Other nucleic acid-programmable DNA binding proteins are also within the scope of this disclosure, although they may not be specifically described in this disclosure.

[0211] In some embodiments, napDNAbp is an "RNA-programmable nuclease" or "RNA-guided nuclease." The terms are used interchangeably herein to refer to a nuclease that forms a complex with (e.g., binds to or associates with) one or more RNA(s) that are not the target for cleavage. In some embodiments, when an RNA-programmable nuclease is complexed with an RNA, it can be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) is referred to as a guide RNA (gRNA). A gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), although "gRNA" is also used to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology to the target nucleic acid (i.e., directs binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. In some embodiments, domain (2) is identical to or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in US Provisional Patent Application USSN 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases And Uses Thereof," and US Provisional Patent Application USSN 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases," the entire contents of each of which are incorporated herein by reference.In some embodiments, the gRNA comprises two or more domains (1) and (2) and may be referred to as an "extended gRNA." For example, an extended gRNA will bind to two or more Cas9 proteins and bind to a target nucleic acid at two or more separate regions as described herein. The gRNA comprises a nucleotide sequence complementary to a target site, which mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex. In some embodiments, the RNA-programmable nuclease is a CRISPR-associated system. C RISPR- a ssociated system))Cas9 and Cas9(Csn1) of Streptococcus pyogenes "Complete genome sequence of an M1 strain of." Streptococcus pyogenes." Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G, Lyon K, Primeaux C, Sezate S, Suvorov AN, Kenton S, Lai HS, Lin SP, Qian Y, Jia HG, Najar FZ, Ren Q, Zhu H, Song L, White J, Yuan X, Clifton SW, Roe BA, McLaughlin RE, Proc Natl Sci. USA 98:4658–4663 (2001); J., Charpentier E., Nature 471:602–607 (2011); (2012) reported that the thermodynamic properties of these specimens were determined by the thermodynamic range.

[0212] In this case, the sgRNA was labeled with napDNAb A sgRNA scaffold with a sgRNA scaffold generated by p., Jinek M, Chylinski K, Fonfara I, Hauer M, Doudna JA, and Charpentier E (2012) A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science, 337, 816-812; View Article PubMed Google Scholar Mali P, Esvelt KM, Church GM (2013) Cas9 as a versatile tool for engineering biology. Nature Methods, 10, 957-963; Li JF, Norville JE, Aach J, McCromack M, Zhang D, Bush J, Church GM, and Sheen J (2013) Multiplex and homologous recombination-mediated genome editing in Arabidopsis and Nicotiana benthamiana using guide RNA and Cas9. Nature Biotech, 31, 688-691; Hwang WY, Fu Y, Reyon D, Maeder ML, Tsai SQ, Sander JD, Peterson RT, Yeh JRJ, Joung JK (2013) Efficient in vivo genome editing using RNA-guided nucleases. Nat Biotechnol, 31, 227-229; Cong L, Ran FA, Cox D, Lin S, Barretto R, Habib N, Hsu PD, Wu X, Jiang W, Marraffini LA, Zhang F (2013) Multiplex genome engineering using CRIPSR / Cas systems.Science, 339, 819-823; Cho SW, Kim S, Kim JM, Kim JS (2013) Targeted genome engineering in human cells with the Cas9 RNA-guided endonuclease. Nat Biotechnol, 31, 230-232; Jinek MJ, East A, Cheng A, Lin S, Ma E, Doudna J (2013) RNA-programmed genome editing in human cells. eLIFE, 2:e00471; DiCarlo JE, Norville JE, Mali P, Rios, Aach J, Church GM (2013) Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucl Acids Res, 41, 4336-4343; Briner AE, Donohoue PD, Gomaa AA, Selle K, Slorach EM, Nye CH, Haurwitz RE, Beisel CL, May AP, and Barrangou R (2014) Guide RNA functional modules direct Cas9 activity and orthogonality. Mol Cell, 56, 333-339; the contents of each of which are incorporated herein by reference. In some embodiments, any of the gRNAs (e.g., sgRNAs) provided herein includes the nucleic acid sequence of GTAATTTCTACTAAGTGTAGAT (SEQ ID NO: 741), i.e., GUAAUUUCUACUAAGUGUAGAU, where each T in SEQ ID NO: 741 is uracil (U), or the sequence GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUU-3' (SEQ ID NO: 618).

[0213] Because RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins are, in principle, capable of targeting any sequence specified by a guide RNA. Methods for using RNA-programmable nucleases, such as Cas9, for site-specific cleavage (e.g., to modify a genome) are known in the art (e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acids Research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).

[0214] The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nuclease comprising the Cas9 protein or a fragment thereof (e.g., a protein comprising an active, inactive, or partially active DNA cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9). Cas9 nuclease is sometimes also referred to as casn1 nuclease or CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to ancestral mobile elements, which target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. The Cas9 / crRNA / tracrRNA then endolytically cleaves linear or circular dsDNA targets complementary to the spacer. Target strands not complementary to the crRNA are first endolytically cleaved and then exolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require proteins and both RNAs. However, single-guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference.Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeats, helping it distinguish self from non-self. Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);"CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure. Such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA-cleavage domain, i.e., the Cas9 is a nickase.

[0215] Nuclease-inactivated Cas9 proteins can be interchangeably referred to as "dCas9" proteins (as nuclease-dead Cas9). Methods for creating Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28;152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28;152(5): 1173-83 (2013)). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology to Cas9 or fragments thereof. For example, the Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9.In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9.

[0216] In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length. In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1, SEQ ID NO: 1 (nucleotide); SEQ ID NO: 2 (amino acid)). [ka] [ka] (SEQ ID NO: 2) (single underline: HNH domain; double underline: RuvC domain)

[0217] In some embodiments, the wild-type Cas9 corresponds to or comprises SEQ ID NO:3 (nucleotide) and / or SEQ ID NO:4 (amino acid): [ka] [ka] (SEQ ID NO: 4) (single underline: HNH domain; double underline: RuvC domain)

[0218] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_002737.2, SEQ ID NO:5 (nucleotide); and Uniport Reference Sequence: Q99ZW2, SEQ ID NO:6 (amino acid). [ka] (SEQ ID NO: 6) (single underline: HNH domain; double underline: RuvC domain)

[0219] In some embodiments, Cas9 is: Corynebacterium ulcerans (NCBI Ref: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Ref: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia(NCBI Ref:NC_017861.1);Spiroplasma taiwanense(NCBI Ref:NC_021846.1);Streptococcus iniae(NCBI Ref:NC_021314.1);Belliella baltica(NCBI Ref:NC_018010.1);Psychroflexus torquis I(NCBI Ref:NC_018721.1);Streptococcus "Cas9" refers to Cas9 from S. thermophilus (NCBI Ref:YP_820832.1), Listeria innocua (NCBI Ref:NP_472073.1), Campylobacter jejuni (NCBI Ref:YP_002344900.1), or Neisseria meningitidis (NCBI Ref:YP_002342100.1), or any of the organisms listed in Example 5.

[0220] In some embodiments, the dCas9 corresponds to or comprises a Cas9 amino acid sequence with one or more mutations that partially or completely inactivate the Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises a D10A and / or H840A mutation. dCas9 (D10A and H840A): [ka] (SEQ ID NO: 7) (single underline: HNH domain; double underline: RuvC domain).

[0221] In some embodiments, the Cas9 domain contains a D10A mutation, but the residue at position 840 remains a histidine in the amino acid sequence provided by SEQ ID NO:6 or in any of the amino acid sequences provided by another Cas9 domain, e.g., in the corresponding position of any of the Cas9 proteins provided herein. Without wishing to be bound by any particular theory, the presence of the catalytic residue H840 restores Cas9 activity to cleave the unedited (e.g., undeaminated) strand containing a G opposite the target C. Restoration of H840 (e.g., from A840) does not result in cleavage of the target strand containing a C. Such Cas9 variants are capable of creating a single-stranded DNA break (nick) at a specific location based on the target sequence defined by the gRNA, leading to repair of the unedited strand and ultimately to a G→A change in the unedited strand. A schematic representation of this process is shown in Figure 108. Briefly, the C in a CG base pair can be deaminated to U by a deaminase, such as an APOBEC deaminase. Nicking the unedited strand with a G facilitates removal of the G by the mismatch repair mechanism. UGI inhibits UDG, which prevents removal of the U.

[0222] In other embodiments, dCas9 variants are provided that have mutations other than D10A and H840A, e.g., that result in a nuclease-inactivated Cas9 (dCas9). By way of example, such mutations include other amino acid substitutions at D10 and H820, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 (e.g., variants of SEQ ID NO: 6) are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to SEQ ID NO: 6. In some embodiments, variants of dCas9 (e.g., variants of SEQ ID NO: 6) are provided having amino acid sequences that are about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more, shorter or longer than SEQ ID NO: 6.

[0223] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of a Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not comprise the full-length Cas9 sequence, but only a fragment thereof. For example, in some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 fragment, which binds to crRNA and tracrRNA or sgRNA but does not comprise a functional nuclease domain, e.g., because it comprises only a truncated version of the nuclease domain or no nuclease domain at all. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those of skill in the art.

[0224] These include Cas9: Corynebacterium ulcerans(NCBI Ref:NC_015683.1; NC_017317.1);Corynebacterium diphtheria(NCBI). Ref:NC_016782.1、NC_016786.1);Spiroplasma syrphidicola(NCBI Ref:NC_021284.1);Prevotella intermedia(NCBI Ref:NC_017861.1); Ref:NC_021314.1);Belliella baltica(NCBI Ref:NC_018010.1);Psychroflexus torquis I(NCBI Ref:NC_018721.1);Streptococcus thermophilus(NCBI Ref:YP_820832.1);Listeria harmless(NCBI Ref:NP_472073.1);Campylobacter jejuni(NCBI Ref:YP_002344900.1); meningitidis (NCBI Ref:YP_002342100.1) This is Cas9.

[0225] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase, which catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase domain, which catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism, which does not occur in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase from an organism.

[0226] As used herein, the term "effective amount" refers to the amount of a biologically active agent that is sufficient to induce a desired biological response. For example, in some embodiments, an effective amount of a nuclease can refer to the amount of nuclease that is sufficient to induce cleavage of a target site specifically bound and cleaved by the nuclease. In some embodiments, an effective amount of a fusion protein provided herein, such as a fusion protein comprising a nuclease-inactive Cas9 domain and a nucleic acid editing domain (e.g., a deaminase domain), can refer to the amount of fusion protein that is sufficient to induce editing of a target site specifically bound and edited by the fusion protein. As will be understood by those skilled in the art, the effective amount of an agent, such as a fusion protein, nuclease, deaminase, recombinase, hybrid protein, protein dimer, protein (or protein dimer) and polynucleotide complex, or polynucleotide, can vary depending on various factors, such as the desired biological response, e.g., the specific allele, genome, or target site to be edited, the cell or tissue to be targeted, and the agent to be used.

[0227] As used herein, the term "linker" refers to a chemical group or molecule that connects two molecules or moieties, such as two domains of a fusion protein, such as a nuclease-inactive Cas9 domain and a nucleic acid editing domain (e.g., a deaminase domain). A linker can be, for example, an amino acid sequence, peptide, or polymer of any length and composition. In some embodiments, a linker connects the gRNA binding domain of an RNA-programmable nuclease, including a Cas9 nuclease domain, to the catalytic domain of a nucleic acid editing protein. In some embodiments, a linker connects dCas9 and a nucleic acid editing protein. Typically, a linker is located between or flanked by two groups, molecules, or other moieties, and is connected to each one via a covalent bond, thereby connecting the two. In some embodiments, a linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, a linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 1 to 100 amino acids in length, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.

[0228] As used herein, the term "mutation" refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues in a sequence. Typically, mutations are described herein by identifying the original residue, then the position of the residue in the sequence, and the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012).

[0229] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound comprising a nucleobase and an acidic moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester linkages. In some embodiments, "nucleic acid" refers to an individual nucleic acid residue (e.g., a nucleotide and / or a nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids can occur naturally, e.g., in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule may be a non-naturally occurring molecule, such as a recombinant DNA or RNA, an artificial chromosome, an artificial genome, or a fragment thereof, or a synthetic DNA, RNA, or DNA / RNA hybrid, or may include non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms encompass nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced and optionally purified using recombinant expression systems, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, e.g., analogs having chemically modified bases or sugars, as well as backbone modifications. Unless otherwise indicated, nucleic acid sequences are presented in the 5' to 3' direction.In some embodiments, nucleic acids are selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylur ... The base may be or contain: cytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); intercalating bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0230] As used herein, the term "nucleic acid editing domain" refers to a protein or enzyme that can make one or more modifications (e.g., deamination of a cytidine residue) to a nucleic acid (e.g., DNA or RNA). Exemplary nucleic acid editing domains include, but are not limited to, deaminase, nuclease, nickase, recombinase, methyltransferase, methylase, acetylase, acetyltransferase, transcriptional activator, or transcriptional repressor domains. In some embodiments, the nucleic acid editing domain is a deaminase (e.g., a cytidine deaminase, such as APOBEC or AID deaminase).

[0231] As used herein, the term "proliferative disorder" refers to any disorder in which cell or tissue homeostasis is disturbed, in that a cell or cell population exhibits an abnormally elevated proliferation rate. Proliferative disorders include hyperproliferative disorders, such as preneoplastic hyperplastic conditions, and neoplastic disorders. Neoplastic disorders are characterized by abnormal cell proliferation and include both benign and malignant neoplasms. Malignant neoplasms are also referred to as cancer.

[0232] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids in length. A protein, peptide, or polypeptide can refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide can be modified, for example, by the addition of a chemical entity, such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification. A protein, peptide, or polypeptide can be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide can be simply a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide containing protein domains from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) or carboxy-terminal (C-terminal) end of the fusion protein, thus forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. The protein can contain different domains, such as a nucleic acid-binding domain (e.g., the gRNA-binding domain of Cas9, which directs the protein to bind to a target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. In some embodiments, the protein includes an amino acid sequence constituting a protein portion, e.g., a nucleic acid-binding domain, and an organic compound, e.g., a compound that can act as a nucleic acid cleavage agent. In some embodiments, the protein is in a complex with or associated with a nucleic acid, e.g., RNA. Any of the proteins provided herein can be produced by any method known in the art.For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4). th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012), the entire contents of which are incorporated herein by reference.

[0233] As used herein, the term "subject" refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, a cattle, a cat, or a dog. In some embodiments, the subject is a vertebrate, an amphibian, a reptile, a fish, an insect, a fly, or a nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically modified, e.g., a genetically modified non-human subject. The subject can be of either sex and at any stage of development.

[0234] The term "target site" refers to a sequence within a nucleic acid molecule that is deaminated by a deaminase or a fusion protein that includes a deaminase (e.g., a dCas9-deaminase fusion protein provided herein).

[0235] The terms "treatment," "treat," and "treating," as described herein, refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof. As used herein, the terms "treatment," "treat," and "treating," as described herein, refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay their recurrence.

[0236] The term "recombinant," as used herein in the context of a protein or nucleic acid, refers to a protein or nucleic acid that does not occur in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0237] As used herein, the term "pharmaceutical composition" refers to a composition that can be administered to a subject in the context of treating a disease or disorder. In some embodiments, the pharmaceutical composition comprises an active ingredient, such as a nuclease or a nucleic acid encoding a nuclease, and a pharmaceutically acceptable excipient.

[0238] As used herein, the term "base editor (BE)" or "nucleobase editor (NBE)" refers to an agent comprising a polypeptide that can make modifications to bases (e.g., A, T, C, G, or U) in a nucleic acid sequence (e.g., DNA or RNA). In some embodiments, a base editor can deaminate a base of a nucleic acid. In some embodiments, a base editor can deaminate a base of a DNA molecule. In some embodiments, a base editor can deaminate a cytosine (C) in DNA. In some embodiments, a base editor is a fusion protein comprising a nucleic acid-programmable DNA binding protein (napDNAbp) fused to a cytidine deaminase domain. In some embodiments, a base editor comprises Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpfl, C2c1, C2c2, C2c3, or an Argonaute protein fused to a cytidine deaminase. In some embodiments, the base editor comprises a Cas9 nickase (nCas9) fused to a cytidine deaminase. In some embodiments, the base editor comprises a nuclease-inactive Cas9 (dCas9) fused to a cytidine deaminase. In some embodiments, the base editor is fused to an inhibitor of base excision repair, such as a UGI domain. In some embodiments, the base editor comprises a CasX protein fused to a cytidine deaminase. In some embodiments, the base editor comprises a CasY protein fused to a cytidine deaminase. In some embodiments, the base editor comprises a Cpf1 protein fused to a cytidine deaminase. In some embodiments, the base editor comprises a C2c1 protein fused to a cytidine deaminase. In some embodiments, the base editor comprises a C2c2 protein fused to a cytidine deaminase. In some embodiments, the base editor comprises a C2c3 protein fused to a cytidine deaminase. In some embodiments, the base editor comprises an Argonaute protein fused to a cytidine deaminase.

[0239] As used herein, the term "uracil glycosylase inhibitor" or "UGI" refers to a protein that is capable of inhibiting the uracil DNA glycosylase base excision repair enzyme.

[0240] As used herein, the term "Cas9 nickase" refers to a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase comprises a D10A mutation, a histidine at position H840 of SEQ ID NO: 6, or a corresponding mutation in another Cas9 domain, such as any of the Cas9 proteins provided herein. For example, the Cas9 nickase may comprise the amino acid sequence defined by SEQ ID NO: 8. Such a Cas9 nickase has an active HNH nuclease domain and is capable of cleaving the non-target strand of DNA, i.e., the strand bound by the gRNA. Furthermore, such a Cas9 nickase has an inactive RuvC nuclease domain and is incapable of cleaving the target strand of DNA, i.e., the strand where base editing is desired.

[0241] An exemplary Cas9 nickase (cloning vector pPlatTET-gRNA2; Accession No. BAV54124). DETAILED DESCRIPTION OF THE INVENTION

[0242] Detailed Description of Certain Aspects of the Invention Some aspects of the present disclosure provide fusion proteins, which comprise a domain capable of binding to a nucleotide sequence (e.g., a Cas9 or Cpf1 protein) and an enzymatic domain, e.g., a DNA editing domain, such as a deaminase domain. Deamination of nucleic acid bases by a deaminase can lead to point mutations of the respective residues, a process referred to herein as nucleic acid editing. Therefore, fusion proteins comprising a Cas9 variant or domain and a DNA editing domain can be used for targeted editing of nucleic acid sequences. Such fusion proteins are useful for targeted editing of DNA in vitro, e.g., for creating mutant cells or animals; for introducing targeted mutations, e.g., for correcting genetic defects in ex vivo cells, e.g., cells obtained from a subject that are then reintroduced into the same or another subject; and for introducing targeted mutations, e.g., for correcting genetic defects in a subject or introducing deactivating mutations in disease-associated genes. Typically, the Cas9 domain of the fusion proteins described herein does not have any nuclease activity and instead is a Cas9 fragment or dCas9 protein or domain. Another aspect of the present invention provides fusion proteins that include (i) a domain capable of binding to a nucleic acid sequence (e.g., a Cas9 or Cpf1 protein); (ii) an enzymatic domain, such as a DNA editing domain (e.g., a deaminase domain); and (iii) one or more uracil glycosylase inhibitor (UGI) domains. The presence of at least one UGI domain increases base editing efficiency compared to a fusion protein without a UGI domain. Fusion proteins containing two UGI domains further increase base editing efficiency and product purity compared to fusion proteins with one UGI domain or no UGI domain. Methods for using the Cas9 fusion proteins described herein are also provided.

[0243] Nucleic acid-programmable DNA-binding proteinsSome aspects of the present disclosure provide nucleic acid-programmable DNA-binding proteins, which can be used to guide proteins such as base editors to specific nucleic acid (e.g., DNA or RNA) sequences. It should be understood that any of the fusion proteins (e.g., base editors) provided herein can include any nucleic acid-programmable DNA-binding protein (napDNAbp). For example, any of the fusion proteins described herein that include a Cas9 domain can use another napDNAbp, such as CasX, CasY, Cpf1, C2c1, C2c2, C2c3, and Argonaute, in place of the Cas9 domain. Nucleic acid-programmable DNA-binding proteins include, without limitation, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2C3, and Argonaute. One example of a nucleic acid-programmable DNA-binding protein with PAM specificity distinct from Cas9 is Clustered Regularly Interspaced Short Palindromic Repeats from Prevotella and Francisella 1 (Cpf1). Similar to Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 mediates robust DNA interference and has been shown to have distinct properties from Cas9. Cpf1 is an endonuclease guided by a single RNA lacking tracrRNA, which utilizes a T-rich protospacer adjacent motif (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA by staggered DNA double-strand breaks. Of the 16 Cpf1 family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome editing activity in human cells. Cpf1 proteins are known in the art and have been previously described.See, for example, Yamano et al., "Crystal structure of Cpfl in complex with guide RNA and target DNA." Cell (165) 2016, pp. 949-962, the entire contents of which are incorporated herein by reference.

[0244] Nuclease-inactive Cpf1 (dCpf1) variants, which can be used as a programmable DNA-binding protein domain with a guide nucleotide sequence, are also useful in the present compositions and methods. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9, but does not have the HNH endonuclease domain, and the N-terminus of Cpf1 lacks the alpha-helical recognition lobe of Cas9. Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference) showed that the RuvC-like domain of Cpf1 is responsible for cleaving both DNA strands, and inactivation of the RuvC-like domain inactivates Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 (SEQ ID NO: 15) inactivate Cpf1 nuclease activity. In some embodiments, the dead Cpf1 (dCpf1) comprises mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A in SEQ ID NO: 9. It should be understood that any mutation that inactivates the RuvC domain of Cpf1, e.g., a substitution mutation, a deletion, or an insertion, can be used in accordance with the present disclosure.

[0245] In some embodiments, the DNA-binding protein (napDNAbp) programmable by any of the nucleic acids of the fusion proteins provided herein is a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactive Cpf1 (dCpf1). In some embodiments, the Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 9-24. In some embodiments, dCpf1 comprises an amino acid sequence at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs:9-16, and includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A of SEQ ID NO:9. In some embodiments, the dCpf1 protein comprises the amino acid sequence of any one of SEQ ID NOs:9-16. It should be understood that Cpf1 from other species may also be used in accordance with the present disclosure. Wild-type Francisella novicida Cpf1 (SEQ ID NO: 9) (D917, E1006, and D1255 are bold and underlined) [ka] [ka] (SEQ ID NO: 9) Francisella novicida Cpf1 D917A (SEQ ID NO: 10) (A917, E1006, and D1255 are bold and underlined) [ka] [ka] (SEQ ID NO: 10) Francisella novicida Cpf1 E1006A (SEQ ID NO: 11) (D917, A1006, and D1255 are bold and underlined) [ka] [ka] (SEQ ID NO: 11) Francisella novicida Cpf1 D1255A (SEQ ID NO: 12) (D917, E1006, and A1255 are bold and underlined) [ka] (SEQ ID NO: 12) Francisella novicida Cpf1 D917A / E1006A (SEQ ID NO: 13) (A917, A1006, and D1255 are bold and underlined) [ka] [ka] (SEQ ID NO: 13) Francisella novicida Cpf1 D917A / D1255A (SEQ ID NO: 14) (A917, E1006, and A1255 are bold and underlined) [ka] [ka] (SEQ ID NO: 14) Francisella novicida Cpf1 E1006A / D1255A (SEQ ID NO: 15) (D917, A1006, and A1255 are bold and underlined) [ka] [ka] (SEQ ID NO: 15) Francisella novicida Cpf1 D917A / E1006A / D1255A (SEQ ID NO: 16) (A917, A1006, and A1255 are bold and underlined) [ka] (SEQ ID NO: 16)

[0246] In some embodiments, the nucleic acid-programmable DNA-binding protein is a Cpf1 protein from Acidaminococcus species (AsCpf1). Forms of Cpf1 proteins from Acidaminococcus species have been previously described and will be apparent to those skilled in the art. Exemplary Acidaminococcus Cpf1 proteins (AsCpf1) include, without limitation, any of the AsCpf1 proteins provided herein.

[0247] Wild-type AsCpf1-residue R912 is indicated in bold and underlined, and residues 661-667 are indicated in italics and underlined. [ka] (SEQ ID NO: 17)

[0248] AsCpf1(R912A)—Residue A912 is indicated in bold and underlined, and residues 661-667 are indicated in italics and underlined.

[0249] [ka] [ka] (SEQ ID NO: 19)

[0250] In some embodiments, the nucleic acid-programmable DNA-binding protein is a Cpf1 protein (LbCpf1) from a Lachnospiraceae species. Forms of Cpf1 proteins from Lachnospiraceae species have been previously described and will be apparent to those skilled in the art. Exemplary Lachnospiraceae Cpf1 proteins (LbCpf1) include, without limitation, any of the LbCpf1 proteins provided herein.

[0251] In some embodiments, LbCpf1 is a nickase. In some embodiments, the LbCpf1 nickase comprises a R836X mutation relative to SEQ ID NO: 18, where X is any amino acid except R. In some embodiments, the LbCpf1 nickase comprises a R836A mutation relative to SEQ ID NO: 18. In some embodiments, LbCpf1 is nuclease-inactive LbCpf1 (dLbCpf1). In some embodiments, dLbCpf1 comprises a D832X mutation relative to SEQ ID NO: 18, where X is any amino acid except D. In some embodiments, dLbCpf1 comprises a D832A mutation relative to SEQ ID NO: 18. Additional dCpf1 proteins are described in the art, for example, in Li et al. "Base editing with a Cpfl-cytidine deaminase fusion" Nature Biotechnology; March 2018 DOI: 10.1038 / nbt.4102, the entire contents of which are incorporated herein by reference. In some embodiments, the dCpf1 contains one, two, or three of the Cpf1 point mutations D832A, E1006A, D1125A described in Li et al.

[0252] Wild-type LbCpf1-residues R836 and R1138 are indicated in bold and underlined. [ka] (SEQ ID NO: 18)

[0253] LbCpf1(R836A) - residue A836 is indicated in bold and underlined. [ka] [ka] (SEQ ID NO: 20)

[0254] LbCpf1(R1138A) - residue A1138 is indicated in bold and underlined. [ka] [ka] (SEQ ID NO: 21)

[0255] In some embodiments, the Cpfl protein is a disabled Cpfl protein. As used herein, a "disabled Cpfl" protein is a Cpfl protein that has reduced nuclease activity compared to a wild-type Cpfl protein. In some embodiments, the disabled Cpfl protein preferentially cleaves the target strand more efficiently than the non-target strand. For example, the Cpfl protein preferentially cleaves the strand of a double-stranded nucleic acid molecule in which the nucleotide to be edited is located. In some embodiments, the disabled Cpfl protein preferentially cleaves the non-target strand more efficiently than the target strand. For example, the Cpfl protein preferentially cleaves the strand of a double-stranded nucleic acid molecule in which the nucleotide to be edited is not located. In some embodiments, the disabled Cpfl protein preferentially cleaves the target strand at least 5% more efficiently than it cleaves the non-target strand. In some embodiments, the disabled Cpf1 protein preferentially cleaves the target strand at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 50%, 60%, 70%, 80%, 90%, or at least 100% more efficiently than it cleaves the non-target strand.

[0256] In some embodiments, the disabled Cpf1 protein is a non-naturally occurring Cpf1 protein. In some embodiments, the disabled Cpf1 protein comprises one or more mutations relative to a wild-type Cpf1 protein. In some embodiments, the disabled Cpf1 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 mutations relative to a wild-type Cpf1 protein. In some embodiments, the disabled Cpf1 protein comprises a R836A mutation at the corresponding amino acid defined by SEQ ID NO: 18 or in another Cpf1 protein. It should be understood that a Cpf1 comprising a residue homologous to R836A in SEQ ID NO: 18 (e.g., the corresponding amino acid) can also be mutated to achieve similar results. In some embodiments, the disabled Cpf1 protein comprises an R1138A mutation at the corresponding amino acid defined by SEQ ID NO: 18 or in another Cpf1 protein. In some embodiments, the disabled Cpf1 protein comprises an R912A mutation at the corresponding amino acid defined by SEQ ID NO: 17 or in another Cpf1 protein. Without wishing to be bound by any particular theory, residue R836 of SEQ ID NO: 18 (LbCpf1) and residue R912 of SEQ ID NO: 17 (AsCpf1) are examples of corresponding (e.g., homologous) residues. For example, some alignment between SEQ ID NOs: 17 and 18 indicates that R912 and R836 are corresponding residues. [ka]

[0257] In some embodiments, any of the Cpf1 proteins provided herein contain one or more amino acid deletions. In some embodiments, any of the Cpf1 proteins provided herein contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid deletions. Without wishing to be bound by any particular theory, Cpf1 contains a helical region that may disrupt the function of a deaminase (e.g., an APOBEC) fused to Cpf1, which encompasses residues 661-667 of AsCpf1 (SEQ ID NO: 17). This region contains the amino acid sequence KKTGDQK. Accordingly, aspects of the present disclosure provide Cpf1 proteins containing mutations (e.g., deletions) that disrupt this helical region of Cpf1. In some embodiments, the Cpfl protein comprises a deletion of one or more of the following residues of SEQ ID NO: 17, or one or more corresponding deletions in another Cpfl protein: K661, K662, T663, G664, D665, Q666, and K667. In some embodiments, the Cpfl protein comprises the T663 and D665 deletion of SEQ ID NO: 17, or a corresponding deletion in another Cpfl protein. In some embodiments, the Cpfl protein comprises the K662, T663, D665, and Q666 deletion of SEQ ID NO: 17, or a corresponding deletion in another Cpfl protein. In some embodiments, the Cpfl protein comprises the K661, K662, T663, D665, Q666, and K667 deletion of SEQ ID NO: 17, or a corresponding deletion in another Cpfl protein.

[0258] AsCpf1 (deletion of T663 and D665)

[0259] AsCpf1 (deleted K662, T663, D665, and Q666)

[0260] AsCpf1 (deleted K661, K662, T663, D665, Q666, and K667)

[0261] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a nucleic acid-programmable DNA-binding protein that does not require a classical (NGG) PAM sequence in the target sequence. In some embodiments, the napDNAbp is an Argonaute protein. One example of such a nucleic acid-programmable DNA-binding protein is the Argonaute protein (NgAgo) from Natronobacterium gregoryi. NgAgo is an endonuclease guided by ssDNA. NgAgo binds to 5'-phosphorylated ssDNA (gDNA) ~24 nucleotides in length, guiding it to the target site and creating a DNA double-strand break at the gDNA site. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer adjacent motif (PAM). Using a nuclease-inactive NgAgo (dNgAgo) can greatly expand the bases that can be targeted. The characterization and use of NgAgo is described in Gao et al., Nat. Biotechnol., 2016 Jul;34(7):768-73. PubMed PMID: 27136078; Swarts et al., Nature 507(7491)(2014):258-61; and Swarts et al, Nucleic Acids Res. 43(10)(2015):5120-9, each of which is incorporated herein by reference. The sequence of Natronobacterium gregoryi Argonaute is provided by SEQ ID NO: 25.

[0262] In some embodiments, the napDNAbp is an Argonaute protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Argonaute protein. In some embodiments, the napDNAbp is a naturally occurring Argonaute protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NO:25. In some embodiments, the napDNAbp comprises any one of the amino acid sequences of SEQ ID NO:25. Wild-type Natronobacterium gregoryi Argonaute (SEQ ID NO: 25) (SEQ ID NO: 25)

[0263] In some embodiments, napDNAbp is a prokaryotic homolog of an Argonaute protein. Prokaryotic homologs of Argonaute proteins are known, for example, as described in Makarova K., et al., "Prokaryotic homologs of Argonaute proteins are predicted to function as key components of a novel system of defense against mobile genetic elements," Biol. Direct. 2009 Aug 25;4:29. doi: 10.1186 / 1745-6150-4-29, incorporated herein by reference. In some embodiments, napDNAbp is a Marinitega piezophila Argunaute (MpAgo) protein. CRISPR-associated Marinitega piezophila Argonaute (MpAgo) protein cleaves single-stranded target sequences using a 5'-phosphorylated guide. The 5' guide is used by all known Argonautes. The crystal structure of the MpAgo-RNA complex shows a guide strand binding site containing residues that block the 5' phosphate interaction. This data suggests the evolution of an Argonaute subclass with noncanonical specificity for 5'-hydroxylated guides. See, e.g., Kaya et al., "A bacterial Argonaute with noncanonical guide RNA specificity," Proc Natl Acad Sci U S A. 2016 Apr 12;113(15):4057-62, the entire contents of which are incorporated herein by reference. It should be understood that other Argonaute proteins can be used in any of the fusion proteins (e.g., base editors) described herein, for example, to guide a deaminase (e.g., cytidine deaminase) to a target nucleic acid (e.g., ssRNA).

[0264] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a single effector of the microbial CRISPR-Cas system. Single effectors of the microbial CRISPR-Cas system include, without limitation, Cas9, Cpf1, C2c1, C2c2, and C2c3. Typically, microbial CRISPR-Cas systems are divided into class 1 and class 2 systems. Class 1 systems have multi-subunit effector complexes, while class 2 systems have single protein effectors. Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three separate class 2 CRISPR-Cas systems (C2c1, C2c2, and C2c3) have been described by Shmakov et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems", Mol. Cell, 2015 Nov. 5;60(3):385-397, the entire contents of which are incorporated herein by reference. The effectors of two of the systems, C2c1 and C2c3, contain a RuvC-like endonuclease domain closely related to Cpf1. The third system, C2c2, contains an effector with two predicted HEPN RNase domains. Unlike C2c1-mediated CRISPR RNA production, mature CRISPR RNA production is tracrRNA-independent. C2c1 depends on both CRISPR RNA and tracrRNA for DNA cleavage. Bacterial C2c2 has been shown to possess a unique RNase activity for CRISPR RNA maturation that is separate from its RNA-activated single-stranded RNA degradation activity. These RNase functions are distinct from each other and from the CRISPR RNA processing behavior of Cpf1.See, e.g., East-Seletsky, et al., "Two distinct RNase activities of CRISPR-C2c2 enable guide-RNA processing and RNA detection," Nature, 2016 Oct 13; 538(7624): 270-273, the entire contents of which are incorporated herein by reference. In vitro biochemical analysis of Leptotrichia shahii C2c2 indicates that C2c2 can be guided by a single CRISPR RNA and programmed to cleave ssRNA targets with a complementary protospacer. Two conserved catalytic residues in the HEPN domain mediate cleavage. Mutation of the catalytic residues creates a catalytically inactive RNA-binding protein. See, e.g., Abudayyeh et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science, 2016 Aug 5; 353(6299), the entire contents of which are incorporated herein by reference.

[0265] The crystal structure of Alicyclobacillus acidoterrastris C2c1 (AacC2c1) has been reported in complex with a chimeric single-molecule guide RNA (sgRNA). See, for example, Liu et al., "C2cl-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism," Mol. Cell, 2017 Jan 19;65(2):310-322, which is incorporated herein by reference. A crystal structure has also been reported for Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex. See, for example, Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease," Cell, 2016 Dec 15;167(7):1814-1828, the entire contents of which are incorporated herein by reference. Catalytically competent conformations of AacC2c1 with both target and non-target DNA strands are independently located and captured within a single RuvC catalytic pocket, and C2c1-mediated cleavage results in staggered seven-nucleotide cuts in the target DNA. Structural comparisons between the C2c1 ternary complex and its previously identified Cas9 and Cpf1 counterparts demonstrate the diversity of the mechanisms employed by the CRISPR-Cas9 system.

[0266] In some embodiments, the DNA-binding protein (napDNAbp) programmable by the nucleic acid of any of the fusion proteins provided herein is a C2c1, C2c2, or C2c3 protein. In some embodiments, the napDNAbp is a C2c1 protein. In some embodiments, the napDNAbp is a C2c2 protein. In some embodiments, the napDNAbp is a C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring C2c1, C2c2, or C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring C2c1, C2c2, or C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 26-28. In some embodiments, the napDNAbp comprises the amino acid sequence of any one of SEQ ID NOs: 26-28. It should be understood that C2c1, C2c2, or C2c3 from other bacterial species can also be used in accordance with the present disclosure. C2c1(uniprot.org / uniprot / T0D7 A2#) sp|T0D7A2|C2Cl_ALIAG CRISPR-associated endonuclease C2c1 OS=Alicyclobacillus acidoterrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2cl PE=1 SV=1 C2c2(uniprot.org / uniprot / P0DOC6) >sp|P0DOC6|C2C2_LEPSD CRISPR-associated endoribonuclease C2c2 OS=Leptotrichia shahii (strain DSM 19757 / CCUG 47503 / CIP 107916 / JCM 16776 / LB37) GN=c2c2 PE=1 SV=1

[0267]

[0268] C2c3, >CEPX01008730.1 Marine metagenomic genome assembly TARA_037_MES_0.1-0.22, contig TARA_037_MES_0.1-0.22_scaffold22115_l, translated from whole-genome shotgun sequence.

[0269]

[0270] In some embodiments, the DNA-binding protein (napDNAbp) programmable by the nucleic acid of any of the fusion proteins provided herein is Cas9 from Archaea (e.g., Nanoarchaea), which constitutes the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, the napDNAbp is CasX or CasY. This is described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21, which is incorporated herein by reference. Using genome-resolved metagenomics, several CRISPR-Cas systems have been identified, including Cas9, which was first reported in the Archaea domain of living organisms. This divergent Cas9 protein was found in Nanoarchaea as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered, which are among the most compact systems ever discovered. In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid-programmable DNA-binding proteins (napDNAbp) and are within the scope of the present disclosure.

[0271] In some embodiments, the DNA-binding protein (napDNAbp) programmable by the nucleic acid of any of the fusion proteins provided herein is a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 29-31. In some embodiments, the napDNAbp comprises the amino acid sequence of any one of SEQ ID NOs: 29-31. It should be understood that CasX and CasY from other bacterial species can also be used in accordance with the present disclosure. CasX(uniprot.org / uniprot / F0NN87;uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS=Sulfolobus islandicus (strain HVE10 / 4) GN=SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG (SEQ ID NO: 29) >tr|F0NH53|F0NH53_SULIR CRISPR-related protein Casx OS=Sulfolobus islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG (SEQ ID NO: 30) CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-related protein CasY [uncultured Parcubacteria group bacterium]

[0272]

[0273] Nucleobase editor Cas9 domain Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations compared to any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein.In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any Cas9 protein, such as any one of the Cas9 amino acid sequences provided herein.

[0274] In some embodiments, the Cas9 domain is a nuclease-inactivated Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactivated dCas9 domain comprises a D10X mutation and a H840X mutation in the amino acid sequence defined by SEQ ID NO: 6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactivated dCas9 domain comprises a D10A mutation and a H840A mutation in the amino acid sequence defined by SEQ ID NO: 6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In one example, the nuclease-inactivated Cas9 domain comprises the amino acid sequence defined by SEQ ID NO: 32 (cloning vector pPlatTET-gRNA2, Accession No. BAV54124).

[0275] Additional suitable nuclease-inactive dCas9 domains are within the scope of this disclosure and will be apparent to those of skill in the art based on this disclosure and knowledge in the art. Exemplary suitable nuclease-inactive dCas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference). In some embodiments, the dCas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the dCas9 domains provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations compared to any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein.In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical stretches of amino acid residues compared to any Cas9 protein, such as any one of the Cas9 amino acid sequences provided herein.

[0276] In some embodiments, the Cas9 domain is a Cas9 nickase. The Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of a double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is base-paired (complementary to) the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase comprises a D10A mutation, a histidine at position 840 of SEQ ID NO: 6, or a mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. For example, the Cas9 nickase can comprise the amino acid sequence defined by SEQ ID NO: 8. In some embodiments, the Cas9 nickase cleaves the non-target strand of a double-stranded nucleic acid molecule that is not base-edited, meaning that the Cas9 nickase cleaves the strand that is not base-paired to the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase comprises an H840A mutation, an aspartic acid residue at position 10 of SEQ ID NO:6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases are within the scope of this disclosure and will be apparent to those of skill in the art based on this disclosure and knowledge in the art.

[0277] Cas9 domains with reduced PAM exclusivity Some aspects of the present disclosure provide Cas9 domains with different PAM specificities. Typically, Cas9 proteins, such as Cas9 from S. pyogenes (spCas9), require a canonical NGG PAM sequence to bind to specific nucleic acid regions. This can limit their ability to edit desired bases in a genome. In some embodiments, the base-editing fusion proteins provided herein may require precise placement, for example, where the target base is placed within a 4-base region (e.g., a "deamination window") approximately 15 bases upstream of the PAM. See Komor, AC, et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. Thus, in some embodiments, any of the fusion proteins provided herein may contain a Cas9 domain capable of binding to nucleotide sequences that do not contain a canonical (e.g., NGG) PAM sequence. Cas9 domains that bind to non-classical PAM sequences have been described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-classical PAM sequences are described in Kleinstiver, BP, et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities" Nature 523, 481-485 (2015); and Kleinstiver, BP, et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition" Nature Biotechnology 33, 1293-1298 (2015); the entire contents of each are incorporated herein by reference.

[0278] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is nuclease-active SaCas9, nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 comprises the amino acid sequence SEQ ID NO: 33. In some embodiments, the SaCas9 comprises an N579X mutation of SEQ ID NO: 33, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid except N. In some embodiments, the SaCas9 comprises an N579A mutation of SEQ ID NO: 33, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence with a non-canonical PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of the E781X, N967X, and R1014X mutations of SEQ ID NO: 33, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of the E781K, N967K, and R1014H mutations of SEQ ID NO: 33, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises the E781K, N967K, or R1014H mutations of SEQ ID NO: 33, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein.

[0279] In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 33-36. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises the amino acid sequence of any one of SEQ ID NOs: 33-36. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein consists of the amino acid sequence of any one of SEQ ID NOs: 33-36. Exemplary SaCas9 Sequences [ka] (SEQ ID NO: 33) Residue N579 of SEQ ID NO: 33, which is underlined and bold, can be mutated (e.g., to A579) to generate a SaCas9 nickase. Exemplary SaCas9d Sequences [ka] (SEQ ID NO: 34) Residue D10 of SEQ ID NO: 34, which is underlined and bold, can be mutated (e.g., to A10) to generate a nuclease-inactive SaCas9d. Exemplary SaCas9n Sequences [ka] [ka] (SEQ ID NO: 35). Residue A579 of SEQ ID NO: 35 is underlined and bold, and can be mutated from N579 of SEQ ID NO: 33 to generate SaCas9 nickase. Exemplary SaKKH Cas9 [ka] (SEQ ID NO: 36). Residue A579 of SEQ ID NO: 36 is underlined and bold, and can be mutated from N579 of SEQ ID NO: 36 to generate SaCas9 nickase. Residues K781, K967, and H1014 of SEQ ID NO: 36 are underlined and italicized, and can be mutated from E781, N967, and R1014 of SEQ ID NO: 36 to generate SaKKH Cas9.

[0280] In some embodiments, the Cas9 domain is a Cas9 domain from Streptococcus pyogenes (SpCas9). In some embodiments, the SpCas9 domain is nuclease-active SpCas9, nuclease-inactive SpCas9 (SpCas9d), or SpCas9 nickase (SpCas9n). In some embodiments, the SpCas9 comprises the amino acid sequence SEQ ID NO: 37. In some embodiments, the SpCas9 comprises a D9X mutation of SEQ ID NO: 37, or a corresponding mutation of any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid except D. In some embodiments, the SpCas9 comprises a D9A mutation of SEQ ID NO: 37, or a corresponding mutation of any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain can bind to a nucleic acid sequence with a non-canonical PAM. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain can bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134E, R1334Q, and T1336R mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1134E, R1334Q, and T1336R mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein.In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, R1334Q, and T1336R mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1134V, R1334Q, and T1336R mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, G1217X, R1334X, and T1336X mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, G1217R, R1334Q, and T1336R mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1134V, G1217R, R1334Q, and T1336R mutations of SEQ ID NO: 37, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein.

[0281] In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 37-41. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises the amino acid sequence of any one of SEQ ID NOs: 37-41. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein consists of the amino acid sequence of any one of SEQ ID NOs: 37-41. Exemplary SpCas9 Exemplary SpCas9n Exemplary SpEQR Cas9 [ka] (SEQ ID NO: 39) Residues E1134, Q1334, and R1336 of SEQ ID NO: 39 are underlined and bold and can be mutated from D1134, R1334, and T1336 of SEQ ID NO: 39 to generate SpEQR Cas9. Exemplary SpVQR Cas9 [ka] [ka] (SEQ ID NO: 40) Residues V1134, Q1334, and R1336 of SEQ ID NO:40 are underlined and bold and can be mutated from D1134, R1334, and T1336 of SEQ ID NO:40 to generate SpVQR Cas9. Exemplary SpVRER Cas9 [ka] [ka] (SEQ ID NO: 41) Residues V1134, R1217, Q1334, and R1336 of SEQ ID NO:41 are underlined and bold and can be mutated from D1134, G1217, R1334, and T1336 of SEQ ID NO:41 to generate SpVRER Cas9.

[0282] The following are exemplary fusion proteins (e.g., base editing proteins) that can bind to nucleic acid sequences with non-classical (e.g., non-NGG) PAM sequences: Exemplary SaBE3 (rAPOBEC1-XTEN-SaCas9n-UGI-NLS) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHH ADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPES KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEEASKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG SGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV (SEQ ID NO: 42) Exemplary SaKKH-BE3 (rAPOBEC1-XTEN-SaCas9n-UGI-NLS) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHH ADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPES KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEEASKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYKNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG SGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV (SEQ ID NO: 43) Exemplary EQR-BE3 (rAPOBEC1-XTEN-Cas9n-UGI-NLS) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPES DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFESPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD SGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV (Sequence number 44) VQR-BE3 (rAPOBEC1-XTEN-Cas9n-UGI-NLS) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPES DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD SGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV (Sequence number 45) VRER-BE3 (rAPOBEC1-XTEN-Cas9n-UGI-NLS) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHH ADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGSETPGTSESATPES DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKEYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD SGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV (SEQ ID NO: 46)

[0283] High-Fidelity Base Editors Some aspects of the present disclosure provide Cas9 fusion proteins (e.g., any of the fusion proteins provided herein) comprising a Cas9 domain with high fidelity. Additional aspects of the present disclosure provide Cas9 fusion proteins (e.g., any of the fusion proteins provided herein) comprising a Cas9 domain with reduced electrostatic interactions between the Cas9 domain and the sugar-phosphate backbone of DNA compared to a wild-type Cas9 domain. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain) comprises one or more mutations that reduce the binding between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, any of the Cas9 fusion proteins provided herein comprise one or more N497X, R661X, Q695X, and / or Q926X mutations in the amino acid sequence provided by SEQ ID NO: 6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, where X is any amino acid. In some embodiments, any of the Cas9 fusion proteins provided herein comprises one or more of the N497A, R661A, Q695A, and / or Q926A mutations in the amino acid sequence provided by SEQ ID NO:6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the Cas9 domain comprises a D10A mutation in the amino acid sequence provided by SEQ ID NO:6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the Cas9 domain (e.g., of any of the fusion proteins provided herein) comprises the amino acid sequence set forth in SEQ ID NO:47. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO:48. Cas9 domains with high fidelity are known in the art and will be apparent to those of skill in the art.For example, high-fidelity Cas9 domains are described in Kleinstiver, BP, et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016); and Slaymaker, IM, et al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015); the entire contents of each are incorporated herein by reference.

[0284]

[0013] It should be understood that base editors provided herein, e.g., base editor 2 (BE2) or base editor 3 (BE3), can be converted to high-fidelity base editors by modifying the Cas9 domain as described herein to create high-fidelity base editors, e.g., high-fidelity base editor 2 (HF-BE2) or high-fidelity base editor 3 (HF-BE3). In some embodiments, base editor 2 (BE2) comprises a deaminase domain, a dCas9, and a UGI domain. In some embodiments, base editor 3 (BE3) comprises a deaminase domain, an nCas9 domain, and a UGI domain. Cas9 domain with mutations in bold and underlined relative to Cas9 of SEQ ID NO: 6 [ka] [ka] (SEQ ID NO: 47) HF-BE3

[0285] Cas9 fusion protein Any of the Cas9 domains disclosed herein (e.g., a nuclease-active Cas9 protein, a nuclease-inactive dCas9 protein, or a Cas9 nickase protein) can be fused to a second protein; thus, the fusion proteins provided herein comprise a Cas9 domain provided herein and a second protein or "fusion partner." In some embodiments, the second protein is fused to the N-terminus of the Cas9 domain. However, in other embodiments, the second protein is fused to the C-terminus of the Cas9 domain. In some embodiments, the second protein fused to the Cas9 domain is a nucleic acid editing domain. In some embodiments, the Cas9 domain and the nucleic acid editing domain are fused via a linker, while in other embodiments, the Cas9 domain and the nucleic acid editing domain are fused directly to each other. In some embodiments, the Cas9 domain and the nucleic acid editing domain are fused via a linker of any length or composition. For example, the linker can be a bond, one or more amino acids, a peptide, or a polymer of any length and composition. In some embodiments, the linker is (GGGS) n (SEQ ID NO: 613), (GGGGS) n (SEQ ID NO: 607), (G) n (SEQ ID NO: 608), (EAAAK) n (SEQ ID NO: 609), (GGS) n (SEQ ID NO: 610), (SGGS) n (SEQ ID NO: 606), SGSETPGTSESATPES (SEQ ID NO: 604), SGGS (GGS) n (SEQ ID NO: 612), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605), or (XP) n (SEQ ID NO: 611), or any combination thereof, wherein n is independently an integer between 1 and 30, and X is any amino acid. In some embodiments, the linker comprises a (GGS) nIn some embodiments, the linker comprises the (GGS) motif, where n is 1, 3, or 7. n (SEQ ID NO: 610) motif, where n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises the amino acid sequence SGGS (GGS) n (SEQ ID NO: 612), where n is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, the linker comprises the amino acid sequence SGGS (GGS) n(SEQ ID NO: 612), where n is 2. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604), also referred to in the examples as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605), also referred to in the examples as a 32-amino acid linker. As illustrated in the examples, the length of the linker can affect the bases to be edited. For example, a 3 amino acid long linker (e.g., (GGS)1) can provide a 2-5, 2-4, 2-3, or 3-4 base editing window relative to the PAM sequence, while a 9 amino acid linker (e.g., (GGS)3 (SEQ ID NO: 610)) can provide a 2-6, 2-5, 2-4, 2-3, 3-6, 3-5, 3-4, 4-6, 4-5, or 5-6 base editing window relative to the PAM sequence. A 16-amino acid linker (e.g., an XTEN linker) can provide a 2-7, 2-6, 2-5, 2-4, 2-3, 3-7, 3-6, 3-5, 3-4, 4-7, 4-6, 4-5, 5-7, 5-6, or 6-7 base window relative to the PAM sequence, with exceptionally strong activity, while a 21-amino acid linker (e.g., (GGS)7 (SEQ ID NO: 610)) can provide a 3-8, 3-7, 3-6, 3-5, 3-4, 4-8, 4-7, 4-6, 4-5, 5-8, 5-7, 5-6, 6-8, 6-7, or 7-8 base editing window relative to the PAM sequence. The novel finding that varying linker lengths can allow the disclosed dCas9 fusion proteins to edit nucleobases at different distances from the PAM sequence provides significant clinical importance. Because the PAM sequence can be at various distances from the disease-causing mutation in the gene to be corrected, it should be understood that the linker lengths listed herein as examples are not meant to be limiting.

[0286] In some embodiments, the second protein comprises an enzymatic domain. In some embodiments, the enzymatic domain is a nucleic acid editing domain. Such a nucleic acid editing domain can be, without limitation, a nuclease, a nickase, a recombinase, a deaminase, a methyltransferase, a methylase, an acetylase, or an acetyltransferase. Non-limiting exemplary binding domains that can be used in accordance with the present disclosure include transcriptional activator domains and transcriptional repressor domains.

[0287] Deaminase domain In some embodiments, the second protein comprises a nucleic acid editing domain. In some embodiments, the nucleic acid editing domain can catalyze a C→U base change. In some embodiments, the nucleic acid editing domain is a deaminase domain. In some embodiments, the deaminase is a cytidine deaminase or a cytidine deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is an APOBEC1 deaminase. In some embodiments, the deaminase is an APOBEC2 deaminase. In some embodiments, the deaminase is an APOBEC3 deaminase. In some embodiments, the deaminase is an APOBEC3A deaminase. In some embodiments, the deaminase is an APOBEC3B deaminase. In some embodiments, the deaminase is an APOBEC3C deaminase. In some embodiments, the deaminase is an APOBEC3D deaminase. In some embodiments, the deaminase is an APOBEC3E deaminase. In some embodiments, the deaminase is an APOBEC3F deaminase. In some embodiments, the deaminase is an APOBEC3G deaminase. In some embodiments, the deaminase is an APOBEC3H deaminase. In some embodiments, the deaminase is an APOBEC4 deaminase. In some embodiments, the deaminase is an activation-induced deaminase (AID). In some embodiments, the deaminase is a vertebrate deaminase. In some embodiments, the deaminase is an invertebrate deaminase. In some embodiments, the deaminase is a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse deaminase. In some embodiments, the deaminase is a human deaminase. In some embodiments, the deaminase is a rat deaminase, e.g., rAPOBEC1. In some embodiments, the deaminase is an activation-induced cytidine deaminase (AID). In some embodiments, the deaminase is a vertebrate deaminase. In some embodiments, the deaminase is an invertebrate deaminase. In some embodiments, the deaminase is a human deaminase. In some embodiments, the deaminase is a rat deaminase, e.g., rAPOBEC1. In some embodiments, the deaminase is an activation-induced cytidine deaminase (AID). In some embodiments, the deaminase is cytidine deaminase 1 (CDA1).In some embodiments, the deaminase is Petromyzon marinus cytidine deaminase 1 (pmCDA1). In some embodiments, the deaminase is human APOBEC3G (SEQ ID NO: 60). In some embodiments, the deaminase is a fragment of human APOBEC3G (SEQ ID NO: 83). In some embodiments, the deaminase is a human APOBEC3G variant comprising a D316R_D317R mutation (SEQ ID NO: 82). In some embodiments, the deaminase is a fragment of human APOBEC3G and comprises a mutation corresponding to the D316R_D317R mutation of SEQ ID NO: 60 (SEQ ID NO: 84).

[0288] In some embodiments, the nucleic acid editing domain is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the deaminase domain of any one of SEQ ID NOs: 49-84. In some embodiments, the nucleic acid editing domain comprises the amino acid sequence of any one of SEQ ID NOs: 49-84.

[0289] Deaminase domains regulate the editing window of base editors Some aspects of the present disclosure are based on the recognition that modulating the catalytic activity of the deaminase domain of any of the fusion proteins provided herein, for example, by making point mutations in the deaminase domain, affects the processivity of the fusion protein (e.g., a base editor). For example, mutations that reduce, but do not eliminate, the catalytic activity of the deaminase domain within a base editing fusion protein can make the deaminase domain less likely to catalyze the deamination of residues adjacent to a target residue, thereby narrowing the deamination window. The ability to narrow the deamination window can prevent unwanted deamination of residues adjacent to a specific target residue, which can reduce or prevent off-target effects.

[0290] In some embodiments, any of the fusion proteins provided herein comprise a deaminase domain (e.g., a cytidine deaminase domain) with reduced deaminase catalytic activity. In some embodiments, any of the fusion proteins provided herein comprise a deaminase domain (e.g., a cytidine deaminase domain) with reduced deaminase catalytic activity compared to a suitable control. For example, a suitable control can be the deaminase activity of a deaminase prior to introducing one or more mutations into the deaminase. In other embodiments, a suitable control can be a wild-type deaminase. In some embodiments, a suitable control is a wild-type apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, a suitable control is APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, or APOBEC3H deaminase. In some embodiments, a suitable control is activation-induced deaminase (AID). In some embodiments, a suitable control is cytidine deaminase 1 (pmCDA1) from Petromyzon marinus. In some embodiments, the deaminase domain can have at least 1%, at least 5%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% less deaminase catalytic activity compared to a suitable control.

[0291] In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of H121X, H122X, R126X, R126X, R118X, W90X, W90X, and R132X of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations of another APOBEC deaminase, where X is any amino acid. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations of another APOBEC deaminase.

[0292] In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of D316X, D317X, R320X, R320X, R313X, W285X, W285X, R326X of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations of another APOBEC deaminase, where X is any amino acid. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of D316R, D317R, R320A, R320E, R313A, W285A, W285Y, R326E of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations of another APOBEC deaminase.

[0293] In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the H121R and H122R mutations of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R126A mutation of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R126E mutation of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R118A mutation of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the W90A mutation of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the W90Y mutation of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R132E mutation of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the W90Y and R126E mutations of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase.In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R126E and R132E mutations of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the W90Y and R132E mutations of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the W90Y, R126E, and R132E mutations of rAPOBEC1 (SEQ ID NO: 76), or one or more corresponding mutations in another APOBEC deaminase.

[0294] In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the D316R and D317R mutations of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R320A mutation of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R320E mutation of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R313A mutation of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising a W285A mutation in hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising a W285Y mutation in hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising a R326E mutation in hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising a W285Y and R320E mutations in hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase.In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the R320E and R326E mutations of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the W285Y and R326E mutations of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase comprising the W285Y, R320E, and R326E mutations of hAPOBEC3G (SEQ ID NO: 60), or one or more corresponding mutations in another APOBEC deaminase.

[0295] Some aspects of the present disclosure provide fusion proteins comprising (i) a nuclease-inactivated Cas9 domain; and (ii) a nucleic acid editing domain. In some embodiments, the nuclease-inactivated Cas9 domain (dCas9) comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of Cas9 provided by any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein, and includes a mutation that inactivates the nuclease activity of Cas9. Mutations that inactivate the nuclease domain of Cas9 are well known in the art. For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)). In some embodiments, the dCas9 of the present disclosure comprises a D10A mutation in the amino acid sequence provided by SEQ ID NO: 6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, the dCas9 of the present disclosure comprises a H840A mutation in the amino acid sequence provided by SEQ ID NO: 6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. In some embodiments, a dCas9 of the present disclosure comprises both the D10A and H840A mutations of the amino acid sequence provided by SEQ ID NO:6, or the corresponding mutations of any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein.In some embodiments, the Cas9 further comprises a histidine residue at position 840 of the amino acid sequence provided by SEQ ID NO: 6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. The presence of catalytic residue H840 restores the activity of Cas9 to cleave the unedited strand containing a G opposite the target C. Restoration of H840 does not result in cleavage of the target strand containing a C. In some embodiments, the dCas9 comprises the amino acid sequence of SEQ ID NO: 32. It will be understood that other mutations that inactivate the nuclease domain of Cas9 can also be encompassed in the dCas9 of the present disclosure.

[0296] The Cas9 or dCas9 domain containing the mutations disclosed herein can be a full-length Cas9 or a fragment thereof. In some embodiments, a protein containing Cas9 or a fragment thereof is referred to as a "Cas9 variant." A Cas9 variant shares homology to Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9, e.g., a Cas9 comprising the amino acid sequence of SEQ ID NO:6.

[0297] Any of the Cas9 fusion proteins of the present disclosure may further comprise a nucleic acid editing domain (e.g., an enzyme capable of modifying a nucleic acid, such as a deaminase). In some embodiments, the nucleic acid editing domain is a DNA editing domain. In some embodiments, the nucleic acid editing domain has deaminase activity. In some embodiments, the nucleic acid editing domain comprises or consists of a deaminase or deaminase domain. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is an APOBEC1 family deaminase. In some embodiments, the deaminase is activation-induced cytidine deaminase (AID). Some nucleic acid editing domains, and Cas9 fusion proteins comprising such domains, are described in detail herein. Additional suitable nucleic acid editing domains will be apparent to those of skill in the art based on this disclosure and knowledge in the art.

[0298] Some aspects of the present disclosure provide fusion proteins comprising a Cas9 domain fused to a nucleic acid editing domain, wherein the nucleic acid editing domain is fused to the N-terminus of the Cas9 domain. In some embodiments, the Cas9 domain and the nucleic acid editing-editing domain are fused via a linker. In some embodiments, the linker is (GGGS) n (SEQ ID NO: 613), (GGGGS) n (SEQ ID NO: 607), (G) n (SEQ ID NO: 608), (EAAAK) n (SEQ ID NO: 609), (GGS) n (SEQ ID NO: 610), (SGGS) n(SEQ ID NO: 606), SGSETPGTSESATPES (SEQ ID NO: 604) motif (see, e.g., Guilinger JP, Thompson DB, Liu DR. Fusion of catalytically inactive Cas9 to Fokl nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014;32(6):577-82; the entire contents of which are incorporated herein by reference), or (XP) n (SEQ ID NO: 611) motif, or any combination thereof, wherein n is independently an integer between 1 and 30. In some embodiments, n is independently 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30, or any combination thereof when more than one linker or more than one linker motif is present. In some embodiments, the linker is (GGS) n (SEQ ID NO: 610) motif, where n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises a (GGS) n(SEQ ID NO: 610) motif, where n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 604). Additional suitable linker motifs and linker configurations will be apparent to those of skill in the art. In some embodiments, suitable linker motifs and configurations include those described in Chen et al., Fusion protein linkers: property, design and functionality. Adv Drug Deliv Rev. 2013;65(10): 1357-69, the entire contents of which are incorporated herein by reference. Additional suitable linker sequences will be apparent to those of skill in the art based on this disclosure. In some embodiments, the general architecture of exemplary Cas9 fusion proteins provided herein has the structure: [NH2]-[nucleic acid editing domain]-[Cas9]-[COOH] or [NH2]-[nucleic acid editing domain]-[linker]-[Cas9]-[COOH] wherein NH2 is the N-terminus of the fusion protein and COOH is the C-terminus of the fusion protein.

[0299] The fusion proteins of the present disclosure may include one or more additional features. For example, in some embodiments, the fusion protein includes a nuclear localization sequence (NLS). In some embodiments, the NLS of the fusion protein is located between the nucleic acid editing domain and the Cas9 domain. In some embodiments, the NLS of the fusion protein is located at the C-terminus of the Cas9 domain.

[0300] Other exemplary features that may be present are localization sequences, such as cytoplasmic localization sequences, export sequences, such as nuclear export sequences, or other localization sequences, and sequence tags useful for solubilizing, purifying, or detecting the fusion protein. Suitable protein tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, polyhistidine tags, also referred to as histidine tags or His tags, maltose-binding protein (MBP) tags, nus tags, glutathione-S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag 1, Softag 3), strep tags, biotin ligase tags, FlAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein comprises one or more His tags.

[0301] In some embodiments, the nucleic acid editing domain is a deaminase. For example, in some embodiments, the general architecture of an exemplary Cas9 fusion protein having a deaminase domain has the structure: [NH2]-[NLS]-[deaminase]-[Cas9]-[COOH], [NH2]-[Cas9]-[deaminase]-[COOH], [NH2]-[deaminase]-[Cas9]-[COOH], or [NH2]-[deaminase]-[Cas9]-[NLS]-[COOH] wherein the NLS is a nuclear localization sequence, NH2 is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein. Nuclear localization sequences are known in the art and would be apparent to one of ordinary skill in the art. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 614) or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 615). In some embodiments, a linker is inserted between Cas9 and the deaminase. In some embodiments, the NLS is positioned at the C-terminus of the Cas9 domain. In some embodiments, the NLS is positioned at the N-terminus of the Cas9 domain. In some embodiments, the NLS is positioned between the deaminase and the Cas9 domain. In some embodiments, the NLS is positioned at the N-terminus of the deaminase domain. In some embodiments, the NLS is positioned C-terminal to the deaminase domain.

[0302] One exemplary suitable type of nucleic acid editing domain is a cytidine deaminase from the APOBEC family. The apolipoprotein B mRNA editing complex (APOBEC) family of cytidine deaminase enzymes encompasses 11 proteins that serve to initiate mutagenesis in a controlled and beneficial manner. 29 One family member, activation-induced cytidine deaminase (AID), is responsible for antibody maturation by converting cytosine to uracil in ssDNA in a transcription-dependent, strand-biased manner. 30 Apolipoprotein B editing complex 3 (APOBEC3) enzymes provide human cells with protection against certain HIV-1 strains by deaminating cytosines in reverse-transcribed viral ssDNA. 31 All of these proteins contain Zn. 2+ Coordination motif (His-X-Glu-X 23-26 -Pro-Cys-X 2-4hAPOBEC3F requires a Glu residue (-Cys; SEQ ID NO: 616) and a bound water molecule for catalytic activity. The Glu residue acts to activate the water molecule to zinc hydroxide for nucleophilic attack in the deamination reaction. Each family member preferentially deaminates at its own specific "hot spot," ranging from WRC (W is A or T, R is A or G) in hAID to TTC in hAPOBEC3F. 32 The recent crystal structure of the catalytic domain of APOBEC3G revealed a secondary structure containing a five-stranded β-sheet core flanked by six α-helices, which is believed to be conserved throughout the family. 33 The active loop has been shown to be responsible for both ssDNA binding and determining the identity of the "hot spot" 34 Overexpression of these enzymes has been linked to genomic instability and cancer, therefore highlighting the importance of sequence-specific targeting. 35 .

[0303] Some aspects of the present disclosure relate to the recognition that the activity of cytidine deaminase enzymes, such as APOBEC enzymes, can be directed to specific sites in genomic DNA. Without wishing to be bound by any particular theory, advantages of using Cas9 as a recognition agent include: (1) the sequence specificity of Cas9 can be easily altered by simply changing the sgRNA sequence; and (2) Cas9 binds to its target sequence by denaturing dsDNA, resulting in a stretch of DNA that is single-stranded and therefore a viable substrate for the deaminase. It should be understood that other catalytic domains, or catalytic domains from other deaminases, can also be used to create fusion proteins with Cas9, although the present disclosure is not limited in this regard.

[0304] Some aspects of the present disclosure are based on the recognition that Cas9:deaminase fusion proteins can efficiently deaminate nucleotides at positions 3-11 according to the numbering scheme of Figure 3. Given the results provided herein regarding the nucleotides that can be targeted by Cas9:deaminase fusion proteins, one of skill in the art will be able to design a suitable guide RNA for targeting the fusion protein to a target sequence containing the nucleotide to be deaminated.

[0305] In some embodiments, the deaminase domain and the Cas9 domain are fused to each other via a linker. Various linker lengths and flexibilities between the deaminase domain (e.g., AID) and the Cas9 domain can be employed to achieve the optimal length for deaminase activity for a particular application (e.g., in the form (GGGGS) n (SEQ ID NO: 607), (GGS) n (SEQ ID NO: 610), and (G) n (SEQ ID NO: 608) from a very flexible linker (EAAAK) n (SEQ ID NO: 609), (SGGS) n (SEQ ID NO: 606), SGSETPGTSESATPES (SEQ ID NO: 604) (see, e.g., Guilinger JP, Thompson DB, Liu DR. Fusion of catalytically inactive Cas9 to Fokl nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82; the entire contents of which are incorporated herein by reference), and (XP) n (ranging from SEQ ID NO: 611 to a more rigid linker) 36 In some embodiments, the linker is (GGS) n (SEQ ID NO: 610) motif, where n is 1, 3, or 7. In some embodiments, the linker comprises a (SGSETPGTSESATPES (SEQ ID NO: 604) motif.

[0306] Some exemplary suitable nucleic acid editing domains, such as deaminases and deaminase domains, that can be fused to a Cas9 domain in accordance with aspects of the present disclosure are provided below. It should be understood that in some embodiments, the active domain of each sequence can be used, for example, a domain without a localization signal (nuclear localization sequence, nuclear export signal, cytoplasmic localization signal).

[0307] Human AID: [ka] (SEQ ID NO: 49) (Underline: nuclear localization sequence; double underline: nuclear export signal)

[0308] Mouse AID: [ka] (SEQ ID NO: 51) (Underline: nuclear localization sequence; double underline: nuclear export signal)

[0309] Dog AID: [ka] [ka] (SEQ ID NO: 52) (Underline: nuclear localization sequence; double underline: nuclear export signal)

[0310] Bovine AID: [ka] (SEQ ID NO: 53) (Underline: nuclear localization sequence; double underline: nuclear export signal)

[0311] Rat: AID: [ka] (SEQ ID NO: 54) (Underline: nuclear localization sequence; double underline: nuclear export signal)

[0312] Mouse APOBEC-3: [ka] (SEQ ID NO: 55) (Italics: nucleic acid editing domain)

[0313] Rat APOBEC-3: [ka] (SEQ ID NO: 56) (Italics: nucleic acid editing domain)

[0314] Rhesus APOBEC-3G: [ka] (SEQ ID NO: 57) (Italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0315] Chimpanzee APOBEC-3G: [ka] (SEQ ID NO: 58) (Italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0316] Green monkey APOBEC-3G: [ka] (SEQ ID NO: 59) (Italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0317] Human APOBEC-3G: [ka] [ka] (SEQ ID NO: 60) (Italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0318] Human APOBEC-3F: [ka] (SEQ ID NO: 61) (Italics: nucleic acid editing domain)

[0319] Human APOBEC-3B: [ka] (SEQ ID NO: 62) (Italics: nucleic acid editing domain)

[0320] Rat APOBEC-3B: MQPQGLGPNAGMGPVCLGCSHRRPYSPIRNPLKKLYQQTFYFHFKNVRYAWGRKNNFLCYEVNGMDCALPVPLRQGVFRKQGHIHAELCFIYWFHDKVLRV LSPMEEFKVTWYMSWSPCSKCAEQVARFLAAHRNLSLAIFSSRLYYYLRNPNYQQKLCRLIQEGVHVAAMDLPEFKKCWNKFVDNDGQPFRPWMRLRINFS FYDCKLQEIFSRMNLLREDVFYLQFNNSHRVKPVQNRYYRRKSYLCYQLERANGQEPLKGYLLYKKGEQHVEILFLEKMRSMELSQVRITCYLTWSPCPNCARQLAAFKKDHPDLILRIYTSRLYFYWRKKFQKGLCTLWRSGIHVDVMDLPQFADCWTNFVNPQRPFRPWNELEKNSWRIQRRLRRIKESWGL (SEQ ID NO: 63)

[0321] Bovine APOBEC-3B: DGWEVAFRSGTVLKAGVLGVSMTEGWAGSGHPGQGACVWTPGTRNTMNLLREVLFKQQFGNQPRVPAPYYRRKTYLCYQLKQRNDLTLDRGCFRNKKQRHAEIRFIDKINSLDLNPSQSYKIICYITWSPCPNCANELVNFITRNNHLKLEIFASRLYFHWIKSFKMGLQDLQNAGISVAVMTHTEFEDCWEQFVDNQSRPFQPWDKLEQYSASIRRRLQRILTAPI (SEQ ID NO: 64)

[0322] Chimpanzee APOBEC-3B: MNPQIRNPMEWMYQRTFYYNFENEPILYGRSYTWLCYEVKIRRGHSNLLWDTGVFRGQMYSQPEHHAEMCFLSWFCGNQLSAYKCFQITWFVSWTPCPDCVAKLAKFLAEHPNVTLTISAARLY YYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYNEGQPFMPWYKFDDNYAFLHRTLKEIIRHLMDPDTFTFNFNNDPLVLRRHQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGF YGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGQVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQVRASSLCMVPHRPPPPPQSPGPCLPLCSEPPLGSLLPTGRPAPSLPFLLTASFSFPPPASLPPLPSLSLSPGHLPVPSFHSLTSCSIQPPCSSRIRETEGWASVSKEGRDLG (SEQ ID NO: 65)

[0323] Human APOBEC-3C: [ka] (SEQ ID NO: 66) (Italics: nucleic acid editing domain)

[0324] Gorilla APOBEC3C: [ka] (SEQ ID NO: 67) (Italics: nucleic acid editing domain)

[0325] Human APOBEC-3A: [ka] (SEQ ID NO: 68) (Italics: nucleic acid editing domain)

[0326] Rhesus APOBEC-3A: [ka] (SEQ ID NO: 69) (Italics: nucleic acid editing domain)

[0327] Bovine (Bovien) APOBEC-3A: [ka] (SEQ ID NO: 70) (Italics: nucleic acid editing domain)

[0328] Human APOBEC-3H: [ka] (SEQ ID NO: 71) (Italics: nucleic acid editing domain)

[0329] Rhesus APOBEC-3H: MALLTAKTFSLQFNNKRRVNKPYYPRKALLCYQLTPQNGSTPTRGHLKNKKKDHAEIRFINKIKSMGLDETQCYQVTCYLTWSPCPSCAGELVDFIKAHRHLNLRIFASRLYYHWRPNYQEGLLLLCGSQVPVEVMGLPEFTDCWENFVDHKEPPSFNPSEKLEELDKNSQAIKRRLERIKSRSVDVLENGLRSLQLGPVTPSSSIRNSR (SEQ ID NO: 72)

[0330] Human APOBEC-3D: [ka] (SEQ ID NO: 73) (Italics: nucleic acid editing domain)

[0331] Human APOBEC-1: MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR (SEQ ID NO: 74)

[0332] Mouse APOBEC-1: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNTRCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYPPSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK (SEQ ID NO: 75)

[0333] Rat APOBEC-1: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO: 76)

[0334] Human APOBEC-2: MAQKEEAAVATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKGGQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEEPEIQAALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK (SEQ ID NO: 77)

[0335] Mouse APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEVQSKGGQAQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK (SEQ ID NO: 78)

[0336] Rat APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYLWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK (SEQ ID NO: 79)

[0337] Bovine APOBEC-2: MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEPEIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK (SEQ ID NO: 80)

[0338] Petromyzon marinus CDA1(pmCDA1) MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAV (SEQ ID NO: 81)

[0339] Human APOBEC3G D316R_D317R MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCWSKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN (SEQ ID NO: 82)

[0340] Human APOBEC3G A chain MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQ (SEQ ID NO: 83)

[0341] Human APOBEC3G A chain D120R_D121R MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQ (SEQ ID NO: 84)

[0342] In some embodiments, the fusion proteins provided herein comprise the full-length amino acid sequence of a nucleic acid editing enzyme, such as one of the sequences provided above. However, in other embodiments, the fusion proteins provided herein comprise only a fragment of a nucleic acid editing enzyme, rather than the full-length sequence. For example, in some embodiments, the fusion proteins provided herein comprise a Cas9 domain and a fragment of a nucleic acid editing enzyme, e.g., the fragment comprises a nucleic acid editing domain. Exemplary amino acid sequences of nucleic acid editing domains are shown in italicized letters in the sequences above, and additional suitable sequences of such domains will be apparent to those skilled in the art.

[0343] Additional suitable nucleic acid editing enzyme sequences, such as deaminase enzyme and domain sequences, that can be used in accordance with aspects of the present invention, e.g., fused to a nuclease-inactive Cas9 domain, will be apparent to those of skill in the art based on this disclosure. In some embodiments, such additional enzyme sequences include deaminase enzyme or deaminase domain sequences that are at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% similar to the sequences provided herein. Additional suitable Cas9 domains, variants, and sequences will also be apparent to those of skill in the art. Examples of such additional suitable Cas9 domains include, but are not limited to, D10A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference). In some embodiments, Cas9 comprises a histidine residue at position 840 of the amino acid sequence provided by SEQ ID NO: 6, or a corresponding mutation in any Cas9 protein, e.g., any one of the Cas9 amino acid sequences provided herein. The presence of catalytic residue H840 restores the activity of Cas9 to cleave the unedited strand containing a G opposite the target C. Restoration of H840 does not result in cleavage of the target strand containing a C.

[0344] Additional suitable strategies for creating fusion proteins comprising a Cas9 domain and a deaminase domain will be apparent to those skilled in the art based on the present disclosure in combination with general knowledge in the art. Based on the present disclosure and knowledge in the art, suitable strategies for creating fusion proteins according to aspects of the present disclosure, with or without the use of a linker, will also be apparent to those skilled in the art. For example, Gilbert et al., CRISPR-mediated modular RNA-guided regulation of transcription in eukaryotes. Cell. 2013; 154(2): 442-51, showed that a C-terminal fusion of Cas9 with VP64 using two NLSs as linkers (SPKKKRKVEAS, SEQ ID NO: 617) can be employed for transcriptional activation. Mali et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nat Biotechnol. 2013; 31(9): 833-8, report that C-terminal fusions with VP64 without a linker can be employed for transcriptional activation. And Maeder et al., CRISPR RNA-guided activation of endogenous human genes. Nat Methods. 2013; 10: 977-979, report that C-terminal fusions with VP64 using a Gly4Ser (SEQ ID NO: 613) linker can be used as transcriptional activators.Recently, dCas9-Fokl nuclease fusions have been successfully created, which show improved enzymatic specificity compared to the parental Cas9 enzyme (Guilinger JP, Thompson DB, Liu DR. Fusion of catalytically inactive Cas9 to Fokl nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82 and Tsai SQ, Wyvekens N, Khayter C, Foden JA, Thapar V, Reyon D, Goodwin MJ, Aryee MJ, Joung JK. Dimeric CRISPR RNA-guided Fokl nucleases for highly specific genome editing. Nat Biotechnol. 2014; 32(6): 569-76. PMID: In 24770325, SGSETPGTSESATPES (SEQ ID NO: 604) or GGGGS (SEQ ID NO: 607) linkers were used in the FokI-dCas9 fusion proteins, respectively).

[0345] Some aspects of the present disclosure provide fusion proteins comprising (i) a Cas9 enzyme or domain (e.g., a first protein); and (ii) a nucleic acid-editing enzyme or domain (e.g., a second protein). In some aspects, the fusion proteins provided herein further include (iii) a programmable DNA-binding protein, such as a zinc finger domain, a TALE, or a second Cas9 protein (e.g., a third protein). Without wishing to be bound by any particular theory, fusing a programmable DNA-binding protein (e.g., a second Cas9 protein) to a fusion protein comprising (i) a Cas9 enzyme or domain (e.g., a first protein); and (ii) a nucleic acid-editing enzyme or domain (e.g., a second protein) can be useful to improve the specificity of the fusion protein for a target nucleic acid sequence, or to improve the specificity or binding affinity of the fusion protein to bind to a target nucleic acid sequence that does not contain a canonical PAM (NGG) sequence. In some embodiments, the third protein is a Cas9 protein (e.g., a second Cas9 protein). In some embodiments, the third protein is any of the Cas9 proteins provided herein. In some embodiments, the third protein is fused to the fusion protein at the N-terminus of the Cas9 protein (e.g., the first protein). In some embodiments, the third protein is fused to the fusion protein at the C-terminus of the Cas9 protein (e.g., the first protein). In some embodiments, the Cas9 domain (e.g., the first protein) and the third protein (e.g., the second Cas9 protein) are fused via a linker (e.g., a second linker). In some embodiments, the linker is (GGGGS) n (SEQ ID NO: 607), (G) n (SEQ ID NO: 608), (EAAAK) n (SEQ ID NO: 609), (GGS) n (SEQ ID NO: 610), (SGGS) n (SEQ ID NO: 606), SGSETPGTSESATPES (SEQ ID NO: 604), SGGS (GGS) n (SEQ ID NO: 612), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605), or (XP)n (SEQ ID NO: 611) motif, or any combination thereof, wherein n is independently an integer between 1 and 30. In some embodiments, the general architecture of exemplary napDNAbp fusion proteins provided herein comprises the structure: [NH2]-[nucleic acid editing enzyme or domain]-[napDNAbp]-[third protein]-[COOH]; [NH2]-[third protein]-[napDNAbp]-[nucleic acid editing enzyme or domain]-[COOH]; [NH2]-[napDNAbp]-[nucleic acid editing enzyme or domain]-[third protein]-[COOH]; [NH2]-[third protein]-[nucleic acid editing enzyme or domain]-[napDNAbp]-[COOH]; [NH2]-[UGI]-[nucleic acid editing enzyme or domain]-napDNAbp]-[third protein]-[COOH]; [NH2]-[UGI]-[third protein]-[napDNAbp]-[nucleic acid editing enzyme or domain]-[COOH]; [NH2]-[UGI]-[napDNAbp]-[nucleic acid editing enzyme or domain]-[third protein]-[COOH]; [NH2]-[UGI]-[third protein]-[nucleic acid editing enzyme or domain]-[napDNAbp]-[COOH]; [NH2]-[nucleic acid editing enzyme or domain]-[napDNAbp]-[third protein]-[UGI]-[COOH]; [NH2]-[third protein]-[napDNAbp]-[nucleic acid editing enzyme or domain]-[UGI]-[COOH]; [NH2]-[NapDNAbp]-[nucleic acid editing enzyme or domain]-[third protein]-[UGI]-[COOH]; [NH2]-[third protein]-[nucleic acid editing enzyme or domain]-[NapDNAbp]-[UGI]-[COOH]; or [NH2]-[nucleic acid editing enzyme or domain]-[NapDNAbp]-[first UGI domain]-[second UGI domain]-[COOH] wherein NH2 is the N-terminus of the fusion protein and COOH is the C-terminus of the fusion protein. In some embodiments, the "]-[" used in the above general architecture indicates the presence of an optional linker sequence. In other examples, the general architecture of exemplary NapDNAbp fusion proteins provided herein comprises the structure: [NH2]-[nucleic acid editing enzyme or domain]-[NapDNAbp]-[second NapDNAbp protein]-[COOH]; [NH2]-[second NapDNAbp protein]-[NapDNAbp]-[nucleic acid editing enzyme or domain]-[COOH]; [NH2]-[NapDNAbp]-[nucleic acid editing enzyme or domain]-[second NapDNAbp protein]-[COOH]; [NH2]-[second NapDNAbp protein]-[nucleic acid editing enzyme or domain]-[NapDNAbp]-[COOH]; [NH2]-[UGI]-[nucleic acid editing enzyme or domain]-[NapDNAbp]-[second NapDNAbp protein]-[COOH], [NH2]-[UGI]-[second NapDNAbp protein]-[NapDNAbp]-[nucleic acid editing enzyme or domain]-[COOH]; [NH2]-[UGI]-[NapDNAbp]-[nucleic acid editing enzyme or domain]-[second NapDNAbp protein]-[COOH]; [NH2]-[UGI]-[second NapDNAbp protein]-[nucleic acid editing enzyme or domain]-[NapDNAbpHCOOH]; [NH2]-[nucleic acid editing enzyme or domain]-[NapDNAbp]-[second NapDNAbp protein]-[UGI]-[COOH]; [NH2]-[second NapDNAbp protein]-[NapDNAbp]-[nucleic acid editing enzyme or domain]-[UGI]-[COOH]; [NH2]-[NapDNAbp]-[nucleic acid editing enzyme or domain]-[second NapDNAbp protein]-[UGI]-[COOH]; or [NH2]-[second NapDNAbp protein]-[nucleic acid editing enzyme or domain]-[NapDNAbp]-[UGI]-[COOH] wherein NH2 is the N-terminus of the fusion protein and COOH is the C-terminus of the fusion protein. In some embodiments, the "]-[" used in the above general architectures indicates the presence of an optional linker sequence. In some embodiments, the second NapDNAbp is a dCas9 protein. In some instances, the general architectures of exemplary Cas9 fusion proteins provided herein include the structure shown in Figure 3. It should be understood that any of the proteins provided by any of the general architectures of exemplary Cas9 fusion proteins can be connected by one or more of the linkers provided herein. In some embodiments, the linkers are the same. In some embodiments, the linkers are different. In some embodiments, one or more of the proteins provided by any of the general architectures of exemplary Cas9 fusion proteins are not fused via a linker. In some embodiments, the fusion protein further comprises a nuclear targeting sequence, e.g., a nuclear localization sequence. In some embodiments, the fusion proteins provided herein further comprise a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the third protein. In some embodiments, the NLS is fused to the C-terminus of the third protein. In some embodiments, the NLS is fused to the N-terminus of the Cas9 protein. In some embodiments, the NLS is fused to the C-terminus of the Cas9 protein. In some embodiments, the NLS is fused to the N-terminus of the nucleic acid editing enzyme or domain. In some embodiments, the NLS is fused to the C-terminus of the nucleic acid editing enzyme or domain. In some embodiments, the NLS is fused to the N-terminus of the UGI protein. In some embodiments, the NLS is fused to the C-terminus of the UGI protein. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker.

[0346] Uracil glycosylase inhibitor fusion protein Some aspects of the present disclosure relate to fusion proteins comprising a uracil glycosylase inhibitor (UGI) domain. In some embodiments, any of the fusion proteins provided herein comprising a Cas9 domain (e.g., a nuclease-active Cas9 domain, a nuclease-inactive dCas9 domain, or a Cas9 nickase) can be further fused to a UGI domain, either directly or via a linker. Some aspects of the present disclosure provide deaminase-dCas9 fusion proteins, deaminase-nuclease-active Cas9 fusion proteins, and deaminase-Cas9 nickase fusion proteins with increased nucleobase editing efficiency. Without wishing to be bound by any particular theory, a cellular DNA repair response to the presence of U:G heteroduplex DNA may be responsible for the reduced nucleobase editing efficiency in cells. For example, uracil DNA glycosylase (UDG) catalyzes the removal of U from cellular DNA, which can initiate base excision repair, the most common outcome of which is the reversion of a U:G pair to a C:G pair. As demonstrated in the examples below, uracil DNA glycosylase inhibitor (UGI) can inhibit human UDG activity. Therefore, the present disclosure further contemplates a fusion protein comprising a dCas9-nucleic acid editing domain fused to a UGI domain. The present disclosure also contemplates a fusion protein comprising a Cas9 nickase-nucleic acid editing domain fused to a UGI domain. It should be understood that the use of a UGI domain can increase the editing efficiency of a nucleic acid editing domain capable of catalyzing a C→U change. For example, a fusion protein comprising a UGI domain may be more efficient at deaminating C residues. In some embodiments, the fusion protein has the structure: [deaminase]-[any linker sequence]-[dCas9]-[any linker sequence]-[UGI]; [deaminase]-[any linker sequence]-[UGI]-[any linker sequence]-[dCas9]; [UGI]-[any linker sequence]-[deaminase]-[any linker sequence]-[dCas9]; [UGI]-[any linker sequence]-[dCas9]-[any linker sequence]-[deaminase]; [dCas9]-[any linker sequence]-[deaminase]-[any linker sequence]-[UGI]; [dCas9]-[any linker sequence]-[UGI]-[any linker sequence]-[deaminase]; [deaminase]-[optional linker sequence]-[dCas9]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]; [deaminase]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[dCas9]; [first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[deaminase]-[optional linker sequence]-[dCas9]; [first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[dCas9]-[optional linker sequence]-[deaminase]; [dCas9]-[optional linker sequence]-[deaminase]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]; or [dCas9]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[deaminase] Includes:

[0347] In other embodiments, the fusion protein has the structure: [deaminase]-[optional linker sequence]-[Cas9 nickase]-[optional linker sequence]-[UGI]; [deaminase]-[optional linker sequence]-[UGI]-[optional linker sequence]-[Cas9 nickase]; [UGI]-[optional linker sequence]-[deaminase]-[optional linker sequence]-[Cas9 nickase]; [UGI]-[optional linker sequence]-[Cas9 nickase]-[optional linker sequence]-[deaminase]; [Cas9 nickase]-[any linker sequence]-[deaminase]-[any linker sequence]-[UGI]; [Cas9 nickase]-[optional linker sequence]-[UGI]-[optional linker sequence]-[deaminase] [deaminase]-[optional linker sequence]-[Cas9 nickase]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]; [deaminase]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[Cas9 nickase]; [first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[deaminase]-[optional linker sequence]-[Cas9 nickase]; [first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[Cas9 nickase]-[optional linker sequence]-[deaminase]; [Cas9 nickase]-[optional linker sequence]-[deaminase]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]; or [Cas9 nickase]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]-[optional linker sequence]-[deaminase] Includes:

[0348] It should be understood that any of the fusion proteins described above can comprise (i) a nucleic acid-programmable DNA-binding protein (napDNAbp); (ii) a cytidine deaminase domain; and (iii) two or more UGI domains, where the two or more UGI domains can be adjacent to each other in the construct (e.g., [first UGI]-[second UGI], where "-" is an optional linker), or the two or more UGI domains can be separated by (i) napDNAbp and / or (ii) cytidine deaminase domains (e.g., [first UGI]-[deaminase]-[second UGI], [first UGI]-[napDNAbp]-[second UGI], [first UGI]-[deaminase]-[napDNAbp]-[second UGI], etc., where "-" is an optional linker).

[0349] In another aspect, the fusion protein comprises: (i) a Cas9 enzyme or domain; (ii) a nucleic acid-editing enzyme or domain (e.g., a second protein) (e.g., a cytidine deaminase domain); (iii) a first uracil glycosylase inhibitor domain (UGI) (e.g., a third protein); and (iv) a second uracil glycosylase inhibitor domain (UGI) (e.g., a fourth protein). The first and second uracil glycosylase inhibitor domains (UGI) can be the same or different. In some embodiments, the Cas9 domain (e.g., a first protein) and the deaminase (e.g., a second protein) are fused via a linker. In some embodiments, the Cas9 domain is fused to the C-terminus of the deaminase. In some embodiments, the Cas9 protein (e.g., a first protein) and the first UGI domain (e.g., a third protein) are fused via a linker (e.g., a second linker). In some embodiments, the first UGI domain is fused to the C-terminus of the Cas9 protein. In some embodiments, the first UGI domain (e.g., a third protein) and the second UGI domain (e.g., a fourth protein) are fused via a linker (e.g., a third linker). In some embodiments, the second UGI domain is fused to the C-terminus of the first UGI domain. In some embodiments, the linker is (GGGGS) n (SEQ ID NO: 607), (G) n (SEQ ID NO: 608), (EAAAK) n (SEQ ID NO: 609), (GGS) n (SEQ ID NO: 610), (SGGS) n (SEQ ID NO: 606), SGSETPGTSESATPES (SEQ ID NO: 604), SGGS (GGS) n (SEQ ID NO: 612), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605), or (XP) n(SEQ ID NO: 611) motif, or any combination thereof, wherein n is independently an integer between 1 and 30. In some embodiments, the first linker comprises an amino acid sequence of 1 to 50 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 1 to 40 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 1 to 35 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 1 to 30 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 1 to 20 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 10 to 20 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 30 to 40 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 14, 16, or 18 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 16 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 30, 32, or 34 amino acids. In some embodiments, the first linker comprises an amino acid sequence of 32 amino acids. In some embodiments, the first linker comprises the SGSETPGTSESATPES (SEQ ID NO: 604) motif. In some embodiments, the first linker comprises the SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605) motif. In some embodiments, the second linker comprises an amino acid sequence of 1 to 50 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 1 to 40 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 1 to 35 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 1 to 30 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 1 to 20 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 2 to 20 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 2 to 10 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 10 to 20 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 2, 4, or 6 amino acids.In some embodiments, the second linker comprises an amino acid sequence of 7, 9, or 11 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 14, 16, or 18 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 4 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 9 amino acids. In some embodiments, the second linker comprises an amino acid sequence of 16 amino acids. In some embodiments, the second linker is (SGGS). n (SEQ ID NO: 606) motif, where n is an integer between 1 and 30. In some embodiments, the second linker comprises (SGGS) n (SEQ ID NO: 606) motif, wherein n is 1. In some embodiments, the second linker comprises an SGGS (GGS) n (SEQ ID NO: 612) motif, where n is an integer between 1 and 30. In some embodiments, the second linker comprises an SGGS (GGS) n(SEQ ID NO: 612) motif, wherein n is 2. In some embodiments, the third linker comprises an amino acid sequence of 1 to 50 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 1 to 40 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 1 to 35 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 1 to 30 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 1 to 20 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 2 to 20 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 2 to 10 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 10 to 20 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 2, 4, or 6 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 7, 9, or 11 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 14, 16, or 18 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 4 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 9 amino acids. In some embodiments, the third linker comprises an amino acid sequence of 16 amino acids. In some embodiments, the third linker comprises an amino acid sequence of (SGGS) n (SEQ ID NO: 606) motif, where n is an integer between 1 and 30. In some embodiments, the third linker comprises (SGGS) n (SEQ ID NO: 606) motif, and n is 1. In some embodiments, the third linker comprises SGGS (GGS) n (SEQ ID NO: 612) motif, where n is an integer between 1 and 30. In some embodiments, the third linker comprises SGGS (GGS) n (SEQ ID NO: 612) motif, wherein n is 2.

[0350] In some embodiments, the fusion protein has the structure: [deaminase]-[optional linker sequence]-[dCas9]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]; [deaminase]-[optional linker sequence]-[Cas9 nickase]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI]; or [deaminase]-[optional linker sequence]-[Cas9]-[optional linker sequence]-[first UGI]-[optional linker sequence]-[second UGI] Includes:

[0351] In another aspect, the fusion protein comprises: (i) a Cas9 enzyme or domain; (ii) a nucleic acid-editing enzyme or domain (e.g., a second protein) (e.g., a cytidine deaminase domain); and (iii) more than two uracil glycosylase inhibitor (UGI) domains.

[0352] In some embodiments, the fusion proteins provided herein do not include a linker sequence. In some embodiments, one or both of the optional linker sequences are present. In some embodiments, one, two, or three of the optional linker sequences are present.

[0353] In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker sequence. In some embodiments, the fusion protein comprising a UGI further comprises a nuclear targeting sequence, e.g., a nuclear localization sequence. In some embodiments, the fusion proteins provided herein further comprise a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the UGI protein. In some embodiments, the NLS is fused to the C-terminus of the UGI protein. In some embodiments, the NLS is fused to the N-terminus of the Cas9 protein. In some embodiments, the NLS is fused to the C-terminus of the Cas9 protein. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, the NLS is fused to the C-terminus of the deaminase. In some embodiments, the NLS is fused to the N-terminus of a second Cas9. In some embodiments, the NLS is fused to the C-terminus of a second Cas9. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises the amino acid sequence of any one of the NLS sequences provided or referenced herein. In some embodiments, the NLS comprises the amino acid sequence defined by SEQ ID NO: 614 or SEQ ID NO: 615.

[0354] In some embodiments, the UGI domain comprises wild-type UGI or UGI defined by SEQ ID NO: 134. In some embodiments, UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. For example, in some embodiments, the UGI domain comprises a fragment of the amino acid sequence defined by SEQ ID NO: 134. In some embodiments, the UGI fragment comprises an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence defined by SEQ ID NO: 134. In some embodiments, UGI comprises an amino acid sequence homologous to the amino acid sequence defined by SEQ ID NO: 134, or an amino acid sequence homologous to a fragment of the amino acid sequence defined by SEQ ID NO: 134. In some embodiments, proteins comprising UGI or a fragment of UGI or a homolog of UGI or a UGI fragment are referred to as "UGI variants." UGI variants share homology to UGI or a fragment thereof. For example, UGI variants are at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type UGI or the UGI defined by SEQ ID NO: 134. In some embodiments, UGI variants include fragments of UGI such that the fragments are at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to the corresponding fragment of wild-type UGI or the UGI defined by SEQ ID NO: 134. In some embodiments, UGI comprises the following amino acid sequence: >sp|P14739|UNGI_BPPB2 Uracil DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 134)

[0355] Suitable UGI protein and nucleotide sequences are provided herein; additional suitable UGI sequences will be known to those of skill in the art, see, e.g., Wang et al., Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase. J. Biol. Chem. 264: 1163-1171 (1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419 (1997); Ravishankar et al., X-ray analysis of a complex of Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887 (1998); and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J. Mol. Biol. 287:331-346 (1999), the entire contents of each of which are incorporated herein by reference.

[0356] It should be understood that additional proteins can be uracil glycosylase inhibitors. For example, other proteins capable of inhibiting (e.g., sterically blocking) uracil DNA glycosylase base excision repair enzymes are within the scope of this disclosure. In addition, any protein that blocks or inhibits base excision repair is also within the scope of this disclosure. In some embodiments, the fusion proteins described herein comprise one UGI domain. In some embodiments, the fusion proteins described herein comprise more than one UGI domain. In some embodiments, the fusion proteins described herein comprise two UGI domains. In some embodiments, the fusion proteins described herein comprise more than two UGI domains. In some embodiments, a protein that binds to DNA is used. In other embodiments, a UGI surrogate is used. In some embodiments, the uracil glycosylase inhibitor is a protein that binds to single-stranded DNA. For example, the uracil glycosylase inhibitor can be Erwinia tasmaniensis single-stranded binding protein. In some embodiments, the single-stranded binding protein comprises the amino acid sequence (SEQ ID NO: 135). In some embodiments, the uracil glycosylase inhibitor is a protein that binds to uracil. In some embodiments, the uracil glycosylase inhibitor is a protein that binds to uracil in DNA. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA glycosylase protein. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA glycosylase protein that does not excise uracil from DNA. For example, the uracil glycosylase inhibitor is UdgX. In some embodiments, UdgX comprises the amino acid sequence (SEQ ID NO: 136). As another example, the uracil glycosylase inhibitor is catalytically inactive UDG. In some embodiments, the catalytically inactive UDG comprises the amino acid sequence (SEQ ID NO: 137). It should be understood that other uracil glycosylase inhibitors are within the scope of this disclosure and will be apparent to those of skill in the art.In some embodiments, the uracil glycosylase inhibitor is a protein that is homologous to any one of SEQ ID NOs: 135-137 or 143-148. In some embodiments, the uracil glycosylase inhibitor is a protein that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical to any one of SEQ ID NOs: 135-137 or 143-148. Erwinia tasmaniensis SSB (Thermostable single-stranded DNA binding protein) MASRGVNKVILVGNLGQDPEVRYMPNGGAVANITLATSESWRDKQTGETKEKTEWHRVVLFGKLAEVAGEYLRKGSQVYIEGALQTRKWTDQAGVEKYTTEVVVNVGGTMQMLGGRSQGGGASAGGQNGGSNNGWGQPQQPQGGNQFSGGAQQQARPQQQPQQNNAPANNEPPIDFDDDIP (SEQ ID NO: 135) UdgX (binds to uracil in DNA but does not excise it) MAGAQDFVPHTADLAELAAAAGECRGCGLYRDATQAVFGAGGRSARIMMIGEQPGDKEDLAGLPFVGPAGRLLDRALEAADIDRDALYVTNAVKHFKFTRAAGGKRRIHKTPSRTEVVACRPWLIAEMTSVEPDVVVLLGATAAKALLGNDFRVTQHRGEVLHVDDVPGDPALVATVHPSSLLRGPKEERESAFAGLVDDLRVAADVRP (SEQ ID NO: 136) UDG (catalytically inactive human UDG. Binds to but does not cleave off uracil in DNA) MIGQKTLYSFFSPSPARKRHAPSPEPAVQGTGVAGVPEESGDAAAIPAKKAPAGQEEPGTPPSSPLSAEQLDRIQRNKAAALLRLAARNVPVGFGESWKKHLSGEFGKPYFIKLMGFVAEERKHYTVYPPPHQVFTWTQMCDIKDVKVVILGQEPYHGPNQAHGLCFSVQRPVPPPPSLENIYKELSTDIEDFVHPGHGDLSGWAKQGVLLLNAVLTVRAHQANSHKERGWEQFTDAVVSWLNQNSNGLVFLLWGSYAQKKGSAIDRKRHHVLQTAHPSPLSVYRGFFGCRHFSKTNELLQKSGKKPIDWKEL (SEQ ID NO: 137)

[0357] Additional single-stranded DNA-binding proteins that can be used as UGIs are listed below. Other single-stranded DNA-binding proteins can also be used as UGIs, such as Dickey TH, Altschuler SE, Wuttke DS. Single-stranded DNA-binding proteins: multiple domains for multiple functions. Structure. 2013 Jul 2;21(7): 1074-84. doi: 10.1016 / j.str.2013.05.013. Review.; Marceau AH. Functions of single-strand DNA-binding proteins in DNA replication, recombination, and repair. Methods Mol Biol. 2012;922: 1-21. doi: 10.1007 / 978-1-62703-032-8_l.; Mijakovic, Ivan, et al., Bacterial single-stranded DNA-binding proteins are phosphorylated on tyrosine. Nucleic Acids Res 2006; 34 (5): 1588-1596. doi: 10.1093 / nar / gkj514; [ PubMed ] Mumtsidu E, Makhov AM, Konarev PV, Svergun DI, Griffith JD, Tucker PA. Structural features of the single-stranded DNA-binding protein of Epstein-Barrvirus. J Structure Biol. 2008 Feb;161(2):172-87. Epub 2007 Nov 1; Nowak M, Olszewski M, Spibida M, Kur J. Characterization of single-stranded DNA-binding proteins from the psychrophilic bacteria Desulfotalea psychrophila, Flavobacterium psychrophilum, Psychrobacter arcticus, Psychrobacter cryohalolentis, Psychromonas ingrahamii, and Psychroflexus torquis Photobacterium profundum. BMC Microbiol. 2014 Apr 14;14:91. doi: 10.1186 / 1471-2180-14-91; Tone T, Takeuchi A, Makino O. Single-stranded DNA binding protein Gp5 of Bacillus subtilis phage Φ29 is required for viral DNA replication in growth-temperature dependent fashion. Biosci Biotechnol Biochem. 2012;76(12):2351-3. Epub 2012 Dec 7; Wold. REPLICATION PROTEIN A: A Heterotrimeric, Single-Stranded DNA-Binding Protein Required for Eukaryotic DNA Metabolism. Annual Review of Biochem. 1997; 66:61-92. doi: 10.1146 / annurev.biochem.66.1.61; Wu Y, Lu J, Kang T. Human single-stranded DNA binding proteins: guardians of genome stability. Acta Biochim Biophys Sin (Shanghai). 2016 Jul;48(7):671-7. doi: 10.1093 / abbs / gmw044. Epub 2016 May 23. It should be understood that those described in the reviews can be used; the entire contents of each are incorporated herein by reference. mtSSB - SSBP1 Single-stranded DNA-binding protein 1 [Homo sapiens (human)] (UniProtKB: Q04837; NP_001243439.1) MFRRPVLQVLRQFVRHESETTTSLVLERSLNRVHLLGRVGQDPVLRQVEGKNPVTIFSLATNEMWRSGDSEVYQLGDVSQKTTWHRISVFRPGLRDVAYQYVKKGSRIYLEGKIDYGEYMDKNNVRRQATTIIADNIIFLSDQTKEKE (SEQ ID NO: 138) Single-stranded DNA binding protein 3 isoform A [Mus musculus] (UniProtKBQ9D032-1; NCBI Ref: NP_076161.2) MFAKGKGSAVPSDGQAREKLALYVYEYLLHVGAQKSAQTFLSEIRWEKNITLGEPPGFLHSWWCVFWDLYCAAPERRDTCEHSSEAKAFHDYSAAAAPSPVLGNIPPNDGMPGGPIPPGFFQGPPGSQPSPHAQPPPHNPSSMMGPHSQPFMSPRYAGGPRPPIRMGNQPPGGVPGTQPLLPNSMDPTRQQGHPNMGGSMQRMNPPRGMGPMGPGPQNYGSGMRPPPNSLGPAMPGINMGPGAGRPWPNPNSANSIPYSSSSPGTYVGPPGGGGPPGTPIMPSPADSTNSSDNIYTMINPVPPGGSRSNFPMGPGSDGPMGGMGGMEPHHMNGSLGSGDIDGLPKNSPNNISGISNPPGTPRDDGELGGNFLHSFQNDNYSPSMTMSV (SEQ ID NO: 139) RPA1 - Replication protein A 70kDa DNA-binding subunit (UniProtKB: P27694; NCBI Ref: NM_002945.3) MVGQLSEGAIAAIMQKGDTNIKPILQVINIRPITTGNSPPRYRLLMSDGLNTLSSFMLATQLNPLVEEEQLSSNCVCQIHRFIVNTLKDGRRVVILMELEVLKSAEAVGVKIGNPVPYNEGLGQPQVAPPAPAASPAASSRPQPQNGSSGMGSTVSKAYGASKTFGKAAGPSLSHTSGGTQSKVVPIASLTPYQSKWTICARVTNKSQIRTWSNSRGEGKLFSLELVDESGEIRATAFNEQVDKFFPLIEVNKVYYFSKGTLKIANKQFTAVKNDYEMTFNNETSVMPCEDDHHLPTVQFDFTGIDDLENKSKDSLVDIIGICKSYEDATKITVRSNNREVAKRNIYLMDTSGKVVTATLWGEDADKFDGSRQPVLAIKGARVSDFGGRSLSVLSSSTIIANPDIPEAYKLRGWFDAEGQALDGVSISDLKSGGVGGSNTNWKTLYEVKSENLGQGDKPDYFSSVATVVYLRKENCMYQACPTQDCNKKVIDQQNGLYRCEKCDTEFPNFKYRMILSVNIADFQENQWVTCFQESAEAILGQNAAYLGELKDKNEQAFEEVFQNANFRSFIFRVRVKVETYNDESRIKATVMDVKPVDYREYGRRLVMSIRRSALM(SEQ ID NO: 140) RPA2 - Replication Protein A 32 kDa Subunit (UniProtKB: P15927; NCBI Ref: NM_002946) MWNSGFESYGSSSYGGAGGYTQSPGGFGSPAPSQAEKKSRARAQHIVPCTISQLLSATLVDEVFRIGNVEISQVTIVGIIRHAEKAPTNIVYKIDDMTAAPMDVRQWVDTDDTSSENTVVPPETYVKVAGHLRSFQNKKSLVAFKIMPLEDMNEFT...

Claims

1. A fusion protein comprising (i) a nucleic acid-programmable DNA-binding protein (napDNAbp) domain; (ii) a cytidine deaminase domain; and (iii) a uracil glycosylase inhibitor (UGI) domain, wherein the napDNAbp domain is CasX, CasY, Cpf1, dCpf1, Cpf1 nickase, Cpf1 mutant protein, C2c1, C2c2, C2c3, or an Argonaut (Ago) protein.

2. The napDNAbp domain: (a) a CasY protein, optionally having an amino acid sequence that is at least 85%, at least 90%, at least 95%, or at least 98% identical to SEQ ID NO: 31, or (ii) the amino acid sequence of SEQ ID NO: 31; (b) Cpf1, dCpf1, Cpf1 nickase, or Cpf1 mutant protein, wherein the Cpf1 nickase or mutant protein optionally comprises (i) an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 10-24, or (ii) an amino acid sequence that is at least one of SEQ ID NOs: 10-24; (c) a C2c1 protein, wherein the C2c1 protein optionally has an amino acid sequence that is at least 85%, at least 90%, at least 95%, or at least 98% identical to (i) SEQ ID NO: 26, or (ii) the amino acid sequence of SEQ ID NO: 26; (d) A C2c2 protein, wherein any C2c2 protein has an amino acid sequence that is at least 85%, at least 90%, at least 95%, or at least 98% identical to SEQ ID NO: 27, or (ii) comprises the amino acid sequence of SEQ ID NO: 27; (e) a C2c3 protein, wherein the C2c3 protein optionally has an amino acid sequence that is at least 85%, at least 90%, at least 95%, or at least 98% identical to (i) SEQ ID NO: 28, or (ii) comprises the amino acid sequence of SEQ ID NO: 28; (f) an Argonaut (Ago) protein, optionally such Ago protein having (i) an amino acid sequence that is at least 85%, at least 90%, at least 95%, or at least 98% identical to SEQ ID NO: 25, or (ii) an amino acid sequence of SEQ ID NO: 25; or (g) A CasX protein, wherein the CasX protein optionally comprises (i) an amino acid sequence that is at least 85%, at least 90%, at least 95%, or at least 98% identical to SEQ ID NO: 29 or 30, or (ii) the amino acid sequence of SEQ ID NO: 29 or 30, according to claim 1.

3. The fusion protein according to claim 1 or 2, wherein the napDNAbp domain is a Cpf1 nickase (nCpf1), and optionally the Cpf1 nickase is a Cpf1 nickase derived from the genus Lachnospiraceae (LbCpf1) or the genus Acidaminococcus (AsCpf1).

4. The cytidine deaminase domain: (i) A deaminase of the apolipoprotein B mRNA editing complex (APOBEC) family, optionally selected from APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, and APOBEC3H; (ii) an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of sequence numbers 49-84; (iii) Containing one of the amino acid sequences of sequence numbers 49 to 84; (iv) rat APOBEC1 (rAPOBEC1), comprising one or more mutations of W90Y, R126E, or R132E in SEQ ID NO: 76, or corresponding mutations in other APOBEC deaminases; (v) Human APOBEC1 (hAPOBEC1) comprising one or more mutations of W90Y, Q126E, and R132E in SEQ ID NO: 74, or corresponding mutations in other APOBEC deaminases; (vi) Human APOBEC3G (hAPOBEC3G) comprising one or more mutations of W285Y, R320E, and R326E in SEQ ID NO: 60, or corresponding mutations in other APOBEC deaminases; (vii) an activation-inducing deaminase (AID); or (viiii) A fusion protein according to any one of claims 1 to 3, wherein the CDA1 (pmCDA1) is derived from Petromyzon marinus.

5. The fusion protein according to any one of claims 1 to 4, wherein the UGI domain comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to SEQ ID NO: 134, or the amino acid sequence of SEQ ID NO:

134.

6. The fusion protein has the following structure: NH 2 - [cytidine deaminase domain] - [napDNAbp domain] - [UGI domain] - COOH; NH 2 - [UGI domain] - [cytidine deaminase domain] - [napDNAbp domain] - COOH; NH 2 - [cytidine deaminase domain] - [UGI domain] - [napDNAbp domain] - COOH; NH 2 - [napDNAbp domain] - [UGI domain] - [cytidine deaminase domain] - COOH; NH 2 -[napDNAbp domain]-[cytidine deaminase domain]-[UGI domain]-COOH; or NH 2 - [UGI domain] - [napDNAbp domain] - [cytidine deaminase domain] - COOH; A fusion protein according to any one of claims 1 to 5, wherein each case of ```-`['' includes an arbitrary linker.

7. Any one of the linkers is SGGS (GGS) n (SEQ ID NO: 612), (GGGS) n (SEQ ID NO: 613), (GGGGS) n (SEQ ID NO: 607), (G) n (SEQ ID NO: 608), (EAAAK) n (SEQ ID NO: 609), (GGS) n (SEQ ID NO: 610), (SGGS) n (SEQ ID NO: 606), SGSETPGTSESATPES (SEQ ID NO: 604), SGGSSGGS SGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 605), (XP) n (SEQ ID NO: 611), or a combination thereof, wherein n is an integer from 1 to 30, and X is any amino acid, the fusion protein according to claim 6.

8. A fusion protein according to any one of claims 1 to 7, further comprising a nuclear localization sequence (NLS), wherein the NLS is optionally a bifid NLS, and further optionally the NLS comprises the amino acid sequence of SEQ ID NO: 614 (PKKKRKV), SEQ ID NO: 615 (MDSLLMNNRRKFLYQFKNVRWAKGRRETYLC), or SEQ ID NO: 740 (KRTADGSEFEPKKKRKV).

9. The fusion protein according to any one of claims 1 to 8, wherein (1) a cytidine deaminase domain and a napDNAbp domain are linked via a linker, and / or (2) a UGI domain and an NLS are linked via a linker, and / or (3) a napDNAbp domain and a UGI domain are linked via a linker.

10. A fusion protein according to any one of claims 1 to 9, wherein a napDNAbp domain and a UGI domain are linked via a linker containing Sequence ID No. 606 (SGGS).

11. The napDNAbp domain and UGI domain are related to sequence number 612 (SGGS(GGS)) n The fusion protein according to any one of claims 1 to 9, which is linked via a linker containing ), where n is 2.

12. A fusion protein according to any one of claims 1 to 11, wherein a cytidine deaminase domain and a napDNAbp domain are linked via a linker comprising SEQ ID NO: 604 (SGSETGTSESAATPES) or SEQ ID NO: 605 (SGGSGGGSGSSETGTSESAATPESGGGSGGGS).

13. A complex comprising a fusion protein according to any one of claims 1 to 12 and a guide nucleic acid bound to the napDNAbp domain of the fusion protein, wherein the guide nucleic acid is optionally a guide RNA containing the nucleic acid sequence of SEQ ID NO:

618.

14. The complex according to claim 13, wherein the guide nucleic acid is 10 to 100 base pairs long and includes a sequence of at least 10 consecutive base pairs complementary to the target sequence, the target sequence is optionally a DNA sequence, the target sequence is optionally located within the genome of a prokaryote or a eukaryote, the target sequence is optionally located within the genome of a mammal, and the mammal is optionally a human, mouse, or rat.

15. A polynucleotide encoding a fusion protein according to any one of claims 1 to 12, wherein the polynucleotide optionally further encodes a guide nucleic acid.

16. A vector comprising a polynucleotide according to claim 15, wherein the vector optionally comprises a heterolog promoter that drives the expression of the polynucleotide.

17. A cell comprising a fusion protein according to any one of claims 1 to 12, a complex according to claim 13 or 14, a polynucleotide according to claim 15, or a vector according to claim 16.

18. A pharmaceutical composition comprising a fusion protein according to any one of claims 1 to 12, a complex according to claim 13 or 14, a polynucleotide according to claim 15, a vector according to claim 16, or a cell according to claim 17, further comprising optionally pharmaceutically acceptable additives.

19. A fusion protein according to any one of claims 1 to 12, a complex according to claim 13 or 14, a polynucleotide according to claim 15, a vector according to claim 16, a cell according to claim 17, or a pharmaceutical composition according to claim 18, for use as a drug.

20. A fusion protein according to any one of claims 1 to 12, a complex according to claim 13 or 14, a polynucleotide according to claim 15, a vector according to claim 16, a cell according to claim 17, or a pharmaceutical composition according to claim 18, for use in the manufacture of a drug for the treatment of a disease, abnormality, or condition, wherein the disease, abnormality, or condition is in a human subject, and further optionally the disease, abnormality, or condition is cystic fibrosis, phenylketonuria, exfoliative keratosis (EHK), Charcot-Marie-Tooth type 4J The fusion protein, complex, polynucleotide, vector, cell, or pharmaceutical composition is a neobiotic disease associated with a mutant PI3KCA protein, mutant CTNNB1 protein, mutant HRAS protein, or mutant p53 protein, such as neuroblastoma (NB), von Willebrand disease (vWD), congenital myotonia, hereditary renal amyloidosis, dilated cardiomyopathy (DCM), hereditary lymphedema, familial Alzheimer's disease, HIV, prion disease, chronic infantile neurocutaneous arthral syndrome (CINCA), desmin-related myopathy (DRM), or a neobiotic disease associated with a mutant PI3KCA protein, mutant CTNNB1 protein, mutant HRAS protein, or mutant p53 protein.

21. A method comprising contacting a nucleic acid molecule with a fusion protein and guide nucleic acid according to any one of claims 1 to 12, wherein the guide nucleic acid comprises a sequence of at least 10 bases complementary to a target sequence in the genome of an organism containing a target base pair; or contacting a nucleic acid molecule with a complex according to claim 13 or 14, The nucleic acid molecule is optionally DNA, further optionally double-stranded DNA, further optionally includes a target sequence associated with a disease or abnormality, further optionally includes a point mutation associated with a disease or abnormality, and further optionally the action of the fusion protein or complex results in the modification of the point mutation. The method, wherein the method is carried out in vitro, ex vivo, or in vivo in a non-human animal.

22. The method according to claim 21, wherein the target sequence includes a T→C point mutation associated with a disease or abnormality, deamination of the mutated C base results in a sequence not associated with a disease or abnormality, optionally the target sequence codes for a protein, the point mutation is located within a codon and results in a change in the amino acid encoded by the mutated codon compared to a wild-type codon, further optionally deamination of the mutated C results in a change in the amino acid encoded by the mutated codon, and further optionally deamination of the mutated C results in a codon encoding a wild-type amino acid.

23. A method for generating a ribonucleoprotein (RNP) complex, the method comprising (i) complexing a nucleic acid base editing factor with RNA in aqueous solution to form a complex comprising the nucleic acid base editing factor and RNA; and (ii) contacting the complex with a cationic lipid, The method wherein the nucleic acid base editing factor is a fusion protein according to any one of claims 1 to 12, or the complex of (i) is the complex according to claim 13 or 14.

24. The method according to claim 23, wherein the nuclear base editing enzyme and RNA in (i) are complexed in a molar ratio of 1:1 to 1:1.5, the complex in aqueous solution of (i) is optionally in contact with the cationic lipid of (ii) in a volume ratio of 1:2 to 2:1, the nucleic acid base editing enzyme is optionally present in the aqueous solution at a concentration of 10 to 100 μM, the RNA is optionally sgRNA, and the cationic lipid is optionally Lipofectamine® selected from Lipofectamine® 2000, Lipofectamine® 3000, Lipofectamine® MessengerMAX, Lipofectamine® RLTX, or Lipofectamine® RNAiMAX.

25. A method for delivering a fusion protein according to any one of claims 1 to 12, a complex according to claim 13 or 14, a polynucleotide according to claim 15, a vector according to claim 16, or a pharmaceutical composition according to claim 18 to the inner ear of a non-human animal.

26. A method for editing nucleic acid base pairs in a double-stranded DNA sequence, wherein the method is A complex comprising a nucleic acid base editing factor and a guide nucleic acid is brought into contact with a target region of the double-stranded DNA sequence, wherein the target region contains the target nucleic acid base pair; thereby To induce chain separation in the target region; Converting the first nucleic acid base of the target nucleic acid base pair in a single strand of the target region to a second nucleic acid base; To sever only one strand of the chain in the target region; Substituting the third nucleic acid base, which was complementary to the first nucleic acid base, with a fourth nucleic acid base, which is complementary to the second nucleic acid base; and Herein, the method causes less than 20% indel formation in the double-stranded DNA sequence; Furthermore, the nucleic acid base editing factor, (i) CasX, CasY, Cpf1, dCpf1, Cpf1 nickase, C2c1, C2c2, C2c3, or algonaut domain; (ii) Cytidine deaminase domain; and (iii) Uracilglycosylase inhibitor (UGI) domain Includes, The method, wherein the method is carried out in vitro, ex vivo, or in vivo in a non-human animal.

27. The method according to claim 26, wherein indel formation is 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or less than 1%, and optionally a second nucleic acid base is substituted with a fifth nucleic acid base complementary to a fourth nucleic acid base to generate the desired edited base pair, and optionally the efficiency of generating the desired edited base pair is at least 5%, or at least 10%, at least 20%, at least 30%, at least 40%, or at least 50%, and optionally the ratio of the desired edited base pair to the unintended edited base pair is 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, or 8:

1.

28. A cleaved single strand hybridizes with a guide nucleic acid, optionally the cleaved single strand faces a strand containing a first nucleic acid base, optionally the first nucleic acid base is cytosine, optionally the second nucleic acid base is uracil, and optionally the target edited base pair is (a) A optionally intended edited base pair located upstream of the PAM site is located 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream of the PAM site; or (b) A optionally intended edited base pair located downstream of the PAM site, where the PAM site is located 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream. The method according to claim 26 or 27, wherein the method optionally does not require a canonical PAM site, optionally the target region optionally includes a target window containing a target nucleic acid base pair, and optionally the target window is 1 to 10 nucleotides long, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long.

29. The method according to any one of claims 26 to 28, wherein the nucleic acid base editing factor comprises the fusion protein described in any one of claims 1 to 12.