Evolved cytidine deaminases and methods for using same to edit DNA

JP2025531669A5Pending Publication Date: 2026-08-25THE BROAD INST INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025508853
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-08-15
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing cytidine deaminases exhibit high Cas-independent off-target editing and are too large for efficient delivery, limiting their application in precise gene editing.

Method used

Development of evolved adenosine deaminases, such as TadA-CD, which preferentially deaminate cytidine on DNA with reduced off-target effects and smaller size, allowing for efficient delivery using AAV vectors.

Benefits of technology

TadA-CD variants maintain high on-target activity while significantly reducing Cas-independent off-target editing, enabling precise gene editing in mammalian cells, including primary human T cells and hematopoietic stem progenitor cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to evolved cytidine deaminases derived from cytidine deaminases and methods for editing DNA using the same. In some aspects, the disclosure describes directed evolution of a TadA-derived adenosine deaminase (TadA-CD) to perform cytidine deamination. In some embodiments, the TadA-CD comprises multiple mutations compared to a parent TadA variant. In some embodiments, the TadA-CD is fused to a programmable DNA-binding protein. Another aspect of the disclosure generally relates to a cytosine base editor (CBE) comprising a programmable DNA-binding protein and TadA-CD. In some embodiments, the disclosed cytosine base editor has improved conversion efficiency and reduced off-target editing frequency compared to naturally occurring CBEs. Polynucleotides, vectors, and kits useful for generating and delivering CBEs are also provided. Cells containing such vectors and CBEs are also provided. Additionally, methods of treatment comprising administering a CBE are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. USSN 63 / 398,483, filed August 16, 2022, entitled "Evolved Cytosine Deaminases and Methods of Editing DNA Using the Same," by David R. Liu et al., and U.S. Provisional Application No. USSN 63 / 380,523, filed October 21, 2022, entitled "Evolved Cytosine Deaminases and Methods of Editing DNA Using the Same," by David R. Liu et al., both of which are incorporated by reference herein in their entireties.

[0002] Government support

[0003] This invention was made with government support under grant numbers RM1HG009490, R01EB027793, R01EB031172, R35GM118062, and U01AI142756 awarded by the National Institutes of Health. The government has certain rights in this invention. Electronic Sequence Listing Reference

[0004] The contents of the electronic sequence listing (B119570170WO00-SEQ-JQM.xml; size 457,025 bytes; created August 15, 2023) are incorporated herein by reference in their entirety. [Background technology]

[0005] Base editors (BEs) are useful tools for performing in vivo forward genetic mutagenesis screening and have the potential to correct pathogenic point mutations by enabling precise incorporation of targeted point mutations in genomic DNA. BEs contain fusions between Cas proteins and base-modifying enzymes (e.g., deaminases). Cytosine base editors (CBEs) convert C·G base pairs to T·A base pairs, and adenine base editors (ABEs) convert A·T base pairs to G·C base pairs. Collectively, CBEs and ABEs can mediate all four possible transition mutations (e.g., C to T, A to G, T to C, and G to A). Reference is made to International Patent Application No. PCT / US2017 / 045381 published February 8, 2018, International Patent Application No. PCT / US2018 / 056146 published April 25, 2019 as WO 2019 / 079347, Koblan et al., Nat Biotechnol (2018), and Gaudelli et al., Nature 551, 464-471 (2017).

[0006] Highly active cytidine deaminases that naturally modify DNA, such as APOBEC family enzymes, can deaminate transiently exposed single-stranded DNA fragments other than those on R-loops generated by the Cas protein domain of CBEs, leading to low-level but widespread Cas-independent modifications of the genome. 13-15,19 Similarly, highly active cytidine deaminases that can tightly engage RNA can also mediate unwanted RNA deamination that is independent of guide RNA hybridization. 21 The significant Cas-independent off-target DNA and RNA editing observed in existing CBEs may limit their use in applications where off-target editing should be minimized. 15Existing CBEs include BE3, which contains the structure NH2-[NLS]-[rAPOBEC1 deaminase]-[Cas9 nickase (D10A)]-[UGI domain]-[NLS]-COOH; BE4, which contains the structure NH2-[NLS]-[rAPOBEC1 deaminase]-[Cas9 nickase (D10A)]-[UGI domain]-[UGI domain]-[NLS]-COOH; and BE4max, a version of BE4 in which the codons in the construct encoding the base editor have been codon-optimized for expression in human cells. Cas-independent off-target effects arise from stochastic binding of the base editor to DNA sites due to the intrinsic affinity of the overexpressed base editor for DNA. Cas-independent off-target DNA editing has been found to be undetectable or much less frequent with some TadA*-based ABEs. 13 , low levels of RNA deamination can be detected from overexpression of some ABEs 8,9,34 .

[0007] There is a need in the art for novel cytidine deaminases and cytosine base editors that maintain high on-target activity while exhibiting lower Cas-independent off-target editing. There is also a need in the art for CBEs of smaller size, e.g., small enough to be encoded by a single adeno-associated virus (AAV) vector (e.g., approximately 4.7 kb packing capacity). Summary of the Invention

[0008] The present disclosure provides the first directed evolution of a deaminase for selectively deaminating different bases. The present disclosure provides variants of adenosine deaminase engineered to preferentially deaminate cytidine on DNA. Accordingly, the present disclosure provides cytidine deaminases that are variants of adenosine deaminases (e.g., wild-type or engineered tRNA adenosine deaminase (TadA)). The present disclosure provides cytosine base editors that include a deaminase variant domain that preferentially deaminates cytidine on DNA and a nucleic acid programmable binding protein (napDNAbp) domain, where the adenosine deaminase variant can deaminate cytidine on nucleic acid molecules to the same or similar extent as existing cytidine deaminases. In some aspects, the present disclosure provides size-minimized deaminase variants that provide base editors with relatively reduced off-target effects while maintaining the high editing efficiency of existing cytosine base editors (CBEs). In some aspects, the disclosure provides base editors, complexes, nucleic acids, vectors, cells, compositions, methods, kits, and uses that utilize the deaminases and base editors provided herein.

[0009] The present disclosure is based at least in part on the hypothesis that adenosine deaminases can be further evolved to recognize cytosine as a substrate.This evolution can result in a new class of highly selective cytidine deaminases and CBEs with high editing efficiency and lower off-target Cas-independent DNA and RNA editing (compared to naturally occurring cytidine deaminases).Wild-type TadA is evolutionarily closely related to cytidine deaminases.In addition, low levels of cytidine deamination have been reported in evolved ABE variants. 11,31,32 Furthermore, the introduction of a mutation into TadA7.10 (TadA-7.10 P48R) disrupted adenosine selectivity and eliminated the 5'-T at protospacer position 6 of the editing window. CIt has been shown to increase cytidine deamination in the context of (counting the SpCas9 protospacer adjacent motif PAM as positions 21-23) 32 Although adenosine deamination in other contexts and locations remains preferred, finally, adenosine deaminases acting on RNA (ADARs) have evolved to perform both cytidine and adenosine deamination on RNA. 33 .

[0010] The present disclosure generally relates to base editors (BEs) for gene editing. Base editors reported to date include, inter alia, a programmable DNA-binding protein domain (e.g., Cas9) fused to a deaminase (e.g., a "base"-modifying domain). In some cases, the BE may also include additional domains that modulate the cellular DNA repair process to increase the efficiency, incorporation, and / or stability of the resulting single-nucleotide change. The programmable DNA-binding domain guides the deaminase to directly convert one base to another at a target site programmed by a guide RNA. Two main classes of BEs have been developed to date: cytidine BEs (CBEs), which convert C·G to T·A, and adenine BEs (ABEs), which convert A·T to G·C. Collectively, CBEs and ABEs enable the correction of all four types of transition mutations (C to T, G to A, A to G, and T to C). Because half of the known disease-associated genetic variants are point mutations and transition mutations account for approximately 60% of known pathogenic point mutations, BE has been widely used to study and treat genetic diseases in a variety of cell types and organisms, including animal models of human genetic diseases.

[0011] CBEs and ABEs may include any programmable DNA binding domain known to those skilled in the art. CBEs further include a deaminase configured to deaminate cytidine, and ABEs include a deaminase configured to deaminate adenosine. Without wishing to be bound by any particular theory, it is generally believed that current CBEs include naturally occurring deaminases or variants thereof configured to deaminate cytidine to uracil. Conversely, ABEs include tRNA-specific adenosine deaminases that have been evolved (e.g., mutated using laboratory techniques such as PACE and PANCE) to accept DNA substrates to enable A·T to G·C editing. Examples include those described in International Patent Application No. PCT / US2021 / 016827, filed February 5, 2021, which is incorporated herein by reference. Already in clinical trials 2 or approved for clinical trials 1 All ABEs reported to date, including 4,8-10 The present invention uses TadA7.10 or an evolved or engineered variant of this deaminase. TadA7.10 is the adenosine deaminase of the prior art ABE ABE7.10, which is disclosed in International Publication No. WO 2018 / 027078, published August 2, 2018. TadA7.10 is also the deaminase domain of ABEmax, which is a variant of ABE7.10 that has been codon-optimized for expression in human cells. For example, the current generation ABE variant ABE8e (which contains a TadA-8e mutant adenosine deaminase) typically achieves higher editing efficiency than existing CBEs, despite the strong tRNA substrate preference of wild-type TadA. 9,11,12 TadA-8e and ABE8e are described in International Publication No. WO 2021 / 158921, published August 12, 2021.

[0012] ABEs have several advantages relative to their CBE counterparts. For example, compared to most CBE deaminases, TadA enzymes are less processive and therefore typically allow for greater single-nucleotide editing precision. 3,7,8,11 ABE also provides lower levels of Cas-independent off-target editing compared to CBE. 8,9,13-15 This advantage is probably due to the fact that it is different from that of wild-type TadA (K for the tRNA stem). m = 830 nM), compared to the tighter independent binding of commonly used cytidine deaminases to nucleic acid substrates (Michaelis constant K for APOBEC1 binding to mRNA). m is 0.21 nM). It likely also arises from the inability of wild-type TadA to process DNA and the fact that TadA-8e evolved exclusively using TadA7.10 in a Cas-dependent manner. Genome mining 19 Protein engineering has provided alternative cytidine deaminases with lower Cas-independent DNA and RNA editing, but to date, these variants suffer from reduced on-target editing activity and / or larger size. 15,20-24 .

[0013] The evolved TadA adenosine deaminase, 166 amino acids in length, is a homolog of APOBEC1 (227 amino acids) and AID (182 amino acids). 25 , CDA (207 amino acids) 7 , or APOBEC3A (198 amino acids) 26 TadA is substantially smaller than commonly used cytidine deaminases such as TadA, making it easier to deliver TadA-derived base editors into cells using size-constrained methods and systems such as AAV. Indeed, the small size of TadA has enabled ABEs, but not CBEs, to be delivered to animal tissues in vivo using a single AAV. 27,28 .

[0014] The inventors of the present disclosure hypothesized that directed evolution of adenosine deaminases to perform cytidine deamination could yield CBEs that maintain high on-target activity but inherit the lower Cas-independent off-target editing and smaller size (e.g., making them easier to deliver into cells by size-constrained methods such as AAV) of current ABEs. Thus, in some embodiments, the present disclosure provides CBEs comprising a mutated adenosine deaminase (which preferentially deaminates cytidine on DNA) and a napDNAbp domain (e.g., Cas9 nickase). Cytidine deaminases evolved from the TadA deaminase described herein are referred to as "TadA-CD," and CBEs disclosed herein containing TadA-CD are referred to herein as "TadCBEs."

[0015] Therefore, aspects of the present disclosure relate to CBEs that include a programmable DNA-binding protein (e.g., Cas9) and an evolved deaminase that preferentially deaminates pyrimidines, particularly cytidine, on DNA. For example, the disclosed TadA-CD deaminase variants exhibit a cytidine deamination to adenine deamination ratio of about 10:1, 15:1, 20:1, or greater than 20:1. In certain embodiments, the disclosed deaminase variants exhibit a cytidine deamination to adenine deamination ratio of about 20:1. One or more TadA-CD deaminases described herein include multiple mutations located on a loop near the active site that are essential for switching the preference for adenosine to cytidine. These mutations confer TadA-CDs that retain activity against cytidines at target regions of DNA, while possessing the additional benefit of low off-target editing frequencies, unlike adenosine deaminases used in existing ABEs, such as TadA-8e. They also have the advantage of being minimized in size (e.g., <4.7 kb), which confers the ability to encode TadCBEs containing these deaminase variants on a single AAV vector rather than between two intein-mediated split AAV vectors, or alternatively, using engineered viral-lipid particles (e.g., those described herein). In some embodiments, the TadCBEs further comprise any napDNAbp domain useful for cytidine base editing activity, and a uracil glycosylase inhibitor (UGI) domain. These TadA-CD variants were generated through continuous and / or discontinuous evolution methodologies, including PACE experiments on TadA-8e substrates (or starting points).

[0016] Another aspect of the present disclosure relates to phage-assisted evolutionary selection systems (e.g., PACE and / or PANCE) for enhancing the substrate specificity of the adenosine deaminase domain of ABE toward cytosine (wherein the ABE contains Cas9 or a Cas9 orthologue). In some embodiments, the selection technique includes a vector system for PACE evolution, including a low-stringency vector and a high-stringency vector. Additional aspects relate to cells containing either of these vectors or the disclosed vector systems. For example, in some embodiments, the highly active adenosine deaminase TadA-8e is evolved (e.g., mutated) to perform cytidine deamination via PACE. The evolved TadA-CD contains mutations located on a loop near the deaminase's active site that are essential for switching the preference for adenosine to cytidine.

[0017] Compared to the most commonly used naturally occurring CBEs, such as BE4max and its variants, the disclosed TadCBEs offer comparable or higher on-target activity, smaller size, and / or substantially lower Cas-independent DNA and RNA off-target editing activity, both of which can be further suppressed without reducing on-target editing by introducing the V106W mutation. As demonstrated herein, these TadCBEs can be used for single or multiple base editing at therapeutically relevant genomic loci in mammalian cells, such as primary human T cells and hematopoietic stem progenitor cells. Other cell types are also possible and are disclosed elsewhere herein. The creation of TadCBEs expands the utility of cytosine base editors for gene editing.

[0018] In some embodiments, the evolved TadA-CD can contain mutations at residues E27, V28, and H96, and can further contain at least one mutation at a residue selected from R26, M61, Y73, I75, M151, Q154, and A158 on the amino acid sequence of SEQ ID NO: 41 (i.e., TadA-8e deaminase), or a corresponding mutation in a homologous adenosine deaminase. Exemplary homologous deaminases include TadA deaminases from any of Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, and Bacillus subtilis. Thus, in some embodiments, the evolved TadA-CD can contain one or more mutations in any of SEQ ID NOs: 317-323, 354, and 355 that confer cytidine activity. In some embodiments, the evolved TadA-CD can include one or more mutations in any of SEQ ID NOs: 34-40, 42-54, 33, 315, and 326 that confer cytidine activity. The deaminases of the present disclosure can be evolved from any adenosine deaminase reported to date to have adenosine deaminase activity.

[0019] In some embodiments, the disclosed TadA-CD variants comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to the amino acid sequence of TadA-8e (SEQ ID NO: 41), wherein the amino acid corresponding to residue 27 of SEQ ID NO: 41 is any amino acid except E.

[0020] In some embodiments, the TadA-CD variant comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to the amino acid sequence of SEQ ID NO: 41, wherein the amino acid corresponding to residue 28 of SEQ ID NO: 41 is any amino acid except V.

[0021] In other embodiments, the TadA-CD variant comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to the amino acid sequence of SEQ ID NO: 41, wherein the amino acid corresponding to residue 96 of SEQ ID NO: 41 is any amino acid except H.

[0022] In some embodiments, the disclosed TadA-CD variants comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to the amino acid sequence of any one of SEQ ID NOs: 34-40. In other embodiments, the TadA-CD variants comprise the amino acid sequence of any one of SEQ ID NOs: 34-40.

[0023] The disclosed TadA-CD variants may further comprise a V106W mutation. In some embodiments, the V106W mutation results in adenine base editing of less than or equal to 1.5%, less than or equal to 1%, less than or equal to 0.75%, less than or equal to 0.5%, less than or equal to 0.25%, less than or equal to 0.1%, less than or equal to 0.05%, or less than or equal to 0.01% among the evaluated targets (the editing frequencies indicated above may represent average or maximum).

[0024] Other aspects of the present disclosure relate to base editors comprising a programmable DNA-binding domain (e.g., napDNAbp) and the disclosed evolved TadA-CD domains. In some embodiments, the napDNAbp of the base editor is a Cas9 protein, such as a Cas9 nickase. In some embodiments, the napDNAbp of the base editor is an Nme2Cas9 protein (e.g., eNme2Cas9 nickase) or an Nme2Cas9 variant. In some embodiments, the napDNAbp of the base editor is any of the proteins listed in Table 6. In some embodiments, the base editor further comprises a UGI domain. In some embodiments, the base editor further comprises a nuclear localization domain. As such, provided herein are TadCBEs. In another aspect, the present disclosure describes complexes comprising any of the disclosed base editors and a guide RNA bound to the napDNAbp domain of the base editor.

[0025] In some aspects, the present disclosure relates to TadA-derived cytidine deaminases that provide efficient conversion of target cytosine to thymine and target adenine to guanine (referred to herein as "TadA dual" deaminases and base editors). TadA dual deaminases can edit C and A bases within a protospacer, particularly within the editing window of the protospacer. These editors incorporate both A to G and C to T editing with roughly equivalent efficiency (e.g., TadA dual, a base editor comprising SEQ ID NO: 39).

[0026] In some embodiments, the TadA dual deaminase is mutated relative to TadA-8e (SEQ ID NO: 41). In some embodiments, the TadA dual deaminase comprises a cytidine deaminase comprising one, two, three, four, or five mutations selected from R26G, V28A, A48R, Y73S, and H96N (e.g., TadA-CDf, SEQ ID NO: 39).

[0027] In some embodiments, the TadA dual deaminase is mutated relative to TadA-CDf (SEQ ID NO: 39). In some embodiments, the TadA dual deaminase comprises a mutation at position N46 of the amino acid sequence of SEQ ID NO: 39. In some embodiments, the Tad dual deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, at least 99.5%, or at least 99.8% identical to the sequence identity of SEQ ID NOs: 39-54.

[0028] In some embodiments, the TadA dual deaminase has an increased affinity for cytosine relative to adenosine. For example, in some embodiments, the dual editor provides A to G and C to T editing at a ratio of 0.7:1, 0.8:1, 0.9:1, 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, or 1.5:1. However, in some embodiments, the TadA dual deaminase has a higher specificity for cytosine than adenosine.

[0029] In other embodiments, the TadA dual (e.g., SEQ ID NO: 39) deaminase can be further mutated (using, e.g., PANCE and / or PACE) to produce a cytidine deaminase with increased affinity for cytosine relative to adenosine. For example, in some embodiments, the ratio of adenosine deaminating activity to cytidine deaminating activity of the deaminase is at least about 0.001:1, 0.005:1, 0.007:1, 0.01:1, 0.05:1, 0.07:1, or 0.1:1.

[0030] Additional aspects of the present disclosure relate to polynucleotides, vectors, and cells encoding napDNAbp, cytidine deaminases, and fusion proteins thereof. In some embodiments, the base editors of the present disclosure can be encoded on the polynucleotides disclosed herein. In some embodiments, the deaminase variants of the present disclosure can be encoded on the polynucleotides disclosed herein. In certain embodiments, the disclosed vectors comprise a polynucleotide encoding any one of the base editors of the present disclosure. In other embodiments, the present disclosure provides cells and compositions comprising any one of the deaminase variants, base editors, complexes, nucleic acids, or vectors described herein. Also provided herein are AAV vectors encoding any of the disclosed base editors and, optionally, a guide RNA.

[0031] Another aspect of the present disclosure provides a pharmaceutical composition comprising any one of the cytidine deaminase or variants thereof, base editors, complexes, viruses, nucleic acids, and / or vectors described herein.

[0032] In some aspects, the present disclosure encompasses methods comprising contacting a nucleic acid molecule (e.g., DNA) with any one of the base editors or complexes described herein. For example, in some embodiments, the method comprises contacting DNA with any one of the BEs described herein having an sgRNA. The contacting in these methods can be in vivo, in vitro, or ex vivo.

[0033] Other embodiments describe methods of using the base editors described herein. In some embodiments, the methods involve using (a) any of the base editors of the invention and (b) a guide RNA that targets the base editor of (a) to a target C:G nucleobase pair on a double-stranded DNA molecule in DNA editing. In other embodiments, the methods involve using a base editor, complex, or pharmaceutical composition of the invention as a medicament. In certain embodiments, the methods involve using a base editor, complex, or pharmaceutical composition of the invention as a medicament for treating a disease, disorder, or condition, such as sickle cell disease or HIV / AIDS.

[0034] In some embodiments, the present disclosure provides methods for selecting (e.g., evolving, engineering, etc.) a cytosine base editor. These methods can include evolving an adenosine base editor through several sequential rounds of PACE and / or PANCE evolution. In certain embodiments, the methods include a selection phage encoding a mutated TadA-8e protein fused to an NpuN intein, a first plasmid encoding an NpuC intein fused to dCas9-UGI, a second plasmid encoding a gIII driven by a T7 or proT7 promoter and encoding an sgRNA, and a third plasmid encoding a T7 RNA polymerase-degron fusion.

[0035] In another aspect, the disclosure encompasses methods of producing one or more of the base editors described herein using any of the vectors described herein.

[0036] Further aspects of the present disclosure also relate to kits comprising a nucleic acid construct comprising (a) a nucleic acid sequence encoding any one of the base editors described herein and (b) a nucleic acid sequence encoding a guide RNA. In some embodiments, the nucleic acid construct further comprises one or more heterologous promoters driving expression of the sequence of (a) and / or the sequence of (b).

[0037] In some aspects, the base editors described herein can be administered to a subject to treat a disease or disorder. Thus, methods are provided in which the described TadCBEs are administered to a subject to edit a target sequence in the subject's genome. The target sequence can include a mutant C:G base pair, for example, a mutant C:G base pair associated with a disease or disorder. In various embodiments of these methods, the degree of cytidine deamination by the base editor exceeds the degree of adenosine deamination by 10, 15, 20, or more than 20 times (a ratio of 10:1, 15:1, 20:1, or more than 20:1).

[0038] The present disclosure further provides the use of any one of the base editors described herein and a guide RNA that targets the base editor to a target C:G base pair on a nucleic acid molecule in the manufacture of a kit or composition for nucleic acid editing, wherein nucleic acid editing comprises contacting the nucleic acid molecule with the base editor and guide RNA under conditions suitable for deamination of a cytosine (C) in the C:G nucleobase pair. The present disclosure further provides the use of any one of the base editors described herein and a guide RNA that targets the base editor to a target C:G base pair on a nucleic acid molecule in the manufacture of a kit for assessing off-target effects of a base editor.

[0039] Other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments of the present disclosure when considered in conjunction with the accompanying figures. [Brief explanation of the drawings]

[0040] Non-limiting aspects of the present disclosure are described, by way of example, with reference to the accompanying figures, which are schematic and not intended to be drawn to scale. In the figures, each identical or nearly identical component illustrated is typically represented by a single numeral. Where illustration is not necessary to allow those skilled in the art to understand the present disclosure, for purposes of clarity, not every component is labeled in every figure, and not every component of every aspect of the present disclosure is shown. In the figures:

[0041] [Figure 1A] Phage-assisted evolution of cytidine deaminase from TadA-8e. [Figure 1B-1C] (Figure 1A) Evolutionary path of TadA-based cytidine deaminases from tRNA deaminase TadA. (Figure 1B) Overview of PACE. Selection phage (purple) encodes the evolving protein. The E. coli host (gray) contains 1) a mutagenesis plasmid (red) for phage diversification, and 2) a plasmid system controlling the expression of pIII (blue, encoded by gIII). Only variants with the desired activity induce pIII production and propagate. Phages without the desired activity cannot propagate and are diluted out of the lagoon. (Figure 1C) [Figure 1D-1E]Selection circuit for cytidine deamination. The TadA-8e variant is encoded on a selection phage (SP, purple). In addition to the mutagenesis plasmid, E. coli contains four accessory plasmids that establish the selection circuit: P1 harbors the base editor Cas9-UGI component. Upon phage infection, the full-length base editor is reconstituted through a split Npu intein system (yellow). P2 encodes the guide RNA and gIII, which is under the transcriptional control of the T7 promoter. P3 contains T7 RNA polymerase, which is inactivated by fusion to a degron tag. C·G to T·A editing activity inserts a stop codon between the T7 RNAP and the degron, resulting in active T7 RNAP, which leads to gIII transcription and phage propagation. (Figure 1D) Two versions of the CBE circuit described herein. In both cases, C·G to T·A editing inserts a stop codon before the degron tag, resulting in active T7 RNAP. A less stringent circuit requires a C·G to T·A edit on the non-coding strand (top) and can tolerate one unwanted A to G edit. A more stringent circuit requires a C·G to T·A edit on the coding strand and cannot tolerate any unwanted A·T to G·C edits. (Figure 1E) Phage-assisted non-sequential evolution of cytidine deaminase from TadA-8e. The ProD (stronger, less stringent) or ProA (weaker, more stringent) promoter used for each PANCE passage is shown. At each passage, phage is diluted 1:50 unless otherwise indicated. After several rounds of evolution, the phage titer stabilized despite increasing dilution rates between passages, suggesting evolution of cytidine deamination activity.

[0042] [Figure 2A-2B]The evolved TadA* variants catalyze cytidine deamination. (Figure 2A) Overview of the TadA-8e variants evolved and characterized herein. The variants represent conserved mutations after 9 passages of PACE or 159 hours of PACE. See Figures 13-14 for a complete list of mutations. (Figure 2B) [Figures 2C-2D] Method for assessing base editing of a target plasmid in E. coli. Cells are cotransformed with the target plasmid (blue) and the base editor plasmid (purple). Base editor expression is induced with arabinose. After 16 hours, cells are harvested and the target plasmid is analyzed by high-throughput sequencing. (Figure 2C) Base editing in E. coli of a protospacer matching the selection circuit target site. C·G to T·A edits are shown in blue. A·T to G·C edits are shown in magenta. Points represent individual biological replicates, and bars represent the mean ± SD from four independent biological replicates. (Figure 2D) Location of evolved mutations in the cryo-EM structure of ABE8e (PDB: 6VPC).

[0043] [Figure 3-1] Characterization of evolved TadCBEs with SpCas9 domains in mammalian cells. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. HEK293T site 3 is abbreviated as HEK3, and HEK293T site 4 is abbreviated as HEK4. [Figure 3-2]Characterization of evolved TadCBEs with SpCas9 domains in mammalian cells. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. HEK293T site 3 is abbreviated as HEK3, and HEK293T site 4 is abbreviated as HEK4. [Figure 3-3] Characterization of evolved TadCBEs with SpCas9 domains in mammalian cells. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. HEK293T site 3 is abbreviated as HEK3, and HEK293T site 4 is abbreviated as HEK4.

[0044] [Figure 4-1]Characterization of evolved deaminases with evolved eNme2-C Cas9 domains in mammalian cells. The designated base editors, using the eNme2-C Cas9 nickase domain (PAM=N4CN) in the ABE8e or BE4max architecture with 2xUGI, were transfected along with each of six guide RNAs targeting the protospacer shown in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are shown in shades of blue. A·T to G·C base edits are shown in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 4-2] Characterization of evolved deaminases with evolved eNme2-C Cas9 domains in mammalian cells. The designated base editors, using the eNme2-C Cas9 nickase domain (PAM=N4CN) in the ABE8e or BE4max architecture with 2xUGI, were transfected along with each of six guide RNAs targeting the protospacer shown in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are shown in shades of blue. A·T to G·C base edits are shown in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0045] [Figure 5A-5B]Characterization of base editing windows and Cas-independent off-target DNA and RNA editing by TadCBE. (Figure 5A) Base editing activity windows for ABE8e, TadCBEa, and TadCBEa-V106W with 2xUGI at nine different target genomic sites in HEK293T. Points represent the average editing across all sites containing the specified base at the indicated position within the protospacer. Individual data points used in this analysis are in Figures 2A-2D, Figure 14, and Figures 16A-16B. (Figure 5B) [Figure 5C-5D] Method for measuring Cas-independent off-target DNA editing by the orthogonal R-loop assay. (Figure 5C) Average Cas-independent off-target editing across all cytosines within six orthogonal R-loops (SaR1–SaR6) generated by dead S. aureus Cas9. (Figure 5D) Off-target RNA editing. RNA was harvested from HEK293T cells 48 hours after transfection with the indicated base editors. After cDNA synthesis, CTNNB1, IP90, and RSL1D1 were amplified and analyzed by high-throughput sequencing. For Figures 5C–5D, points represent individual biological replicates, and bars represent the mean ± SD of three independent biological replicates.

[0046] [Figure 6]Base editing at therapeutically relevant loci by TadCBE in primary human T cells and hematopoietic stem and progenitor cells. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into human T cells (n = 4 donors) along with two synthetic guide RNAs targeting (top) CXCR4 or (middle) CCR5 at the indicated protospacers. Targeted cytosines are blue, targeted adenines are magenta, and PAM sequences are underlined. Three days later, genomic DNA was harvested from T cell lysates and analyzed by high-throughput sequencing. Gray boxes indicate the desired locations of stop codon incorporation in CXCR4 and CCR5. Targeted cytidines that generate TAG (CXCR4) and TAA (CCR5) stop codons by cytosine base editing are underlined. The bottom graph shows that mRNA encoding the indicated base editor or GFP as a negative control was electroporated into hematopoietic stem and progenitor cells along with a synthetic guide RNA targeting the BCL11A enhancer. Three days later, genomic DNA was harvested from cell lysates and analyzed by high-throughput sequencing. C·G to T·A base edits are indicated by blue hues, and A·T to G·C base edits are indicated by magenta hues. Points represent individual biological replicates, and bars represent the mean ± SD from n=4 donors (top and middle rows) or n=3 donors (bottom row).

[0047] [Figure 7] Basis for deamination-selective selection in the PACE and PANCE cycles. In cycle 1, formation of a stop codon is prevented only if the base editor deaminates both A7 and A8. Therefore, cycle 1 tolerates mild levels of A deamination. In cycle 2, deamination of a single adenine, A6, would prevent formation of a stop codon and disrupt cycle activation and phage propagation. Therefore, cycle 2 is more stringent in selecting against adenosine deamination.

[0048] [Figure 8A-1] PANCE titers and evolved TadA-CD genotypes (Figure 8A). [Figure 8A-2] PANCE titers and evolved TadA-CD genotypes (Figure 8A). [Figure 8B] Phage titers during PANCE for lagoons 1–7. Stringency was adjusted by increasing promoter strength from ProD (strongest, least stringent) to ProA (weakest, most stringent), increasing the dilution factor, and switching from circuit 1 to circuit 2. Lagoons 1–6 were inoculated with phage encoding TadA8e-NpuN, and lagoon 7 was inoculated with phage encoding TadA8e A48R-NpuN. (Figure 8B) Genotypes from various PANCE lagoons (L1–L7) after PANCE.

[0049] [Figure 9A-9B] PACE titers and evolved TadA-CD genotypes. (Figure 9A) Phage titers and lagoon flow rates during PACE. Lagoon 1 showed activity-independent growth in S2060 cells after t = 43 h, indicating phage evolved selection-independent replication and did not continue (Figure 9B). [Figure 9C] Genotype of evolved TadA* variants from lagoon 1 at t = 43 h, before the emergence of selection-independent growth. (Fig. 9C) Genotypes of evolved TadA* variants from lagoon 2 at various time points.

[0050] [Figures 10A-10C]AlphaFold model of TadA-CDa. (Figure 10A) The cryo-EM structure of ABE8e bound to DNA containing the 8-azanebularine (8Az) substrate mimic of adenosine (PDB ID 6VPC) is shown. Val 28 (magenta) supports the proper positioning of the adenine substrate relative to the catalytic zinc. (Figure 10B) 8Az was replaced by cytidine using the "Swapna" function in Chimera software. In the resulting model, C4 of cytosine, targeted for nucleophilic attack during deamination, is approximately 1 Å away from the target carbon of 8Az and may therefore require a shift of the DNA substrate for productive catalysis. Val 28 may hinder this shift of the DNA substrate deeper into the TadA-8e pocket. (Figure 10C) AlphaFold3 was used to generate a model of the evolved TadA-CDa. The ABE8e structure was superimposed to generate a model with a DNA substrate R-loop from 6VPC. The evolved enzyme is not predicted to adopt any obvious differences in secondary structure compared to TadA8e. The evolved replacement of Val28 on TadA-8e with the smaller Ala or Gly residue found in TadA-CD may relieve steric constraints predicted to interfere with productive positioning of the target C4 on cytosine relative to the catalytic zinc ion.

[0051] [Figure 11-1] Indel and C·G to G·C editing by SpCas9 variants at nine genomic target sites. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected into HEK293T cells along with each of the nine guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target editing data can be found in Figure 3. [Figure 11-2]Indel and C·G to G·C editing by SpCas9 variants at nine genomic target sites. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected into HEK293T cells along with each of the nine guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target editing data can be found in Figure 3.

[0052] [Figure 12] Indel and C·G to G·C editing by eNme2-C Cas9 variants at six genomic target sites. The designated base editors, using the eNme2-C Cas9 nickase domain in the ABE8e or BE4max architecture with 2×UGI, were transfected into HEK293T cells along with each of the six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data are in Figure 4.

[0053] [Figure 13] V106W proximal to TadA-CD mutations. Mutations generated during the evolution of TadA-CD are shown in blue. Residue V106 is shown in red. Addition of V106W to TadA-7.10, TadA-8e, and TadA-8.17 reduces off-target editing activity4-6. Addition of V106W to TadCBEa-e increases selectivity for deaminating cytidine over adenosine, also reducing off-target editing activity.

[0054] [Figure 14]Base editing by the V106W variant at six genomic target sites. The designated base editors, using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of the six guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0055] [Figure 15] Indels and C·G to G·C edits due to the V106W variant at six genomic target sites. The indicated base editors, using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of the six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0056] [Figure 16A] Base editing, indel formation, and C·G to G·C editing by the TadA-CD(V106W) variant at three additional genomic target sites. The designated base editors, using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2×UGI, were transfected into HEK293T cells along with each of the three guide RNAs targeting the protospacer indicated in each graph (Figure 16A). [Figure 16B]Targeted cytosines are blue, targeted adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are shown in shades of blue. A·T to G·C base edits are shown in shades of magenta (Figure 16B). C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0057] [Figure 17-1] Base editing activity windows of CBEs at nine genomic target sites. Points represent the average editing across all sites containing the specified base at the indicated position within the protospacer. Individual data points used for this analysis are shown in Figures 2A-2D, 14, and 16A-16B. [Figure 17-2] Base editing activity windows of CBEs at nine genomic target sites. Points represent the average editing across all sites containing the specified base at the indicated position within the protospacer. Individual data points used for this analysis are shown in Figures 2A-2D, 14, and 16A-16B.

[0058] [Figure 18] On-target editing of EMX1 in a Cas-independent R-loop editing experiment. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2xUGI were transfected into HEK293T cells along with SpCas9 guide RNAs targeting EMX1 and the indicated SaCas9 sgRNAs. The average on-target C·G to T·A base edits at C5 and C6 on EMX1 are shown for the indicated base editor. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding Cas-dependent off-target data are shown in Figure 5C, Figure 19, and Figure 20.

[0059] [Figure 19-1]Cas-independent off-target C·G to T·A editing at individual sites within six orthogonal R-loops generated by SaCas9. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture 7. Cells were transfected with one SpCas9 sgRNA targeting the base editor and EMX1 locus in conjunction with the orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to SaCas sites 1-6. Points represent individual biological replicates, and bars represent the mean ± SD from three independent biological replicates. Corresponding on-target data are in Figure 18. [Figure 19-2] Cas-independent off-target C·G to T·A editing at individual sites within six orthogonal R-loops generated by SaCas9. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture 7. Cells were transfected with one SpCas9 sgRNA targeting the base editor and EMX1 locus in conjunction with the orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to SaCas sites 1-6. Points represent individual biological replicates, and bars represent the mean ± SD from three independent biological replicates. Corresponding on-target data are in Figure 18.

[0060] [Figure 20-1] Cas-independent off-target C·G to T·A editing by TadCBEe V106W at individual sites within six orthogonal R-loops generated by SaCas9. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture. Cells were transfected with one SpCas9 sgRNA targeting the base editor and EMX1 locus in conjunction with the orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to SaCas sites 1-6. Points represent individual biological replicates, and bars represent the mean ± SD from three independent biological replicates. [Figure 20-2]Cas-independent off-target C·G to T·A editing by TadCBEe V106W at individual sites within six orthogonal R-loops generated by SaCas9. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture. Cells were transfected with one SpCas9 sgRNA targeting the base editor and EMX1 locus in conjunction with the orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to SaCas sites 1-6. Points represent individual biological replicates, and bars represent the mean ± SD from three independent biological replicates.

[0061] [Figures 21A-21C] Cas-independent off-target DNA editing by TadCBEe V106W in six genomic SaCas9 R-loops. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture. Cells were transfected with an orthogonal dead SaCas9 and one SpCas9 sgRNA targeting the base editor and the EMX1 locus (on-target), along with one SaCas9 sgRNA corresponding to Sa sites 1-6 (SaR1-SaR6). Figure 21A shows on-target editing at the EMX1 locus. Figure 21B shows the average C·G to T·A base editing at all adenines within the indicated protospacer. (Figure 21C) The average A·T to G·C base editing at all adenines within the indicated protospacer is depicted on the graph. Points represent individual biological replicates, and bars represent the mean ± SD from three independent biological replicates.

[0062] [Figure 22]Cas-independent off-target DNA editing in six genomic SaCas9 R-loops. Orthogonal R-loop assays were performed on CBE variants in the BE4max architecture. Cells were transfected with one SpCas9 sgRNA targeting the base editor and the EMX1 locus (on-target) along with the orthogonal dead SaCas9 and one SaCas9 sgRNA corresponding to Sa sites 1-6 (SaR1-SaR6). Average A·T to G·C base editing across all adenines within the indicated protospacer is depicted on the graph. Points represent individual biological replicates, and bars represent the mean ± SD from three independent biological replicates.

[0063] [Figures 23A-23B] Cas-independent off-target RNA editing of all cytosines and adenines examined in three transcripts for TadCBEe V106W. Total RNA was harvested from HEK293T cells 48 hours after transfection with the indicated base editors. After cDNA synthesis, CTNNB1, IP90, and RSL1D1 were amplified and analyzed by high-throughput sequencing. Simultaneously, genomic DNA was harvested from the other parallel transfected plate. Genomic DNA was analyzed for on-target editing of EMX1 as a control for base editor activity. Figure 23A shows on-target editing of EMX1 in samples corresponding to the RNA editing analysis. Figure 23B shows the average C to U (blue shades) or A to I (magenta shades). Points represent individual biological replicates, and bars represent the mean ± SD of three independent biological replicates.

[0064] [Figure 24]On-target editing of EMX1 in an RNA off-target editing experiment. The indicated base editors were transfected into HEK293T cells on two parallel plates. In one plate, RNA was harvested from HEK293T cells 48 hours after transfection with the indicated base editor and analyzed as described in Figures 23A-23B. Simultaneously, genomic DNA was harvested from the other parallel transfected plate. Genomic DNA was analyzed for on-target editing of EMX1 as a control for base editor activity. Points represent individual biological replicates, and bars represent the mean ± SD of three independent biological replicates.

[0065] [Figure 25-1] Cas-dependent editing of known off-target sites for HEK3. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2xUGI were transfected into HEK293T cells along with guide RNA targeting HEK293T site 3 (HEK3). Seventy-two hours after transfection, genomic DNA was harvested and known off-target sites were amplified using primers listed in Tables 2A–2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 25-2]Cas-dependent editing of known off-target sites for HEK3. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2xUGI were transfected into HEK293T cells along with guide RNA targeting HEK293T site 3 (HEK3). Seventy-two hours after transfection, genomic DNA was harvested and known off-target sites were amplified using primers listed in Tables 2A–2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0066] [Figure 26-1] Cas-dependent editing of known off-target sites for HEK4. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2x UGI were transfected into HEK293T cells along with guide RNA targeting HEK293T site 4 (HEK4). Seventy-two hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Tables 2A–2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 26-2]Cas-dependent editing of known off-target sites for HEK4. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2x UGI were transfected into HEK293T cells along with guide RNA targeting HEK293T site 4 (HEK4). Seventy-two hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Tables 2A–2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0067] [Figure 27A-1] Cas-dependent editing of known off-target sites for EMX1. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected into HEK293T cells along with guide RNAs targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using primers from Tables 2A-2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 35. [Figure 27A-2]Cas-dependent editing of known off-target sites for EMX1. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected into HEK293T cells along with guide RNAs targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using primers from Tables 2A-2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 35. [Figure 27B-1] Cas-dependent editing of known off-target sites for EMX1. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected into HEK293T cells along with guide RNAs targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using primers from Tables 2A-2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 35. [Figure 27B-2]Cas-dependent editing of known off-target sites for EMX1. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected into HEK293T cells along with guide RNAs targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using primers from Tables 2A-2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 35.

[0068] [Figure 28] Cas-dependent editing of known off-target sites in BCL11A. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2x UGI were transfected into primary human CD34+ hematopoietic stem and progenitor cells (n = 3 donors) along with guide RNAs targeting BCL11A. Seventy-two hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Tables 2A–2E. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0069] [Figure 29A]C·G to G·C edits and indels for T cell experiments targeting CXCR4 and CCR5. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into primary human T cells (n = 4 donors) along with two synthetic guide RNAs targeting (Figure 29) CXCR4 or (Figure 29B) CCR5 at the indicated protospacers. Three days later, genomic DNA was harvested from T cell lysates and analyzed by high-throughput sequencing. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 29B] C·G to G·C edits and indels for T cell experiments targeting CXCR4 and CCR5. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into primary human T cells (n = 4 donors) along with two synthetic guide RNAs targeting (Figure 29) CXCR4 or (Figure 29B) CCR5 at the indicated protospacers. Three days later, genomic DNA was harvested from T cell lysates and analyzed by high-throughput sequencing. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0070] [Figure 30]Cas-dependent off-target editing in T cell experiments targeting CXCR4 and CCR5. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into primary human T cells (n = 4 donors) along with two synthetic guide RNAs targeting CXCR4 or CCR5 at the indicated protospacers. Three days later, genomic DNA was harvested from T cell lysates, and known off-target sites were amplified using the primers in Tables 2A–2E. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0071] [Figure 31A] C·G to G·C edits, indels, and Cas-dependent off-target edits of BCL11A in hematopoietic stem and progenitor cells. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into CD34-positive human hematopoietic stem and progenitor cells (n = 3 donors) along with a synthetic guide RNA targeting BCL11A at the indicated protospacer. Three days later, genomic DNA was harvested from cell lysates and analyzed by high-throughput sequencing. (Figure 31A) C·G to G·C base edits are shown in blue. Indels are shown in gray. (Figure 31B) Known Cas-dependent off-target sites were amplified with primers listed in Tables 2A–2E. C·G to G·C base edits are shown in blue. A·T to G·C base edits are shown in magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 31B]C·G to G·C edits, indels, and Cas-dependent off-target edits of BCL11A in hematopoietic stem and progenitor cells. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into CD34-positive human hematopoietic stem and progenitor cells (n = 3 donors) along with a synthetic guide RNA targeting BCL11A at the indicated protospacer. Three days later, genomic DNA was harvested from cell lysates and analyzed by high-throughput sequencing. (Figure 31A) C·G to G·C base edits are shown in blue. Indels are shown in gray. (Figure 31B) Known Cas-dependent off-target sites were amplified with primers listed in Tables 2A–2E. C·G to G·C base edits are shown in blue. A·T to G·C base edits are shown in magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0072] [Figure 32A] Characterization of TadCBE using a genome-integrated mESC target sequence library. Figure 32A shows the overall efficiency and selectivity of base editors analyzed throughout library editing. Data show the average proportion of edited sequencing reads across all library members between protospacer positions -9 and -20, where positions 21-23 are PAMs. [Figure 32B] Figure 32B shows the editing profiles of BE4max, TadCBEa-d, TadCBEd V106W, and the dual base editor TadDE at 10,683 genome-integrated target sites. The editing window is defined as the protospacer position where the average editing efficiency is ≥ 30% of the average peak editing efficiency. Window plots for all variants tested in the library experiment can be found in Figure 39. [Figure 32C]Figure 32C shows the sequence motifs of TadCBEd and TadCBEd V106W for cytosine and adenine base editing outcomes as determined by regression against editing efficiency. The opacity of the sequence motif is proportional to the test R for the holdout set of sequences. The complete sequence motif plots for all variants are shown in Figures 41A and 41B.

[0073] [Figure 33] Individual mutations are tested in TadCBE. Base editing in E. coli of a protospacer matching the selection circuit target site. Cells are cotransformed with the target and base editor plasmids. Base editor expression is induced with arabinose. After 16 hours, cells are harvested and the target plasmid is analyzed by high-throughput sequencing. (Top graph) Addition of individual mutations identified through evolution to ABE8e is insufficient to generate CBE. (Middle graph) Analysis of mutations in TadCBEa-c and TadCBEe. Mutations in the loop region of ABE8e confer selectivity for cytidine deamination, and complementary mutations boost activity. (Bottom graph) Analysis of mutations in TadCBEd. C·G to T·A edits are shown in blue. A·T to G·C edits are shown in magenta. Points represent individual biological replicates, and bars represent the mean ± SD from four independent biological replicates.

[0074] [Figure 34]TadCBE reversion analysis. Base editing in E. coli of a protospacer matching the target site in the selection circuit. Cells are cotransformed with the target and base editor plasmids. Base editor expression is induced with arabinose. After 16 hours, cells are harvested and the target plasmid is analyzed by high-throughput sequencing. (Top graph) Mutations shown are relative to TadCBEa (gray box). (Bottom graph) Mutations shown are relative to TadCBEe (gray box). C·G to T·A edits are shown in blue. A·T to G·C edits are shown in magenta. Points represent individual biological replicates, and bars represent the mean ± SD from four independent biological replicates.

[0075] [Figure 35] On-target editing of EMX1. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI were transfected into HEK293T cells along with guide RNAs targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Table 1. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding off-target data are shown in Figures 27A and 27B.

[0076] [Figure 36-1]On-target and off-target editing of EMX1 by TadCBE V106W. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2x UGI were transfected into HEK293T cells along with guide RNA targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Table 4. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 36-2] On-target and off-target editing of EMX1 by TadCBE V106W. Designated base editors using the SpCas9 nickase domain in the ABE8e or BE4max architecture with a 2x UGI were transfected into HEK293T cells along with guide RNA targeting EMX1. 72 hours after transfection, genomic DNA was harvested and known off-target sites were amplified using the primers in Table 4. C·G to T·A base edits are indicated by shades of blue. A·T to G·C base edits are indicated by shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0077] [Figure 37] Schematic of the mESC library experiment. Thousands of pairs of sgRNAs and corresponding target sites are integrated into mESCs and treated with a base editor. Base editor-containing cells are enriched by antibiotic selection, and the library cassettes are amplified for high-throughput sequencing.

[0078] [Figure 38-1] Correlation between replicates in the mESC library experiment. Uncorrected C·G to T·A editing efficiency at each target site for each replicate. The dashed red line is the global least-squares regression line. [Figure 38-2] Correlation between replicates in the mESC library experiment. Uncorrected C·G to T·A editing efficiency at each target site for each replicate. The dashed red line is the global least-squares regression line. [Figure 38-3] Correlation between replicates in the mESC library experiment. Uncorrected C·G to T·A editing efficiency at each target site for each replicate. The dashed red line is the global least-squares regression line. [Figure 38-4] Correlation between replicates in the mESC library experiment. Uncorrected C·G to T·A editing efficiency at each target site for each replicate. The dashed red line is the global least-squares regression line.

[0079] [Figure 39-1] Editing window of the TadCBE V106W variant in an mESC library editing experiment. The editing window is defined as a position within the protospacer where the average proportion of converted bases at that position is at least 30% of the average edit at the most edited position. C·G to T·A base edits are shown in blue. A·T to G·C base edits are shown in red. [Figure 39-2] Editing window of the TadCBE V106W variant in an mESC library editing experiment. The editing window is defined as a position within the protospacer where the average proportion of converted bases at that position is at least 30% of the average edit at the most edited position. C·G to T·A base edits are shown in blue. A·T to G·C base edits are shown in red.

[0080] [Figure 40A-40B]Effect of V106W on peak editing in mESC library experiments. Figure 40A shows the C·G to T·A editing efficiency by TadCBEd (with or without the V106W substitution) for each library member containing a cytosine at protospacer position 6. The red dashed line is the global least-squares regression line. Figure 40B shows the A·T to G·C editing efficiency by TadCBEd (with or without V106W) for each library member containing an adenine at protospacer position 6. The red dashed line is the global least-squares regression line.

[0081] [Figure 41A-1] Sequence motifs for context preference of TadCBE. Sequence motifs for base editing activity from regression on editing efficiency. Logo opacity is proportional to R for the holdout validation set. Plots are provided for C·G to T·A base editing (Figure 41A) and A·T to G·C base editing (Figure 41B). [Figure 41A-2] Sequence motifs for context preference of TadCBE. Sequence motifs for base editing activity from regression on editing efficiency. Logo opacity is proportional to R for the holdout validation set. Plots are provided for C·G to T·A base editing (Figure 41A) and A·T to G·C base editing (Figure 41B). [Figure 41B-1] Sequence motifs for context preference of TadCBE. Sequence motifs for base editing activity from regression on editing efficiency. Logo opacity is proportional to R for the holdout validation set. Plots are provided for C·G to T·A base editing (Figure 41A) and A·T to G·C base editing (Figure 41B). [Figure 41B-2] Sequence motifs for context preference of TadCBE. Sequence motifs for base editing activity from regression on editing efficiency. Logo opacity is proportional to R for the holdout validation set. Plots are provided for C·G to T·A base editing (Figure 41A) and A·T to G·C base editing (Figure 41B).

[0082] [Figure 42-1] Characterization of evolved deaminases with evolved eNme2-C Cas9 domains. The designated base editors, using the eNme2-C Cas9 nickase domain (PAM=N4CN) in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of six guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 42-2] Characterization of evolved deaminases with evolved eNme2-C Cas9 domains. The designated base editors, using the eNme2-C Cas9 nickase domain (PAM=N4CN) in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of six guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0083] [Figure 43-1]Indel and C·G to G·C editing by eNme2-C Cas9 variants at six genomic target sites. The designated base editors, using the eNme2-C Cas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of the six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 42. [Figure 43-2] Indel and C·G to G·C editing by eNme2-C Cas9 variants at six genomic target sites. The designated base editors, using the eNme2-C Cas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of the six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 42.

[0084] [Figure 44-1]Characterization of evolved deaminases with SaCas9 domains. The designated base editors, using the Sas9 nickase domain (PAM = NNGRRT) in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293-T cells along with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 44-2] Characterization of evolved deaminases with SaCas9 domains. The designated base editors, using the Sas9 nickase domain (PAM = NNGRRT) in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293-T cells along with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 44-3]Characterization of evolved deaminases with SaCas9 domains. The designated base editors, using the Sas9 nickase domain (PAM = NNGRRT) in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293-T cells along with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0085] [Figure 45-1] Indel and C·G to G·C editing by SaCas9 variants at nine genomic target sites. The designated base editors, using the SaCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 44. [Figure 45-2] Indel and C·G to G·C editing by SaCas9 variants at nine genomic target sites. The designated base editors, using the SaCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 44. [Figure 45-3] Indel and C·G to G·C editing by SaCas9 variants at nine genomic target sites. The designated base editors, using the SaCas9 nickase domain in the ABE8e or BE4max architecture with 2xUGI, were transfected into HEK293T cells along with each of six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is shown in Figure 44.

[0086] [Figure 46-1] Characterization of SpCas9-mediated TadDE in mammalian cells. The indicated base editors using the SpCas9 nickase domain (PAM=NGG) in the ABE8e or BE4max architecture with 2xUGI were transfected along with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 46-2]Characterization of SpCas9-mediated TadDE in mammalian cells. The indicated base editors using the SpCas9 nickase domain (PAM=NGG) in the ABE8e or BE4max architecture with 2xUGI were transfected along with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 46-3] Characterization of SpCas9-mediated TadDE in mammalian cells. The indicated base editors using the SpCas9 nickase domain (PAM=NGG) in the ABE8e or BE4max architecture with 2xUGI were transfected along with each of the nine guide RNAs targeting the protospacer indicated in each graph. Target cytosines are blue, target adenines are magenta, and PAM sequences are underlined. C·G to T·A base edits are indicated in shades of blue. A·T to G·C base edits are indicated in shades of magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0087] [Figure 47] Indel and C·G to G·C editing by SpCas9 variants at nine genomic target sites. The designated base editors, using the SpCas9 nickase domain in the ABE8e or BE4max architecture with 2×UGI, were transfected into HEK293T cells along with each of six guide RNAs targeting the protospacer indicated in each graph. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Dots represent individual values, and bars represent the mean ± SD of three independent biological replicates. Corresponding on-target data is in Figure 46.

[0088] [Figure 48] On-target editing of the V106W variant for T cell experiments targeting CXCR4 and CCR5. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into human T cells (n = 4 donors) along with two synthetic guide RNAs targeting (top) CXCR4 or (bottom) CCR5 at the indicated protospacers. Targeted cytosines are blue, targeted adenines are magenta, and PAM sequences are underlined. Three days later, genomic DNA was harvested from T cell lysates and analyzed by high-throughput sequencing. Gray boxes indicate the desired locations of stop codon incorporation in CXCR4 and CCR5. Targeted cytidines that result in TAG (CXCR4) and TAA (CCR5) stop codons by cytosine base editing are underlined. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0089] [Figure 49] C·G to G·C edits and indels for a cellular experiment targeting CXCR4 and CCR5 with the TadCBEe V106W variant. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into primary human T cells (n = 4 donors) along with two synthetic guide RNAs targeting (top) CXCR4 or (bottom) CCR5 at the indicated protospacers. Three days later, genomic DNA was harvested from T cell lysates and analyzed by high-throughput sequencing. C·G to G·C base edits are shown in shades of blue. Indels are shown in gray. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0090] [Figure 50-1]Cas-dependent off-target editing in T cell experiments targeting CXCR4 and CCR5 with the TadCBEe V106W variant. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into primary human T cells (n = 4 donors) along with two synthetic guide RNAs targeting (top) CXCR4 or (bottom) CCR5 at the indicated protospacers. Three days later, genomic DNA was harvested from T cell lysates, and known off-target sites were amplified using the primers in Table 4. C·G to T·A base edits are indicated in blue. A·T to G·C base edits are indicated in magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates. [Figure 50-2] Cas-dependent off-target editing in T cell experiments targeting CXCR4 and CCR5 with the TadCBEe V106W variant. mRNA encoding the indicated base editor or GFP as a negative control was electroporated into primary human T cells (n = 4 donors) along with two synthetic guide RNAs targeting (top) CXCR4 or (bottom) CCR5 at the indicated protospacers. Three days later, genomic DNA was harvested from T cell lysates, and known off-target sites were amplified using the primers in Table 4. C·G to T·A base edits are indicated in blue. A·T to G·C base edits are indicated in magenta. Points represent individual values, and bars represent the mean ± SD of three independent biological replicates.

[0091] [Figure 51A-51B] Predictive use of an active and selective cytosine base editor for stop codon incorporation at disease-relevant sites. Editing the residual A to G prevents correct stop codon incorporation (Figure 51A). Schematic of the evolution of cytosine base editors from the TadA dual base editor (TadA-DE) (Figure 51B). [Figure 51C]Diagram illustrating phage-assisted continuous evolution, or PACE (left), and a selection circuit (right) used according to some embodiments. In some embodiments, a continuous stream of E. coli host cells harboring the selection circuit and mutagenesis plasmid (red) is infected with a selection phage (SP) encoding a partial deaminase. In this particular embodiment of the selection circuit, phage propagation is linked to expression of gIII (P2), which can only be transcribed by active T7 RNA polymerase. In some embodiments, T7 RNA polymerase (P3) is fused to a C-terminal degron, and the deaminase must perform C-to-U editing to incorporate a stop codon before the degron, yielding active T7 RNA polymerase. In the event of phage infection, the full-length deaminase is completed using a split intein system (P1), and mutations can occur on the deaminase. Beneficial mutations lead to phage proliferation and concentration in the lagoon, while less adapted phages are unable to proliferate and are subsequently washed away by constant runoff (Figure 51C). [Figure 51D] Evolutionary trajectory of active and selective cytosine base editors from TadA-DE. Phage-assisted non-linear evolution (PANCE) was performed on TadA-DE until the phage titer increased, indicating that beneficial mutations had occurred, despite higher stringency from dilution factors and promoter strength. The resulting variants identified conserved mutations at position N46 of the deaminase. Therefore, an NNK library was constructed at position N46, and PANCE was performed on these variants. To further increase stringency, PACE was performed for >100 hours on the resulting variants from both PANCEs. The dilution factors are indicated on the right y-axis (Figure 51D). [Figure 51E] Cryo-EM structure of ABE8e with new conserved mutations labeled (PDB: 6VPC) (Figure 51E).

[0092] [Figure 52A]Genotypes from the PANCE lagoon (L1-L2) after PANCE (Figure 52B). [Fig. 52B-52C] Genotypes from PANCE lagoons (L1-L3) after PANCE using the NNK library at N46 (Figure 52B). Genotypes from PACE lagoons (L1) after PACE using the NNK library at N46 (Figure 52C). [Fig. 52D-52E] Genotypes at various time points from the PACE lagoon (L1) after PACE using the NNK library at N46 (Figure 52D). Genotypes at various time points from the PACE lagoon (L2) after PACE using the NNK library at N46 (Figure 52E). Selected sequences are shown in Figure 52A.

[0093] [Figure 53] Profiling the activity and sequence context specificity of TadCBE in E. coli. Bars indicate the average activity of CBE variants when tested against a library of substrates designed to contain a target base (A or C) at protospacer position 6 with the 5' and 3' bases varied as A, T, C, or G. Each point represents the percentage of sequencing reads containing the specified edit for a given sequence context. Points are colored according to the 5' context of the base (A, red; C, green; G, blue; T, yellow). Mutations in the newly evolved variants are listed relative to TadDE. TadDE = TadA8e R26G V28A A48 Y73S H96N. eTdCBEmax, CBET-1.52, TadCBEd, TadDE N46I Y73P, and TadDE N46C Y73P exhibited lower activity depending on the sequence context, whereas the evolved variants TadDE N46V Y73P and TadDE N46L Y73P exhibited over 80% editing regardless of the sequence context. [Figure 54A]Comparison of the evolved active and selective cytosine base editor with existing cytosine base editors in mammalian cells. HEK293T cells were transfected with guide RNAs targeting the three protospacers alongside existing cytosine base editors carrying SpCas9 nickase in the BE4max architecture. The TadDE N46 variant exhibits comparable on-target activity without residual A-to-G editing. Dots represent individual values ​​from independent biological replicates. PAM sequences are underlined. HEK293T site 2 is abbreviated as HEK2, and HEK293T site 4 is abbreviated as HEK4. HEK293T cells were transfected with guide RNAs targeting the two protospacers alongside existing cytosine base editors carrying eNme-Cas9 nickase in the BE4max architecture. TadDE N46 variants exhibit higher or comparable on-target activity without residual A to G editing. Dots represent individual values ​​from independent biological replicates. PAM sequences are underlined. [Figure 54B]Comparison of the evolved active and selective cytosine base editor with existing cytosine base editors in mammalian cells. HEK293T cells were transfected with guide RNAs targeting the three protospacers alongside existing cytosine base editors carrying SpCas9 nickase in the BE4max architecture. The TadDE N46 variant exhibits comparable on-target activity without residual A-to-G editing. Dots represent individual values ​​from independent biological replicates. PAM sequences are underlined. HEK293T site 2 is abbreviated as HEK2, and HEK293T site 4 is abbreviated as HEK4. HEK293T cells were transfected with guide RNAs targeting the two protospacers alongside existing cytosine base editors carrying eNme-Cas9 nickase in the BE4max architecture. TadDE N46 variants exhibit higher or comparable on-target activity without residual A to G editing. Dots represent individual values ​​from independent biological replicates. PAM sequences are underlined.

[0094] [Figure 55A-55B] Cas9-independent RNA off-target editing by TadCBE. Average Cas9-independent off-target editing at all cytosines for four orthogonal R-loops (SaR1-SaR4) generated by dead S. aureus Cas9. Mutations in newly evolved mutations are listed relative to TadDE. TadDE N46 variants show similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates (Figure 55A). Off-target RNA editing. TadDE N46 variants show similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates (Figure 55B).

[0095] [Figure 56]Stop codon incorporation at a therapeutically relevant locus by TadCBE in HEK293T. TadCBE was used to incorporate a stop codon into PCSK9, a therapeutic strategy being explored for lowering blood cholesterol. The gray box indicates the desired location of stop codon incorporation. Mutations in the newly evolved mutations are listed relative to TadDE. Editing of the residual A to G from TadCBEd results in the erasure of the stop codon, demonstrating that the lack of the residual A to G in the TadDE N46 variant is essential for stop codon incorporation. Dots represent individual values ​​from independent biological replicates. The PAM sequence is underlined.

[0096] [Figure 57-1] On-target and Cas-dependent editing of known off-target sites for HEK3. The TadDE N46 variant, along with an existing cytosine base editor with SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with a guide RNA targeting HEK3. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant shows similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates. [Figure 57-2] On-target and Cas-dependent editing of known off-target sites for HEK3. The TadDE N46 variant, along with an existing cytosine base editor with SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with a guide RNA targeting HEK3. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant shows similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates.

[0097] [Figure 58-1]On-target and Cas-dependent editing of a known off-target site in HEK4. The TadDE N46 variant, along with an existing cytosine base editor harboring the SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with a guide RNA targeting HEK4. Mutations in the newly evolved variants are listed relative to TadDE. The TadDE N46 variant exhibits similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates. [Figure 58-2] On-target and Cas-dependent editing of a known off-target site in HEK4. The TadDE N46 variant, along with an existing cytosine base editor harboring the SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with a guide RNA targeting HEK4. Mutations in the newly evolved variants are listed relative to TadDE. The TadDE N46 variant exhibits similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates.

[0098] [Figure 59] On-target and Cas-dependent editing of a known off-target site in EMX1. The TadDE N46 variant, along with an existing cytosine base editor harboring the SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with a guide RNA targeting EMX1. Mutations in the newly evolved variants are listed relative to TadDE. The TadDE N46 variant exhibits similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates.

[0099] [Figure 60]On-target and Cas-dependent editing of a known off-target site in BCL11a. The TadDE N46 variant, along with an existing cytosine base editor with SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with a guide RNA targeting BCL11a. Mutations in the newly evolved mutations are listed relative to TadDE. The TadDE N46 variant shows similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates.

[0100] [Figure 61-1] On-target editing in EMX1 correlated with Cas-independent off-target editing. The TadDE N46 variant, along with an existing cytosine base editor harboring the SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with an SpCas9 guide RNA targeting EMX1, along with an SaCas9 guide RNA. Mutations in the newly evolved mutations are listed relative to TadDE. Points represent individual values ​​from independent biological replicates. [Figure 61-2] On-target editing in EMX1 correlated with Cas-independent off-target editing. The TadDE N46 variant, along with an existing cytosine base editor harboring the SpCas9 nickase in the BE4max architecture, was transfected into HEK293T cells with an SpCas9 guide RNA targeting EMX1, along with an SaCas9 guide RNA. Mutations in the newly evolved mutations are listed relative to TadDE. Points represent individual values ​​from independent biological replicates.

[0101] [Figure 62]On-target editing in EMX1 correlated with RNA off-target editing. HEK293T cells were transfected with the TadDE N46 variant in conjunction with an existing cytosine base editor with SpCas9 nickase in the BE4max architecture in two plates. RNA was harvested from one plate 48 hours post-transfection, and genomic DNA was harvested from the other plate. The genomic DNA was analyzed for on-target editing of EMX1. Mutations in the newly evolved variants are listed relative to TadDE. The TadDE N46 variant exhibits similar off-target editing compared to TadCBEd. Points represent individual values ​​from independent biological replicates.

[0102] [Figure 63] Continuation of Figure 53. Each graph in Figure 53 is also represented in Figure 63; however, each data point (represented as a dot) in Figure 53 is shown as a bar in Figure 63. DETAILED DESCRIPTION OF THE INVENTION

[0103] definition As used in the specification and claims, the singular forms "a," "an," and "the" include singular and plural references unless the context clearly dictates otherwise. Thus, for example, reference to an "agent" includes a single agent as well as a plurality of such agents.

[0104] AAV "Adeno-associated virus" or "AAV" is a virus that infects humans and several other primate species. The wild-type AAV genome is a single-stranded deoxyribonucleic acid (ssDNA) of either positive or negative sense. The genome contains two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs), rep and cap, located between the ITRs. The rep ORF contains four overlapping genes encoding the Rep proteins required for the AAV life cycle. The cap ORF contains overlapping genes encoding the capsid proteins VP1, VP2, and VP3, which interact together to form the viral capsid. VP1, VP2, and VP3 are translated from a single mRNA transcript, which can be spliced ​​in two different ways. Either longer or shorter introns can be excised, resulting in the formation of two mRNA isoforms: approximately 2.3 kb and approximately 2.6 kb in length. The capsid is a supramolecular assembly of approximately 60 individual capsid protein subunits in a non-enveloped T-1 icosahedral lattice, capable of protecting the AAV genome. The mature capsid is composed of VP1, VP2, and VP3 (molecular masses of approximately 87, 73, and 62 kDa, respectively) in an approximately 1:1:10 ratio.

[0105] rAAV particles can include a nucleic acid vector (e.g., a recombinant genome) that can minimally include: (a) one or more heterologous nucleic acid regions comprising a sequence encoding a protein or polypeptide of interest (e.g., a split Cas9 or split nucleic acid base) or an RNA of interest (e.g., a gRNA), or one or more nucleic acid regions comprising a sequence encoding a Rep protein; and (b) one or more nucleic acid regions comprising inverted terminal repeat (ITR) sequences (e.g., wild-type ITR sequences or engineered ITR sequences) flanking the one or more nucleic acid regions (e.g., heterologous nucleic acid regions). In some embodiments, the nucleic acid vector is between 4 kb and 5 kb in size (e.g., 4.2-4.7 kb in size). In some embodiments, the nucleic acid vector further comprises a region encoding a Rep protein. In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, a double-stranded nucleic acid vector can be a self-complementary vector, e.g., containing a region of the nucleic acid vector that is complementary to another region of the nucleic acid vector and initiates the formation of a double-stranded state of the nucleic acid vector.

[0106] Deaminase The term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine (or adenine) deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine to inosine on deoxyribonucleic acid (DNA). In other embodiments, the deaminase is a cytidine (or cytosine) deaminase, which catalyzes the hydrolytic deamination of cytidine or cytosine.

[0107] The deaminase provided herein can be from any organism, such as bacteria.In some embodiments, deaminase or deaminase domain is the variant of naturally occurring deaminase from organism.In some embodiments, deaminase or deaminase domain does not exist in nature.For example, in some embodiments, deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical with naturally occurring deaminase.

[0108] As used herein, the term "adenosine deaminase" or "adenosine deaminase domain" refers to a protein or enzyme that catalyzes the deamination of adenosine (or adenine). The terms "adenosine" and "adenine" are used interchangeably for purposes of this disclosure. For example, for purposes of this disclosure, reference to an "adenine base editor" (ABE) refers to the same entity as an "adenosine base editor" (ABE). Similarly, for purposes of this disclosure, reference to an "adenine deaminase" refers to the same entity as an "adenosine deaminase." However, one of skill in the art will understand that "adenine" refers to a purine base, and "adenosine" refers to a larger nucleoside molecule that includes a purine base (adenine) and a sugar moiety (e.g., either ribose or deoxyribose). In certain embodiments, the present disclosure provides base editor fusion proteins comprising one or more adenosine deaminase domains. For example, the adenosine deaminase domain can comprise a heterodimer of a first adenosine deaminase domain and a second deaminase domain connected by a linker.The adenosine deaminase provided herein (for example, engineered adenosine deaminase or evolved adenosine deaminase) can be an enzyme that converts adenine (A) to inosine (I) on DNA or RNA.Such adenosine deaminase can lead to A:T to G:C base pair conversion.In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism.In some embodiments, the deaminase does not occur in nature. For example, in some embodiments, the deaminase is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase.

[0109] In some embodiments, the adenosine deaminase is derived from a bacterium, such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated E. coli TadA deaminase. For example, the truncated ecTadA may lack one or more N-terminal amino acids relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the ecTadA deaminase does not include an N-terminal methionine. Reference is made to U.S. Patent Publication No. 2018 / 0073012, published March 15, 2018, which is incorporated herein by reference.

[0110] As used herein, the term "cytidine deaminase" or "cytidine deaminase domain" refers to a protein or enzyme that catalyzes the deamination of cytidine or cytosine. The terms "cytidine" and "cytosine" are used interchangeably for purposes of this disclosure. For example, for purposes of this disclosure, reference to a "cytosine base editor" (CBE) refers to the same entity as a "cytosine base editor" (CBE). Similarly, for purposes of this disclosure, reference to a "cytidine deaminase" refers to the same entity as a "cytosine deaminase." However, those skilled in the art will understand that "cytosine" refers to a pyrimidine base, and "cytidine" refers to a larger nucleoside molecule that includes a pyrimidine base (cytosine) and a sugar moiety (e.g., either ribose or deoxyribose). Cytidine deaminase is encoded by the CDA gene and is an enzyme that catalyzes the removal of amine groups from cytidine (i.e., the base cytosine when attached to a ribose ring, i.e., a nucleoside called cytidine) to uridine (C to U) and deoxyuridine (C to U). A non-limiting example of a cytidine deaminase is APOBEC1 ("apolipoprotein B mRNA editing enzyme, catalytic polypeptide 1"). Another example is AID ("activation-induced cytidine deaminase"). Under standard Watson-Crick hydrogen bond pairing, the cytosine base hydrogen bonds to the guanine base. When cytidine is converted to uridine (or cytidine is converted to deoxyuridine), the uridine (or the uracil base of uridine) undergoes hydrogen bond pairing with the base adenine. Therefore, the conversion of "C" to uridine ("U") by cytidine deaminase will result in the insertion of "A" instead of "G" during cellular repair and / or replication processes. Because adenine "A" pairs with thymine "T," cytidine deaminase coordinates with DNA replication to convert C·G pairing to T·A pairing in double-stranded DNA molecules.

[0111] antisense strand In genetics, the "antisense" strand of a given fragment of double-stranded DNA is the template strand, which is considered to extend in the 3' to 5' direction. In contrast, the "sense" strand is the fragment of double-stranded DNA extending 5' to 3', which is complementary to the antisense or template strand of DNA extending 3' to 5'. In the case of a DNA fragment encoding a protein, the sense strand is the strand of DNA with the same sequence as the mRNA. This takes the antisense strand as its template during transcription and ultimately (typically, but not always) undergoes translation into protein. Thus, the antisense strand carries the RNA that is subsequently translated into protein, while the sense strand possesses a structure nearly identical to that of the mRNA. Note that for each fragment of dsDNA, there are potentially two sets: sense and antisense, depending on which direction it is read (since sense and antisense are relative to the perspective). Ultimately, it is the gene product, or mRNA, that defines which strand of a single piece of dsDNA is said to be sense or antisense.

[0112] BasesEdit "Base editing" refers to genome editing technologies that involve converting specific nucleic acid bases at targeted genomic loci to another. In certain embodiments, this can be achieved without requiring double-stranded DNA breaks (DSBs) or single-strand breaks (i.e., nicking). To date, other genome editing technologies, including CRISPR-based systems, begin with the introduction of a DSB at the locus of interest. Cellular DNA repair enzymes then repair the break, typically resulting in a random insertion or deletion of a base (indel) at the site of the DSB. However, when the introduction or correction of a point mutation at a target locus is desired rather than the stochastic disruption of the entire gene, these genome editing techniques are unsuitable because the correction rate is low (e.g., typically 0.1% to 5%) and the primary genome editing product is an indel. To increase the efficiency of gene correction without simultaneously introducing random indels, the inventors previously modified the CRISPR / Cas9 system to directly convert one DNA base to another without DSB formation. See Komor, AC, et al., Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference.

[0113] Base Editor As used herein, the term "base editor (BE)" refers to an agent comprising a polypeptide that can modify a base (e.g., A, T, C, G, or U) in a nucleic acid sequence (e.g., DNA or RNA) to convert one base to another (e.g., A to G, A to C, A to T, C to T, C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G). In some embodiments, a base editor can deaminate a base in a nucleic acid, such as a base in a DNA molecule. In the case of an adenine base editor, the base editor can deaminate adenine (A) on DNA. Such a base editor can include a nucleic acid-programmable DNA-binding protein (napDNAbp) fused to adenosine deaminase. Some base editors include CRISPR-mediated fusion proteins utilized in the base editing methods described herein. In some embodiments, the base editor comprises a nuclease-inactive Cas9 (dCas9) fused to a deaminase, which binds to nucleic acids via R-loop formation in a manner programmed by a guide RNA but does not cleave the nucleic acid. For example, as described in PCT / US2016 / 058344, published April 27, 2017 as WO 2017 / 070632 and incorporated herein by reference in its entirety, the dCas9 domain of the fusion protein can contain D10A and H840A mutations (which allow Cas9 to cleave only one strand of a nucleic acid duplex). The DNA cleavage domain of S. pyogenes Cas9 contains two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA (the "targeted strand," or the strand on which editing or deamination occurs), and the RuvC1 subdomain cleaves the non-complementary strand containing the PAM sequence (the "unedited strand").The RuvC1 mutant D10A generates nicks in the target strand, and the HNH mutant H840A generates nicks in the non-edited strand (see Jinek et al., Science, 337:816-821 (2012); Qi et al., Cell. 28;152(5):1173-83 (2013)).

[0114] As used herein, the terms cytidine, cytosine and deoxycytidine are all synonymous and refer to the cytidine that can be edited using CBE.Similarly, the terms adenosine, adenine and deoxyadenine are all synonymous and refer to the adenine that can be edited using ABE.Furthermore, the terms cytidine base editor, cytosine base editor and the like are synonymous.Similarly, the terms adenosine base editor, adenine base editor and the like are synonymous.

[0115] In some embodiments, a nucleobase editor is a polymer or polymer complex that primarily (e.g., greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 99%, greater than 99.9%, or 100%) effects the conversion of nucleobases on a polynucleic acid sequence to another nucleobase (i.e., transition or transversion) and uses a combination of 1) a nucleotide-, nucleoside-, or nucleobase-modifying enzyme; and 2) a nucleic acid-binding protein that can be programmed to bind to a specific nucleic acid sequence.

[0116] In some embodiments, the nucleobase editor comprises a DNA-binding domain (e.g., a programmable DNA-binding domain such as dCas9 or nCas9) that guides it to the target sequence. In some embodiments, the nucleobase editor comprises a nucleobase-modifying enzyme fused to the programmable DNA-binding domain (e.g., dCas9 or nCas9). A "nucleobase-modifying enzyme" is an enzyme that can modify a nucleobase, converting one nucleobase to another (e.g., a deaminase such as cytidine deaminase or adenosine deaminase). In some embodiments, the nucleobase editor can target a cytosine (C) base on a nucleic acid sequence and convert the C to a thymine (T) base. In some embodiments, editing of a C to a T is performed by a deaminase, e.g., a cytidine deaminase. Base editors that can perform other types of base conversions (e.g., adenosine (A) to guanine (G), C to G) are also contemplated.

[0117] In some embodiments, the nucleobase editor that converts C to T comprises a cytidine deaminase. "Cytidine deaminase" refers to an enzyme that catalyzes the chemical reaction "cytosine + HO → uracil + NH" or "5-methylcytosine + HO → thymine + NH." As may be apparent from the reaction formula, such a chemical reaction results in a C to U / T nucleobase change. In the context of a gene, such a nucleotide change or mutation can in turn lead to an amino acid change in the protein, which may affect the function of the protein. Examples include loss-of-function and gain-of-function. In some embodiments, the C to T nucleobase editor comprises dCas9 or nCas9 fused to a cytidine deaminase. In some embodiments, the cytidine deaminase domain is fused to the N-terminus of dCas9 or nCas9. In some embodiments, the nucleobase editor further comprises a domain that inhibits uracil glycosylase and / or a nuclear localization signal.Such nucleobase editors have been described in the art, for example, in Rees & Liu, Nat Rev Genet. 2018;19(12):770-788 and Koblan et al., Nat Biotechnol. 2018;36(9):843-846; and in U.S. Patent Publication No. 2018 / 0073012, published March 15, 2018, issued October 30, 2018 as U.S. Patent No. 10,113,163; U.S. Patent Publication No. 2017 / 0121693, published May 4, 2017, issued January 1, 2019 as U.S. Patent No. 10,167,457; International Publication No. WO 2017 / 070633, published April 27, 2017; and U.S. Patent Publication No. 2017 / 0121693, published May 4, 2017, published January 1, 2019 as U.S. Patent No. 10,167,457. 2015 / 0166980; U.S. Patent No. 9,840,699, issued December 12, 2017; U.S. Patent No. 10,077,453, issued September 18, 2018; International Publication No. WO 2019 / 023680, published January 31, 2019; International Publication No. WO 2018 / 0176009, published September 27, 2018; International Application No. PCT / US2019 / 033848, filed May 23, 2019; International Application No. PCT / US2019 / 47996, filed August 23, 2019; International Application No. PCT / US2019 / 049793, filed September 5, 2019; U.S. Provisional Application No. PCT / US2019 / 049793, filed April 17, 2019. 62 / 835,490; International Application No. PCT / US2019 / 61685, filed November 15, 2019; International Application No. PCT / US2019 / 57956, filed October 24, 2019; U.S. Provisional Application No. 62 / 858,958, filed June 7, 2019; and International Publication No. PCT / US2019 / 58678, filed October 29, 2019, the contents of each of which are incorporated herein by reference in their entirety.

[0118] In some embodiments, the nucleobase editor converts A to G. In some embodiments, the nucleobase editor comprises an adenosine deaminase. "Adenosine deaminase" is an enzyme involved in purine metabolism. It is required for the breakdown of adenosine from food and for the turnover of nucleic acids in tissues. Its primary function in humans is the development and maintenance of the immune system. Adenosine deaminase catalyzes the hydrolytic deamination of adenosine in the context of DNA (forming inosine, which base pairs as G). There are no known adenosine deaminases that act on DNA. Instead, known adenosine deaminase enzymes act only on RNA (tRNA or mRNA). Evolved adenosine deaminase enzymes that accept DNA and deaminate dA to deoxyinosine are described, by way of example, in PCT Application No. PCT / US2017 / 045381, filed August 3, 2017, published as WO 2018 / 027078, and PCT Application No. PCT / US2019 / 033848, published as WO 2019 / 226953, each of which is incorporated herein by reference.

[0119] Exemplary adenine base editors (ABEs) (or "adenosine base editors") and cytosine base editors (CBEs) (or "cytosine base editors") are also described in: Rees & Liu, Base editing: precision chemistry on the genome and transcriptome of living cells, Nat. Rev. Genet. 2018;19(12):770-788; and U.S. Patent Publication No. 2018 / 0073012, published March 15, 2018, issued October 30, 2018 as U.S. Patent No. 10,113,163; U.S. Patent Publication No. 2017 / 0121693, published May 4, 2017, issued January 1, 2019 as U.S. Patent No. 10,167,457; and International Publication No. WO 2017 / 0121693, published April 27, 2017. US Patent Publication No. 2017 / 070633, published June 18, 2015; US Patent Publication No. 2015 / 0166980, published December 12, 2017; and US Patent No. 10,077,453, published September 18, 2018, the contents of each of which are incorporated herein by reference in their entirety.

[0120] Cas9 The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nuclease containing a Cas9 domain or a fragment thereof (e.g., a protein containing the active or inactive DNA cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9). As used herein, a "Cas9 domain" refers to a protein fragment containing the active or inactive cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9. A "Cas9 protein" refers to a full-length Cas9 protein. Cas9 nuclease is sometimes also referred to as casn1 nuclease or CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection from mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain a complementary sequence, a spacer, to the ancestral mobile element and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endolytically cleaves linear or circular dsDNA targets complementary to the spacer. Target strands not complementary to the crRNA are first endolytically cut and then 3'-5' exolytically trimmed. In nature, DNA binding and cleavage typically require proteins and both RNAs. However, single-guide RNAs ("sgRNAs," or simply "gRNAs") can be engineered to incorporate aspects of both the crRNA and tracrRNA into a single RNA species.See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) on the CRISPR repeat sequence to help distinguish self from non-self.Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao See Jinek M., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in a variety of species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease contains one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0121] A Cas9 domain with inactivated nucleases may be interchangeably referred to as a "dCas9" protein (from nuclease "dead" Cas9). Methods for generating a Cas9 domain (or a fragment thereof) with an inactive DNA cleavage domain are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28;152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83(2013)). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology to Cas9 or fragments thereof.For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200). In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200). In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 200).

[0122] As used herein, the term "nCas9" or "Cas9 nickase" refers to a Cas9 or variant thereof that cleaves or nicks only one strand at a target cut site, thereby introducing a nick into a double-stranded DNA molecule rather than creating a double-stranded break. This can be achieved by introducing an appropriate mutation into wild-type Cas9 that inactivates one of Cas9's two endonuclease activities. Any suitable mutation that inactivates one Cas9 endonuclease activity while leaving the other intact is contemplated. For example, either the D10A or H840A mutation in the wild-type S. pyogenes Cas9 amino acid sequence, or the D10A mutation in the wild-type S. aureus Cas9 amino acid sequence, can be used to form nCas9.

[0123] cDNA The term "cDNA" refers to a strand of DNA copied from an RNA template to which the cDNA is complementary.

[0124] cyclic permutant As used herein, the term "circular permutation" refers to a protein or polypeptide (e.g., Cas9) that contains a circular permutation, which is a change in the structural organization of a protein that involves a change in the order in which amino acids appear in the amino acid sequence of the protein. In other words, a circular permutation is a protein with altered N- and C-termini compared to its wild-type counterpart. For example, the wild-type C-terminal half of a protein becomes a new N-terminal half. Circular permutation (or CP) is essentially a topological rearrangement of a protein's primary sequence, often connecting its N- and C-termini with a peptide linker and simultaneously splitting the sequence at a different position to create new adjacent N- and C-termini. The result is a protein structure with different connectivity but often the same overall similar three-dimensional (3D) shape, which may potentially include improved or altered characteristics, including reduced proteolytic susceptibility, improved catalytic activity, altered substrate or ligand binding, and / or improved thermal stability. Circular permutation proteins can occur in nature (e.g., concanavalin A and lectins). Additionally, circular permutations can occur as a result of post-translational modifications or can be engineered using recombinant techniques (see, e.g., Oakes et al., "Protein Engineering of Cas9 for enhanced function," Methods Enzymol, 2014, 546: 491-511 and Oakes et al., "CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification," Cell, January 10, 2019, 176: 254-267, each of which is incorporated herein by reference).

[0125] CRISPR CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of a preceding infection by a virus that has invaded a prokaryote. The snippets of DNA are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with various CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNAs, they effectively constitute the prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of the pre-crRNA requires a small trans-encoded RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endolytically cleaves linear or circular dsDNA targets complementary to the RNA. Specifically, target strands not complementary to the crRNA are first endolytically cut and then 3'-5' exolytically trimmed. In nature, DNA binding and cleavage typically require proteins and both RNAs. However, single-guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into the guide RNA of a single RNA species. For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) on the CRISPR repeat sequence to help distinguish self from non-self.CRISPR biology and Cas9 nuclease sequence and construction are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM, Gonzales See Jinek M., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.

[0126] In certain types of CRISPR systems (e.g., type II CRISPR systems), proper processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endolytically cleaves linear or circular nucleic acid targets complementary to the RNA. Specifically, target strands not complementary to the crRNA are first endolytically cut and then 3'-5 exolytically trimmed. In nature, DNA binding and cleavage typically require both proteins and RNAs. However, single-guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both the crRNA and tracrRNA into the guide RNA of a single RNA species.

[0127] Generally, a "CRISPR system" collectively refers to the transcripts and other elements involved in the expression or activity of CRISPR-associated ("Cas") genes, and includes sequences encoding Cas genes, tracr (transactivating CRISPR) sequences (e.g., tracrRNA or an active partial tracrRNA), tracr mate sequences (which, in the context of an endogenous CRISPR system, encompass "direct repeats" and partial direct repeats processed by the tracrRNA), guide sequences (also referred to as "spacers" in the context of an endogenous CRISPR system), or other sequences and transcripts from the CRISPR locus. The tracrRNA of the system is complementary (fully or partially) to the tracr mate sequence present on the guide RNA.

[0128] Degron The term "degron" or "degron domain" refers to a portion of a polypeptide that affects, controls, directs, or otherwise regulates the degradation rate of the polypeptide. Degrons can be highly variable and can include short amino acid sequences, structural motifs, and / or exposed amino acids. Degrons can also be located anywhere within a polypeptide (e.g., at the N-terminus, C-terminus, or internal position within the primary structure). The specific mechanism of polypeptide degradation regulated by a degron is not limited and can include ubiquitin-dependent degradation (i.e., degradation involving proteasome-based degradation) or ubiquitin-independent degradation. For example, the four-amino acid tail of NH3-EMLA-COOH (SEQ ID NO: 384), encoded by exon 8 of the SMN2 gene, functions as a degron and induces degradation of SMN2.

[0129] Effective dose As used herein, the term "effective amount" refers to an amount of a biologically active agent that is sufficient to elicit a desired biological response. For example, in some embodiments, an effective amount of a base editor can refer to an amount of the base editor that is sufficient to edit a target site nucleotide sequence, e.g., a genome. In some embodiments, an effective amount of a base editor provided herein, e.g., a base editor including a Cas9 nickase domain and a nucleobase-modifying domain (e.g., a deaminase domain), can refer to an amount of the base editor that is sufficient to induce editing of a target site that is specifically bound and edited by the base editor. In some embodiments, an effective amount of a base editor provided herein can refer to an amount of the base editor that is sufficient to induce editing having the following characteristics: >50% product purity, <5% indels in the region immediately surrounding the target sequence, and / or an editing window of 2-8 nucleotides. In other embodiments, an effective amount of a base editor can refer to an amount of the base editor that is sufficient to induce >45% product purity, <10% indels, a ratio of intended point mutations to indels that is at least 5:1, and / or an editing window of 2-10 nucleotides. As will be understood by one of skill in the art, the effective amount of an agent, e.g., a base editor, nuclease, deaminase, hybrid protein, protein-polynucleotide complex, or polynucleotide (e.g., gRNA), can vary depending on various factors, e.g., the desired biological response, e.g., the particular allele, genome, or target site to be edited, the target cell or tissue (i.e., the cell or tissue to be edited), and the agent to be used.

[0130] Off-target and on-target editing As used herein, the term "off-target editing" refers to the introduction of unintended modifications (e.g., deamination) to nucleotides (e.g., cytosine) in sequences outside the classical base editor binding window (i.e., from one protospacer position to another, typically 2-8 nucleotides in length). Off-target editing can result from weak or nonspecific binding of the gRNA sequence to the target sequence. Off-target editing can also result from intrinsic binding of the base editor's nucleotide-modifying domain (e.g., deaminase domain) to nucleobases in loci unrelated to the target sequence.

[0131] The term "Cas9-dependent off-target editing" refers to the introduction of unintended modifications resulting from weak or nonspecific binding of a Cas9-gRNA complex (e.g., a complex between a gRNA and a Cas9 domain of a base editor) to a nucleic acid site that has a fairly high sequence identity to the target sequence (e.g., greater than 60% relative thereto, or fewer than six mismatches thereto). In contrast, the term "Cas9-independent off-target editing" refers to the introduction of unintended modifications resulting from weak binding of a base editor (e.g., a nucleotide-modifying domain) to a nucleic acid site that does not have a high sequence identity to the target sequence (e.g., less than about 60% relative thereto, or six to eight or more mismatches thereto). These bindings occur independently of any hybridization between the Cas9-gRNA complex and the nucleic acid site of interest, and are therefore referred to as "Cas9-independent."

[0132] As used herein, the term "on-target editing" refers to the introduction of a intended modification (e.g., deamination) into a nucleotide (e.g., cytosine) on a target sequence, for example, using a base editor described herein.

[0133] As used herein, the terms "on-target editing frequency" and "on-target editing efficiency" refer to the number or percentage of intended base pairs that are edited. For example, if a base editor edits 10% of the base pairs that it is intended to target (e.g., in a cell or a population of cells), the base editor can be described as 10% efficient. Some aspects of editing efficiency encompass the modification (e.g., deamination) of specific nucleotides in DNA without generating a large number or percentage of insertions or deletions (i.e., indels). It is generally accepted that editing while generating less than 5% indels in the region immediately surrounding the target sequence (measured in the total target nucleotide substrate) corresponds to high editing efficiency. The generation of more than 20% indels is generally accepted as poor or low editing efficiency.

[0134] As used herein, the term " off-target editing frequency " refers to the number or rate of unintended base pairs that are edited.On-target and off-target editing frequency can be measured by the methods and assays described herein, taking into account the techniques known in the art, including high-throughput sequencing reads.High-throughput sequencing used herein involves the hybridization of nucleic acid primers (e.g., DNA primers) that have complementarity with the nucleic acid (e.g., DNA) region immediately upstream or downstream of the target sequence or off-target sequence in question.Since the DNA target sequence and Cas9-independent off-target sequence are known a priori in the method disclosed herein, nucleic acid primers that have sufficient complementarity with the upstream or downstream region of the target sequence and Cas9-independent off-target sequence in question can be designed using techniques known in the art, such as PhusionU PCR Kit (Life Technologies), Phusion HS II Kit (Life Technologies), and Illumina MiSeq Kit. Since many Cas9-dependent off-target sites have high sequence identity with the target site in question, nucleic acid primers with sufficient complementarity to the upstream or downstream region of the Cas9-dependent off-target site can also be designed using techniques and kits known in the art. These kits utilize polymerase chain reaction (PCR) amplification to produce amplicons as intermediate products. Target and off-target sequences can include genomic loci that further include protospacers and PAMs. Therefore, the term "amplicon" used herein can refer to a nucleic acid molecule that corresponds to the assembly of genomic loci, protospacers, and PAMs. The high-throughput sequencing technology used herein can also include Sanger sequencing and / or whole genome sequencing (WGS).Off-target effects of the disclosed base editors can be measured using the assays and methods disclosed in International Application No. PCT / US2020 / 624628, filed November 25, 2020, which is incorporated herein by reference.

[0135] functional equivalent The term "functional equivalent" refers to a second biomolecule that is functionally equivalent to a first biomolecule, but not necessarily structurally equivalent. For example, a "Cas9 equivalent" refers to a protein that has the same or substantially the same function as Cas9, but not necessarily the same amino acid sequence. In the context of this disclosure, the specification refers throughout to "protein X or a functional equivalent thereof." In this context, a "functional equivalent" of protein X encompasses any homolog, paralog, fragment, naturally occurring, engineered, circularly permuted, mutated, or synthetic version of protein X that has equivalent function.

[0136] Fusion proteins The term "fusion protein" as used herein refers to a hybrid polypeptide containing protein domains from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) portion of the fusion protein, thus forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. The protein can contain different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9, which guides the protein to bind to the target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. Another example includes Cas9 fused to adenosine deaminase or its equivalent. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced via recombinant protein expression and purification. This is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0137] Guide nucleic acid The terms "guide nucleic acid" or "napDNAbp-programming nucleic acid molecule" or, equivalently, "guide sequence" refer to one or more nucleic acid molecules that bind to a napDNAbp protein and direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to one or more nucleic acid molecules (or a portion or region thereof) bound to the protein, thereby causing the napDNAbp protein to bind to the nucleotide sequence at the specific target site. A non-limiting example is a guide RNA for a Cas protein of a CRISPR-Cas genome editing system.

[0138] Guide RNA is a specific type of guide nucleic acid that is most commonly associated with the Cas protein of CRISPR-Cas9, binds to Cas9, and guides the Cas9 protein to a specific sequence on a DNA molecule that contains the complementarity of the guide RNA's protospacer sequence. As used herein, "guide RNA" refers to a synthetic fusion of endogenous bacterial crRNA and tracrRNA, which provides both the targeting specificity and backbone and / or binding ability of the Cas9 nuclease to the target DNA. This synthetic fusion does not occur in nature and is also commonly referred to as sgRNA. However, this term also encompasses equivalent guide nucleic acid molecules that bind to Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and otherwise program the Cas9 equivalent to localize to a specific target nucleotide sequence. Cas9 equivalents include other napDNAbp from any type of CRISPR system (e.g., type II, V, VI), and may include Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Additional Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein. Additionally, methods for designing suitable guide RNA sequences are provided herein.

[0139] Guide RNA ("gRNA") As used herein, the term "guide RNA" generally refers to a specific type of guide nucleic acid that associates with Cas9 and directs the Cas9 protein to a specific sequence on a DNA molecule that contains a complementary sequence to the protospacer sequence of the guide RNA. However, this term also encompasses equivalent guide nucleic acid molecules that associate with Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and otherwise program the Cas9 equivalent to localize to a specific target nucleotide sequence. Cas9 equivalents include other napDNAbp from any type of CRISPR system (e.g., type II, V, VI), and may include Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein.

[0140] A guide RNA may contain various structural elements, including, but not limited to: (a) a spacer sequence, which is a sequence (approximately 20 nt in length) on the guide RNA that binds to the complementary strand of the target DNA (and has the same sequence as the DNA protospacer), and (b) a gRNA core (or gRNA backbone or main chain sequence), which refers to the sequence within the gRNA responsible for Cas9 binding, excluding the approximately 20 bp spacer sequence used to guide Cas9 to the target DNA.

[0141] As used herein, "guide RNA target sequence" refers to approximately 20 nucleotides that are complementary to the protospacer sequence on the PAM strand. The target sequence is the sequence that anneals to or is targeted by the spacer sequence of the guide RNA. The spacer sequence of the guide RNA and the protospacer have the same sequence (except that the spacer sequence is RNA and the protospacer is DNA).

[0142] As used herein, "guide RNA backbone sequence" refers to the sequence within the gRNA responsible for Cas9 binding. It does not include the 20 bp spacer / targeting sequence used to guide Cas9 to the target DNA.

[0143] host cell As used herein, the term "host cell" refers to a cell that can host, replicate, and transfer a phage vector useful in the continuous evolution process provided herein. In embodiments where the vector is a viral vector, a suitable host cell is one that can infect, replicate, and package the viral vector into viral particles that can infect new host cells. A cell can host a viral vector if it supports the expression of the viral vector's genes, replication of the viral genome, and / or production of viral particles. One criterion for determining whether a cell is a suitable host cell for a given viral vector is whether the cell can support the viral life cycle of the wild-type viral genome from which the viral vector is derived. For example, if the viral vector is a modified M13 phage genome provided in some embodiments described herein, a suitable host cell would be any cell that can support the wild-type M13 phage life cycle. Suitable host cells for viral vectors useful in continuous evolution processes are well known to those of skill in the art, and the present disclosure is not limited in this respect. In some embodiments, the viral vector is a phage and the host cell is a bacterial cell. In some embodiments, the host cell is an E. coli cell. Suitable E. coli host strains will be apparent to those of skill in the art, including, but not limited to, New England Biolabs (NEB) Turbo, Top10F', DH12S, ER2738, ER2267, and XL1-Blue MRF'. These strain designations are art-recognized, and the genotypes of these strains are well characterized. It should be understood that the above strains are exemplary only, and the invention is not limited in this respect. The term "fresh," which is used interchangeably herein with the terms "uninfected" or "uninfected" in the context of host cells, refers to host cells that have not been infected with a viral vector containing the gene of interest used in the continuous evolution process provided herein.However, the new host cells may already be infected with a viral vector unrelated to the vector to be evolved, or with a vector of the same or similar type but not carrying the gene of interest.

[0144] In some embodiments, the host cell is a prokaryotic cell, such as a bacterial cell. In some embodiments, the host cell is an E. coli cell. In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, an insect cell, or a mammalian cell. The type of host cell will, of course, depend on the viral vector employed, and suitable host cell / vector combinations will be readily apparent to one skilled in the art.

[0145] Inteins and split inteins As used herein, the term "intein" refers to a self-processing polypeptide domain found in organisms from all domains of life. Inteins (intervening proteins) carry out a unique self-processing event known as protein splicing, in which they excise themselves from a larger precursor polypeptide through the cleavage of two peptide bonds and, in the process, ligate flanking extein (external protein) sequences through the formation of new peptide bonds. Because intein genes are found embedded in-frame within other protein-coding genes, this rearrangement occurs post-translationally (or potentially co-translationally). Furthermore, intein-mediated protein splicing is spontaneous; it requires no external factors or energy sources, only the folding of the intein domain. This process is also known as cis-protein splicing, in contrast to the natural process of trans-protein splicing by "split inteins."

[0146] Split inteins are a subcategory of inteins. Unlike more common continuous inteins, split inteins are transcribed and translated as two separate polypeptides, the N-intein and the C-intein, each fused to an extein. Upon translation, the intein fragments spontaneously and noncovalently assemble into the classical intein structure to carry out protein splicing in trans.

[0147] Inteins and split inteins are the protein equivalents of self-splicing RNA introns (see Perler et al., Nucleic Acids Res. 22:1125-1127 (1994)), which catalyze their own excision from precursor proteins with concomitant fusion of flanking protein sequences known as exteins (reviewed in Perler et al., Curr. Opin. Chem. Biol. 1:292-299 (1997); Perler, F. B Cell 92(1):1-4 (1998); Xu et al., EMBO J. 15(19):5146-5153 (1996)).

[0148] As used herein, the term "protein splicing" refers to the process by which an internal region (intein) of a precursor protein is excised and flanking regions (extein) of the protein are ligated to form a mature protein. This natural process has been observed in numerous proteins from both prokaryotes and eukaryotes (Perler, FB, Xu, MQ, Paulus, H. Current Opinion in Chemical Biology 1997, 1, 292-299; Perler, FB, Nucleic Acids Research 1999, 27, 346-347). Intein units contain the components required to catalyze protein splicing and often contain an endonuclease domain involved in intein mobility (Perler, FB, Davis, EO, Dean, GE, Gimble, FS, Jack, WE, Neff, N., Noren, CJ, Thomer, J., Belfort, M. Nucleic Acids Research 1994, 22, 1127-1127). However, the resulting proteins are linked and not expressed as separate proteins. Protein splicing can also occur in trans. Here, split inteins expressed on separate polypeptides spontaneously combine to form a single intein, which then undergoes the protein splicing process to join the separate proteins.

[0149] Elucidation of the mechanism of protein splicing has led to a number of intein-based applications (Comb, et al., US Pat. No. 5,496,714; Comb, et al., US Pat. No. 5,834,247; Camarero and Muir, J. Amer. Chem. Soc., 121:5597-5598 (1999); Chong, et al., Gene, 192:271-281 (1997); Chong, et al., Nucleic Acids Res., 26:5109-5115 (1998); Chong, et al., J. Biol. Chem., 273:10567-10577 (1998); Cotton, et al. J. Am. Chem. Soc., 121:1100-1101 (1999);Evans, et al., J. Biol. Chem., 274:18359-18363 (1999);Evans, et al., J. Biol. Chem., 274:3923-3926 (1999);Evans, et al., Protein Sci., 7:2256-2264 (1998);Evans, et al., J. Biol. Chem., 275:9091-9094 (2000);Iwai and Pluckthun, FEBS Lett. 459:166-172 (1999);Mathys, et al., Gene, 231:1-13 (1999);Mills, et al., Proc. Natl. Acad. Sci. USA 95:3543-3548 (1998);Muir, et al., Proc. Natl. Acad. Sci. USA 95:6705-6710 (1998);Otomo, et al., Biochemistry 38:16040-16044 (1999);Otomo, et al., J. Biolmol. NMR 14:105-114 (1999);Scott, et al., Proc. Natl. Acad. Sci. USA 96:13638-13643 (1999); Severinov and Muir, J. Biol. Chem., 273:16205-16209 (1998);Shingledecker, et al., Gene, 207:187-195 (1998);Southworth, et al., EMBO J. 17:918-926 (1998);Southworth, et al., Biotechniques, 27:110-120 (1999);Wood, et al., Nat. Biotechnol., 17:889-892 (1999);Wu, et al., Proc. Natl. Acad. Sci. USA 95:9226-9231 (1998a);Wu, et al., Biochim Biophys Acta 1387:422-432 (1998b);Xu, et al., Proc. Natl. Acad. Sci. USA 96:388-393 (1999); Yamazaki, et al., J. Am. Chem. Soc., 120:5591-5592 (1998)). Each reference is incorporated herein by reference.

[0150] Linker As used herein, the term "linker" refers to a chemical group or molecule that connects two molecules or domains, such as dCas9 and deaminase. Typically, the linker is located between or flanked by two groups, molecules, or other domains, and is connected to each other via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical domain. Chemical groups include, but are not limited to, disulfide, hydrazone, and azide domains. In some embodiments, the linker is 5-100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the linker is an XTEN linker. In some embodiments, the linker is a 32-amino acid linker. In other embodiments, the linker is a 30, 31, 33, or 34 amino acid linker.

[0151] mutation The term "mutation" as used herein refers to the substitution of a residue in a sequence, for example, a nucleic acid or amino acid sequence, with another residue; the deletion or insertion of one or more residues in a sequence; or the substitution of a residue in a genomic sequence in a target to be corrected. Mutations are typically described herein by identifying the original residue, followed by the position of the residue in the sequence and the identity of the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art, and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). Mutations can include various categories, such as single nucleotide polymorphisms, microduplication regions, indels, and inversions, and are not intended to be limiting in any way. Mutations can include "loss-of-function" mutations, which are mutations that reduce or eliminate protein activity. Most loss-of-function mutations are recessive. This is because in heterozygotes, the second chromosomal copy carries a non-mutated version of the gene encoding a fully functional protein, compensating for the effect of the mutation. There are some exceptions where loss-of-function mutations are dominant. One example is haploinsufficiency, where an organism cannot tolerate the approximately 50% reduction in protein activity that heterozygotes suffer. This explains a small number of genetic disorders in humans, including Marfan syndrome, which results from mutations in the gene for a connective tissue protein called fibrillin. Mutations also encompass "gain-of-function" mutations, which confer abnormal activity to a protein or cell that is otherwise not present under normal conditions. Many gain-of-function mutations are in regulatory sequences rather than coding regions and therefore can have several consequences. For example, a mutation can lead to one or more genes being expressed in the wrong tissues, and these tissues acquire a function that they normally lack.Alternatively, the mutations may lead to overexpression of one or more genes involved in cell cycle control, thus leading to uncontrolled cell division and therefore cancer. By their nature, gain-of-function mutations are usually dominant.

[0152] napDNAbp The term "napDNAbp," meaning "nucleic acid programmable DNA binding protein," refers to any protein that can bind to (e.g., form a complex with) one or more nucleic acid molecules (i.e., which may be broadly referred to as "nucleic acid molecules that program the napDNAbp," and which include, e.g., guide RNAs in the case of a Cas system) that direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to one or more nucleic acid molecules (or a portion or region thereof) bound to the protein, thereby causing the protein to bind to the nucleotide sequence at the specific target site. The term napDNAbp encompasses CRISPR-Cas9 proteins and Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and includes Cas9 equivalents from any type of CRISPR system (e.g., Type II, V, VI), and can include Cpf1 (Type V CRISPR-Cas system), C2c1 (Type V CRISPR-Cas system), C2c2 (Type VI CRISPR-Cas system), C2c3 (Type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14, Argonaute, and nCas9. Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353 (6299), the contents of which are incorporated herein by reference. However, nucleic acid programmable DNA binding proteins (napDNAbp) that can be used in connection with the present invention are not limited to CRISPR-Cas systems. The present invention encompasses any such programmable protein, such as the Argonaute protein (NgAgo) from Natronobacterium gregoryi, which can also be used for DNA-guided genome editing.The NgAgo-guided DNA system does not require a PAM sequence or a guide RNA molecule. This means that genome editing can be performed on any genome sequence simply by expressing a generic NgAgo protein and introducing synthetic oligonucleotides. See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7):768-73, which is incorporated herein by reference.

[0153] In some embodiments, napDNAbp is an RNA-programmable nuclease, and when complexed with RNA, it can be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) is referred to as a guide RNA (gRNA). A gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), although "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with a target nucleic acid (e.g., and directs binding of the Cas9 (or equivalent) complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is homologous to tracrRNA as depicted in Figure 1E of Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Patent No. 9,340,799, entitled "mRNA-Sensing Switchable gRNAs," and International Patent Application No. PCT / US2014 / 054247, filed September 6, 2013, published as WO 2015 / 035136, entitled "Delivery System For Functional Nucleases," the entire contents of each of which are incorporated herein by reference. In some embodiments, a gRNA comprises two or more of domains (1) and (2) and may be referred to as an "extended gRNA." For example, as described herein, an extended gRNA may bind, e.g., to two or more Cas9 proteins and bind to a target nucleic acid at two or more distinct regions.The gRNA comprises a nucleotide sequence complementary to a target site, which mediates binding of the nuclease / RNA complex to said target site and provides sequence specificity for the nuclease:RNA complex. In some embodiments, the RNA-programmable nuclease is a (CRISPR-related system) Cas9 endonuclease, e.g., Cas9 (Csn1) from Streptococcus pyogenes (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes," Ferretti JJ et al., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E. et al., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Jinek M. et al., Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference.

[0154] napDNAbp nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, and these proteins can in principle be targeted to any sequence specified by a guide RNA. Methods of using napDNAbp nucleases such as Cas9 for site-specific cleavage (e.g., to modify genomes) are known in the art (see, e.g., Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature Biotechnology 31, 227-229 (2013); Jinek, M. et al. RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013); Jiang, W. et al. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).

[0155] Nickase The term "nickase" refers to a napDNAbp that has only a single nuclease activity that cuts only one strand of target DNA rather than both strands (e.g., one of the two nuclease domains is inactivated). Therefore, a nickase-type napDNAbp does not leave a double-strand break. In some embodiments, any of the disclosed base editors or vectors can include an S. pyogenes Cas9 nickase containing a D10A mutation (SpCas9n or nCas9). In some embodiments, any of the disclosed base editors can include an Nme2Cas9 nickase containing a D16A mutation (Nme2Cas9n).

[0156] Nuclear localization signal A nuclear localization signal or sequence (NLS) is an amino acid sequence that tags, recommends, or otherwise marks a protein for import into the cell nucleus via nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface. Different nuclear-localized proteins may share the same NLS. NLSs have the opposite function of nuclear export signals (NESs), which target proteins outside the nucleus. Therefore, a single nuclear localization signal can direct the entity to which it is attached into the nucleus of the cell. Such sequences can be of any size and composition, e.g., more than 25, 25, 15, 12, 10, 8, 7, 6, 5, or 4 amino acids, but preferably contain at least 4-8 amino acids known to function as nuclear localization signals (NLSs).

[0157] nucleic acid molecule As used herein, the term "nucleic acid molecule" refers to RNA and single- and / or double-stranded DNA. Nucleic acid molecules can occur naturally, for example, in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. Alternatively, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, or DNA / RNA hybrids, or can contain non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms encompass nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced and optionally purified using recombinant expression systems, chemically synthesized, etc. Where appropriate, for example, in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as analogs with chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise indicated. In some embodiments, nucleic acids are selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, adenosine, deoxythymidine, deoxyguanosine, and cytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, C5-methylcytidine, C5-methyl- ... adenosine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, inosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); intercalating bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0158] PACE As used herein, the term "phage-assisted continuous evolution (PACE)" refers to continuous evolution that employs phages as viral vectors. The general concept of PACE technology is described, for example, in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, published March 11, 2010 as WO 2010 / 028347; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, published June 28, 2012 as WO 2012 / 088381; US ​​Patent No. 9,023,594, issued May 5, 2015; International PCT Application No. PCT / US2015 / 012022, filed January 20, 2015, published September 11, 2015 as WO 2015 / 134121; and US Patent No. and International PCT Application PCT / US2016 / 027795, filed April 15, 2016, published October 20, 2016 as PCT International Application No. PCT / US2016 / 027795, the entire contents of each of which are incorporated herein by reference.

[0159] promoter The term "promoter" is art-recognized and refers to a nucleic acid molecule having a sequence that can be recognized by a cell's transcriptional machinery and initiate transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is active only in the presence of specific conditions. For example, a conditional promoter may be active only in the presence of a specific protein that connects a protein associated with a control element on the promoter to the basal transcription machinery, or in the absence of an inhibitory molecule. A subclass of conditionally active promoters is inducible promoters, which require the presence of a small molecule "inducer" for activity. Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. A variety of constitutive, conditional, and inducible promoters are well known to those of skill in the art, and one of skill in the art will be able to ascertain a variety of such promoters useful in carrying out the present invention. This is not limiting in this regard. In various embodiments, the present disclosure provides vectors having a suitable promoter for driving expression of a nucleic acid sequence encoding a fusion protein (or one or more individual components thereof).

[0160] Product Purity As used herein, the term "product purity" refers to the percentage of the desired product relative to the total product of a base editing reaction. For example, the product purity of CBE can be measured as the percentage of total sequencing reads in which the target C is edited to T in the nucleic acid in question (reads in which the target C is converted to a different base). Product purity encompasses the absence of indels and the desired product of base conversion.

[0161] The term "R-loop" refers to a triplex structure in which two strands of double-stranded DNA are separated by a stretch of nucleotides and held apart by a single-stranded RNA molecule (e.g., gRNA). R-loop formation can be induced by hybridization of a gRNA with complementarity to DNA bound to a napDNAbp protein or domain (e.g., Cas9). Two R-loops are said to be "orthogonal" when the mechanisms that generate their formation (e.g., napDNAbp-gRNA complex) function independently of each other.

[0162] Protospacer As used herein, the term "protospacer" refers to a sequence (approximately 20 bp) on DNA adjacent to the PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one of its strands, i.e., the "target strand" versus the "non-target strand" of the target DNA sequence). For Cas9 to function, it also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease from S. pyogenes recognizes a PAM sequence of NGG, which is found directly downstream of the target sequence on genomic DNA on the non-target strand. Those skilled in the art will understand that prior art literature sometimes refers to the approximately 20-nt target-specific guide sequence on the guide RNA itself as a "protospacer" rather than referring to it as a "spacer." Therefore, in some cases, the term "protospacer" as used herein can be used interchangeably with the term "spacer." The descriptive context surrounding the appearance of either "protospacer" or "spacer" will help inform the reader as to whether the term refers to a gRNA or a DNA target.

[0163] Protospacer adjacent motif (PAM) As used herein, the term "protospacer adjacent sequence" or "PAM" refers to a DNA sequence of approximately 2-6 base pairs that is the key targeting component of the Cas9 nuclease. Typically, the PAM sequence is located on either strand, downstream in the 5' to 3' direction from the Cas9 cut site. The classic PAM sequence (i.e., the PAM sequence associated with Streptococcus pyogenes Cas9 nuclease or SpCas9) is 5'-NGG-3', where "N" is any nucleobase followed by two guanine ("G") nucleobases. Different PAM sequences may be associated with different Cas9 nucleases or equivalent proteins from different organisms. Additionally, any given Cas9 nuclease, e.g., SpCas9, can be modified to modulate the nuclease's PAM specificity by allowing the nuclease to recognize alternative PAM sequences.

[0164] For example, referring to the classical SpCas9 amino acid sequence of SEQ ID NO: 200, the PAM sequence can be modified by introducing one or more mutations, including: (a) a D1135V, R1335Q, and T1337R "VQR variant" that modulates PAM specificity to NGAN or NGNG; (b) a D1135E, R1335Q, and T1337R "EQR variant" that modulates PAM specificity to NGAG; and (c) a D1135V, G1218R, R1335E, and T1337R "VRER variant" that modulates PAM specificity to NGCG. In addition, the D1135E variant of classical SpCas9 still recognizes NGG, but it is more selective compared to the wild-type SpCas9 protein.

[0165] It will also be understood that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) can have different PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. Additionally, Cas9 from Neisseria meningitis (NmCas) recognizes NNNNGATT. In another example, Cas9 from Streptococcus thermophilis (StCas9) recognizes NNAGAAW. In yet another example, Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These are examples and are not meant to be limiting. It will further be understood that non-SpCas9s bind to a variety of PAM sequences, making them useful when a suitable SpCas9 PAM sequence is not present at the desired target cut site. Furthermore, non-SpCas9s may have other characteristics that make them more useful than SpCas9s. For example, Cas9 from Staphylococcus aureus (SaCas9) is approximately 1 kilobase smaller than SpCas9. Therefore, it can be packaged into adeno-associated virus (AAV). Further reference may be made to Shah et al., "Protospacer recognition motifs: mixed identities and functional diversity," RNA Biology, 10(5): 891-899, which is incorporated herein by reference.

[0166] Sense strand In genetics, a "sense" strand is a segment of double-stranded DNA that extends 5' to 3' and is complementary to the antisense or template strand of DNA, which extends 3' to 5'. In the case of a DNA segment encoding a protein, the sense strand is the strand of DNA with the same sequence as the mRNA. This takes the antisense strand as its template during transcription and ultimately (typically, but not always) undergoes translation into protein. Thus, the antisense strand carries the RNA that is subsequently translated into protein, while the sense strand possesses a structure nearly identical to that of the mRNA. Note that for each segment of dsDNA, there will potentially be two sets: sense and antisense, depending on which direction it is read (since sense and antisense are relative to the perspective). Ultimately, it is the gene product or mRNA that defines which strand of a segment of dsDNA is said to be sense or antisense.

[0167] Spacer sequence The term "spacer sequence" as used herein with respect to a guide RNA refers to a portion of the guide RNA of approximately 20 nucleotides that contains a nucleotide sequence complementary to a protospacer sequence on the target DNA sequence. The spacer sequence anneals to the protospacer sequence to form an ssRNA / ssDNA hybrid structure at the target site and a corresponding R-loop ssDNA structure on the endogenous DNA strand that is complementary to the protospacer sequence.

[0168] subject As used herein, the term "subject" refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject can be of either sex and at any stage of development. In some embodiments, the subject is a plant.

[0169] target site The term "target site" refers to a sequence within a nucleic acid molecule that is edited by a fusion protein (e.g., a dCas9-deaminase fusion protein provided herein). Target site also refers to a sequence within a nucleic acid molecule to which a complex of a fusion protein and a gRNA binds.

[0170] transcription terminator A "transcription terminator" is a nucleic acid sequence that causes transcription to stop. A transcription terminator can be unidirectional or bidirectional. It consists of a DNA sequence involved in the specific termination of an RNA transcript by an RNA polymerase. A transcription terminator sequence prevents transcriptional activation of a downstream nucleic acid sequence by an upstream promoter. A transcription terminator may be necessary in vivo to achieve a desired expression level or to avoid transcription of certain sequences. A transcription terminator is considered "operably linked" to a nucleotide sequence when it is capable of terminating transcription of the sequence to which it is linked.

[0171] The most commonly used type of terminator is a forward terminator. When placed downstream of a nucleic acid sequence that is normally transcribed, a forward transcription terminator will cause transcription to cease. In some embodiments, a bidirectional transcription terminator is provided, which normally causes transcription to terminate on both the forward and reverse strands. In some embodiments, a reverse transcription terminator is provided, which normally terminates transcription only on the reverse strand.

[0172] In prokaryotic systems, terminators are typically classified into two categories: (1) rho-independent terminators and (2) rho-dependent terminators. Rho-independent terminators are generally composed of palindromic sequences that form a GC base pair-rich stem-loop followed by several T bases. Without wishing to be bound by theory, the conventional model of transcription termination is that the stem-loop causes the RNA polymerase to pause, and transcription of the poly(A) tail causes the RNA:DNA duplex to unwind and dissociate from the RNA polymerase.

[0173] In eukaryotic systems, terminator regions can contain specific DNA sequences that expose polyadenylation sites, allowing site-specific cleavage of nascent transcripts. This signals specialized endogenous polymerases to add a stretch of approximately 200 A residues (polyA) to the 3' end of the transcript. RNA molecules modified with this polyA tail appear to be more stable and are translated more efficiently. Therefore, in some embodiments involving eukaryotes, terminators can contain signals for RNA cleavage. In some embodiments, terminator signals promote polyadenylation of the message. Terminator and / or polyadenylation site elements can serve to enhance output nucleic acid levels and / or minimize readthrough between nucleic acids.

[0174] Terminators for use in accordance with the present disclosure include any transcription terminator described herein or known to those skilled in the art. Examples of terminators include, without limitation, gene termination sequences, such as the bovine growth hormone terminator, and viral termination sequences, such as the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA, and arcA terminators. In some embodiments, the termination signal may be a sequence that cannot be transcribed or translated, such as one resulting from sequence truncation.

[0175] Transitions As used herein, "transition" refers to the interchanging of purine nucleobases (A⇔G) or pyrimidine nucleobases (C⇔T). This class of interchanging involves nucleobases of similar shape. The compositions and methods disclosed herein can induce one or more transitions on a target DNA molecule. The compositions and methods disclosed herein can also induce both transitions and transversions in the same target DNA molecule. These changes involve A⇔G, G⇔A, C⇔T, or T⇔C. In the context of double-stranded DNA with Watson-Crick paired nucleobases, transversion refers to the following base pair exchanges: A:T⇔G:C, G:G⇔A:T, C:G⇔T:A, or T:A⇔C:G. The compositions and methods disclosed herein can induce one or more transitions on a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing other nucleotide changes, including both transitions and transversions, as well as deletions and insertions, in the same target DNA molecule.

[0176] treatment The terms "treatment," "treat," and "treating," as described herein, refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof. As used herein, the terms "treatment," "treat," and "treating," as described herein, refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay their recurrence.

[0177] Uracil glycosylase inhibitors As used herein, the term "uracil glycosylase inhibitor" or "UGI" refers to a protein capable of inhibiting the uracil DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or the UGI set forth in SEQ ID NO: 272. In some embodiments, the UGI proteins provided herein encompass fragments of UGI and proteins homologous to UGI or UGI fragments. For example, in some embodiments, the UGI domain comprises a fragment of the amino acid sequence set forth in SEQ ID NO: 272. In some embodiments, the UGI fragment comprises an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence set forth in SEQ ID NO: 272. In some embodiments, UGI comprises an amino acid sequence homologous to the amino acid sequence set forth in SEQ ID NO: 272, or an amino acid sequence homologous to a fragment of the amino acid sequence set forth in SEQ ID NO: 272. In some embodiments, a protein comprising UGI or a fragment of UGI, or a homolog of UGI or a UGI fragment, is referred to as a "UGI variant." UGI variants share homology to UGI or a fragment thereof. For example, UGI variants are at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type UGI or the UGI set forth in SEQ ID NO: 272.In some embodiments, the UGI variant comprises a fragment of UGI, such that the fragment is at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to the corresponding fragment of wild-type UGI or UGI set forth in SEQ ID NO: 272. In some embodiments, UGI comprises the following amino acid sequence: MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 272) (P14739|UNGI_BPPB2 uracil DNA glycosylase inhibitor).

[0178] variant As used herein, the term "variant" refers to a protein that has characteristics that deviate from those found in nature, while retaining at least one of its functional (i.e., binding, interaction, or enzymatic) ability and / or therapeutic properties. A "variant" is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to a wild-type protein. For example, a Cas9 variant may include a Cas9 with one or more amino acid residue changes compared to the wild-type Cas9 amino acid sequence. As another example, a deaminase variant may include a deaminase with one or more amino acid residue changes compared to the wild-type deaminase amino acid sequence, for example, after reconstruction of the deaminase's ancestral sequence. These changes include chemical modifications, including substitution of different amino acid residues, truncations, covalent additions (e.g., of tags), and any other mutations. The term also encompasses circular permutations, mutants, truncations, or domains of a reference sequence that exhibit the same or substantially the same functional activity(ies) as the reference sequence. The term also encompasses fragments of the wild-type protein.

[0179] The level or degree to which properties are retained may be reduced relative to the wild-type protein, but typically will be the same or similar. Generally, variants will be overall closely similar, and in many regions, identical, to the amino acid sequences of the proteins described herein. Those skilled in the art will understand how to make and use variants that retain all or at least some of the functional capabilities or properties.

[0180] A variant protein can comprise, or alternatively consist of, an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of a wild-type protein or any protein provided herein (e.g., an SMN protein), for example.

[0181] By a polypeptide having an amino acid sequence that is at least, for example, 95% "identical" to a query amino acid sequence, it is intended that the amino acid sequence of the subject polypeptide is identical to the query sequence, except that the subject polypeptide sequence may contain up to 5 amino acid modifications per 100 amino acids of the query amino acid sequence. In other words, to obtain a polypeptide having an amino acid sequence that is at least 95% identical to the query amino acid sequence, up to 5% of the amino acid residues in the subject sequence may be inserted, deleted, or substituted with another amino acid. These modifications of the reference sequence may occur at the amino or carboxy terminal positions of the reference amino acid sequence or anywhere between these terminal positions, either individually between residues on the reference sequence or in one or more consecutive groups within the reference sequence.

[0182] In practice, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of a protein, such as the SMN protein, can be conventionally determined using known computer programs. A preferred method for determining the best overall match between a query sequence (a sequence of the present invention) and a subject sequence, also referred to as global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)). In a sequence alignment, the query and subject sequences are either both nucleotide sequences or both amino acid sequences. The results of such global sequence alignments are expressed as percent identity. Preferred parameters used in FASTDB amino acid alignments are: Matrix=PAM 0, k-tuple=2, Mismatch Penalty=1, Joining Penalty=20, Randomization Group Length=0, Cutoff Score=1, Window Size=sequence length, Gap Penalty=5, Gap Size Penalty=0.05, Window Size=500 or the length of the amino acid sequence of interest, whichever is shorter.

[0183] If the subject sequence is shorter than the query sequence due to N- or C-terminal deletions rather than internal deletions, a manual correction must be made to the results. This is because the FASTDB program does not take into account N- and C-terminal truncations of the subject sequence when calculating the global percent identity. For subject sequences that are truncated at the N- and C-termini relative to the query sequence, the percent identity is corrected by calculating the number of query sequence residues at the N- and C-termini of the subject sequence that are not matched / aligned with the corresponding subject residues as a percentage of the total bases in the query sequence. Whether a residue matches / aligns is determined by the results of the FASTDB sequence alignment. This percentage is then subtracted from the percent identity calculated by the FASTDB program above using the specified parameters to arrive at a final percent identity score. This final percent identity score is what is used for the purposes of the present invention. Only residues at the N- and C-termini of the subject sequence that are not matched / aligned with the query sequence are considered for the purposes of manually adjusting the percent identity score. That is, only query residue positions outside the farthest N- and C-terminal residues of the subject sequence.

[0184] vector As used herein, the term "vector" refers to a nucleic acid that can be modified to encode a gene of interest, enter a host cell, mutate and replicate within the host cell, and then transfer the replicated form of the vector to another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phages, as well as conjugative transfer plasmids. Additional suitable vectors will be apparent to those skilled in the art based on this disclosure.

[0185] Wild type The term "wild-type" as used herein is a term of art understood by those skilled in the art and means the typical form of an organism, strain, gene, or characteristic as it occurs in nature, as distinguished from mutant or variant forms.

[0186] Detailed Description of the Invention The present disclosure provides a cytosine base editor comprising an evolutionarily directed adenosine deaminase domain (e.g., a variant of the adenosine deaminase TadA that preferentially deaminates cytidine on DNA as described herein) and a napDNAbp domain (e.g., a Cas9 protein) capable of binding to a specific nucleotide sequence, wherein the adenosine deaminase variant provides a base editor (TadCBE) with a smaller size and fewer off-target effects while maintaining the high editing efficiency of existing CBEs. Deamination of cytidine by TadCBE can lead to a cytosine (C) to (T) point mutation, thus converting a C·G base pair to a T·A base pair, a process referred to herein as nucleic acid editing. Such base editors are particularly useful for targeted editing of nucleic acid sequences, such as DNA molecules. Such base editors can be used for targeted editing of DNA in vitro, for example, for generating mutant cells or animals. Such base editors can be used for introducing targeted mutations in living mammalian cells. Such base editors can also be used for introducing targeted mutations in ex vivo cells, e.g., to correct genetic defects in cells obtained from a subject that are subsequently reintroduced into the same or another subject, or for multiplexed editing of multiple genes in a genome. These base editors can then be used for introducing targeted mutations in vivo, e.g., to correct genetic defects in a subject or to introduce deactivating mutations in disease-associated genes, or for multiplexed editing of a genome. The cytosine base editors described herein can be used for targeted editing of T to C mutations (e.g., targeted genome editing). The present invention provides deaminases, base editors, nucleic acids, vectors, cells, compositions, methods, kits, and uses that utilize the deaminases and base editors provided herein.

[0187] Here, PACE and PANCE were utilized to modulate the substrate specificity of TadA-8e, resulting in a new class of selective cytidine deaminases (TadA-CDs) and cytosine base editors (Figure 1A). To enable cytidine deamination, TadA-CD variants acquired mutations in residues that interact with the DNA backbone near the active site. The disclosed TadA-CD cytosine base editor (TadCBE) is highly active and exhibits comparable or higher C·G to T·A conversion efficiency at various sites in mammalian cells compared to the current BE4max, evoAPOBEC1-BE4max (evoA), and evoFERNY-BE4max (evoFERNY) CBEs. TadA-CD is also compatible with both SpCas9 (PAM=NGG) and evolved eNme2-C Cas9 (PAM=N4CN) variants, facilitating broad target accessibility. Off-target analysis reveals that TadCBE induces less Cas-independent off-target DNA and RNA editing than widely used APOBEC-based CBE variants. 9,34The addition of TadCBEs further reduces off-target editing by TadCBEs, enhances their editing window, and improves selectivity from C·G to T·A while maintaining peak on-target editing efficiency. Herein, evolved TadCBEs were extensively characterized using a library of 10,638 genome-integrated, highly variable target sites in mouse embryonic stem cells (mESCs) to determine the selectivity and sequence context preference of TadCBEs. TadA-CD is also compatible with both SpCas9 and evolved eNme2-C Cas9 variants, facilitating broad target accessibility. The disclosed TadCBEs can be used for efficient cytosine base editing in human cells at therapeutically relevant loci, including multiple edits, particularly for cytosine editing at therapeutically relevant sites in primary human hematopoietic stem and progenitor cells (HSPCs). These disclosed TadCBEs exhibit a more precise editing window with fewer bystander edits than existing CBEs, for example, in the CXCR5 and CCR5 genes in primary human T cells. This disclosure provides a new family of small CBEs with high on-target activity, a well-defined editing window that facilitates precise base editing, and low off-target activity, establishing the potential of adenosine deaminases to evolve into selective cytidine deaminases.

[0188] In some aspects, the present disclosure relates to an adenosine deaminase (e.g., TadA-CD) with targeted cytosine activity. In some embodiments, TadA-CD is evolved from an E. coli tRNA adenosine deaminase (e.g., TadA-8e) previously engineered to act on single-stranded DNA (as opposed to RNA) for adenosine base editing applications. Those skilled in the art will appreciate that PACE and PANCE methodologies can be used to introduce additional mutations into the TadA-8e domain that modulate the enzyme's substrate specificity to generate TadA-CD. In some embodiments, TadA-CD (e.g., a mutated TadA-8e deaminase) contains between 80% and 99.5% sequence identity with the parent TadA-8e. In some cases, the TadA-CD deaminase contains mutations at E27, V28, and H96, relative to the parent TadA-8e, and further contains at least one mutation at a residue selected from R26, M61, Y73, I76, M151, Q154, and A158.

[0189] In some embodiments, TadA-CD variants have enhanced selectivity and deamination activity for cytosine relative to adenosine compared to the parent TadA-8e variant. For example, in some embodiments, TadA-CD deaminase converts between 85% and 92% of CT base pairs to TA base pairs at the C4 and C5 positions of a target sequence (depending on the type of variant), with less than 2% adenine editing; whereas base editors comprising TadA-8e deaminase convert approximately 90% of AT base pairs to GC base pairs at the A6 position of a target sequence, with less than 2% CG to TA base pair editing (see Example 2). This represents a greater than 3000-fold change in the adenine-to-cytosine base editing ability of the TadA-CD variant relative to TadA-8e.

[0190] In some aspects, the present disclosure relates to a cytosine base editor (CBE) comprising a nucleic acid programmable DNA binding protein (e.g., Cas9) domain fused to a TadA-CD deaminase with cytidine activity (e.g., TadCBE). In some embodiments, the napDNAbp domain comprises a Cas homolog, paralog, ortholog, or analog. The napDNAb domain is a member of Cas9, Cas9n (e.g., SpCas9n), dCas9, CasX, CasY, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, Cas9-NG, LbCas12a, enAsCas12a, SaCas9, SaCas9-KKH, circularly permuted Cas9, Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpCas9-VRQR, SpCas9-NRRH, SpaCas9-NRTH, SpCas9-NRCH, eNme2Cas9, and eNme2-C. The napDNAbp domain may be selected from Cas9, enCjCas9, SauriCas9, Cas9-NG-VRQR, and / or variants thereof. In certain embodiments, the napDNAbp domain is or includes a Cas9 domain or a Cas12a domain from S. pyogenes or S. aureus. In some cases, the napDNAbp domain is an Nme2Cas9 domain from Neisseria meningitidis. In some embodiments, the napDNAbp domain includes a nuclease-dead Cas9 (dCas9) domain, a Cas9 nickase (nCas9) domain, or a nuclease-active Cas9 domain. In some cases, the napDNAbp domain is CjCas9. In various embodiments, the napDNAbp domain is a nickase.

[0191] The disclosed CBEs exhibit low levels of unwanted editing, for example, low Cas9-independent off-target editing.The disclosed CBEs exhibit fewer insertions and / or deletions (indels) and unwanted editing of RNA molecules after their use in the method of editing target sequences on nucleic acids.The disclosed CBEs also exhibit editing efficiencies that exceed the efficiency of the most commonly used CBEs for some therapeutically relevant sites and cell types.

[0192] In some aspects, TadA-CD exhibits a narrower editing window than native cytosine base editors while maintaining comparable or higher maximum editing efficiencies. Collectively, the small size of TadCBEs, their compatibility with eNme2Cas9 (and eNme2-C Cas9), their narrower editing window, and their high editing efficiency and selectivity for cytosine base editing over adenine demonstrate their suitability for various high-precision cytosine base editing applications.

[0193] Other aspects of the present disclosure relate to compositions comprising a TadCBE described herein and one or more guide RNAs, e.g., single guide RNAs ("sgRNAs"). Additionally, the present disclosure provides for nucleic acid molecules encoding and / or expressing a TadCBE described herein, as well as expression vectors or constructs (e.g., AAV vectors) for expressing the TadCBE and / or gRNA described herein, host cells comprising the nucleic acid molecules and expression vectors, and compositions for delivering and / or administering the nucleic acid-based embodiments described herein. In particular, the present disclosure provides improved methods for delivery of the disclosed base editors, e.g., to a subject. Delivery of the disclosed TadCBE variants as RNPs rather than DNA plasmids typically increases the on-target:off-target DNA editing ratio. Delivery of the disclosed TadCBE variants as mRNA molecules (e.g., using electroporation) may increase editing efficiency. CBEs with apparent in vivo on-target editing efficiencies of about 50% are described in International Publication No. WO / 2019 / 226953, published November 28, 2019, and Komor et al., Sci. Adv. 2017; 3:eaao4774, each of which is incorporated herein by reference. The disclosed CBEs may exhibit higher on-target editing efficiencies for target cytosine bases.

[0194] Further provided herein are methods of contacting any of the disclosed TadCBEs with a nucleic acid molecule, e.g., a nucleic acid molecule (e.g., DNA) comprising a target sequence. In some embodiments of the disclosed methods, low off-target DNA and / or RNA editing effects are observed. In some embodiments, the nucleic acid molecule comprises DNA, e.g., single-stranded DNA or double-stranded DNA. The target sequence of the nucleic acid molecule may comprise a target nucleic acid base pair containing cytosine (C). The target sequence may be contained within a genome, e.g., a human genome. The target sequence may comprise a sequence associated with a disease or disorder, such as sickle cell disease or HIV / AIDS, e.g., a target sequence with a point mutation. In other embodiments, the target nucleotide sequence is located in the genome of a rodent, such as a mouse or rat. In other embodiments, the target nucleotide sequence is located in the genome of a livestock animal, such as a horse, cat, dog, or rabbit. In some embodiments, the target nucleotide sequence is located in the genome of a research animal. In some embodiments, the target nucleotide sequence is located in the genome of a genetically engineered non-human subject. In some embodiments, the target nucleotide sequence is located in the genome of a plant. In some embodiments, the target nucleotide sequence is located in the genome of a microorganism, such as a bacterium.

[0195] Additionally, the present disclosure provides for methods of generating the TadCBEs described herein, and methods of using base editors or nucleic acid molecules encoding any of these base editors in applications involving editing nucleic acid molecules, e.g., genomes. In certain embodiments, the methods of engineering base editors provided herein involve a phage-assisted continuous evolution (PACE) system or a non-continuous system (e.g., PANCE). These can be utilized to evolve one or more components of a base editor (e.g., a deaminase domain). In certain embodiments, after successful evolution of one or more components of a base editor (e.g., a deaminase domain), methods of making the base editor include recombinant protein expression methodologies and techniques known to those skilled in the art. Exemplary base editors are made by fusing or linking an adenosine deaminase domain to any of the various napDNAbp domains disclosed herein, such as a Cas9 domain.

[0196] Without wishing to be bound by any particular theory, the TadCBE described herein induces editing in nucleic acid substrates by using a TadA variant to deaminate C bases, resulting in the mutation of C to T via uracil formation. It is believed that fusing one or more uracil DNA glycosylase inhibitors to the deaminase and napDNAbp domains of the CBE inhibits the innate DNA repair process. This, when combined with a nucleic acid-programmable DNA-binding protein (e.g., dCas9) engineered to nick the unedited DNA strand (e.g., the strand containing the G of the original CG target base pair), results in the conversion of the original C·G base pair to a T·A base pair. Without wishing to be bound by any particular theory, it is believed that mutations at residues 26-28 of the disclosed deaminase (relative to TadA8e deaminase) facilitate the "sliding" of the backbone of the DNA substrate to allow the binding pocket of this adenosine deaminase to accept cytosine.

[0197] In some embodiments, the TadCBEs described herein have been engineered to exhibit highly targeted and efficient editing capabilities. Such TadCBEs can be used to target and reverse single nucleotide polymorphisms (SNPs) in genes associated with diseases, such as those associated with sickle cell disease and HIV / AIDS. However, in some cases, the TadCBEs described herein may permit targeted C substitutions with a mixture of T, A, and G. For example, TadCBEs lacking the UGI domain may be useful, for example, as screening platforms for targeted random in vivo mutagenesis. More specifically, they can be used as forward genetics tools to screen for gain-of-function and / or loss-of-function variants at base resolution.

[0198] Deaminase domain The present disclosure provides a cytidine base editor (TadCBE) evolved from the adenosine deaminase domain of an existing adenosine base editor (ABE). The adenosine deaminase used herein was evolved using standard methodologies for converting adenosine (A) to inosine (I) on mammalian DNA. Such adenosine deaminases can cause A:T to G:C base pair conversions. A prior art ABE is ABE7.10, which is disclosed in International Publication No. WO 2018 / 027078, published August 2, 2018. A more recently generated ABE is ABE8e, which contains an adenosine deaminase domain containing a single deaminase variant known as TadA8e, as described in International Publication No. WO 2021 / 158921, published August 12, 2021. TadA8e contains nine mutations relative to the adenosine deaminase TadA7.10 of ABE7.10, which is also the deaminase domain of ABEmax, a variant of ABE7.10 that has been codon-optimized for expression in human cells.

[0199] In some embodiments, the adenosine deaminase is a variant of the known adenosine deaminase TadA7.10 that contains the following mutations compared to wild-type ecTadA (SEQ ID NO: 325): W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, R152P, E155V, I156F, and K157N. In some embodiments, the disclosed adenosine deaminase is a variant of TadA from a species other than E. coli, such as Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis.

[0200] The substrate for the evolution experiments disclosed herein was TadA-8e, which contains the following mutations relative to TadA7.10: A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and D167N. For disclosures of phage-assisted evolution experimental methods, the following references are made: International Publication No. WO 2018 / 027078; International Publication No. WO 2019 / 079347, published April 25, 2019; International Publication No. WO 2019 / 226593, published November 28, 2019; US Patent Publication No. 2018 / 0073012, published March 15, 2018, issued October 30, 2018 as US Patent No. 10,113,163; US Patent Publication No. 2017 / 0121693, published May 4, 2017, issued January 1, 2019 as US Patent No. 10,167,457; and International Publication No. WO 2018 / 027078, published October 22, 2020. No. 2020 / 214842, and International Patent Application No. PCT / US2020 / 033873, filed May 20, 2020, International Publication No. WO 2020 / 236982, published November 26, 2020, and International Publication No. WO 2021 / 158921, the contents of each of which are incorporated herein by reference in their entirety.

[0201] Exemplary, non-limiting embodiments of adenosine deaminases used in evolution are provided herein. In some embodiments, the adenosine deaminase domain of any of the disclosed base editors comprises a single adenosine deaminase or monomer. In some embodiments, the adenosine deaminase domain comprises two, three, four, or five adenosine deaminases. In some embodiments, the adenosine deaminase domain comprises two adenosine deaminases or a dimer. In some embodiments, the deaminase domain comprises a dimer of an engineered (or evolved) deaminase and a wild-type deaminase, such as a wild-type E. coli-derived deaminase. It should be understood that the mutations provided herein (e.g., mutations in ecTadA) can be applied to other adenosine base editors, such as adenosine deaminases, such as those provided in: International Publication No. WO 2018 / 027078, published August 2, 2018; International Publication No. WO 2019 / 079347, published April 25, 2019; International Application No. PCT / US2019 / 033848, filed May 23, 2019, published November 28, 2019 as International Publication No. WO 2019 / 226593; US Patent Publication No. 2018 / 0073012, published March 15, 2018, published October 30, 2018 as US Patent No. 10,113,163; US Patent No. U.S. Patent Publication No. 2017 / 0121693, published May 4, 2017, issued January 1, 2019 as U.S. Patent Publication No. 10,167,457; International Publication No. WO 2017 / 070633, published April 27, 2017; U.S. Patent Publication No. 2015 / 0166980, published June 18, 2015; U.S. Patent No. 9,840,699, published December 12, 2017; and U.S. Patent No. 10,077,453, issued September 18, 2018, and International Patent Application No. PCT / US2020 / 28568, filed April 16, 2020; all of which are incorporated by reference herein in their entirety.

[0202] Exemplary adenosine deaminase substrates that can be evolved into cytidine deaminases according to the present disclosure are disclosed below. Exemplary TadA deaminases from Bacillus subtilis (full length submitted as SEQ ID NO: 318), S. aureus (SEQ ID NO: 317), and S. pyogenes (SEQ ID NO: 354) are provided. Amino acid substitutions in E. coli TadA-8e and homologous mutations in B. subtilis, S. aureus, and S. pyogenes TadA deaminases are shown. Thus, one skilled in the art would be able to generate mutations in any naturally occurring adenosine deaminase that correspond to any of the mutations described herein, such as any of the mutations identified in ecTadA (e.g., with homology to ecTadA). In some embodiments, the adenosine deaminase is derived from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from E. coli. One skilled in the art would be able to identify corresponding residues on any homologous proteins and on their respective encoding nucleic acids by methods well known in the art, such as by sequence alignment and determination of homologous residues.

[0203] In some embodiments, the adenosine deaminase substrate comprises TadA9 or a variant thereof. TadA9 contains V82S and Q154R substitutions relative to TadA-8e (in other words, TadA9 contains Y147R, Q154R, and I76Y mutations relative to TadA7.10). In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA9 (SEQ ID NO: 33). TadA9 may be referred to in the art as TadA*8.9. An ABE containing TadA9 deaminase is referred to herein as ABE9. TadA9 is described in additional detail in Gaudelli et al., Nat Biotechnol. 2020 Jul;38(7):892-900 and PCT Publication No. WO 2021 / 050571, published March 18, 2021, each of which is incorporated herein by reference.

[0204] In some embodiments, the adenosine deaminase substrate comprises TadA20, TadA-8.17-m (TadA17), or a variant thereof. TadA20 contains I76Y, V82S, Y123H, Y147R, and Q154R substitutions relative to TadA7.10. TadA17 contains V82S and Q154R substitutions relative to TadA7.10. Additional details about TadA20 and TadA17 are described in Gaudelli et al., Nat Biotechnol. 2020 Jul;38(7):892-900 and WO 2021 / 050571, published March 18, 2021. TadA20 may be referred to in the art as TadA*8.20. In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA20 (SEQ ID NO: 326). An ABE containing TadA20 deaminase is referred to herein as ABE20. It may also be referred to in the art as ABE8.20, ABE8.20-d, or ABE8.20-m. An ABE containing TadA17 deaminase is referred to herein as ABE17. It may also be referred to in the art as ABE8.17 or ABE8.17-m.

[0205] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences of SEQ ID NOs: 317-323.

[0206] In certain embodiments, the adenosine deaminase domain comprises an adenosine deaminase having a sequence with at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% sequence identity to one of the following:

[0207] TadA 7.10 (E. coli): MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 315)

[0208] TadA-8e (E. coli): SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 350)

[0209] Tad1: SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 1)

[0210] Tad2: SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 2)

[0211] Tad3: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYGLIDATLYVTFEPCVMCAGAIIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 3)

[0212] Tad4: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 4)

[0213] Tad6: SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYGLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 5)

[0214] Tad6-SR: SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYGLIDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRRVFNAQKKAQSSIN SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYGLIDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRRVFNAQKKAQSSIN (SEQ ID NO: 6)

[0215] TadA9: SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILANECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 33)

[0216] TadA20 SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLYDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSTD (SEQ ID NO: 326)

[0217] Staphylococcus aureus TadA: MGSHMTNDIYFMTLAIEEAAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFKNLRANKKSTN (sequence number 317)

[0218] Bacillus subtilis TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE (sequence number 318)

[0219] Salmonella typhimurium(S. typhimurium)TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRWADEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV(sequence number 319)

[0220] Shewanella putrefaciens(S. putrefaciens)TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE(sequence number 320)

[0221] Haemophilus influenzae F3031(H. influenzae)TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK(sequence number 321)

[0222] Caulobacter crescentus(C. crescentus)TadA: MRTDESEDQDHRMMRLALDAARAAAEEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI(sequence number 322)

[0223] Geobacter sulfurreducens (G. sulfurreducens)TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP(sequence number 323)

[0224] Streptococcus pyogenes (S. pyogenes) TadA: MPYSLEEQTYFMQEALKEAEKSLQKAEIPIGCVIVKDGEIIGRGHNAREESNQAIMHAEIMAINEANAHEGNWRLLDTTLFVTIEPCVMCSGAIGLARIPHVIYGASNQKFGGADSLYQILTDERLNHRVQVERGLLAADCANIMQTFFRQGRERKKIAKHLIKEQSDPFD (SEQ ID NO: 354)

[0225] Aquifex aeolicus(A. aeolicus)TadA: MGKEYFLKVALREAKRAFEKGEVPVGAIIVKEGEIISKAHNSVEELKDPTAHAEMLAIKEACRRLNTKYLEGCELYVTLEPCIMCSYALVLSRIEKVIFSALDKKHGGVVSVFNILDEPTLNHRVKWEYYPLEEASELLSEFFKKLRNNII (SEQ ID NO: 355)

[0226] In some embodiments, the TadA deaminase is a full-length E. coli TadA deaminase (ecTadA). For example, in certain embodiments, the adenosine deaminase domain comprises a deaminase comprising the following amino acid sequence:

[0227] MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 325)

[0228] TadA-derived cytidine deaminase (TadA-CD)

[0229] Aspects of the present disclosure relate to evolved adenosine deaminases with enhanced cytosine specificity and cytidine deamination activity. According to certain embodiments, the evolved deaminases can deaminate cytidine on DNA. In some embodiments, the deaminases are evolved from parent adenosine deaminases using sequential and / or non-sequential laboratory directed methods (e.g., PACE and PANCE). In some embodiments, the parent adenosine deaminases evolved using PACE and / or PANCE have cytidine deaminase activity. Deaminases of the present disclosure can be evolved from any adenosine deaminase reported to date to have adenosine deaminase activity, such as those described in the following: International Patent Application No. PCT / US2017 / 045381, filed August 3, 2017; International Patent Application No. PCT / US2020 / 028568, filed April 16, 2020; International Patent Application No. PCT / US2021 / 016827, filed February 5, 2021; and PCT / US2022 / 073781, filed July 15, 2022; all of which are incorporated by reference herein in their entireties. In some cases, the parent deaminase comprises E. coli tRNA adenosine deaminase (TadA). The deaminases of the present application can be evolved from a previously mutated (i.e., evolved) parent TadA variant, such as that described in International Patent Application No. PCT / US2021 / 016827, filed February 5, 2021, published August 12, 2021 as WO 2021 / 158921. For example, in some embodiments, the parent adenosine deaminase is TadA7.10. In other embodiments, the parent adenosine deaminase is a TadA8e variant, which contains eight additional mutations relative to TadA7.10: A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and D167N. Other parent adenosine deaminase substrates are also possible.

[0230] In some embodiments, the TadA-derived cytidine deaminase of the present application is derived from a parent adenosine deaminase (e.g., TadA-8e) using a combination of phage-assisted continuous evolution (PACE) and non-continuous evolution (PANCE). According to certain embodiments, the parent adenosine deaminase comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41. In some cases, the parent adenosine deaminase comprises the sequence of SEQ ID NO: 41.

[0231] In some embodiments, the evolved TadA-derived cytidine deaminase is at least partially homologous to the parent TadA-8e variant. For example, according to certain embodiments, the TadA-derived cytidine deaminase (e.g., TadA-CD) comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41, wherein residue 27 of SEQ ID NO: 41 is any amino acid except E (glutamic acid). TadA-CDs with other sequence homologies are also possible. For example, in certain embodiments, a TadA-derived cytidine deaminase (e.g., TadA-CD) comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41, where residue 28 of SEQ ID NO: 41 is any amino acid except V (valine). In another exemplary embodiment, a TadA-derived cytidine deaminase is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO: 41, where residue 96 of SEQ ID NO: 41 is any amino acid except H (histidine).

[0232] As will be appreciated by those skilled in the art, TadA-derived cytidine deaminases (e.g., TadA-CD) can contain multiple mutations relative to the parent adenosine deaminase (e.g., TadA-8e). In some embodiments, the deaminases of the present application (e.g., TadA-CD) contain mutations at residues E27, V28, and H96. In some embodiments, the disclosed deaminases further contain at least one mutation at a residue selected from R26, M61, Y73, I76, M151, Q154, and A158 on the amino acid sequence of SEQ ID NO: 41, or a corresponding mutation in a homologous adenosine deaminase.

[0233] In some embodiments, the deaminase comprises at least one mutation selected from E27A, E27K, V28G, V28A, and H96N, and further comprises at least one mutation at a residue selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S on the amino acid sequence of SEQ ID NO: 41, or a corresponding mutation in a homologous adenosine deaminase. Other mutations are also possible. For example, in certain embodiments, the TadA-CD enzyme comprises a mutation selected from E27A, V28G, and H96N, and further comprises at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S on the amino acid sequence of SEQ ID NO: 41, or a corresponding mutation in a homologous adenosine deaminase.

[0234] Other exemplary embodiments may include: (1) a deaminase comprising the mutations E27K, V28G, and H96N, and further comprising at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S in the amino acid sequence of SEQ ID NO: 41, or corresponding mutations in a homologous adenosine deaminase; (2) a deaminase comprising the mutations E27A, V28A, and H96N, and further comprising at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S in the amino acid sequence of SEQ ID NO: 41; (3) a deaminase comprising the mutations E27K, 28A, and H96N and further comprising at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S in the amino acid sequence of SEQ ID NO: 41, or the corresponding mutations in a homologous adenosine deaminase.

[0235] In some embodiments, the TadA-derived cytidine deaminase (TadA-CD) comprises at least two mutations at residues selected from R26, M61, Y73, 176, M151, Q154, and A158 (relative to the parent deaminase). In other embodiments, the TadA-CD comprises at least two mutations at residues selected from R26G, M61I, Y73H, 176F, M151I, Q154H, Q154R, and A158S.

[0236] In some aspects, TadA-derived cytidine deaminases are provided that can retain some A-to-G base editing activity. Without wishing to be bound by any particular theory, it has been determined through regression analysis that residues 26-28 of TadA-8e deaminase (as set forth in SEQ ID NO:41), located on a loop near the active site, are critical for switching selectivity from adenosine to cytidine. Furthermore, it is believed that substrate positioning in the active site is a key determinant of deamination selectivity, and that sequence context can influence selective deamination of target bases, as interactions between TadA-CD and the 5' and 3' residues can impact substrate positioning in the active site.

[0237] Again, without wishing to be bound by theory, it is believed that residual A to T editing is highest when an adenine is centered in the editing window (e.g., in SpCas9, protospacer position 5 or 6 with the PAM at positions 21-23) and preceded by a T or a C. In some embodiments, the addition of the V106W mutation improves selectivity by suppressing A deamination to a greater extent than C deamination.

[0238] In some aspects, TadA-derived cytidine deaminases are provided that provide efficient conversion of targeted cytosine to thymine and targeted adenine to guanine (referred to herein as "TadA dual" deaminases and base editors). TadA dual deaminases can edit C and A bases within protospacers, particularly within the editing window of the protospacer. These editors incorporate both A-to-G and C-to-T editing with roughly equivalent efficiency.

[0239] For example, the disclosed TadA dual deaminases incorporate A-to-G editing and C-to-T editing at a ratio of approximately 1.1:1. In some embodiments, the dual editor provides A-to-G and C-to-T editing at ratios of 0.7:1, 0.8:1, 0.9:1, 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, or 1.5:1. Other ranges, including ratios greater than 1.5:1, are also possible. These evolved TadA deaminases and "dual" editors containing these deaminases, which can edit A·T to G·C with substantially the same efficiency as C·G to T·A, may be useful in screening applications, such as methods for screening novel Cas homolog domains and other napDNAbp domains for editing activity against various target sequences. These deaminases and dual editors may also be useful in mutagenesis applications, such as in vivo forward genetic or targeted random mutagenesis screens. These dual editors can also be useful for multiple editing applications.

[0240] Dual Editor In some embodiments, the TadA-based dual editor comprises a cytidine deaminase comprising one, two, three, four, or five mutations selected from R26G, V28A, A48R, Y73S, and H96N. This dual editor is referred to herein as TadDE, and the dual editing deaminase is referred to herein as TadA-CDf (e.g., TadA dual), which has the amino acid sequence set forth in SEQ ID NO: 39.

[0241] Therefore, in some embodiments, deaminases are provided herein that include mutations at residues R26, V28, A48, and Y73 on the amino acid sequence of SEQ ID NO: 41, or corresponding mutations in homologous adenosine deaminases. Additionally, deaminases are provided herein that include mutations at residues R26, E27, V28, A48, and Y73 on the amino acid sequence of SEQ ID NO: 41 (i.e., further include a mutation at E27). In certain embodiments, these deaminases include mutations R26G, V28A, A48R, Y73S, and H96N. In some embodiments, these deaminases include mutations R26G, V28G, A48R, and Y73C.

[0242] As described above and herein, preferred Tad-A derived cytidine deaminases evolved using the PACE and PANCE approaches may contain one or more mutations. For example, TadA-CD variants may contain at least one mutation selected from the following: R26G, E27A, V28G, I76F, H96N, and M151I (e.g., TadA-CDa, SEQ ID NO: 34); R26G, E27A, V28G, I76F, H96N, and A158S (e.g., TadA-CDb, SEQ ID NO: 35); R26G, E27A, V28G, I76F, H96N, Q154R, and A158S (e.g., TadA-CDb, SEQ ID NO: 36). Dc, sequence number 36); E27A, V28G, Y73H, H96N, Q154H, and A158S (for example, TadA-CDd, sequence number 37); R26G, V28A, A48R, Y73S, and H96N (for example, TadA-CDe, sequence number 38); V28A, A48R, and Y73S (for example, TadA-CDf, sequence number 39), and R26G, V28G, A48R, and Y73C (for example, TadA-CDg, sequence number 40).

[0243] In some preferred embodiments, the deaminase may comprise the following mutations: R26G, E27A, V28G, I76F, H96N, and A158S (e.g., TadA-CDa, SEQ ID NO: 34), R26G, E27A, V28G, I76F, H96N, Q154R, and A158S (e.g., TadA-CDb, SEQ ID NO: 35), R26G, E27A, V28G, I76F, H96N, and M151I (e.g., TadA-CDc, SEQ ID NO: 36). 36), E27K, V28A, M61I, and H96N (for example, TadA-CDd, SEQ ID NO: 37), E27A, V28G, Y73H, H96N, Q154H, and A158S (for example, TadA-CDe, SEQ ID NO: 38), R26G, V28A, A48R, Y73S, and H96N (for example, TadA-CDf, SEQ ID NO: 39), and R26G, V28G, A48R, and Y73C (for example, TadA-CDg, SEQ ID NO: 40).

[0244] Those skilled in the art will understand that the evolved deaminases described herein may exhibit varying specificity and / or deamination activity for cytosine and / or adenosine bases due to different types and combinations of inherited mutations. In some embodiments, the cytidine deamination activity of TadA-CD exceeds the cytidine deamination activity of TadA-8e. For example, the cytidine deamination activity of a TadA-CD variant can be equal to or greater than 10×, 20×, 40×, 80×, 100×, 200×, 400×, 800×, 1000×, 2000×, 3000×, or 4000× that of the cytidine deamination activity of TadA-8e. In other embodiments, the cytidine deamination activity of the TadA-CD variant is equal to or less than 4000×, 2000×, 1000×, 800×, 800×, 400×, 200×, 100×, 80×, 40×, 20×, or 10× the cytidine deamination activity of TadA-8e.

[0245] In some embodiments, the adenosine deamination activity of the TadA-CD deaminase is less than the deaminase activity of TadA-8e. For example, in some cases, the adenosine deamination activity of the TadA-CD variant is equal to or less than 4000x, 2000x, 1000x, 800x, 800x, 400x, 200x, 100x, 80x, 40x, 20x, or 10x the adenosine deamination activity of TadA-8e.

[0246] In some embodiments, the TadA-CD variants described above and herein can also include a V106W mutation. Recently, adenosine deaminase TadA variants containing the V106W mutation, such as those described in International Patent Publication Nos. WO 2021 / 214842 and WO 2021 / 158921, each of which is incorporated herein by reference, have been discovered to have reduced Cas-independent off-target editing of DNA and RNA while maintaining high levels of on-target adenosine deaminase activity. In some embodiments, TadA-CD variants containing the V106W mutation have average peak editing efficiencies of 50% or greater, 60% or greater, 70% or greater, 80% or greater, and 90% or greater. In other embodiments, TadA-CD variants containing the V106W mutation have an average peak editing efficiency of less than or equal to 90%, less than or equal to 80%, less than or equal to 70%, less than or equal to 60%, or less than or equal to 50%. ABEs containing only a single TadA deaminase domain rather than a single-chain dimer allow for reduced editor size. 30,31Moreover, although SaCas9 is small enough (1053 amino acids in length, SEQ ID NO: 347) to provide a base editor compatible with a single AAV, its utility is severely limited by the rarity of its NNGRRT PAM. Because base editing requires the presence of a suitable PAM to position the target nucleotide within the editing window, TadCBE, which collectively offers broad PAM compatibility along with simple and efficient in vivo delivery, will advance the in vivo application of base editing.

[0247] In some embodiments, any one of the deaminases listed in Table 10 can further comprise a V106W mutation. In some embodiments, the TadA-CD variant comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% of any of the amino acid sequences listed in Table 10, wherein any one of the sequences listed in Table 10 further comprises a V106W mutation.

[0248] In some embodiments, the TadA variant comprises at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identity to any of the amino acid sequences listed in Table 10. Table 10. [Table A-1] [Table A-2]

[0249] In some embodiments, the dual editor deaminase of the TadDE dual editor (e.g., TadA-CDf or TadA dual, SEQ ID NO: 39) can be further evolved, e.g., using the PACE and / or PANCE assays further described below and elsewhere herein. In some embodiments, the TadA dual deaminase (e.g., TadA-CDf, SEQ ID NO: 39) is further evolved to enhance its specificity for cytosine bases and reduce its specificity for adenosine bases. For example, Figure 51E shows a table listing evolved TadA dual deaminases (e.g., TadDE-1 through TadDE-5) along with their mutations relative to the unmutated TadA dual deaminase and its parent TadA-8e deaminase.

[0250] In some embodiments, the TadA dual deaminase is mutated using PACE, as shown in Figure 51C. In some embodiments, phage-assisted continuous evolution, or PACE (Figure 51C, left), is used in conjunction with a selection circuit (Figure 51C, right). In some embodiments, a continuous stream of E. coli host cells is infected with a selection phage (SP) encoding a partial deaminase. One of skill in the art will understand that the E. coli host cells must also contain the plasmids defining the selection circuit and the mutagenesis plasmid. In the selection circuit, phage propagation is linked to expression of gIII (P2), which can only be transcribed by active T7 RNA polymerase. In some embodiments, T7 RNA polymerase (P3) is fused to a C-terminal degron, and the deaminase must perform C-to-U editing to incorporate a stop codon before the degron, generating active T7 RNA polymerase. In the event of phage infection, the full-length deaminase is completed using a split intein system (P1), and mutations can occur on the deaminase. Beneficial mutations lead to phage proliferation and concentration in the lagoon, while less adapted phages are unable to propagate and are subsequently washed away by constant outflow.

[0251] In some embodiments, the TadA dual deaminase is mutated using phage-assisted non-sequential evolution (PANCE), as shown in Figure 51D. In some embodiments, PANCE is performed on the TadA dual deaminase (SEQ ID NO: 39) until the phage titer increases despite higher stringency from dilution factors and promoter strength, indicating that a beneficial mutation has occurred. In some embodiments, the beneficial mutation comprises a mutation at position N46 of the deaminase.

[0252] In some embodiments, PANCE is performed on the NNK library at position N46 to further identify beneficial mutations. In some embodiments, a combination of mutagenesis assays can be performed. For example, in some embodiments, PACE can be performed for more than 100 hours on the resulting variants from the PANCE study.

[0253] In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46I, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-1, Figure 51E, PANCE, SEQ ID NO: 42). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46T, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-2, Figure 51E, PANCE, SEQ ID NO: 43). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46T, A48R, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-3, Figure 51E, PANCE, SEQ ID NO: 44). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-4, PANCE against NNK library at N46, SEQ ID NO: 45). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-5, PANCE against NNK library at N46, SEQ ID NO: 46). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46L, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-6, PANCE against NNK library at N46, SEQ ID NO: 47). In some embodiments, the evolved TadA dual deaminase comprises the mutations V28A, N46L, A48P, and Y73P relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-7, PANCE against NNK library at N46, SEQ ID NO: 48).In some embodiments, the evolved TadA dual deaminase comprises the mutations V28A, N46C, A48P, and Y73P relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-8, PANCE against NNK library at N46, SEQ ID NO: 49). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-9, Figure 51E, PACE, SEQ ID NO: 50). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Q71H, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-10, Figure 51E, PACE, SEQ ID NO: 51). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46L, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-11, Figure 51E, PACE, SEQ ID NO: 52). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46C, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-12, Figure 51E, PACE, SEQ ID NO: 53). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46C, A48R, Y73P, H96N, and A162V relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-13, Figure 51E, PACE, SEQ ID NO: 54).

[0254] In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46I, A48R, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-14, Figure 51E, PANCE, SEQ ID NO: 359). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, A48R, Q71S, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-15, Figure 51E, PANCE, SEQ ID NO: 360). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46L, A48R, and Y73P relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-16, Figure 51E, PANCE, SEQ ID NO: 361). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46L, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-17, Figure 51E, PANCE, SEQ ID NO: 362). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-18, Figure 51E, PANCE, SEQ ID NO: 363). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-19, Figure 51E, PANCE, SEQ ID NO: 364). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-20, Figure 51E, PANCE, SEQ ID NO: 365). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G and N46L relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-21, Figure 51E, PANCE, SEQ ID NO: 366).In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46I, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-22, Figure 51E, PANCE, SEQ ID NO: 367). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-23, Figure 51E, PANCE, SEQ ID NO: 368). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, A48P, Y73H, T79P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-24, Figure 51E, PANCE, SEQ ID NO: 369). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, N46I, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-25, Figure 51E, PANCE, SEQ ID NO: 370).

[0255] In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-26, Figure 51E, PANCE, SEQ ID NO: 371). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46L, A48R, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-27, Figure 51E, PANCE, SEQ ID NO: 372). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46C, A48R, H96N, and A162V relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-28, Figure 51E, PANCE, SEQ ID NO: 373). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Q71H, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-29, Figure 51E, PANCE, SEQ ID NO: 374). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46C, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-30, Figure 51E, PANCE, SEQ ID NO: 375). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46C, A48R, Y73P, H96N, and A162V relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-31, Figure 51E, PANCE, SEQ ID NO: 376).

[0256] In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-32, Figure 51E, PANCE, SEQ ID NO: 377). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48R, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-33, Figure 51E, PANCE, SEQ ID NO: 378). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46V, A48P, Y73S, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-34, Figure 51E, PANCE, SEQ ID NO: 379). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46C, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-35, Figure 51E, PANCE, SEQ ID NO: 380). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, L34M, N46L, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-36, Figure 51E, PANCE, SEQ ID NO: 381). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46L, A48R, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-37, Figure 51E, PANCE, SEQ ID NO: 382). In some embodiments, the evolved TadA dual deaminase comprises the mutations R26G, V28A, N46L, A48P, R64K, Y73P, and H96N relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-38, Figure 51E, PANCE, SEQ ID NO: 383).

[0257] In some embodiments, the evolved TadA dual deaminase comprises the mutations N46I, S73P, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-1, Figure 51E, PANCE, SEQ ID NO: 42). In some embodiments, the evolved TadA dual deaminase comprises the mutation N46T relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-2, Figure 51E, PANCE, SEQ ID NO: 43). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46T and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-3, Figure 51E, PANCE, SEQ ID NO: 44). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-4, PANCE against NNK library at N46, SEQ ID NO: 45). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V, S73P, G105S, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-5, PANCE against NNK library at N46, SEQ ID NO: 46). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46L, S73P, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-6, PANCE against NNK library at N46, SEQ ID NO: 47). In some embodiments, the evolved TadA dual deaminase comprises the mutations G26R N46L, R48P, S73P, N96H, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-7, PANCE against NNK library at N46, SEQ ID NO: 48). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46C, N96H, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-8, PANCE against NNK library at N46, SEQ ID NO: 49).In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V, S73P, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-9, Figure 51E, PACE, SEQ ID NO: 50). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V, Q71H, S73P, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-10, Figure 51E, PACE, SEQ ID NO: 51). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46L and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-11, Figure 51E, PACE, SEQ ID NO: 52). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46C, S73P, and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-12, Figure 51E, PACE, SEQ ID NO: 53). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46C, S73P, H154Q, and A162V relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-13, Figure 51E, PACE, SEQ ID NO: 54).

[0258] In some embodiments, the evolved TadA dual deaminase comprises the mutations N46I and H154Q relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-14, Figure 51E, PACE, SEQ ID NO: 359). In some embodiments, the evolved TadA dual deaminase comprises the mutations Q71S and H154Q relative to the amino acid sequence of SEQ ID NO: 41 (TadA-CD-15, Figure 51E, PANCE, SEQ ID NO: 360). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46L, S73P, N79T, and N96H relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-16, Figure 51E, PANCE, SEQ ID NO: 361). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46L, S73P, N79T relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-17, Figure 51E, PANCE, SEQ ID NO: 362). In some embodiments, the evolved TadA dual deaminase comprises the mutations R48A, S73P, and N79T relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-18, Figure 51E, PANCE, SEQ ID NO: 363). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V and N79T relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-19, Figure 51E, PANCE, SEQ ID NO: 364). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V, S73P, and N79T relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-20, Figure 51E, PANCE, SEQ ID NO: 365). In some embodiments, the evolved TadA dual deaminase comprises the mutations A28V, N46L, R48A, S73Y, N79T, and N96H relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-21, Figure 51E, PANCE, SEQ ID NO: 366). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46I, S73P, and N79T relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-22, Figure 51E, PANCE, SEQ ID NO: 367). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V, S73P, N79T, and G106S relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-23, Figure 51E, PANCE, SEQ ID NO: 368).In some embodiments, the evolved TadA dual deaminase comprises the mutations R48P, S73H, and N79P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-24, Figure 51E, PANCE, SEQ ID NO: 369). In some embodiments, the evolved TadA dual deaminase comprises the mutations A28V, N46I, R48A, S73Y, and N79T relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-25, Figure 51E, PANCE, SEQ ID NO: 370).

[0259] In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V and S73P relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-26, Figure 51E, PANCE, SEQ ID NO:371). In some embodiments, the evolved TadA dual deaminase comprises the mutation N46L relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-27, Figure 51E, PANCE, SEQ ID NO:372). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46C, S73Y, and A162V relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-28, Figure 51E, PANCE, SEQ ID NO:373). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V, Q71H, and S73P relative to the amino acid sequence of SEQ ID NO:39 (TadA-CD-29, Figure 51E, PANCE, SEQ ID NO:374). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46C and S73P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-30, Figure 51E, PANCE, SEQ ID NO: 375). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46C, S73P, and A162V relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-31, Figure 51E, PANCE, SEQ ID NO: 376). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V and S73P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-32, Figure 51E, PANCE, SEQ ID NO: 377). In some embodiments, the evolved TadA dual deaminase comprises the mutation N46V relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-33, Figure 51E, PANCE, SEQ ID NO: 378). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46V and R48P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-34, Figure 51E, PANCE, SEQ ID NO: 379).In some embodiments, the evolved TadA dual deaminase comprises the mutations N46CV and S73P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-35, Figure 51E, PANCE, SEQ ID NO: 380). In some embodiments, the evolved TadA dual deaminase comprises the mutations L34M, N46L, and S73P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-36, Figure 51E, PANCE, SEQ ID NO: 381). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46L and S73P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-37, Figure 51E, PANCE, SEQ ID NO: 382). In some embodiments, the evolved TadA dual deaminase comprises the mutations N46L, r48P, R64K, and S73P relative to the amino acid sequence of SEQ ID NO: 39 (TadA-CD-38, Figure 51E, PANCE, SEQ ID NO: 383).

[0260] In some embodiments, TadA-CD deaminases evolved from TadA dual deaminases have improved specificity for cytosine bases. In some embodiments, evolved TadA-CD deaminases exhibit cytosine on-target activity similar to other evolved deaminases described herein. In some embodiments, evolved deaminases evolved from TadA dual deaminases have increased specificity for cytosine bases and decreased specificity for adenosine bases. In some embodiments, deaminases evolved from TadA dual deaminases do not exhibit residual A to G base editing (e.g., TadA-CD-1 through TadA-CD-38).

[0261] In some embodiments, TadA-CD-1 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-2 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-3 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-4 exhibits no residual A to G base editing when incorporated into the BE4max architecture. In some embodiments, TadA-CD-5 exhibits no residual A to G base editing when incorporated into the BE4max architecture. The above description is not intended to be limiting in any way, and the evolved TadA-CD deaminases described herein can be used in any suitable architecture known to those of skill in the art.

[0262] In some embodiments, the TadA-CD evolved from the TadA dual comprises at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identity to any of the amino acid sequences listed in Table 11.

[0263] In some embodiments, any one of the deaminases listed in Table 11 can further comprise a V106W mutation. In some embodiments, the TadA-CD variant comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% of any of the amino acid sequences listed in Table 10, wherein any one of the sequences listed in Table 11 further comprises a V106W mutation.

[0264]

[0265] Table 11. List of exemplary mutated TadA-CD relative sequences derived from the TadA dual (SEQ ID NO: 39). The sequences of TadA-8e and the TadA dual are provided as references. [Table B-1] [Table B-2] [Table B-3] [Table B-4] [Table B-5] [Table B-6]

[0266] napDNAbp domain The base editors described herein contain a nucleic acid-programmable DNA-binding (napDNAbp) domain. The napDNAbp binds to at least one guide nucleic acid (e.g., a guide RNA), which localizes the napDNAbp to a DNA sequence containing a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid or a portion thereof (e.g., a protospacer of the guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp domain to localize and bind to a complementary sequence on the target strand. Binding of the napDNAbp domain to the complementary sequence allows the nucleobase-modifying domain (i.e., an adenosine deaminase domain) of the base editor to access and enzymatically deaminate the target base on the target strand.

[0267] napDNAbp may be a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. As outlined above, CRISPR is an adaptive immune system that provides protection from mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain a spacer, a sequence complementary to an ancestral mobile element, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of the pre-crRNA requires a small trans-encoded RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first endolytically cut and then 3'-5' exolytically trimmed. In nature, DNA binding and cleavage typically require both a protein and an RNA. However, single-guide RNAs ("sgRNAs," or simply "gRNAs") can be engineered to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, for example, Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference.

[0268] The description below of various napDNAbps that can be used in conjunction with the disclosed adenosine deaminases is not meant to be limiting in any way. Base editors can include any variant Cas9 protein, including classical SpCas9, or any orthologous Cas9 protein, or any naturally occurring variant, mutant, or otherwise engineered version of Cas9. This may be known, or it may be created or evolved through a directed evolutionary or other mutagenesis process. In various embodiments, the napDNAbp has nickase activity, i.e., it cleaves only one strand of the target DNA sequence. In other embodiments, the napDNAbp has an inactive nuclease, e.g., a "dead" protein. Other variant Cas9 proteins that can be used are those that have a smaller molecular weight than classical SpCas9 (e.g., for easier delivery) or have an altered or rearranged primary amino acid sequence (e.g., a circularly permuted form). The base editors described herein can also include Cas9 equivalents, including Cas12a / Cpf1 proteins. As used herein, napDNAbp (e.g., SpCas9, SaCas9, or SaCas9 or SpCas9 variants) can also contain various modifications that modulate / enhance their PAM specificity. The present disclosure contemplates any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to any of the Cas9 proteins disclosed herein. In some embodiments, the napDNAbp domain comprises a nickase variant of wild-type Cas9. In some embodiments, the napDNAbp domain comprises any of the Cas9 nickases disclosed herein.

[0269] In some embodiments, the napDNAbp directs cleavage of one or both strands at the location of the target sequence, e.g., within the target sequence and / or within the complement of the target sequence. In some embodiments, the napDNAbp directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. For example, an aspartate to alanine substitution (D10A) on the RuvC I catalytic domain of Cas9 from S. pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that render Cas9 a nickase include, without limitation, H840A, N854A, and N863A with reference to the classical SpCas9 sequence, H588A and D16A with reference to the Nme2Cas9 sequence, and equivalent amino acid positions in other Cas9 variants or Cas9 equivalents.

[0270] As used herein, the term "Cas protein" refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence that differs from a naturally occurring Cas protein, or any fragment of a Cas protein that still retains all or a significant amount of the essential basic functions required for the disclosed methods, namely, (i) possession of nucleic acid-programmable binding of the Cas protein to target DNA and (ii) the ability to nick a target DNA sequence on one strand. Cas proteins contemplated herein encompass CRISPR Cas9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9)), homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and include Cas9 equivalents from any type of CRISPR system (e.g., Type II, V, VI), including Cpf1 (Type V CRISPR-Cas system), C2c1 (Type V CRISPR-Cas system), C2c2 (Type VI CRISPR-Cas system), and C2c3 (Type V CRISPR-Cas system). Additional Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299). No. 6,239,999, the contents of which are incorporated herein by reference.

[0271] The term "Cas9" or "Cas9 domain" encompasses any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of naturally occurring or engineered Cas9. The term Cas9 is not meant to be particularly limiting and may be referred to as "Cas9 or equivalent." Exemplary Cas9 proteins are further described herein and / or in the art and are incorporated herein by reference. The present disclosure is not limited with respect to the particular napDNAbp employed in the base editors of the present disclosure.

[0272] As used herein, the terms "compact Cas9 protein," "compact napDNAbp," and "compact variant (of a Cas protein)" refer to a Cas9 protein or variant having an amino acid length of less than about 1250 amino acids. In some embodiments, the compact Cas9 protein or compact napDNAbp contains less than 1250, 1240, 1230, 1220, 1210, 1200, 1190, 1180, 1170, 1160, 1150, 1140, 1130, 1120, 1110, 1100, 1050, 1000, 950, 900, 850, 800, 750, 700, 650, 600, 550, or 500 amino acids in length. These terms also encompass any Cas9 protein or variant encoded by a nucleic acid sequence having a length of less than about 3750 nucleotides. The base editors of the present disclosure may include a compact napDNAbp and / or a compact Cas9 protein. In some embodiments, the compact Cas9 protein is about 350 amino acids shorter than SpCas9. In some embodiments, the compact Cas9 protein is about 1000 amino acids in length. In some embodiments, the compact protein is a compact variant of S. pyogenes Cas9 (SpCas9), Cpfl, CasX, CasY, C2cl, C2c2, C2c3, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Cas3, or CasΦ. A "compact variant" may refer to a Cas9 protein that has one or more truncations or one or more deletions relative to a wild-type Cas9 protein, such as wild-type SpCas9 or Cpfl.

[0273] The complete genome sequence of an M1 strain of Streptococcus pyogenes is readily available from the Cas9 strain of the species.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G, Lyon K, Primeaux C, Sezate S, Suvorov AN, Kenton S, Lai HS, Lin SP, Qian Y, Jia HG, Najar FZ, Ren Q, Zhu H, Song L, White J, Yuan X, Clifton SW, Roe BA, McLaughlin RE, Proc Natl. Acad 471:602-607(2011);および“A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012). All of them are covered in a smooth, smooth surface.

[0274] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not meant to be limiting. The base editors of the present disclosure can use any suitable napDNAbp, including any suitable Cas9 or Cas9 equivalent.

[0275] In some embodiments, the Cas9 comprises or is derived from a wild-type SaCas9 (e.g., Staphylococcus aureus, 1053AA, 123kDa). In some embodiments, the wild-type SaCas9 comprises the following amino acid sequence:

[0276]

[0277] In some embodiments, the sequence of SaCas9 comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, at least 99.5%, or at least 99.8% sequence identity to SEQ ID NO: 347.

[0278] In some embodiments, the Cas9 comprises or is derived from a wild-type SpCas9 (e.g., SpCas9, Streptococcus pyogenes M1, SwissProt Accession No. Q99ZW2, wild-type). In some embodiments, the wild-type SaCas9 comprises the following amino acid sequence:

[0279]

[0280] In some embodiments, the sequence of SpCas9 comprises at least at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, at least 99.5%, or at least 99.8% sequence identity to SEQ ID NO:200. Cas nickase

[0281] In some embodiments, the disclosed base editors can comprise a napDNAbp domain that includes a Cas nickase. In some embodiments, the base editors described herein comprise a Cas9 nickase. In some embodiments, any of the disclosed base editors or vectors can comprise an S. pyogenes Cas9 nickase (SpCas9n or nCas9) containing a D10A mutation. In some embodiments, any of the disclosed base editors can comprise an Nme2Cas9 nickase (Nme2Cas9n) or an eNme2-C Cas9 nickase (eNme2-C Cas9n), each of which contains a D16A mutation.

[0282] The term "Cas9 nickase" in "nCas9" refers to a variant of Cas9 that can introduce single-strand breaks into double-stranded DNA molecular targets. In some embodiments, the Cas9 nickase contains only a single functional nuclease domain. Wild-type Cas9 (e.g., classical SpCas9) contains two distinct nuclease domains: the RuvC domain (which cleaves non-protospacer DNA strands) and the HNH domain (which cleaves protospacer DNA strands). In one embodiment, the Cas9 nickase contains a mutation in the RuvC domain that inactivates RuvC nuclease activity. For example, mutations at aspartic acid (D) 10, histidine (H) 983, aspartic acid (D) 986, or glutamic acid (E) 762 have been reported as loss-of-function mutations in the RuvC nuclease domain and generate functional Cas9 nickases (see, e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, incorporated herein by reference). Thus, nickase mutations on the RuvC domain can include D10X, H983X, D986X, or E762X, where X is any amino acid other than the wild-type amino acid. In certain embodiments, the nickase can be D10A, H983A, D986A, or E762A, or a combination thereof.

[0283] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises S. pyogenes Cas9 nickase (SpCas9n). In some embodiments, the napDNAbp domain of any of the disclosed base editors is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to or comprises SEQ ID NO: 343. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 343.

[0284] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises S. aureus Cas9 nickase (SaCas9n). In some embodiments, the napDNAbp domain of any of the disclosed base editors is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to or comprises SEQ ID NO: 351. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 351.

[0285] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises N. meningitidis Cas9 nickase (Nme2Ca9n) or a variant thereof. In some embodiments, the napDNAbp domain of any of the disclosed base editors is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to or comprises SEQ ID NO: 352 or 353. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 352 or 353. In some embodiments, the napDNAbp domain comprises the amino acid sequence of SEQ ID NO: 353. eNme2 -C The Cas9 (SEQ ID NO: 353) variant exhibits a preference for targeting the NNNNCN (N4CN) PAM. Base editors containing this eNme2-C variant produced base edits at the N4CC PAM with approximately 60% or higher efficiency in human cells. This represents a two-fold improvement over base editors containing wild-type Nme2Cas9.

[0286] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises a wild-type Nme2Cas9 nuclease (SEQ ID NO: 349).

[0287] In various embodiments, the Cas nickase can have a mutation in the RuvC nuclease domain and can have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto: [Table C-1] [Table C-2] [Table C-3] [Table C-4] [Table C-5]

[0288] Cas9 equivalents In some embodiments, the base editors described herein may encompass any Cas9 equivalent. As used herein, the term "Cas9 equivalent" is a broad term encompassing any napDNAbp that performs the same function as Cas9 on the base editor, even though its primary amino acid sequence and / or its three-dimensional structure may differ and / or may be evolutionarily unrelated. Thus, a Cas9 equivalent encompasses any evolutionarily related Cas9 ortholog, homolog, mutant, or variant described or encompassed herein, but also encompasses proteins that may have evolved through convergent evolution processes to have the same or similar function as Cas9 but do not necessarily share any similarity in amino acid sequence and / or three-dimensional structure. Although a Cas9 equivalent may be based on a protein that has arisen through convergent evolution, the base editors described herein encompass any Cas9 equivalent that would provide the same or similar function as Cas9. For example, where Cas9 refers to a Type II enzyme of a CRISPR-Cas system, a Cas9 equivalent could refer to a Type V or Type VI enzyme of a CRISPR-Cas system.

[0289] For example, Cas12e (CasX) is a Cas9 equivalent that reportedly has the same function as Cas9 but evolved through convergent evolution. Thus, the Cas12e (CasX) protein described in Liu et al., "CasX enzymes comprise a distinct family of RNA-guided genome editors," Nature, 2019, Vol. 566: 218-223, is contemplated for use in the base editors described herein. Additionally, any variants or modifications of Cas12e (CasX) are envisioned and within the scope of the present disclosure.

[0290] Cas9 is a bacterial enzyme that has evolved in a wide variety of species, however, Cas9 equivalents contemplated herein may also be obtained from Archaea, which represent a domain and kingdom of unicellular prokaryotic microorganisms distinct from Bacteria.

[0291] In some embodiments, Cas9 equivalents may refer to Cas12e (CasX) or Cas12d (CasY). These are described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21, the entire contents of which are incorporated herein by reference. Using genome-resolution metagenomics, several CRISPR-Cas systems have been identified, including the first reported Cas9 in the archaeal domain of life. This divergent Cas9 protein was found in the little-studied nanoarchaea as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-Cas12e and CRISPR-Cas12d, have been discovered. These are among the most compact systems discovered to date. In some embodiments, Cas9 refers to Cas12e or a variant of Cas12e. In some embodiments, Cas9 refers to Cas12d or a variant of Cas12d. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid-programmable DNA-binding proteins (napDNAbp) and are within the scope of this disclosure. See also Liu et al., "CasX enzymes comprise a distinct family of RNA-guided genome editors," Nature, 2019, Vol. 566: 218-223. Any of these Cas9 equivalents are contemplated.

[0292] In some embodiments, the Cas9 equivalent comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp is a naturally occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a wild-type Cas moiety or any Cas moiety provided herein.

[0293] In various embodiments, nucleic acid programmable DNA binding proteins include, without limitation, Cas9 (e.g., dCas9 and nCas9), C2C3Cas12e (CasX), Cas12d (CasY), Cas12a (Cpf1), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), and Argonaute. One example of a nucleic acid programmable DNA binding protein with a PAM specificity different from Cas9 is Clustered Regularly Interspaced Short Palindromic Repeats 1 from Prevotella and Francisella (i.e., Cas12a (Cpf1)). Similar to Cas9, Cas12a (Cpf1) is also a class 2 CRISPR effector, but it is a member of the type V subgroup of enzymes rather than the type II subgroup. Cas12a (Cpf1) has been shown to mediate robust DNA interference with characteristics distinct from those of Cas9. Cas12a (Cpf1) is a single RNA-guided endonuclease lacking tracrRNA, which utilizes a T-rich protospacer adjacent motif (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA through staggered DNA double-strand breaks. Of the 16 Cpf1 family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome editing activity in human cells. Cpf1 proteins are known in the art and have been previously described, for example, in Yamano et al., "Crystal structure of Cpf1 in complex with guide RNA and target DNA," Cell (165) 2016, pp. 949-962, the entire contents of which are incorporated herein by reference.

[0294] In still other embodiments, the Cas protein may include any CRISPR-associated protein, including but not limited to: Cas12a, Cas12b1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Cs m3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, preferably comprising a nickase mutation (e.g., a mutation corresponding to the D10A mutation in the wild-type Cas9 polypeptide of SEQ ID NO: 200).

[0295] In various other embodiments, the napDNAbp can be any of the following proteins: Cas9, C2c3Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, a circularly permuted Cas9, or an Argonaute (Ago) domain, or a variant thereof.

[0296] Cas9 variants with altered PAM specificity The base editors of the present disclosure can also include Cas9 variants with altered PAM specificity. For example, the base editors described herein can utilize any naturally occurring or engineered variant of SpCas9 with extended and / or relaxed PAM specificity. These are described in literature, including Nishimasu et al., "Engineered CRISPR-Cas9 nuclease with expanded targeting space," Science, 2018, 361: 1259-1262; Chatterjee et al., "Robust Genome Editing of Single-Base PAM Targets with Engineered ScCas9 Variants," BioRxiv, April 26, 2019. Some aspects of the present disclosure provide Cas9 proteins that exhibit activity against target sequences that do not contain a classical PAM at their 3' end (5'-NGG-3', where N is A, C, G, or T). In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NGG-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NNG-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NNA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NNC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NNT-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NGT-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NGA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NGC-3' PAM sequence at its 3' end.In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAT-3' PAM sequence at its 3' end. In still other embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAG-3' PAM sequence at its 3' end.

[0297] The above description of various napDNAbps that can be used in connection with the disclosed base editors is not meant to be limiting in any way. The base editor can include any variant Cas9 protein, including a classical SpCas9, or any orthologous Cas9 protein, or any naturally occurring variant, mutant, or otherwise engineered version of Cas9. This may be known, or it may be created or evolved through a directed evolutionary or other mutagenesis process. In various embodiments, the Cas9 or Cas9 variant has nickase activity, i.e., it cleaves only one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variant has an inactive nuclease, i.e., it is a "dead" Cas9 protein. Other variant Cas9 proteins that can be used are those that have a smaller molecular weight than classical SpCas9 (e.g., for easier delivery) or have an altered or rearranged primary amino acid structure (e.g., a circular permutation format) The base editors described herein may also include Cas9 equivalents, including Cas12a / Cpf1 and Cas12b proteins, which are the result of convergent evolution. As used herein, napDNAbp (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) may also contain various modifications that modulate / enhance their PAM specificity. Finally, the present application contemplates any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as a reference classical SpCas9 sequence or a reference Cas9 equivalent (e.g., Cas12a / Cpf1).

[0298] In some embodiments, SpCas9(H840A) comprises a sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or at least 99.5% identical to the amino acid sequence of SEQ ID NO:480.

[0299] SpCas9(H840A)

[0300]

[0301] In certain embodiments, the Cas9 variant with expanded PAM capability is SpCas9(H840A)VRQR, which has the following amino acid sequence (V, R, Q, R substitutions are shown in bold and underlined relative to SpCas9(H840A) of SEQ ID NO:480. Additionally, the methionine residue on SpCas9(H840) has been removed in SpCas9(H840A)VRQR) ("SpCas9-VRQR"). This SpCas9 variant possesses altered PAM specificity, recognizing a 5'-NGA-3' PAM instead of the classic 5'-NGG-3' PAM:

[0302]

[0303] In another specific embodiment, the Cas9 variant with extended PAM capability is SpCas9(H840A)VQR, having the following amino acid sequence (V, Q, R substitutions are shown in bold and underlined relative to SpCas9(H840A) of SEQ ID NO:480. Additionally, the methionine residue on SpCas9(H840) has been removed in SpCas9(H840A)VRQR) ("SpCas9-VQR"). This SpCas9 variant possesses altered PAM specificity, recognizing a 5'-NGA-3' PAM instead of the classic 5'-NGG-3' PAM:

[0304]

[0305] In another specific embodiment, the Cas9 variant with expanded PAM capability is SpCas9(H840A)VRER, having the following amino acid sequence (V,R,E,R substitutions are shown in bold and underlined relative to SpCas9(H840A) of SEQ ID NO:480. Additionally, the methionine residue on SpCas9(H840) has been removed in SpCas9(H840A)VRER) ("SpCas9-VRER"). This SpCas9 variant possesses altered PAM specificity, recognizing the 5'-NGCG-3' PAM instead of the 5'-NGG-3' classical PAM.

[0306]

[0307] In another embodiment, the Cas9 variant with expanded PAM capability is SpCas9-NG, as reported in Nishimasu et al., "Engineered CRISPR-Cas9 nuclease with expanded targeting space," Science, 2018, 361: 1259-1262, which is incorporated herein by reference. SpCas9-NG (VRVRFRR) has the following amino acid sequence substitutions relative to the classic SpCas9 sequence (SEQ ID NO: 200): R1335V, L1111R, D1135V, G1218R, E1219F, A1322R, and T1337R. This SpCas9 has relaxed PAM specificity, i.e., it has activity against the PAM of NGH (where H = A, T, or C). See Nishimasu et al., "Engineered CRISPR-Cas9 nuclease with expanded targeting space," Science, 2018, 361: 1259-1262, which is incorporated herein by reference.

[0308]

[0309] Additionally, any available method for obtaining or constructing a variant or mutant Cas9 protein can be utilized. As used herein, the term "mutation" refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues within a sequence. Mutations are typically described herein by identifying the original residue, followed by the position of the residue within the sequence and the identity of the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art, and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). Mutations can encompass various categories, such as single nucleotide polymorphisms, microduplication regions, indels, and inversions, and are in no way intended to be limiting. Mutations can include "loss-of-function" mutations, which are the normal result of a mutation that reduces or eliminates protein activity. Most loss-of-function mutations are recessive. This is because in heterozygotes, the second chromosomal copy carries a fully functional, unmutated version of the gene encoding the protein, whose presence compensates for the effect of the mutation. Mutations also encompass "gain-of-function" mutations, which confer abnormal activity to a protein or cell that would not otherwise be present under normal conditions. Many gain-of-function mutations are in regulatory sequences rather than coding regions and therefore can have several consequences. For example, a mutation can lead to one or more genes being expressed in the wrong tissues, and these tissues gain a function they normally lack. By their nature, gain-of-function mutations are usually dominant.

[0310] Mutations can be introduced into a reference Cas9 protein using site-directed mutagenesis. Older methods of site-directed mutagenesis known in the art rely on subcloning the sequence to be mutated into a vector, such as an M13 bacteriophage vector, which allows for the isolation of a single-stranded DNA template. In these methods, a mutagenesis primer (i.e., a primer that can anneal to the site to be mutated but has one or more mismatched nucleotides at the site to be mutated) is annealed to the single-stranded template, and then the complement of the template is polymerized starting from the 3' end of the mutagenesis primer. The resulting duplex is then transformed into host bacteria, and plaques are screened for the desired mutation. More recently, site-directed mutagenesis has adopted PCR methodology, which has the advantage of not requiring a single-stranded template. In addition, methods that do not require subcloning have been developed. When PCR-based site-directed mutagenesis is performed, several issues must be considered. First, in these methods, it is desirable to reduce the number of PCR cycles to prevent the extension of unwanted mutations introduced by the polymerase. Second, selection must be employed to reduce the number of unmutated parent molecules that persist in the reaction. Third, extended-length PCR methods are preferred to allow the use of a single PCR primer set. Fourth, due to the template-independent end-extension activity of some thermostable polymerases, it is often necessary to incorporate an end-polishing step into the procedure prior to blunt-end ligation of the mutant products generated by PCR.

[0311] Mutations can also be introduced by directed evolution processes such as phage-assisted continuous evolution (PACE) or phage-assisted non-continuous evolution (PANCE). As used herein, the term "phage-assisted continuous evolution (PACE)" refers to continuous evolution that employs phages as viral vectors. The general concept of PACE technology is described, for example, in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, published March 11, 2010 as WO 2010 / 028347; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, published June 28, 2012 as WO 2012 / 088381; US ​​Patent No. 9,023,594, issued May 5, 2015; International PCT Application No. PCT / US2015 / 012022, filed January 20, 2015, published September 11, 2015 as WO 2015 / 134121; and US Patent No. International PCT Application PCT / US2016 / 027795, filed April 15, 2016, published October 20, 2016 as PCT / US2016 / 168631, the entire contents of each of which are incorporated herein by reference. Variant Cas9s can also be obtained by phage-assisted non-sequential evolution (PANCE). As used herein, this refers to non-sequential evolution employing phages as viral vectors. PANCE is a simplified technique for rapid in vivo directed evolution, using serial flask transfers of evolving "selection phages" (SPs) containing the gene of interest to be evolved between new E. coli host cells, thereby allowing the gene within the host E. coli to remain constant while the gene contained in the SP is serially evolved. Serial flask transfers have long served as a widely accessible approach for laboratory evolution of microorganisms. More recently, a similar approach for bacteriophage evolution has been developed. The PANCE system is characterized by lower stringency than the PACE system. Compact Cas9 variants with altered PAM specificity

[0312] In some embodiments, the napDNAbp comprises a compact Cas protein, such as Cas9 from C. jejuni, S. auricularis, N. meningitidis, or S. aureus. In exemplary embodiments, the napDNAbp comprises a CjCas9 nickase, a SauriCas9 nickase, an Nme2Cas9 nickase, a SaCas9 nickase, or a SaKKH-Cas9 nickase. In some embodiments, the napDNAbp is not an Nme2Cas9 protein or nickase. In some embodiments, the napDNAbp is not a SaCas9 protein or nickase.

[0313] In some embodiments, the disclosed base editors comprise a napDNAbp domain, which comprises a Cas9 orthologue from Neisseria meningitidis (Nme or Nme2). In some embodiments, the napDNAbp domain comprises Nme2Cas9. In other embodiments, the napDNAbp domain is an Nme2Cas9 domain. In some embodiments, the disclosed base editors comprise an Nme2Cas9 nickase. Nme2Cas9 recognizes simple dinucleotide PAMs, NNNNCC or N4CC (where N is any nucleotide), as described in Edraki et al., Molecular Cell 73, 714-726, incorporated herein by reference. In other embodiments, the napDNAbp domain comprises an Nme2Cas9 variant. Nme2Cas9 variants can recognize a wider variety of PAMs. In some embodiments, the disclosed Nme2Cas9 variants recognize single-nucleotide pyrimidine PAMs. In some embodiments, the Nme2Cas9 variant recognizes a PAM of the sequence NYN, where Y is any pyrimidine (i.e., C, T, or U). In other embodiments, the Nme2Cas9 variant recognizes a PAM of the sequence NNNNCN or N4CN. In some embodiments, the Nme2Cas9 variant is an eNme2Cas9 nickase (SEQ ID NO: 439). In some embodiments, the Nme2Cas9 variant is an eNme2-C Cas9 nickase (SEQ ID NO: 353).

[0314] The sequence of wild-type Nme2Cas9 is submitted as SEQ ID NO: 349. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 349. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 5. This protein may be referred to herein as engineered Nme2Cas9 or eNme2Cas9. In various embodiments, any of the disclosed TadCBEs comprises Nme2Cas9 or a variant of Nme2Cas9.

[0315] Wild type Nme2Cas9

[0316] The "e" at the beginning of the Nme2Cas9 variants described herein denotes an "evolved" Nme2 variant. Amino acid substitutions relative to wild-type Nme2Cas9 are indicated in bold and underlined.

[0317] eNme2-C Cas9 [ka] (SEQ ID NO: 353)

[0318] In some embodiments, the disclosed base editor comprises a napDNAbp comprising the compact Cas9 orthologue from Campylobacter jejuni (CjCas9). In some embodiments, the napDNAbp comprises CjCas9. In some embodiments, the disclosed base editor comprises a CjCas9 nickase. CjCas9 recognizes the NNNNACA and NNNNACAC PAMs. See Kim et al., Nature Communications 8(14500):1-12 (2017), incorporated herein by reference. The sequence of CjCas9 (nickase) is submitted as SEQ ID NO: 348. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 348. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 348. The protein is 984 amino acids in length. This protein may be referred to herein as engineered CjCas9 or enCjCas9. Rationally engineered CjCas9 variants (enCjCas9) are described in Nakagawa, et al., Communications Biology, (2022) 5:211, which is incorporated herein by reference. In various embodiments, any of the disclosed TadCBEs comprises a variant of CjCas9 or enCjCas9 (SEQ ID NO: 348).

[0319] MARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRLARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPVVLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKPERKKPSGALHEETFRKEEEFYQSYGGKEGVLKALELGKIRKVNGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKDMQEPEFVYYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKSIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKK(SEQ ID NO: 348)

[0320] Base editors of the present disclosure can also include Cas9 variants with altered PAM specificity. Some aspects of the present disclosure provide Cas9 proteins that exhibit activity toward target sequences that do not contain a classical PAM (5'-NGG-3', where N is A, C, G, or T) at their 3' end. In some embodiments, the Cas9 protein exhibits activity toward target sequences that contain a 5'-NGG-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity toward target sequences that contain a 5'-NNG-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity toward target sequences that contain a 5'-NNA-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity toward target sequences that contain a 5'-NNC-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity toward target sequences that contain a 5'-NNT-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NGT-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NGA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NGC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAT-3' PAM sequence at its 3' end. In still other embodiments, the Cas9 protein exhibits activity toward a target sequence comprising a 5'-NAG-3' PAM sequence at its 3' end.

[0321] In some embodiments, the disclosed base editors comprise a napDNAbp domain comprising SpCas9-NG with a PAM corresponding to NGN. In some embodiments, the disclosed base editors comprise a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SpCas9-NG. The sequence of SpCas9-NG is illustrated below:

[0322] In some embodiments, the disclosed base editors comprise a napDNAbp domain comprising the S. aureus Cas9 nickase KKH or SaCas9-KKH or SaKKH-Cas9, which has a PAM corresponding to NNNRRT or NNGRRT. This Cas9 variant contains the amino acid substitutions D10A, E782K, N968K, and R1015H ("KKH") relative to the wild-type SaCas9 submitted as SEQ ID NO: 347. In some embodiments, the disclosed base editors comprise a napDNAbp domain having a sequence at least 90%, at least 95%, at least 98%, or at least 99% identical to SaCas9-KKH. SaCas9 (and SaKKH-Cas9) is 1053 amino acids in length. The sequence of SaCas9-KKH (nickase) is illustrated below:

[0323] S. aureus Cas9 nickase KKH (SaCas9-KKH)

[0324] In some embodiments, the disclosed base editor comprises a napDNAbp comprising the Cas9 protein from Staphylococcus Auricularis (S. auri Cas9 or SauriCas9). In some embodiments, the disclosed base editor comprises a SauriCas9 nickase. SauriCas9 recognizes NNGG and NNNGG PAMs. The sequence of SauriCas9 (nickase) is submitted as SEQ ID NO: 358. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 358. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 358. The protein is 1061 amino acids in length.

[0325] In some embodiments, the napDNAbp comprises a SauriCas9-KKH variant or a SauriCas9-KKH nickase variant. SauriCas9-KKH contains the corresponding triple KKH mutations: Q788K, Y973K, and R1020H. See Hu et al. (2020) PLoS Biol. 18(3): e3000686, incorporated herein by reference.

[0326] In some embodiments, the disclosed base editors comprise a napDNAbp domain comprising an S. pyogenes Cas9 nickase KKH or SpCas9-KKH, which has a PAM corresponding to an NNNRRT.

[0327] In some embodiments, the Cas variant is a variant of SpRY with a mutation that confers high fidelity. Such variants are known as SprY-HF or SprY-HF1. A high-fidelity variant of SpRY, or any of the Cas variants provided herein, can include one or more of the following mutations relative to SEQ ID NO: 74: N497A, R661A, Q695A, and / or Q926A, or the corresponding mutations in any of the Cas9s provided herein. Cas9 variants with high fidelity are known in the art and will be apparent to those skilled in the art. For example, high-fidelity Cas9 domains are described in Kleinstiver, BP, et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016); and Slaymaker, IM, et al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015); each of which is incorporated herein by reference.

[0328] In some embodiments, the disclosed Cas variants include variants of Cas9 derived from Streptococcus macacae, e.g., Streptococcus macacae NCTC 11558, or SmacCas9. In some embodiments, the Cas variants include a hybrid variant of SmacCas9 incorporating an SmacCas9 domain and an SpCas9 domain, known as Spy-macCas9, or variants thereof. In some embodiments, the Cas variants include a hybrid variant of SmacCas9 incorporating an enhanced nucleolytic variant of the SpCas9 (iSpy Cas9) domain, known as iSpy-macCas9. Relative to Spymac-Cas9, iSpyMac-Cas9 contains two mutations, R221K and N394K, identified by deep mutation scanning of Spy Cas9 that increase the rate of protein modification in most targets. See Jakimo et al., bioRxiv, A Cas9 with Complete PAM Recognition for Adenine Dinucleotides (Sep 2018), incorporated herein by reference. Jakimo et al. showed that hybrid Spy-macCas9 and iSpy-macCas9 recognize short 5'-NAA-3' PAM sequences, recognize all adenine dinucleotide PAM sequences evaluated, and possess robust editing efficiency in human cells. Liu et al. engineered a base editor containing Spy-mac Cas9 and demonstrated that a cytidine and adenine base editor containing the Spymac domain can induce efficient C to T and A to G conversions in vivo. Additionally, Liu et al. suggested that the PAM range of Spy-mac Cas9 may be 5'-TAAA-3' rather than 5'-NAA-3' as reported by Jakimo et al. (See Liu et al. Cell Discovery (2019) 5:58, incorporated herein by reference).

[0329] Any of the above-noted references regarding Cas9 variants or Cas9 equivalents are hereby incorporated by reference in their entirety, even if not already claimed as such.

[0330] The following table provides a comparison of the PAM preferences targeted by the Nme2Cas variants of the present disclosure relative to those targeted by other Cas homologs (R = any purine, Y = any pyrimidine, N = any nucleotide). Table 6 - PAM preferences of exemplary Cas homologs [Table D]

[0331] PAM / protospacer sequence Base editing requires the presence of a protospacer adjacent motif (PAM) located approximately 15 base pairs from the target nucleotide(s) of a classical (i.e., S. pyogenes Cas9-derived) base editor. Each programmable DNA-binding protein domain recognizes a different PAM sequence. Only about one-quarter of pathogenic transition point mutations have a suitably positioned classical PAM "NGG" sequence that is compatible with S. pyogenes Cas9 (SpCas9)-derived base editors. Naturally occurring cytidine deaminases are also known as S. aureus Cas9 (SaCas9). 98 , SaCas9-KKH 8 , Cas12a(Cpf1) 9,10 , SpCas9-NG 11 , and circularly permuted CP-Cas9s 7 It has shown broad compatibility with many Cas homologs, including , greatly expanding their targeting range.

[0332] In some embodiments, the napDNAbp comprises a PAM sequence and a protospacer positioned upstream of the PAM sequence. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence TGG. In other embodiments, the protospacer sequence is upstream of a PAM having the sequence GGG. In still other embodiments, the protospacer sequence is upstream of a PAM having the sequence AGG. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence CGG. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence AGACCC. In other embodiments, the protospacer sequence is upstream of a PAM having the sequence ACCTCA. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence GGGGCG. In other embodiments, the protospacer sequence is upstream of a PAM having the sequence CAGCCG. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence GCGGCT. In still other embodiments, the protospacer sequence is upstream of a PAM having the sequence GGGGCA. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence AAGGGT. In other embodiments, the protospacer sequence is upstream of a PAM having the sequence TCGGGT. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence GAGAGT. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence CAGAAT. In some embodiments, the protospacer sequence is upstream of a PAM having the sequence CTGGGT.

[0333] In some embodiments, the intended base pair to be edited is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream of the PAM site. In some embodiments, the intended base pair to be edited is downstream of the PAM site. In some embodiments, the intended base pair to be edited is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream of the PAM site. In some embodiments, the method does not require a classical (e.g., NGG) PAM site. In some embodiments, the target region comprises a target window, wherein the target window comprises the target nucleic acid base pair.

[0334] Protospacer sequences of the present disclosure may include, but are not limited to, the following sequences: [Table E-1] [Table E-2]

[0335] Edit window The base editors of the present disclosure can possess a variable target region of a target window (e.g., an editing window or a deamination window) that includes the target nucleic acid base pair where the nucleotide change is to be incorporated. In some embodiments, a TadA-CD base editor has a C to T base editing window, which corresponds to protospacer positions 2-12 of a protospacer. In some embodiments, a TadA-CD base editor has a C to T base editing window, which corresponds to protospacer positions 2-12. In particular embodiments, a TadA-CD base editor has a C to T base editing window, which corresponds to protospacer positions 3-8. The base editors of the present disclosure can have particularly high editing activity for cytosines between protospacer positions 5-7.

[0336] In some embodiments, the target window (e.g., editing window) comprises 1 to 10 nucleotides. In some embodiments, the editing window is 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, or 1 nucleotide in length. In some embodiments, the target window (e.g., editing window) is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. In some embodiments, the intended base pair to be edited is within the editing window. In some embodiments, the editing window includes the intended base pair to be edited.

[0337] In certain cases, the TadA-CD base editing window starts after position 2, after position 3, after position 4, after position 5, after position 6, after position 7, after position 8, after position 9, after position 10, and after position 11 of the protospacer. In some embodiments, the editing window ends before position 12, before position 11, before position 10, before position 9, before position 8, before position 7, before position 6, before position 5, before position 4, and before position 3 of the protospacer.

[0338] In some embodiments, a TadA-CD base editor comprising the V106W mutation has a narrower editing window relative to a TadA-CD base editor lacking said mutation. For example, the base editing window of TadA-CDa (SEQ ID NO: 34) is between approximately position 4 and approximately position 9 of the protospacer. In certain embodiments, TadA-CD base editors comprising the V106W mutation (e.g., TadA-CDa V106W and TadA-CDd V106W) possess a C to T base editing window between position 3 and position 9 of the protospacer, or any combination thereof. For example, the editor can incorporate a C to T substitution at position 3, position 4, position 5, position 6, position 7, position 8, or position 9 of the protospacer, or any combination thereof.

[0339] In some cases, the TadA-CD V106W base editing window starts after position 2, after position 4, after position 5, after position 6, after position 7, after position 8, or after position 9 of the protospacer. In some embodiments, the editing window ends before position 10, before position 9, before position 8, before position 7, before position 6, before position 5, or before position 4 of the protospacer.

[0340] In some embodiments, the TadA-CD base editor has an A to G base editing window between about positions 4 and 7 of the protospacer. In some cases, the TadA-CD base editor incorporates A to G editing at position 4, position 5, position 6, or position 7 of the protospacer, or any combination thereof. According to some embodiments, the A to G base editing properties of TadA-CD can be narrowed to between positions 5 and 7 of the protospacer by including the V106W mutation.

[0341] Those skilled in the art will appreciate that the TadA-CD base editor described above and herein has a narrower C to T base editing window than some existing cytidine deaminases, such as rAPOBEC1, evoAPOBEC1 (evoA), evoFERNY, and YE1. For example, BE4max and evoABE4max exhibit a C to T editing window ranging from position 1 to position 14 of the protospacer, evoFERNY-BE4 exhibits a C to T editing window from position 1 to about position 10; and YE1-BE4 exhibits a C to T editing window from position 3 to position 9 of the protospacer (see Figure 3).

[0342] Those skilled in the art will also appreciate that the TadA-CD base editor described above and herein possesses a narrower A to G and a wider C to T base editing window compared to the parent adenosine deaminase (e.g., TadA-8e) from which it was evolved. For example, TadA-8e exhibits an A to G base editing window between positions 1 and 15 of the protospacer, and a C to T base editing window between positions 4 and 7 of the protospacer.

[0343] The TadA-CD base editor of the present disclosure can convert one or more targeted cytosines within a protospacer sequence to thymines. For example, in some embodiments, TadA-CB can convert two cytosines, three cytosines, four cytosines, or five cytosines within a protospacer sequence.

[0344] Editing Efficiency Aspects of the present disclosure relate to the efficiency of the cytosine base editors described herein for editing a DNA target sequence within a target region of a target window containing a target nucleic acid base pair. In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the intended base pairs are edited. In some embodiments, the C to T conversion efficiency of any of the disclosed base editors or methods using these base editors is at least 80% across all sequencing reads. In certain embodiments, TadCBEa achieved an average conversion efficiency of 51-60% for targeted cytosines.

[0345] In some embodiments, any of the disclosed base editors or methods of using these base editors provide an average of 70% cytosine conversion efficiency in clinically relevant genes, such as the CXCR5 and CCR5 genes involved in HIV / AIDS.

[0346] In some embodiments, the cytidine deaminating activity of the disclosed deaminases (and therefore the cytosine editing activity of the disclosed base editors) exceeds the adenosine deaminating activity of the deaminase by a significant ratio. For example, the ratio of cytidine deaminating activity to adenosine deaminating activity of the disclosed Tad-CD deaminases is at least about 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 17:1, 19:1, 20:1, 21:1, 23:1, 25:1, 30:1, or greater than 30:1. In some embodiments, the ratio of cytidine deaminating activity to adenosine deaminating activity of the deaminase is at least about 10:1. In some embodiments, the ratio is at least about 20:1. In some embodiments, the ratio is about 5:1 to 7.5:1, 7.5:1 to 9.5:1, 5:1 to 10:1, 10:1 to 15:1, 15:1 to 20:1, 10 to 17:1, 12:1 to 17:1, 20:1 to 21:1, 21:1 to 25:1, 20:1 to 30:1, 25:1 to 35:1, 30:1 to 35:1, 30:1 to 40:1, 40:1 to 42:1, 21:1 ~42:1, 25:1~40:1, 10:1~40:1, 25~45:1, 30:1~50:1, 45:1~50:1, 50:1~60:1, 55:1~65:1, 60:1~70:1, 70:1~80:1, 80:1~85:1, 10:1~80:1, 40:1~80:1, 20:1~60:1, 20:1~80:1, or 75:1~85:1.

[0347] In some embodiments, the peak editing efficiency of TadA-CD is comparable to that of a native cytosine base editor (e.g., a BE4max editor containing APOBEC1, evoFERNY, or evoA deaminase). In some embodiments, the editing efficiency of the TadA-CD base editor is higher relative to that of a native cytosine base editor. For example, in some embodiments, TadA-CDa, TadA-CDb, and TadA-CDc edit the Nme50 gene at positions 3-8 of the protospacer with efficiencies between 5 and 48%. In some cases, TadA-CD contains a V106W substitution, which narrows the editing window of the base editor while maintaining editing efficiency.

[0348] In some embodiments, the disclosed TadCBEs and editing methods comprising contacting DNA with any of the disclosed TadCBEs result in an on-target DNA (C to T) base editing efficiency of at least about 20%, 21%, 25%, 30%, 35%, 40%, 50%, 60%, 70%, 80%, 85%, or greater than 85% at the target nucleic acid base pair in all sequencing reads. The contacting step can result in a C to T base editing efficiency of at least about 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 52%, 55%, 60%, 62%, 65%, 70%, 72%, 75%, 80%, 82%, 85%, or greater than 85%. In particular, the contacting step results in an on-target base editing efficiency of greater than 75%. In certain embodiments, ba...

Claims

[Claim 1] The invention described herein.