AAV vectors encoding base editors and uses thereof

JP2025515503A5Pending Publication Date: 2026-05-12THE BROAD INST INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
THE BROAD INST INC
Filing Date
2023-04-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Current methods for delivering base editors, such as adenine and cytosine base editors, into cells using adeno-associated virus (AAV) vectors are limited by the large size of these editors, which exceeds the packaging capacity of AAVs, necessitating dual AAV vectors and complicating in vivo delivery.

Method used

Development of minimized size AAV vectors that package base editors and associated regulatory elements, allowing for single AAV delivery of complete base editors without the need for transsplicing inteins, and incorporating optimized regulatory components and guide RNA encoding to maintain editing efficiency.

Benefits of technology

The single AAV base editing system achieves similar or improved editing efficiency in various tissues at multiple doses, reduces the required AAV dose, and simplifies application, characterization, and manufacturing, while minimizing unnecessary byproducts and toxicity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Described herein are nucleic acid molecules, compositions, recombinant AAV (rAAV) particles, kits, and methods for delivering base editors (or "nucleobase editors") to cells, for example via AAV vectors. In particular, the present disclosure provides compositions, methods, and uses for the delivery of adenine and cytosine base editors in a single AAV vector (or genome). Further described herein are improved AAV vectors that contain size-minimized regulatory components that allow for packaging of base editors, for example. Provided herein are methods and compositions for delivering base editor proteins to cells or tissues in a single recombinant AAV (rAAV) vector. Contemplated herein are improved methods and compositions for delivering these base editors in vivo in a single rAAV particle. Further provided herein are base editors, as well as compositions and cells that include these base editors.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims priority under 35 USC § 119(e) to U.S. Provisional Patent Application No. 63 / 336,064, filed April 28, 2022, and U.S. Provisional Patent Application No. 63 / 389,796, filed July 15, 2022, each of which is incorporated by reference herein.

[0002] Government Assistance Provisions This invention was funded by grant numbers UG3AI150551, U01AI142756, R35GM118062, RM1HG009490, R01EY009339, R01HL148769, and R35HL145203 awarded by the National Institutes of Health. The Government has certain rights in this invention.

[0003] Electronic Sequence Listing Reference The contents of the electronic sequence listing (B119570158WO00-SEQ-JQM.XML; size: 447,661 bytes; creation date: April 28, 2023) are incorporated herein by reference in their entirety. [Background technology]

[0004] 2. Background of the Invention Gene editing offers the potential for clinically validated treatment of a wide variety of genetic disorders for which only a few therapeutic options are available. Because the study and treatment of most genetic disorders through gene editing requires in vivo editing, a clinically useful method for mediating efficient delivery of high-precision gene editing agents into cells of tissues in animals, such as mammals, is required. 1,74 continues to play a vital role in advancing the field of technology.

[0005] Adeno-associated viruses (AAV) are animal models of human disease 3,4 , Clinical Trials 5 , and FDA-approved drugs 6,7AAV has been used to deliver genes encoding many therapeutic proteins in humans. AAV has become a popular in vivo delivery method due to its clinical validation, its ability to target various clinically relevant tissues, and its relatively well-understood and favorable safety profile. Summary of the Invention [Means for solving the problem]

[0006] Summary of the Invention In some aspects, described herein are nucleic acid molecules, compositions, recombinant AAV (rAAV) particles, kits, and methods for delivering complete base editors (or "nucleobase editors") to cells, for example via a single AAV vector (or genome). In particular, the present disclosure provides compositions, methods, and uses for the delivery of size-minimized adenine and cytosine base editors in a single AAV vector, where the adenine base editor and associated regulatory elements have a length shorter than the packaging capacity of AAV of approximately 4.9 kilobases (kb). Further described herein are improved AAV vectors containing size-minimized regulatory components that allow packaging of base editors. The present disclosure provides host cells and compositions comprising the disclosed rAAV particles. The present disclosure further provides improved AAV vectors containing size-minimized regulatory components that allow packaging of larger transgenes other than base editors.

[0007] Base Editor 8,9 (BE) can efficiently incorporate targeted mutations in a variety of therapeutically relevant cell types in vitro and in animal models of human genetic disease. 1,10BE can also efficiently incorporate targeted mutations into various therapeutically relevant tissues in subjects, such as human subjects, including liver tissue. Unlike nuclease-mediated gene editing, base editing does not require double-stranded DNA breaks and therefore produces minimal unwanted indel by-products, chromosomal translocations, and other mutations. 11 , chromosomal aneuploidy 12 , large deletion 13,14 , p53 activation 15,16 , or Chromosipsis 17 Although base editors can correct point mutations that cause a variety of genetic diseases, their delivery to subjects in vivo is complicated by their large size (approximately 5.2 kb), which typically exceeds the maximum packaging capacity of adeno-associated viruses (AAVs), which is approximately 4.9 kb between the inverted terminal repeats (ITRs). 18,19 Since the optimal packaging capacity of AAV is approximately 4.7 kb, delivery of base editors that can act on target DNA in a single AAV particle presents a major obstacle. For example, it has not been feasible to package a BE containing the commonly used Streptococcus pyogenes Cas9 protein (SpCas9, which is approximately 4.2 kb in length) and a single guide RNA (sgRNA) on a single (single) vector. Although technically feasible, this approach leaves little room for customized expression and control elements, such as nuclear localization sequences (NLSs). In addition to the base editor itself, the AAV that delivers the base editor must also include promoters that drive the guide RNA, base editor and sgRNA expression, as well as cis-regulatory elements.

[0008] Previous research 20-26In, AAV was used to deliver base editors by splitting the base editor into two halves between two AAV nucleic acid vectors. See U.S. Patent Application Publication No. 2018 / 0127780, published May 10, 2018; International Publication No. WO 2020 / 236982, published November 26, 2020; Levy, JM, et al. Nat Biomed Eng 4, 97-110 (2020); Chen, Y., et al. Development of Highly Efficient Dual-AAV Split Adenosine Base Editor for In Vivo Gene Therapy. Small Methods 4, 2000309 (2020); and Villiger, L. et al. Nature Medicine 24, 1519-1525 (2018), each of which is incorporated herein by reference. In a dual AAV approach, each "split" part of the base editor transgene is fused to a small trans-splicing intein. 27 or each part is expressed as a trans-spliced ​​mRNA 28 Typically, each of the two AAV vectors is packaged into a separate AAV viral particle (or virion). The dual AAV approach relies on the incorporation of trans-splicing inteins, which mediate the reconstitution of the full-length BE from the split portions in cells following delivery and transduction of the AAV particles. Since two AAV particles are required to deliver a single base editor, two successful and relatively simultaneous transductions of target cells are necessary.

[0009] While dual AAV delivery of base editors has supported therapeutic-level editing, including in mouse models of human disease, the development of a single AAV base editing system would further increase its potential impact by simplifying the application, characterization, and manufacture of base editors, and potentially increasing editing efficiency by obviating the need for co-transduction of multiple AAVs. The single AAV base editing system disclosed herein also reduces the required dose of AAV, an important advancement since clinical application of AAV is often constrained by dose-limiting toxicity. 29 Therefore, there is a need in the art for single AAV in vivo base editor delivery that requires only a single transduction of target cells that retains high editing potential. Single AAV delivery would provide advantages for research and clinical use, especially in mammalian tissues that are more difficult to effectively transduce, such as cardiac, CNS, and muscle tissues.

[0010] Adenine base editors (ABEs) are a particularly useful class of editing agents because they incorporate A·T to G·C transversions that correct approximately half of all known pathogenic SNPs. 9 Phage-assisted continuous evolution (PACE) of the adenine base editor ABE7.10 recently yielded the deoxyadenosine deaminase TadA-8e with expanded compatibility with Cas domains other than SpCas9 and increased activity. 30ABE7.10, which contains TadA7.10 deaminase, can perform clean and efficient A·T to G·C conversion in DNA with very low levels of unwanted by-products, such as small insertions or deletions (indels), in cultured cells, adult mice, plants, and other organisms. Additional details about TadA-8e and TadA7.10 deaminase can be found in WO 2021 / 158921, published August 12, 2021; WO 2018 / 027078, published February 8, 2018; WO 2019 / 079347, published April 25, 2019; Koblan et al., Nat Biotechnol 36, 843-846 (2018); and Gaudelli et al., Nature 551, 464-471 (2017). Each of these is incorporated herein by reference. ABEs containing only a single TadA deaminase domain rather than a single-chain dimer allow for a reduction in editor size. 30,31 Moreover, although SaCas9 is small enough (1053 amino acids in length, SEQ ID NO: 377) to provide a base editor compatible with a single AAV, its utility is severely limited by the rarity of its NNGRRT PAM. Because base editing requires the presence of a suitable PAM to place the target nucleotide within the editing window, an ABE that collectively offers broad PAM compatibility combined with simple and efficient in vivo delivery would advance the in vivo application of base editing.

[0011] Similarly, cytosine base editors (CBEs) that offer broad PAM compatibility combined with simple and efficient in vivo delivery would advance in vivo applications of base editing. Current CBEs contain a uracil glycosylase inhibitor domain, which is approximately 84 bp in length. Although not very large, these additional 84 base pairs make delivery of CBEs in a single AAV vector more difficult than ABEs.

[0012] Ran et al. engineered a size-minimized S. aureus Cas9 (SaCas9) for delivery in a single AAV vector in vivo to incorporate a double-stranded break in the target genomic DNA. See Ran, FA, et al. (2015) Nature 520(7546): 186-191, incorporated herein by reference. Ran's AAV cassette contained a sgRNA driven by a U6 promoter and a SaCas9 transgene driven by a cytomegalovirus (CMV) promoter or a thyroxine-binding globulin (TBG) promoter. Recently, Tran et al. engineered a single AAV vector containing a size-minimized SaCas9 ABE, microABE I744, in which an ABE7.10 TadA deaminase monomer is fitted (i.e., inserted) into the SaCas9 domain. See Tran et al., Nat.Commun. 11, 4871 (2020), incorporated herein by reference. However, this AAV-encoded ABE only showed <0.25% in vitro editing and was not evaluated in vivo. More recently, Zhang, Sontheimer et al. generated a single AAV vector encoding an ABE containing the N. meningitidis 2 Cas9 (Nme2Cas9) protein. See Zhang et al., Adenine Base Editing in vivo with a Single Adeno-Associated Virus Vector. bioRxiv 2021.12.13.472434, incorporated herein by reference. However, Zhang's base editor showed significant editing at bystander adenines and a maximum editing efficiency of 35% after AAV delivery to liver tissue. However, single AAV delivery of base editors containing other compact Cas9 protein domains that exhibit highly efficient base editing, similar to the currently used ABE and CBE, has not been disclosed, and in addition, single AAV delivery of base editors into cardiac and muscle tissue has not been disclosed.

[0013] The present disclosure provides a size-minimized AAV vector with a length of less than about 4.90 kb between ITRs.This single AAV base editing platform provides similar or improved editing efficiency in various tissues at multiple doses compared to dual AAV ABE8e when delivered systemically to mice.The exemplary AAV vector of the present disclosure does not incorporate or rely on the use of trans-splicing intein for successful delivery.

[0014] The AAV vectors of the present disclosure are based, at least in part, on advances in genetic engineering that have produced vectors that contain size-minimized components required for efficient expression in target cells and editing of target bases in vivo. Exemplary target cells include muscle cells, neurons, liver cells, neuromuscular cells, and cardiac cells. These vectors are smaller than the vectors disclosed in US Patent Application Publication No. 2018 / 0127780, published May 10, 2018, and International Publication No. WO 2020 / 236982, published November 26, 2020, and are therefore compatible with incorporation into a single AAV particle. In particular, the disclosed AAV vectors are based, in part, on the discovery that post-transcriptional response elements such as WPRE in transcription terminators (or polyadenylation signals) are not required for successful expression of base editors in target tissues in vivo. Thus, the disclosed AAV vectors contain shorter (or size-minimized) terminators. The disclosed AAV vectors further contain other size-minimized regulatory elements, such as short promoters.

[0015] The disclosed AAV vector is also based in part on the discovery that a guide RNA compatible with any of the disclosed Cas proteins can be encoded at the 3' end of the vector and maintain a total length of ITR to ITR of less than 4.9kb, less than 4.8kb, less than 4.7kb, less than 4.6kb, or less than 4.5kb. Thus, through the genetic engineering techniques described herein, the base editor and its guide can be efficiently packaged into a single rAAV particle for in vivo delivery. The guide RNA can be encoded in the opposite orientation (3' to 5') to that of the promoter (and its replication origin) that drives the expression of the base editor and base editor transgene on the vector (i.e., 5' to 3').

[0016] The present disclosure also provides size-minimized base editors. These base editors were developed to enable efficient in vivo base editing mediated by a single AAV particle. The disclosed AAV-encoded base editors may include size-minimized Cas proteins. These Cas9 proteins are about 1000-1050 amino acids in length, which is about 350 amino acids shorter than the SpCas9 protein. These size-minimized Cas proteins include, but are not limited to, S. aureus Cas9 (SaCas9), Nme2Cas9, C. jejuni Cas9 (CjCas9), S. auricularis Cas9 (SauriCas9), and variants of any of these Cas9 proteins. The disclosed base editors may contain any of these Cas9 proteins, or evolved or mutated variants of any of the Cas9 proteins disclosed herein.

[0017] Thus, in various embodiments, the disclosed AAV nucleic acid molecules do not include an intein, such as a trans-splicing intein (e.g., does not include a trans-splicing intein from Nostoc punctiforme or Npu). In various embodiments, the disclosed AAV nucleic acid molecules include a transcription terminator that does not include a post-transcriptional response element. In some embodiments, the disclosed AAV nucleic acid molecules do not include an intein or a post-transcriptional response element. In some embodiments, the nucleic acid molecule includes a first nucleic acid fragment including: (i) a 5' inverted terminal repeat (ITR); (ii) a sequence encoding a base editor operably linked to a first promoter, where the base editor includes a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain; and a polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a second promoter; and (iv) a 3'ITR.

[0018] In some aspects, provided herein is an rAAV vector with a regulatory element whose size is minimized, allowing large transgene packaging.Thus, provided herein is an rAAV nucleic acid molecule, comprising, in the order of 5' to 3', (i) a 5' inverted terminal repeat (ITR); (ii) a transgene operably linked to a first promoter, wherein the first promoter has a length of less than 300 nucleotides; and a first nucleic acid fragment comprising a transcription terminator that does not contain a post-transcriptional response element; (iii) a second nucleic acid fragment operably linked to a second promoter, wherein the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the first nucleic acid fragment; and (iv) a 3'ITR.In various embodiments, the length between the 5'ITR and the 3'ITR is less than about 4.90 kb. In some embodiments, the length, between the 5'ITR and the 3'ITR, is less than about 4.85 kb, less than about 4.80 kb, less than about 4.75 kb, less than about 4.725 kb, less than about 4.70 kb, or less than 4.65 kb.

[0019] In some embodiments, the first nucleic acid fragment encodes a base editor and the second nucleic acid fragment encodes a gRNA. In some embodiments, the first nucleic acid fragment encodes a protein that is not a base editor.

[0020] Any of the disclosed base editors may include (i) a napDNAbp domain; and (ii) a deaminase domain. The disclosed base editors may include a wild-type napDNAbp domain (e.g., wild-type SaKKH-Cas9). The disclosed base editors may include a napDNAbp domain with nickase activity (e.g., SaKKH-Cas9 nickase or "SaKKH"). In some embodiments, the napDNAbp domain is a Cas9 nickase domain. In some embodiments, the napDNAbp domain is a SaKKH-Cas9 nickase. The napDNAbp domain may be selected from S. aureus Cas9 (SaCas9), N. meningitidis 2 Cas9 (Nme2Cas9), C. jejuni Cas9 (CjCas9), or S. auricularis (SauriCas9) domains, and variants thereof. These Cas proteins have broader PAM compatibility than standard SpCas9 proteins. Thus, the disclosed single AAV-encoded base editors can potentially target the majority of adenines in a genome or the majority of cytosines in a genome. In exemplary embodiments, the napDNAbp domain is a SaCas9 domain, a SaCas9 nickase domain, a SaKKH domain, or a SaKKH nickase domain.

[0021] The ABEs encoded by the AAVs of the present disclosure contain an adenosine deaminase domain that contains a single deaminase, i.e., a deaminase monomer (such as a TadA-8e monomer), rather than an adenosine deaminase dimer (i.e., two adenosine deaminases). The use of a deaminase monomer facilitates the generation of size-minimized base editors. The TadA monomer is approximately 166 amino acids in length.

[0022] Any of the disclosed adenine base editors may comprise an adenosine deaminase domain that is a variant of E. coli TadA deaminase. In some embodiments, the adenosine deaminase is selected from TadA-8e, TadA-8e(V106W), TadA9, TadA20, and TadA7.10 deaminase. In some embodiments, any of the disclosed base editors comprises an adenosine deaminase fused to the N-terminus of a napDNAbp domain, such as a Cas9 nickase. In some embodiments, the adenosine deaminase is TadA-8e.

[0023] In some aspects, the present disclosure provides size-minimized ABE8e variants. Each variant is compatible with single AAV delivery, and the three such variants collectively provide sufficient PAM compatibility to target 87% and edit 82% of the adenines on the human genome. These three variants are Sauri-ABE8e, SaKKH-ABE8e, and SaABE8e. Each contains TadA-8e adenosine deaminase and a nickase variant of each of SauriCas9, SaKKH-Cas9, and SaCas9. The present disclosure further provides ABE8e variants CjCas9-ABE8e and Nme2Cas9-ABE8e. The present disclosure further provides ABE variants SaKKH-ABE8e(V106W), SauriCas9-ABE8e(V106W), CjCas9-ABE8e(V106W), Nme2Cas9-ABE8e(V106W), and SaCas9-ABE8e(V106W); SaKKH-ABE9, SauriCas9-ABE9, CjCas9-ABE9, Nme2Cas9-ABE9, and and SaCas9-ABE9; SaKKH-ABE20, SauriCas9-ABE20, CjCas9-ABE20, Nme2Cas9-ABE20, and SaCas9-ABE20; and SaKKH-ABE7.10, SauriCas9-ABE7.10, CjCas9-ABE7.10, Nme2Cas9-ABE7.10, and SaCas9-ABE7.10. In any of these disclosed base editors, wild-type or nickase variants of SauriCas9, SaKKH-Cas9, SaCas9, CjCas9, and Nme2Cas9 can be used.

[0024] The present disclosure further provides size-minimized CBE variants, particularly size-minimized BE3.9 variants. Examples of these variants include CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9. In some embodiments, the base editor further comprises a uracil glycosylase inhibitor (UGI) domain. The size of an exemplary cytidine deaminase (e.g., rAPOBEC1 deaminase) of the disclosed CBE is about 229 amino acids.

[0025] By integrating these developments, a single AAV-delivered base editor of the present disclosure was used to treat a mouse model of high cholesterol (which is implicated in cardiovascular disease), resulting in correction of casual mutations in cardiac tissue and increased lifespan of the animals.

[0026] In the examples of the present disclosure, single AAV delivery of ABE achieved 66%, 33%, and 22% editing in hepatic liver, heart, and muscle tissues at doses lower or similar to those used in recent preclinical and clinical trials of AAV particles targeting these tissues (see, for example, clinical trial numbers NCT02122952 (SMA treatment) and NCT03375164 (DMD treatment)). In addition, it was demonstrated that a single AAV8 ABE can efficiently edit therapeutically relevant targets in mice with efficiencies of 60-85%. These edits produced disruptions at the native splice acceptor sites in the targeted human and mouse Pcsk9 and mouse Angptl3 genes. This resulted in near complete (average 93%) knockdown of these genes at doses lower than previously reported, resulting in substantial reductions in plasma cholesterol and triglycerides. The hepatic protein proprotein convertase subtilisin / kexin type 9 (PCSK9) is a secreted, globular, autoactivating serine protease. It acts as a protein-binding adaptor in endosomal vesicles to bridge pH-dependent interactions with the low-density lipoprotein receptor (LDL-R) during endocytosis of LDL particles, preventing recycling of LDL-R to the cell surface, leading to reduced LDL-cholesterol clearance. Angiopoietin-like 3 protein (Angptl3) is an endogenous inhibitor of lipoprotein lipase (LPL), the main enzyme involved in the hydrolysis of triglyceride-rich lipoproteins. Base editors targeting disease-causing mutations on the Pcsk9 gene are disclosed in U.S. Patent Application Publication No. 2018 / 0237787, published August 23, 2018, which is incorporated herein by reference.

[0027] By minimizing the size of the adenine base editor and AAV components, a series of single AAV adenine base editor systems have been developed that have broad targeting capabilities due to their overall PAM compatibility and support robust editing in vivo. The single AAV BE vectors of the present disclosure facilitate base editing for research and therapeutic applications by simplifying production and characterization, as well as by reducing the total dose of AAV required to achieve a desired level of editing. These single AAV BEs offer several potential advantages over dual AAV approaches for clinical use: clinical-scale production of a single vector rather than two; increased titers, especially at low doses; and reduced complexity from a simpler construct that negates the need to use trans-splicing inteins. For these reasons, in vivo editing approaches compatible with single AAV delivery may be more readily applied to large animal models and human therapeutics, where systemic delivery is commonly used. The development of smaller promoters that provide sufficient expression of the base editor allows further minimization of the elements of the single AAV ABE. This should facilitate translation to the clinic by increasing the proportion of full-length packaged AAV genomes.

[0028] In another aspect, a host cell is provided that comprises the composition described herein.The disclosed cell can comprise any of the disclosed nucleic acid molecules, rAAV vectors, or rAAV particles described herein.In still another aspect, a kit is provided that comprises any of the disclosed rAAV particles and instructions for delivery to cells (such as host cells).

[0029] Yet another aspect of the present disclosure provides a method comprising contacting a target nucleic acid molecule with any of the compositions described herein.In various embodiments, the target nucleic acid molecule is in a cell, such as a eukaryotic cell (e.g., a mammalian cell).In some embodiments, the target nucleic acid molecule is genomic DNA.In some embodiments, the genomic DNA is in a cell or tissue of a subject, such as a human subject.Therefore, a method is contemplated comprising contacting a cell with any of the rAAV particles or compositions disclosed herein.

[0030] Yet another aspect of the present disclosure provides a method comprising administering a therapeutically effective amount of any of the compositions (or rAAV particles) described herein to a subject in need thereof. In some embodiments, the subject has a disease or disorder (e.g., a genetic disease). In some embodiments, the disease or condition is a cardiovascular disease. In some embodiments, the composition is administered to the liver tissue, heart tissue, or skeletal muscle tissue of the subject.

[0031] Still other aspects of the disclosure provide methods of making any of the disclosed rAAV particles and compositions.

[0032] Details of certain aspects of the invention are presented in the detailed description, as set forth below. Other features, objects, and advantages of the invention will be apparent from the definition, examples, figures, and claims. [Brief description of the drawings]

[0033] [Figure 1A-1C]Figure 1A shows the AAV constructs that were evaluated in vivo. An sgRNA targeting the Pcsk9 gene with the W8 mutation was delivered together with EGFP in one AAV. This was co-injected with either one or two additional AAVs encoding either intact or intein-split SaABE8e. In total, two AAVs were used to deliver intact SaABE8e and the sgRNA, and three AAVs were used to deliver intein-split SaABE8e. Black boxes represent the ITRs, the EFS promoter is EF1α short, W3 is the truncated WPRE, bGH is the bovine growth hormone polyadenylation signal, purple boxes are sgRNA cassettes driven by the U6 promoter in the orientation indicated by the arrow, the NpuN and NpuC split inteins from Nostoc punctiforme are shown in brown, and protein coding regions are indicated for EGFP and SaABE8e. Figure 1B shows the in vivo editing efficiency from injection of AAVs encoding intein split and intact SaABE8e. The total dose of base editor AAV administered to each mouse is shown. Figure 1C shows a comparison of in vivo editing efficiency from injection of AAV9 encoding intact SaABE8e in five different AAV architectures when administered at the doses indicated. In all cases, the editor AAV dose was either 4 x 1011 vector genomes (vg) or 4 x 1010 vg, and the sgRNA EGFP AAV dose was either 4 x 1011 vg or 4 x 1010 vg at a 1:1 ratio of base editor AAV to sgRNA AAV. The size of the delivered editor AAV construct (encompassing the ITRs) is shown in the legend. C57BL / 6J mice aged 6-7 weeks weighing 20-25 g were injected systemically by retro-orbital injection. Dots represent individual mice. Values ​​and error bars represent the mean±SEM of n=3 different mice.

[0034] [Figure 2A-2C]Development and characterization of single AAV SaABE8e. Figure 2A shows a schematic of the single AAV SaABE8e genome (5,064 bp encompassing the ITRs). The arrow indicates the orientation of the U6 sgRNA cassette. Figure 2B shows a comparison of dual vs. single SaABE8e. Dual (SaABE8e with intact editor on one genome, sgRNA and EGFP on the second genome) or single (SaABE8e and guide RNA on a single vector) AAV vectors containing either base editing activity both incorporate Pcsk9 W8R editing, are packaged in AAV9, and have matched promoters and terminators (polyA). Base editor AAVs were administered at the doses indicated in the legend (dual AAVs delivered at the indicated doses of editor AAV and sgRNA AAV, single AAVs delivered at the indicated doses) via retro-orbital injection into 6-8 week old C57BL / 6J mice weighing 20-25 grams, tissues were harvested 3 weeks after injection and analyzed by high-throughput DNA sequencing (HTS). Figure 2C shows the dose titration of single AAV SaABE8e. Dots represent individual mice, error bars represent the mean ± SEM of n=3 different mice.

[0035] [Figure 3A-3D]Characterization of Nme2ABE8e (Figure 3A), CjABE8e (Figure 3B), and SauriABE8e (Figure 3C) in HEK293T cells. Editing of each targeted adenine within the protospacer is shown. Target sites are indicated, and the sequence of each targeted protospacer and PAM is listed in Table 1. Targeted adenines are numbered for the standard protospacer length of each editor (22 nucleotides for CjABE8e, 24 nucleotides for Nme2ABE8e, and 21 nucleotides for SauriABE8e). Points represent values, and error bars represent the mean ± SEM of n=3 replicates. Figure 3D shows the percent of genomic adenines on the hg38 human reference genome that are targetable by size-minimized ABEs, both independently and collectively. A schematic showing a representative portion of the genome that is targetable by size-minimized ABEs is shown. Targetable PAMs are indicated by colored lines and targetable adenines are indicated by colored dots. The pie chart at the right end of the schematic indicates that 87% of single nucleotide adenine polymorphisms are collectively targetable (and 82% are editable) by all of the ABEs encoded by the small AAVs of the present disclosure.

[0036] [Figure 4A-4D]Characterization of Nme2ABE8e (Figure 4A), CjABE8e (Figure 4B), and SauriABE8e (Figure 4C) in HEK293T cells. For each size-minimized ABE, the base editing activity window from 8-9 genomic target sites is shown. Positions that were not present in either tested site are shaded gray. Target adenines are numbered for the standard protospacer length for each editor (24 nt for Nme2ABE8e, 22 nt for CjABE8e, and 21 nt for SauriABE8e). Points represent values, error bars represent mean ± SEM of n=3 replicates for each site, and each position represents 1-3 genomic sites. Dotted lines correspond to 25% of the average peak activity of each base editor, defining possible activity thresholds within the editing window. Individual sites containing 3 or more adenines within the protospacer were included in this analysis. Sites with only one or two adenines within the protospacer were excluded from the summary analysis, but data from all sites analyzed are shown in Figures 3A-3C. Figure 4D shows the percent of genomic adenines on the hg38 human reference genome that are targetable by size-minimized ABEs, either independently or collectively. An example is shown showing a representative portion of the human genome that is targetable by size-minimized ABEs. Targetable PAMs are indicated by colored lines and targetable adenines are indicated by colored dots.

[0037] [Figure 5A-5K]Evaluation of genome editing and plasma lipids when targeting PCSK9, Pcsk9, and Angptl3 in vivo by single AAV. Figure 5A shows the strategy for evaluating base editing and plasma analytes in AAV-treated mice, generated using BioRender. Figure 5B shows bulk liver editing efficiency in human PCSK9, as well as mouse Pcsk9 and Angptl3 (n=3-5 mice. Each point represents one mouse, error bars represent SEM). Human PCSK9 editing was performed using humanized PCSK9 mice, and mouse Pcsk9 and Angptl3 editing was performed at the endogenous mouse locus in wild-type C57BL / 6J mice. AAV was administered by retro-orbital injection at a dose of 1×1011vg / mouse at 6-8 weeks of age. Figure 5C shows dose-dependent base editing for dual SpABE8e and single SaKKH-ABE8e in mouse Pcsk9 exon 1 splice donor. Total AAV dose administered is indicated in vg / mouse below each set of bars. Total AAV dose administered is indicated in vg / mouse below each set of bars. Dual SpABE8e editing data was reported in another publication from our group2. Each point represents a different mouse (n=5). Figure 5D shows a direct comparison of editing efficiency of dual AAV8 intein split SaKKH-ABE8e and single AAV8 SaKKH-ABE8e targeting Pcsk9 exon 1 donor site in bulk liver at two doses. Total AAV dose administered is indicated in vg / mouse below each set of bars. ***P=0.0004. Each point represents a different mouse (n=4). Figure 5E shows plasma PCSK9 protein in humanized mice treated with 1x1011vg single AAV8 SaKKH-ABE8e. Figure 5F shows plasma Pcsk9 protein in C57BL / 6J mice treated with either 1x1011vg single AAV8 SaKKH-ABE8e, dual AAV8 SpABE8e, or non-targeting control. ***P=0.0001.Figure 5G shows plasma Angptl3 protein in C57BL / 6J mice treated with 1x10vg single AAV8 SaKKH-ABE8e or non-targeted control. **P=0.0027. Figure 5H shows plasma total cholesterol in humanized mice treated with 1x10vg single AAV8 SaKKH-ABE8e. Figure 5I shows plasma total cholesterol in C57BL / 6J mice treated with either 1x10vg single AAV8 SaKKH-ABE8e, dual AAV8 SpABE8e, or non-targeted control. ***P=0.0007. Figure 5J shows plasma total cholesterol***P=0.0007 and plasma triglycerides*P=0.0118. Figure 5K shows C57BL / 6J mice treated with 1x10vg single AAV8 SaKKH-ABE8e or non-targeting control. In Figures 5E-5K, points represent the mean and error bars represent SEM for n=5 different mice. Significance was calculated using a two-way unpaired t-test for Figures 5C-5D. Significance for Figures 5F-5G and Figures 5I-5K was calculated using a two-way repeated measures ANOVA with Tukey or Sidak multiple comparisons as appropriate and is shown for the 4 week time point for all graphs, with the exception of Figure 5F where the 3 week significance is shown as the 4 week protein level did not reach statistical significance. In all cases, the non-targeting control is dual AAV8 SpABE7.10 with sgRNA targeting mouse Dnmt1 at an unrelated site on the mouse genome, administered at the same time point, route, and dose.

[0038] [Figure 6] Figure 6 shows validation of SaABE targets in mouse Neuro-2A and 3T3 cells. Base editors and PAMs are annotated below each set of bars. Points represent independent biological replicates (n=3) and error bars show SEM.

[0039] [Figure 7]Figure 7 shows titration of sgRNA cassette AAV in vivo. A fixed dose of 4x1011vg of full-length SaABE8e editor AAV was delivered to C57BL / 6J mice by retro-orbital injection along with various ratios of sgRNA cassette-containing AAV. Tissues were harvested 3 weeks after injection and analyzed by HTS. Points represent individual mice (n=3) and error bars show SEM.

[0040] [Figure 8A-8B] Figure 8A shows SaABE8e activity window in Pcsk9 W8R in liver. SaABE8e maintains a wide editing window in vivo consistent with observations in cultured cells. Figure 8B shows that indels remain low under all conditions, reaching 2.4%, 1.6%, and 1.1% indels in male and female heart, muscle, and liver tissues at the high dose of 8x10vg single AAV SaABE8e. Points represent individual mice (n=3) and error bars show SEM.

[0041] [Figure 9A-9C] Validated guides targeting PCSK9 in HEK293T cells (Figure 9A), and Pcsk9 (Figure 9B) and Angptl3 (Figure 9C) in mouse Neuro-2a cells. For each sgRNA, the protospacer positions that disrupt the indicated targets are indicated for the exon, type of target (start codon, splice donor, or splice acceptor), ABE8e variant (SaABE8e or SaKKH-ABE8e), and 22 nt protospacer length. Edits at the indicated target-disrupting positions are plotted. Points represent independent replicates (n=2) and error bars indicate standard deviation (SD).

[0042] [Figure 10A-10B]Figures 10A and 10B show editing in Neuro-2A cells by SauriABE8e at the Pcsk9 exon 1 splice donor target. Figure 10A shows targeting of mouse Pcsk9 exon 1 splice donor by SauriABE8e and SaKKH-ABE8e in mouse Neuro-2A cells. The target adenine is A9 for the Sauri protospacer and A5 for the SaKKH protospacer. Figure 10B shows a comparison of the SaCas9 guide RNA backbone33,74 for editing activity at mouse Pcsk9 exon 1 splice donor. Because the native sgRNA of SauriCas9 is unknown, the SaCas9 sgRNA backbone is used with the homologous SauriCas9 protein.

[0043] [Figure 11A-11B] In vivo editing of control AAV for lipid modification experiments. 6-8 week old C57BL / 6J mice were injected by retro-orbital injection and whole livers were analyzed by HTS 4 weeks later. Figure 11A shows editing in Dnmt1 A41A (silent editing) by dual SpABE7.10 at a dose of 1x1011vg dual AAV8. Figure 11B shows incorporation of Pcsk9 W8R substitution using single SaKKH-ABE8e at a dose of 1x1011vg single AAV8.

[0044] [Figures 12A-12D]Dose response of single AAV8 SaKKH-ABE8e and dual AAV8 SpABE8e (with sgRNA targeting Pcsk9 exon 1 splice donor site) on plasma Pcsk9 and total cholesterol. Figure 12A shows circulating Pcsk9 protein and Figure 12B shows total cholesterol from plasma taken weekly. Normalized to baseline. Figure 12C shows circulating Pcsk9 protein and Figure 12D shows total cholesterol from raw (not normalized) weekly plasma. Points represent mean values ​​and error bars represent SEM of n=5 different mice. All mice were administered systemically at 6-8 weeks of age by retro-orbital injection with the total dose of AAV8 indicated in the legend and blood samples were removed sequentially over 4 weeks.

[0045] [Figures 13A-13E] Raw (non-normalized) levels of plasma analytes from either single AAV ABE or non-targeting control dual AAV ABE mice for human PCSK9 and mouse Angptl3 targets. Figure 13A shows ELISA of human PCSK9 in plasma from humanized mice. Figure 13B shows total plasma cholesterol in humanized PCSK9 mice. Figure 13C shows ELISA of mouse Angptl3 in plasma from C57BL / 6J mice. Figure 13D shows total plasma cholesterol in C57BL / 6J mice. Figure 13E shows plasma triglycerides from C57BL / 6J mice. Points represent mean values ​​and error bars represent SEM of n=5 different mice. Non-targeting control is dual AAV ABE7.10 targeting Dnmt1. All mice were administered a dose of 1x10vg AAV8 systemically by retro-orbital injection at 6-8 weeks of age and blood samples were removed sequentially over 4 weeks.

[0046] [Figure 14]14 is a schematic diagram showing the construction of a single AAV vector expressing a cytosine base editor that includes a uracil glycosylase domain. The Cas9 domain of the cytosine base editor is CjCas9. The sgRNA controlled by U6 is located in the reverse orientation at the 3' end. The construct has a total length of 5.012 kb including the ITRs.

[0047] [Figure 15] FIG. 15 shows a base editor-matched comparison of guide-dependent on-target editing between single and dual AAV SaKKH-ABE8e at the Pcsk9 exon 1 splice donor site.

[0048] [Figures 16A-16C] Figures 16A-16C show guide-dependent off-target DNA editing analysis in vivo and in culture between single AAV8 SaKKH and each half of AAV8 intein split SaKKH-ABE8e. Figure 16A shows in vivo editing in liver tissue from single AAV ABE-treated mice. The top three predicted off-target ("OT") sites of SaKKH-ABE8e targeting the Pcsk9 exon 1 splice donor were sequenced from liver tissue. Figure 16B shows that editing observed at OT2 is dose-dependent. OT, off-target; NT, dose-matched non-targeting dual AAV8 ABE7.10 targeting an unrelated gene Dnmt1. Figure 16C shows editing in cell culture. After plasmid transfection of N2A cells with sgRNA targeting the Pcsk9 exon 1 donor site and full-length or intein split SpABE8e, on- and off-target sites were sequenced. Full-length and intein-split SpABE8e did not differ significantly in efficiency in on- or off-target editing by unpaired multiple t-tests with the Holm-Sidak method to correct for multiple comparisons.

[0049] [Figure 17]Figure 17 shows that in vivo mRNA off-target editing is undetectable in single AAV ABE8e-treated mice. Dots represent single adenines on each amplicon for n=4 mice.

[0050] [Figure 18] FIG. 18 shows a comparison of the editing windows of exemplary single AAV-encoded ABEs in the liver.

[0051] [Figure 19A-19B] FIG. 19A, histopathological evaluation by hematoxylin and eosin staining of liver from untreated mice and FIG. 19B, mice 4 weeks after treatment with 1×10 vg of single AAV8 SaKKH-ABE8e targeting human PCSK9. Representative images are shown. Scale bar, 50 μm.

[0052] [Figure 20] Figure 20 shows quantification of AAV genomes from tissues encoding SaABE8e dual AAV (SaABE8e with intact editor on one genome, sgRNA and EGFP on the second genome) or single AAV (SaABE8e and guide RNA, all-in-one). Both incorporate Pcsk9 W8R and are packaged into AAV9. Editors were packaged into AAV9 and administered by retro-orbital injection at doses indicated in legend at 6-8 weeks of age (dual AAVs were delivered at doses indicated in legend for editor AAV and sgRNA AAV, single AAVs were delivered at total doses indicated). Tissues were harvested 3 weeks after injection and then editor AAV genomes were quantified by ddPCR using SaCas9 primers and probes and normalized to Gapdh. Points represent individual mice (n=1-3) and error bars indicate SEM.

[0053] [Figure 21] FIG. 21 shows alkaline gel electrophoresis of the packaged AAV genome.

[0054] [Fig. 22A-22D] Figures 22A-22D show off-target mRNA editing. RNA was extracted from single AAV8 SaKKH-ABE8e treated and untreated mouse livers, reverse transcribed, and cDNA amplicons from Aars, Canx, Ctnnb, and Usp38 mRNA were analyzed by HTS. Dots represent individual adenines on the sequenced amplicons (n=3 mice). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0055] Detailed Description of the Invention Provided herein are methods and compositions for delivering base editor proteins to cells or tissues in a single recombinant AAV (rAAV) vector. Contemplated herein are improved methods and compositions for delivering these base editors in vivo (such as to a subject in need thereof) in a single rAAV particle. Delivery in a single rAAV particle has the advantage of requiring fewer injections and lower doses for delivery to the target tissue, as well as reducing by half the number of successful transductions of the target tissue required for expression of the base editor in the target cell. These rAAV vectors contain size-minimized base editors and regulatory components that allow the vector to have a length within the 4.7 kb to 4.9 kb packaging capacity of the rAAV particle. Also provided are rAAV particles containing any of the disclosed rAAV vectors and capsid proteins, as well as compositions and cells comprising same. Methods of administering such compositions and cells to a subject are further provided. Additionally provided are base editors, as well as compositions and cells comprising these base editors. The disclosed single AAV adenine base editors provide similar or enhanced editing efficiency compared to dual AAV editors in various tissues in vivo.

[0056] Recombinant AAV vectors (or AAV genomes) are widely used for transgene delivery. Transgenes are inserted into the AAV genome between inverted terminal repeat (ITR) sequences and packaged into AAV viral particles, which are used to transduce host cells (e.g., mammalian cells, human cells). AAV has been used to deliver genes encoding many therapeutic proteins in animal models of human disease, in clinical trials, and in FDA-approved drugs. A range of available AAV serotypes provides access to a variety of clinically relevant cell types in mice, non-human primates, and humans.

[0057] The present disclosure provides rAAV vectors with size-minimized regulatory elements that allow for packaging of larger transgenes than prior art vectors. In some embodiments, the transgene encodes a base editor, such as an adenine base editor. In some embodiments, the transgene is not a base editor. In some embodiments, the base editor contains a compact protein napDNAbp domain, such as S. aureus Cas9 (SaCas9), N. meningitidis 2 Cas9 (Nme2Cas9), C. jejuni Cas9 (CjCas9), or S. auricularis (SauriCas9) domain, or a variant thereof.

[0058] In some aspects, provided herein is a rAAV vector comprising a first nucleic acid fragment comprising: (i) a 5'ITR; (ii) a sequence encoding a base editor operably linked to a first promoter, where the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain; and a polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a second promoter; and (iv) a 3'ITR, where the length between the 5'ITR and the 3'ITR is less than about 4.90 kb. In some embodiments, the rAAV vector consists essentially of components (i)-(iv).

[0059] In some embodiments, the nucleic acid vector is the genome of adeno-associated virus packaged in rAAV particle.In some embodiments, the first and / or second nucleic acid fragment is operably linked to a first promoter.In some embodiments, the first promoter is a constitutive promoter.In some embodiments, the first promoter is an inducible promoter.

[0060] In various embodiments, the first promoter is a short promoter. In various embodiments, the first promoter has a length of less than 325 nucleotides, less than 300 nucleotides, less than 285 nucleotides, less than 270 nucleotides, less than 265 nucleotides, or less than 250 nucleotides. This short length ensures optimal packaging capacity for the base editor and additional regulatory elements. In some embodiments, the first promoter has a length between 200 and 250 nucleotides, between 225 and 255 nucleotides, between 250 and 275 nucleotides, between 275 and 300 nucleotides, or between 300 and 325 nucleotides. In certain embodiments, the first promoter has a length of 280 nucleotides. In some embodiments, the first promoter has a length of 229 nucleotides, 253 nucleotides, or 324 nucleotides.

[0061] In some embodiments, the first promoter is a tissue-specific promoter. The first promoter can be a cardiac tissue-specific promoter, a muscle tissue-specific promoter, or a neuronal tissue-specific promoter. The first promoter can be active in tissues other than liver, muscle, and neuron, such as eye tissue. In some embodiments, the first promoter is active in neuromuscular tissue.

[0062] In some embodiments, the first promoter is the EF-1α (short) ("EFS") promoter, which is an intronless form of EF-1α. In some embodiments, the first promoter is the MeCP2 promoter, which is active in neuronal tissue. In some embodiments, the first promoter is the P3 promoter, which is active in liver tissue. In some embodiments, the first promoter is the U1a promoter, which is active in liver tissue. Additional details about the MeCP2 promoter, P3 promoter, and U1a promoter can be found in Gray et al., Human Gene Therapy.Sep 2011.1143-1153; Viecelli et al., Hepatology, 60: 1035-1043 (2014); and Ibraheim et al., Genome Biol. 19, 137 (2018). Each of these is incorporated herein by reference.

[0063] In some embodiments, the first nucleic acid fragment comprises a transcription terminator. In various embodiments, the first nucleic acid fragment does not contain a post-transcriptional response element (e.g., W3). In some embodiments, the transcription terminator is a polyA signal selected from bovine growth hormone (bGH) signal, human growth hormone (hGH) signal, or SV40 signal. In some embodiments, the terminator is a bGH polyA signal. In some embodiments, the terminator is an SV40 late polyA signal.

[0064] In some embodiments, the first nucleic acid fragment comprises a minimal virus of mice (MVM) intron. In some embodiments, the MVM is located 5' of the promoter and 3' of the transgene.

[0065] In some embodiments, the second nucleic acid fragment comprises a nucleotide sequence encoding a gRNA operably linked to a second promoter. In some embodiments, the second promoter is a constitutive promoter. In some embodiments, the second promoter is an inducible promoter. In some embodiments, the second promoter is a U6 promoter, such as a human U6 promoter. In some embodiments, the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the first nucleic acid fragment.

[0066] The disclosed rAAV vector, which contains a first nucleic acid fragment containing a promoter and a terminator and a second nucleic acid fragment that can encode a guide RNA, has a packaging capacity between 5'ITR and 3'ITR of a length that fits a transgene encoding one or more of the disclosed base editors.For example, the disclosed AAV vector can contain a length between 5'ITR and 3'ITR of between 4.7kb and 4.9kb. The disclosed AAV vector can contain a length between 5'ITR and 3'ITR of between 4.7kb and 5.1kb. The disclosed AAV vector can contain a length between 5'ITR and 3'ITR of between 4.6kb and 4.9kb, or between 4.6kb and 4.8kb. In certain embodiments, the length between the 5'ITR and the 3'ITR is about 4.60 kb, about 4.65 kb, about 4.70 kb, about 4.725 kb, about 4.75 kb, about 4.80 kb, about 4.825 kb, about 4.85 kb, about 4.90 kb, or about 4.95 kb.

[0067] In various embodiments, the length between the 5'ITR and the 3'ITR is less than about 4.90 kb. In some embodiments, the length between the 5'ITR and the 3'ITR is about 4.80 kb. In some embodiments, the length between the ITRs is 4.804 kb (4804 bp). In some embodiments, the length between the ITRs is 4.828 kb. In some embodiments, the length between the ITRs is 4.722 kb.

[0068] The present disclosure provides rAAV vectors that contain a size-minimized adenine base editor, and rAAV vectors that contain a size-minimized cytosine base editor. Exemplary AAV vectors of the disclosure are shown in Figures 2A and 14.

[0069] In some aspects, provided herein are rAAV vectors comprising: (i) a 5' ITR; (ii) a first nucleic acid fragment comprising a sequence encoding a SaKKH-ABE8e, SauriCas9-ABE8e, CjCas9-ABE8e base editor, or Nme2Cas9 base editor operably linked to a first promoter, where the first promoter is selected from an EFS, MeCP2, P3, and U1A promoter; and a bGH polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, where the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the rAAV vector contains a first nucleic acid fragment, comprising (i) a 5'ITR; (ii) a sequence encoding the Sauri-ABE8e base editor operably linked to an EFS promoter; and a first nucleic acid fragment comprising a polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the nucleic acid molecule; and (iv) a 3'ITR. In various embodiments, the length between the 5'ITR and the 3'ITR is less than about 4.90 kb. In some embodiments, the poly(A) signal is a bGH poly(A) signal.

[0070] In certain embodiments, the rAAV vector comprises, from 5' to 3': (i) a 5' ITR; (ii) a first nucleic acid fragment comprising a sequence encoding the SaKKH-ABE8e base editor operably linked to an EFS promoter; and a bGH polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, where the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the rAAV vector contains a first nucleic acid fragment, comprising (i) a 5'ITR; (ii) a sequence encoding a Sauri-ABE8e base editor operably linked to an EFS promoter; and a bGH polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the nucleic acid molecule; and (iv) a 3'ITR. In some embodiments, the base editor is SauriCas9-ABE8e. In some embodiments, the base editor is CjCas9-ABE8e.

[0071] In some embodiments, the rAAV vector encodes a CBE and comprises, from 5' to 3': (i) a 5' ITR; (ii) a first nucleic acid fragment comprising a sequence encoding a CjCas9-FERNY-BE3.9 or CjCas9-evoFERNY-BE3 base editor operably linked to a first promoter that is an EFS promoter; and a polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, where the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the rAAV vector encodes a CBE and includes, from 5' to 3': (i) a 5' ITR; (ii) a first nucleic acid fragment comprising a sequence encoding a CjCas9-FERNY-BE3.9 or CjCas9-evoFERNY-BE3 base editor operably linked to a first promoter, where the first promoter is selected from an EFS, MeCP2, P3, and U1A promoter; and a polyadenylation (polyA) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, where the direction of transcription of the second nucleic acid fragment is reversed relative to the direction of transcription of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the poly(A) signal is a bGH poly(A) signal. In some embodiments, the poly(A) signal is an SV40 poly(A) signal.

[0072] In various embodiments, any of the disclosed rAAV vectors are encapsidated into an AAV8 capsid. In some embodiments, the disclosed rAAV vectors are encapsidated into an AAV9 capsid.

[0073] Additionally, provided herein is a demonstration of regulatory elements on the disclosed AAV vectors for high expression levels of the encoded base editors.

[0074] definition As used in this specification and claims, the singular forms "a," "an," and "the" include both the singular and the plural unless the context clearly dictates otherwise. Thus, for example, reference to an "agent" includes a single agent as well as a plurality of such agents.

[0075] "Adeno-associated virus" or "AAV" is a virus that infects humans and several other primate species. The wild-type AAV genome is a single-stranded deoxyribonucleic acid (ssDNA) of either positive or negative sense. The genome contains two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs): rep and cap, between the ITRs. The rep ORF contains four overlapping genes that code for the Rep proteins required for the AAV life cycle. The cap ORF contains overlapping genes that code for the capsid proteins VP1, VP2, and VP3, which interact together to form the viral capsid. VP1, VP2, and VP3 are translated from one mRNA transcript, which can be spliced ​​in two different ways. Either a longer or shorter intron can be excised, resulting in the formation of two isoforms of mRNA: approximately 2.3 kb and approximately 2.6 kb long mRNA isoforms. The capsid is capable of forming a supramolecular assembly of approximately 60 individual capsid protein subunits in a non-enveloped T-1 icosahedral lattice to protect the AAV genome. The mature capsid is composed of VP1, VP2, and VP3 (with molecular masses of approximately 87, 73, and 62 kDa) in a ratio of approximately 1:1:10.

[0076] The rAAV particle may include a nucleic acid vector (e.g., a recombinant genome), which may include, at a minimum: (a) one or more heterologous nucleic acid regions (or transgenes) including sequences encoding a protein or polypeptide of interest (e.g., a base editor) or an RNA of interest (e.g., a gRNA); and (b) one or more regions including inverted terminal repeat (ITR) sequences (e.g., wild-type ITR sequences or engineered ITR sequences) flanking the one or more nucleic acid regions (e.g., heterologous nucleic acid regions). In some embodiments, the nucleic acid vector is between 4 kb and 5 kb in size (e.g., 4.2-4.7 kb in size). In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, the double-stranded nucleic acid vector may be a self-complementary vector, e.g., containing a region of the nucleic acid vector that is complementary to another region of the nucleic acid vector and initiates the formation of a double-stranded nature of the nucleic acid vector.

[0077] As used herein, the term "adenosine deaminase" or "adenosine deaminase domain" refers to a protein or enzyme that catalyzes the deamination reaction of adenosine (or adenine). The terms are used interchangeably. In certain embodiments, the present disclosure provides a base editor that includes one or more adenosine deaminase domains. For example, the adenosine deaminase domain can include a heterodimer of a first adenosine deaminase domain and a second deaminase domain connected by a linker. The adenosine deaminase provided herein (e.g., an engineered adenosine deaminase or an evolved adenosine deaminase) can be an enzyme that potentially converts adenine (A) to inosine (I) on DNA or RNA. Such an adenosine deaminase can lead to an A:T to G:C base pair conversion. In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase is not naturally occurring. For example, in some embodiments, the deaminase is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase.

[0078] In some embodiments, the adenosine deaminase is derived from bacteria such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated E. coli TadA deaminase. For example, the truncated ecTadA may lack one or more N-terminal amino acids relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the ecTadA deaminase does not include an N-terminal methionine. Reference is made to U.S. Patent Application Publication No. 2018 / 0073012, published March 15, 2018, which is incorporated herein by reference.

[0079] In genetics, the "antisense" strand of a given fragment in double-stranded DNA is the template strand, which is considered to run in the 3' to 5' direction. In contrast, the "sense" strand is the fragment in double-stranded DNA running 5' to 3', which is complementary to the antisense or template strand of DNA running 3' to 5'. In the case of a DNA fragment that codes for a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription and ultimately (typically, but not always) undergoes translation into a protein. Thus, the antisense strand carries the RNA that is later translated into a protein, while the sense strand possesses a structure nearly identical to that of the mRNA. Note that for each fragment of dsDNA, there will potentially be two sets, sense and antisense, depending on which way it is read (since sense and antisense are relative to the way it is viewed). Ultimately, it is the gene product, or mRNA, that defines which strand of a piece of dsDNA is said to be sense or antisense.

[0080] "Base editing" refers to genome editing technologies that involve the conversion of a specific nucleic acid base to another at a targeted genomic locus. In certain embodiments, this can be accomplished without requiring a double-stranded DNA break (DSB) or single-stranded break (i.e., nicking). To date, other genome editing technologies, including CRISPR-based systems, begin with the introduction of a DSB at the locus in question. Subsequently, cellular DNA repair enzymes repair the break, usually resulting in a random insertion or deletion (indel) of a base at the site of the DSB. However, these genome editing technologies are unsuitable when the introduction or correction of a point mutation at a target locus is desired rather than the stochastic disruption of the entire gene. This is because the correction rate is low (e.g., typically 0.1%-5%) and the major genome editing product is an indel. To increase the efficiency of gene correction without simultaneously introducing random indels, we previously modified the CRISPR / Cas9 system to directly convert one DNA base to another without DSB formation. See Komor, AC, et al., Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference.

[0081] As used herein, the term "base editor (BE)" refers to an agent that includes a polypeptide that can modify bases (e.g., A, T, C, G, or U) in a nucleic acid sequence (e.g., DNA or RNA), converting one base to another (e.g., A to G, A to C, A to T, C to T, C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G). In some embodiments, the base editor can deaminate a base in a nucleic acid, such as a base in a DNA molecule. In the case of an adenine base editor, the base editor can deaminate an adenine (A) on DNA. Such base editors can include a nucleic acid programmable DNA binding protein (napDNAbp) fused to adenosine deaminase. Some base editors include CRISPR-mediated fusion proteins that are utilized in the base editing methods described herein. In some embodiments, the base editor comprises a nuclease-inactive Cas9 (dCas9) fused to a deaminase, which binds to nucleic acid through the formation of an R-loop in a manner programmed by a guide RNA, but does not cleave the nucleic acid. For example, as described in PCT / US2016 / 058344, published on April 27, 2017 as International Publication No. WO 2017 / 070632, which is incorporated herein by reference in its entirety, the dCas9 domain of the fusion protein can include D10A and H840A mutations, which allow Cas9 to cleave only one strand of a nucleic acid duplex. The DNA cleavage domain of S. pyogenes Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA (the "targeted strand," or the strand where editing or deamination occurs), and the RuvC1 subdomain cleaves the non-complementary strand, which contains the PAM sequence (the "non-edited strand").The RuvC1 mutant D10A generates a nick in the targeted strand, and the HNH mutant H840A generates a nick on the non-edited strand (see Jinek et al., Science, 337:816-821 (2012); Qi et al., Cell. 28;152(5):1173-83 (2013), each of which is incorporated herein by reference).

[0082] In some embodiments, the base editor is a polymer or polymer complex that primarily (e.g., greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 99%, greater than 99.9%, or 100%) effects the conversion of a nucleobase on a polynucleic acid sequence to another nucleobase (i.e., transition or transversion) and uses a combination of 1) a nucleotide, nucleoside, or nucleobase modifying enzyme and 2) a nucleic acid binding protein that can be programmed to bind to a specific nucleic acid sequence.

[0083] In some embodiments, the base editor comprises a DNA binding domain (e.g., a programmable DNA binding domain such as dCas9 or nCas9) that guides it to a target sequence. In some embodiments, the base editor comprises a nucleobase-modifying enzyme fused to a programmable DNA binding domain (e.g., dCas9 or nCas9). A "nucleobase-modifying enzyme" is an enzyme that can modify a nucleobase and convert one nucleobase to another (e.g., a deaminase such as adenosine deaminase). Base editors that perform certain types of base conversions (e.g., adenosine (A) to guanine (G), C to G) are contemplated.

[0084] In some embodiments, the base editor converts A to G. In some embodiments, the base editor comprises an adenosine deaminase. "Adenosine deaminase" is an enzyme involved in purine metabolism. It is required for the degradation of adenosine from food and for the turnover of nucleic acids in tissues. Its primary function in humans is the development and maintenance of the immune system. Adenosine deaminase catalyzes the hydrolytic deamination of adenosine in the context of DNA (forming inosine, which base pairs as G). There are no known naturally occurring adenosine deaminases that act on DNA. Instead, known adenosine deaminase enzymes act only on RNA (tRNA or mRNA).Evolved deoxyadenosine deaminase enzymes that accept a DNA substrate and deaminate dA to deoxyinosine are described, for example, in PCT Application No. PCT / US2017 / 045381, filed August 3, 2017, published as WO 2018 / 027078, and PCT Application No. PCT / US2019 / 033848, filed May 23, 2019, published as WO 2019 / 226953 on November 28, 2019, and PCT Application No. PCT / US2019 / 033848, filed October 30, 2018, published as WO 2019 / 226953 on October 30, 2018; U.S. Pat. No. 10,113,163, filed May 23, 2019; U.S. Patent Application Publication No. 2018 / 0073012 published March 15, 2018; U.S. Patent Application Publication No. 2017 / 0121693 published May 4, 2017, which published January 1, 2019 as U.S. Patent No. 10,167,457; International Publication No. 2017 / 070633 published April 27, 2017; U.S. Patent Application Publication No. 2015 / 0166980 published June 18, 2015; U.S. Patent Application Publication No. 9,840,699 published December 12, 2017; U.S. Patent No. 9,840,699 published September 18, 2018 No. 10,077,453; International Publication No. WO 2019 / 023680 published on January 31, 2019; International Application No. PCT / US2019 / 033848 filed May 23, 2019, published on November 28, 2019 as WO 2019 / 226593; International Publication No. WO 2018 / 0176009 published on September 27, 2018, International Publication No. WO 2020 / 041751 published on February 27, 2020; International Publication No. WO 2020 / 051360 published on March 12, 2020; International Publication No. WO 2020 / 052024 published on May 22, 2020 No. 0 / 102659, published April 30, 2020; WO 2020 / 086908, published September 10, 2020; WO 2020 / 181180, published September 10, 2020; WO 2020 / 214842, published October 22, 2020; WO 2020 / 092453, published May 7, 2020; WO 2020 / 236982, published November 26, 2020; WO 2021 / 108717, published June 3, 2021, and WO 2021 / 158921, published August 12, 2021, the contents of each of which are incorporated herein by reference in their entirety.

[0085] In some embodiments, the base editor converts C to T. In some embodiments, the base editor comprises a cytidine deaminase. "Cytosine deaminase" or "cytidine deaminase" refers to the chemical reaction "cytosine + H 2 O → Uracil + NH 3 " or "5-methylcytosine + H 2 O → Thymine + NH 3" refers to an enzyme that catalyzes the reaction. As may be apparent from the reaction equation, such a chemical reaction results in a change of the nucleobase from C to U / T. In the context of a gene, such a nucleotide change or mutation may in turn lead to an amino acid change on the protein, which may affect the function of the protein. Examples include loss of function or gain of function. In some embodiments, the cytosine base editor comprises dCas9 or nCas9 fused to a cytidine deaminase. In some embodiments, the cytidine deaminase domain is fused to the N-terminus of dCas9 or nCas9. In some embodiments, the base editor further comprises a domain that inhibits uracil glycosylase and / or a nuclear localization signal.Such base editors have been described in the art, e.g., in Rees & Liu, Nat Rev Genet. 2018;19(12):770-788, Rees, et al. Sci. Advances 5, eaax5717 (2019), and Koblan et al., Nat Biotechnol. 2018;36(9):843-846; and in U.S. Patent Application Publication No. 2018 / 0073012, published March 15, 2018, which issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Application Publication No. 2017 / 01, published May 4, 2017, which issued as U.S. Patent No. 10,167,457 on January 1, 2019. No. 21693, published April 27, 2017; U.S. Patent Application Publication No. 2017 / 070633, published June 18, 2015; U.S. Patent Application Publication No. 2015 / 0166980, published December 12, 2017; U.S. Patent No. 9,840,699, published September 18, 2018; U.S. Patent No. 10,077,453, published September 18, 2018; U.S. Patent Application Publication No. 2019 / 01366980, published June 18, 2015; No. 019 / 023680; International Publication No. 2018 / 0176009 published September 27, 2018; PCT Application No. PCT / US2019 / 033848 filed May 23, 2019; PCT Application No. PCT / US2019 / 47996 filed August 23, 2019; PCT Application No. PCT / US2020 / 028568 filed April 17, 2020 No. PCT Application No. PCT / US2019 / 61685, filed November 15, 2019; PCT Application No. PCT / US2019 / 57956, filed October 24, 2019; PCT Publication No. PCT / US2019 / 58678, filed October 29, 2019; and WO 2021 / 108717, published June 3, 2021, the contents of each of which are incorporated herein by reference in their entirety.

[0086] Exemplary adenine and cytosine base editors are described in Rees & Liu, Base editing: precision chemistry on the genome and transcriptome of living cells, Nat. Rev. Genet. 2018;19(12):770-788; and U.S. Patent Application Publication No. 2018 / 0073012, published March 15, 2018, which issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Application Publication No. 2017 / 0121693, published May 4, 2017, which issued as U.S. Patent No. 10,167,457 on January 1, 2019; No. 2015 / 0166980, published June 18, 2015; U.S. Pat. No. 9,840,699, issued December 12, 2017; and U.S. Pat. No. 10,077,453, issued September 18, 2018, the contents of each of which are incorporated herein by reference in their entirety.

[0087] The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nuclease that includes a Cas9 domain or a fragment thereof (e.g., a protein that includes an active or inactive DNA cleavage domain of Cas9 and / or a gRNA binding domain of Cas9). As used herein, a "Cas9 domain" is a protein fragment that includes an active or inactive cleavage domain of Cas9 and / or a gRNA binding domain of Cas9. A "Cas9 protein" is a full-length Cas9 protein. Cas9 nuclease is sometimes referred to as casn1 nuclease or CRISPR (clustered regularly interspaced short palindromic repeats ( C Lustered R egularly I Interspaced S hort P alindromic RCRISPR is also referred to as epeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection from mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to the ancestral mobile elements, to target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA requires a small trans-encoded RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 domains. The tracrRNA acts as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cut and then exonucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require proteins and both RNAs. However, single guide RNAs ("sgRNAs", or simply "gRNAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. For an example, see Jinek M., et al. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.Cas9 nuclease sequences and structures are known to those skilled in the art (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);"CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III."Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in a variety of species, including but not limited to S. pyogenes and S. thermophilus (e.g., StCas9 or St1Cas9). Additional suitable Cas9 nucleases and sequences will be apparent to one of skill in the art based on the present disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0088] The nuclease-inactivated Cas9 domain may be interchangeably referred to as a "dCas9" protein (from nuclease "dead" Cas9). Methods for generating a Cas9 domain (or a fragment thereof) with an inactive DNA cleavage domain are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28;152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83 (2013)). In some embodiments, proteins are provided that include fragments of Cas9. For example, in some embodiments, the protein includes one of the two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, the protein that includes Cas9 or a fragment thereof is referred to as a "Cas9 variant." The Cas9 variant shares homology with Cas9 or a fragment thereof.For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO:74). In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 74). In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to a corresponding fragment of a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO:74). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO:74).

[0089] As used herein, the term "nCas9" or "Cas9 nickase" refers to Cas9 or a variant thereof that cuts or nicks only one of the strands at the target cut site, thereby introducing a nick into a double-stranded DNA molecule rather than creating a double-stranded break. This can be achieved by introducing a suitable mutation in wild-type Cas9 that inactivates one of the two endonuclease activities of Cas9. Any suitable mutation that inactivates one Cas9 endonuclease activity but leaves the other intact is contemplated. For example, one of the D10A or H840A mutations in the wild-type S. pyogenes Cas9 amino acid sequence or the D10A mutation in the wild-type S. aureus Cas9 amino acid sequence can be used to form nCas9.

[0090] The term "cDNA" refers to a strand of DNA copied from an RNA template, to which the cDNA is complementary.

[0091] CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of preceding infection by viruses that have invaded prokaryotes. The snippets of DNA are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with various CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNAs, effectively constitute a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of the pre-crRNA requires a small trans-encoded RNA (tracrRNA), an endogenous ribonuclease 3 (rnc), and a Cas9 protein. The tracrRNA acts as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves the linear or circular dsDNA target that is complementary to the RNA. Specifically, the target strand that is not complementary to the crRNA is first endonucleolytically cut, and then exonucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require proteins and both RNAs. However, single guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into the guide RNA of a single RNA species. For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (the PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.CRISPR biology and Cas9 nuclease sequences and structures are known in the art (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);"CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III."Deltcheva E., Chylinski K., Sharma CM, See Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in a variety of species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to one of skill in the art based on the present disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.

[0092] The term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is adenosine (or adenine) deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine to inosine on deoxyribonucleic acid (DNA). In other embodiments, the deaminase is cytidine (or cytosine) deaminase, which catalyzes the hydrolytic deamination of cytidine to uracil.

[0093] The deaminase described herein can be from any organism, such as bacteria.In some embodiments, the deaminase or deaminase domain is a variant of the naturally occurring deaminase from an organism.In some embodiments, the deaminase or deaminase domain is not naturally occurring.For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to naturally occurring deaminase.

[0094] As used herein, the term "DNA binding protein" or "DNA binding protein domain" refers to any protein that localizes and binds to a specific target DNA nucleotide sequence (e.g., a genomic locus). This term encompasses RNA programmable proteins that associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., this includes, e.g., guide RNAs in the case of Cas systems) that guide or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a DNA sequence) that is complementary to one or more nucleic acid molecules (or a portion or region thereof) associated with the protein. Exemplary RNA programmable proteins are CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), including Cas9 equivalents from any type of CRISPR system (e.g., types II, V, VI), Cpf1 (type V CRISPR-Cas system) (now known as Cas12a), Cas9, argonaute, and nCas9. Additional Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference.

[0095] As used herein, the term "DNA editing efficiency" refers to the number or percentage of intended base pairs that are edited. For example, if a base editor edits 10% of the base pairs that it is intended to target (e.g., in a cell or cell population), then the base editor can be described as 10% efficient. Some aspects of editing efficiency encompass the modification (e.g., deamination) of specific nucleotides in DNA without generating a large number or percentage of insertions or deletions (i.e., indels). It is generally accepted that editing while generating less than 5% indels (measured across the total target nucleotide substrate) is a high editing efficiency. The generation of more than 20% indels is generally recognized as poor or low editing efficiency. Indel formation can be measured by techniques known in the art, including high-throughput screening of sequencing reads.

[0096] As used herein, the term "off-target editing frequency" refers to the number or percentage of unintended base pairs, e.g. DNA base pairs, that are edited. On-target and off-target editing frequency can be measured by the methods and assays described herein, further considering the techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves the hybridization of a nucleic acid primer (e.g. DNA primer) that has complementarity to the nucleic acid (e.g. DNA) region immediately upstream or downstream of the target sequence or off-target sequence in question. Since the DNA target sequence and the Cas9-independent off-target sequence are known a priori in the methods disclosed herein, a nucleic acid primer that has sufficient complementarity to the upstream or downstream region of the target sequence and the Cas9-independent off-target sequence in question can be designed using techniques known in the art, such as PhusionU PCR kit (Life Technologies), Phusion HS II kit (Life Technologies), and Illumina MiSeq kit. The number of off-target DNA edits can be measured by techniques known in the art, including high-throughput screening of sequencing reads, EndoV-Seq, GUIDE-Seq, CIRCLE-Seq, and Cas-OFFinder. Many Cas9-dependent off-target sites have high sequence identity to the target site in question, so nucleic acid primers with sufficient complementarity to the upstream or downstream regions of the Cas9-dependent off-target site can be designed using techniques and kits known in the art as well. These kits utilize polymerase chain reaction (PCR) amplification to produce amplicons as intermediate products. Target and off-target sequences can include genomic loci that further include protospacers and PAMs. Thus, as used herein, the term "amplicon" can refer to a nucleic acid molecule that comprises a collection of genomic loci, protospacers, and PAMs.High throughput sequencing techniques as used herein may further include Sanger sequencing and Illumina-based next generation genomic sequencing (NGS).

[0097] As used herein, the term "on-target editing" refers to the introduction of an intended modification (e.g., deamination) to a nucleotide (e.g., adenine) on a target sequence, for example, using a base editor described herein. As used herein, the term "off-target DNA editing" refers to the introduction of an unintended modification (e.g., deamination) to a nucleotide (e.g., adenine) on a sequence outside of a classical base editor binding window (i.e., from one protospacer position to another, typically 2-8 nucleotides in length). Off-target DNA editing can result from weak or non-specific binding of the gRNA sequence to the target sequence. As used herein, the term "bystander editing" refers to a synonymous off-target point mutation at a nucleobase close (adjacent) to the target base that does not change the outcome of the intended mutation (e.g., intended disruption of a splice acceptor site, incorporation of a premature stop codon, or reversion of a mutated codon). Bystander editing can encompass non-silent mutations at the relevant codon of a transcript that do not result in a different translated protein.

[0098] The terms "purity" and "product purity" of a base editor as used herein refer to the average percentage of edited sequencing reads (reads in which the target nucleobase is converted to a different base) in which the intended target conversion occurs (e.g., target A, and only target A is converted to G). See Komor et al., Sci Adv 3 (2017).

[0099] As used herein, the terms "upstream" and "downstream" are relative terms that define the linear location of at least two elements located on a nucleic acid molecule (whether single-stranded or double-stranded) in a 5' to 3' direction. In particular, where a first element is located somewhere 5' of a second element, the first element is upstream of the second element on the nucleic acid molecule. For example, if the SNP is 5' of the nick site, the SNP is upstream of the nick site induced by Cas9. Conversely, where a first element is located somewhere 3' of a second element, the first element is downstream of the second element on the nucleic acid molecule. For example, if the SNP is 3' of the nick site, the SNP is downstream of the nick site induced by Cas9. The nucleic acid molecule can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. The analysis is the same for single-stranded nucleic acid molecules and double-stranded molecules. The terms upstream and downstream refer only to a single strand of a nucleic acid molecule, except that it is necessary to select which strand of a double-stranded molecule is being considered. In many cases, the strand of a double-stranded DNA that can be used to determine the relative position of at least two elements is the "sense" or "coding" strand. In genetics, a "sense" strand is a segment within a double-stranded DNA that extends from 5' to 3' and is complementary to the antisense or template strand of DNA that extends from 3' to 5'. Thus, by way of example, if a SNP nucleobase is 3' to the promoter on the sense or coding strand, the SNP nucleobase is "downstream" of the promoter sequence on genomic DNA (which is double-stranded).

[0100] As used herein, the term "effective amount" refers to an amount of a biologically active agent that is sufficient to induce a desired biological response. For example, in some embodiments, an effective amount of a base editor can refer to an amount of the editor that is sufficient to edit a target site nucleotide sequence, e.g., a genome. In some embodiments, an effective amount of a base editor described herein, e.g., a base editor comprising a nickase Cas9 domain and a guide RNA, can refer to an amount of the base editor that is sufficient to induce editing of a target site that is specifically bound and edited by the base editor. As will be understood by those skilled in the art, the effective amount of an agent, e.g., a base editor, a nuclease, a hybrid protein, a protein dimer, a complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide, can vary and depend on various factors, such as, for example, a desired biological response, e.g., a particular allele, genome, or target site to be edited, a cell or tissue to be targeted, and an agent to be used.

[0101] The term "functional equivalent" refers to a second biomolecule that is functionally equivalent to a first biomolecule, but not necessarily structurally equivalent. For example, a "Cas9 equivalent" refers to a protein that has the same or substantially the same function as Cas9, but not necessarily the same amino acid sequence. In the context of this disclosure, the specification refers throughout to "protein X or a functional equivalent thereof." In this context, a "functional equivalent" of protein X encompasses any homolog, paralog, fragment, naturally occurring, engineered, circularly permuted, mutated, or synthetic version of protein X that has an equivalent function.

[0102] As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) portion of the protein, thus forming a hybrid "amino-terminal fusion protein" or a "carboxy-terminal fusion protein". The protein may contain different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9, which directs the binding of the protein to a target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. Another example includes Cas9 or its equivalent fused to adenosine deaminase. Any of the proteins described herein may be produced by any method known in the art. For example, the proteins described herein may be produced via recombinant protein expression and purification. This is particularly suitable for fusion proteins that contain peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012), the entire contents of which are incorporated herein by reference.

[0103] The term "guide nucleic acid" or "nucleic acid molecule that programs napDNAbp" or equivalently "guide sequence" refers to one or more nucleic acid molecules that bind to napDNAbp protein and guide or program it to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to one or more nucleic acid molecules (or a portion or region thereof) bound to the protein, thereby causing the napDNAbp protein to bind to the nucleotide sequence at the specific target site. A non-limiting example is the guide RNA of the Cas protein of the CRISPR-Cas genome editing system. Chemically, the guide nucleic acid can be all RNA, all DNA, or a chimera of RNA and DNA. The guide nucleic acid can also include nucleotide analogs. The guide nucleic acid can be expressed as a transcription product or can be synthesized.

[0104] As used herein, "guide RNA" or "gRNA" refers to a synthetic fusion of endogenous bacterial crRNA and tracrRNA, which provides both targeting specificity and backbone and / or binding ability of Cas9 nuclease to target DNA. This synthetic fusion is not naturally occurring and is also commonly referred to as sgRNA. However, the term guide RNA also encompasses equivalent guide nucleic acid molecules, whether naturally occurring or not (e.g. engineered or recombinant), that associate with Cas9 equivalents, homologs, orthologs, or paralogs and otherwise program the Cas9 equivalent to localize to a specific target nucleotide sequence. Cas9 equivalents can include other napDNAbp from any type of CRISPR system (e.g. types II, V, VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein. Additionally, methods for designing suitable guide RNA sequences are provided herein.

[0105] A guide RNA is a specific type of guide nucleic acid that is most commonly associated with the Cas protein of CRISPR-Cas9 and associates with Cas9 to guide the Cas9 protein to a specific sequence on a DNA molecule that contains a sequence complementary to the protospacer sequence of the guide RNA. Functionally, a guide RNA associates with Cas9 and guides (or programs) the Cas9 protein to a specific sequence on a DNA molecule that contains a sequence complementary to the protospacer sequence of the guide RNA.

[0106] As used herein, a "spacer sequence" is a sequence (approximately 20 nt in length) of a guide RNA that has the same sequence (with the exception of uridine bases instead of thymine bases) as the protospacer of the PAM strand of a target (DNA) sequence and is complementary to the target strand (or non-PAM strand) of the target sequence.

[0107] As used herein, "target sequence" refers to approximately 20 nucleotides on the target DNA sequence that have complementarity to the protospacer sequence on the PAM strand. The target sequence is the sequence that anneals to or is targeted by the spacer sequence of the guide RNA. The spacer sequence of the guide RNA and the protospacer have the same sequence (with the exception that the spacer sequence is RNA and the protospacer is DNA).

[0108] As used herein, the terms "guide RNA core," "guide RNA scaffold sequence," and "backbone sequence" refer to the sequence within the gRNA that is responsible for Cas9 binding, which does not include the 20 bp spacer sequence used to guide Cas9 to the target DNA.

[0109] As used herein, the term "host cell" refers to a cell that can host and replicate a vector encoding a base editor, a guide RNA, and / or a combination thereof, as described herein. In some embodiments, the host cell is a mammalian cell, such as a human cell. Provided herein are methods of transducing and transfecting a host cell, such as a human cell, e.g., a human cell in a subject, with one or more vectors provided herein, e.g., one or more viral (e.g., rAAV) vectors provided herein.

[0110] It should be understood that any of the base editors, guide RNAs, and or combinations thereof described herein can be introduced into a host cell in any suitable manner, either stably or transiently. In some embodiments, the base editor can be transfected into the host cell. In some embodiments, the host cell can be transduced or transfected with a nucleic acid construct encoding the base editor. For example, the host cell can be transduced with a nucleic acid encoding the base editor or a translated base editor (e.g., by a viral particle encoding the base editor). As an additional example, the host cell can be transfected with a nucleic acid encoding the base editor (e.g., a plasmid) or a translated base editor. Such transduction or transfection can be stable or transient. In some embodiments, the host cell expressing or containing the base editor can be transduced or transfected with one or more gRNA molecules, for example when the base editor comprises a Cas9 (e.g., nCas9) domain. In some embodiments, a plasmid expressing a base editor can be introduced into a host cell through electroporation, transient transfection (e.g., lipofection, such as with Lipofectamine 3000®), stable genomic integration (e.g., piggybac), viral transduction, or other methods known to one of skill in the art.

[0111] Host cells for packaging viral particles are also provided herein.In an embodiment where the vector is a viral vector, a suitable host cell is one that can be infected by the viral vector, replicate it, and package it into viral particles that can infect new host cells.A cell can be a host for a viral vector if it supports the expression of the gene of the viral vector, the replication of the viral genome, and / or the production of viral particles.In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, an insect cell, or a mammalian cell.The type of host cell will of course depend on the vector employed, and suitable host cell / vector combinations will be readily apparent to those skilled in the art.

[0112] An "intein" is a fragment of a protein that can excise itself and splice the remaining part (the extein) together by a peptide bond in a process known as protein splicing. Inteins are also referred to as "protein introns." The process by which an intein excises itself and splices together the remaining part of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the inteins of a precursor protein (the intein-containing protein that precedes intein-mediated protein splicing) come from two genes. Such inteins are referred to herein as split inteins. For example, in cyanobacteria, DnaE, ​​the catalytic subunit alpha of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene is referred to herein as "intein-N." The intein encoded by the dnaE-c gene is referred to herein as "intein-C." In various embodiments of the disclosed nucleic acid molecules, the nucleic acid molecule does not contain an intein. In various embodiments of the disclosed nucleic acid molecules, the nucleic acid molecule does not contain a trans-splicing intein.

[0113] Other intein systems can also be used. For example, synthetic inteins based on dnaE intein, Cfa-N and Cfa-C intein pairs have been described (e.g., Stevens et al., J Am Chem Soc. 2016 Feb 24;138(7):2162-5, incorporated herein by reference). As another example, Nostoc punctiforme (Npu) intein pair, synthetic intein based on dnaE intein, has been described (see Zettler, J., Schutz, V. & Mootz, HD, The naturally split Npu DnaE intein exhibits an extraordinarily high rate in the protein trans-splicing reaction. FEBS letters 583, 909-914 (2009), incorporated herein by reference). In some embodiments, the intein is a fast-splicing gp41 intein, such as gp41-8 intein. Reference is made to Carvajal-Vallejos et al, Journal of Biological Chemistry 287(34): 28686-28696 (2012) and Pinto, Thornton & Wang, Nat.Comm. (2020) 11:1529, each of which is incorporated herein by reference. Non-limiting examples of intein pairs that can be used according to the present disclosure include: Cfa DnaE intein, Npu DnaE intein, gp41-8 intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Pat. No. 8,394,604, which is incorporated herein by reference).

[0114] As used herein, the term "linker" refers to a chemical group or molecule that connects two molecules or domains, for example dCas9 and deaminase. Typically, a linker is located between or flanked by two groups, molecules, or other domains, and is connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (for example, a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical domain. Chemical groups include, but are not limited to, disulfide, hydrazone, and azide domains. In some embodiments, the linker is 5-100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the linker is an XTEN linker that is 32 amino acids in length. In some embodiments, the linker is a 32 amino acid linker. In other embodiments, the linker is a 30, 31, 33, or 34 amino acid linker.

[0115] As used herein, the term "mutation" refers to the replacement of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue; the deletion or insertion of one or more residues in a sequence; or the replacement of a residue in a genomic sequence in a subject to be modified. Mutations are typically described herein by identifying the original residue, followed by the position of the residue in the sequence and the identity of the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and can be found, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012). Mutations can encompass a variety of categories, such as single nucleotide polymorphisms, microduplication regions, indels, and inversions, and are in no way meant to be limiting. Mutations can encompass "loss-of-function" mutations, which are mutations that reduce or eliminate protein activity. Most loss-of-function mutations are recessive, since in heterozygotes, the second chromosomal copy carries an unmutated version of the gene that codes for a fully functional protein, and its presence compensates for the effect of the mutation. There are some exceptions where loss-of-function mutations are dominant. One example is haploinsufficiency, where the organism cannot tolerate the roughly 50% reduction in protein activity suffered by heterozygotes. This accounts for a small number of genetic diseases in humans, including Marfan syndrome, which results from a mutation in the gene for a connective tissue protein called fibrillin. Mutations also encompass "gain-of-function" mutations, which confer abnormal activity to a protein or cell that would not otherwise be present under normal conditions. Many gain-of-function mutations are in regulatory sequences rather than coding regions and therefore can have a number of consequences. By their nature, gain-of-function mutations are usually dominant. Many loss-of-function mutations are recessive, autosomal recessive, etc. Many of the USH2A mutations that the base editing method of the present disclosure aims to correct are autosomal recessive.

[0116] The term "napDNAbp", meaning a "nucleic acid programmable DNA binding protein", refers to any protein that can bind to (e.g., form a complex with) one or more nucleic acid molecules (i.e., which may be broadly referred to as "nucleic acid molecules that program the napDNAbp", which in the case of a Cas system, includes, for example, guide RNAs) that direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to one or more nucleic acid molecules (or a portion or region thereof) bound to the protein, thereby causing the protein to bind to the nucleotide sequence at the specific target site. The term napDNAbp encompasses CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and includes Cas9 equivalents from any type of CRISPR system (e.g., types II, V, VI), and may include Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14, Argonaute, and nCas9. Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353 (6299), the contents of which are incorporated herein by reference. However, the nucleic acid programmable DNA binding proteins (napDNAbp) that may be used in connection with the present invention are not limited to CRISPR-Cas systems.The present invention encompasses any such programmable protein, such as the Argonaute protein (NgAgo) from Natronobacterium gregoryi, which can also be used for DNA-guided genome editing. The NgAgo-guided DNA system does not require a PAM sequence or a guide RNA molecule. This means that genome editing can be performed on any genome sequence simply by expressing a generic NgAgo protein and introducing synthetic oligonucleotides. See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7):768-73, incorporated herein by reference.

[0117] In some embodiments, napDNAbp is an RNA programmable nuclease, and when in a complex with RNA, it can be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) is referred to as a guide RNA (gRNA). gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (and, for example, directs binding of the Cas9 (or equivalent) complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is homologous to tracrRNA as shown in Figure 1E of Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (including, for example, domain 2) can be found in U.S. Patent No. 9,340,799, entitled "mRNA-Sensing Switchable gRNAs," and PCT Application No. PCT / US2014 / 054247, entitled "Delivery System For Functional Nucleases," published as International Publication WO 2015 / 035136, filed September 6, 2013, the entire contents of each of which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more domains (1) and (2), and can be referred to as an "extended gRNA." For example, as described herein, an extended gRNA will bind, for example, two or more Cas9 proteins and bind to a target nucleic acid at two or more distinct regions. The gRNA contains a nucleotide sequence complementary to a target site, which mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex.In some embodiments, the RNA programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 (Csn1) from Streptococcus pyogenes (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti JJ et al., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M. et al., Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference.

[0118] The napDNAbp nucleases (Cas9 for example) use RNA:DNA hybridization to target DNA cleavage sites, and these proteins can in principle be targeted to any sequence specified by the guide RNA. Methods of using napDNAbp nucleases such as Cas9 for site-specific cleavage (e.g., to modify genomes) are known in the art (see, e.g., Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature Biotechnology 31, 227-229 (2013); Jinek, M. et al. RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013); Jiang, W. et al. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013), the entire contents of each of which are incorporated herein by reference.

[0119] The term "nickase" refers to a napDNAbp (e.g., Cas9) that has only a single nuclease activity that cuts only one strand of the target DNA rather than both strands. Thus, a nickase-type napDNAbp does not leave a double-stranded break. Exemplary nickases include SpCas9 and SaCas9 nickases. Exemplary nickases include sequences that have at least 99% or 100% identity to the amino acid sequence of SEQ ID NO: 107.

[0120] A nuclear localization signal or sequence (NLS) is an amino acid sequence that tags, selects, or otherwise marks a protein for import into the cell nucleus by nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface. Different nuclear localization proteins may share the same NLS. NLSs have the opposite function to nuclear export signals (NESs), which target proteins out of the nucleus. Hence, a single nuclear localization signal may direct the entity to which it is attached into the nucleus of the cell. Such sequences may be of any size and composition, e.g., more than 25, 25, 15, 12, 10, 8, 7, 6, 5, or 4 amino acids, but preferably will contain at least 4-8 amino acid sequences known to function as nuclear localization signals (NLSs).

[0121] As used herein, the term "nucleic acid molecule" refers to RNA and single-stranded and / or double-stranded DNA.Nucleic acid molecule can be naturally occurring, for example, in the context of genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule.On the other hand, nucleic acid molecule can be a non-naturally occurring molecule, for example, recombinant DNA or RNA, artificial chromosome, engineered genome or its fragment, or synthetic DNA, RNA, DNA / RNA hybrid, or can include non-naturally occurring nucleotide or nucleoside.

[0122] Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms encompass nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, optionally purified, chemically synthesized, and the like. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as analogs having chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise indicated. In some embodiments, the nucleic acid is selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C ...bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C -aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, inosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases, e.g., 2'-O-methylated bases); intercalation bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0123] As used herein, the term "phage-assisted continuous evolution (PACE)" refers to continuous evolution employing phages as viral vectors. The general concept of PACE technology can be seen, for example, in PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, published as WO 2010 / 028347 on March 11, 2010; PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; PCT Application No. PCT / US2011 / 066747, filed May 5, 2015, published as WO 2012 / 088381 on May 5, 2015; No. 9,023,594, filed Jan. 20, 2015, published as WO 2015 / 134121 on Sep. 11, 2015, and PCT Application No. PCT / US2015 / 012022, filed Apr. 15, 2016, published as WO 2016 / 168631 on Oct. 20, 2016, the entire contents of each of which are incorporated herein by reference.

[0124] The term "promoter" is recognized in the art and refers to a nucleic acid molecule having a sequence that can be recognized by the transcriptional machinery of a cell and initiate transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is only active in the presence of a particular condition. For example, a conditional promoter can be active only in the presence of a specific protein that connects a protein associated with a regulatory element on the promoter to the basal transcriptional machinery or in the absence of an inhibitory molecule. A subclass of conditionally active promoters are inducible promoters, which require the presence of a small molecule "inducer" for activity. Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. A variety of constitutive, conditional, and inducible promoters are known to those of skill in the art, and the skilled artisan will be able to ascertain a variety of such promoters useful in carrying out the present invention. This is not limiting in this respect. In various embodiments, the disclosure provides vectors with a suitable promoter to drive expression of a nucleic acid sequence encoding a base editor (or one or more individual components thereof).

[0125] As used herein, the term "protospacer" refers to a sequence (eg, a sequence of approximately 20 bp) on DNA adjacent to a PAM (protospacer adjacent motif) sequence, which shares the same sequence as the spacer sequence of the guide RNA, which is complementary to the target sequence of the non-PAM strand. The spacer sequence of the guide RNA anneals to the target sequence located on the non-PAM strand. For Cas9 to function, it also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease from S. pyogenes recognizes a PAM sequence of NGG, which is found directly downstream of the protospacer sequence on genomic DNA on the non-target strand. Those skilled in the art will understand that the prior art literature in this field sometimes refers to "protospacer" as the approximately 20 nt target-specific guide sequence on the guide RNA itself, rather than calling it a "spacer" (and that the protospacer (DNA) and the spacer (RNA) have the same sequence). Therefore, as used herein, the term "protospacer" may be used interchangeably with the term "spacer". The written context surrounding the appearance of either "protospacer" or "spacer" will help inform the reader as to whether the term is a reference to a gRNA or DNA sequence. Both uses of these terms are permissible, as the current state of the art uses both terms in each of these ways.

[0126] As used herein, the term "protospacer adjacent sequence" or "PAM" refers to a DNA sequence of approximately 2-6 base pairs that is a key targeting component of the Cas9 nuclease. Typically, the PAM sequence is on either strand and is downstream in the 5' to 3' direction of the Cas9 cut site. The classical PAM sequence (i.e., the PAM sequence associated with the Cas9 nuclease of Streptococcus pyogenes or SpCas9) is 5'-NGG-3', where "N" is any nucleobase followed by two guanine ("G") nucleobases. Different PAM sequences may be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., SpCas9, may be modified to modulate the PAM specificity of the nuclease by allowing the nuclease to recognize alternative PAM sequences.

[0127] For example, referring to the classical SpCas9 amino acid sequence being SEQ ID NO: 74, the PAM sequence can be modified by introducing one or more mutations, including: (a) D1135V, R1335Q, and T1337R "VRQR variants" that modulate PAM specificity to NGAN or NGNG, (b) D1135E, R1335Q, and T1337R "EQR variants" that modulate PAM specificity to NGAG, and (c) D1135V, G1218R, R1335E, and T1337R "VRER variants" that modulate PAM specificity to NGCG. In addition, the D1135E variant of classical SpCas9 still recognizes NGG, but it is more selective compared to the wild-type SpCas9 protein.

[0128] It will also be appreciated that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologues) may have different PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. In addition, Cas9 from Neisseria meningitis (NmeCas) recognizes NNNNGATT. Cas9 from Staphylococcus auricularis (SauriCas9) recognizes NNGG and NNNGG. Cas9 from Streptococcus thermophilis (StCas9) recognizes NNAGAAW. Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These are examples and are not meant to be limiting. It will further be appreciated that non-SpCas9 bind to a variety of PAM sequences, making them useful when a suitable SpCas9 PAM sequence is not present at the desired target cut site. Furthermore, non-SpCas9 may have other properties that make them more useful than SpCas9. For example, Cas9 from Staphylococcus aureus (SaCas9) is approximately 1 kilobase smaller than SpCas9. Therefore, it can be packaged into an adeno-associated virus (AAV). Further reference may be made to Shah et al., "Protospacer recognition motifs: mixed identities and functional diversity," RNA Biology, 10(5): 891-899, which is incorporated herein by reference.

[0129] The terms "protein", "peptide" and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids in length. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids on a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity, such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide may also be a single molecule or may be a multi-molecular complex. A protein, peptide, or polypeptide may be simply a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. It is to be understood that the present disclosure provides any of the polypeptide sequences provided herein without the N-terminal methionine (M) residue.

[0130] In genetics, a "sense" strand is a segment in a double-stranded DNA that runs 5' to 3' and is complementary to the antisense or template strand of DNA that runs 3' to 5'. In the case of a DNA segment that codes for a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription and ultimately (typically, but not always) undergoes translation into a protein. Thus, the antisense strand carries the RNA that is later translated into a protein, while the sense strand possesses a structure nearly identical to that of the mRNA. Note that for each segment of dsDNA, there will potentially be two sets, sense and antisense, depending on which way it is read (since sense and antisense are relative to the way it is viewed). Ultimately, it is the gene product or mRNA that defines which strand of a segment of dsDNA is said to be sense or antisense.

[0131] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein provided as an N-terminal portion (interchangeably referred to herein as the N-terminal half) and a C-terminal portion (interchangeably referred to herein as the C-terminal half) encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be combined (joined together) to form a complete Cas9 protein. The Cas9 protein is known to consist of a bilobal structure linked by a disordered linker (e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, which is incorporated herein by reference). In some embodiments, the "split" occurs between the two lobes, generating two portions of the Cas9 protein, each containing one lobe.

[0132] As used herein, the term "subject" refers to an individual organism, for example, an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, for example a genetically engineered non-human subject. The subject can be of either sex and at any developmental stage. In some embodiments, the subject is a livestock animal. In some embodiments, the subject is a plant.

[0133] The term "target site" refers to the sequence in a nucleic acid molecule that is edited by the base editor (BE) disclosed herein. The term "target site" in the context of a single strand can also refer to the "target strand" that anneals or binds to the spacer sequence of the guide RNA. In certain embodiments, the target site can refer to a double-stranded DNA fragment that includes the PAM strand (or non-target strand) and the protospacer on the target strand (i.e., the strand of the target site that has the same nucleotide sequence as the spacer sequence of the guide RNA), which is similarly complementary to the protospacer and the spacer, and anneals to the spacer of the guide RNA, thereby targeting or programming the Cas9 base editor to target the target site.

[0134] A "transcription terminator" is a nucleic acid sequence that causes transcription to stop. A transcription terminator can be unidirectional or bidirectional. It is composed of a DNA sequence that is involved in the specific termination of an RNA transcript by an RNA polymerase. A transcription terminator sequence prevents transcriptional activation of a downstream nucleic acid sequence by an upstream promoter. A transcription terminator may be necessary in vivo to achieve a desired expression level or to avoid transcription of a particular sequence. A transcription terminator is considered to be "operably linked to" a nucleotide sequence when it is capable of terminating the transcription of the sequence to which it is linked.

[0135] In eukaryotic systems, the terminator region may contain a specific DNA sequence that exposes a polyadenylation site to allow site-specific cleavage of the new transcript. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3' end of the transcript. RNA molecules modified with this polyA tail (signal) appear to be more stable and are translated more efficiently. Therefore, in some embodiments involving eukaryotes, the terminator may contain a signal for cleavage of the RNA. In some embodiments, the terminator signal promotes polyadenylation of the message. The terminator and / or polyadenylation site elements may act to enhance output nucleic acid levels and / or to minimize read-through between nucleic acids.

[0136] In some embodiments, the transcription terminator contains a post-transcriptional response element, a sequence that creates a tertiary structure that enhances expression when transcribed. In some embodiments, the post-transcriptional response element is derived from Woodchuck Hepatitis Virus (WHV), i.e., WPRE. In some embodiments, the terminator contains the gamma subunit or W3 of WPRE, which was first reported in Choi, JH, et al. (2014), Mol. Brain 7: 17, which is incorporated herein by reference. WPRE also has alpha and beta subunits. Typically, the post-transcriptional response element is inserted 5' of the transcription terminator. In certain embodiments, the WPRE is a truncated WPRE sequence. In certain embodiments, the WPRE is a full-length WPRE.

[0137] Terminators for use according to the present disclosure include any transcription terminator described herein or known to those skilled in the art.Examples of terminators include, without limitation, gene termination sequences, such as bovine growth hormone terminator, and viral termination sequences, such as SV40 terminator and bGH terminator.In some embodiments, termination signals can be sequences that cannot be transcribed or translated, such as those resulting from sequence truncation.

[0138] As used herein, "transition" refers to the mutual exchange of purine nucleobases (A⇔G) or pyrimidine nucleobases (C⇔T). This class of mutual exchanges involves nucleobases of similar shape. The compositions and methods disclosed herein can induce one or more transitions on a target DNA molecule. The compositions and methods disclosed herein can also induce both transitions and transversions in the same target DNA molecule. These changes involve A⇔G, G⇔A, C⇔T, or T⇔C. In the context of double-stranded DNA with Watson-Crick paired nucleobases, transitions refer to the following base pair exchanges: A:T⇔G:C, G:G⇔A:T, C:G⇔T:A, or T:A⇔C:G. The compositions and methods disclosed herein can induce one or more transitions on a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing other nucleotide changes, including both transitions and transversions, as well as deletions and insertions, in the same target DNA molecule.

[0139] As used herein, "transversion" refers to the interchange of purine nucleobases with pyrimidine nucleobases or vice versa, and therefore involves the interchange of nucleobases with dissimilar shapes. These changes involve T⇔A, T⇔G, C⇔G, C⇔A, A⇔T, A⇔C, G⇔C, and G⇔T. In the context of double-stranded DNA with Watson-Crick paired nucleobases, transversion refers to the following base pair exchanges: T:A⇔A:T, T:A⇔G:C, C:G⇔G:C, C:G⇔A:T, A:T⇔T:A, A:T⇔C:G, G:C⇔C:G, and G:C⇔T:A.

[0140] The terms "treatment", "treat" and "treating" refer to clinical interventions aimed at reversing, alleviating, delaying the onset or inhibiting the progression of a disease or disorder or one or more symptoms thereof, as described herein. As used herein, the terms "treatment", "treat" and "treating" refer to clinical interventions aimed at reversing, alleviating, delaying the onset or inhibiting the progression of a disease or disorder or one or more symptoms thereof, as described herein. In some embodiments, treatment can be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment can be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or to inhibit the onset or progression of a disease. For example, treatment can be administered to a predisposed individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other predisposing factors). Treatment can also be continued after symptoms have resolved, e.g., to prevent or delay their recurrence.

[0141] As used herein, the terms "upstream" and "downstream" are relative terms that define the linear location of at least two elements located on a nucleic acid molecule (whether single-stranded or double-stranded) in a 5' to 3' direction. In particular, where a first element is located somewhere 5' of a second element, the first element is upstream of the second element on the nucleic acid molecule. For example, if the SNP is 5' of the nick site, the SNP is upstream of the nick site induced by Cas9. Conversely, where a first element is located somewhere 3' of a second element, the first element is downstream of the second element on the nucleic acid molecule. For example, if the SNP is 3' of the nick site, the SNP is downstream of the nick site induced by Cas9. The nucleic acid molecule can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. The analysis is the same for single-stranded nucleic acid molecules and double-stranded molecules. The terms upstream and downstream refer only to a single strand of a nucleic acid molecule, except that it is necessary to select which strand of a double-stranded molecule is being considered. In many cases, the strand of a double-stranded DNA that can be used to determine the relative position of at least two elements is the "sense" or "coding" strand. In genetics, a "sense" strand is a segment within a double-stranded DNA that extends from 5' to 3' and is complementary to the antisense or template strand of DNA that extends from 3' to 5'. Thus, by way of example, if a SNP nucleobase is 3' to the promoter on the sense or coding strand, the SNP nucleobase is "downstream" of the promoter sequence on the genomic DNA (which is double-stranded).

[0142] As used herein, the term "variant" refers to a protein that has a characteristic that deviates from that of naturally occurring proteins that retain at least one of its functional, i.e., binding, interaction, or enzymatic capabilities and / or therapeutic properties. A "variant" is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to a wild-type protein. For example, a variant of Cas9 can include a Cas9 that has one or more changes in amino acid residues compared to the wild-type Cas9 amino acid sequence. As another example, a variant of a deaminase can include a deaminase that has one or more changes in amino acid residues compared to the wild-type deaminase amino acid sequence, for example after reconstruction of the deaminase's ancestral sequence. These changes include chemical modifications, including substitution of different amino acid residues, truncations, covalent additions (e.g., of tags), and any other mutations. The term also encompasses circular permutations, mutants, truncations, or domains of the reference sequence that exhibit the same or substantially the same functional activity(ies) as the reference sequence. The term also encompasses fragments of the wild-type protein.

[0143] The level or degree to which the properties are retained may be reduced relative to the wild-type protein, but typically is the same or similar. In general, the variants are generally very similar and in many regions identical to the amino acid sequences of the proteins described herein. Those skilled in the art will understand how to make and use variants that retain all or at least some of the functional capabilities or properties. A variant protein may, for example, comprise or alternatively consist of an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of the wild-type protein or any protein provided herein.

[0144] By a polypeptide having an amino acid sequence that is at least, for example, 95% "identical" to a query amino acid sequence, it is intended that the amino acid sequence of the subject polypeptide is identical to the query sequence, except that the subject polypeptide sequence may include up to 5 amino acid modifications per 100 amino acids of the query amino acid sequence. In other words, to obtain a polypeptide having an amino acid sequence that is at least 95% identical to the query amino acid sequence, up to 5% of the amino acid residues on the subject sequence may be inserted, deleted, or replaced by another amino acid. These modifications of the reference sequence may occur at the amino or carboxy terminal positions of the reference amino acid sequence or anywhere between these terminal positions, either individually between the residues on the reference sequence or in one or more consecutive groups within the reference sequence.

[0145] In practice, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to the amino acid sequence of, for example, a fusion protein can be conventionally determined using known computer programs.A preferred method for determining the best overall match between a query sequence (sequence of the present invention) and a subject sequence, also referred to as global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)).In sequence alignment, the query sequence and the subject sequence are either both nucleotide sequences or both amino acid sequences.The result of the global sequence alignment is expressed as percent identity. Preferred parameters used in FASTDB amino acid alignments are Matrix=PAM 0, k-tuple=2, Mismatch Penalty=1, Joining Penalty=20, Randomization Group Length=0, Cutoff Score=1, Window Size=sequence length, Gap Penalty=5, Gap Size Penalty=0.05, Window Size=500 or the length of the subject amino acid sequence, whichever is shorter.

[0146] If the subject sequence is shorter than the query sequence due to N- or C-terminal deletions, but not because of internal deletions, manual correction must be made to the results. This is because the FASTDB program does not take into account the N- and C-terminal truncations of the subject sequence when calculating the global identity percentage. For subject sequences that are truncated at the N- and C-terminus relative to the query sequence, the percent identity is corrected by calculating the number of residues of the query sequence that are N- and C-terminal of the subject sequence that do not match / align with the corresponding subject residues as a percentage of the total bases of the query sequence. Whether a residue matches / aligns is determined by the results of the FASTDB sequence alignment. This percentage is then subtracted from the percent identity calculated by the FASTDB program above using the specified parameters to arrive at a final percent identity score. This final percent identity score is what is used for the purposes of the present invention. Only the residues at the N- and C-terminus of the subject sequence that do not match / align with the query sequence are considered for the purpose of manually adjusting the identity percentage score. That is, only query residue positions outside the farthest N- and C-terminal residues of the subject sequence.

[0147] As used herein, the term "vector" refers to a nucleic acid that can be modified to code for a gene of interest and can enter a host cell, replicate within the host cell, and then transfer the replicated form of the vector to another host cell. Exemplary suitable vectors include viral vectors, such as AAV vectors, or bacteriophages and filamentous phages, as well as conjugative plasmids. Additional suitable vectors will be apparent to those skilled in the art based on this disclosure.

[0148] As used herein, the term "wild-type" is a term of art understood by those of skill in the art and means the typical form of an organism, strain, gene, or trait as it occurs in nature, as distinguished from mutant or variant forms.

[0149] napDNAbp domain The base editors described herein include a nucleic acid programmable DNA binding (napDNAbp) domain. Each napDNAbp binds to at least one guide nucleic acid (e.g., a guide RNA), which localizes the napDNAbp to a DNA sequence that includes a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid or a portion thereof (e.g., a protospacer of the guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp domain to localize and bind to a complementary sequence of the target strand. The binding of the napDNAbp domain to the complementary sequence allows the nucleobase-modifying domain (i.e., an adenosine deaminase domain) of the base editor to access and enzymatically deaminate the target base on the target strand.

[0150] napDNAbp may be a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. As outlined above, CRISPR is an adaptive immune system that provides protection from mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain sequences complementary to ancestral mobile elements, spacers, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA requires a small trans-encoded RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein. tracrRNA acts as a guide for ribonuclease 3-assisted processing of pre-crRNA. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets that are complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cut and then exonucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically requires a protein and both RNAs. However, single guide RNAs ("sgRNAs", or simply "gRNAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference.

[0151] The description below of various napDNAbps that can be used with respect to the disclosed adenosine deaminase is not meant to be limiting in any way. The base editor can include any variant Cas9 protein, including classical SpCas9, or any orthologous Cas9 protein, or any naturally occurring variant, mutant, or otherwise engineered version of Cas9. This is known, or it can be created or evolved through a directed evolutionary or otherwise mutagenic process. In various embodiments, the napDNAbp has nickase activity, i.e., it cuts only one strand of the target DNA sequence. In other embodiments, the napDNAbp has an inactive nuclease, e.g., a "dead" protein. Other variant Cas9 proteins that can be used are those that have a smaller molecular weight than classical SpCas9 (e.g., for easier delivery) or have a modified or rearranged primary amino acid sequence (e.g., a circularly permuted form). The base editors described herein can also include Cas9 equivalents, including Cas12a / Cpf1 proteins. The napDNAbp (e.g., SpCas9, SaCas9, or SaCas9 variant or SpCas9 variant) used herein may also contain various modifications that modulate / enhance their PAM specificity. The present disclosure contemplates any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to any of the Cas9 proteins disclosed herein. In some embodiments, the napDNAbp domain comprises a nickase variant of wild-type Cas9. In some embodiments, the napDNAbp domain comprises any of the Cas9 nickases disclosed herein.

[0152] In some embodiments, the napDNAbp directs cleavage of one or both strands, for example at the location of the target sequence within the target sequence and / or within the complement of the target sequence. In some embodiments, the napDNAbp directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. For example, an aspartic acid to alanine substitution (D10A) on the RuvC I catalytic domain of Cas9 from S. pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that make Cas9 a nickase include, without limitation, H840A, N854A, and N863A, with reference to the equivalent amino acid positions on the classical SpCas9 sequence, or other Cas9 variants or Cas9 equivalents.

[0153] As used herein, the term "Cas protein" refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence that differs from a naturally occurring Cas protein, or any fragment of a Cas protein that still retains all or a significant amount of the essential basic functions required for the disclosed methods, namely, (i) possession of nucleic acid programmable binding of the Cas protein to target DNA and (ii) the ability to nick a target DNA sequence on one strand. The Cas proteins contemplated herein encompass CRISPR Cas9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9)) homologs, orthologs, or paralogs, whether naturally occurring or not (e.g., engineered or recombinant), and include Cas9 equivalents from any type of CRISPR system (e.g., types II, V, VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Additional Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299). No. 6,399,433, the contents of which are incorporated herein by reference.

[0154] The term "Cas9" or "Cas9 domain" encompasses any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of naturally occurring or engineered Cas9. The term Cas9 is not meant to be particularly limiting and may be referred to as "Cas9 or equivalent." Exemplary Cas9 proteins are further described herein and / or in the art and are incorporated herein by reference. The present disclosure is not limited with respect to the specific napDNAbp employed in the base editors of the present disclosure.

[0155] As used herein, the terms "compact Cas9 protein," "compact napDNAbp," and "compact variant [of a Cas protein]" refer to a Cas9 protein or variant having an amino acid length of less than about 1250 amino acids. In some embodiments, a compact Cas9 protein or compact napDNAbp contains less than 1250, 1240, 1230, 1220, 1210, 1200, 1190, 1180, 1170, 1160, 1150, 1140, 1130, 1120, 1110, 1100, 1050, 1000, 950, 900, 850, 800, 750, 700, 650, 600, 550, or 500 amino acids in length. These terms also encompass any Cas9 protein or variant encoded by a nucleic acid sequence having a length of less than about 3750 nucleotides. The base editor of the present disclosure may include a compact napDNAbp and / or a compact Cas9 protein. In some embodiments, the compact Cas9 protein is about 350 amino acids shorter than SpCas9. In some embodiments, the compact Cas9 protein is about 1000 amino acids in length. In some embodiments, the compact protein is a compact variant of S. pyogenes Cas9 (SpCas9), Cpf1, CasX, CasY, C2c1, C2c2, C2c3, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Cas3, or CasΦ. A "compact variant" may refer to a Cas9 protein that has one or more truncations or one or more deletions relative to a wild-type Cas9 protein, e.g., wild-type SpCas9 or Cpf1.

[0156] The Cas9 strain of the strain is specifically designated as "Complete genome sequence of an M1 strain of Streptococcus." pyogenes."Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G, Lyon K, Primeaux C, Sezate S, Suvorov AN, Kenton S, Lai HS, Lin SP, Qian Y, Jia HG, Najar FZ, Ren Q, Zhu H, Song L, White J, Yuan X, Clifton SW, Roe BA, McLaughlin RE, Proc Natl. Sci. USA 98:4658–4663(2001); J., Charpentier E., Nature 471:602-607(2011);および"A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity."Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012). It is based on a slightly less expensive background.

[0157] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not meant to be limiting. The base editors of the present disclosure may use any suitable napDNAbp, including any suitable Cas9 or Cas9 equivalent.

[0158] napDNAbp nickase In some embodiments, the disclosed base editor may comprise a napDNAbp domain that comprises a nickase. In some embodiments, the base editor described herein comprises a Cas9 nickase. The term "Cas9 nickase" in "nCas9" refers to a variant of Cas9 that can introduce a single-stranded break into a double-stranded DNA molecular target. In some embodiments, the Cas9 nickase comprises only a single functional nuclease domain. Wild-type Cas9 (e.g., classical SpCas9) comprises two separate nuclease domains, the RuvC domain (which cleaves the non-protospacer DNA strand) and the HNH domain (which cleaves the protospacer DNA strand). In one embodiment, the Cas9 nickase comprises a mutation on the RuvC domain that inactivates RuvC nuclease activity. For example, mutations at aspartic acid (D) 10, histidine (H) 983, aspartic acid (D) 986, or glutamic acid (E) 762 have been reported as loss-of-function mutations in the RuvC nuclease domain and generate functional Cas9 nickases (see, e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, which is incorporated herein by reference). Thus, the nickase mutations on the RuvC domain can include D10X, H983X, D986X, or E762X, where X is any amino acid other than the wild-type amino acid. In certain embodiments, the nickase can be D10A, H983A, or D986A, or E762A, or a combination thereof.

[0159] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises S. pyogenes Cas9 nickase (SpCas9n). In some embodiments, the napDNAbp domain of any of the disclosed base editors is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to or comprises SEQ ID NO: 365 or 370. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 365. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 370.

[0160] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises S. aureus Cas9 nickase (SaCas9n). In some embodiments, the napDNAbp domain of any of the disclosed base editors is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to or comprises SEQ ID NO: 438. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 438.

[0161] In various embodiments, the Cas9 nickase may have a mutation in the RuvC nuclease domain and may have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. [Table A-1] [Table A-2] [Table A-3] [Table A-4] [Table A-5] [Table A-6]

[0162] Compact Cas9 variants with altered PAM specificity In some embodiments, the napDNAbp comprises a compact Cas protein, such as Cas9 from C. jejuni, S. auricularis, N. meningitidis, or S. aureus. In exemplary embodiments, the napDNAbp comprises a CjCas9 nickase, a SauriCas9 nickase, an Nme2Cas9 nickase, a SaCas9 nickase, or a SaKKH-Cas9 nickase. In some embodiments, the napDNAbp is not an Nme2Cas9 protein or nickase. In some embodiments, the napDNAbp is not a SaCas9 protein or nickase.

[0163] The base editors of the present disclosure may also include Cas9 variants with altered PAM specificity. Some aspects of the present disclosure provide Cas9 proteins that exhibit activity against target sequences that do not include a classical PAM (5'-NGG-3', where N is A, C, G, or T) at their 3' end. In some embodiments, the Cas9 protein exhibits activity against target sequences that include a 5'-NGG-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity against target sequences that include a 5'-NNG-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity against target sequences that include a 5'-NNA-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity against target sequences that include a 5'-NNC-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity against target sequences that include a 5'-NNT-3' PAM sequence at their 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NGT-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NGA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NGC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NAA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NAC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NAT-3' PAM sequence at its 3' end. In still other embodiments, the Cas9 protein exhibits activity against a target sequence that comprises a 5'-NAG-3' PAM sequence at its 3' end.

[0164] In some embodiments, the disclosed base editors comprise a napDNAbp domain that comprises SpCas9-NG with a PAM corresponding to NGN. In some embodiments, the disclosed base editors comprise a napDNAbp domain that has a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SpCas9-NG. The sequence of SpCas9-NG is illustrated below:

[0165] In some embodiments, the disclosed base editor comprises a napDNAbp domain comprising S. aureus Cas9 nickase KKH or SaCas9-KKH, which has a PAM corresponding to NNNRRT or NNGRRT. This Cas9 variant contains amino acid substitutions D10A, E782K, N968K, and R1015H ("KKH") relative to the wild-type SaCas9 submitted as SEQ ID NO: 377. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SaCas9-KKH. SaCas9 (and SaKKH-Cas9) is 1053 amino acids in length. The sequence of SaCas9-KKH (nickase) is illustrated below:

[0166] S. aureus Cas9 nickase KKH (SaCas9-KKH)

[0167] In some embodiments, the disclosed base editor comprises a napDNAbp comprising a Cas9 protein from Staphylococcus Auricularis (S. auri Cas9 or SauriCas9). In some embodiments, the disclosed base editor comprises a SauriCas9 nickase. SauriCas9 recognizes NNGG and NNNGG PAMs. The sequence of SauriCas9 (nickase) is submitted as SEQ ID NO: 479. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 479. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 479. The length of this protein is 1061 amino acids.

[0168] In some embodiments, the napDNAbp comprises a SauriCas9-KKH variant or a SauriCas9-KKH nickase variant. SauriCas9-KKH contains the corresponding triple KKH mutations: Q788K, Y973K, and R1020H. See Hu et al. (2020) PLoS Biol. 18(3): e3000686, which is incorporated herein by reference.

[0169] In some embodiments, the disclosed base editors comprise a napDNAbp domain that includes a S. pyogenes Cas9 nickase KKH or SpCas9-KKH, which has a PAM corresponding to an NNNRRT.

[0170] In some embodiments, the disclosed base editor comprises a napDNAbp, which comprises a compact Cas9 orthologue from Neisseria meningitidis (Nme or Nme2). In some embodiments, the napDNAbp comprises Nme2Cas9. In some embodiments, the disclosed base editor comprises an Nme2Cas9 nickase. Nme2Cas9 is a simple dinucleotide PAM, NNNNCC, or N, as described in Edraki et al., Molecular Cell 73, 714-726, which is incorporated herein by reference. 4 It recognizes CC, where N is any nucleotide. The sequence of Nme2Cas9 is submitted as SEQ ID NO:5. In some embodiments, the disclosed base editors comprise a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:5. In some embodiments, the disclosed base editors comprise a napDNAbp comprising SEQ ID NO:5. The length of this protein is 1082 amino acids.

[0171] The amino acid sequence of NmeCas9 is provided below as SEQ ID NO: 6. In some embodiments, the disclosed base editors comprise a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 6. In some embodiments, the disclosed base editors comprise a napDNAbp comprising SEQ ID NO: 6. The protein is 1083 amino acids in length.

[0172] In some embodiments, the disclosed base editor comprises a napDNAbp comprising a compact Cas9 orthologue from Campylobacter jejuni (CjCas9). In some embodiments, the napDNAbp comprises CjCas9. In some embodiments, the disclosed base editor comprises a CjCas9 nickase. CjCas9 recognizes NNNNACA and NNNNACAC PAMs. See Kim et al., Nature Communications 8(14500):1-12 (2017), which is incorporated herein by reference. The sequence of CjCas9 (nickase) is submitted as SEQ ID NO: 379. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 379. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 379. The length of this protein is 984 amino acids. MARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRLARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPVVLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKPERKKPSGALHEETFRKEEEFYQSYGGKEGVLKALELGKIRKVNGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKDMQEPEFVYYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKSIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKK(SEQ ID NO: 379)

[0173] Additional compact Cas proteins In various embodiments, the nucleic acid programmable DNA binding protein includes, without limitation, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, circularly permuted Cas9 domains, e.g., CP1012, CP1028, CP1041, CP1249, and These include CP1300, Argonaute (Ago) domain, Cas9-KKH, SmacCas9, SpRY, SpRY-HF1, Spy-macCas9, SpCas9-VRQR, SpCas9-VRER, SpCas9-VQR, SpCas9-EQR, SpCas9-NRRH, SpaCas9-NRTH, SpCas9-NRCH, LbCas12a, AsCas12a, CeCas12a, MbCas12a, CasΦ, SpCas9-NG-CP1041, and compact variants of SpCas9-NG-VRQR.

[0174] In still other embodiments, the napDNAbp can include a compact Cas9 orthologue from Staphylococcus lugdunensis Cas9 (SlugCas9), Staphylococcus lutrae Cas9 (SlutrCas9), or Staphylococcus haemolyticus Cas9 (ShaCas9). See Hu et al., Nucleic Acids Research, 49(7), April 2021, 4008-4019, which is incorporated herein by reference. SlugCas9, SlutrCas9, and ShaCas9 proteins recognize the NNGG, NNGG / NNGA, and NNGG PAMs.

[0175] In still other embodiments, the Cas protein may include any CRISPR-associated protein: Cas12a, Cas12b, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6 , Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, preferably including a nickase mutation (e.g., a mutation corresponding to the D10A mutation in the wild-type SpCas9 polypeptide of SEQ ID NO: 326).

[0176] In certain embodiments, the base editor contemplated herein may include a Cas9 protein that is smaller in molecular weight than the classical SpCas9 sequence. In some embodiments, the smaller size of the Cas9 variant may facilitate delivery to cells, for example, by AAV vectors, expression vectors, or other delivery means. The classical SpCas9 protein is 1368 amino acids in length and has a predicted molecular weight of 158 kilodaltons. As used herein, the term "small size Cas9 variant" refers to a Cas9 variant having less than about 1300 amino acids, or at least less than 1290 amino acids, or less than 1280 amino acids, or less than 1270 amino acids, or less than 1260 amino acids, or less than 1250 amino acids, or less than 1240 amino acids, or less than 1230 amino acids, or less than 1220 amino acids, or less than 1210 amino acids, or less than 1200 amino acids, or less than 1190 amino acids, or less than 1180 amino acids, or less than 1170 amino acids, or less than 1160 amino acids, or less than 1150 amino acids, or less than 1140 amino acids, or less than 1130 amino acids. "Cas9 variants" refers to Cas9 variants, either naturally occurring, engineered or otherwise, that are less than 1,120 amino acids, or less than 1,110 amino acids, or less than 1,100 amino acids, or less than 1,050 amino acids, or less than 1,000 amino acids, or less than 950 amino acids, or less than 900 amino acids, or less than 850 amino acids, or less than 800 amino acids, or less than 750 amino acids, or less than 700 amino acids, or less than 650 amino acids, or less than 600 amino acids, or less than 550 amino acids, or less than 500 amino acids, but greater than at least about 400 amino acids, and that retain a required function of the Cas9 protein.

[0177] In various embodiments, the base editors disclosed herein can include one of the small Cas9 variants described as follows, or a Cas9 variant thereof having at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to any reference small Cas9 protein. Exemplary small Cas9 variants include, but are not limited to, SauriCas9, SaCas9, CjCas9, Nme2Cas9, AsCas12a, and LbCas12a.

[0178] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises LbCas12a, such as wild-type LbCas12a. In some embodiments, the napDNAbp domain of any of the disclosed base editors is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to or comprises SEQ ID NO: 381. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 381.

[0179] In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises AsCas12a, such as wild-type AsCas12a. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises mutant AsCas12a, such as engineered AsCas12a or enAsCas12a. In some embodiments, the napDNAbp domain of any of the disclosed base editors is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to or includes SEQ ID NO: 383. In some embodiments, the napDNAbp domain of any of the disclosed base editors comprises the amino acid sequence of SEQ ID NO: 383. [Table B-1] [Table B-2] [Table B-3]

[0180] Additional exemplary Cas9 equivalent protein sequences may include: [Table C-1] [Table C-2] [Table C-3] [Table C-4] [Table C-5] [Table C-6] [Table C-7] [Table C-8] [Table C-9]

[0181] The base editor described herein can also include Cas12a / Cpf1 (dCpf1) variants that can be used as guide nucleotide sequence programmable DNA binding protein domain.Cas12a / Cpf1 protein has RuvC-like endonuclease domain that is similar to RuvC domain of Cas9, but does not have HNH endonuclease domain, and N-terminus of Cpf1 does not have alpha-helix recognition lobe of Cas9.It was shown in Zetsche et al., Cell, 163, 759-771, 2015 (which is incorporated herein by reference) that RuvC-like domain of Cpf1 is responsible for cleaving both DNA strands, and inactivation of RuvC-like domain inactivates Cpf1 nuclease activity.

[0182] Cytidine deaminase domain In some embodiments, the base editor comprises a deaminase that is a cytosine deaminase. In some embodiments, the cytosine deaminase domain is fused to the N-terminus of the napDNAbp.

[0183] In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is an APOBEC1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, an APOBEC3H deaminase, or an APOBEC4 deaminase. In some embodiments, the deaminase is an activation-induced deaminase (AID). In some embodiments, the deaminase is a lamprey CDA1 (pmCDA1) deaminase. In some embodiments, the deaminase is from a human, a chimpanzee, a gorilla, a monkey, a cow, a dog, a rat, or a mouse. In some embodiments, the deaminase is from a human. In some embodiments, the deaminase is from a rat. In some embodiments, the deaminase is human APOBEC1 deaminase. In some embodiments, the deaminase is pmCDA1. In some embodiments, the deaminase is human APOBEC3G. In some embodiments, the deaminase is human APOBEC3G variant. In some embodiments, the deaminase is rat APOBEC1.

[0184] In certain embodiments, the cytidine deaminase domain is a "FERNY" polypeptide having an amino acid sequence according to SEQ ID NO:393, or an amino acid sequence that is at least 80%, 85%, 90%, 95, 98%, 99%, or 99.5% identical to SEQ ID NO:393, as follows: MFERNYDPRELRKETYLLYEIKWGKSGKLWRHWCQNNRTQHAEVYFLENIFNARRFNPSTHCSITWYLSWSPCAECSQKIVDFLKEHPNVNLEIYVARLYYHEDERNRQGLRDLVNSGVTIRIMDLPDYNYCWKTFVSDQGGDEDYWPGHFAPWIKQYSLKL (SEQ ID NO: 393)

[0185] In certain other embodiments, the cytidine deaminase domain is a domain evolved from a wild-type domain, e.g., evolved through phage-assisted continuous evolution (PACE). In some embodiments, the cytidine deaminase domain is an "evoFERNY" polypeptide having an amino acid sequence according to SEQ ID NO:394, or an amino acid sequence that is at least 85%, 90%, 95, 98%, 99%, or 99.5% identical to SEQ ID NO:394, which contains H102P and D104N substitutions relative to SEQ ID NO:393, as follows: [ka] The FERNY and evoFERNY deaminase domains are 162 amino acids in length. Therefore, these domains are shorter than any one of SEQ ID NOs: 276-277, 281, 133-134, 292-295, and 487.

[0186] In exemplary embodiments, the disclosed CBE comprises a FERNY or evoFERNY deaminase domain. In some embodiments, the disclosed CBE comprises a deaminase domain comprising the amino acid sequence of SEQ ID NO: 393 or 394.

[0187] The state-of-the-art cytosine base editor BE3.9 contains the rat APOBEC1 (rAPOBEC1) cytidine deaminase domain. They have high overall activity, severely impaired activity editing GC targets, and high editing in TC targets. Alternative deaminases have been demonstrated as base editors. Both AID and CDA work well on GC targets, but generally have lower activity than APOBEC1. APOBEC3G works less well than all of these (see Komor, AC et al. Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity. Sci Adv 3, eaao4774 (2017), which is incorporated herein by reference). An implementation of TARGET-AID base editing uses CDA (Nishida, K. et al. Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science 353, aaf8729-aaf8729 (2016), which is incorporated herein by reference). "FERNY" is an N- and C-terminally truncated ancestral sequence reconstruction based on the APOBEC family phylogenetic tree. rAPOBEC1: 229aa; FERNY: 161aa. Sequence similarity to rAPOBEC1 is 55%. The evolved FERNY genotype also has high GC activity and is as active as APOBEC, despite being a shorter protein. The evolved FERNY is described in further detail in WO 2019 / 023680, published January 31, 2019, which is incorporated herein by reference. Additional exemplary cytidine deaminases are disclosed in WO 2021 / 108717, published June 3, 2021.

[0188] Non-limiting examples of suitable cytosine deaminase domains are provided below as SEQ ID NOs: 276-277, 281, 133-134, 292-295, and 487. In some embodiments, the deaminase is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences submitted below.

[0189] Human AID MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 276)

[0190] Mouse AID MDSLLMKQKKFLYHFKNVRWAKGRHETYLCYVVKRRDSATSCSLDFGHLRNKSGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVAEFLRWNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIGIMTFKDYFYCWNTFVENRERTFKAWEGLHENSVRLTRQLRRILLPLYEVDDLRDAFRMLGF (SEQ ID NO: 277)

[0191] Rat APOBEC-3 (rAPOBEC3) MGPFCLGCSHRKCYSPIRNLISQETFKFHFKNLRYAIDRKDTFLCYEVTRKDCDSPVSLHHGVFKNKDNIHAEICFLYWFHDKVLKVLSPREEFKITWYMSWSPCFECA EQVLRFLATHHNLSLDIFSSRLYNIRDPENQQNLCRLVQEGAQVAAMDLYEFKKCWKKFVDNGGRRFRPWKKLLTNFRYQDSKLQEILRPCYIPVPSSSSSTLSNICLTK GLPETRFCVERRRVHLLSEEEFYSQFYNQRVKHLCYYHGVKPYLCYQLEQFNGQAPLKGCLLSEKGKQHAEILFLDKIRSMELSQVIITCYLTWSPCPNCAWQLAAFKRDRPDLILHIYTSRLYFHWKRPFQKGLCSLWQSGILVDVMDLPQFTDCWTNFVNPKRPFWPWKGLEIISRRTQRRLHRIKESWGLQDLVNDFGNLQLGPPMS (SEQ ID NO: 281)

[0192] Human APOBEC-3G MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCWSKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN (SEQ ID NO: 133)

[0193] Human APOBEC-3F MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPRLDAKIFRGQVYSQPEHHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYSEGQPFMPWYKFDDNYAFLHRTLKEILRNPMEAMYPHIFYFHFKNLRKAYGRNESWLCFTMEVVKHHSPVSWKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPEECAGEVAEFLARHSNVNLTIFTARLYYFWDTDYQEGLRSLSQEGASVEIMGYKDFKYCWENFVYNDDEPFKPWKGLKYNFLFLDSKLQEILE (SEQ ID NO: 134)

[0194] Human APOBEC-1 MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR (SEQ ID NO: 292)

[0195] Mouse APOBEC-1 MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNTRCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYPPSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK (sequence number 293)

[0196] Rat APOBEC-1 (rAPOBEC1) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO: 294)

[0197] Petromyzon marinus CDA1(pmCDA1) MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAV (SEQ ID NO: 295)

[0198] Evolved pmCDA1 (evoCDA1) MTDAEYVRIHEKLDIYTFKKQFSNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWVCKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMFQVKILHTTKSPAV (SEQ ID NO: 487)

[0199] Adenosine deaminase domain The present disclosure provides adenosine deaminase variants that have activity on deoxyadenosine nucleosides on DNA.Therefore, the variants provided herein are deoxyadenosine deaminases.In some embodiments, the disclosed adenosine deaminase is a variant of known adenosine deaminase TadA7.10, which comprises the following mutations compared with wild-type ecTadA (SEQ ID NO: 325): W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, R152P, E155V, I156F and K157N. In some embodiments, the disclosed adenosine deaminase is a variant of TadA from a species other than E. coli, such as Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis.

[0200] In some embodiments, the disclosed adenosine deaminase domain comprises a variant TadA-8e of E. coli TadA 7.10. TadA-8e (submitted as SEQ ID NO: 433) contains the following substitutions relative to TadA7.10 (SEQ ID NO: 315): T111, D119, F149, R26, V88, A109, H122, T166, and D167. TadA-8e is disclosed in PCT Publication No. WO 2021 / 158921, published August 12, 2021, which is incorporated herein by reference. In some embodiments, the adenosine deaminase domain comprises TadA-8e (V106W), which contains a V106W substitution relative to TadA-8e.

[0201] In various embodiments, the disclosed adenosine deaminases hydrolytically deaminate targeted adenosines on a nucleic acid of interest to inosine, which is read as guanosine (G) by a DNA polymerase enzyme.

[0202] These variants can include a domain of any of the disclosed base editors (i.e., the adenosine deaminase domain of an adenine base editor). In some embodiments, any of the disclosed adenine base editors can deaminate an adenosine on a nucleic acid sequence (e.g., DNA or RNA). The disclosed adenine base editors can further deaminate an adenine on DNA.

[0203] Exemplary non-limiting embodiments of adenosine deaminase are provided herein. In some embodiments, the adenosine deaminase domain of any of the disclosed base editors comprises a single adenosine deaminase or monomer. In some embodiments, the adenosine deaminase domain comprises two, three, four or five adenosine deaminases. In some embodiments, the adenosine deaminase domain comprises two adenosine deaminases or a dimer. In some embodiments, the deaminase domain comprises a dimer of an engineered (or evolved) deaminase and a wild-type deaminase, such as a wild-type E. coli derived deaminase. The mutations provided herein (e.g., mutations on ecTadA) may be used in combination with other adenine base editors, such as those described in WO 2018 / 027078 published August 2, 2018; WO 2019 / 079347 published April 25, 2019; International Application No. PCT / US2019 / 033848 filed May 23, 2019, which published as WO 2019 / 226593 on November 28, 2019; U.S. Patent Application Publication No. 2018 / 0073012 published March 15, 2018, which published as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Application Publication No. 2018 / 0073012 published on January 1, 2019 as U.S. Patent No. 10,167,457 published on January 1, 2019 as U.S. Patent No. 10,167,457. It should be understood that the methods and techniques described herein may be applied to adenosine deaminase as provided in U.S. Patent Application Publication No. 2017 / 0121693, published May 4, 2017; International Publication No. 2017 / 070633, published April 27, 2017; U.S. Patent Application Publication No. 2015 / 0166980, published June 18, 2015; U.S. Patent No. 9,840,699, published December 12, 2017; and U.S. Patent No. 10,077,453, published September 18, 2018; PCT Application No. PCT / US2020 / 28568, filed April 16, 2020, and International Publication No. WO 2021 / 158921, published August 12, 2021, all of which are incorporated herein by reference in their entirety.

[0204] In some embodiments, any of the adenosine deaminases provided herein can deaminate adenine, for example, deaminate adenine on deoxyadenosine nucleosides of DNA. The adenosine deaminase can be derived from any suitable organism, for example, E. coli. In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations corresponding to any of the mutations provided herein, for example, mutations on ecTadA. Those skilled in the art will be able to identify the corresponding residues on any homologous protein and on the corresponding coding nucleic acid by methods well known in the art, for example, by sequence alignment and determining homologous residues. Exemplary TadA deaminases are provided from Bacillus subtilis (submitted in full as SEQ ID NO: 318), S. aureus (SEQ ID NO: 317), and S. pyogenes (SEQ ID NO: 448). Amino acid substitutions of E. coli TadA-8e and homologous mutations on B. subtilis, S. aureus, and S. pyogenes TadA deaminase are shown. Thus, one skilled in the art will be able to generate mutations on any naturally occurring adenosine deaminase that correspond to any of the mutations described herein, for example, any of the mutations identified on ecTadA (for example, with homology to ecTadA). In some embodiments, the adenosine deaminase is from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from E. coli.

[0205] In some embodiments, the adenosine deaminase comprises TadA9 or a variant thereof. TadA9 contains V82S and Q154R substitutions relative to TadA-8e. (In other words, TadA9 contains Y147R, Q154R, and I76Y mutations relative to TadA7.10.) In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA9 (SEQ ID NO: 33). TadA9 may be referred to in the art as TadA*8.9. ABE containing TadA9 deaminase is referred to herein as ABE9. TadA9 is described in additional detail in Gaudelli et al., Nat Biotechnol. 2020 Jul;38(7):892-900 and WO 2021 / 050571, published March 18, 2021, each of which is incorporated herein by reference.

[0206] In some embodiments, the adenosine deaminase comprises TadA20 or a variant thereof. TadA20 contains I76Y, V82S, Y123H, Y147R, and Q154R substitutions relative to TadA7.10. Additional details of TadA20 are described in Gaudelli et al., Nat Biotechnol. 2020 Jul;38(7):892-900 and WO 2021 / 050571 published March 18, 2021. TadA20 may be referred to in the art as TadA*8.20. In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA20 (SEQ ID NO: 326). An ABE containing TadA20 deaminase is referred to herein as ABE20. It may also be referred to in the art as ABE8.20, ABE8.20-d, or ABE8.20-m.

[0207] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences of SEQ ID NOs: 33, 315, 317-326, 433, or 448-449.

[0208] It should be understood that the adenosine deaminases provided herein can include one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides adenosine deaminases having a certain percent identity plus any of the mutations described herein or combinations thereof. Any of the adenosine deaminases described herein can be truncated variants of any of the other adenosine deaminases described herein, e.g., any of the adenosine deaminases of SEQ ID NO: 33, 315, 317-326, 433, or 448-449.

[0209] Exemplary truncated adenosine deaminases may include truncations of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more than 15 amino acids from the N-terminus. Other exemplary truncated adenosine deaminases may include truncations of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more than 15 amino acids from the C-terminus. In some embodiments, the adenosine deaminase domain comprises a truncated version of wild-type ecTadA as set forth in SEQ ID NO: 324. Any of the adenosine deaminases described herein may include an N-terminal methionine (M) amino acid residue.

[0210] It should be understood that any of the mutations provided herein (based on the ecTadA amino acid sequence of SEQ ID NO: 315, for example) can be introduced into other adenosine deaminases, such as S. aureus TadA (saTadA), A. aeolicus TadA (AaTadA), or another adenosine deaminase (for example, another bacterial adenosine deaminase), such as the sequences provided below. It will be clear to one skilled in the art how to identify amino acid residues from other adenosine deaminases that are homologous to the mutated residues on ecTadA. Thus, any of the identified mutations on ecTadA can be made on other adenosine deaminases that have homologous amino acid residues. Any of the mutations provided herein can be made on ecTadA or another adenosine deaminase individually or in any combination. Any of the mutated deaminases provided herein can be used in the context of an adenine base editor. Any of the deaminases provided herein can have a sequence beginning with a methionine ("M") before the first amino acid shown in the sequence below.

[0211] Exemplary adenosine deaminase variants of the present disclosure are described below. In certain embodiments, the adenosine deaminase domain comprises an adenosine deaminase having a sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% sequence identity to one of the following:

[0212] TadA 7.10 (E. coli) SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 315)

[0213] TadA-8e (E. coli) SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 433)

[0214] TadA9 SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILANECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 33)

[0215] TadA20 SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLYDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSTD (SEQ ID NO: 326)

[0216] Staphylococcus aureus TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN (SEQ ID NO: 317)

[0217] Bacillus subtilis TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE (SEQ ID NO: 318)

[0218] Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV (SEQ ID NO: 319)

[0219] Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE (SEQ ID NO: 320)

[0220] Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK(SEQ ID NO: 321)

[0221] Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI(SEQ ID NO: 322)

[0222] Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP(SEQ ID NO: 323)

[0223] Streptococcus pyogenes (S. pyogenes) TadA MPYSLEEQTYFMQEALKEAEKSLQKAEIPIGCVIVKDGEIIGRGHNAREESNQAIMHAEIMAINEANAHEGNWRLLDTTLFVTIEPCVMCSGAIGLARIPHVIYGASNQKFGGADSLYQILTDERLNHRVQVERGLLAADCANIMQTFFRQGRERKKIAKHLIKEQSDPFD (SEQ ID NO: 448)

[0224] Aquifex aeolicus(A. aeolicus)TadA MGKEYFLKVALREAKRAFEKGEVPVGAIIVKEGEIISKAHNSVEELKDPTAHAEMLAIKEACRRLNTKYLEGCELYVTLEPCIMCSYALVLSRIEKVIFSALDKKHGGVVSVFNILDEPTLNHRVKWEYYPLEEASELLSEFFKKLRNNII (SEQ ID NO: 449)

[0225] In some embodiments, the adenosine deaminase domain comprises an N-terminally truncated E. coli TadA. In certain embodiments, the adenosine deaminase comprises the following amino acid sequence: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 324).

[0226] In some embodiments, the TadA deaminase is a full-length E. coli TadA deaminase (ecTadA). For example, in certain embodiments, the adenosine deaminase domain comprises a deaminase comprising the following amino acid sequence: MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 325)

[0227] In some embodiments, the base editor comprises an adenosine deaminase monomer. In other aspects, the base editor comprises an adenosine deaminase dimer. The base editor may comprise a heterodimer of a first adenosine deaminase and a second adenosine deaminase. In some embodiments, the first adenosine deaminase is N-terminal to the second adenosine deaminase on the base editor. In some embodiments, the first adenosine deaminase is C-terminal to the second adenosine deaminase on the base editor. In some embodiments, the first adenosine deaminase and the second deaminase are fused to each other directly or via a linker. In some embodiments, the first adenosine deaminase is fused to the N-terminus of the napDNAbp via a linker and the second deaminase is fused to the C-terminus of the napDNAbp domain via a linker. In other embodiments, the second adenosine deaminase is fused at the N-terminus of the napDNAbp domain via a linker and the first deaminase is fused at the C-terminus of the napDNAbp via a linker.

[0228] Exemplary Base Editors Adenine Base Editor In some aspects, the base editing methods of the present disclosure include the use of an adenine base editor. Exemplary adenine base editors of the present disclosure include monomeric and dimeric versions of the following editors: Sauri-ABE8e, SaKKH-ABE8e, SaABE8e, CjCas9-ABE8e, and Nme2Cas9-ABE8e; SaKKH-ABE8e(V106W), SauriCas9-ABE8e(V106W), CjCas9-ABE8e(V106W), Nme2Cas9-ABE8e(V106W), and SaCas9-ABE8e(V106W); SaKKH -ABE9, SauriCas9-ABE9, CjCas9-ABE9, Nme2Cas9-ABE9, and SaCas9-ABE9; SaKKH-ABE20, SauriCas9-ABE20, CjCas9-ABE20, Nme2Cas9-ABE20, and SaCas9-ABE20; and SaKKH-ABE7.10, SauriCas9-ABE7.10, CjCas9-ABE7.10, Nme2Cas9-ABE7.10, and SaCas9-ABE7.10. In some embodiments, the ABE is Sauri-ABE8e, SaKKH-ABE8e, or SaABE8e. These base editors are 1298, 1291, and 1291 amino acids in length. The additional base editors are 1221 amino acids in length (CjABE8e) and 1319 amino acids in length (Nme2ABE8e). In an exemplary embodiment, the ABE is SaKKH-ABE8e. An exemplary ABE contains an adenosine deaminase domain that includes TadA8e and does not include a second adenosine deaminase (i.e., the adenosine deaminase domain consists of a deaminase monomer).

[0229] Additional exemplary adenine base editors are disclosed in WO 2020 / 051360, published March 12, 2020; WO 2021 / 050571; WO 2020 / 168132, published August 20, 2020; and WO 2021 / 158921, published August 12, 2021, each of which is incorporated herein by reference.

[0230] ABE8e may be referred to in the art as "ABE8" or "ABE8.0". The ABE8e base editor and variants thereof may comprise an adenosine deaminase domain containing a TadA-8e adenosine deaminase monomer (monomeric form) or a TadA-8e adenosine deaminase homodimer or heterodimer (dimeric form). In some embodiments, the architecture of a base editor comprising an adenosine deaminase domain and napDNAbp is as follows: NH 2 -[adenosine deaminase]-[napDNAbp domain]-COOH; or NH 2 -[napDNAbp domain]-[adenosine deaminase]-COOH. In certain embodiments, the base editor is 2 The ABE8e monomer architecture comprises: -[NLS]-[adenosine deaminase]-[napDNAbp domain]-[NLS]-COOH, where "NLS" is a nuclear localization sequence.

[0231] In some aspects, the present disclosure provides a complex of an adenine base editor and a guide RNA. Exemplary disclosed complexes include any of the following ABEs together with a guide RNA (e.g., a single guide RNA): Sauri-ABE8e, SaKKH-ABE8e, SaABE8e, CjCas9-ABE8e, and Nme2Cas9-ABE8e; SaKKH-ABE8e(V106W), SauriCas9-ABE8e(V106W), CjCas9-ABE8e(V106W), Nme2Cas9-ABE8e(V106W), and SaCas9-ABE8e(V106W); S aKKH-ABE9, SauriCas9-ABE9, CjCas9-ABE9, Nme2Cas9-ABE9, and SaCas9-ABE9; SaKKH-ABE20, SauriCas9-ABE20, CjCas9-ABE20, Nme2Cas9-ABE20, and SaCas9-ABE20; and SaKKH-ABE7.10, SauriCas9-ABE7.10, CjCas9-ABE7.10, Nme2Cas9-ABE7.10, and SaCas9-ABE7.10. Other ABEs can be used to deaminate A nucleobases according to the disclosed complexes.

[0232] In some embodiments, the ABE is CjCas9-ABE8e. During adenine editing evaluation, this editor showed a higher context preference for pyrimidines at the nucleotide position 5' of the targeted adenine base ( Y A>> RA). As used herein, "preference" and "context preference" refer to a product purity of greater than 40% for the target adenosine. Thus, in some aspects, the present disclosure provides ABEs with pyrimidine ("Y") context preference, where "context" refers to the presence of a pyrimidine or purine ("R") immediately 5' of the adenine base (or target adenine base) to be edited. For example, CjABE8e editor and variants thereof. These ABEs can have a preference for editing an adenosine on a target nucleic acid sequence of 5'-YAN-3', where Y is C or T; N is A, T, C, G, or U; and A is the target adenosine. Thus, in some embodiments, an ABE is provided that has a context preference for deaminating an adenosine on a target nucleic acid sequence of 5'-YAN-3', where Y is C or T and N is A, T, C, G, or U; and A is the target adenosine.

[0233] An exemplary AAV-encoded adenine base editor construct is shown in Figure 2A. This construct contains the SaABE8e base editor operably controlled by the EFS promoter and bGH polyA sequence. It also contains a guide RNA encoded in the reverse orientation, as indicated by the arrow pointing away from the 3' end.

[0234] The disclosed ABE complex may have an on-target editing efficiency of more than 50% after contacting with a nucleic acid molecule containing a target sequence. Further exemplary ABE complexes have an on-target editing efficiency of more than 60% after contacting with a nucleic acid molecule containing a target sequence. Further exemplary ABEs have an on-target editing efficiency of more than 65%, more than 70%, more than 75%, more than 80%, more than 82.5%, or more than 85% after contacting with a nucleic acid molecule containing a target sequence. The disclosed ABE complexes may exhibit an indel frequency of less than 2.5%, less than 2.4%, less than 2.2%, less than 2.0%, less than 1.75%, less than 1.5%, less than 1.3%, less than 1.1%, or less than 1.0% after contacting with a nucleic acid molecule containing a target sequence.

[0235] In some aspects, the present disclosure provides a base editor comprising a napDNAbp domain and an adenosine deaminase domain as described herein. The Cas9 domain can be any of the Cas9 domains or Cas9 proteins provided herein (e.g., Cas9 nickase or nCas9). In some embodiments, any of the Cas9 domains or Cas9 proteins provided herein (e.g., nCas9) can be fused with any of the adenosine deaminase domains provided herein.

[0236] In some embodiments, the base editor comprising adenosine deaminase and napDNAbp (e.g., Cas9 domain) does not include a linker sequence. In some embodiments, a linker is present between the adenosine deaminase domain and / or adenosine deaminase and napDNAbp. In some embodiments, the "]-[" used in the general architecture above indicates the presence of an optional linker. In some embodiments, the adenosine deaminase domain and napDNAbp domain are fused via any of the linkers provided herein. For example, in some embodiments, the adenosine deaminase domain (which may include one or more adenosine deaminases) and napDNAbp are fused via any of the linkers provided below in the section entitled "Linkers."

[0237] In some embodiments, the adenine base editor comprises an adenosine deaminase that comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% sequence identity to SEQ ID NO: 433 (TadA-8e). In some embodiments, the adenine base editor comprises the sequence of SEQ ID NO:433.

[0238] In some embodiments, an adenine base editor of the disclosure comprises the sequence of SEQ ID NO: 181. In some embodiments, an adenine base editor of the disclosure comprises the sequence of SEQ ID NO: 182. In other embodiments, an adenine base editor of the disclosure comprises the sequence of SEQ ID NO: 183. In other embodiments, an adenine base editor of the disclosure comprises the sequence of SEQ ID NO: 171 or 172.

[0239] In some embodiments, any of the adenine base editors described herein can include an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more than 30 amino acids that differ relative to the amino acid sequence of any of SEQ ID NOs: 171-172 and 181-183. These differences can include inserted, deleted, or substituted amino acids relative to the reference sequence. In some embodiments, the disclosed adenosine deaminase domains contain a stretch of about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 300, about 400, about 500, or more than 500 contiguous amino acids in common with either of SEQ ID NOs: 171-172 and 181-183.

[0240] Exemplary base editors include sequences that are at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% identical to any of the following amino acid sequences (SEQ ID NOs: 171-172 and 181-183). In some embodiments, the disclosed base editors have a sequence that includes any of the following amino acid sequences:

[0241] [ka]

[0242] [ka]

[0243] [ka]

[0244] [ka] [ka]

[0245] [ka]

[0246] Cytosine Base Editor In some aspects, the present disclosure provides cytosine base editors (CBEs). Examples of these CBEs include CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9. The CBEs of the present disclosure can be any of the variants of CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9.

[0247] In some embodiments, the disclosed cytosine base editors (CBEs) can comprise a fusion protein that includes: (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain; (ii) a cytidine deaminase domain; and (iii) a uracil glycosylase inhibitor domain (UGI). In various embodiments, the disclosed CBEs contain a single UGI domain. The disclosed CBEs can be structurally arranged in a variety of configurations, including, but not limited to, the following: NH 2 -[cytidine deaminase domain]-[napDNAbp domain]-[UGI]-COOH; NH 2 -[cytidine deaminase domain]-[UGI]-[napDNAbp domain]-COOH; NH 2 -[napDNAbp domain]-[UGI]-[cytidine deaminase domain]-COOH; NH 2 -[napDNAbp domain]-[cytidine deaminase domain]-[UGI]-COOH; NH 2 -[UGI]-[cytidine deaminase domain]-[napDNAbp domain]-COOH; or NH 2-[UGI]-[napDNAbp domain]-[cytidine deaminase domain]-COOH, where each occurrence of "]-[" includes an optional linker.

[0248] An exemplary construct encoding a single AAV-CBE is shown in FIG. 14. This CBE contains the BE3.9 architecture, with the deaminase FERNY (or evoFERNY) as the cytidine deaminase domain located 5' to the napDNAbp domain. In this construct, the cytosine base editor contains the CjCas9 napDNAbp domain. A human U6-controlled ("hU6") sgRNA is located at the 3' end, in the reverse orientation relative to the base editor. This construct contains an EFS promoter driving the base editor, and an SV40 late polyA sequence. The construct has a length of 5.012 kb. Where indicated, "BE3. 9" refers to the BE3.9 CBE architecture, i.e., NH 2 -[first nuclear localization sequence]-[cytosine deaminase domain]-[32 aa linker]-[napDNAbp domain]-[9 aa linker]-[first UGI domain]-[second nuclear localization sequence]-COOH. The disclosed CBEs can further comprise one or more nuclear localization signals (NLS).

[0249] Additional exemplary CBEs are disclosed in International Publication No. WO 2019 / 023680, published January 31, 2019, and International Publication No. WO 2021 / 108717, published June 3, 2021, each of which is incorporated herein by reference. Additional CBEs are disclosed in Villiger, L. et al. Nature Medicine 24, 1519-1525 (2018), which is incorporated herein by reference. Villiger et al. developed an intein-split S. aureus CBE. The disclosed single AAV CBE demonstrates comparable, if not improved, activity relative to that shown by Villiger.

[0250] The disclosed CBEs can comprise modified (or evolved) cytosine deaminase domains, e.g., deaminase domains that recognize extended PAM sequences, have improved efficiency in deaminating 5'-GC targets, and / or edit on a narrower target window. In some embodiments, the disclosed cytidine nucleobase editors comprise evolved nucleic acid programmable DNA binding proteins (napDNAbp), e.g., evolved Cas9.

[0251] In some aspects, the disclosure provides a complex of a cytosine base editor and a guide RNA, e.g., any of CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9, and an sgRNA.

[0252] Non-limiting examples of CBEs are provided below as SEQ ID NOs: 19 and 20, which are the FERNY-BE3.9 editor. These editors contain SpCas9 as the napDNAbp domain. The exemplary AAV-encoded CBE of the disclosure contains a CjCas9 domain (SEQ ID NO: 379) in place of the SpCas9 domain of the FERNY-BE3.9 editor described below.

[0253] The following base editor (SEQ ID NO: 19) contains a wild-type FERNY that can be used as a reference base editor. This base editor was evolved to generate the evoFERNY base editor shown as SEQ ID NO: 20. These base editors contain a bpNLS and a single UGI domain. [ka] [ka]

[0254] The following base editor contains evoFERNY, which was evolved based on the base editor provided above (SEQ ID NO: 19). [ka]

[0255] In some embodiments, the disclosed base editor comprises CjCas9-FERNY-BE3.9, which is provided below as SEQ ID NO: 21. In some embodiments, the disclosed base editor comprises CjCas9-evoFERNY-BE3.9, which is provided below as SEQ ID NO: 22. Any of the disclosed base editors may comprise a sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identity to any of SEQ ID NOs: 21 and 22. The base editor may comprise a sequence of SEQ ID NO: 21 or 22. These base editors contain a bpNLS and a single UGI domain. FERNY deaminase is in italics.

[0256] [ka]

[0257] [ka] [ka]

[0258] Recombinant adeno-associated virus (rAAV) vectors Aspects of the present disclosure relate to the use of recombinant adeno-associated virus vectors for delivery of any of the disclosed nucleic acid molecules. The rAAV particles of the present disclosure include an rAAV vector (i.e., a recombinant genome of rAAV) encapsidated by a viral capsid protein. See U.S. Patent Application Publication No. 2018 / 0127780, published May 10, 2018, and International Publication No. WO 2020 / 236982, published November 26, 2020. The disclosures of each of these are incorporated herein by reference.

[0259] In some embodiments, the AAV nucleic acid vector is single stranded. In some embodiments, the AAV nucleic acid vector is self-complementary. In various embodiments, the rAAV vector of the present disclosure does not contain any inteins.

[0260] In some embodiments, the integration-facilitating viral sequence comprises an inverted terminal repeat (ITR) sequence. In some embodiments, the nucleic acid molecule is flanked on both sides by ITR sequences. In some embodiments, the nucleic acid vector further comprises a region encoding an AAV Rep protein as described herein, contained either within or outside the region flanked by the ITRs. The ITR sequence can be derived from any AAV serotype (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) or from more than one serotype. In some embodiments, the ITR sequence is derived from AAV8 or AAV9. In some embodiments, in the method of packaging any of the disclosed rAAV particles, a nucleic acid plasmid, such as a helper plasmid, is provided that comprises a region encoding a Rep protein and / or a Cap (capsid) protein.

[0261] Thus, in some embodiments, the rAAV particles disclosed herein include rAAV2, rAAV6, rAAV8, rPHP.B, rPHP.eB, or rAAV9 particles, or variants thereof. In certain embodiments, the disclosed rAAV particles are rAAV8 or rAAV9 particles.

[0262] Exemplary rAAV particles provided herein include, but are not limited to, rAAV8-Sauri-ABE8e, rAAV9-Sauri-ABE8e, rAAV8-SaKKH-ABE8e, rAAV9-SaKKH-ABE8e, rAAV8-CjCas9-ABE8e, rAAV9-CjCas9-ABE8e, rAAV8-Nme2Cas9-ABE8e, or rAAV9-Nme2Cas9-ABE8e particles.In certain embodiments, the rAAV particles comprise rAAV8-SaKKH-ABE8e particles.In some embodiments, the rAAV particles comprise rAAV9-CjBE3.9 particles or rAAV8-CjBE3.9 particles.

[0263] ITR sequences and plasmids containing ITR sequences are known in the art and are commercially available (see, e.g., Vector Biolabs, Philadelphia, PA; Cellbiolabs, San Diego, CA; Agilent Technologies, Santa Clara, Ca; and Addgene, Cambridge, MA; and Gene delivery to skeletal muscle results in sustained expression and systemic delivery of a therapeutic protein. Kessler PD, Podsakoff GM, Chen X, McQuiston SA, Colosi PC, Matelis LA, Kurtzman GJ, Byrne BJ. Proc Natl Acad Sci USA. 1996 Nov 26;93(24):14082-7; and Curtis A. Machida. Methods in Molecular Medicine TM. Viral Vectors for Gene Therapy Methods and Protocols. 10.1385 / 1-59259-304-6:201 (Copyright) Humana Press Inc. 2003. Chapter 10. Targeted Integration by Adeno-Associated Virus. Matthew D. Weitzman, Samuel M. Young Jr., Toni Cathomen and Richard Jude Samulski; see products and services available from U.S. Patent Nos. 5,139,941 and 5,962,313, all of which are incorporated herein by reference).

[0264] In some embodiments, the rAAV vector of the present disclosure comprises one or more regulatory elements (e.g., promoter, transcription terminator, and / or other regulatory elements) for controlling the expression of the heterologous nucleic acid region. In some embodiments, the first and / or second nucleotide sequence is operably linked to one or more (e.g., 1, 2, 3, 4, 5, or more) transcription terminators. Non-limiting examples of transcription terminators that can be used according to the present disclosure include the transcription terminators (or polyadenylation signals) of bovine growth hormone gene (bGH), human growth hormone gene (hGH), SV40, CW3, Φ, or combinations thereof. In an exemplary embodiment, the transcription terminator is an SV40 polyadenylation signal. In an exemplary embodiment, the transcription terminator does not contain a post-transcriptional response element such as a WPRE element.

[0265] In some aspects, provided herein are methods of making (or manufacturing or packaging) any of the disclosed rAAV particles. rAAV particles can be produced according to any method known in the art. Packaging methods are known in the art and reagents are commercially available (see, e.g., Zolotukhin et al. Production and purification of serotype 1, 2, and 5 recombinant adeno-associated viral vectors. Methods 28 (2002) 158-167; and U.S. Patent Application Publication Nos. 2007-0015238 and 2012-0322861, which are incorporated herein by reference; and plasmids and kits available from ATCC and Cell Biolabs, Inc.). For example, a plasmid containing the gene of interest can be combined with one or more helper plasmids containing, for example, the rep genes (encoding, for example, Rep78, Rep68, Rep52, and Rep40) and the cap genes (encoding VP1, VP2, and VP3, including a VP2 region modified as described herein) and transfected into a recombinant cell such that rAAV particles can be packaged and subsequently purified.

[0266] Packaging cells (or host cells) are typically used to form viral particles that can infect host cells. Such cells include 293 cells, which package adenovirus, and ψ2 or PA317 cells, which package retrovirus. Viral vectors used in gene therapy are usually produced by producing cell lines that package nucleic acid vectors into viral particles. The vector typically contains the minimum viral sequences required for packaging and subsequent integration into the host, with other viral sequences being replaced by an expression cassette for the polynucleotide(s) to be expressed. Missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically possess only the ITR sequences from the AAV genome that are required for packaging and integration into the host genome. Viral DNA is packaged in a cell line that contains a helper plasmid that encodes the other AAV genes, namely rep and cap, but lacks the ITR sequences. The cell line can also be infected with adenovirus as a helper. Helper virus promotes the replication of AAV vector and the expression of AAV genes from helper plasmid.Helper plasmid is not packaged in significant amounts due to lack of ITR sequence.Contamination by adenovirus can be reduced by heat treatment, for example, to which adenovirus is more sensitive than AAV.Additional methods for the delivery of nucleic acid to cells are known to those skilled in the art, for example, those disclosed in US Patent Application Publication No. 2003 / 0087817 published May 8, 2003, International Publication No. WO 2016 / 205764 published December 22, 2016, and International Publication No. WO 2018 / 071868 published April 19, 2018.

[0267] In various embodiments, the base editor construct can be engineered for delivery in one or more rAAV vectors. The rAAV involved in any of the methods and compositions provided herein can be of any serotype, including any derivative or pseudotype (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 2 / 1, 2 / 5, 2 / 8, 2 / 9, 3 / 1, 3 / 5, 3 / 8, or 3 / 9). The rAAV can contain the gene payload to be delivered to the cell (i.e., a recombinant nucleic acid vector expressing a transgene of interest, such as all of the base editors carried into the cell by the rAAV). The rAAV can be chimeric.

[0268] As used herein, the serotype of a rAAV refers to the serotype of the capsid protein of the recombinant virus. Non-limiting examples of derivatives and pseudotypes include rAAV2 / 1, rAAV2 / 5, rAAV2 / 8, rAAV2 / 9, AAV2-AAV3 hybrid, AAVrh.10, AAVrh.74, AAVhu.14, AAV3a / 3b, AAVrh32.33, AAV-HSC15, AAV-HSC17, AAVhu.37, AAVrh.8, CHt-P6, AAV2.5, AAV6.2, AAV2i8, AAV-HSC15 / 17, AAVM41, AAV9.45, AAV6(Y445F / Y731F), AAV2.5T, AAV-HAE1 / 2, AAV clone 32 / 83, AAVShH10, AAV2 (Y->F), AAV8 (Y733F), AAV2.15, AAV2.4, AAVM41, and AAVr3.45. A non-limiting example of a derivative and pseudotype having a chimeric VP1 protein is rAAV2 / 5-1VP1u, which has the genome of AAV2, the capsid backbone of AAV5, and the VP1u of AAV1. Other non-limiting examples of derivatives and pseudotypes having a chimeric VP1 protein are rAAV2 / 5-8VP1u, rAAV2 / 9-1VP1u, and rAAV2 / 9-8VP1u. In some embodiments, the capsid of the disclosed rAAV particles is AAV8 (serotype 8). In some embodiments, the capsid of the disclosed rAAV particles is AAV8 (serotype 9). In some embodiments, the capsid is serotype 2, 6, PHP.B, or PHP.eB.

[0269] Compositions comprising any of the disclosed rAAV particles are provided herein. Exemplary compositions contain any of the disclosed rAAV8 particles, rAAV9 particles, rAAV2 particles, rAAV6 particles, rAAVPHP.B particles, and rAAVPHP.eB particles.

[0270] In some embodiments, methods of administering compositions of rAAV8 or rAAV9 particles by intravenous administration are disclosed. In other embodiments, tissues that are not sufficiently transduced by intravenous AAV9 injection can be transduced by other existing AAV variants, such as AAV4 transduction of the lung, or by a different delivery route, such as AAV9 transduction of kidney cells by retrograde urinary injection.

[0271] AAV derivatives / pseudotypes and methods for producing such derivatives / pseudotypes are known in the art (see, for example, Mol. Ther. 2012 Apr;20(4):699-708. doi: 10.1038 / mt.2011.287. Epub 2012 Jan 24. The AAV vector toolkit: poised at the clinical crossroads. Asokan A1, Schaffer DV, Samulski RJ.). Methods for producing and using pseudotyped rAAV vectors are known in the art (see, e.g., Duan et al., J. Virol., 75:7662-7671, 2001; Halbert et al., J. Virol., 74:1524-1532, 2000; Zolotukhin et al., Methods, 28:158-167, 2002; and Auricchio et al., Hum. Molec. Genet., 10:3075-3081, 2001).

[0272] In some aspects, the present disclosure provides a composition that contains a plurality of any of the disclosed rAAV particles.In some aspects, the present disclosure provides a host cell that contains a plurality of any of the disclosed rAAV particles.In some embodiments, the host cell is a mammalian cell, such as a human cell.In other embodiments, the host cell is a yeast cell, a plant cell, or a bacterial cell.

[0273] The disclosed rAAV particles and the compositions comprising rAAV particles and methods for delivering any of the host cells to target cells or target tissues are known in the art.In some embodiments, any of the disclosed rAAV particles, host cells, or compositions are delivered to a subject, such as a mammalian subject.In some embodiments, the rAAV particles are delivered to a human subject.

[0274] In some embodiments, the disclosed rAAV particles and compositions are administered to a subject in a single injection, such as a single systemic injection. In some embodiments, the disclosed rAAV particles and compositions are administered to a subject in multiple injections. Although rAAV particles are known to transduce target tissues within a few days, typically 3-4 weeks are allowed to complete transduction, genome integration, and clearance from the cells. Thus, in some aspects, any of the disclosed rAAV particles or compositions are administered to a subject for a period of 3 weeks. In some aspects, any of the disclosed rAAV particles or compositions are administered to a subject for a period of between 3 and 4 weeks.

[0275] In some embodiments, any of the disclosed rAAV particles or compositions comprises about 10 15 , about 10 14 , about 10 13 , about 10 12 , about 10 11 , or about 10 11 In some embodiments, the rAAV particles are administered to a subject or target tissue in a therapeutically effective amount of vector genome (vg) of less than 10 per kg. 15 and 10 14 Between 14 and 10 13 Between 13 and 10 1 Between 12 and 10 11 Between or 10 12 and 10 11 In some embodiments, the rAAV particles are administered in an amount between 10 vg per kg. 14 and 1011 In some embodiments, any of the disclosed rAAV particles or compositions are administered to a target tissue of a subject at a lower dose than is conventional for dual AAV particle delivery, such as that described in WO 2020 / 236982 published November 26, 2020, and Levy, JM, et al. Nat Biomed Eng 4, 97-110 (2020).

[0276] In some aspects, the disclosure provides single AAV vector delivery of base editors to target tissues such as liver, neuronal, cardiac, muscular, or ocular tissue. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the liver (liver) tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the cardiac tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the neuronal tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the muscle or neuromuscular tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the ocular tissue of a subject.

[0277] In some embodiments, the disclosed rAAV particles provide transduction of a target tissue to achieve expression and translation of a payload or transgene, e.g., a base editor according to the present disclosure, for a sufficient duration to incorporate a desired mutation on the genome of the target cell. In some embodiments, the desired mutation is an A to G mutation. In some embodiments, the desired mutation is a C to T mutation. In some embodiments, the disclosed rAAV particles provide sufficient expression and translation of a base editor transgene for a sufficient duration to incorporate a desired (on-target) mutation on the genome with a tolerable degree of off-target effects, e.g., bystander editing. In some embodiments, the disclosed rAAV particles provide sufficient expression and translation of a base editor transgene for a sufficient duration to incorporate a desired mutation on the genome without significant off-target editing. In some embodiments, the disclosed rAAV particles provide sufficient expression and translation of a base editor transgene for a sufficient duration to incorporate a desired mutation on the genome without significant bystander editing.

[0278] Suitable routes of administration of the disclosed compositions of rAAV particles include, without limitation, topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, intradental, intracochlear, transtympanic, intravisceral, epidural, intrathecal, intramuscular, intravenous, systemic, intravascular, intraosseous, periocular, intratumoral, intracerebral, parenteral, and intraventricular administration. In some embodiments, the route of administration is systemic (intravenous). In some embodiments, the pharmaceutical compositions described herein are administered locally to the site of disease.

[0279] In some aspects, a pharmaceutical composition is provided herein that includes any of the disclosed compositions and a pharma- ceutically acceptable carrier. In some embodiments, the pharmaceutical composition is formulated for delivery to a subject, for example for base editing on a genome. In some embodiments, the disclosed compositions are formulated according to routine procedures as compositions suitable for intravenous or subcutaneous administration to a subject, for example a human. In some embodiments, the compositions provided herein are formulated for delivery to a subject, for example a human subject, to achieve targeted genome modification in the subject. In some embodiments, cells are obtained from a subject and contacted with any of the pharmaceutical compositions provided herein. In some embodiments, the cells removed from a subject and contacted with the pharmaceutical composition ex vivo are reintroduced into the subject, optionally after desired genome modification is achieved or detected in the cells. Subjects to which administration of the pharmaceutical compositions is contemplated include, but are not limited to, humans and / or other primates; mammals, livestock, pets, and mammals of commercial interest, such as cows, pigs, horses, sheep, cats, dogs, mice, and / or rats; and / or birds, including birds of commercial interest, such as chickens, ducks, geese, and / or turkeys.

[0280] The formulations of the pharmaceutical compositions described herein can be prepared by any method known or hereafter developed in the art of pharmacology. In general, such preparation methods include the steps of combining the active ingredient(s) with an excipient and / or one or more other accessory ingredients, and then shaping and / or packaging the product into the desired single or multi-dose unit, if necessary and / or desired. Pharmaceutical formulations may additionally include pharmaceutically acceptable excipients. As used herein, this includes any and all solvents, dispersion media, diluents, or other liquid carriers, dispersion or suspension aids, surfactants, isotonicity agents, thickening or emulsifying agents, preservatives, solid binders, lubricants, and the like, suitable for the specific dosage form desired.

[0281] In some embodiments, the pharmaceutical composition of rAAV particles for administration by injection is a solution in a sterile isotonic aqueous buffer. Where necessary, the pharmaceutical may also include a solubilizing agent and a local anesthetic such as lignocaine to ease pain at the injection site. Generally, the ingredients are either supplied separately or mixed together in a unit dosage form, for example, as a dry lyophilized powder or water-free concentrate in an airtight container such as an ampoule or sachet indicating the quantity of active agent. Where the pharmaceutical is to be administered by infusion, it may be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. Where the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline may be provided so that the ingredients can be mixed prior to administration.

[0282] The term "pharmaceutically acceptable carrier" as used herein means a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc, magnesium, calcium, or zinc stearate, or stearic acid), or solvent encapsulating material, involved in carrying or transporting a compound from one site (e.g., a delivery site) in the body to another site (e.g., an organ, tissue, or part of the body). A pharmaceutically acceptable carrier is "acceptable" in the sense of being compatible with other ingredients of the formulation and not toxic to the tissue of the subject (e.g., physiologically compatible, sterile, physiological pH, etc.). Some examples of materials which may act as pharma- ceutically acceptable carriers include: (1) sugars, such as lactose, glucose, and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose and its derivatives, such as sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants, such as magnesium stearate, sodium lauryl sulfate, and talc; (8) excipients, such as cocoa butter and suppository wax; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (10) glycols, such as propylene glycol; (11) cellulose acetate, cellulose acetate, cellulose ether, cellulose acetate ... (1) polyols, such as glycerin, sorbitol, mannitol, and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffers, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer's solution; (19) ethyl alcohol; (20) pH buffer solutions; (21) polyesters, polycarbonates, and / or polyanhydrides; (22) bulking agents, such as polypeptides and amino acids; (23) serum components, such as serum albumin, HDL, and LDL; (22) C2-C12 alcohols, such as ethanol; and (23) other non-toxic compatible substances employed in pharmaceutical formulations.Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, perfuming agents, preservatives, and antioxidants may also be present in the formulation. Terms such as "excipient," "carrier," or "pharmaceutical acceptable carrier," and the like, are used interchangeably herein.

[0283] The rAAV particle pharmaceutical compositions of the present disclosure can be administered or packaged, for example, as unit doses. The term "unit dose" when used with reference to pharmaceutical compositions of the present disclosure refers to physically discrete units suitable as unitary dosages for subjects, each unit containing a predetermined quantity of active material calculated to produce a desired therapeutic effect, in association with the required diluent, i.e., carrier or base.

[0284] rAAV vector sequence In some aspects, the disclosure provides rAAV vector nucleic acid sequences as presented below. In some embodiments, the disclosed vectors comprise a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% sequence identity to any of SEQ ID NOs: 100-102. In certain embodiments, the disclosed vectors are any of the nucleic acid sequences of SEQ ID NOs: 100-102. Each sequence component is indicated along with the length of the ITR to ITR.

[0285] In some embodiments, any of the vectors described herein may contain a nucleic acid sequence having 1-5, 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, 45-50, or more than 50 nucleotides that differ relative to any of SEQ ID NOs: 100-102. These differences may include inserted, deleted, or substituted nucleotides relative to any of SEQ ID NOs: 100-102. In some embodiments, the disclosed vectors contain a stretch of about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 300, about 400, about 500, or more than 500 contiguous nucleotides in common with any of SEQ ID NOs: 100-102.

[0286] [ka] [ka] [ka]

[0287] [ka] [ka]

[0288] [ka] [ka] [ka]

[0289] Methods for editing a target nucleic acid molecule In some aspects, methods are provided herein for contacting any of the disclosed AAV-encoded base editors with a nucleic acid molecule, e.g., a nucleic acid molecule (e.g., DNA) that includes a target sequence. In some embodiments, the nucleic acid molecule includes DNA, e.g., single-stranded DNA or double-stranded DNA. The target sequence of the nucleic acid molecule may include a target nucleic acid base pair that contains adenine (A). The target sequence of the nucleic acid molecule may include a target nucleic acid base pair that contains cytosine (C). The target sequence may be a genomic sequence, e.g., a human genomic sequence. The target sequence may include a sequence associated with a disease or disorder, e.g., a target sequence with a point mutation. The target sequence with a point mutation may be associated with a cardiovascular disease.

[0290] In some embodiments, the target nucleotide sequence is a DNA sequence on a genome, for example a eukaryotic genome. In certain embodiments, the target nucleotide sequence is on a mammalian (for example human) genome. In certain embodiments, the target nucleotide sequence is on a human genome. In other embodiments, the target nucleotide sequence is on a rodent genome, such as a mouse or a rat. In other embodiments, the target nucleotide sequence is on a domesticated genome, such as a horse, a cat, a dog, or a rabbit. In some embodiments, the target nucleotide sequence is on a research animal genome. In some embodiments, the target nucleotide sequence is on a genome of a genetically engineered non-human subject. In some embodiments, the target nucleotide sequence is on a plant genome. In some embodiments, the target nucleotide sequence is on a microorganism genome, such as a bacterium.

[0291] In some embodiments, the disclosed AAV-encoded base editors exhibit low off-target effects, e.g., low off-target editing frequency. In some embodiments, the disclosed AAV-encoded base editors exhibit low off-target editing frequency while exhibiting high on-target editing efficiency. In some embodiments, the use of TadA-8e deaminase or TadA-8e(V106W) deaminase in any of the disclosed AAV-encoded adenine base editors can exhibit an off-target editing frequency of 0.32% or less while maintaining an on-target editing efficiency of about 80% or more. See International Publication WO 2021 / 158921, published August 12, 2021.

[0292] The disclosed base editors can provide (or produce) an on-target editing efficiency of greater than 50% or greater than 60% (e.g., greater than 70%, greater than 75%, greater than 80%, or greater than 85%) at the target nucleic acid base pair of one or more base editors under evaluation. Any of the disclosed editing methods can produce an on-target editing efficiency of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, or at least about 85%.

[0293] In some embodiments, the editing method comprising the disclosed BE and contacting a cell comprising a target DNA sequence with any of the disclosed BEs results in an actual or average off-target DNA editing frequency of about 2.0% or less, 1.75% or less, 1.5% or less, 1.2% or less, 1% or less, 0.9% or less, 0.8% or less, 0.75% or less, 0.7% or less, 0.65% or less, or 0.6% or less. These off-target editing frequencies can be obtained in sequences with any level of sequence identity to the target sequence. The modifier "average" when used herein to refer to off-target DNA editing frequency refers to the average value from all editing events detected (for example, detected by high-throughput sequencing) at sites other than a given target nucleic acid base pair.

[0294] In various embodiments, the disclosed editing method results in an on-target DNA base editing efficiency of at least about 35%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 98%, or 99% at the target nucleic acid base pair. The contacting step can result in a DNA base editing efficiency of at least about 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, or 75%. In particular, the contacting step can result in an on-target base editing efficiency of greater than 75%. The contacting step can result in a DNA base editing efficiency between 60 and 85%. In certain embodiments, a base editing efficiency of greater than 85% can be achieved.

[0295] These editing efficiencies can be achieved following administration of any of the disclosed AAV-encoded base editors in any target tissue, such as liver tissue, heart tissue, muscle tissue, neuronal tissue, or eye tissue.

[0296] Administration of any of the disclosed rAAV vectors or particles to cardiac or muscle tissue (e.g., skeletal muscle tissue) can achieve editing efficiencies of at least about 20%, at least about 22%, at least about 24%, at least about 27%, at least about 30%, at least about 33%, or at least about 36%. These editing efficiencies represent a 2- to 2.5-fold increase relative to the editing efficiencies reported for dual AAV vectors in cardiac and muscle tissue. Administration of any of the disclosed rAAV vectors or particles to cardiac or muscle tissue (e.g., cardiac tissue) can achieve indel rates of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less in cardiac or muscle cells.

[0297] In various embodiments, the disclosed editing methods result in a ratio of on-target:off-target editing of about 25:1, 50:1, 65:1, 75:1, 80:1, 85:1, 90:1, 95:1, 100:1, 110:1, 125:1, or more than 125:1. In various embodiments, the disclosed editing methods result in a ratio of on-target:off-target editing of about 150:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 1100:1, 1200:1, 1250:1, 1275:1, 1300:1, 1325:1, 1350:1, 1400:1, 1500:1, or more than 1500:1. As used herein, the ratio of on-target:off-target editing is equivalent to the ratio of sequencing reads that reflect on-target deamination relative to the deamination of known or predicted off-target sites or candidate off-target sites.Candidate off-target sites can be identified, and thus the ratio of on-target:off-target editing can be measured using experimental assays or computational algorithms (e.g., Cas-OFFinder).For example, candidate off-target sites can be identified using experimental assays such as EndoV-Seq, GUIDE-Seq, or CIRCLE-Seq.In some embodiments, the ratio of on-target:off-target editing relies on the use of EndoV-Seq.

[0298] In some embodiments, the disclosed editing methods result in a minimal degree of bystander editing (i.e., synonymous off-target point mutations at nucleobases close to the target base that do not change the outcome of the intended editing method), which the disclosed base editors generate. In some embodiments, the disclosed editing methods result in less than 10, less than 9, less than 8, less than 7, less than 6, less than 5, less than 4, less than 3, less than 2, less than 1, or less than 0 non-silent bystander editing. For example, an editing method using the disclosed single AAV-encoded SaKKH-ABE8e editor in liver tissue may result in only a small number (i.e., minimal) of non-silent bystander editing.

[0299] Some aspects of the present disclosure are based on the recognition that any of the adenine base editors provided herein can modify specific DNA bases without generating a significant proportion of indels. As used herein, "indel" refers to the insertion or deletion of a nucleotide base in a DNA substrate. Such insertion or deletion can lead to frameshift mutations in the coding region of a gene. In some embodiments, it is desirable to generate an adenine base editor that efficiently modifies (e.g., mutates or deaminates) specific nucleotides in DNA without generating a large number of insertions or deletions (i.e., indels) on nucleic acid (while at the same time having a lower RNA editing effect than existing adenine base editors).

[0300] In some embodiments, the disclosed editing methods using the disclosed BEs can result in less than 20%, 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1.5%, 1%, 0.5%, 0.2%, or 0.1% indel formation in a nucleic acid (e.g., DNA) comprising a target sequence. In some embodiments, the disclosed editing methods result in an indel rate of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less. See Figures 11A, 11B, and 15.

[0301] In some embodiments, the disclosed editing methods result in a base editing:indel ratio of at least about 5:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, or greater than about 15:1.

[0302] Some aspects of the disclosure are based on the recognition that any of the base editors provided herein can efficiently generate intended mutations, such as point mutations on DNA (e.g., DNA in a subject's genome) without generating a significant number of unintended mutations, e.g., unintended point mutations. In some embodiments, the intended mutation is a mutation generated by a specific base editor bound to a gRNA specifically designed to generate the intended mutation (e.g., deamination). In some embodiments, the intended mutation is a mutation associated with a disease or disorder, such as sickle cell disease. In some embodiments, the intended mutation is an adenine (A) to guanine (G) point mutation associated with a disease or disorder. In some embodiments, the intended mutation is a thymine (T) to cytosine (C) point mutation associated with a disease or disorder. In some embodiments, the intended mutation is an adenine (A) to guanine (G) point mutation in the coding region of a gene. In some embodiments, the intended mutation is a thymine (T) to cytosine (C) point mutation in the coding region of a gene.

[0303] In some embodiments, the intended mutation is a deamination that generates a stop codon, such as a premature stop codon in the coding region of a gene. In some embodiments, the intended mutation is a mutation that eliminates the stop codon. In some embodiments, the intended mutation eliminates the stop codon, including the nucleic acid sequence 5'-TAG-3', 5'-TAA-3', or 5'-TGA-3'.

[0304] In some embodiments, the intended mutation is a deamination that modulates the regulatory sequence of a gene (e.g., a gene promoter or a gene repressor). In some embodiments, the intended mutation is a deamination introduced into a gene promoter. In certain embodiments, the deamination introduced into a gene promoter leads to a decrease in the transcription of a gene operably linked to the gene promoter. In other embodiments, the deamination leads to an increase in the transcription of a gene operably linked to the gene promoter.

[0305] In some embodiments, the intended mutation is a deamination that modulates the gene sequence or splicing of the gene. Thus, in some embodiments, the intended deamination results in the introduction of a splice site on the gene. In other embodiments, the intended deamination results in the removal of a splice site. In some embodiments, the intended deamination results in the introduction of a stop codon on the gene. In other embodiments, the intended deamination results in the removal of a stop codon.

[0306] In some embodiments, any of the base editors provided herein can generate a ratio of intended to unintended mutations (e.g., intended point mutations:unintended point mutations) that is greater than 1:1. In some embodiments, the base editors provided herein are capable of generating a ratio of intended mutations to unintended mutations (e.g., intended point mutations:unintended point mutations) that is at least 1.5:1, at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 150:1, at least 200:1, at least 250:1, at least 500:1, or at least 1000:1, or more. It should be understood that the properties of the base editors described in this section and in the following sections of this disclosure can apply to any of the base editors or methods of using the base editors provided herein.

[0307] The base editors encoded by the AAVs of the present disclosure disclosed herein have reduced and / or low RNA editing effects. In some embodiments, the base editors are evolved or engineered to have reduced RNA editing effects. As used herein, the term "RNA editing effect" refers to the introduction of nucleotide modifications (e.g., deamination) in cellular RNA, e.g., messenger RNA (mRNA). The key goal of DNA base editing efficiency is the modification of a specific nucleotide in DNA (e.g., deamination) without introducing a similar nucleotide modification in RNA. The RNA editing effect is "low" or "reduced" when the detected mutation is introduced into the RNA molecule at a frequency of 0.3% or less.

[0308] The present disclosure further provides methods of administering the disclosed base editors, where the methods result in reduced and / or low RNA editing effects. The present disclosure further provides base editors (such as the disclosed ABEs) that induce (or result, provide, or cause) low and / or undetectable RNA editing effects (see FIG. 17). In some embodiments, the base editor provides an average Adenosine (A) to Inosine (I) (A to I) editing frequency on cellular mRNA transcripts of 0.3% or less. In some embodiments, the base editor provides an average Adenosine (A) to Inosine (I) (A to I) actual and / or consistent editing frequency on RNA of about 0.3% or less. A base editor may provide an actual or average A to I editing frequency on the RNA of about 0.5% or less, 0.4% or less, 0.35% or less, 0.25% or less, 0.2% or less, 0.15% or less, 0.12% or less, 0.1% or less, 0.08% or less, or 0.075% or less.

[0309] Guide sequence (e.g. guide RNA) The present disclosure further provides guide RNA for use according to the disclosed editing method. The present disclosure provides guide RNA designed to recognize a target sequence. Such gRNA can be designed to have a guide sequence (or "spacer") that has complementarity to a protospacer in the target sequence.

[0310] Guide RNAs are also provided for use with one or more of the disclosed base editors in the disclosed methods of editing nucleic acid molecules, by way of example. Such gRNAs can be designed to have a guide sequence that has complementarity to a protospacer in the target sequence to be edited, and to have a backbone sequence that specifically interacts with the napDNAbp domain of any of the disclosed base editors, such as the Cas9 nickase domain of the disclosed base editors. Guide RNAs according to the disclosed editing methods can have complementarity to any of the protospacer sequences listed in Table 1 (SEQ ID NOs: 430-565).

[0311] In various embodiments, the base editor can be complexed, bound, or otherwise associated with one or more guide sequences (e.g., via any type of covalent or non-covalent bond). The guide sequence becomes bound or associated with the base editor, directing its localization to a particular target sequence that has complementarity to the guide sequence or a portion thereof. The specific design aspects of the guide sequence will depend on the nucleotide sequence of the genomic target sequence (i.e., the desired site to be edited) and the type of napDNAbp present on the base editor (e.g., the type of Cas9 protein), among other factors, such as, for example, the location of the PAM sequence, the percent G / C content of the target sequence, the degree of microhomology regions, secondary structure, etc.

[0312] In general, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of napDNAbp (e.g., Cas9 or Cas9 variant) to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more, or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more, when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences. Non-limiting examples of this include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transform (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).

[0313] In some embodiments, the guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more nucleotides in length, or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more nucleotides in length. In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length, or fewer. In some embodiments, the guide RNA is between 15-300, 25-300, 50-300, 25-250, 25-200, 15-200, or 15-100 nucleotides in length. In some embodiments, the guide RNA is between about 25 and 200 nucleotides in length. In some embodiments, the guide RNA is between about 15 and 200, or between 15 and 100 nucleotides in length.

[0314] The ability of a guide sequence to direct sequence-specific binding of a base editor to a target sequence can be assessed by any suitable assay. For example, a base editor component comprising the guide sequence to be tested can be provided to a host cell having a corresponding target sequence, e.g., by transfection with a vector encoding a base editor component disclosed herein, followed by assessment of preferential cleavage within the target sequence. Similarly, cleavage of a target polynucleotide sequence can be assessed in situ by providing the target sequence, a base editor component comprising the guide sequence to be tested, and a control guide sequence that differs from the test guide sequence, and comparing the rate of binding or cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible and will occur to those skilled in the art.

[0315] The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of the cell. Exemplary target sequences include those that are unique on the target genome.

[0316] In some embodiments, the guide sequence is selected to reduce the degree of secondary structure in the guide sequence. The secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. One such algorithm is mFold, described by Zuker & Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example folding algorithm is the online web server RNAfold, developed at the Institute for Theoretical Chemistry, University of Vienna, which uses a centroid structure prediction algorithm (see, for example, AR Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr & GM Church, 2009, Nature Biotechnology 27(12): 1151-62). Additional algorithms can be found in Chuai, G. et al., DeepCRISPR: optimized CRISPR guide RNA design by deep learning, Genome Biol. 19:80 (2018), as well as U.S. Patent Application Publication No. 61 / 836,080 and U.S. Patent No. 8,871,445, issued October 28, 2014, each of which is incorporated herein by reference in its entirety.

[0317] The guide sequence of the gRNA is linked to a tracr mate (also known as a "backbone") sequence, which in turn hybridizes with the tracr sequence. The tracr mate sequence encompasses any sequence that has sufficient complementarity to the tracr sequence to promote one or more of the following: (1) excision of the guide sequence flanked by the tracr mate sequence in cells containing the corresponding tracr sequence; and (2) formation of a complex in the target sequence, where the complex includes the tracr mate sequence hybridized with the tracr sequence. In general, the degree of complementarity is by reference to the optimal alignment of the tracr mate sequence and the tracr sequence along the length of the shorter of the two sequences. The optimal alignment can be determined by any suitable alignment algorithm, and can further take into account secondary structures such as self-complementarity either within the tracr sequence or the tracr mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and the tracr mate sequence along the shorter length of the two when optimally aligned is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more, or greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more. In some embodiments, the tracr sequence is about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 nucleotides or more in length, or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 nucleotides or more in length. In some embodiments, the tracr sequence and the tracr mate sequence are contained within a single transcript such that hybridization between the two produces a transcript with a secondary structure such as a hairpin. The preferred loop-forming sequence for use in hairpin structures is 4 nucleotides in length, and most preferably has the sequence GAAA. However, longer or shorter loop sequences can be used as alternative sequences.The sequence preferably includes a nucleotide triplet (e.g., AAA) and an additional nucleotide (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAG. In some embodiments of the present invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins. In certain embodiments, the transcript has two, three, four, or five hairpins. In further embodiments of the present invention, the transcript has at most five hairpins. In some embodiments, the single transcript further includes a transcription termination sequence; preferably, this is a poly-T sequence, e.g., six T nucleotides.

[0318] In some embodiments, the guide RNA for use according to the disclosed editing method comprises a synthetic single guide RNA (sgRNA) containing modified ribonucleotides. In some embodiments, the guide RNA contains modifications such as 2'-O-methylated nucleotides and phosphorothioate linkages. In some embodiments, the guide RNA contains 2'-O-methyl modifications in the first 3 nucleotides and the last 3 nucleotides, and phosphorothioate linkages between the first 3 and the last 3 nucleotides. Exemplary modified synthetic sgRNAs are disclosed in Hendel A. et al., Nat. Biotechnol. 33, 985-989 (2015), which is incorporated herein by reference.

[0319] In some embodiments, the guide RNA for use according to the editing method of the present disclosure comprises a backbone structure recognized by N. meningitis Cas9 protein or domain, such as Nme2Cas9 domain. The backbone structure (or backbone) recognized by Nme2Cas9 protein can comprise the sequence provided below: 5'-[guide sequence]-gttgtagctccctttctcatttcggaaacgaaatgagaaccgttgctacaataaggccgtctgaaaagatgtgccgcaacgctctgccccttaaagcttctgctttaaggggcatcgttta-3' (SEQ ID NO: 719). This backbone sequence is recognized by NmeCas9, Nme1Cas9, Nme2Cas9, and Nme3Cas9 proteins. Exemplary guide RNAs for editing by base editors containing Nme2Cas9 domain and variants thereof are described in Edraki et al., Molecular Cell 73, 714-726, which is incorporated herein by reference.

[0320] In other embodiments, the guide RNA for use according to the disclosed editing methods comprises a backbone structure recognized by an S. pyogenes Cas9 protein or domain, such as the SpCas9 domain of the disclosed base editor. The backbone structure recognized by the SpCas9 protein can comprise the sequence 5'-[guide sequence]-guuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuu-3' (SEQ ID NO: 339), where the guide sequence comprises a sequence that is complementary to the protospacer of the target sequence. See U.S. Patent Application Publication No. 2015 / 0166981, published June 18, 2015, the disclosure of which is incorporated herein by reference. The guide sequence is typically 20 nucleotides in length.

[0321] In other embodiments, a guide RNA for use according to the disclosed editing methods comprises a backbone structure recognized by a C. jejuni Cas9 protein or domain, such as the CjCas9 domain of a disclosed base editor. The backbone structure recognized by the CjCas9 protein can comprise the sequence 5'-[guide sequence]-gttttagtccctgaaaagggactaaaataaagagtttgcgggactctgcggggttacaatcccctaaaaccgcttttttt-3' (SEQ ID NO: 340), where the guide sequence comprises a sequence that is complementary to the protospacer of the target sequence.

[0322] In other embodiments, the guide RNA for use according to the disclosed editing method comprises a backbone structure recognized by S. aureus Cas9 protein. The backbone structure recognized by SaCas9 protein can comprise the sequence 5'-[guide sequence]-guuuuaguacucuguaaugaaaauuacagaaucuacuaaaacaaggcaaaaugccguguuuaucucgucaacuuguuggcgagauuuuuuu-3' (SEQ ID NO: 78). This is also the backbone structure recognized by SaKKH-Cas9 and the SaCas9 orthologue, SauriCas9.

[0323] Guide RNAs for use in accordance with the disclosed editing methods may include the backbone structures listed in Table 2 (SEQ ID NOs: 566-571).

[0324] Based on this disclosure, the sequence of a suitable guide RNA for targeting the disclosed BE to a specific genomic target site will be clear to one of skill in the art. Such a suitable guide RNA sequence typically includes a guide sequence that is complementary to a nucleic acid sequence within 50 nucleotides upstream or downstream of the target nucleic acid base pair to be edited. Some exemplary guide RNA sequences suitable for targeting any of the provided BEs to a specific target sequence are provided herein. Additional guide sequences are known in the art and can be used with the base editors described herein. Additional exemplary guide sequences are described, for example, in Jinek M., et al., Science 337:816-821(2012);Mali P, Esvelt KM & Church GM (2013) Cas9 as a versatile tool for engineering biology, Nature Methods, 10, 957-963;Li JF et al., (2013) Multiplex and homologous recombination-mediated genome editing in Arabidopsis and Nicotiana benthamiana using guide RNA and Cas9, Nature Biotechnology, 31, 688-691;Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system, Nature Biotechnology 31, 227-229 (2013);Cong L et al., (2013) Multiplex genome engineering using CRIPSR / Cas systems, Science, 339, 819-823;Cho SW et al., (2013) Targeted genome engineering in human cells with the Cas9 RNA-guided endonuclease, Nature Biotechnology, 31, 230-232;Jinek, M. et al., RNA-programmed genome editing in human cells, eLife 2, e00471 (2013);Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013);Briner AE et al., (2014) Guide RNA functional modules direct Cas9 activity and orthogonality, Mol Cell, 56, 333-339, the entire contents of each of which are incorporated herein by reference.

[0325] Methods for generating Cas variants and base editors The invention further relates in various aspects to methods of making the disclosed improved base editors by various engineering modes, including, but not limited to, codon optimization to achieve greater expression levels in a cell, and the use of a nuclear localization sequence (NLS), preferably at least two NLSs, e.g., two bipartite NLSs, to increase localization of the expressed base editor within the cell nucleus.

[0326] Preparation of base editors for increased expression in cells Base editors contemplated herein can include modifications that result in increased expression, for example through codon optimization.

[0327] In some embodiments, the base editor (or a component thereof) is codon-optimized for expression in a specific cell, such as a eukaryotic cell. The eukaryotic cell may be of or derived from a specific organism, such as a mammal, including but not limited to a human, mouse, rat, rabbit, dog, or non-human primate. In general, codon optimization refers to the process of modifying a nucleic acid sequence for enhanced expression in a host cell of interest by maintaining the native amino acid sequence while replacing at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 codons or more, or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 codons or more) of the native sequence with a codon that is frequently or most frequently used by genes of that host cell. Different species show a particular bias towards certain codons for certain amino acids. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of translation of messenger RNA (mRNA). This, in turn, is believed to depend, among other things, on the properties of the codon to be translated and on the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell is generally a reflection of the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example in the "Codon Usage Database," and these tables can be adapted in a number of ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available for codon-optimizing specific sequences for expression in a particular host cell. For example, Gene Forge (Aptagen; Jacobus, PA) is also available.In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more or all codons) on a sequence encoding a CRISPR enzyme correspond to the most frequently used codon for a particular amino acid.

[0328] Nuclear localization sequences and additional base editor components In some embodiments, the base editors provided herein further comprise one or more nuclear targeting sequences, e.g., nuclear localization sequences (NLSs). In some embodiments, the NLS comprises an amino acid sequence that facilitates the import of a protein comprising the NLS into a cell nucleus (e.g., by nuclear transport). In some embodiments, any of the base editors provided herein further comprises one or more nuclear localization sequences (NLSs). In certain embodiments, any of the base editors comprises two NLSs. In some embodiments, one or more of the NLSs is a bipartite NLS ("bpNLS"). In certain embodiments, the disclosed base editors comprise two bipartite NLSs. In some embodiments, the disclosed base editors comprise more than two bipartite NLSs.

[0329] In some embodiments, the NLS is fused to the N-terminus of the base editor. In some embodiments, the NLS is fused to the C-terminus of the base editor. In some embodiments, the NLS is fused to the C-terminus of the napDNAbp. In some embodiments, the NLS is fused to the N-terminus of the adenosine deaminase. In some embodiments, the NLS is fused to the C-terminus of the adenosine deaminase. In some embodiments, the NLS is fused to the base editor via one or more linkers. In some embodiments, the NLS is fused to the base editor without a linker.

[0330] In some embodiments, the NLS comprises the amino acid sequence of any one of the NLS sequences provided or referenced herein. In some embodiments, the NLS comprises the amino acid sequence presented in SEQ ID NO: 408 or SEQ ID NO: 409. Additional nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in PCT / EP2000 / 011690 to Plank et al., the contents of which are incorporated herein by reference for disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 408), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 409), KRTADGSEFESPKKKRKV (SEQ ID NO: 410), or KRTADGSEFEPKKKRKV (SEQ ID NO: 411). In other embodiments, the NLS comprises the amino acid sequence NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 482), PAAKRVKLD (SEQ ID NO: 483), RQRRNELKRSF (SEQ ID NO: 484), or NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 485).

[0331] In some embodiments, the base editor comprises a bpNLS. The bpNLS may comprise an amino acid sequence selected from the group consisting of KRTADGSEFEPKKKRKV (SEQ ID NO: 398), KRPAATKKAGQAKKKK (SEQ ID NO: 344), KKTELQTTNAENKTKKL (SEQ ID NO: 345), KRGINDRNFWRGENGRKTR (SEQ ID NO: 346), and RKSGKIAAIVVKRPRK (SEQ ID NO: 347). In certain embodiments, the bpNLA comprises the amino acid sequence presented in SEQ ID NO: 344 or 398.

[0332] In some embodiments, the base editors provided herein do not include a linker. In some embodiments, a linker is present between one or more of the domains or proteins (e.g., deaminase, napDNAbp, and / or NLS). In some embodiments, the "]-[" used in the general architecture above indicates the presence of an optional linker.

[0333] In some embodiments, the general architecture of an exemplary base editor having a first adenosine deaminase, a second adenosine deaminase, and a napDNAbp domain comprises any one of the following structures, where the NLS is a nuclear localization sequence (e.g., any NLS provided herein), and NH 2 is the N-terminus of the base editor and COOH is the C-terminus of the base editor.

[0334] An exemplary base editor comprising a deaminase, a napDNAbp domain, and an NLS (by way of example, any of the NLSs provided herein) may have the following architecture: NH 2 -[deaminase domain]-[napDNAbp domain]-[NLS]-COOH; NH 2 -[napDNAbp domain]-[deaminase domain]-[NLS]-COOH; NH 2 -[NLS]-[deaminase domain]-[napDNAbp domain]-COOH; or NH 2 -[NLS]-[napDNAbp domain]-[deaminase domain]-COOH.

[0335] In certain embodiments, the disclosed base editors comprise the following ABE architecture, where TadA-8e is an adenosine deaminase domain: NH 2 -[bpNLS]-[TadA-8e]-[napDNAbp domain]-[bpNLS]-COOH; NH 2 -[bpNLS]-[napDNAbp domain]-[TadA-8e]-[bpNLS]-COOH; NH 2 -[bpNLS]-[TadA-8e]-[napDNAbp domain]-[bpNLS]-COOH; or NH 2 -[bpNLS]-[napDNAbp domain]-[TadA-8e]-[bpNLS]-COOH.

[0336] An exemplary base editor comprising a cytidine deaminase, a napDNAbp domain, a UGI domain, and an NLS (e.g., any NLS provided herein) may have the following architecture: NH 2 -[napDNAbp domain]-[cytidine deaminase domain]-[UGI domain]-[bpNLS]-COOH; or NH 2 -[bpNLS]-[napDNAbp domain]-[cytidine deaminase domain]-[UGI domain]-COOH;NH 2 -[cytidine deaminase domain]-[napDNAbp domain]-[UGI domain]-[upNLS]-COOH; or NH 2 -[bpNLS]-[cytidine deaminase domain]-[napDNAbp domain]-[UGI domain]-COOH.

[0337] A representative nuclear localization signal is a peptide sequence that directs a protein to the nucleus of the cell in which the sequence is expressed. Nuclear localization signals are predominantly basic and can be located at almost any position on the amino acid sequence of a protein, generally comprising a short sequence of 4 amino acids (Autieri & Agrawal, (1998) J. Biol. Chem. 273: 14731-37, incorporated herein by reference) to 8 amino acids, and typically rich in lysine and arginine residues (Magin et al., (2000) Virology 274: 11-16, incorporated herein by reference). Nuclear localization signals often contain proline residues. A variety of nuclear localization signals have been identified and used to achieve the transport of biological molecules from the cytoplasm to the nucleus of a cell. See, e.g., Tinland et al., (1992) Proc. Natl. Acad. Sci. USA 89:7442-46; Moede et al., (1999) FEBS Lett. 461:229-34, which are incorporated herein by reference. It is currently believed that transport involves nuclear pore proteins.

[0338] Most NLSs can be classified into three general groups: (i) monoglottal NLSs, exemplified by the SV40 large T antigen NLS (PKKKRKV (SEQ ID NO: 408)); (ii) biglottal motifs consisting of two basic domains separated by a variable number of spacer amino acids, exemplified by the Xenopus nucleoplasmin NLS (KRXXXXXXXXXXKKKL (SEQ ID NO: 486)); and (iii) non-classical sequences, such as M9 of the hnRNP A1 protein, influenza virus nucleoprotein NLS, and yeast Gal4 protein NLS (Dingwall and Laskey, Trends Biochem Sci. 1991 Dec;16(12):478-81).

[0339] Nuclear localization signals appear at various points on the amino acid sequence of a protein. NLSs have been identified at the N-terminus, C-terminus, and central region of a protein. Thus, the present specification provides base editors that can be modified by one or more NLSs at the C-terminus, N-terminus, and internal regions of the base editor. Residues of longer sequences that do not function as component NLS residues should be selected so that they do not interfere, for example rigidly or sterically, with the nuclear localization signal itself. Thus, there is no strict limit to the composition of sequences that include NLS, but in reality, such sequences may be functionally limited in length and composition.

[0340] The present disclosure contemplates any suitable means for modifying a fusion protein (or base editor) to include one or more NLSs. In one aspect, a base editor can be engineered to express a fusion protein that is translationally fused at its N-terminus or its C-terminus (or both) to one or more NLSs, i.e., to form a fusion protein-NLS fusion construct. In other embodiments, a nucleotide sequence encoding a fusion protein can be engineered to incorporate a reading frame encoding one or more NLSs into an internal region of the encoded fusion protein. In addition, the NLS can include various amino acid linker or spacer regions encoded between the fusion protein and the N-terminal, C-terminal, or internally attached NLS amino acid sequence. Thus, the present disclosure also enables nucleotide constructs, vectors, and host cells for expressing fusion proteins and base editors that include one or more NLSs.

[0341] The base editors described herein may also include a nuclear localization signal that is linked to the fusion protein through one or more linkers, e.g., polymer, amino acid, polysaccharide, chemical, or nucleic acid linker elements. In certain embodiments, the NLS is linked to the fusion protein using the XTEN linker set forth in SEQ ID NO: 412. Linkers within the contemplated scope of the present disclosure are not intended to have any limitation and may be any suitable type of molecule (e.g., polymer, amino acid, polysaccharide, nucleic acid, lipid, or any synthetic chemical linker domain) and may be tethered to the fusion protein by any suitable strategy that achieves the formation of a bond (e.g., covalent linkage, hydrogen bond) between the fusion protein and one or more NLSs.

[0342] The base editors described herein can also include one or more additional elements, hi certain embodiments, the additional elements can include effectors of base repair, such as inhibitors of base repair.

[0343] In some embodiments, the base editors described herein can include one or more heterologous protein domains (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, in addition to the base editor component). The base editor can include any additional protein sequences, and optionally a linker sequence between any two domains. Other exemplary features that can be present are localization sequences, e.g., cytoplasmic localization sequences, export sequences, e.g., nuclear export sequences, or other localization sequences, and sequence tags.

[0344] Examples of heterologous protein domains that can be fused to a base editor or a component thereof (e.g., napDNAbp domain, nucleotide-modifying domain, or NLS domain) include, without limitation, epitope tags and reporter gene sequences. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, autofluorescent proteins, including glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and blue fluorescent protein (BFP). Base editors can be fused to genetic sequences encoding proteins or fragments of proteins that bind to DNA molecules or other cellular molecules, including, but not limited to, maltose binding protein (MBP), S-tags, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that can form part of a base editor are described in U.S. Patent Application Publication No. 2011 / 0059502, published March 10, 2011, which is incorporated herein by reference in its entirety.

[0345] In some aspects of the present disclosure, reporter genes, including but not limited to glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent protein, including blue fluorescent protein (BFP), can be introduced into cells to code gene products that act as markers for measuring the modulation or alteration of the expression of gene products.In some specific embodiments of the present disclosure, the gene product is luciferase.In further embodiments of the present disclosure, the expression of gene products is reduced.

[0346] Other exemplary features that may be present are tags that are useful for solubilizing, purifying, or detecting the base editor. Suitable protein tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tag, myc tag, calmodulin tag, FLAG tag, hemagglutinin (HA) tag, bgh-polyA tag, polyhistidine tag, also referred to as histidine tag or His tag, maltose binding protein (MBP) tag, nus tag, glutathione-S-transferase (GST) tag, green fluorescent protein (GFP) tag, thioredoxin tag, S tag, Softag (e.g., Softag 1, Softag 3), strep tag, biotin ligase tag, FlAsH tag, V5 tag, and SBP tag. Additional suitable sequences will be apparent to those skilled in the art. In some embodiments, the base editor comprises one or more His tags.

[0347] Linker In certain embodiments, a linker may be used to link either the peptide or peptide domain or domains of the base editor (for example, a napDNAbp domain is covalently linked to an adenosine deaminase domain, which is covalently linked to an NLS domain). The base editors described herein may comprise a linker that is 32 amino acids in length.

[0348] The linker can be as simple as a covalent bond, or it can be a polymeric linker many atoms in length. In certain embodiments, the linker is a polypeptide or based on amino acids. In other embodiments, the linker is not peptide-like. In certain embodiments, the linker is a covalent bond (e.g., carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In certain embodiments, the linker is a carbon-nitrogen bond of an amide linkage. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker. In certain embodiments, the linker is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbon ring moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In other embodiments, the linker comprises an amino acid. In certain embodiments, the linker comprises a peptide. In certain embodiments, the linker comprises an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker may include a functionalized moiety to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0349] In some embodiments, the linker is 5-100 amino acids in length, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-110, 110-120, 120-130, 130-140, 140-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the linker is 32 amino acids in length. In exemplary embodiments, the linker comprises the 32 amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 412), also known as the XTEN linker. In some embodiments, the linker comprises the 9 amino acid sequence SGGSGGSGGS (SEQ ID NO: 413). In some embodiments, the linker comprises the 4 amino acid sequence SGGS (SEQ ID NO: 414).

[0350] In some embodiments, the linker comprises the amino acid sequence (GGGGS) n (SEQ ID NO: 415), (G) n (SEQ ID NO:416), (EAAAK) n (SEQ ID NO: 417), (GGS) n (SEQ ID NO: 418), (SGGS) n (SEQ ID NO: 419), (XP) n (SEQ ID NO:420), or any combination thereof, where n is independently an integer between 1 and 30, and where X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS): n (SEQ ID NO:421), where n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO:422).

[0351] In some embodiments, the linker comprises SGSETPGTSESATPES (SEQ ID NO: 422) and SGGS (SEQ ID NO: 414). In some embodiments, the linker comprises SGGSSGSETPGTSESATPESSGGS (SEQ ID NO: 423). In some embodiments, the linker comprises SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 412). In some embodiments, the linker comprises GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (SEQ ID NO: 424). In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES (SEQ ID NO: 425). In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSS (SEQ ID NO: 426). In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 427). In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS (SEQ ID NO: 428). It should be understood that any of the linkers provided herein can be used to link a first adenosine deaminase and a second adenosine deaminase; an adenosine deaminase domain (including, for example, a first and / or second adenosine deaminase) and napDNAbp; napDNAbp and an NLS; or an adenosine deaminase domain and an NLS.

[0352] In some embodiments, any of the base editors provided herein comprises an adenosine deaminase and a napDNAbp fused to each other via a linker. In some embodiments, any of the base editors provided herein comprises a first adenosine deaminase and a second adenosine deaminase fused to each other via a linker. In some embodiments, any of the base editors provided herein comprises an NLS, which can be fused to an adenosine deaminase (e.g., a first and / or a second adenosine deaminase) and a nucleic acid programmable DNA binding protein (napDNAbp). To achieve the optimal length for deaminase activity for a particular application, various linker lengths and flexibilities can be employed between the adenosine deaminase (e.g., engineered ecTadA) and napDNAbp (e.g., Cas9 domain) and / or between the first and second adenosine deaminase (e.g., highly flexible linkers in the form of SEQ ID NOs: 119, 121-124 (see, e.g., Guilinger JP, Thompson DB, Liu DR. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82; the entire contents of which are incorporated herein by reference) and (XP) n (SEQ ID NO: 420). In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker is (GGS) n(SEQ ID NO: 421) motif, where n is 1, 3, or 7. In some embodiments, the adenosine deaminase and napDNAbp, and / or the first adenosine deaminase and the second adenosine deaminase of any of the base editors provided herein are fused via a linker comprising an amino acid sequence selected from SEQ ID NOs: 119-132. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence (SGGS) 2 -SGSETPGTSESATPES-(SGGS) 2 (SEQ ID NO: 412), which includes (SGGS) 2 -XTEN-(SGGS) 2 (SEQ ID NO: 429). In some embodiments, the linker comprises an amino acid sequence where n is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker is 92 amino acids in length.

[0353] The above description is meant to be non-limiting on creating base editors with increased expression and thereby increasing editing efficiency.

[0354] Methods of Treatment and Use Another aspect of the present disclosure provides a method for delivering a base editor into a cell to form a complete and functional Cas9 protein or nucleobase editor. For example, in some embodiments, a cell is contacted with a composition described herein (e.g., a composition comprising a nucleotide sequence encoding a base editor, or an AAV particle containing a nucleic acid vector comprising such a nucleotide sequence). In some embodiments, the contact results in the delivery of such a nucleotide sequence into the cell, where the N-terminal portion of the Cas9 protein or nucleobase editor and the C-terminal portion of the Cas9 protein or nucleobase editor are expressed in the cell and joined together to form a complete Cas9 protein or a complete nucleobase editor.

[0355] It should be understood that any rAAV particle, nucleic acid molecule, or composition provided herein can be introduced into a cell in any suitable manner, either stably or transiently. In some embodiments, the disclosed proteins can be transfected into a cell. In some embodiments, a cell can be transduced or transfected by a nucleic acid molecule. For example, a cell can be transduced (e.g., by a virus encoding a protein) or transfected (e.g., by a plasmid encoding a protein) by a rAAV particle containing a nucleic acid molecule encoding a protein, or a viral genome encoding one or more nucleic acid molecules. Such transduction can be stable or transient transduction. In some embodiments, a cell expressing or containing a protein can be transduced or transfected by one or more guide RNA sequences, for example, in the delivery of a base editor. In some embodiments, a plasmid expressing a protein can be introduced into a cell through electroporation, transient (e.g., lipofection) and stable genome integration (e.g., nucleofection or piggybac) and viral transduction, or other methods known to those skilled in the art.

[0356] In some aspects, the invention provides methods that include delivering a polynucleotide encoding one or more base editors, one or more transcripts thereof, and / or one or more proteins transcribed therefrom to a cell using a non-viral delivery method. Methods of non-viral delivery of nucleic acids include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid:nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787; and 4,897,355) and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids that are suitable for efficient receptor-recognition lipofection of polynucleotides include those of Feigner, WO 1991 / 17424; WO 1991 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or to target tissues (e.g., in vivo administration).

[0357] In certain embodiments, the compositions provided herein comprise lipids and / or polymers.In certain embodiments, lipids and / or polymers are cationic.The preparation of such lipid particles is well known.See, for example, U.S. Patent No. 4,880,635; U.S. Patent No. 4,906,477; U.S. Patent No. 4,911,928; U.S. Patent No. 4,917,951; U.S. Patent No. 4,920,016; U.S. Patent No. 4,921,757; and U.S. Patent No. 9,737,604.Each of these is incorporated herein by reference.

[0358] In some embodiments, the target nucleotide sequence is a DNA sequence on a genome, e.g., a eukaryotic genome. In certain embodiments, the target nucleotide sequence is on a mammalian (e.g., human) genome.

[0359] The target nucleotide sequence may include a target sequence (e.g., a point mutation) associated with a disease, disorder, or condition. The target sequence may include a T to C (or A to G) point mutation associated with a disease, disorder, or condition, where deamination of the mutant C base results in mismatch repair-mediated correction to a sequence not associated with the disease, disorder, or condition. The target sequence may include a G to A (or C to T) point mutation otherwise associated with a disease, disorder, or condition, where deamination of the mutant A base results in mismatch repair-mediated correction to a sequence not associated with the disease, disorder, or condition. The target sequence may code for a protein, where the point mutation is on a codon, resulting in a change in the amino acid encoded by the mutant codon compared to the wild-type codon. The target sequence may also be at a splice site, where the point mutation results in a change in splicing of the mRNA transcript compared to the wild-type transcript. In addition, the target may be at a non-coding sequence of a gene, such as a promoter, where the point mutation results in increased or decreased expression of the gene.

[0360] Thus, in some aspects, deamination of mutant C results in a change in the amino acid encoded by the mutant codon, which may in some cases result in expression of the wild-type amino acid. In other aspects, deamination of mutant A results in a change in the amino acid encoded by the mutant codon, which may in some cases result in expression of the wild-type amino acid.

[0361] The methods described herein involving contacting a cell with a composition or rAAV particle can occur in vitro, ex vivo, or in vivo. In certain embodiments, the step of contacting the cell occurs in a subject. In certain embodiments, the subject has been diagnosed with a disease, disorder, or condition. In some embodiments, the step of contacting the cell occurs ex vivo or outside of the subject.

[0362] In some embodiments, the methods disclosed herein involve contacting a mammalian cell with a composition or rAAV particle. In certain embodiments, the methods involve contacting a retinal cell, a cortical cell, or a cerebellar cell.

[0363] The compositions described herein can be administered to a subject in need thereof in a therapeutically effective amount to treat and / or prevent a disease or disorder from which the subject suffers. Any disease or disorder that can be treated and / or prevented using CRISPR / Cas9-based genome editing technology can be treated by the base editors described herein. It should be understood that in cases where the nucleotide sequence encoding the base editor does not further encode a gRNA, a separate nucleic acid vector encoding a gRNA can be administered together with the compositions described herein.

[0364] Exemplary suitable diseases, disorders, or conditions include, without limitation, cardiovascular disease, cystic fibrosis, phenylketonuria, epidermolytic hyperkeratosis (EHK), chronic obstructive pulmonary disease (COPD), Charcot-Marie-Tooth disease type 4J, neuroblastoma (NB), von Willebrand disease (vWD), congenital myotonia, hereditary renal amyloidosis, dilated cardiomyopathy, hereditary lymphedema, familial Alzheimer's disease, prion disease, chronic infantile neurological cutaneous and articular syndrome (CINCA), congenital hearing loss, Niemann-Pick disease type C (NPC) disease, and desmin-related myopathy (DRM). In some embodiments, the disease or condition is a cardiovascular disease. In some embodiments, the disease or condition is Niemann-Pick disease type C (NPC) disease.

[0365] In some embodiments, the disease, disorder, or condition is associated with a point mutation that introduces a stop codon, e.g., a premature stop codon, in the coding region of the gene. In some embodiments, the desired base edit removes the stop codon in the coding region of the gene. In some embodiments, the desired base edit disrupts a splice acceptor site or a splice donor site.

[0366] In some embodiments, the desired base edit is associated with disruption of a splice acceptor or splice donor site on the PCSK9 gene or the Angptl3 gene. In certain embodiments, the desired base edit is associated with disruption of a splice acceptor site at W8 on the PCKS9 gene. In some embodiments, the desired base edit is an A to G edit that disrupts the splice acceptor site at residue 8, generating a W8R substitution.

[0367] Additional exemplary diseases, disorders, and conditions include: cystic fibrosis (see, e.g., Schwank et al., Functional repair of CFTR by CRISPR / Cas9 in intestinal stem cell organoids of cystic fibrosis patients. Cell stem cell. 2013; 13: 653-658; and Wu et. al., Correction of a genetic disease in mouse via use of CRISPR-Cas9. Cell stem cell. 2013; 13: 659-662, none of which use a deaminase fusion protein to correct the genetic defect); phenylketonuria - e.g., a phenylalanine to serine mutation (T>C mutation) at position 835 (mouse) or 240 (human) or a homologous residue in the phenylalanine hydroxylase gene - e.g., McDonald et al., Genomics. 1997; 39:402-405; Bernard-Soulier syndrome (BSS) - for example, a phenylalanine to serine mutation at position 55 or a homologous residue of platelet membrane glycoprotein IX, or a cysteine ​​to arginine at residue 24 or a homologous residue (T>C mutation) - for example, see Noris et al., British Journal of Haematology. 1997; 97: 312-320 and Ali et al., Hematol. 2014; 93: 381-384; epidermolytic hyperkeratosis (EHK) - for example, a leucine to proline mutation at position 160 or 161 (counting the initiator methionine) or a homologous residue of keratin 1 (T>C mutation) - for example, see Chipev et al., Cell. 1992; 70: 821-828. See also the UNIPROT database at www[dot]uniprot[dot]org, accession number P04264; Chronic Obstructive Pulmonary Disease (COPD) - e.g., α 1- a leucine to proline mutation (T>C mutation) at position 54 or 55 (counting the initiator methionine) or homologous residues in the processed form of antitrypsin or at residue 78 or homologous residues in the unprocessed form - see for example Poller et al., Genomics. 1993; 17: 740-743. see also UNIPROT database accession number P01011;Charcot-Marie-Tooth disease type 4J - for example, an isoleucine to threonine mutation at position 41 or a homologous residue in FIG4 (T>C mutation) - for example, see Lenk et al., PLoS Genetics. 2011; 7: e1002104;Neuroblastoma (NB) - for example, a leucine to proline mutation at position 197 or a homologous residue in caspase-9 (T>C mutation) - for example, see Kundu et al., 3 Biotech. 2013, 3:225-234;von Willebrand disease (vWD) - for example, a cysteine ​​to arginine mutation at position 509 or a homologous residue in the processed form of von Willebrand factor or at position 1272 or a homologous residue in the unprocessed form of von Willebrand factor (T>C mutation) - for example, see Lavergne et al., J. Med. 2013, 2:1011-1021; et al., Br. J. Haematol. 1992. See also UNIPROT database accession number P04275;82: 66-72;myotonia congenita - for example, a cysteine ​​to arginine mutation (T>C mutation) at position 277 or a homologous residue in the muscle chloride channel gene CLCN1 - for example, see Weinberger et al., The J. of Physiology. 2012;590: 3449-3464;hereditary renal amyloidosis - for example, a stop codon to arginine mutation (T>C mutation) at position 78 or a homologous residue in the processed form of apolipoprotein AII or at position 101 or a homologous residue in the unprocessed form - for example, see Yazaki et al., Kidney Int.2003; 64: 11-16;dilated cardiomyopathy (DCM) - e.g., a tryptophan to arginine mutation at position 148 or a homologous residue (T>C mutation) of the FOXD4 gene, see e.g., Minoretti et. al., Int. J. of Mol. Med. 2007; 19: 369-372;hereditary lymphedema - e.g., a histidine to arginine mutation at position 1035 or a homologous residue (A>G mutation) of the VEGFR3 tyrosine kinase, see e.g., Irrthum et al., Am. J. Hum. Genet. 2000; 67: 295-301;familial Alzheimer's disease - e.g., an isoleucine to valine mutation at position 143 or a homologous residue (A>G mutation) of presenilin 1, see e.g., Gallo et. al., J. Alzheimer's disease. 2011; 25: 425-431; prion diseases - for example, a methionine to valine mutation (A>G mutation) at position 129 or a homologous residue in the prion protein - for example, see Lewis et. al., J. of General Virology. 2006; 87: 2443-2449; chronic infantile neurological cutaneous articular syndrome (CINCA) - for example, a tyrosine to cysteine ​​mutation (A>G mutation) at position 570 or a homologous residue in cryopyrin - for example, see Fujisawa et. al. Blood. 2007; 109: 2903-2911; and desmin-related myopathy (DRM) - for example, an arginine to glycine mutation (A>G mutation) at position 120 or a homologous residue in αβ-crystallin - for example, see Kumar et al., J. Biol. Chem. 1999; 274: 24137-24141. The entire contents of all references and database entries are incorporated herein by reference.

[0368] Treating a disease or disorder includes delaying the onset or progression of the disease, or reducing the severity of the disease. Treating a disease does not necessarily require a curative outcome.

[0369] As used herein, "delaying" the onset of disease means to suspend, hinder, slow down, delay, stabilize, and / or postpone the progression of disease. This delay can be of various lengths of time, depending on the history of the disease and / or the individual being treated. A method of "delaying" or mitigating the onset of disease or delaying the onset of disease is a method that reduces the probability of developing one or more symptoms of disease in a given time frame and / or reduces the severity of symptoms in a given time frame when compared to not using the method. Such comparisons are typically based on clinical trials that use a sufficient number of subjects to provide statistically significant results.

[0370] "Development" or "progression" of a disease refers to the initial manifestation of the disease and / or its subsequent progression. The development of a disease may be detectable and can be evaluated using standard clinical techniques known in the art. However, development also refers to progression that may be undetectable. For the purposes of this disclosure, development or progression refers to the biological process of a condition. "Development" includes occurrence, recurrence, and onset.

[0371] As used herein, "onset" or "occurrence" of a disease includes initial onset and / or recurrence. Depending on the type of disease or site of disease to be treated, conventional methods known to those skilled in the art of medicine can be used to administer the isolated polypeptide or pharmaceutical composition to a subject.

[0372] In some aspects, the disclosure provides the use of any one of the disclosed base editors described herein and a guide RNA targeting the nucleobase editor to a target in the manufacture of a medicament. In some aspects, the use of any one of the nucleobase editors described herein and a guide RNA in the manufacture of a kit for base editing is provided. Wherein, the base editing comprises contacting a nucleic acid molecule with the base editor and the guide RNA under conditions suitable for replacing an adenine (A) of an A:T nucleobase pair on the target with a guanine (G) or a cytosine (C) of a C:T nucleobase pair on the target with a thymine (T). In some embodiments, the contacting step induces separation of double-stranded DNA in the target region. In some embodiments, the contacting step further comprises nicking one strand of the double-stranded DNA, where the one strand constitutes a non-mutated strand.

[0373] In some embodiments of the described uses, the contacting step is performed in vitro. In other embodiments, the contacting step is performed in vivo. In some embodiments, the contacting step is performed in a subject (e.g., a human subject or a non-human animal subject). In some embodiments, the contacting step is performed in a human or non-human animal cell. In some embodiments, the contacting step is performed in a plant cell.

[0374] The present disclosure also provides the use of any one of the nucleobase editors or any one of the complexes of a nucleobase editor and a guide RNA described herein as a medicament. The present disclosure also provides the use of any of the described pharmaceutical compositions or cells comprising any of the nucleobase editors or complexes disclosed herein, and the encoding vectors or rAAV particles as a medicament. In some embodiments, the medicament is for the treatment of cardiovascular disease.

[0375] In some aspects, the disclosure also provides the use of any one of the base editors described herein and a guide RNA that targets the base editor to a target base pair on a nucleic acid molecule in the manufacture of a kit for nucleic acid editing. Wherein, the nucleic acid editing comprises contacting the nucleic acid molecule with the base editor and the guide RNA under conditions suitable for the desired base editing. In some embodiments, the desired base editing is the replacement of adenine (A) of the target A:T base pair with guanine (G). In some embodiments of these uses, the nucleic acid molecule is a double-stranded DNA molecule. In some embodiments, the contacting step induces separation of the double-stranded DNA in the target region. In some embodiments, the contacting step comprises nicking one strand of the double-stranded DNA, where the one strand constitutes the unmutated strand that includes the T of the target A:T nucleic acid base pair.

[0376] In some embodiments of the described uses, the contacting step is performed in vitro. In other embodiments, the contacting step is performed in vivo. In some embodiments, the contacting step is performed in a subject (e.g., a human subject or a non-human animal subject). In some embodiments, the contacting step is performed in a human or non-human animal cell. In some embodiments, the contacting step is performed in a plant cell.

[0377] The present disclosure also provides the use of any one of the adenine base editors described herein as a medicament.The present disclosure also provides the use of any one of the complexes of the adenine base editor and guide RNA described herein as a medicament.

[0378] kit The compositions of the present disclosure can be assembled into a kit. In some embodiments, the kit comprises a nucleic acid vector for expressing the nucleobase editor described herein. In some embodiments, the kit further comprises a suitable guide nucleotide sequence (e.g., gRNA) for targeting the Cas9 protein or nucleobase editor to a desired target sequence, or a nucleic acid vector for expressing such a guide nucleotide sequence.

[0379] The kits described herein may include one or more containers that contain components for carrying out the methods described herein, and optionally instructions for use. Any of the kits described herein may further include components required for carrying out assay methods. Each component of the kit may be provided in liquid form (e.g., in solution) or in solid form (e.g., dry powder), where applicable. In certain cases, some of the components may be reconstituted or otherwise processable (e.g., into active form) by, for example, adding a suitable solvent or other species (e.g., water), which may or may not be provided with the kit.

[0380] In some embodiments, the kit may optionally include instructions and / or promotion for the use of the components provided. As used herein, "instructions" defines the components of instructions and / or promotion, and typically involves written instructions on or associated with the packaging of the present disclosure. Instructions may also include any oral or electronic instructions provided in any format that clearly identifies to the user that the instructions should be associated with the kit, such as audiovisual (e.g., videotape, DVD, etc.), Internet, and / or web-based communication, etc. Written instructions may be in a form dictated by a governmental agency that regulates the manufacture, use, or sale of pharmaceutical or biological products. These may also reflect the approval by the authorities of manufacture, use, or sale for animal administration. As used herein, the term "promoted" includes all methods of doing business related to the present disclosure, including education, hospital and other clinical instructions, scientific research, drug discovery or development, academic research, pharmaceutical sales, pharmaceutical industry activities, and any methods of advertising or other promotional activities, including any form of written, oral, and electronic communication. In addition, the kit may include other components, depending on the particular application, as described herein.

[0381] The kit may contain any one or more of the components described herein in one or more containers. The components may be aseptically prepared, packaged in syringes, and shipped frozen. Alternatively, it may be contained in a vial or other container for storage. A second container may have other components prepared aseptically. Alternatively, the kit may include active agents that are premixed and shipped in a vial, tube, or other container.

[0382] The kits may have various forms, such as blister pouches, shrink-wrap pouches, vacuum-sealed pouches, sealed thermoformed trays, or similar pouch or tray forms, with the accessories loosely packed in a pouch, one or more tubes, containers, boxes, or bags. The kits may be sterilized after the accessories are added, thereby allowing the individual accessories in the container to be otherwise unpackaged. The kits may be sterilized using any suitable sterilization technique, such as radiation sterilization, heat sterilization, or other sterilization methods known in the art. The kits may also include other components, such as containers, cell culture media, salts, buffers, reagents, syringes, needles, fabrics such as gauze for applying or removing disinfectants, disposable gloves, drug supports prior to administration, etc., depending on the specific application.

[0383] host cell Cells that may contain any of the compositions described herein include prokaryotic and eukaryotic cells. The methods described herein are used to deliver Cas9 proteins or nucleobase editors into eukaryotic cells (e.g., mammalian cells, such as human cells). In some embodiments, the cells are in vitro (e.g., cultured cells. In some embodiments, the cells are in vivo (e.g., in a subject, such as a human subject). In some embodiments, the cells are ex vivo (e.g., isolated from a subject and can be administered again to the same or a different subject).

[0384] Mammalian cells of the present disclosure include human cells, primate cells (e.g., vero cells), rat cells (e.g., GH3 cells, OC23 cells), or mouse cells (e.g., MC3T3 ("3T3") cells or mouse neuroblastoma neuro-2A ("N2A") cells). There are various human cell lines, including, without limitation, human embryonic kidney (HEK or HEK293T) cells, HeLa cells, cancer cells from the National Cancer Institute 60 cancer cell line (NCI60), DU145 (prostate cancer) cells, Lncap (prostate cancer) cells, MCF-7 (breast cancer) cells, MDA-MB-438 (breast cancer) cells, PC3 (prostate cancer) cells, T47D (breast cancer) cells, THP-1 (acute myeloid leukemia) cells, U87 (glioblastoma) cells, SHSY5Y human neuroblastoma cells (cloned from a myeloma), and Saos-2 (bone cancer) cells. In some embodiments, the rAAV vector is delivered into human embryonic kidney (HEK) cells (e.g., HEK 293 or HEK 293T cells). In some embodiments, the rAAV vector is delivered into stem cells (e.g., human stem cells), such as pluripotent stem cells (e.g., human pluripotent stem cells, including human induced pluripotent stem cells (hiPSCs)). Stem cells refer to cells that have the ability to divide indefinitely in culture to give rise to specialized cells. Pluripotent stem cells refer to a type of stem cell that can differentiate into all tissues of an organism, but cannot support the development of a complete organism by itself. Human induced pluripotent stem cells refer to somatic cells (e.g., mature or adult) that are reprogrammed into an embryonic stem cell-like state by being forced to express genes and factors that are important for maintaining the characteristics that define embryonic stem cells (see, e.g., Takahashi and Yamanaka, Cell 126 (4):663-76, 2006, which is incorporated herein by reference). Human induced pluripotent stem cells express stem cell markers and are capable of generating cells characteristic of all three germ layers (ectoderm, endoderm, and mesoderm).

[0385] The 293-T, 293-T, 3T3, and 293-T, 3T3 N2A, 4T1, 721, 9L, A-549, A172, A20, A253 A2780, A2780ADR, A2780cis, A431, ALC, B16, B35, BCP-1, BEAS-2B, bEnd.3, BHK-21, BR 293. BxPC3, C2C12, C3H-10T1 / 2, C6, C6 / 36, Cal-27, CGR8, CHO, CML T1, CMT, COR-L23, COR-L23 / 5010, COR-L23 / CPR, COR-L23 / R23, COS-7, COV-434, CT26, D17, DH82, DU145, Du CaP, E14Tg2a, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, Hepa1c1c7, High Five Ranger HL-60 HMEC HT-29 HUVEC J558L Ranger Jurkat JY Ranger K562 Ranger KCL22 KG1 Ku812 KYO1 LNCap Ma-Mel 1, 2, 3....48, MC-38, MCF-10A, MCF-7, MDA-MB-231, MDA-MB-435, MDA-MB-468, MDCK II, MG63, MONO-MAC 6 MOR / 0.2R MRC5 MTD-1A MyEnd NALM-1 NCI-H69 / CPR NCI-H69 / LX10、NCI-H69 / LX20、NCI-H69 / LX4、NIH-3T3、NW-145、OPCN / OPCT Peer、PNT-1A / PNT 2. PTK2, Raji, RBL, RenCa, RIN-5F, RMA / RMAS, S2, Saos-2, Sf21, Sf9, SiHa, SKBR3, SKO V-3, T-47D, T2, T84, THP1, U373, U87, U937, VC aP, WM39, WT-49, X63, YAC-1, and the YAR converters.

[0386] Without further elaboration, it is believed that one skilled in the art can utilize the present disclosure to its fullest extent based on the above description. Therefore, the following specific embodiments are to be construed as merely illustrative, and in no way limiting of the remainder of the disclosure. All publications cited herein are incorporated by reference for the purposes or subject matter referenced herein. EXAMPLES

[0387] Example 1 In this study, we constructed small, highly active ABE8e variants and identified the minimal required cis-acting components on the AAV genome to develop highly efficient single AAV vectors with broad in vivo targeting capabilities. ABE8e variants using the compact CjCas9, Nme2Cas9, and SauriCas9 domains wer...

Claims

1. (i) 5' inverted terminal repeat (ITR); (ii) A first nucleic acid fragment comprising a sequence encoding a base editor operably linked to a first promoter, and a polyadenylation (polyA) signal, Here, the base editor includes a nucleic acid programmable DNA-binding protein (napDNAbp) domain and a deaminase domain; Here, the first promoter is the EF-1α short (EFS) promoter; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably ligated to a second promoter; and (iv) 3'ITR A nucleic acid molecule containing, Here, the length between the 5'ITR and the 3'ITR is less than approximately 4.90 kb. The nucleic acid molecule.

2. The nucleic acid molecule according to claim 1, wherein the nucleic acid molecule does not contain a post-transcriptional response element or intein.

3. The nucleic acid molecule according to claim 1, wherein the first promoter has a length of less than 280 nucleotides, less than 270 nucleotides, less than 265 nucleotides, or less than 250 nucleotides.

4. The nucleic acid molecule according to claim 1, wherein the second promoter is an EF-1α short (EFS) promoter, a MeCP2 promoter, a P3 promoter, a U1A promoter, or a U6 promoter.

5. The nucleic acid molecule according to claim 4, wherein the second promoter is the U6 promoter.

6. The nucleic acid molecule according to claim 1, wherein the napDNAbp domain includes a Cas9 domain.

7. The napDNAbp domain (1) S. aureus Cas9 (SaCas9) domain, N. meningitidis 2 Cas9 (Nme2Cas9) domain, C. jejuni Cas9 (CjCas9) domain, S. auricularis (SauriCas9) domain, SaKKH domain, or variants thereof; (2) S. pyogenes Cas9 (SpCas9), Cpf1, CasX, CasY, C2c1, C2c2, C2c3, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Cas3, or compact variants of CasΦ; or (3) SaCas9 nickased domain, SaKKH nickased domain, CjCas9 nickased domain, or variants thereof A nucleic acid molecule according to claim 1, comprising:

8. The nucleic acid molecule according to claim 1, wherein the napDNAbp domain has nickase activity.

9. The nucleic acid molecule according to claim 1, wherein the deaminase domain comprises an adenosine deaminase selected from the group consisting of TadA-8e, TadA-8e(V106W), TadA9, TadA20, and TadA7.10 deaminases.

10. The nucleic acid molecule according to claim 1, wherein the deaminase domain comprises a cytidine deaminase domain selected from the group consisting of FERNY and evolved FERNY (evoFERNY).

11. The nucleic acid molecule according to claim 1, wherein the base editor further comprises one or more nuclear localization sequences (NLS).

12. The nucleic acid molecule according to claim 1, wherein the base editor is ABE8e, ABE8e(V106W), ABE9, ABE20, ABE7.10, SaKKH-ABE8e, SauriCas9-ABE8e, CjCas9-ABE8e, Nme2Cas9-ABE8e, SaCas9-ABE8e, BE3.9, FERNY-BE3.9, or a variant thereof.

13. The nucleic acid molecule according to claim 11, wherein the base editor further comprises a uracil glycosylase inhibitor (UGI) domain.

14. The nucleic acid molecule according to claim 1, wherein the base editor comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 92.5%, at least 95%, at least 97%, at least 98%, or at least 99% identity with either SEQ ID NO: 21 or 22.

15. The nucleic acid molecule according to claim 14, comprising an amino acid sequence in which the base editor is identical to any one of sequence numbers 171-172 and 181-183.

16. The nucleic acid molecule according to claim 1, wherein the base editor comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 92.5%, at least 95%, at least 97%, at least 98%, or at least 99% identity with any one of sequence numbers 171-172 and 181-183.

17. The nucleic acid molecule according to claim 16, comprising an amino acid sequence in which the base editor is identical to any one of sequence numbers 171-172 and 181-183.

18. The nucleic acid molecule according to claim 1, wherein the transcription direction of the second nucleic acid fragment is reversed relative to the transcription direction of the first nucleic acid fragment.

19. The nucleic acid molecule according to claim 1, wherein the nucleic acid molecule comprises a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% identity with any one of sequence numbers 100 to 102.

20. The nucleic acid molecule according to claim 19, wherein the nucleic acid molecule contains a nucleic acid sequence that is identical to any one of the nucleic acid sequences of sequence numbers 100 to 102.

21. (i) 5' inverted terminal repeat (ITR); (ii) (a) A sequence encoding a SaKKH-ABE8e, SauriCas9-ABE8e, CjCas9-ABE8e, or Nme2Cas9 base editor operably linked to a first promoter, where the first promoter is selected from the group consisting of EFS, MeCP2, P3, and U1A promoters; and (b) bGH polyadenylation (polyA) signal The first nucleic acid fragment containing; (iii) A second nucleic acid fragment encoding a guide RNA (gRNA) operably ligated to the U6 promoter, wherein the transcription direction of the second nucleic acid fragment is reversed relative to the transcription direction of the first nucleic acid fragment; and (iv) 3'ITR The nucleic acid molecule according to claim 1, comprising from 5' to 3'.

22. (i) 5' inverted terminal repeat (ITR); (ii) (a) a sequence encoding a CjCas9-FERNY-BE3.9 or CjCas9-evoFERNY-BE3.9 base editor operably linked to a first promoter, where the first promoter is selected from the group consisting of EFS, MeCP2, P3, and U1A promoters; and (b) bGH polyadenylation (polyA) signal The first nucleic acid fragment containing; (iii) A second nucleic acid fragment encoding a guide RNA (gRNA) operably ligated to the U6 promoter, wherein the transcription direction of the second nucleic acid fragment is reversed relative to the transcription direction of the first nucleic acid fragment; and (iv) 3'ITR The nucleic acid molecule according to claim 1, comprising from 5' to 3'.

23. (i) 5' inverted terminal repeat (ITR); (ii) A first nucleic acid fragment comprising a transgene operably coupled to a first promoter, wherein the first promoter is an EF-1α short (EFS) promoter, and a transcription terminator that does not contain a post-transcriptional response element; (iii) a second nucleic acid fragment operably linked to a second promoter, wherein the transcription direction of the second nucleic acid fragment is reversed relative to the transcription direction of the first nucleic acid fragment; and (iv) 3'ITR An AAV nucleic acid molecule containing in the order from 5' to 3', Here, the length between the 5'ITR and the 3'ITR is less than approximately 4.90 kb. The AAV nucleic acid molecule.

24. Recombinant AAV (rAAV) particles comprising a nucleic acid molecule as described in claim 1, capsidized to a capsid, wherein optionally the capsid is serotype 2, 6, 8, 9, PH.B, or PH.eB.

25. A composition comprising the nucleic acid molecule described in claim 1.

26. A pharmaceutical composition comprising the nucleic acid molecule described in claim 1 and a pharmaceutically acceptable carrier.

27. A cell comprising the nucleic acid molecule described in claim 1, wherein the cell is optionally a bacterial cell or a eukaryotic cell.

28. A method for editing a target nucleic acid molecule, comprising contacting it with a cell comprising a nucleic acid molecule according to any one of claims 1 to 22, an AAV nucleic acid molecule according to claim 23, rAAV particles according to claim 24, a composition according to claim 25, a pharmaceutical composition according to claim 26, or a cell according to claim 27.

29. The method according to claim 28, wherein the step of making contact is performed on an object, the object being not a human.

30. The method according to claim 29, wherein the subject has been diagnosed with a disease or disability.

31. The method according to claim 28, wherein the target sequence is located within the PCSK9 gene or the ANGPTL3 gene.

32. The step of making contact is (i) editing efficiency of at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, or at least approximately 85%; and / or (ii) Base editing:indel ratios greater than at least approximately 5:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, or approximately 15:1; and / or (iii) Indel rates of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less; and / or (iv) Minimal bystander editing The method according to claim 28, which brings about the following.

33. A method comprising contacting a cell with a nucleic acid molecule according to any one of claims 1 to 22, an AAV nucleic acid according to claim 23, rAAV particles according to claim 24, a composition according to claim 25, or a pharmaceutical composition according to claim 26, wherein the contact results in the delivery of the nucleic acid molecule into the cell.

34. The method according to claim 33, wherein the cells are human cells, and optionally the cells are cardiac cells or muscle cells.

35. The method according to claim 33, wherein the contact step results in an editing efficiency of at least about 20%, at least about 22%, at least about 24%, at least about 27%, at least about 30%, at least about 33%, or at least about 36% in the cells.

36. The method according to claim 33, wherein the contact step results in an indel rate of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less in the cells.

37. Use in the manufacture of a pharmaceutical product of a nucleic acid molecule according to any one of claims 1 to 22, AAV nucleic acid according to claim 23, rAAV particles according to claim 24, the composition according to claim 25, the pharmaceutical composition according to claim 26, or the cells according to claim 27.

38. Use of a nucleic acid molecule according to any one of claims 1 to 22, the AAV nucleic acid according to claim 23, the rAAV particles according to claim 24, the composition according to claim 25, the pharmaceutical composition according to claim 26, or the cells according to claim 27, for treating a disease or disorder in a subject that requires it.

39. The use according to claim 38, wherein the subject is a human.

40. Use of a nucleic acid molecule according to any one of claims 1 to 22, the AAV nucleic acid according to claim 23, the rAAV particles according to claim 24, the composition according to claim 25, the pharmaceutical composition according to claim 26, or the cells according to claim 27 for editing a target nucleic acid molecule in a subject that requires it.