AAV vectors encoding base editor and uses thereof

By developing AAV vectors with minimal size, the problem of delivering base editors in a single AAV vector is solved, efficient tissue editing is achieved, and delivery complexity and dose requirements are reduced. It is suitable for mammalian tissues that are difficult to transduce, such as the heart, CNS and muscle.

CN120456933APending Publication Date: 2025-08-08THE BROAD INST INC +1

Patent Information

Application Number
CN202380049631.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-07-15
Filing Date
2023-04-28
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art has difficulty delivering base editors efficiently in a single AAV vector, resulting in high dose demand and delivery complexity, especially in inefficient in difficult-to-transduce mammalian tissues such as the heart, CNS and muscle tissue.

Method used

AAV vector with minimization size is developed, including a minimization size regulatory component and a base editor, which can package and effectively express a base editor in a single AAV vector, avoiding the use of trans-splicing integrative peptides and simplifying the delivery process.

Benefits of technology

A similar or improved editing efficiency to dual AAV delivery in multiple tissues is achieved, reducing the required dose of AAV, simplifying the production and characterization process, improving editing efficiency, and suitable for tissues that are more difficult to transduce.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456933A_ABST
    Figure CN120456933A_ABST
Patent Text Reader

Abstract

Described herein are nucleic acid molecules, compositions, recombinant AAV (rAAV) particles, kits, and methods for delivering a base editor (or a "nucleobase editor") to a cell, for example via an AAV vector. In particular, the present disclosure provides compositions, methods and uses for delivery of adenine base editors and cytosine base editors in a single AAV vector (or genome). Further described herein are improved AAV vectors comprising a minimised size regulatory component, for example, a regulatory component that can enable packaging of a base editor. Provided herein are methods and compositions for delivering a base editor protein to a cell or tissue in a single recombinant AAV (rAAV) vector. Improved methods and compositions for in vivo delivery of these base editors in a single rAAV particle are contemplated herein. Further provided herein are base editors, and compositions and cells comprising these base editors.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority under 35 USC §119(e) to U.S. Provisional Application USSN 63 / 336064, filed April 28, 2022, and U.S. Provisional Application USSN 63 / 389796, filed July 15, 2022, each of which is incorporated herein by reference.

[0003] Government support clauses

[0004] This invention was supported by the National Institutes of Health under Grant Nos. UG3AI150551, U01AI142756, R35GM118062, RM1HG009490, R01EY009339, R01HL148769, and R35HL145203. The government has certain rights in this invention.

[0005] References to electronic sequence listings

[0006] The contents of the electronic Sequence Listing (B119570158WO00-SEQ-JQM.XML; size: 447,661 bytes; and creation date: April 28, 2023) are incorporated herein by reference in their entirety. Background Art

[0007] Gene editing offers the potential for clinically validated treatments of a wide range of genetic disorders for which few therapeutic options are available. Because most genetic disorders studied and treated by gene editing require in vivo editing, clinically useful methods for efficiently delivering precise gene editing agents into tissue cells of animals, such as mammals, are essential. 1,74 Continue to play an important role in promoting the development of this field.

[0008] Adeno-associated virus (AAV) has been used in animal models of human disease 3,4 In clinical trials 5 and in FDA-approved drugs 6,7 AAV has become a popular in vivo delivery method for delivering genes encoding many therapeutic proteins because it is clinically validated, can target a variety of clinically relevant tissues, and has a relatively well-understood and favorable safety profile. SUMMARY OF THE INVENTION

[0010] In some aspects, nucleic acid molecules, compositions, recombinant AAV (rAAV) particles, kits and methods for delivering complete base editors (or "nucleobase editors") to cells, for example, via a single AAV vector (or genome) are described herein. In particular, the present disclosure provides compositions, methods and uses for delivering adenine base editors and cytosine base editors with minimized size in a single AAV vector, wherein the adenine base editor and related regulatory elements have a length shorter than the packaging capacity of AAV, which is ~4.9 kilobases (kb). Improved AAV vectors are further described herein, which contain minimized-size regulatory components that enable packaging of base editors. The present disclosure provides host cells and compositions comprising the disclosed rAAV particles. Improved AAV vectors are further provided herein, which contain minimized-size regulatory components that enable packaging of larger transgenes other than base editors.

[0011] base editors 8,9 (BE) can efficiently install targeted mutations in a variety of therapeutically relevant cell types in vitro and in animal models of human genetic diseases 1,10 BE can also effectively install targeted mutations in a variety of therapeutically relevant tissues (including liver tissue) in subjects (e.g., human subjects). Unlike nuclease-mediated gene editing, base editing does not require double-stranded DNA breaks and therefore generates minimal undesirable indel byproducts, chromosomal translocations, or other chromosomal abnormalities. 11 , chromosomal aneuploidy 12 , big missing 13,14 , p53 activation 15,16 Chromosome fragmentation 17 Base editors can correct point mutations that cause various genetic diseases, but their delivery to subjects in vivo is complicated by their large size (~5.2 kb), which often exceeds the maximum packaging capacity of adeno-associated virus (AAV), which is ~4.9 kb between inverted terminal repeats (ITRs). 18,19 . Because the optimal packaging capacity of AAV is ~4.7 kb, there are significant obstacles to delivering base editors capable of acting on target DNA in a single AAV particle. For example, it is not feasible to package a BE containing the commonly used Streptococcus pyogenes Cas9 protein (SpCas9, which is about 4.2 kb in length) and a single guide RNA (sgRNA) in a single vector. Although technically feasible, this approach leaves little room for customized expression and control elements such as nuclear localization sequences (NLS). In addition to the base editor itself, the AAV that delivers the base editor must also include a guide RNA, a promoter that drives the expression of the base editor and sgRNA, and cis regulatory elements.

[0012] In previous studies20-26 , AAV is used to deliver a base editor by splitting the base editor into two halves in two AAV nucleic acid vectors. See U.S. Patent Publication No. 2018 / 0127780, published on May 10, 2018, PCT Publication No. WO 2020 / 236982, published on November 26, 2020, Levy, JM, et al. Nat Biomed Eng 4, 97-110 (2020), Chen, Y., et al. Development of Highly Efficient Dual-AAV Split Adenosine Base Editor for In Vivo Gene Therapy. Small Methods 4, 2000309 (2020) and Villiger, L. et al. Nature Medicine 24, 1519-1525 (2018), each of which is incorporated herein by reference. In the dual AAV method, each “split” part of the base editing transgene is fused to a small trans-splicing intein 27 , or each part is expressed as an mRNA that undergoes trans-splicing 28 . Typically, the two AAV vectors are each packaged in a separate AAV viral particle (or virion). The dual AAV approach relies on the incorporation of a trans-splicing intein that, following delivery and transduction of the AAV particles, mediates the reconstruction of the full-length BE from the cleaved parts in the cell. Because two AAV particles are required to deliver a single base editor, two successful and relatively simultaneous transductions of the target cells are necessary.

[0013] While dual AAV delivery of base editors has supported therapeutic levels of editing, including in mouse models of human disease, the development of a single AAV base editing system will further increase its potential impact by simplifying the application, characterization, and preparation of base editors, as well as potentially increasing editing efficiency by eliminating the need for simultaneous transduction of multiple AAVs. The single AAV base editing system disclosed herein also reduces the required dose of AAV, an important advance because the clinical application of AAV is often limited by dose-limiting toxicities. 29 Therefore, there is a need in the art for single AAV in vivo base editor delivery that requires only a single transduction of target cells while maintaining high editing efficiency. Single AAV delivery would offer advantages for research and clinical use, particularly in mammalian tissues that are more difficult to transduce efficiently, such as heart, CNS, and muscle tissue.

[0014] Adenine base editors (ABEs) are a particularly useful class of editing agents because they install A•T-to-G•C transitions, which can correct approximately half of all known disease-causing SNPs. 9 Phage-assisted continuous evolution (PACE) of the adenine base editor ABE7.10 recently generated TadA-8e, a deoxyadenosine deaminase with increased activity and broader compatibility with Cas domains other than SpCas9. 30 ABE7.10, containing the TadA7.10 deaminase, can perform clean and efficient A•T-to-G•C conversion in DNA in cultured cells, adult mice, plants, and other organisms, with very low levels of undesirable side products such as small insertions or deletions (indels). Additional details about TadA-8e and TadA7.10 deaminases can be found in PCT Publication No. WO 2021 / 158921, published on August 12, 2021; PCT Publication No. WO 2018 / 027078, published on February 8, 2018, PCT Patent Publication No. WO 2019 / 079347, published on April 25, 2019; Koblan et al., Nat Biotechnol 36, 843-846 (2018); and Gaudelli et al., Nature 551, 464-471 (2017), each of which is incorporated herein by reference. ABEs containing only a single TadA deaminase domain rather than a single-chain dimer allow for reduced editor size 30,31 . In addition, although SaCas9 is small enough (1053 amino acids in length, SEQ ID NO: 377) to provide a single AAV-compatible base editor, its utility is greatly limited by the rarity of its NNGRRT PAM. Since base editing requires the presence of a suitable PAM to place the target nucleotide within the editing window, ABEs that jointly provide broad PAM compatibility and simple and efficient in vivo delivery will facilitate the application of base editing in vivo.

[0015] Similarly, cytosine base editors (CBEs) that offer broad PAM compatibility and simple and efficient in vivo delivery will facilitate the application of base editing in vivo. Current CBEs contain a uracil glycosylase inhibitor domain that is approximately 84 bp in length. Although not large, these additional 84 base pairs make CBEs more difficult to deliver than ABEs in a single AAV vector.

[0016] Ran and colleagues engineered a size-minimized Staphylococcus aureus Cas9 (SaCas9) for in vivo delivery in a single AAV vector to install double-strand breaks into target genomic DNA. See Ran, FA, et al. (2015) Nature 520(7546): 186-191, which is incorporated herein by reference. Ran's AAV cassette contains a U6 promoter-driven sgRNA and a cytomegalovirus (CMV) promoter-driven SaCas9 transgene. Recently, Tran and colleagues engineered a single AAV vector containing a size-minimized SaCas9 ABE, microABE I744, in which the ABE7.10 TadA deaminase monomer was embedded (i.e., inserted) within the SaCas9 domain. See Tran et al., Nat. Commun. 11, 4871 (2020), which is incorporated herein by reference. However, this AAV-encoded ABE showed only <0.25% editing in vitro and was not evaluated in vivo. Recently, Zhang, Sontheimer and colleagues generated a single AAV vector encoding an ABE containing the Neisseria meningitidis 2 Cas9 (Nme2Cas9) protein. See Zhang et al., Adenine Base Editing in vivo with a Single Adeno-AssociatedVirus Vector. bioRxiv 2021.12.13.472434, which is incorporated herein by reference. However, after AAV delivery to liver tissue, Zhang's base editor showed significant editing at the bystander adenine, with a maximum editing efficiency of 35%. However, single AAV delivery of base editors containing other compact Cas9 protein domains that exhibit efficient base editing has not yet been disclosed, as is the case with the currently used ABEs and CBEs. In addition, single AAV delivery of base editors to heart and muscle tissue has not yet been disclosed.

[0017] The present disclosure provides a minimized size AAV vector having a length of less than about 4.90 kb between ITRs. This single AAV base editing platform provides similar or improved editing efficiency compared to dual AAV ABE8e in multiple tissues at multiple doses when delivered systemically to mice. The exemplary AAV vectors of the present disclosure do not contain or rely on the use of trans-splicing inteins to achieve successful delivery.

[0018] The AAV vectors disclosed herein are based at least in part on advances in genetic engineering that have generated vectors containing size-minimized components necessary for efficient expression in target cells and editing target bases in vivo. Exemplary target cells include muscle cells, neurons, hepatocytes, neuromuscular cells, and cardiac cells. These vectors are smaller than those disclosed in U.S. Patent Publication No. 2018 / 0127780, published on May 10, 2018, and PCT Publication No. WO 2020 / 236982, published on November 26, 2020, and are therefore suitable for incorporation into single AAV particles. In particular, the disclosed AAV vectors are based in part on the discovery that post-transcriptional response elements such as WPRE in transcription terminators (or polyadenylation signals) are not necessary for the successful expression of base editors in target tissues in vivo. Therefore, the disclosed AAV vectors contain shorter (or size-minimized) terminators. The disclosed AAV vectors further contain other size-minimized regulatory elements, such as short promoters.

[0019] The disclosed AAV vectors are also based in part on the following findings: a guide RNA compatible with any of the disclosed Cas proteins can be encoded at the 3' end of the vector and maintain a total ITR-to-ITR length of less than 4.9 kb, less than 4.8 kb, less than 4.7 kb, less than 4.6 kb, or less than 4.5 kb. Thus, through the genetic engineering techniques described herein, base editors and their guides can be effectively packaged in a single rAAV particle for in vivo delivery. The guide RNA can be encoded in the vector in an opposite direction (3' to 5') to the base editor (ie, 5' to 3') and the promoter driving the expression of the base editor transgene (and its replication origin).

[0020] The present disclosure also provides base editors with minimized size. These base editors are developed to achieve effective in vivo base editing mediated by a single AAV particle. The disclosed AAV-encoded base editors may include Cas proteins with minimized size. The length of these Cas9 proteins is approximately 1000-1050 amino acids, which is about 350 amino acids shorter than the SpCas9 protein. These size-minimized Cas proteins include but are not limited to Staphylococcus aureus Cas9 (SaCas9), Nme2Cas9, Campylobacter jejuni Cas9 (CjCas9), Staphylococcus aureus Cas9 (SauriCas9) and variants of any of these Cas9 proteins. The disclosed base editors may contain any of these Cas9 proteins or an evolved or mutated variant of any of the Cas9 proteins disclosed herein.

[0021] Thus, in various embodiments, the disclosed AAV nucleic acid molecules do not comprise an intein, such as a trans-splicing intein (e.g., does not comprise a trans-splicing intein derived from Nostoc punctata or Npu). In various embodiments, the disclosed AAV nucleic acid molecules comprise a transcription terminator that does not contain a post-transcriptional response element. In some embodiments, the disclosed AAV nucleic acid molecules do not comprise an intein or a post-transcriptional response element. In some embodiments, the nucleic acid molecule comprises a first nucleic acid fragment comprising: (i) a 5' inverted terminal repeat (ITR); (ii) a first nucleic acid fragment comprising a sequence encoding a base editor operably linked to a first promoter, wherein the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain; and a polyadenylation (poly A) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a second promoter; and (iv) a 3' ITR.

[0022] In some aspects, provided herein are rAAV vectors with size-minimized regulatory elements that allow packaging of large transgenics. Accordingly, provided herein are rAAV nucleic acid molecules that comprise, in 5' to 3' order: (i) 5' inverted terminal repeats (ITRs); (ii) a first nucleic acid fragment comprising a transgenic operably linked to a first promoter, wherein the first promoter has a length of less than 300 nucleotides; and a transcription terminator that does not contain a post-transcriptional response element; (iii) a second nucleic acid fragment that is operably linked to a second promoter, wherein the transcription direction of the second nucleic acid fragment is reverse relative to the transcription direction of the first nucleic acid fragment; and (iv) 3' ITR. In various embodiments, the length between 5' ITR and 3' ITR is less than about 4.90 kb. In some embodiments, the length between 5' ITR and 3' ITR is less than about 4.85 kb, less than about 4.80 kb, less than about 4.75 kb, less than about 4.725 kb, less than about 4.70 kb, or less than 4.65 kb.

[0023] In some embodiments, the first nucleic acid fragment encodes a base editor and the second nucleic acid fragment encodes a gRNA. In some embodiments, the first nucleic acid fragment encodes a protein that is not a base editor.

[0024] Any disclosed base editor may include (i) a napDNAbp domain; and (ii) a deaminase domain. The disclosed base editor may include a wild-type napDNAbp domain (e.g., wild-type SaKKH-Cas9). The disclosed base editor may include a napDNAbp domain with nickase activity (e.g., SaKKH-Cas9 nickase or "SaKKH"). In some embodiments, the napDNAbp domain is a Cas9 nickase domain. In some embodiments, the napDNAbp domain is a SaKKH-Cas9 nickase. The napDNAbp domain may be selected from Staphylococcus aureus Cas9 (SaCas9), Neisseria meningitidis 2 Cas9 (Nme2Cas9), Campylobacter jejuni Cas9 (CjCas9), or Staphylococcus auris (SauriCas9) domains and variants thereof. These Cas proteins have a wider PAM compatibility than standard SpCas9 proteins. Thus, the disclosed single AAV-encoded base editor can potentially target the vast majority of adenines throughout the genome or the vast majority of cytosines throughout the genome. In exemplary embodiments, the napDNAbp domain is a SaCas9 domain, a SaCas9 nickase domain, a SaKKH domain, or a SaKKH nickase domain.

[0025] The AAV-encoded ABEs disclosed herein comprise an adenosine deaminase domain containing a single deaminase, i.e., a deaminase monomer (e.g., TadA-8e monomer), rather than an adenosine deaminase dimer (i.e., two adenosine deaminases). The use of a deaminase monomer facilitates the generation of a base editor with minimized size. The TadA monomer is approximately 166 amino acids long.

[0026] Any disclosed adenine base editor may include an adenosine deaminase domain that is a variant of the Escherichia coli TadA deaminase. In some embodiments, the adenosine deaminase is selected from TadA-8e, TadA-8e (V106W), TadA9, TadA20, and TadA7.10 deaminase. In some embodiments, any base editor disclosed herein includes an adenosine deaminase fused to the N-terminus of the napDNAbp domain, such as a Cas9 nickase. In some embodiments, the adenosine deaminase is TadA-8e.

[0027] In some aspects, the present disclosure provides ABE8e variants with minimized size. Each variant is compatible with single AAV delivery, and three such variants together provide PAM compatibility sufficient to target 87% of adenines and edit 82% of adenines in the human genome. The three variants are Sauri-ABE8e, SaKKH-ABE8e, and SaABE8e. Each contains a TadA-8e adenosine deaminase and a nickase variant of SauriCas9, SaKKH-Cas9, and SaCas9, respectively. The present disclosure further provides ABE8e variants CjCas9-ABE8e and Nme2Cas9-ABE8e. The present disclosure further provides ABE variants SaKKH-ABE8e(V106W), SauriCas9-ABE8e(V106W), CjCas9-ABE8e(V106W), Nme2Cas9-ABE8e(V106W) and SaCas9-ABE8e(V106W); SaKKH-ABE9, SauriCas9-ABE9, CjCas9-ABE9, Nme2Cas9-ABE9 and SaCas9-ABE9; SaKKH-ABE20, SauriCas9-ABE20, CjCas9-ABE20, Nme2Cas9-ABE20 and SaCas9-ABE20; and SaKKH-ABE7.10, SauriCas9-ABE7.10, CjCas9-ABE7.10, Nme2Cas9-ABE7.10 and SaCas9-ABE7.10. In any of these disclosed base editors, wild-type or nickase variants of SauriCas9, SaKKH-Cas9, SaCas9, CjCas9, and Nme2Cas9, respectively, can be used.

[0028] The present disclosure further provides CBE variants with minimized size, in particular BE3.9 variants with minimized size. Examples of these variants include CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9. In some embodiments, the base editor further comprises a uracil glycosylase inhibitor (UGI) domain. An exemplary cytidine deaminase (e.g., rAPOBEC1 deaminase) of the CBE disclosed herein has a size of approximately 229 amino acids.

[0029] By integrating these advances, the disclosed single AAV-delivered base editors were used to treat a mouse model of high cholesterol, which is associated with cardiovascular disease, resulting in the correction of an occasional mutation in heart tissue and an increase in the animals' lifespan.

[0030] In the examples disclosed herein, single AAV delivery of ABEs achieved 66%, 33%, and 22% editing in liver, heart, and muscle tissues, respectively, at doses lower than or similar to those used in recent preclinical and clinical studies of AAV particles targeting these tissues (see, e.g., clinical trial numbers NCT02122952 (SMA treatment) and NCT03375164 (DMD treatment)). Furthermore, it was demonstrated that a single AAV8 ABE can effectively edit therapeutically relevant targets in mice with 60%-85% efficiency. These edits produced disruptions at the native splice acceptor sites in the target human Pcsk9 and mouse Pcsk9 and mouse Angptl3 genes, which resulted in nearly complete (average 93%) knockout of these genes at doses lower than previously reported, leading to significant reductions in plasma cholesterol and triglycerides. The liver protein proprotein convertase subtilisin / Kexin type 9 (PCSK9) is a secreted, globular, self-activating serine protease that acts as a protein-binding adaptor within endosomal vesicles to bridge pH-dependent interactions with low-density lipoprotein receptors (LDL-Rs) during endocytosis of LDL particles, preventing LDL-Rs from recycling to the cell surface and leading to reduced LDL cholesterol clearance. Angiopoietin-like 3 protein (Angptl3) is an endogenous inhibitor of lipoprotein lipase (LPL), the main enzyme involved in the hydrolysis of triglyceride-rich lipoproteins. Base editors targeting pathogenic mutations in the Pcsk9 gene are disclosed in U.S. Publication No. 2018 / 0237787, published on August 23, 2018, which is incorporated herein by reference.

[0031] By minimizing the size of adenine base editors and AAV components, a set of single AAV adenine base editor systems was developed that have broad targeting capabilities and support robust in vivo editing due to their collective PAM compatibility. The single AAV BE vectors disclosed herein facilitate base editing for research and therapeutic applications by simplifying production and characterization and by reducing the total dose of AAV required to achieve the desired editing level. These single AAV BEs offer several potential advantages over dual AAV approaches for clinical use: clinical-scale production of a single vector instead of two vectors; increased efficacy, especially at lower doses; and reduced complexity from a simpler construct that eliminates the need to use trans-splicing inteins. For these reasons, in vivo editing methods compatible with single AAV delivery can be more easily applied to large animal models and human treatments where systemic delivery is typically used. The development of smaller promoters that provide sufficient expression of base editors allows further minimization of the elements of single AAV ABEs, which can facilitate clinical translation by increasing the proportion of full-length packaged AAV genomes.

[0032] In other aspects, host cells comprising the compositions described herein are provided. The disclosed cells can comprise any disclosed nucleic acid molecules, rAAV vectors, or rAAV particles described herein. In other aspects, kits comprising any disclosed rAAV particles and instructions for delivery to cells (such as host cells) are provided.

[0033] Other aspects of the present disclosure provide methods comprising contacting a target nucleic acid molecule with any composition described herein. In various embodiments, the target nucleic acid molecule is in a cell, such as a eukaryotic cell (e.g., a mammalian cell). In some embodiments, the target nucleic acid molecule is genomic DNA. In some embodiments, the genomic DNA is in a cell or tissue of a subject (e.g., a human subject). Therefore, methods comprising contacting a cell with any rAAV particle or composition disclosed herein are contemplated.

[0034] Other aspects of the present disclosure provide methods comprising administering a therapeutically effective amount of any of the compositions described herein (or rAAV particles) to a subject in need thereof. In some embodiments, the subject suffers from a disease or condition (e.g., a genetic disease). In some embodiments, the disease or condition is cardiovascular disease. In some embodiments, the composition is administered to the subject's liver tissue, heart tissue, or skeletal muscle tissue.

[0035] Other aspects of the present disclosure provide methods of making any of the disclosed rAAV particles and compositions.

[0036] The details of certain embodiments of the invention are set forth in the detailed description of certain embodiments, as described below. Other features, objects and advantages of the invention will be apparent from the definitions, examples, drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figures 1A-1C . Figure 1A AAV constructs evaluated in vivo are shown. An sgRNA targeting the Pcsk9 gene with a W8 mutation was delivered along with EGFP in one AAV, which was co-injected with one or two additional AAVs encoding intact SaABE8e or split-intein SaABE8e, respectively. A total of two AAVs were used to deliver intact SaABE8e and the sgRNA, and three AAVs were used to deliver split-intein SaABE8e. Black boxes represent ITRs, the EFS promoter is a short EF1α, W3 is a truncated WPRE, bGH is the bovine growth hormone polyadenylation signal, the purple box is the sgRNA cassette driven by the U6 promoter in the direction indicated by the arrow, the NpuN and NpuC split inteins from Nostoc punctata are shown in brown, and the protein-coding regions are indicated for EGFP and SaABE8e. Figure 1BFigure 3. In vivo editing efficiencies of AAVs encoding intein-cleaved SaABE8e and intact SaABE8e injected. The total dose of base editor AAVs administered to each mouse is shown. Figure 1C Figure 5 shows the in vivo editing efficiency of AAV9 encoding the complete SaABE8e injected in five different AAV architectures when administered at the indicated doses. In all cases, the editor AAV dose was 4x10 11 vector genomes (vg) or 4x10 10 The dose of AAV for vg and sgRNA EGFP is 4x10 11 VG or 4x10 10 vg, with a 1:1 ratio of base editor AAV to sgRNA AAV. The sizes of the delivered editor AAV constructs (including ITRs) are shown in the figure legend. 6-7 week old C57BL / 6J mice weighing 20-25 g were injected systemically via retroorbital injection. Dots represent individual mice. Values and error bars represent the mean ± SEM of n = 3 different mice.

[0038] Figures 2A-2C Development and characterization of a single AAV SaABE8e. Figure 2A A schematic diagram of a single AAV SaABE8e genome (5,064 bp, including ITRs) is shown. The arrow indicates the direction of the U6 sgRNA cassette. Figure 2B Figure 3 shows a comparison of dual SaABE8e and single SaABE8e. Base editing activity of dual AAV vectors (with the complete editor on one genome and SaABE8e with sgRNA and EGFP on the second genome) or single AAV vectors (SaABE8e and guide RNA in a single vector), both of which installed Pcsk9 W8R editing and were packaged in AAV9 with matching promoters and terminators (poly A). Base editor AAVs were administered via retroorbital injection at the doses indicated in the figure legend (dual AAVs were delivered at the indicated doses of editor AAV and sgRNA AAV, while single AAVs were delivered at the indicated doses) to 6- to 8-week-old C57BL / 6J mice weighing 20 to 25 grams, and tissues were harvested three weeks after injection and analyzed by high-throughput DNA sequencing (HTS). Figure 2C Shown is a dose titration of a single AAV SaABE8e. Points represent individual mice and error bars represent mean ± SEM of n=3 different mice.

[0039] Figures 3A-3D Nme2ABE8e( Figure 3A )、CjABE8e( Figure 3B ) and SauriABE8e( Figure 3C) characterization. Editing of each target adenine within the protospacer is shown. The target sites are indicated, with the sequences of each target protospacer and PAM listed in Table 1. Target adenines are numbered relative to the standard protospacer length for each editor (22 nucleotides for CjABE8e, 24 nucleotides for Nme2ABE8e, and 21 nucleotides for SauriABE8e). Dots represent values, and error bars represent mean ± SEM of n = 3 replicates. Figure 3D The percentage of genomic adenines in the hg38 human reference genome that can be targeted individually and collectively by size-minimized ABEs is shown. A schematic diagram of a representative portion of the genome that can be targeted by size-minimized ABEs is shown, with targetable PAMs represented by colored lines and targetable adenines represented by colored dots. The pie chart on the far right of the schematic diagram indicates that 87% of the single nucleotide adenine polymorphisms are targetable by all small AAV-encoded ABEs of the present disclosure (and 82% are editable).

[0040] Figures 4A-4D Nme2ABE8e( Figure 4A )、CjABE8e( Figure 4B ) and SauriABE8e( Figure 4C ). The base editing activity window for 8-9 genomic target sites of each size-minimized ABE is shown. Positions that are not present in any of the tested sites are shaded in gray. The target adenines are numbered relative to the standard protospacer length of each editor (24 nt for Nme2ABE8e, 22 nt for CjABE8e, and 21 nt for SauriABE8e). The dots represent the values, and the error bars represent the mean ± SEM of n = 3 replicates per site, where each position represents 1-3 genomic sites. The dotted line corresponds to 25% of the average peak activity of each base editor and defines the activity threshold considered within the editing window. Single sites containing three or more adenines within the protospacer were included in this analysis. Sites with only one or two adenines within the protospacer were excluded from the pooled analysis, but data from all analyzed sites are shown in Figures 3A-3C middle. Figure 4D Shown are the percentages of genomic adenines in the hg38 human reference genome that are targetable by size-minimized ABEs, either individually or collectively. An example of a representative portion of the human genome that is targetable by size-minimized ABEs is shown, with targetable PAMs represented by colored lines and targetable adenines represented by colored dots.

[0041] Figures 5A-5K Genome editing and plasma lipids were evaluated when PCSK9, Pcsk9, and Angptl3 were targeted in vivo with a single AAV. Figure 5A The strategy used to evaluate base editing and plasma analytes in AAV-treated mice is shown and was created using BioRender. Figure 5B Figure 3 shows bulk liver editing efficiencies for human PCSK9 and mouse Pcsk9 and Angptl3 (n = 3-5 mice, each dot represents one mouse, error bars represent SEM). Human PCSK9 editing was performed using humanized PCSK9 mice, while mouse Pcsk9 and Angptl3 editing was performed at the endogenous mouse locus in wild-type C57BL / 6J mice. AAV was administered by retroorbital injection at 6-8 weeks of age at a dose of 1x10 11 vg / mouse. Figure 5C Dose-dependent base editing of dual SpABE8e and single SaKKH-ABE8e at the mouse Pcsk9 exon 1 splice donor is shown. The total AAV dose administered is indicated below each line in vg / mouse. The total AAV dose administered is indicated below each line in vg / mouse. The dual SpABE8e editing data were reported in another publication from our group. 2 Each dot represents a different mouse (n=5). Figure 5D A direct comparison of the editing efficiency of double AAV8 intein-split SaKKH-ABE8e and single AAV8 SaKKH-ABE8e targeting the Pcsk9 exon 1 donor site in bulk liver at two doses is shown. The total AAV dose administered is indicated below each group line in vg / mouse, ***P=0.0004. Each dot represents a different mouse (n=4). Figure 5E Shown with 1x10 11 Plasma PCSK9 protein in humanized mice treated with single AAV8 SaKKH-ABE8e. Figure 5F Shown with 1x10 11 vg Plasma Pcsk9 protein in C57BL / 6J mice treated with single AAV8 SaKKH-ABE8e, dual AAV8 SpABE8e, or non-targeting control, ***P = 0.0001. Figure 5G Shown with 1x10 11 Plasma Angptl3 protein in C57BL / 6J mice treated with single AAV8 SaKKH-ABE8e or non-targeting control, **P=0.0027. Figure 5H Shown with 1x10 11 vg Plasma total cholesterol in humanized mice treated with single AAV8 SaKKH-ABE8e. Figure 5I Shown with 1x10 11vg Plasma total cholesterol in C57BL / 6J mice treated with single AAV8 SaKKH-ABE8e, dual AAV8 SpABE8e, or non-targeting control, ***P = 0.0007. Figure 5J Shown are plasma total cholesterol ***P=0.0007 and plasma triglycerides *P=0.0118. Figure 5K Shown with 1x10 11 vg single AAV8 SaKKH-ABE8e or non-targeting control treated C57BL / 6J mice. Figures 5E-5K , Dots represent mean values, and error bars represent SEM of n = 5 different mice. Figures 5C-5D The significance was calculated using a two-way unpaired t-test. Figures 5F-5G and Figures 5I-5K The significance of the results was calculated using two-way repeated measures ANOVA with Tukey or Šídák multiple comparisons (as appropriate), and the results are shown except for Figure 5F The week 4 time point is shown for all figures except for Figure 5 , where significance at week 3 is shown because protein levels did not reach statistical significance at week 4. In all cases, the non-targeting control was dual AAV8 SpABE7.10 with sgRNA targeting mouse Dnmt1 (an irrelevant site in the mouse genome), administered at the same time point, route, and dose.

[0042] Figure 6 Validation of SaABE targets in mouse Neuro-2A and 3T3 cells is shown. The base editor and PAM are indicated below each set of lines. Dots represent independent biological replicates (n = 3), and error bars show SEM.

[0043] Figure 7 Shown is the titration of sgRNA cassette AAV in vivo. 4x10 11 A constant dose of the full-length SaABE8e editor AAV containing vg was delivered to C57BL / 6J mice via retroorbital injection. Tissues were harvested three weeks after injection and analyzed by HTS. Dots represent individual mice (n = 3), and error bars show SEM.

[0044] Figure 8A and 8B . Figure 8A Figure 3. SaABE8e activity window at Pcsk9 W8R in liver. SaABE8e maintains a wide in vivo editing window, consistent with observations in cultured cells. Figure 8B It showed that indels remained low under all conditions, with a 11A high dose of vg single AAV SaABE8e resulted in 2.4%, 1.6%, and 1.1% indels in heart, muscle, and liver tissues, respectively. Dots represent individual mice (n=3), and error bars indicate SEM.

[0045] Figures 9A-9C Validation of the guide targeting PCSK9 in HEK293T cells ( Figure 9A ) and validated Pcsk9 in mouse Neuro-2a cells ( Figure 9B ) and Angptl3( Figure 9C For each sgRNA, the exon, target type (start codon, splice donor, or splice acceptor), ABE8e variant (SaABE8e or SaKKH-ABE8e), and protospacer position at which the indicated target was disrupted relative to the 22 nt protospacer length are indicated. Edits at the position at which the indicated target was disrupted are plotted. Dots represent independent replicates (n = 2), and error bars show standard deviation (SD).

[0046] Figure 10A and 10B Editing at the Pcks9 exon 1 splice donor target using SauriABE8e in Neuro-2A cells is shown. Figure 10A Targeting of the mouse Pcsk9 exon 1 splice donor with SauriABE8e and SaKKH-ABE8e in mouse Neuro-2A cells is shown. The target adenine is A9 relative to the Sauri protospacer and A5 relative to the SaKKH protospacer. Figure 10B SaCas9 guide RNA scaffold shown 33,74 Comparison of editing activity at the mouse Pcsk9 exon 1 splice donor. The SaCas9 sgRNA scaffold was used with the cognate SauriCas9 protein, as the native sgRNA for SauriCas9 is unknown.

[0047] Figure 11A and 11B In vivo editing of control AAV for lipid modification experiments. Six- to eight-week-old C57BL / 6J mice were injected via retroorbital injection, and whole livers were analyzed by HTS four weeks later. Figure 11A Shown is the use of dual SpABE7.10 at 1x10 11 vg double dose of AAV8 editing of Dnmt1 A41A (silent editing). Figure 11B shows the use of single SaKKH-ABE8e with 1x10 11 vg single dose of AAV8 to install the Pcsk9 W8R replacement.

[0048] Figures 12A-12DDose responses of single AAV8 SaKKH-ABE8e and dual AAV8 SpABE8e (with sgRNA targeting the Pcsk9 exon 1 splice donor site) on plasma Pcsk9 and total cholesterol. Figure 12A Circulating Pcsk9 protein is shown, Figure 12B Total cholesterol in plasma collected weekly is shown, normalized to baseline. Figure 12C Circulating Pcsk9 protein is shown, Figure 12D Total cholesterol in plasma collected weekly is shown, which is raw (unnormalized). Dots represent the mean, and error bars represent the SEM of n=5 different mice. All mice were systemically administered the total dose of AAV8 indicated in the legend via retroorbital injection at 6-8 weeks of age, and blood samples were drawn continuously for four weeks.

[0049] Figures 13A-13E Raw (unnormalized) levels of plasma analytes in mice with single AAV ABE or non-targeting control double AAV ABE for human PCSK9 and mouse Angptl3 targets. Figure 13A ELISA of human PCSK9 in plasma from humanized mice is shown. Figure 13B Total plasma cholesterol in humanized PCSK9 mice is shown. Figure 13C Shown is an ELISA of mouse Angptl3 in plasma from C57BL / 6J mice. Figure 13D Shown are total plasma cholesterol in C57BL / 6J mice. Figure 13E Plasma triglycerides from C57BL / 6J mice are shown. Dots represent mean values, and error bars represent SEM for n=5 different mice. A non-targeting control was a dual AAV ABE7.10 targeting Dnmt1. All mice were systemically administered 1x10 11 vg AAV8 dose, and serial blood samples were drawn over four weeks.

[0050] Figure 14 Figure 1 is a schematic diagram showing the construction of a single AAV vector expressing a cytosine base editor that includes a uracil glycosylase domain. The Cas9 domain of the cytosine base editor is CjCas9. The U6-controlled sgRNA is located at the 3' end in the opposite orientation. The total length of the construct, including the ITRs, is 5.012 kb.

[0051] Figure 15 Shown is a comparison of base editor matches for guide-dependent on-target editing between single and dual AAV SaKKH-ABE8e at the Pcsk9 exon 1 splice donor site.

[0052] Figures 16A-16C Shown are guide-dependent off-target DNA editing analyses in vivo and in culture between single AAV8 SaKKH and each half of AAV8 intein-split SaKKH-ABE8e. Figure 16A In vivo editing in liver tissue from single AAV ABE-treated mice is shown. The top three predicted off-target ("OT") sites for SaKKH-ABE8e targeting the Pcsk9 exon 1 splice donor were sequenced from liver tissue. Figure 16B Figure 3. Editing observed at OT2 is dose-dependent. OT, off-target; NT, non-targeting dose-matched dual AAV8 ABE7.10 targeting Dnmt1 (an unrelated gene). Figure 16C Editing in cell culture is shown. After N2A cells were transfected with an sgRNA targeting the Pcsk9 exon 1 donor site and full-length or intein-split SpABE8e plasmids, on-target and off-target sites were sequenced. No significant differences in the efficiency of on-target or off-target editing between full-length SpABE8e and intein-split SpABE8e were found using multiple unpaired t-tests and the Holm-Šidák method, corrected for multiple comparisons.

[0053] Figure 17 Figure 5: No detectable in vivo mRNA off-target editing in mice treated with AAV ABE8e alone. Dots represent individual adenines per amplicon from n = 4 mice.

[0054] Figure 18 Shown is a comparison of the editing windows of exemplary single AAV-encoded ABEs in the liver.

[0055] Figures 19A-19B The livers from untreated mice ( Figure 19A ) and use 1x10 11 Four weeks after treatment with a single AAV8 SaKKH-ABE8e targeting human PCSK9 ( Figure 19B ) for histopathological evaluation. Representative images are shown. Scale bar, 50 µm.

[0056] Figure 20Quantification of AAV genomes from tissues encoding SaABE8e dual AAV (SaABE8e with the complete editor on one genome and sgRNA and EGFP on the second genome) or single AAV (SaABE8e and guide RNA integrated), both installed with Pcsk9 W8R packaged in AAV9. The editors were packaged in AAV9 and administered by retroorbital injection at 6-8 weeks of age at the doses indicated in the legend (dual AAVs were delivered at the doses indicated in the legend for editor AAV and sgRNA AAV, while single AAVs were delivered at the indicated total doses). Tissues were harvested 3 weeks after injection and editor AAV genomes were quantified by ddPCR using SaCas9 primers and probes, normalized to Gapdh. Dots represent individual mice (n=1-3) and error bars show SEM.

[0057] Figure 21 Alkaline gel electrophoresis of packaged AAV genomes is shown.

[0058] Figures 22A-22D Off-target mRNA editing is shown. RNA was extracted from the livers of mice treated with a single AAV8 SaKKH-ABE8e and untreated mice, reverse transcribed, and cDNA amplicons from Aars, Canx, Ctnnb, and Usp38 mRNAs were analyzed by HTS. Dots indicate individual adenines on the sequenced amplicons (n = 3 mice). Detailed Description of the Invention

[0060] Provided herein are methods and compositions for delivering base editor proteins to cells or tissues in a single recombinant AAV (rAAV) vector. Contemplated herein are improved methods and compositions for delivering these base editors in vivo (e.g., to subjects in need thereof) in a single rAAV particle. Delivery in a single rAAV particle has the following advantages: fewer injections and lower doses are required for delivery to the target tissue, and the number of successful transductions of the target tissue required for expressing the base editor in the target cell is reduced by half. These rAAV vectors contain base editors and regulatory elements with minimized size, which enable the vector to have a length within the 4.7 kb-4.9 kb packaging capacity of the rAAV particles. Also provided are rAAV particles containing any disclosed rAAV vector and capsid protein, and compositions and cells comprising the same. Further provided are methods for administering such compositions and cells to a subject. Further provided are base editors and compositions and cells comprising these base editors. The disclosed single AAV adenine base editor provides comparable or enhanced editing efficiency in a variety of tissues in vivo compared to the double AAV editor.

[0061] Recombinant AAV vectors (or AAV genomes) are widely used for transgene delivery. Transgenes are inserted into the AAV genome between inverted terminal repeat (ITR) sequences and packaged into AAV viral particles, which are used to transduce host cells (e.g., mammalian cells, human cells). AAV has been used to deliver genes encoding many therapeutic proteins in animal models of human diseases, in clinical trials, and in FDA-approved drugs. A set of available AAV serotypes provides approaches to obtain multiple clinically relevant cell types in mice, non-human primates, and humans.

[0062] The present disclosure provides rAAV vectors with minimized size regulatory elements, which allow packaging of transgenes that are larger than those of the prior art. In some embodiments, the transgenic encodes a base editor, such as an adenine base editor. In some embodiments, the transgenic is not a base editor. In some embodiments, the base editor contains a napDNAbp domain, which is a compact protein, such as Staphylococcus aureus Cas9 (SaCas9), Neisseria meningitidis 2 Cas9 (Nme2Cas9), Campylobacter jejuni Cas9 (CjCas9) or Staphylococcus auris (SauriCas9) domain or variants thereof.

[0063] In some aspects, provided herein is an rAAV vector comprising a first nucleic acid segment comprising: (i) a 5'ITR; (ii) a first nucleic acid segment comprising a sequence encoding a base editor operably linked to a first promoter, wherein the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain; and a polyadenylation (poly A) signal; (iii) a second nucleic acid segment encoding a guide RNA (gRNA) operably linked to a second promoter; (iv) a 3'ITR, wherein the length between the 5'ITR and the 3'ITR is less than about 4.90 kb. In some embodiments, the rAAV vector consists essentially of components (i)-(iv).

[0064] In some embodiments, the nucleic acid vector is a genome of an adeno-associated virus packaged in rAAV particles. In some embodiments, the first and / or second nucleic acid fragment is operably linked to a first promoter. In some embodiments, the first promoter is a constitutive promoter. In some embodiments, the first promoter is an inducible promoter.

[0065] In various embodiments, the first promoter is a short promoter. In various embodiments, the first promoter has a length of less than 325 nucleotides, less than 300 nucleotides, less than 285 nucleotides, less than 270 nucleotides, less than 265 nucleotides, or less than 250 nucleotides. This short length ensures the optimal packaging capacity of the base editor and additional regulatory elements. In some embodiments, the first promoter has a length of 200 to 250 nucleotides, 225 to 255 nucleotides, 250 to 275 nucleotides, 275 to 300 nucleotides, or 300 to 325 nucleotides. In certain embodiments, the first promoter has a length of 280 nucleotides. In some embodiments, the first promoter has a length of 229 nucleotides, 253 nucleotides, or 324 nucleotides.

[0066] In some embodiments, the first promoter is a tissue-specific promoter. The first promoter can be a cardiac tissue-specific promoter, a muscle tissue-specific promoter, or a neuronal tissue-specific promoter. The first promoter can be active in tissues other than liver, muscle, and neurons (e.g., eye tissue). In some embodiments, the first promoter is active in neuromuscular tissue.

[0067] In some embodiments, the first promoter is the EF-1α (short) ("EFS") promoter, which is an intronless form of EF-1α. In some embodiments, the first promoter is the MeCP2 promoter, which is active in neuronal tissue. In some embodiments, the first promoter is the P3 promoter, which is active in liver tissue. In some embodiments, the first promoter is the U1a promoter, which is active in liver tissue. Additional details about the MeCP2 promoter, the P3 promoter, and the U1a promoter can be found in Gray et al., Human Gene Therapy. Sep 2011. 1143-1153; Viecelli et al., Hepatology, 60: 1035-1043 (2014); and Ibraheim et al., Genome Biol. 19, 137 (2018), each of which is incorporated herein by reference.

[0068] In some embodiments, the first nucleic acid fragment comprises a transcription terminator. In various embodiments, the first nucleic acid fragment does not contain a post-transcriptional response element (e.g., W3). In some embodiments, the transcription terminator is a poly A signal selected from a bovine growth hormone (bGH) signal, a human growth hormone (hGH) signal, or an SV40 signal. In some embodiments, the terminator is a bGH poly A signal. In some embodiments, the terminator is an SV40 late poly A signal.

[0069] In some embodiments, the first nucleic acid segment comprises a minute virus of mice (MVM) intron. In some embodiments, the MVM is located 5' to the promoter and 3' to the transgene.

[0070] In some embodiments, the second nucleic acid fragment comprises a nucleotide sequence encoding a gRNA operably linked to a second promoter. In some embodiments, the second promoter is a constitutive promoter. In some embodiments, the second promoter is an inducible promoter. In some embodiments, the second promoter is a U6 promoter, such as a human U6 promoter. In some embodiments, the transcription direction of the second nucleic acid fragment is reversed relative to the transcription direction of the first nucleic acid fragment.

[0071] The disclosed rAAV vectors - which comprise a first nucleic acid segment comprising a promoter and a terminator and a second nucleic acid segment encoding a guide RNA - have a packaging capacity between a 5' ITR and a 3' ITR of length suitable for encoding a transgene of one or more disclosed base editors. For example, the disclosed AAV vectors may contain a length between a 5' ITR and a 3' ITR of between 4.7 kb and 4.9 kb. The disclosed AAV vectors may contain a length between a 5' ITR and a 3' ITR of between 4.7 kb and 5.1 kb. The disclosed AAV vectors may contain a length between a 5' ITR and a 3' ITR of between 4.6 kb and 4.9 kb or between 4.6 kb and 4.8 kb. In certain embodiments, the length between the 5' ITR and the 3' ITR is about 4.60 kb, about 4.65 kb, about 4.70 kb, about 4.725 kb, about 4.75 kb, about 4.80 kb, about 4.825 kb, about 4.85 kb, about 4.90 kb, or about 4.95 kb.

[0072] In various embodiments, the length between the 5' ITR and the 3' ITR is less than about 4.90 kb. In various embodiments, the length between the 5' ITR and the 3' ITR is about 4.80 kb. In some embodiments, the length between the ITRs is 4.804 kb (4804 bp). In some embodiments, the length between the ITRs is 4.828 kb. In some embodiments, the length between the ITRs is 4.722 kb.

[0073] The present disclosure provides rAAV vectors containing adenine base editors with minimized size and rAAV vectors containing cytosine base editors with minimized size. Figure 2A and Figure 14 middle.

[0074] In some aspects, provided herein are rAAV vectors comprising: (i) a 5' ITR; (ii) a first nucleic acid segment comprising a sequence encoding a SaKKH-ABE8e, SauriCas9-ABE8e, CjCas9-ABE8e base editor, or Nme2Cas9 base editor operably linked to a first promoter, wherein the first promoter is selected from the group consisting of EFS, MeCP2, P3, and U1A promoters; and a bGH polyadenylation (poly A) signal; (iii) a second nucleic acid segment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the direction of transcription of the second nucleic acid segment is reverse relative to the direction of transcription of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the rAAV vector contains a first nucleic acid segment comprising: (i) a 5'ITR; (ii) a first nucleic acid segment comprising a Sauri-ABE8e base editor encoding an EFS promoter; and a polyadenylation (poly A) signal; (iii) a second nucleic acid segment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the transcription direction of the second nucleic acid segment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) a 3'ITR. In various embodiments, the length between the 5'ITR and the 3'ITR is less than about 4.90 kb. In some embodiments, the poly (A) signal is a bGH poly (A) signal.

[0075] In certain embodiments, the rAAV vector comprises, from 5' to 3': (i) a 5' ITR; (ii) a first nucleic acid segment comprising a sequence encoding a SaKKH-ABE8e base editor operably linked to an EFS promoter; and a bGH polyadenylation (poly A) signal; (iii) a second nucleic acid segment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the transcription direction of the second nucleic acid segment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the rAAV vector contains a first nucleic acid segment comprising: (i) a 5' ITR; (ii) a first nucleic acid segment comprising a sequence encoding a Sauri-ABE8e base editor operably linked to an EFS promoter; and a bGH polyadenylation (poly A) signal; (iii) a second nucleic acid segment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the transcription direction of the second nucleic acid segment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the base editor is SauriCas9-ABE8e. In some embodiments, the base editor is CjCas9-ABE8e.

[0076] In some embodiments, the rAAV vector encodes CBE and comprises, from 5' to 3', (i) a 5' ITR; (ii) a first nucleic acid fragment comprising a sequence encoding a CjCas9-FERNY-BE3.9 or CjCas9-evoFERNY-BE3 base editor operably linked to a first promoter, the first promoter being an EFS promoter; and a polyadenylation (poly A) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the transcription direction of the second nucleic acid fragment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the rAAV vector encodes CBE and comprises, from 5' to 3': (i) a 5' ITR; (ii) a first nucleic acid segment comprising a sequence encoding a CjCas9-FERNY-BE3.9 or CjCas9-evoFERNY-BE3 base editor operably linked to a first promoter, wherein the first promoter is selected from EFS, MeCP2, P3, and U1A promoters; and a polyadenylation (poly A) signal; (iii) a second nucleic acid segment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the transcription direction of the second nucleic acid segment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) a 3' ITR. In some embodiments, the poly (A) signal is a bGH poly (A) signal. In some embodiments, the poly (A) signal is an SV40 poly (A) signal.

[0077] In various embodiments, any of the disclosed rAAV vectors are encapsulated in an AAV8 capsid. In some embodiments, the disclosed rAAV vectors are encapsulated in an AAV9 capsid.

[0078] Further provided herein are empirical testing of regulatory elements in the disclosed AAV vectors for high expression levels of the encoded base editors.

[0079] definition

[0080] As used herein and in the claims, the singular forms "a," "an," and "the" include both the singular and the plural unless the context clearly dictates otherwise. Thus, for example, reference to "an agent" includes a single agent and a plurality of such agents.

[0081] “Adeno-associated virus” or “AAV” is a virus that infects humans and some other primate species. The wild-type AAV genome is single-stranded deoxyribonucleic acid (ssDNA) and can be either positive-sense or negative-sense. The genome contains two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs): rep and cap, located between the ITRs. The rep ORF contains four overlapping genes encoding the Rep proteins required for the AAV life cycle. The cap ORF contains overlapping genes encoding the capsid proteins: VP1, VP2, and VP3, which interact together to form the viral capsid. VP1, VP2, and VP3 are translated from a single mRNA transcript that can be spliced in two different ways: either a longer or shorter intron can be excised, resulting in two mRNA isoforms: one ~2.3 kb and one ~2.6 kb long. The capsid forms a supramolecular assembly composed of approximately 60 individual capsid protein subunits, forming a non-enveloped T-1 icosahedral lattice that protects the AAV genome. The mature capsid contains VP1, VP2, and VP3 (molecular weights of approximately 87 kDa, 73 kDa, and 62 kDa, respectively) in a ratio of approximately 1:1:10.

[0082] The rAAV particle may comprise a nucleic acid vector (e.g., a recombinant genome) that may comprise at least: (a) one or more heterologous nucleic acid regions (or transgenes) comprising a sequence encoding a protein or polypeptide of interest (e.g., a base editor) or an RNA of interest (e.g., a gRNA); and (b) one or more regions comprising inverted terminal repeat (ITR) sequences (e.g., a wild-type ITR sequence or an engineered ITR sequence), flanked by one or more nucleic acid regions (e.g., heterologous nucleic acid regions). In some embodiments, the nucleic acid vector is 4 kb to 5 kb in size (e.g., 4.2 to 4.7 kb in size). In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, the double-stranded nucleic acid vector can be, for example, a self-complementary vector containing a region of the nucleic acid vector that is complementary to another region of the nucleic acid vector, thereby initiating the formation of double-stranded nucleic acid vectors.

[0083] As used herein, the term "adenosine deaminase" or "adenosine deaminase domain" refers to a protein or enzyme that catalyzes the deamination reaction of adenosine (or adenine). These terms can be used interchangeably. In certain embodiments, the present disclosure provides a base editor comprising one or more adenosine deaminase domains. For example, the adenosine deaminase domain may comprise a heterodimer of a first adenosine deaminase and a second deaminase domain connected by a linker. The adenosine deaminase provided herein (e.g., an engineered adenosine deaminase or an evolved adenosine deaminase) can be an enzyme that converts adenine (A) in DNA or RNA into inosine (I). Such adenosine deaminase can result in a base pair conversion of A:T to G:C. In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase does not exist in nature. For example, in some embodiments, the deaminase is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase.

[0084] In some embodiments, the adenosine deaminase is derived from a bacterium, e.g., Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, or Caulobacter crescentus. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is an Escherichia coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated Escherichia coli TadA deaminase. For example, the truncated ecTadA may lack one or more N-terminal amino acids relative to the full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length ecTadA. In some embodiments, the truncated ecTadA can lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length ecTadA. In some embodiments, the ecTadA deaminase does not contain an N-terminal methionine. Reference is made to U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018, which is incorporated herein by reference.

[0085] In genetics, the "antisense" strand of a fragment in double-stranded DNA is a template strand, and is considered to run in a 3' to 5' direction. In contrast, a "sense" strand is a fragment running from 5' to 3' in double-stranded DNA, and is complementary to the antisense or template strand of the DNA running from 3' to 5'. In the case of a protein-encoding DNA fragment, the sense strand is a DNA strand with the same sequence as mRNA, which uses the antisense strand as its template during transcription and ultimately undergoes (usually, not always) translation into protein. Therefore, the antisense strand is responsible for the RNA that is subsequently translated into protein, while the sense strand has a composition almost identical to that of mRNA. Note that for each fragment of dsDNA, there may be two groups of sense strands and antisense strands, depending on the direction of reading (because sense strands and antisense strands are related to viewing angle). Ultimately, it is the gene product or mRNA that determines which strand of a fragment of dsDNA is called the sense strand or antisense strand.

[0086] "Base editing" refers to a genome editing technique that involves converting a specific nucleic acid base into another at a targeted genomic locus. In certain embodiments, this can be achieved without the need for double-stranded DNA breaks (DSBs) or single-strand breaks (i.e., nicks). So far, other genome editing techniques, including systems based on CRISPR, have all started by introducing DSBs at the locus of interest. Subsequently, cellular DNA repair enzymes patch the break, typically resulting in random insertions or deletions (indels) of the bases at the DSB site. However, when it is desired to introduce or correct point mutations at the target locus rather than randomly destroy the entire gene, these genome editing techniques are unsuitable because the correction rate is low (e.g., typically 0.1% to 5%), and the main genome editing product is an indel. In order to improve the efficiency of gene correction when random indels are not introduced simultaneously, the inventors have previously modified the CRISPR / Cas9 system to directly convert one DNA base into another without forming a DSB. See, Komor, AC, et al., Programmable editing of atarget base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference.

[0087] As used herein, the term "base editor (BE)" refers to a polypeptide that can modify a base (e.g., A, T, C, G, or U) by converting one base to another (e.g., A to G, A to C, A to T, C to T, C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G) within a nucleic acid sequence (e.g., DNA or RNA). In some embodiments, a base editor is capable of deaminating a base within a nucleic acid (such as a base within a DNA molecule). In the case of an adenine base editor, a base editor is capable of deaminating adenine (A) in DNA. Such a base editor can include a nucleic acid programmable DNA binding protein (napDNAbp) fused to an adenosine deaminase. Some base editors include CRISPR-mediated fusion proteins used in the base editing methods described herein. In some embodiments, the base editor comprises a nuclease-inactivated Cas9 (dCas9) fused to a deaminase that binds to nucleic acids in a guide RNA-programmed manner via the formation of an R-loop, but does not cut nucleic acids. For example, the dCas9 domain of the fusion protein can include D10A and H840A mutations (which enable Cas9 to cut only one strand of the nucleic acid duplex), as described in PCT / US2016 / 058344, which was published as WO 2017 / 070632 on April 27, 2017, and is incorporated herein by reference in its entirety. The DNA cleavage domain of Streptococcus pyogenes Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the chain complementary to the gRNA (the "targeting chain," or the chain in which editing or deamination occurs), while the RuvC1 subdomain cuts the non-complementary chain containing the PAM sequence (the "non-editing chain"). The RuvC1 mutant D10A nicks the targeted strand, while the HNH mutant H840A nicks the non-edited strand (see Jinek et al., Science, 337:816-821 (2012); Qi et al., Cell. 28;152(5):1173-83 (2013), each of which is incorporated herein by reference).

[0088] In some embodiments, a base editor is a macromolecule or macromolecular complex that primarily (e.g., greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 99%, greater than 99.9%, or 100%) results in the conversion of a nucleobase to another nucleobase (i.e., a transition or transversion) in a polynucleic acid sequence using a combination of 1) a nucleotide, nucleoside, or nucleobase modifying enzyme and 2) a nucleic acid binding protein that can be programmed to bind to a specific nucleic acid sequence.

[0089] In some embodiments, the base editor comprises a DNA binding domain (e.g., a programmable DNA binding domain, such as dCas9 or nCas9), which guides the base editor to the target sequence. In some embodiments, the base editor comprises a nucleobase modifying enzyme fused to a programmable DNA binding domain (e.g., dCas9 or nCas9). A "nucleobase modifying enzyme" is an enzyme that can modify a nucleobase and convert one nucleobase into another (e.g., a deaminase such as adenosine deaminase). Base editors that perform certain types of base conversions (e.g., adenosine (A) to guanine (G), C to G) are contemplated.

[0090] In some embodiments, the base editor converts A to G. In some embodiments, the base editor comprises an adenosine deaminase. "Adenosine deaminase" is an enzyme involved in purine metabolism. It is required for the breakdown of adenosine from food and the turnover of nucleic acids in tissues. Its main function in the human body is the development and maintenance of the immune system. Adenosine deaminase catalyzes the hydrolytic deamination of adenosine in the context of DNA (forming inosine, whose base pairing is G). There are no known natural adenosine deaminases that act on DNA. In contrast, known adenosine deaminases only act on RNA (tRNA or mRNA).Evolved deoxyadenosine deaminases that accept a DNA substrate and deaminate dA to deoxyinosine have been described, for example, in PCT Application Nos. PCT / US2017 / 045381, filed August 3, 2017 (published as WO 2018 / 027078), and PCT Application No. PCT / US2019 / 033848, filed May 23, 2019 (published as WO 2018 / 027078). 2019 / 226953), U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018 (which issued as U.S. Patent No. 10,113,163); U.S. Patent Publication No. 2017 / 0121693, published on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published on May 4, 2017, which issued as U.S. Patent No. 10,167,457 on January 1, 2019; PCT Publication No. WO 2017 / 050691, published on April 27, 2017 2017 / 070633; U.S. Patent Publication No. 2015 / 0166980, published on June 18, 2015; U.S. Patent No. 9,840,699, issued on December 12, 2017; U.S. Patent No. 10,077,453, issued on September 18, 2018; PCT Publication No. WO 2019 / 023680, published on January 31, 2019; International Application No. PCT / US2019 / 033848, filed on May 23, 2019, published as Publication No. WO 2019 / 226593 on November 28, 2019; PCT Publication No. WO 2018 / 0176009, published on September 27, 2018, and PCT Publication No. WO 2019 / 023680, published on January 31, 2019; International Application No. PCT / US2019 / 033848, filed on May 23, 2019, published as Publication No. WO 2019 / 226593 on November 28, 2019; PCT Publication No. WO 2018 / 0176009, published on September 27, 2018, and PCT Publication No. WO 2019 / 023680, published on February 27, 2020. 2020 / 041751; PCT Publication No. WO 2020 / 051360 published on March 12, 2020; PCT Publication No. WO 2020 / 102659 published on May 22, 2020; PCT Publication No. WO 2020 / 086908 published on April 30, 2020; PCT Publication No. WO 2020 / 181180 published on September 10, 2020; PCT Publication No. WO 2020 / 214842 published on October 22, 2020; PCT Publication No. WO 2020 / 092453 published on May 7, 2020; PCT Publication No. WO 2020 / 236982 published on November 26, 2020; PCT Publication No. WO 2020 / 236982 published on June 3, 2021 2021 / 108717 and PCT Publication No. WO 2021 / 158921 published on August 12, 2021, the contents of each of which are incorporated herein by reference in their entirety.

[0091] In some embodiments, the base editor converts C to T. In some embodiments, the base editor comprises a cytidine deaminase. "Cytosine deaminase" or "cytidine deaminase" refers to an enzyme that catalyzes the chemical reaction "cytosine + H2O Uracil + NH3" or "5-methyl-cytosine + H2O Thymine + NH3". As is apparent from the reaction formula, such chemical reactions result in a C to U / T nucleobase change. In the case of genes, such nucleotide changes or mutations can in turn result in amino acid changes in the protein, which can affect the function of the protein, e.g., loss of function or gain of function. In some embodiments, the cytosine base editor comprises dCas9 or nCas9 fused to a cytidine deaminase. In some embodiments, the cytidine deaminase domain is fused to the N-terminus of dCas9 or nCas9. In some embodiments, the base editor further comprises a domain that inhibits uracil glycosylase and / or a nuclear localization signal. Such base editors have been described in the art, e.g., Rees & Liu, Nat Rev Genet. 2018;19(12):770-788, Rees, et al. Sci. Advances5, eaax5717 (2019), and Koblan et al., Nat Biotechnol. 2018;36(9):843-846; and U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018, which issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published on May 4, 2017, which issued as U.S. Patent No. 10,167,457 on January 1, 2019; PCT Publication No. WO 2017 / 0121693, published on April 27, 2017 2017 / 070633; U.S. Patent Publication No. 2015 / 0166980, published on June 18, 2015; U.S. Patent No. 9,840,699, issued on December 12, 2017; U.S. Patent No. 10,077,453, issued on September 18, 2018; PCT Publication No. WO2019 / 023680, published on January 31, 2019; PCT Publication No. WO 2018 / 0176009, PCT Application No. PCT / US2019 / 033848 filed on May 23, 2019, PCT Application No. PCT / US2019 / 47996 filed on August 23, 2019; PCT Application No. PCT / US2020 / 028568 filed on April 17, 2020; PCT Application No. PCT / US2019 / 61685 filed on November 15, 2019; PCT Application No. PCT / US2019 / 57956 filed on October 24, 2019; PCT Publication No. PCT / US2019 / 58678 filed on October 29, 2019; and PCT Publication No. WO 2021 / 108717 published on June 3, 2021, the contents of each of which are incorporated herein by reference in their entirety.

[0092] Exemplary adenine and cytosine base editors are also described in: Rees & Liu, Base editing: precision chemistry on the genome and transcriptome of living cells, Nat. Rev. Genet. 2018;19(12):770-788; and U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018, which issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published on May 4, 2017, which issued as U.S. Patent No. 10,167,457 on January 1, 2019; PCT Publication No. WO 2017 / 0121693, published on April 27, 2017 2017 / 070633; U.S. Patent Publication No. 2015 / 0166980, published on June 18, 2015; U.S. Patent No. 9,840,699, issued on December 12, 2017; and U.S. Patent No. 10,077,453, issued on September 18, 2018, the contents of each of which are incorporated herein by reference in their entirety.

[0093] The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nucleic acid comprising a Cas9 domain or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9 and / or a gRNA binding domain of Cas9). As used herein, a "Cas9 domain" is a protein fragment comprising an active or inactive cleavage domain of Cas9 and / or a gRNA binding domain of Cas9. A "Cas9 protein" is a full-length Cas9 protein. Cas9 nuclease is sometimes also referred to as casn1 nuclease or CRISPR ( C lustered R egularly I Interspaced S hort P alindromic Repeat, clustered regularly interspaced short palindromic repeats) related nucleases. CRISPR is an adaptive immune system that can provide protection against mobile genetic elements (viruses, transposable elements and conjugated plasmids). The CRISPR cluster contains a spacer, a sequence complementary to the previous mobile element and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), an endogenous ribonuclease 3 (rnc) and a Cas9 domain. TracrRNA serves as a guide for ribonuclease 3 to assist in the processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA cuts the linear or circular dsDNA target complementary to the spacer by endonuclease. The target chain that is not complementary to the crRNA is first cut by an endonuclease, and then trimmed 3'-5' by an exonuclease. In nature, DNA binding and cutting generally require proteins and two types of RNA. However, single guide RNA ("sgRNA", or simply "gRNA") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., et al. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, Streptococcus pyogenes and Streptococcus thermophilus (e.g., StCas9 or St1Cas9).Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0094] Nuclease-inactivated Cas9 domains are interchangeably referred to as "dCas9" proteins (for nuclease-"dead" Cas9). Methods for generating Cas9 domains (or fragments thereof) with inactivated DNA cleavage domains are known (see, e.g., Jinek et al., Science.337:816-821(2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell.28;152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of known Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the chain complementary to the gRNA, while the RuvC1 subdomain cuts the non-complementary chain. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83(2013)). In some embodiments, proteins comprising Cas9 fragments are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or a fragment thereof are referred to as "Cas9 variants". Cas9 variants share homology with Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 74).In some embodiments, the Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 74). In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to a corresponding fragment of wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 74). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 74).

[0095] As used herein, the term "nCas9" or "Cas9 nickase" refers to Cas9 or a variant thereof that cuts or nicks only one strand of the target cleavage site, thereby introducing a nick in the double-stranded DNA molecule rather than generating a double-strand break. This can be achieved by introducing an appropriate mutation into wild-type Cas9 that inactivates one of the two endonuclease activities of Cas9. It is contemplated that any suitable mutation that inactivates one Cas9 endonuclease activity but leaves the other intact, such as one of the D10A or H840A mutations in the wild-type Streptococcus pyogenes Cas9 amino acid sequence, or the D10A mutation in the wild-type Staphylococcus aureus Cas9 amino acid sequence, can be used to form nCas9.

[0096] The term "cDNA" refers to a DNA strand copied from an RNA template. The cDNA is complementary to the RNA template.

[0097] CRISPR is a family of DNA sequences in bacteria and archaea (i.e., CRISPR clusters), which represent fragments of previous infections by viruses that have invaded prokaryotes. DNA fragments are used by prokaryotic cells to detect and destroy DNA from subsequent attacks of similar viruses, and together with arrays of CRISPR-related proteins (including Cas9 and its homologs) and CRISPR-related RNAs, effectively form a prokaryotic immune defense system. In nature, the CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), the correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 domains. TracrRNA is used as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA cuts linear or circular dsDNA targets complementary to RNA by endonuclease. Specifically, the target chain that is not complementary to crRNA is first cut by endonuclease, and then 3'-5' is trimmed by exonuclease. In nature, DNA binding and cleavage typically require proteins and two types of RNA. However, single guide RNA ("sgRNA", or simply "gRNA") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species - the guide RNA. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.CRISPR biology and the Cas9 nuclease sequence and structure are well known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, Streptococcus pyogenes and Streptococcus thermophilus.Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.

[0098] The term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine (or adenine) deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA) to inosine. In some embodiments, the deaminase is a cytidine (or cytosine) deaminase, which catalyzes the hydrolytic deamination of cytidine to uracil.

[0099] The deaminases described herein can be from any organism, such as bacteria. In some embodiments, the deaminases or deaminases domains are variants of naturally occurring deaminases from an organism. In some embodiments, the deaminases or deaminases domains do not occur in nature. For example, in some embodiments, the deaminases or deaminases domains are at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to naturally occurring deaminases.

[0100] As used herein, the term "DNA binding protein" or "DNA binding protein domain" is any protein that is designated to localize to and bind to a specific target DNA nucleotide sequence (e.g., a genomic locus). The term includes RNA programmable proteins that associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., including, for example, guide RNAs in the case of a Cas system) that guide or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a DNA sequence) that is complementary to the one or more nucleic acid molecules (or portions or regions thereof) associated with the protein. Exemplary RNA programmable proteins are CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and can include Cas9 equivalents from any type of CRISPR system (e.g., type II, V, VI), including Cpf1 (type V CRISPR-Cas system) (now called Cas12a), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Nme2Cas9, SaCas9, SaKKH-Cas9, SauriCas9, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14, Argonaute, and nCas9. Further Cas equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), the contents of which are incorporated herein by reference.

[0101] As used herein, the term "DNA editing efficiency" refers to the number or ratio of expected base pairs edited. For example, if a base editor edits 10% of its expected targeted base pairs (for example, within a cell or within a cell population), the efficiency of the base editor can be described as 10%. Certain aspects of editing efficiency include modifying (for example, deaminating) specific nucleotides in DNA without generating a large amount or a high percentage of insertions or deletions (i.e., indels). It is generally accepted that when less than 5% indels are generated (as measured on the total target nucleotide substrate), editing is a high editing efficiency. Generating more than 20% indels is generally considered to be poor or low editing efficiency. The formation of indels can be measured by techniques known in the art, including high-throughput screening of sequencing reads.

[0102] As used herein, the term "off-target editing efficiency" refers to the number or ratio of unintended base pairs (e.g., DNA base pairs) that are edited. The on-target and off-target editing frequencies can be measured by methods and assays described herein, further taking into account techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves hybridization of nucleic acid primers (e.g., DNA primers) complementary to nucleic acid (e.g., DNA) regions just upstream or downstream of the target sequence of interest or off-target sequence. Because DNA target sequences and Cas9-independent off-target sequences are inherently known in the methods disclosed herein, nucleic acid primers with sufficient complementarity can be designed using techniques known in the art (e.g., PhusionU PCR kit (Life Technologies), Phusion HS II kit (Life Technologies), and Illumina MiSeq kit) to target sequences of interest and Cas9-independent off-target sequence upstream or downstream regions. The number of off-target DNA edits can be measured by techniques known in the art, including high-throughput screening, EndoV-Seq, GUIDE-Seq, CIRCLE-Seq, and Cas-OFFinder of sequencing reads. Because many Cas9 dependencies miss the target site and the target site of interest have high sequence identity, therefore equally can use technology known in the art and test kit design and Cas9 dependency miss the target site upstream or downstream region have enough complementary nucleic acid primers.These test kits utilize polymerase chain reaction (PCR) amplification, and it produces amplicon as intermediate product.The target sequence and miss the target sequence can comprise the genomic locus further containing protospacer and PAM.Therefore, as used herein, term "amplicon" can refer to the nucleic acid molecule of the aggregate that constitutes genomic locus, protospacer and PAM.The high throughput sequencing technology used herein can further include Sanger order-checking and next generation genome sequencing (NGS) based on Illumina.

[0103] As used herein, the term “on-target editing” refers to the introduction of an expected modification (e.g., deamination) to a nucleotide (e.g., adenine) in a target sequence, such as using a base editor as described herein. As used herein, the term “off-target DNA editing” refers to the introduction of an unexpected modification (e.g., deamination) to a nucleotide (e.g., adenine) in a sequence outside a typical base editor binding window (i.e., from one protospacer position to another protospacer position, typically 2 to 8 nucleotides long). Off-target DNA editing can be caused by weak or nonspecific binding of the gRNA sequence to the target sequence. As used herein, the term “bystander editing” refers to a synonymous off-target point mutation at a nucleobase close to (near) the target base, and does not change the outcome of the expected mutation (e.g., expected destruction of a splice acceptor site, incorporation of a premature stop codon, or reversal of a mutant codon). Bystander editing can cover non-silent mutations in relevant codons of a transcript that do not produce different translated proteins.

[0104] As used herein, the terms "purity" and "product purity" of a base editor refer to the average percentage of edited sequencing reads (reads in which the target nucleobase has been converted to a different base) in which the intended target conversion occurs (e.g., where target A and only target A is converted to G). See Komor et al., Sci Adv 3 (2017).

[0105] As used herein, the terms "upstream" and "downstream" are relative terms defining the linear position of at least two elements in a nucleic acid molecule (whether single-stranded or double-stranded), which are oriented in a 5' to 3' direction. In particular, in a nucleic acid molecule, a first element is located upstream of a second element, wherein the first element is located somewhere 5' of the second element. For example, if a SNP is located at the 5' side of a nicking site, the SNP is located upstream of the nicking site induced by Cas9. On the contrary, in a nucleic acid molecule, a first element is located downstream of the second element, wherein the first element is located somewhere 3' of the second element. For example, if a SNP is located at the 3' side of a nicking site, the SNP is located downstream of the nicking site induced by Cas9. Nucleic acid molecules can be DNA (double-stranded or single-stranded). RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. The analysis for single-stranded nucleic acid molecules and double-stranded molecules is the same because the terms upstream and downstream refer only to the single strand of nucleic acid molecules, except for the need to consider which chain of the double-stranded molecule to select. Typically, the strand of double-stranded DNA that can be used to determine the positional correlation of at least two elements is the "sense" or "coding" strand. In genetics, the "sense" strand is the segment running from 5' to 3' within the double-stranded DNA and is complementary to the antisense or template strand of the DNA running from 3' to 5'. Thus, as an example, if a SNP nucleobase is located 3' to the 3' side of the promoter on the sense or coding strand, then the SNP nucleobase is located "downstream" of the promoter sequence in the genomic DNA (which is double-stranded).

[0106] As used herein, the term "effective amount" refers to the amount of a bioactive agent sufficient to induce a desired biological response. For example, in some embodiments, the effective amount of a base editor may refer to the amount of an editor sufficient to edit the target site nucleotide sequence (e.g., genome). In some embodiments, the effective amount of a base editor as described herein, for example, an effective amount of a base editor comprising a nickase Cas9 domain and a guide RNA may refer to the amount of a base editor sufficient to induce editing of a target site specifically bound and edited by a base editor. As will be appreciated by those skilled in the art, the effective amount of an agent (e.g., a base editor, a nuclease, a hybrid protein, a protein dimer, a complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide) may vary according to various factors, such as, for example, according to the desired biological response, for example, according to the specific allele to be edited, the genome or the target site, according to the targeted cell or tissue, and according to the agent used.

[0107] The term "functional equivalent" refers to a second biomolecule that is functionally equivalent to a first biomolecule, but not necessarily structurally equivalent. For example, a "Cas9 equivalent" refers to a protein that has the same or substantially the same function as Cas9, but not necessarily the same amino acid sequence. In the context of the present disclosure, reference is made throughout this specification to "protein X or a functional equivalent thereof." Herein, a "functional equivalent" of protein X includes any homolog, paralog, fragment, naturally occurring, engineered, ring-substituted, mutated, or synthetic form of protein X that has an equivalent function.

[0108] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) portion or the carboxyl-terminal (C-terminal) protein of the fusion protein, thereby forming an "amino-terminal fusion protein" or a "carboxyl-terminal fusion protein", respectively. A protein can contain different domains, for example, a nucleic acid binding domain (e.g., a gRNA binding domain of Cas9 that guides the protein to bind to the target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. Another example includes Cas9 fused to adenosine deaminase or its equivalent. Any protein described herein can be produced by any method known in the art. For example, the proteins described herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins comprising peptide linkers. Methods for recombinant protein expression and purification are well known and include Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0109] The term "guide nucleic acid" or "napDNAbp-programmed nucleic acid molecule" or equivalently "guide sequence" refers to one or more nucleic acid molecules that associate with the napDNAbp protein and guide or otherwise program the napDNAbp protein to locate to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to the one or more nucleic acid molecules (or portions or regions thereof) associated with the protein, thereby binding the napDNAbp protein to the nucleotide sequence at the specific target site. A non-limiting example is the guide RNA of the Cas protein of the CRISPR-Cas genome editing system. Chemically, a guide nucleic acid can be all RNA, all DNA, or a chimera of RNA and DNA. The guide nucleic acid can also include nucleotide analogs. The guide nucleic acid can be expressed as a transcript or can be synthesized.

[0110] As used herein, "guide RNA" or "gRNA" refers to a synthetic fusion of endogenous bacterial crRNA and tracrRNA, which provides targeting specificity and scaffold and / or binding ability to the target DNA for the Cas9 nuclease. This synthetic fusion does not exist in nature and is also commonly referred to as sgRNA. However, the term guide RNA also includes the equivalent guide nucleic acid molecules associated with Cas9 equivalents, homologues, straight homologues or paralogues, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and it is otherwise programmed to locate to a specific target nucleotide sequence to the Cas9 equivalent. Cas9 equivalents can include other napDNAbp from any type of CRISPR system (e.g., type II, type V, type VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system) and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing suitable guide RNA sequences are provided herein.

[0111] Guide RNA is a specific type of guide nucleic acid that is typically associated with the Cas protein of CRISPR-Cas9 and binds to Cas9, directing the Cas9 protein to a specific sequence in a DNA molecule that contains complementarity with the protospacer sequence of the guide RNA. Functionally, the guide RNA associates with Cas9, directing (or programming) the Cas9 protein to a specific sequence in a DNA molecule that contains a sequence complementary to the protospacer sequence of the guide RNA.

[0112] As used herein, a "spacer sequence" is a sequence of a guide RNA (~20 nt in length) that has the same sequence as the protospacer of the PAM strand of the target (DNA) sequence (except for uridine bases instead of thymine bases) and that is complementary to the target strand (or non-PAM strand) of the target sequence.

[0113] As used herein, "target sequence" refers to the ~20 nucleotides of the target DNA sequence that have complementarity with the protospacer sequence in the PAM strand. The target sequence is the sequence that anneals to or is targeted by the spacer sequence of the guide RNA. The spacer sequence and the protospacer of the guide RNA have the same sequence (except that the spacer sequence is RNA, while the protospacer is DNA).

[0114] As used herein, the terms "guide RNA core," "guide RNA scaffold sequence," and "backbone sequence" refer to the sequence within the gRNA responsible for Cas9 binding, excluding the 20 bp spacer sequence used to guide Cas9 to the target DNA.

[0115] As used herein, the term "host cell" refers to a cell that can carry and replicate a vector encoding a base editor, a guide RNA, and / or a combination thereof as described herein. In some embodiments, the host cell is a mammalian cell, such as a human cell. Provided herein are methods for transducing and transfecting host cells (e.g., human cells, such as human cells in a subject) with one or more vectors provided herein (e.g., one or more viruses (e.g., rAAV) vectors provided herein).

[0116] It should be understood that any base editor, guide RNA and / or combination thereof described herein can be stably or transiently introduced into a host cell in any suitable manner. In some embodiments, the base editor can be transfected into the host cell. In some embodiments, the host cell can be transduced or transfected with a nucleic acid construct encoding a base editor. For example, the host cell can be transduced with a nucleic acid encoding a base editor or a translated base editor (e.g., with a viral particle encoding a base editor). As another example, the host cell can be transfected with a nucleic acid encoding a base editor (e.g., a plasmid) or a translated base editor. Such transduction or transfection can be stable or transient. In some embodiments, a host cell expressing a base editor or containing a base editor can be transduced or transfected with one or more gRNA molecules, for example, when the base editor comprises a Cas9 (e.g., nCas9) domain. In some embodiments, the plasmid expressing the base editor can be transduced or transfected by electroporation, transient transfection (e.g., lipid transfection, for example with Lipofectamine 3000 ® ), stable genomic integration (e.g., piggybac), viral transduction, or other methods known to those skilled in the art.

[0117] This article also provides host cells for packaging viral particles. In embodiments where the vector is a viral vector, suitable host cells are cells that can be infected by the viral vector, can replicate the viral vector, and can package the viral vector into viral particles that can infect fresh host cells. If the cell supports the expression of viral vector genes, the replication of the viral genome, and / or the generation of viral particles, then the cell can carry the viral vector. In some embodiments, the host cell is a eukaryotic cell, for example, a yeast cell, an insect cell, or a mammalian cell. Of course, the type of host cell will depend on the vector used, and suitable host cell / vector combinations will be apparent to those skilled in the art.

[0118] "Intein" is a fragment of a protein that can excise itself and connect the remaining part (extein) with a peptide bond in a process called protein splicing. Intein is also called "protein intron". The process in which an intein excises itself and connects the remaining part of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In some embodiments, the intein of a precursor protein (a protein containing an intein before intein-mediated protein splicing) comes from two genes. Such inteins are referred to herein as split inteins. For example, in cyanobacteria, DnaE (catalytic subunit α of DNA polymerase III) is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene is referred to herein as "intein-N". The intein encoded by the dnaE-c gene is referred to herein as "intein-C". In various embodiments of the disclosed nucleic acid molecules, the nucleic acid molecules do not contain inteins. In various embodiments of the disclosed nucleic acid molecules, the nucleic acid molecules do not contain trans-splicing inteins.

[0119] Other intein systems can also be used. For example, synthetic intein based on the dnaE intein, the Cfa-N and Cfa-C intein pairs, have been described (e.g., as described in Stevens et al., J Am Chem Soc. 2016 Feb 24; 138(7): 2162-5, incorporated herein by reference). As another example, synthetic intein based on the dnaE intein, the Nostoc punctata (Npu) intein pairs, have been described (see Zettler, J., Schutz, V. & Mootz, HD, The naturally split Npu DnaE intein exhibits an extraordinarily high rate in the protein trans-splicing reaction. FEBS letters 583, 909-914 (2009), incorporated herein by reference). In some embodiments, the intein is a fast-splicing gp41 intein, such as the gp41-8 intein. With reference to Carvajal-Vallejos et al, Journal of Biological Chemistry 287(34): 28686-28696 (2012) and Pinto, Thornton & Wang, Nat. Comm. (2020) 11: 1529, each of which is incorporated herein by reference. Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include: Cfa DnaE intein, Npu DnaE intein, gp41-8 intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, incorporated herein by reference).

[0120] As used herein, the term "linker" refers to a chemical group or molecule that connects two molecules or domains, for example, dCas9 and deaminase. Typically, the linker is located between or flanks two groups, molecules or other domains and is connected to each group, molecule or other domain via a covalent bond, thereby connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer or chemical domain. Chemical groups include but are not limited to disulfide, hydrazone and azide domains. In some embodiments, the length of the linker is 5-100 amino acids, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150 or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the linker is an XTEN linker having a length of 32 amino acids. In some embodiments, the linker is a 32 amino acid linker. In other embodiments, the linker is a 30, 31, 33 or 34 amino acid linker.

[0121] As used herein, the term "mutation" refers to the replacement of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) by another residue; the deletion or insertion of one or more residues within a sequence; or the replacement of a residue within the genomic sequence of a subject to be corrected. Mutations are generally described herein by identifying the original residue, followed by the position of the residue in the sequence, and the identity of the new, replaced residue. Various methods for preparing amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012). Mutations can include a variety of categories, such as single base polymorphisms, microduplications, indels, and inversions, and are not meant to be limiting in any way. Mutations can include "loss-of-function" mutations, which are mutations that reduce or eliminate protein activity. Most loss-of-function mutations are recessive because, in heterozygotes, the second chromosome copy carries an unmutated version of the gene encoding a fully functional protein, the presence of which compensates for the effects of the mutation. There are some exceptions where loss-of-function mutations are dominant, an example being haploinsufficiency, where the organism cannot tolerate the approximately 50% reduction in protein activity experienced by heterozygotes. This is an explanation for some human genetic diseases, including Marfan syndrome, which is caused by mutations in the gene for a connective tissue protein called fibrillin. Mutations also include "gain-of-function" mutations, which are mutations that confer abnormal activities on a protein or cell that are not present under normal conditions. Many gain-of-function mutations are located in regulatory sequences rather than in coding regions and therefore can have many consequences. By their nature, gain-of-function mutations are typically dominant. Many loss-of-function mutations are recessive, such as autosomal recessive. Many of the USH2A mutations that currently published base editing methods are designed to correct are autosomal recessive mutations.

[0122] The term "napDNAbp," which stands for "nucleic acid programmable DNA binding protein," refers to any protein that can associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., which can be broadly referred to as "napDNAbp programming nucleic acid molecules," and include, for example, guide RNAs in the case of Cas systems) that direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to the one or more nucleic acid molecules (or portions or regions thereof) associated with the protein, thereby causing the protein to bind to the nucleotide sequence at the specific target site. The term napDNAbp includes CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and can include Cas9 equivalents from any type of CRISPR system (e.g., type II, V, VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14, Argonaute, and nCas9. Further Cas equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353 (6299), the contents of which are incorporated herein by reference. However, the nucleic acid programmable DNA binding proteins (napDNAbp) that can be used in conjunction with the present invention are not limited to CRISPR-Cas systems. The present invention includes any such programmable proteins, such as the Argonaute protein from NgAgo, which can also be used for DNA-guided genome editing. The NgAgo-guided DNA system does not require a PAM sequence or a guide RNA molecule, which means that genome editing can be simply performed by expressing a universal NgAgo protein and introducing a synthetic oligonucleotide on any genomic sequence.See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7):768-73, which is incorporated herein by reference.

[0123] In some embodiments, the napDNAbp is an RNA programmable nuclease, which, when in a complex with RNA, can be referred to as a nuclease:RNA complex. Typically, the bound RNA is referred to as a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), although "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species comprises two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs the binding of the Cas9 (or equivalent) complex to the target); and (2) a domain that binds the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and comprises a stem-loop structure. For example, in some embodiments, domain (2) is homologous to the tracrRNA depicted in Figure 1E of Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Patent No. 9,340,799, entitled “mRNA-Sensing Switchable gRNAs,” and PCT Application No. PCT / US2014 / 054247, filed on September 6, 2013, published as WO 2015 / 035136, and entitled “Delivery System For Functional Nucleases,” the entire contents of each of which are incorporated herein by reference. In some embodiments, a gRNA comprises two or more of domains (1) and (2) and can be referred to as an “extended gRNA.” For example, an extended gRNA will, for example, bind two or more Cas9 proteins and bind to a target nucleic acid at two or more different regions, as described herein. The gRNA comprises a nucleotide sequence complementary to a target site that mediates binding of the nuclease / RNA complex to the target site, providing sequence specificity for the nuclease:RNA complex.In some embodiments, the RNA programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 from Streptococcus pyogenes (Csn1) (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti JJ et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and hostfactor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference).

[0124] napDNAbp nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, and these proteins can, in principle, target any sequence specified by a guide RNA. Methods for using napDNAbp nucleases such as Cas9 to perform site-specific cleavage (e.g., to modify a genome) are known in the art (see, e.g., Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature Biotechnology 31, 227-229 (2013); Jinek, M. et al. RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013); Jiang, W. et al. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).

[0125] The term "nickase" refers to a napDNAbp (e.g., Cas9) that has only a single nuclease activity that cuts only one strand of the target DNA instead of both strands. Thus, a nickase-type napDNAbp does not leave a double-strand break. Exemplary nickases include SpCas9 and SaCas9 nickases. Exemplary nickases include a sequence having at least 99% or 100% identity to the amino acid sequence of SEQ ID NO: 107.

[0126] Nuclear localization signal or sequence (NLS) is an amino acid sequence that marks, specifies or otherwise identifies a protein to be imported into the nucleus by nuclear transport. Typically, the signal consists of one or more short sequences of positively charged lysine or arginine exposed on the protein surface. Different nuclear localization proteins can share the same NLS. NLS has a function opposite to the nuclear export signal (NES), which targets proteins outside the nucleus. Therefore, a single nuclear localization signal can guide the entity associated with it to the nucleus. Such a sequence can be of any size and composition, for example, more than 25, 25, 15, 12, 10, 8, 7, 6, 5 or 4 amino acids, but preferably comprises at least 4 to 8 amino acid sequences known to function as nuclear localization signals (NLS).

[0127] As used herein, the term "nucleic acid molecule" refers to RNA and single-stranded and / or double-stranded DNA. A nucleic acid molecule can be naturally occurring, for example, in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule can be a non-naturally occurring molecule, for example, a recombinant DNA or RNA, an artificial chromosome, an engineered genome or fragment thereof, or a synthetic DNA, RNA, DNA / RNA hybrid, or include non-naturally occurring nucleotides or nucleosides.

[0128] In addition, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, e.g., analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified, chemically synthesized, and the like. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, e.g., analogs having chemically modified bases or sugars and backbone modifications. Unless otherwise indicated, nucleic acid sequences are presented in a 5' to 3' orientation. In some embodiments, the nucleic acid is or comprises natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C ... -aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, inosineadenosine, 8-oxoguanosine, O(6)-methylguanosine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases, such as 2'-O-methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0129] As used herein, the term "phage-assisted continuous evolution (PACE)" refers to continuous evolution using bacteriophage as a viral vector. The general concept of PACE technology has been described in, for example, PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, published as WO 2010 / 028347 on March 11, 2010; PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; U.S. Application No. 9,023,594, issued May 5, 2015, PCT Application No. PCT / US2015 / 012022, filed January 20, 2015, published as WO 2012 / 088381 on September 11, 2015; 2015 / 134121, and PCT Application No. PCT / US2016 / 027795, filed April 15, 2016, published as WO 2016 / 168631 on October 20, 2016, the entire contents of each of which are incorporated herein by reference.

[0130] The term "promoter" is recognized in the art and refers to a nucleic acid molecule having a sequence that is recognized by the cell transcription machinery and can initiate transcription of downstream genes. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular environment, or conditionally active, meaning that the promoter is only active under specific conditions. For example, a conditional promoter may be active only in the presence of a specific protein that connects a protein associated with a regulatory element in the promoter to the basic transcription machinery, or only in the absence of an inhibitory molecule. A subclass of conditionally active promoters is an inducible promoter, which requires the presence of a small molecule "inducer" to obtain activity. Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. Various constitutive, conditional, and inducible promoters are well known to those skilled in the art, and those skilled in the art will be able to determine various such promoters that can be used to implement the present invention, and the present invention is not limited in this respect. In various embodiments, the present disclosure provides vectors having suitable promoters for driving expression of nucleic acid sequences encoding base editors (or one or more individual components thereof).

[0131] As used herein, the term "protospacer" refers to a sequence in the DNA adjacent to a PAM (protospacer adjacent motif) sequence (e.g., a ~20 bp sequence) that shares the same sequence with the spacer sequence of the guide RNA and is complementary to the target sequence of the non-PAM strand. The spacer sequence of the guide RNA anneals to the target sequence located on the non-PAM strand. In order for Cas9 to function, it also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from Streptococcus pyogenes, recognizes a PAM sequence of NGG on the non-target strand, which is found directly downstream of the protospacer sequence in genomic DNA. The skilled artisan will understand that the literature in the prior art sometimes refers to the "protospacer" as the ~20-nt target-specific guide sequence on the guide RNA itself, rather than as a "spacer" (and the protospacer (DNA) and the spacer (RNA) have the same sequence). Therefore, the term "protospacer" as used herein can be used interchangeably with the term "spacer". The context surrounding the description of the occurrence of "protospacer" or "spacer" will help inform the reader whether the term refers to the gRNA or the DNA sequence. Both uses of these terms are acceptable, as the prior art uses these two terms in each of these ways.

[0132] As used herein, the term "protospacer adjacent sequence" or "PAM" refers to a DNA sequence of about 2-6 base pairs, which is an important targeting component of the Cas9 nuclease. Typically, the PAM sequence is on either chain and downstream of the Cas9 cleavage site in the 5' to 3' direction. A typical PAM sequence (i.e., a PAM sequence associated with the Cas9 nuclease or SpCas9 of Streptococcus pyogenes) is 5'-NGG-3', where "N" is any nucleobase followed by two guanine ("G") nucleobases. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease (e.g., SpCas9) can be modified to change the PAM specificity of the nuclease so that the nuclease recognizes an alternative PAM sequence.

[0133] For example, with reference to the typical SpCas9 amino acid sequence is SEQ ID NO: 74, the PAM sequence can be modified by introducing one or more mutations, including (a) D1135V, R1335Q and T1337R "VRQR variant", which changes the PAM specificity to NGAN or NGNG, (b) D1135E, R1335Q and T1337R "EQR variant", which changes the PAM specificity to NGAG, and (c) D1135V, G1218R, R1335E and T1337R "VRER variant", which changes the PAM specificity to NGCG. In addition, the D1135E variant of the typical SpCas9 still recognizes NGG, but has higher selectivity than the wild-type SpCas9 protein.

[0134] It will also be understood that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) can have different PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. In addition, Cas9 from Neisseria meningitidis (NmeCas) recognizes NNNNGATT. Cas9 from Staphylococcus aureus (SauriCas9) recognizes NNGG and NNNGG. Cas9 from Streptococcus thermophilus (StCas9) recognizes NNAGAAW. Cas9 from Treponema denticola (TdCas) recognizes NAAAC. These examples are not meant to be limiting. It will be further understood that non-SpCas9 binds to a variety of PAM sequences, which makes them useful when there is no suitable SpCas9 PAM sequence at the desired target cleavage site. In addition, non-SpCas9 can have other features that make them more useful than SpCas9. For example, Cas9 from Staphylococcus aureus (SaCas9) is about 1 kilobase smaller than SpCas9, so it can be packaged into adeno-associated virus (AAV). Further reference may be made to Shah et al., “Protospacer recognition motifs: mixed identities and functional diversity,” RNA Biology, 10(5): 891-899 (incorporated herein by reference).

[0135] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids long. A protein, peptide, or polypeptide may refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified, for example, by adding chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications, and the like. A protein, peptide, or polypeptide may also be a single molecule or may be a multimolecular complex. A protein, peptide, or polypeptide may be merely a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. It should be understood that the present disclosure provides any polypeptide sequence provided herein that does not have an N-terminal methionine (M) residue.

[0136] In genetics, the "sense" strand is the fragment running from 5' to 3' in double-stranded DNA, and it is complementary to the antisense strand or template strand of the DNA running from 3' to 5'. In the case of a DNA fragment encoding a protein, the sense strand is a DNA strand with the same sequence as the mRNA, which uses the antisense strand as its template during transcription and ultimately undergoes (usually, not always) translation into protein. Therefore, the antisense strand is responsible for the RNA that is subsequently translated into protein, while the sense strand has a composition almost identical to that of the mRNA. Note that for each fragment of dsDNA, there may be two groups of sense strands and antisense strands, depending on the direction of reading (because sense strands and antisense strands are related to the viewing angle). Ultimately, it is the gene product or mRNA that determines which strand of a fragment of dsDNA is called the sense strand or the antisense strand.

[0137] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein provided as an N-terminal portion (which is interchangeably referred to herein as the N-terminal half) and a C-terminal portion (which is interchangeably referred to herein as the C-terminal half) encoded by two separate nucleotide sequences. The polypeptides corresponding to the N-terminal portion and the C-terminal portion of the Cas9 protein can be combined (joined) to form a complete Cas9 protein. It is known that the Cas9 protein consists of a two-lobed structure connected by an unordered linker (e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935–949, 2014, incorporated herein by reference). In some embodiments, "splitting" occurs between the two lobes, generating two parts of the Cas9 protein, each containing one lobe.

[0138] As used herein, the term "subject" refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, an amphibian, a reptile, a fish, an insect, a fly, or a nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject can be of either sex and can be at any stage of development. In some embodiments, the subject is a domesticated animal. In some embodiments, the subject is a plant.

[0139] The term "target site" refers to a sequence edited by a base editor (BE) disclosed herein within a nucleic acid molecule. In the context of a single strand, the term "target site" may also refer to a "target chain" that anneals or binds to a spacer sequence of a guide RNA. In certain embodiments, the target site may refer to a fragment of double-stranded DNA, comprising a PAM chain (or non-target chain) and a protospacer on the target chain (i.e., a chain of the target site having the same nucleotide sequence as the spacer sequence of the guide RNA), which is complementary to the protospacer and a similar spacer, and anneals to the spacer of the guide RNA, thereby targeting or programming the Cas9 base editor to target the target site.

[0140] A "transcription terminator" is a nucleic acid sequence that causes transcription to stop. A transcription terminator can be unidirectional or bidirectional. It consists of a DNA sequence that participates in the specific termination of RNA transcripts by RNA polymerase. The transcription terminator sequence prevents transcriptional activation of a downstream nucleic acid sequence by an upstream promoter. Transcription terminators may be necessary in vivo to achieve desired expression levels or to prevent transcription of certain sequences. A transcription terminator is considered "operably linked" to a nucleotide sequence when it is capable of terminating transcription of the sequence to which it is linked.

[0141] In eukaryotic systems, terminator regions can include specific DNA sequences that allow site-specific cutting of new transcripts to expose polyadenylation sites. This signals specialized endogenous polymerases to add a fragment of approximately 200 A residues (poly A) to the 3' end of the transcript. RNA molecules modified with this poly A tail (signal) appear to be more stable and translate more effectively. Therefore, in some embodiments relating to eukaryotic organisms, terminator regions can include signals for cutting RNA. In some embodiments, terminator signals promote the polyadenylation of information. Terminator and / or polyadenylation site elements can be used for enhancing output nucleic acid levels and / or minimizing the read-through between nucleic acids.

[0142] In some embodiments, the transcription terminator comprises a post-transcriptional response element, which is a sequence that can produce a tertiary structure that enhances expression when transcribed. In some embodiments, the post-transcriptional response element is derived from woodchuck hepatitis virus (WHV), i.e., WPRE. In some embodiments, the terminator contains the gamma subunit or W3 of WPRE, as first reported in Choi, JH, et al. (2014), Mol. Brain 7: 17, incorporated herein by reference. WPRE also has alpha and beta subunits. Typically, the post-transcriptional response element is inserted 5' of the transcription terminator. In certain embodiments, the WPRE is a truncated WPRE sequence. In certain embodiments, the WPRE is a full-length WPRE.

[0143] Terminators used in accordance with the present disclosure include any transcriptional terminator described herein or known to those of ordinary skill in the art. Examples of terminators include, but are not limited to, gene terminator sequences such as, for example, the bovine growth hormone terminator, and viral terminator sequences such as, for example, the SV40 terminator and the bGH terminator. In some embodiments, the termination signal can be a sequence that cannot be transcribed or translated, such as those produced by sequence truncation.

[0144] As used herein, "transition" refers to the interchange of purine nucleobases (A↔G) or the interchange of pyrimidine nucleobases (C↔T). Such interchanges involve nucleobases of similar shape. The compositions and methods disclosed herein can induce one or more transitions in a target DNA molecule. The compositions and methods disclosed herein can also induce transitions and transversions in the same target DNA molecule. These changes involve A↔G, G↔A, C↔T, or T↔C. In the context of double-stranded DNA with Watson-Crick paired nucleobases, transitions refer to the following base pair exchanges: A:T↔G:C, G:G↔A:T, C:G↔T:A, or T:A↔C:G. The compositions and methods disclosed herein can induce one or more transitions in a target DNA molecule. The compositions and methods disclosed herein can also induce transitions and transversions, as well as other nucleotide changes, including deletions and insertions, in the same target DNA molecule.

[0145] As used herein, "transversion" refers to the interchange of a purine nucleobase with a pyrimidine nucleobase, or vice versa, thus involving the interchange of nucleobases having different shapes. These changes involve T↔A, T↔G, C↔G, C↔A, A↔T, A↔C, G↔C, and G↔T. In the context of double-stranded DNA with Watson-Crick paired nucleobases, transversion refers to the following base pair exchanges: T:A↔A:T, T:A↔G:C, C:G↔G:C, C:G↔A:T, A:T↔T:A, A:T↔C:G, G:C↔C:G, and G:C↔T:A.

[0146] The term "treatment" (treatment, treat, treating) refers to a clinical intervention intended to reverse a disease or disorder or one or more symptoms thereof, alleviate a disease or disorder or one or more symptoms thereof, delay the onset of a disease or disorder or one or more symptoms thereof, or suppress a disease or disorder or one or more symptoms thereof, as described herein. As used herein, the term "treatment" (treatment, treat, treating) refers to a clinical intervention intended to reverse a disease or disorder or one or more symptoms thereof, alleviate a disease or disorder or one or more symptoms thereof, delay the onset of a disease or disorder or one or more symptoms thereof, or suppress a disease or disorder or one or more symptoms thereof, as described herein. In some embodiments, treatment can be applied after one or more symptoms have developed and / or after the disease has been diagnosed. In other embodiments, treatment can be applied in the absence of symptoms, for example, to prevent symptoms or delay the onset of symptoms or suppress the onset or progress of the disease. For example, treatment can be applied to susceptible individuals (for example, according to symptom history and / or according to genetic or other susceptibility factors) before the onset of symptoms. Treatment can also continue after symptoms have subsided, for example, to prevent or delay its recurrence.

[0147] As used herein, the terms "upstream" and "downstream" are relative terms defining the linear position of at least two elements in a nucleic acid molecule (whether single-stranded or double-stranded), which are oriented in a 5' to 3' direction. In particular, in a nucleic acid molecule, a first element is located upstream of a second element, wherein the first element is located somewhere 5' of the second element. For example, if a SNP is located at the 5' side of a nicking site, the SNP is located upstream of the nicking site induced by Cas9. On the contrary, in a nucleic acid molecule, a first element is located downstream of the second element, wherein the first element is located somewhere 3' of the second element. For example, if a SNP is located at the 3' side of a nicking site, the SNP is located downstream of the nicking site induced by Cas9. Nucleic acid molecules can be DNA (double-stranded or single-stranded). RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. The analysis for single-stranded nucleic acid molecules and double-stranded molecules is the same because the terms upstream and downstream refer only to the single strand of nucleic acid molecules, except for the need to consider which chain of the double-stranded molecule to select. Typically, the strand of double-stranded DNA that can be used to determine the positional correlation of at least two elements is the "sense" or "coding" strand. In genetics, the "sense" strand is the segment running from 5' to 3' within the double-stranded DNA and is complementary to the antisense or template strand of the DNA running from 3' to 5'. Thus, as an example, if a SNP nucleobase is located 3' to the 3' side of a promoter on the sense or coding strand, then the SNP nucleobase is "downstream" of the promoter sequence in the genomic DNA (which is double-stranded).

[0148] As used herein, the term "variant" refers to a protein having features that deviate from those present in nature but retains at least one function (i.e., binding, interaction, or enzymatic ability and / or therapeutic properties). The variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the wild-type protein. For example, a variant of Cas9 may include a Cas9 having one or more changes in amino acid residues compared to the wild-type Cas9 amino acid sequence. As another example, a variant of a deaminase may include a deaminase having one or more changes in amino acid residues compared to the wild-type deaminase amino acid sequence, for example, after reconstruction of the ancestral sequence of the deaminase. These changes include chemical modifications, including replacement of different amino acid residues, truncations, covalent additions (e.g., tags), and any other mutations. The term also encompasses circular substitutions, mutants, truncations, or domains of a reference sequence that exhibit the same or substantially the same functional activity or activities as the reference sequence. The term also includes fragments of a wild-type protein.

[0149] The level or degree of retention of a characteristic may be reduced relative to the wild-type protein, but is generally identical or similar in kind. Typically, variants are very similar overall and are identical to the amino acid sequence of proteins described herein in many regions. Those skilled in the art will understand how to prepare and use variants that retain all or at least some functional capabilities or characteristics. Variant proteins can comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical to, for example, the amino acid sequence of the wild-type protein or any protein provided herein, or alternatively consist of it.

[0150] By a polypeptide having an amino acid sequence that is at least, for example, 95% "identical" to a query amino acid sequence, it is meant that the amino acid sequence of the subject polypeptide is identical to the query sequence except that the subject polypeptide sequence may include up to five amino acid changes for every 100 amino acids in the query amino acid sequence. In other words, up to 5% of the amino acid residues in the subject sequence may be inserted, deleted, or substituted with another amino acid to obtain a polypeptide having an amino acid sequence that is at least 95% identical to the query amino acid sequence. These changes in the reference sequence may occur at the amino-terminal or carboxyl-terminal positions of the reference amino acid sequence or at any position between these terminal positions, interspersed individually between residues in the reference sequence or in one or more contiguous groups within the reference sequence.

[0151] In fact, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical with the amino acid sequence of for example fusion rotein, can use known computer program to determine by convention.Be used to determine the preferred method of the best overall matching (also referred to as global sequence alignment) between query sequence (sequence of the present invention) and the subject sequence, can use the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)) to determine.In sequence alignment, query sequence and subject sequence both are nucleotide sequences or both are amino acid sequences.The result of described global sequence alignment is expressed as identity percentage. Preferred parameters for FASTDB amino acid alignments are: matrix=PAM 0, k-tuple=2, mismatch penalty=1, ligation penalty=20, randomization group length=0, cutoff score=1, window size=sequence length, gap penalty=5, gap size penalty=0.05, window size=500 or the length of the subject amino acid sequence, whichever is shorter.

[0152] If due to N-terminal or C-terminal deletion, rather than due to internal deletion, the subject sequence is shorter than the query sequence, then the result must be manually corrected. This is because the FASTDB program does not consider the N-terminal and C-terminal truncations of the subject sequence when calculating the global identity percentage. For the subject sequence that is truncated at the N-terminal and C-terminal ends relative to the query sequence, the percentage of the total bases of the query sequence is corrected by calculating the number of residues (these residues do not match / align with the corresponding subject residues) located at the N-terminal and C-terminal ends of the subject sequence in the query sequence. Whether the residue is matched / aligned is determined by the result of the FASTDB sequence comparison. Then, from the percentage identity calculated using the specified parameters by the above-mentioned FASTDB program, this percentage is subtracted to obtain final percentage identity score. This final percentage identity score is used for the purpose of the present invention. For the purpose of manually adjusting the identity percentage score, only the residues at the N-terminal and C-terminal ends of the subject sequence that do not match / align with the query sequence are considered. That is, only the residue positions outside the farthest N-terminal and C-terminal residues of the query subject sequence are queried.

[0153] As used herein, the term "vector" refers to a nucleic acid that can be modified to encode a gene of interest and is capable of entering and replicating in a host cell, and then transferring the replicated form of the vector to another host cell. Exemplary suitable vectors include viral vectors, such as AAV vectors or bacteriophages and filamentous phages, and conjugated plasmids. Based on this disclosure, additional suitable vectors will be apparent to those skilled in the art.

[0154] As used herein, the term "wild type" is a term understood by those skilled in the art and means the typical form of an organism, strain, gene or trait occurring in nature, as distinguished from mutant or variant forms.

[0155] napDNAbp domain

[0156] The base editors described herein include a nucleic acid programmable DNA binding (napDNAbp) domain. napDNAbp is associated with at least one guide nucleic acid (e.g., a guide RNA), which positions napDNAbp at a DNA sequence comprising a DNA chain (i.e., a target chain) complementary to the guide nucleic acid or a portion thereof (e.g., a protospacer of a guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp domain to locate and bind to the complementary sequence of the target chain. The binding of the napDNAbp domain to the complementary sequence enables the base editor's core base modification domain (i.e., adenosine deaminase domain) to approach the target base in the target chain and enzymatically deaminate it.

[0157] napDNAbp can be a nuclease associated with CRISPR (clustered regularly interspaced short palindromic repeats). As mentioned above, CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugated plasmids). The CRISPR cluster contains a spacer, a sequence complementary to the previously mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein. TracrRNA serves as a guide for ribonuclease 3 to assist in the processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA cuts the linear or circular dsDNA target complementary to the spacer through an endonuclease. The target chain that is not complementary to the crRNA is first cut by an endonuclease, and then the 3'-5' chain is trimmed by an exonuclease. In nature, DNA binding and cutting usually require proteins and RNA. However, single guide RNA ("sgRNA", or simply "gRNA") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, e.g., Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference.

[0158] The various napDNABPs described below that can be used in combination with the disclosed adenosine deaminase are not intended to be limiting in any way. The base editor may include a known typical SpCas9 or any orthologous Cas9 protein or any variant Cas9 protein (including any naturally occurring variant, mutant, or other engineered version of Cas9) that can be prepared or evolved by directed evolution or other mutagenesis processes. In various embodiments, napDNAbp has nickase activity, that is, it only cuts one chain of the target DNA sequence. In other embodiments, napDNAbp has an inactive nuclease, for example, a "dead" protein. Other variant Cas9 proteins that can be used are those with a smaller molecular weight than a typical SpCas9 (for example, for easier delivery) or with a modified or rearranged primary amino acid sequence (for example, a cyclic substitution form). The base editors described herein may also include Cas9 equivalents, including Cas12a / Cpf1 proteins. The napDNAbp used herein (for example, SpCas9, SaCas9, or SaCas9 variants or SpCas9 variants) may also include various modifications that change / enhance its PAM specificity. The present disclosure encompasses any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to any Cas9 protein disclosed herein. In some embodiments, the napDNAbp domain comprises a nickase variant of wild-type Cas9. In some embodiments, the napDNAbp domain comprises any Cas9 nickase disclosed herein.

[0159] In some embodiments, napDNAbp guides the cutting of one or two chains at the position of the target sequence (for example, in the target sequence and / or in the complementary sequence of the target sequence).In some embodiments, napDNAbp guides the cutting of one or two chains in about 1,2,3,4,5,6,7,8,9,10,15,20,25,50,100,200,500 or more base pairs from the first or last nucleotide of the target sequence.For example, the replacement (D10A) of aspartic acid to alanine in the RuvC I catalytic domain of the Cas9 from Streptococcus pyogenes converts Cas9 from the nuclease of cutting two chains into a nickase (cutting single strand).Other examples of the mutation that makes Cas9 become a nickase include but are not limited to H840A, N854A and N863A with reference to typical SpCas9 sequences, or with reference to the equivalent amino acid positions in other Cas9 variants or Cas9 equivalents.

[0160] As used herein, the term "Cas protein" refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence different from that of a naturally occurring Cas protein, or any fragment of a Cas protein that still retains all or a significant amount of the necessary basic functions required for the disclosed methods, i.e., (i) nucleic acid programmable binding of the Cas protein to the target DNA, and (ii) the ability to cleave the target DNA sequence on one strand. The Cas proteins contemplated herein include CRISPR Cas9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactivated Cas9 (dCas9)) homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and can include Cas9 equivalents from any type of CRISPR system (e.g., type II, V, VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), the contents of which are incorporated herein by reference.

[0161] The term "Cas9" or "Cas9 domain" encompasses any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog or paralog from any organism, and any mutant or variant of naturally occurring or engineered Cas9. The term Cas9 is not meant to be particularly restrictive and may be referred to as "Cas9 or equivalent". Exemplary Cas9 proteins are further described herein and / or described in the art and incorporated herein by reference. The present disclosure is not limited with respect to the specific napDNAbp used in the base editors of the present disclosure.

[0162] As used herein, the terms "compact Cas9 protein," "compact napDNAbp," and "compact variant [of a Cas protein]" refer to a Cas9 protein or variant having an amino acid length of less than about 1250 amino acids. In some embodiments, the compact Cas9 protein or compact napDNAbp comprises less than 1250 amino acids, less than 1240 amino acids, less than 1230 amino acids, less than 1220 amino acids, less than 1210 amino acids, less than 1200 amino acids, less than 1190 amino acids, less than 1180 amino acids, less than 1170 amino acids, less than 1160 amino acids, less than 1150 amino acids, less than 1140 amino acids, less than 1130 amino acids, less than 1120 amino acids, less than 1110 amino acids, less than 1050 amino acids, less than 1000 amino acids, less than 950 amino acids, less than 900 amino acids, less than 850 amino acids, less than 800 amino acids, less than 750 amino acids, less than 700 amino acids, less than 650 amino acids, less than 600 amino acids, less than 550 amino acids, less than 600 amino acids, less than 550 amino acids, or less than 500 amino acids in length. These terms also include any Cas9 protein or variant encoded by a nucleic acid sequence having a length of less than about 3750 nucleotides. The base editor of the present disclosure may include a compact napDNAbp and / or a compact Cas9 protein. In some embodiments, the compact Cas9 protein is about 350 amino acids shorter than SpCas9. In some embodiments, the length of the compact Cas9 protein is about 1000 amino acids. In some embodiments, the compact Cas9 protein is a compact variant of Streptococcus pyogenes Cas9 (SpCas9), Cpf1, CasX, CasY, C2c1, C2c2, C2c3, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Cas3 or CasΦ. "Compact variant" may refer to a Cas9 protein having one or more truncations or one or more deletions relative to a wild-type Cas9 protein (e.g., wild-type SpCas9 or Cpf1).

[0163] Additional Cas9 sequences and structures are well known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference), and provided below.

[0164] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not meant to be limiting. The base editors of the present disclosure can use any suitable napDNA bp, including any suitable Cas9 or Cas9 equivalent.

[0165] napDNAbp nickase

[0166] In some embodiments, the disclosed base editor may include a napDNAbp domain containing a nickase. In some embodiments, the base editor described herein includes a Cas9 nickase. The "Cas9 nickase" of the term "nCas9" refers to a Cas9 variant that can introduce single-strand breaks in a double-stranded DNA molecule target. In some embodiments, the Cas9 nickase comprises only a single functional nuclease domain. Wild-type Cas9 (e.g., typical SpCas9) includes two separate nuclease domains, i.e., a RuvC domain (which cuts the non-protospacer DNA chain) and an HNH domain (which cuts the protospacer DNA chain). In one embodiment, the Cas9 nickase includes a mutation in the RuvC domain that inactivates RuvC nuclease activity. For example, mutations in aspartic acid (D) 10, histidine (H) 983, aspartic acid (D) 986, or glutamic acid (E) 762 have been reported as loss-of-function mutations in the RuvC nuclease domain and the generation of functional Cas9 nickases (e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935–949, which is incorporated herein by reference). Thus, nickase mutations in the RuvC domain may include D10X, H983X, D986X, or E762X, where X is any amino acid other than the wild-type amino acid. In certain embodiments, the nickase may be D10A, H983A, or D986A, or E762A, or a combination thereof.

[0167] In some embodiments, the napDNAbp domain of any disclosed base editor comprises Streptococcus pyogenes Cas9 nickase (SpCas9n). In some embodiments, the napDNAbp domain of any disclosed base editor comprises at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 365 or 370. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 365. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 370.

[0168] In some embodiments, the napDNAbp domain of any disclosed base editor comprises Staphylococcus aureus Cas9 nickase (SaCas9n). In some embodiments, the napDNAbp domain of any disclosed base editor comprises at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 438. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 438.

[0169] In various embodiments, the Cas9 nickase can have a mutation in the RuvC nuclease domain and have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181] Compact Cas9 variants with modified PAM specificity

[0182] In some embodiments, the napDNAbp comprises a compact Cas protein, such as Cas9 derived from Campylobacter jejuni, Staphylococcus auris, Neisseria meningitidis, or Staphylococcus aureus. In exemplary embodiments, the napDNAbp comprises a CjCas9 nickase, a SauriCas9 nickase, a Nme2Cas9 nickase, a SaCas9 nickase, or a SaKKH-Cas9 nickase. In some embodiments, the napDNAbp is not a Nme2Cas9 protein or nickase. In some embodiments, the napDNAbp is not a SaCas9 protein or nickase.

[0183] The base editor of the present disclosure may also include a Cas9 variant with a modified PAM specificity. Some aspects of the present disclosure provide Cas9 proteins that exhibit activity to target sequences that do not include a typical PAM (5'-NGG-3', wherein N is A, C, G, or T) at their 3'-end. In some embodiments, the Cas9 protein exhibits activity to target sequences that include a 5'-NGG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to target sequences that include a 5'-NNG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to target sequences that include a 5'-NNA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to target sequences that include a 5'-NNC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to target sequences that include a 5'-NNC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity to target sequences that include a 5'-NNT-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NGT-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NGA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NGC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAA-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAC-3' PAM sequence at its 3' end. In some embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAT-3' PAM sequence at its 3' end. In other embodiments, the Cas9 protein exhibits activity against a target sequence comprising a 5'-NAG-3' PAM sequence at its 3' end.

[0184] In some embodiments, the disclosed base editor comprises a napDNAbp domain comprising SpCas9-NG having a PAM corresponding to NGN. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SpCas9-NG. The sequence of SpCas9-NG is shown below:

[0185]

[0186] In some embodiments, the disclosed base editor comprises a napDNAbp domain comprising the Staphylococcus aureus Cas9 nickase KKH or SaCas9-KKH, which has a PAM corresponding to NNNRRT or NNGRRT. The Cas9 variant contains amino acid substitutions D10A, E782K, N968K, and R1015H ("KKH") relative to wild-type SaCas9, as shown in SEQ ID NO: 377. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SaCas9-KKH. SaCas9 (and SaKKH-Cas9) has a length of 1053 amino acids. The sequence of SaCas9-KKH (nickase) is as follows:

[0187] Staphylococcus aureus Cas9 nickase KKH (SaCas9-KKH)

[0188]

[0189] In some embodiments, the disclosed base editor comprises a napDNAbp comprising a Cas9 protein derived from Staphylococcus auris (S. auri Cas9 or SauriCas9). In some embodiments, the disclosed base editor comprises a SauriCas9 nickase. SauriCas9 recognizes NNGG and NNNGG PAMs. The sequence of SauriCas9 (nickase) is shown in SEQ ID NO: 479. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 479. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 479. The protein is 1061 amino acids in length.

[0190]

[0191] In some embodiments, the napDNAbp comprises a SauriCas9-KKH variant or a SauriCas9-KKH nickase variant. SauriCas9-KKH contains the corresponding triple KKH mutations: Q788K, Y973K, R1020H. See Hu et al. (2020) PLoS Biol. 18(3): e3000686, which is incorporated herein by reference.

[0192] In some embodiments, the disclosed base editors comprise a napDNAbp domain comprising the Streptococcus pyogenes Cas9 nickase KKH or SpCas9-KKH, which nickase has a PAM corresponding to NNGRRT.

[0193] In some embodiments, the disclosed base editor comprises a napDNAbp containing a compact Cas9 ortholog derived from Neisseria meningitidis (Nme or Nme2). In some embodiments, the napDNAbp comprises Nme2Cas9. In some embodiments, the disclosed base editor comprises an Nme2Cas9 nickase. Nme2Cas9 recognizes simple dinucleotides PAM, NNNNCC, or N4CC (wherein N is any nucleotide), as described in Edraki et al., Molecular Cell 73,714-726, incorporated herein by reference. The sequence of Nme2Cas9 is shown in SEQ ID NO: 5. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence identical to SEQ ID NO: 5 of at least 90%, at least 95%, at least 98%, or at least 99%. In some embodiments, the disclosed base editor comprises a napDNAbp containing SEQ ID NO: 5. The protein has a length of 1082 amino acids.

[0194]

[0195] The amino acid sequence of NmeCas9 is provided below, as SEQ ID NO: 6. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 6. In some embodiments, the disclosed base editor comprises a napDNAbp comprising SEQ ID NO: 6. The protein is 1083 amino acids in length.

[0196]

[0197] In some embodiments, the disclosed base editor comprises a napDNAbp containing a compact Cas9 ortholog derived from Campylobacter jejuni (CjCas9). In some embodiments, the napDNAbp comprises CjCas9. In some embodiments, the disclosed base editor comprises a CjCas9 nickase. CjCas9 recognizes NNNNACA and NNNNACACPAM. See Kim et al., Nature Communications 8(14500):1-12 (2017), which is incorporated herein by reference. The sequence of CjCas9 (nickase) is shown in SEQ ID NO: 379. In some embodiments, the disclosed base editor comprises a napDNAbp domain having a sequence that is at least 90%, at least 95%, at least 98% or at least 99% identical to SEQ ID NO: 379. In some embodiments, the disclosed base editor comprises a napDNAbp containing SEQ ID NO: 379. The protein is 984 amino acids long.

[0198] (SEQ ID NO: 379).

[0199] Additional compact Cas proteins

[0200] In various embodiments, nucleic acid programmable DNA binding proteins include, but are not limited to, compact variants of Cas9 (e.g., dCas9 and nCas9) (CasX, CasY, Cpf1, C2c1, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG), loop-replacement Cas9 domains (e.g., CP1012, CP1028, CP1041, CP1249, and Csn2), and Cas9-like structures. P1300), Argonaute (Ago) domain, Cas9-KKH, SmacCas9, SpRY, SpRY-HF1, Spy-macCas9, SpCas9-VRQR, SpCas9-VRER, SpCas9-VQR, SpCas9-E QR, SpCas9-NRRH, SpaCas9-NRTH, SpCas9-NRCH, LbCas12a, AsCas12a, CeCas12a, MbCas12a, CasΦ, SpCas9-NG-CP1041, SpCas9-NG-VRQR.

[0201] In other embodiments, the napDNAbp may comprise a compact Cas9 ortholog from: Staphylococcus lugdunensis Cas9 (SlugCas9), Staphylococcus ottersensis Cas9 (SlutrCas9), or Staphylococcus haemolyticus Cas9 (ShaCas9). See Hu et al., Nucleic Acids Research, 49(7), April 2021, 4008-4019, which is incorporated herein by reference. SlugCas9, SlutrCas9, and ShaCas9 proteins recognize NNGG, NNGG / NNGA, and NNGG PAMs, respectively.

[0202] In other embodiments, the Cas protein can include any CRISPR-associated protein, including but not limited to Cas12a, Cas12b, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2. Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified forms thereof, and preferably comprises a nickase mutation (e.g., a mutation corresponding to the D10A mutation of the wild-type SpCas9 polypeptide of SEQ ID NO: 326).

[0203] In certain embodiments, the base editors contemplated herein may include a Cas9 protein having a smaller molecular weight than a typical SpCas9 sequence. In some embodiments, smaller sized Cas9 variants may facilitate delivery to cells, for example, via AAV vectors, expression vectors, or other modes of delivery. A typical SpCas9 protein is 1368 amino acids in length and has a predicted molecular weight of 158 kilodaltons. As used herein, the term "small-sized Cas9 variant" refers to any Cas9 variant, naturally occurring, engineered, or otherwise, having less than about 1300 amino acids, or at least less than 1290 amino acids, or less than 1280 amino acids, or less than 1270 amino acids, or less than 1260 amino acids, or less than 1250 amino acids, or less than 1240 amino acids, or less than 1230 amino acids, or less than 1220 amino acids, or less than 1210 amino acids, or less than 1200 amino acids, or less than 1190 amino acids, or less than 1180 amino acids, or less than 1170 amino acids, or less than 1160 amino acids. amino acids, or less than 1150 amino acids, or less than 1140 amino acids, or less than 1130 amino acids, or less than 1120 amino acids, or less than 1110 amino acids, or less than 1050 amino acids, or less than 1000 amino acids, or less than 950 amino acids, or less than 900 amino acids, or less than 850 amino acids, or less than 800 amino acids, or less than 750 amino acids, or less than 700 amino acids, or less than 650 amino acids, or less than 600 amino acids, or less than 550 amino acids, or less than 500 amino acids but at least greater than 400 amino acids, and retains the desired function of the Cas9 protein.

[0204] In various embodiments, the base editors disclosed herein can comprise one of the small-sized Cas9 variants described below, or a Cas9 variant thereof that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference small-sized Cas9 protein. Exemplary small-sized Cas9 variants include, but are not limited to, SauriCas9, SaCas9, CjCas9, Nme2Cas9, AsCas12a, and LbCas12a.

[0205] In some embodiments, the napDNAbp domain of any disclosed base editor comprises LbCas12a, such as wild-type LbCas12a. In some embodiments, the napDNAbp domain of any disclosed base editor comprises at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 381. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 381.

[0206] In some embodiments, the napDNAbp domain of any disclosed base editor comprises AsCas12a, such as wild-type AsCas12a. In some embodiments, the napDNAbp domain of any disclosed base editor comprises a mutant AsCas12a, such as an engineered AsCas12a or enAsCas12a. In some embodiments, the napDNAbp domain of any disclosed base editor comprises at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 383. In some embodiments, the napDNAbp domain of any disclosed base editor comprises the amino acid sequence of SEQ ID NO: 383.

[0207]

[0208]

[0209]

[0210]

[0211] Additional exemplary Cas9 equivalent protein sequences may include the following:

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225] The base editor described herein may also include a Cas12a / Cpf1 (dCpf1) variant, which can be used as a guide nucleotide sequence programmable DNA binding protein domain. The Cas12a / Cpf1 protein has a RuvC-like nuclease domain similar to the RuvC domain of Cas9 but does not have an HNH nuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. Zetsche et al., Cell, 163, 759–771, 2015 (which is incorporated herein by reference) shows that the RuvC-like domain of Cpf1 is responsible for cutting two DNA chains, and inactivation of the RuvC-like domain inactivates the Cpf1 nuclease activity.

[0226] Cytidine deaminase domain

[0227] In some embodiments, the base editor comprises a deaminase that is a cytosine deaminase. In some embodiments, the cytosine deaminase domain is fused to the N-terminus of the napDNAbp domain.

[0228] In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is an APOBEC1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, an APOBEC3H deaminase, or an APOBEC4 deaminase. In some embodiments, the deaminase is an activation-inducible deaminase (AID). In some embodiments, the deaminase is a lamprey CDA1 (pmCDA1) deaminase. In some embodiments, the deaminase is from a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase is from a human. In some embodiments, the deaminase is from a rat. In some embodiments, the deaminase is a human APOBEC1 deaminase. In some embodiments, the deaminase is pmCDA1. In some embodiments, the adenosine deaminase is human APOBEC3G. In some embodiments, the deaminase is a human APOBEC3G variant. In some embodiments, the deaminase is rat APOBEC1.

[0229] In certain embodiments, the cytidine deaminase domain is a "FERNY" polypeptide having an amino acid sequence according to SEQ ID NO: 393, or an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to SEQ ID NO: 393, as follows:

[0230] MFERNYDPRELRKETYLLYEIKWGKSGKLWRHWCQNNRTQHAEVYFLENIFNARRFNPSTHCSITWYLSWSPCAECSQKIVDFLKEHPNVNLEIYVARLYYHEDERNRQGLRDLVNSGVTIRIMDLPDYNYCWKTFVSDQGGDEDYWPGHFAPWIKQYSLKL(SEQ ID NO: 393)

[0231] In certain other embodiments, the cytidine deaminase domain is a domain evolved from a wild-type domain, e.g., by phage-assisted continuous evolution (PACE). In some embodiments, the cytidine deaminase domain is an "evoFERNY" polypeptide having an amino acid sequence according to SEQ ID NO: 394, or an amino acid sequence that is at least 85%, 90%, 95%, 98%, 99%, or 99.5% identical to SEQ ID NO: 394, which contains H102P and D104N substitutions relative to SEQ ID NO: 393, as follows:

[0232] MFERNYDPRELRKETYLLYEIKWGKSGKLWRHWCQNNRTQHAEVYFLENIFNARRFNPSTHCSITWYLSWSPCAECSQKIVDFLKEHPNVNLEIYVARLYY P E N ERNRQGLRDLVNSGVTIRIMDLPDYNYCWKTFVSDQGGDEDYWPGHFAPWIKQYSLKL (SEQ ID NO: 394). The FERNY and evoFERNY deaminase domains are 162 amino acids long. Thus, these domains are shorter than any of the domains in SEQ ID NOs: 276-277, 281, 133-134, 292-295, and 487.

[0233] In exemplary embodiments, the disclosed CBE comprises a FERNY or evoFERNY deaminase domain. In some embodiments, the disclosed CBE comprises a deaminase domain comprising an amino acid sequence of SEQ ID NO: 393 or 394.

[0234] The most advanced cytosine base editor BE3.9 contains the rat APOBEC1 (rAPOBEC1) cytidine deaminase domain. They have high overall activity, severely impaired activity in editing GC targets, and high editing of TC targets. Alternative deaminases have been shown to be base editors. Both AID and CDA work well on GC targets, but are generally less active than APOBEC1. APOBEC3G is less effective than all of these (see Komor, AC et al. Improved baseexcision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity. Sci Adv 3, eaao4774 (2017), which is incorporated herein by reference). TARGET-AID base editing was achieved using CDA (Nishida, K. et al. Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science 353, aaf8729–aaf8729 (2016), which is incorporated herein by reference). "FERNY" was reconstructed based on an ancestral sequence with N- and C-terminal truncations of the APOBEC family phylogenetic tree. rAPOBEC1: 229 aa; FERNY: 161 aa. The sequence similarity to rAPOBEC1 is 55%. The evolved FERNY genotype also has high GC activity, and despite being a shorter protein, its activity is comparable to that of APOBEC. Evolved FERNY is described in more detail in PCT Publication No. WO 2019 / 023680, published on January 31, 2019, which is incorporated herein by reference. Additional exemplary cytidine deaminases are disclosed in PCT Publication No. WO 2021 / 108717, published on June 3, 2021.

[0235] Non-limiting examples of suitable cytosine deaminase domains are provided below, such as SEQ ID NOs: 276-277, 281, 133-134, 292-295, and 487. In some embodiments, the deaminase is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences listed below.

[0236] Human AID

[0237] MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL(SEQ ID NO: 276)

[0238] Mouse AID

[0239] MDSLLMKQKKFLYHFKNVRWAKGRHETYLCYVVKRRDSATSCSLDFGHLRNKSGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVAEFLRWNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIGIMTFKDYFYCWNTFVENRERTFKAWEGLHENSVRLTRQLRRILLPLYEVDDLRDAFRMLGF(SEQ ID NO: 277)

[0240] Rat APOBEC-3 (rAPOBEC3)

[0241] MGPFCLGCSHRKCYSPIRNLISQETFKFHFKNLRYAIDRKDTFLCYEVTRKDCDSPVSLHHGVFKNKDNIHAEICFLYWFHDKVLKVLSPREEFKITWYMSWSPCFECAEQVLRFLATHHNLSLDIFSSRLYNIRDPENQQNLCRLVQEGAQVAAMDLYEFKKCWKKFVDNGGRRFRPWKKLLTNFRYQDSKLQEILRPCYIPVPSSSSSTLSNICLTKGLPETRFCVERRRVHLLSEEEFYSQFYNQRVKHLCYYHGVKPYLCYQLEQFNGQAPLKGCLLSEKGKQHAEILFLDKIRSMELSQVIITCYLTWSPCPNCAWQLAAFKRDRPDLILHIYTSRLYFHWKRPFQKGLCSLWQSGILVDVMDLPQFTDCWTNFVNPKRPFWPWKGLEIISRRTQRRLHRIKESWGLQDLVNDFGNLQLGPPMS(SEQ ID NO: 281)

[0242] Human APOBEC-3G

[0243] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCWSKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN(SEQ ID NO: 133)

[0244] Human APOBEC-3F

[0245] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPRLDAKIFRGQVYSQPEHHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYSEGQPFMPWYKFDDNYAFLHRTLKEILRNPMEAMYPHIFYFHFKNLRKAYGRNESWLCFTMEVVKHHSPVSWKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLYYFWDTDYQEGLRSLSQEGASVEIMGYKDFKYCWENFVYNDDEPFKPWKGLKYNFLFLDSKLQEILE(SEQ ID NO: 134)

[0246] Human APOBEC-1<> <>

[0247] MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR(SEQ ID NO: 292)<> <>

[0248] Mouse APOBEC-1<> <>

[0249] It should be noted that there are some angle brackets in the original text that seem to be in an incorrect format. I have translated them as they are while trying to maintain the overall structure. If there are specific requirements or corrections regarding these symbols, it may affect the translation result.MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNTRCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYPPSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK(SEQ ID NO: 293)

[0250] Rat APOBEC-1 (rAPOBEC1)

[0251] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK(SEQ ID NO: 294)

[0252] Lamprey CDA1 (pmCDA1)

[0253] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAV(SEQ ID NO:295)

[0254] Evolved pmCDA1 (evoCDA1)

[0255] MTDAEYVRIHEKLDIYTFKKQFSNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILE WYNQELRGNGHTLKIWVCKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMFQVKILHTTKSPAV(SEQ ID NO:487)

[0256] Adenosine deaminase domain

[0257] The present disclosure provides adenosine deaminase variants that are active against deoxyadenosine nucleosides in DNA. Thus, the variants provided herein are deoxyadenosine deaminases. In some embodiments, the disclosed adenosine deaminase is a variant of the known adenosine deaminase TadA7.10, which comprises the following mutations compared to wild-type ecTadA (SEQ ID NO: 325): W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, R152P, E155V, I156F, and K157N. In some embodiments, the disclosed adenosine deaminase is a variant of TadA derived from species other than Escherichia coli, such as Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis.

[0258] In some embodiments, the disclosed adenosine deaminase domain comprises TadA-8e, a variant of Escherichia coli TadA7.10. TadA-8e (set forth in SEQ ID NO: 433) contains the following substitutions relative to TadA7.10 (SEQ ID NO: 315): T111, D119, F149, R26, V88, A109, H122, T166, and D167. TadA-8e is disclosed in PCT Publication No. WO 2021 / 158921, published on August 12, 2021, which is incorporated herein by reference. In some embodiments, the adenosine deaminase domain comprises TadA-8e(V106W), which contains the V106W substitution relative to TadA-8e.

[0259] In various embodiments, the disclosed adenosine deaminases hydrolytically deaminate targeted adenosine in a nucleic acid of interest to inosine, which is read as guanosine (G) by a DNA polymerase.

[0260] These variants can comprise a domain of any disclosed base editor (i.e., an adenosine deaminase domain of an adenine base editor). In some embodiments, any disclosed adenine base editor is capable of deaminating adenosine in a nucleic acid sequence (e.g., DNA or RNA). The disclosed adenine base editor is further capable of deaminating adenine in DNA.

[0261] Provided herein are exemplary non-limiting embodiments of adenosine deaminases. In some embodiments, the adenosine deaminase domain of any disclosed base editor comprises a single adenosine deaminase or monomer. In some embodiments, the adenosine deaminase domain comprises 2, 3, 4, or 5 adenosine deaminases. In some embodiments, the adenosine deaminase domain comprises two adenosine deaminases or dimers. In some embodiments, the deaminase domain comprises an engineered (or evolved) deaminase and a dimer of a wild-type deaminase (such as a wild-type E. coli-derived deaminase). It should be understood that the mutations provided herein (e.g., mutations in ecTadA) can be applied to adenosine deaminases in other adenine base editors, such as those provided in: PCT Publication No. WO 2018 / 027078, published on August 2, 2018; PCT Publication No. WO 2019 / 079347, published on April 25, 2019; International Application No. PCT / US2019 / 033848, filed on May 23, 2019, which was published as PCT Publication No. WO 2019 / 033848 on November 28, 2019. 2019 / 226593; U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018, which issued as U.S. Patent No. 10,113,163 on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published on May 4, 2017, which issued as U.S. Patent No. 10,167,457 on January 1, 2019; PCT Publication No. WO 2017 / 0121693, published on April 27, 2017 2017 / 070633; U.S. Patent Publication No. 2015 / 0166980, published on June 18, 2015; U.S. Patent No. 9,840,699, issued on December 12, 2017; and U.S. Patent No. 10,077,453, issued on September 18, 2018; PCT Application No. PCT / US2020 / 28568, filed on April 16, 2020, and PCT Publication No. WO 2021 / 158921, published on August 12, 2021; all of which are incorporated herein by reference in their entirety.

[0262] In some embodiments, any adenosine deaminase provided herein can deaminate adenine, for example, deaminate adenine in the deoxyadenosine nucleoside of DNA. Adenosine deaminase can be derived from any suitable organism (e.g., E. coli). In some embodiments, adenosine deaminase is a naturally occurring adenosine deaminase comprising one or more mutations corresponding to any mutation provided herein (e.g., mutations in ecTadA). Those skilled in the art will be able to identify the corresponding residues in any homologous protein and corresponding encoding nucleic acid by methods well known in the art, for example, by sequence alignment and determination of homologous residues. Exemplary TadA deaminases derived from Bacillus subtilis (complete as shown in SEQ ID NO: 318), Staphylococcus aureus (SEQ ID NO: 317) and Streptococcus pyogenes (SEQ ID NO: 448) are provided. Amino acid replacements in E. coli TadA-8e and homologous mutations in Bacillus subtilis, Staphylococcus aureus and Streptococcus pyogenes are shown. Thus, one skilled in the art will be able to generate mutations corresponding to any of the mutations described herein in any naturally occurring adenosine deaminase (e.g., having homology to ecTadA), e.g., any mutation identified in ecTadA. In some embodiments, the adenosine deaminase is derived from a prokaryotic organism. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from Escherichia coli.

[0263] In some embodiments, the adenosine deaminase domain comprises TadA9 or a variant thereof. TadA9 contains the V82S and Q154R substitutions relative to TadA-8e. (In other words, TadA9 contains the Y147R, Q154R, and I76Y mutations relative to TadA7.10.) In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA9 (SEQ ID NO: 33). TadA9 may be referred to in the art as TadA*8.9. ABEs containing the TadA9 deaminase are referred to herein as ABE9. Further details of TadA9 are described in Gaudelli et al., Nat Biotechnol. 2020 Jul;38(7):892-900 and PCT Publication No. WO 2021 / 050571 published on March 8, 2021, each of which is incorporated herein by reference.

[0264] In some embodiments, the adenosine deaminase domain comprises TadA20 or a variant thereof. TadA20 contains I76Y, V82S, Y123H, Y147R, and Q154R substitutions relative to TadA7.10. More details of TadA20 are described in Gaudelli et al., Nat Biotechnol. 2020 Jul;38(7):892-900 and WO 2021 / 050571 published on March 18, 2021. TadA20 may be referred to in the art as TadA*8.20. In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of TadA20 (SEQ ID NO: 326). ABEs containing TadA20 deaminase are referred to herein as ABE20. It may be referred to in the art as ABE8.20, ABE8.20-d or ABE8.20-m.

[0265] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of any one of SEQ ID NOs: 33, 315, 317-326, 433, or 448-449.

[0266] It should be understood that the adenosine deaminases provided herein can include one or more mutations (e.g., any mutation provided herein). The present disclosure provides adenosine deaminases having a certain percentage of identity plus any mutation described herein or a combination thereof. Any adenosine deaminase described herein can be a truncated variant of any other adenosine deaminase described herein, for example, any of the adenosine deaminases of SEQ ID NOs: 33, 315, 317-326, 433, or 448-449.

[0267] Exemplary truncated adenosine deaminases can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more than 15 amino acids truncated from the N-terminus. Other exemplary truncated adenosine deaminases can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more than 15 amino acids truncated from the C-terminus. In some embodiments, the adenosine deaminase domain comprises a truncated version of wild-type ecTadA, as set forth in SEQ ID NO: 324. Any of the adenosine deaminases described herein can include an N-terminal methionine (M) amino acid residue.

[0268] It will be understood that any mutation provided herein (e.g., based on the ecTadA amino acid sequence of SEQ ID NO: 315) can be introduced into other adenosine deaminases, such as Staphylococcus aureus TadA (saTadA), hyperthermophilic TadA (AaTadA), or another adenosine deaminase (e.g., another bacterial adenosine deaminase), such as those sequences provided below. It will be apparent to those skilled in the art how to identify amino acid residues homologous to the mutated residues in ecTadA from other adenosine deaminases. Thus, any mutation identified in ecTadA can be generated in other adenosine deaminases having homologous amino acid residues. Any mutation provided herein can be generated in ecTadA or another adenosine deaminase, alone or in any combination. Any mutant deaminase provided herein can be used in the context of an adenine base editor. Any deaminase provided herein can have a sequence that begins with methionine ("M") before the first amino acid shown in the following sequence.

[0269] Exemplary adenosine deaminase variants of the present disclosure are described below. In certain embodiments, the adenosine deaminase domain comprises an adenosine deaminase having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% sequence identity to one of the following sequences:

[0270] TadA 7.10 (Escherichia coli)

[0271] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD(SEQ ID NO: 315)

[0272] TadA-8e (Escherichia coli)

[0273] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO: 433)

[0274] TadA9

[0275] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIGLYDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMDHRVEITEGILANECAALLCDFYRMPRQVFNAQKKAQSSIN(SEQ ID NO: 33)

[0276] TadA20

[0277] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLYDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSTD(SEQ ID NO: 326)

[0278] Staphylococcus aureus TadA:

[0279] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN(SEQ ID NO: 317)

[0280] Bacillus subtilis TadA:

[0281] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE(SEQ ID NO: 318)​​​​

[0283] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV(SEQ ID NO: 319)

[0284] Shewanella putrefaciens TadA:

[0285] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE(SEQ ID NO: 320)

[0286] Haemophilus influenzae F3031 TadA:

[0287] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK(SEQ ID NO: 321)

[0288] Caulobacter crescentus TadA:

[0289] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI(SEQ ID NO: 322)

[0290] Geobacter sulfurreducens TadA:

[0291] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP (SEQ ID NO: 323)

[0292] Streptococcus pyogenes TadA

[0293] MPYSLEEQTYFMQEALKEAEKSLQKAEIPIGCVIVKDGEIIGRGHNAREESNQAIMHAEIMAINEANAHEGNWRLLDTTLFVTIEPCVMCSGAIGLARIPHVIYGASNQKFGGADSLYQILTDERLNHRVQVERGLLAADCANIMQTFFRQGRERKKIAKHLIKEQSDPFD (SEQ ID NO: 448)

[0294] Aquifex aeolicus TadA

[0295] MGKEYFLKVALREAKRAFEKGEVPVGAIIVKEGEIISKAHNSVEELKDPTAHAEMLAIKEACRRLNTKYLEGCELYVTLEPCIMCSYALVLSRIEKVIFSALDKKHGGVVSVFNILDEPTLNHRVKWEYYPLEEASELLSEFFKKLRNNII (SEQ ID NO: 449)

[0296] In some embodiments, the adenosine deaminase domain comprises an N-terminally truncated E. coli TadA. In certain embodiments, the adenosine deaminase comprises the amino acid sequence: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 324).

[0297] In some embodiments, the TadA deaminase is full-length Escherichia coli TadA deaminase (ecTadA). For example, in certain embodiments, the adenosine deaminase domain comprises a deaminase comprising the amino acid sequence:

[0298] MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD(SEQ ID NO: 325)

[0299] In some embodiments, the base editor comprises an adenosine deaminase monomer. In other aspects, the base editor comprises an adenosine deaminase dimer. The base editor may comprise a heterodimer of a first adenosine deaminase and a second adenosine deaminase. In some embodiments, the first adenosine deaminase is located at the N-terminus of the second adenosine deaminase in the base editor. In some embodiments, the first adenosine deaminase is located at the C-terminus of the second adenosine deaminase in the base editor. In some embodiments, the first adenosine deaminase and the second deaminase are fused directly to each other or via a linker. In some embodiments, the first adenosine deaminase is fused to the N-terminus of napDNAbp via a linker, and the second deaminase is fused to the C-terminus of the napDNAbp domain via a linker. In other embodiments, the second adenosine deaminase is fused to the N-terminus of the napDNAbp domain via a linker, and the first deaminase is fused to the C-terminus of napDNAbp via a linker.

[0300] Exemplary base editors

[0301] Adenine base editor

[0302] In some aspects, the base editing methods of the present disclosure include the use of adenine base editors. Exemplary adenine base editors disclosed herein include monomeric and dimeric forms of the following editors: Sauri-ABE8e, SaKKH-ABE8e, SaABE8e, CjCas9-ABE8e, and Nme2Cas9-ABE8e; SaKKH-ABE8e(V106W), SauriCas9-ABE8e(V106W), CjCas9-ABE8e(V106W), Nme2Cas9-ABE8e(V106W), and SaCas9-ABE8e(V106W); SaKKH-AB E9, SauriCas9-ABE9, CjCas9-ABE9, Nme2Cas9-ABE9, and SaCas9-ABE9; SaKKH-ABE20, SauriCas9-ABE20, CjCas9-ABE20, Nme2Cas9-ABE20, and SaCas9-ABE20; and SaKKH-ABE7.10, SauriCas9-ABE7.10, CjCas9-ABE7.10, Nme2Cas9-ABE7.10, and SaCas9-ABE7.10. In some embodiments, ABE is Sauri-ABE8e, SaKKH-ABE8e, or SaABE8e. The lengths of these base editors are 1298, 1291, and 1291 amino acids, respectively. The lengths of the other base editors are 1221 amino acids (CjABE8e) and 1319 amino acids (Nme2ABE8e). In an exemplary embodiment, the ABE is SaKKH-ABE8e. An exemplary ABE contains an adenosine deaminase domain comprising TadA8e and does not comprise a second adenosine deaminase (ie, the adenosine deaminase domain consists of a deaminase monomer).

[0303] Additional exemplary adenine base editors are disclosed in PCT Publication No. WO 2020 / 051360, published on March 12, 2020; PCT Publication No. WO 2021 / 050571; PCT Publication No. WO 2020 / 168132, published on August 20, 2020; and PCT Publication No. WO 2021 / 158921, published on August 12, 2021, each of which is incorporated herein by reference.

[0304] ABE8e may be referred to in the art as "ABE8" or "ABE8.0". The ABE8e base editor and its variants may comprise an adenosine deaminase domain comprising a TadA-8e adenosine deaminase monomer (monomer form) or a TadA-8e adenosine deaminase homodimer or heterodimer (dimer form). In some embodiments, the architecture of the base editor comprising an adenosine deaminase domain and napDNAbp is as follows: NH2-[adenosine deaminase]-[napDNAbp domain]-COOH; or NH2-[napDNAbp domain]-[adenosine deaminase]-COOH. In certain embodiments, the base editor comprises an ABE8e monomer architecture comprising NH2-[NLS]-[adenosine deaminase]-[napDNAbp domain]-[NLS]-COOH, wherein "NLS" is a nuclear localization sequence.

[0305] In some aspects, the present disclosure provides a complex of an adenine base editor and a guide RNA. Exemplary disclosed complexes include any one of the following ABEs bound to a guide RNA (e.g., a single guide RNA): Sauri-ABE8e, SaKKH-ABE8e, SaABE8e, CjCas9-ABE8e, and Nme2Cas9-ABE8e; SaKKH-ABE8e (V106W), SauriCas9-ABE8e (V106W), CjCas9-ABE8e (V106W), Nme2Cas9-ABE8e (V106W), and SaCas9-ABE8e (V106W); SaKK Other ABEs can be used to deaminize A nucleobases according to the disclosed complexes.

[0306] In some embodiments, the ABE is CjCas9-ABE8e. During adenine editing evaluation, this editor exhibited a higher background preference for pyrimidines in the 5' nucleotide position of the target adenine base ( Y A >> RA). As used herein, "preference" and "background preference" refer to a product purity greater than 40% relative to the target adenosine. Thus, in some aspects, the present disclosure provides ABEs with a background preference for pyrimidines ("Y"), where "background" refers to the presence of pyrimidines or purines ("R") immediately 5' to the adenine base to be edited (or the target adenine base), such as the CjABE8e editor and its variants. These ABEs can have a preference for editing adenosine in a target nucleic acid sequence of 5'-YAN-3', where Y is C or T; N is A, T, C, G, or U; and A is the target adenosine. Thus, in some embodiments, an ABE is provided that has a background preference for deaminating adenosine in a target nucleic acid sequence of 5'-YAN-3', where Y is C or T, and N is A, T, C, G, or U; and A is the target adenosine.

[0307] Exemplary AAV-encoded adenine base editor constructs are shown in Figure 2A This construct contains the SaABE8e base editor operably controlled by the EFS promoter and the bGH poly A sequence. It also contains a reverse-encoded guide RNA, as indicated by the arrow pointing away from the 3' end.

[0308] The disclosed ABE complexes may have an on-target editing efficiency greater than 50% upon contact with a nucleic acid molecule comprising a target sequence. Additional exemplary ABE complexes may have an on-target editing efficiency greater than 60% upon contact with a nucleic acid molecule comprising a target sequence. Additional exemplary ABEs may have an on-target editing efficiency greater than 65%, greater than 70%, greater than 75%, greater than 80%, greater than 82.5%, or greater than 85% upon contact with a nucleic acid molecule comprising a target sequence. The disclosed ABE complexes may exhibit an indel frequency of less than 2.5%, less than 2.4%, less than 2.2%, less than 2.0%, less than 1.75%, less than 1.5%, less than 1.3%, less than 1.1%, or less than 1.0% upon contact with a nucleic acid molecule comprising a target sequence.

[0309] In some aspects, the present disclosure provides a base editor comprising a napDNAbp domain and an adenosine deaminase domain as described herein. The Cas9 domain can be any Cas9 domain or Cas9 protein (e.g., Cas9 nickase or nCas9). In some embodiments, any Cas9 domain or Cas9 protein (e.g., nCas9) provided herein can be fused to any adenosine deaminase domain provided herein.

[0310] In some embodiments, the base editor comprising an adenosine deaminase and a napDNAbp (e.g., a Cas9 domain) does not include a linker sequence. In some embodiments, a linker is present between the adenosine deaminase domains and / or between the adenosine deaminase and the napDNAbp. In some embodiments, the "]-[" used in the general framework above indicates the presence of an optional linker. In some embodiments, the adenosine deaminase domain and the napDNAbp domain are fused via any linker provided herein. For example, in some embodiments, the adenosine deaminase domain (which may include one or more adenosine deaminases) and the napDNAbp are fused via any linker provided in the section entitled "Linker" below.

[0311] In some embodiments, the adenine base editor comprises an adenosine deaminase comprising a sequence having at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% sequence identity to SEQ ID NO: 433 (TadA-8e). In some embodiments, the adenine base editor comprises the sequence of SEQ ID NO: 433.

[0312] In some embodiments, a disclosed adenine base editor comprises a sequence of SEQ ID NO: 181. In some embodiments, a disclosed adenine base editor comprises a sequence of SEQ ID NO: 182. In some embodiments, a disclosed adenine base editor comprises a sequence of SEQ ID NO: 183. In some embodiments, a disclosed adenine base editor comprises a sequence of SEQ ID NO: 171 or 172.

[0313] In some embodiments, any of the adenine base editors described herein may comprise an amino acid sequence that differs from the amino acid sequence of any one of SEQ ID NOs: 171-172 and 181-183 by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more than 30 amino acids. These differences may include amino acids that have been inserted, deleted, or substituted relative to the reference sequence. In some embodiments, the disclosed adenosine deaminase domain contains a fragment that shares about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 300, about 400, about 500, or more than 500 consecutive amino acids with any one of SEQ ID NOs: 171-172 and 181-183.

[0314] Exemplary base editors comprise a sequence that is at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% identical to any of the following amino acid sequences (SEQ ID NOs: 171-172 and 181-183). In some embodiments, the disclosed base editors have a sequence comprising any of the following amino acid sequences:

[0315] SaABE8e: NLS , connector, TadA-8e, SaCas9

[0316] MKRTADGSEFESPKKKRKV SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGS GKR NYILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTD HSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQL ERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYE MLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNE EDIKGYRVTTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQI SNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIK VINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCL YSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNL AKGKGRISKTKKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRK WKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHI KDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNLNLGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQK LKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFD VYLDNGVYKFVTVKNLDVIKKENNYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNR IEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG SGGS CHRISTMAS DAY PKKKRKV (SEQ ID NO: 171)

[0317] SaKKH-ABE8e: NLS , connector, TadA-8e, SaKKH

[0318] MKRTADGSFESPKKKRKV SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGS GKR NYILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTD HSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQL ERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKWYE MLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNE EDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQI SNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIK VINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCL YSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNL AKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRK WKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHI KDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQK LKLIMEQYGDEKNPLYKYYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFD VYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYKNDLIKINGELYRVIGVNNDLLNR IEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG SGGS KRTADGSEFE PKKKRKV (SEQ ID NO: 181)

[0319] SauriABE8e: NLS , linker, TadA-8e, SauriCas9

[0320] MKRTADGSEFSPKKKRKV SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGS DOG QQKQNYILGLAIGITSVGYGLIDSKTREVIDAGVRLFPEADSENNSNRRSKRGARRLKRRRIHRLNRVKDLLADYQ MIDLNNVPKSTDPYTIRVKGLREPLTKEEFAIALLHIAKRRGLHNISVSMGDEEQDNELSTKQQLQKNAQQLQDKY VCELQLERLTNINKVRGEKNRFKTEDFVKEVKQLCETQRQYHNIDDQFIQQYIDLVSTRREYFEGPGNGSPYGWDG DLLKWYEKLMGRCTYFPEELRSVKYAYASADLFNALNDLNNLVVTRDDNPKLEYYEKYHIIENVFKQKKNPTLKQIA KEIGVQDYDIRGYRITKSGKPQFTSFKLYHDLKNIFEQAKYLEDVEMLDEIAKILTIYQDEISIKKALDQLPELLT ESEKSQIAQLTGYTGTHRLSLKCIHIVIDELWESPENQMEIFTRLNLKPKKVEMSEIDSIPTTLVDEFILSPVKR AFIQSIQUINAVINRFGLPEDIIIELAREKNSKDRRKFINKLQKQNEATRKKIEQLLAKYGNTNAKYMIEKIKLHD MQEGKCLYSLEAIPLEDLLSNPTHYEVDHIIPRSVSFDNSLNNKVLVKQSENSKKGNRTPYQYLSSNESKISYNQF KQHILNLSKADRISKKKRDMLLEERDINKFEVQKEFINRNLVDTRYATRELSNLLKTYFSTHDYAVKVKTINGGF TNHLRKVWDFKKHRNHGYKHHAEDALVIANADFLFKTHKALRRTDKILEQPGLEVNDTTVKVDTEEKYQELFETPK QVKNIKQFRDFKYSHRVDKKPNRQLINDTLYSTREIDGTYVVQTLKDLYAKDNEKVKKLFTERPQKILMYQHDPK TFEKLMTILNQYAEAKNPLAAYYEDKGEYVTKYAKKGNGPAIHKIKYIDKKLGSYLDVSNKYPETQNKLVKLSLKS FRFDIYKCEQGYKMVSIGYLDVLKKDNYYYIPKDKYEAEKQKKKIKESDLFVGSFYYNDLIMYEDELFRVIGVNSD INNLVELNMVDITYKDFCEVNNVTGEKRIKKTIGKRVVLIEKYTTDILGNLYKTPLPKKPQLIFKRGEL SGGS KRT ADGSEFEKKKRKV (SEQ ID NO: 172)

[0321] CjABE8e: NLS , linker, TadA-8e, CjCas9

[0322] MKRTADGSFESPKKKRKVSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGS ARI LAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLAPRRLARSARKRLARRKARNLNHLKHLIANEFKLNY EDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKL ANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFSKKFEEVLSVAFY KRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQ TKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDS LSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPVVLRAIKEYR KVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEF CAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLP TKKQKRILDKNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTS ALRHTWGFSAKDRNNHLHHAIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQ KVLDKIDEIFVSKPERKKPSGALHEETFRKEEEFYQSYGGKEGVLKALELGKIRKVNGKIVKNGDMFRVDIFKHKK TNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKDMQEPEFVYYNAFTSST VSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKSIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKK SGGS CRTA DGSEFEPKKKRKV (SEQ ID NO: 182)

[0323] Nme2ABE8e: NLS , linker, TadA-8e, Nme2Cas9

[0324] MKRTADGSFESPKKKRKV SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGS AAF KPNPINYILGLAIGIASVGWAMVEIDEENPIRLIDLGVRVFERAEVPKTGDSLAMARRLARSVRRLTRRRAHRLL RARRLLKREGVLQAADFDENGLIKSLPNTPWQLRAAALDRKLTPLEWSAVLLHLIKHRGYLSQRKNEGETADKELG ALLKGVANNAHALQTGDFRTPAELALNKFEKESGHIRNQRGDYSHTFSRKDLQAELILLFEKQKEFGNPHVSGGLK EGIETLLMTQRPALSGDAVQKMLGHCTFEPAEPKAAKNTYTAERFIWLTKLNNLRILEQGSERPLTDTERATLMDE PYRKSKLTYAQARKLLGLEDTAFKGLRYGKDNAEASTLMEMKAYHAISRALEKEGLKDKKSPLNLSSELQDEIGT AFSLFKTDEDITGRLKDRVQPEILEALLKHISFDKFVQISLKALRRIVPLMEQGKRYDEACAEIYGDHYGKKNTEE KIYLPPIPADEIRNPVVLRALSQARKVINGVVRRYGSPARIHIETAREVGKSFKDRKEIEKRQEENRKDREKAAAK FREYFPNFVGEKPSKDILKLRYEQQHGKCLYSGKEINLVRLNEKGYVEIDHALPFSRTWDDSFNNKVLVLGSENQ NKGNQTPYEYFNGKDNSREWQEFKARVETSRFPRSKKQRILLQKFDEDGFKECNLNDTRYVNRFLCQFVADHILLT GKGKRRVFASNGQITNLLRGFWGLRKVRAENDRHHALDAVVVACSTVAMQQKITRFVRYKEMNAFDGKTIDKETGK VLHQKTHFPQPWEFFAQEVMIRVFGKPDGKPEFEEADTPEKLRTLLAEKLSSRPEAVHEYVTPLFVSRAPNRKMSG AHKDTLRSAKRFVKHNEKISVKRVWLTEIKLADLENMVNYKNGREIELYEALKARLEAYGGNAKQAFDPKDNPFYK KGGQLVKAVRVEKTQESGVLLNKKNAYTIADNGDMVRVDVFCKVDKKGKNQYFIVPIYAWQVAENILPDIDCKGYR IDDSYTFCFSLHKYDLIAFQKDEKSKVEFAYYINCDSSNGRFYLAWHDKGSKEQQFRISTQNLVLIQKYQVNELGK EIRPCRLKKRPPVR SGGS KRTADGSEFEPKKKRKV (SEQ ID NO: 183)

[0325] Cytosine base editors

[0326] In some aspects, the present disclosure provides cytosine base editors (CBEs). Examples of these CBEs include CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9. The CBEs disclosed herein may be variants of any one of CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9.

[0327] In some embodiments, the disclosed cytosine base editor (CBE) may comprise a fusion protein comprising: (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain; (ii) a cytidine deaminase domain; and (iii) a uracil glycosylase inhibitor domain (UGI). In various embodiments, the disclosed CBE contains a single UGI domain. The disclosed CBE can be structurally arranged in a variety of configurations, including but not limited to:

[0328] NH2-[cytidine deaminase domain]-[napDNAbp domain]-[UGI]-COOH;

[0329] NH2-[cytidine deaminase domain]-[UGI]-[napDNAbp domain]-COOH;

[0330] NH2-[napDNAbp domain]-[UGI]-[cytidine deaminase domain]-COOH;

[0331] NH2-[napDNAbp domain]-[cytidine deaminase domain]-[UGI]-COOH;

[0332] NH2-[UGI]-[cytidine deaminase domain]-[napDNAbp domain]-COOH; or

[0333] NH2-[UGI]-[napDNAbp domain]-[cytidine deaminase domain]-COOH, wherein each instance of "]-[" includes an optional linker.

[0334] An exemplary construct encoding a single AAV-CBE is shown in FIG. Figure 14As shown. The CBE contains the BE3.9 architecture, in which the deaminase FERNY (or evoFERNY) serves as the cytidine deaminase domain located 5' of the napDNAbp domain. In this construct, the cytosine base editor contains the CjCas9 napDNAbp domain. The human U6-controlled ("hU6") sgRNA is located at the 3' end, in the opposite direction of the base editor. The construct contains the EFS promoter and the SV40 late poly A sequence that drive the base editor. The length of the construct is 5.012 kb. Where specified, "BE3.9" refers to the BE3.9 CBE architecture, i.e., NH2-[first nuclear localization sequence]-[cytosine deaminase domain]-[32aa linker]-[napDNAbp domain]-[9aa linker]-[first UGI domain]-[second nuclear localization sequence]-COOH. The disclosed CBE may further comprise one or more nuclear localization signals (NLS).

[0335] Additional exemplary CBEs are disclosed in PCT Publication Nos. WO 2019 / 023680, published on January 31, 2019, and WO 2021 / 108717, published on June 3, 2021, each of which is incorporated herein by reference. Additional CBEs are disclosed in Villiger, L. et al. Nature Medicine 24, 1519-1525 (2018), incorporated herein by reference. Villiger and colleagues developed an intein-split Staphylococcus aureus CBE. The disclosed single AAV CBE exhibited comparable or even higher activity than that demonstrated by Villiger.

[0336] The disclosed CBEs may comprise modified (or evolved) cytosine deaminase domains, such as deaminase domains that recognize expanded PAM sequences, have improved efficiency in deaminating 5'-GC targets, and / or edit within a narrower target window. In some embodiments, the disclosed cytosine nucleobase editors comprise evolved nucleic acid programmable DNA binding proteins (napDNAbps), such as evolved Cas9.

[0337] In some aspects, the present disclosure provides a complex of a cytosine base editor and a guide RNA, for example, a complex of any one of CjCas9-BE3.9, CjCas9-FERNY-BE3.9, and CjCas9-evoFERNY-BE3.9 and an sgRNA.

[0338] The following provides non-limiting examples of CBEs, such as SEQ ID NOs: 19 and 20, which are FERNY-BE3.9 editors. These editors contain SpCas9 as the napDNAbp domain. The exemplary AAV-encoded CBEs disclosed herein contain a CjCas9 domain (SEQ ID NO: 379) in place of the SpCas9 domain of the FERNY-BE3.9 editor described below.

[0339] The following base editor (SEQ ID NO: 19) contains wild-type FERNY, which can be used as a reference base editor. This base editor was evolved to generate the evoFERNY base editor shown in SEQ ID NO: 20. These base editors contain bpNLS and a single UGI domain.

[0340] Amino acid sequence of FERNY-BE3.9

[0341]

[0342] The following base editor contains evoFERNY, which was evolved based on the base editor provided above (SEQ ID NO: 19).

[0343] Amino acid sequence of evoFERNY-BE3.9

[0344]

[0345] In some embodiments, the disclosed base editor comprises CjCas9-FERNY-BE3.9, which is provided below as SEQ ID NO: 21. In some embodiments, the disclosed base editor comprises CjCas9-evoFERNY-BE3.9, which is provided below as SEQ ID NO: 22. Any disclosed base editor may comprise a sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identity to any one of SEQ ID NOs: 21 and 22. The base editor may comprise the sequence of SEQ ID NO: 21 or 22. These base editors contain a bpNLS and a single UGI domain. FERNY deaminase is indicated in italics.

[0346] CjCas9-evoFERNY-BE3.9

[0347] <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0348] <h2 style=";text-align:left;direction:ltr"> CjCas9-FERNY-BE3.9<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0349]

[0350] recombinant adeno-associated virus (rAAV) vector

[0351] Aspects of the present disclosure relate to the use of recombinant adeno-associated virus vectors for delivering any disclosed nucleic acid molecules. The rAAV particles disclosed herein include rAAV vectors (i.e., recombinant genomes of rAAV) encapsulated in viral capsid proteins. Referring to U.S. Patent Publication No. 2018 / 0127780 disclosed on May 10, 2018 and PCT Publication No. WO 2020 / 236982 disclosed on November 26, 2020, the disclosures of which are incorporated herein by reference.

[0352] In some embodiments, the AAV nucleic acid vector is single-stranded. In some embodiments, the AAV nucleic acid vector is self-complementary. In various embodiments, the rAAV vector of the present disclosure does not contain any intein.

[0353] In some embodiments, the viral sequences that promote integration comprise inverted terminal repeat (ITR) sequences. In some embodiments, the nucleic acid molecule is flanked by ITR sequences on each side. In some embodiments, the nucleic acid vector further comprises a region encoding an AAV Rep protein as described herein, which is contained within the region flanking the ITR or outside the region flanking the ITR. The ITR sequence can be derived from any AAV serotype (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) or can be derived from more than one serotype. In some embodiments, the ITR sequence is derived from AAV8 or AAV9. In some embodiments, in the method for packaging any of the disclosed rAAV particles, a nucleic acid plasmid, such as a helper plasmid, comprising a region encoding Rep protein and / or Cap (capsid) protein is provided.

[0354] Thus, in some embodiments, the rAAV particles disclosed herein comprise rAAV2 particles, rAAV6 particles, rAAV8 particles, rPHP.B particles, rPHP.eB particles, or rAAV9 particles, or variants thereof. In specific embodiments, the rAAV particles disclosed herein are rAAV8 or rAAV9 particles.

[0355] Exemplary rAAV particles provided herein include, but are not limited to, rAAV8-Sauri-ABE8e, rAAV9-Sauri-ABE8e, rAAV8-SaKKH-ABE8e, rAAV9-SaKKH-ABE8e, rAAV8-CjCas9-ABE8e, rAAV9-CjCas9-ABE8e, rAAV8-Nme2Cas9-ABE8e, or rAAV9-Nme2Cas9-ABE8e particles. In certain embodiments, the rAAV particles comprise rAAV8-SaKKH-ABE8e particles. In some embodiments, the rAAV particles comprise rAAV9-CjBE3.9 particles or rAAV8-CjBE3.9 particles.

[0356] ITR sequences and plasmids containing ITR sequences are known in the art and are commercially available (see, e.g., products and services available from Vector Biolabs, Philadelphia, PA; Cellbiolabs, San Diego, CA; Agilent Technologies, Santa Clara, CA; and Addgene, Cambridge, MA; and Gene delivery to skeletal muscle results in sustained expression and systemic delivery of a therapeutic protein. Kessler PD, Podsakoff GM, Chen X, McQuiston SA, Colosi PC, Matelis LA, Kurtzman GJ, Byrne BJ. Proc Natl Acad Sci USA. 1996 Nov 26;93(24):14082-7; and Curtis A. Machida. Methods in Molecular Medicine™. Viral Vectors for Gene Therapy Methods and Protocols. 10.1385 / 1-59259-304-6:201 © Humana Press Inc. 2003. Chapter 10. Targeted Integration by Adeno-Associated Virus. Matthew D. Weitzman, Samuel M. Young Jr., Toni Cathomen and Richard Jude Samulski; U.S. Patent Nos. 5,139,941 and 5,962,313, which are incorporated herein by reference in their entireties).

[0357] In some embodiments, the rAAV vectors of the present disclosure comprise one or more regulatory elements to control the expression of the heterologous nucleic acid region (e.g., promoters, transcription terminators and / or other regulatory elements). In some embodiments, the first and / or second nucleotide sequence is operably linked to one or more (e.g., 1, 2, 3, 4, 5 or more) transcription terminators. Non-limiting examples of transcription terminators that can be used according to the present disclosure include the transcription terminators (or polyadenylation signals) of the bovine growth hormone gene (bGH), the human growth hormone gene (hGH), SV40, CW3, φ, or a combination thereof. In an exemplary embodiment, the transcription terminator is an SV40 polyadenylation signal. In an exemplary embodiment, the transcription terminator does not contain a post-transcriptional response element, such as a WPRE element.

[0358] In some aspects, provided herein are methods for preparing (or manufacturing, or packaging) any disclosed rAAV particles. rAAV particles can be prepared according to any method known in the art. Packaging methods are known in the art, and reagents are commercially available (see, e.g., Zolotukhin et al. Production and purification of serotype 1, 2, and 5 recombinant adeno-associated viral vectors. Methods 28 (2002) 158–167; and U.S. Patent Publication Nos. US 2007-0015238 and US 2012-0322861, which are incorporated herein by reference; and plasmids and kits available from ATCC and Cell Biolabs, Inc.). For example, a plasmid containing a gene of interest can be combined with one or more helper plasmids (e.g., containing rep genes (e.g., encoding Rep78, Rep68, Rep52, and Rep40) and cap genes (encoding VP1, VP2, and VP3, including modified VP2 regions as described herein)) and transfected into recombinant cells so that rAAV particles can be packaged and subsequently purified.

[0359] Packaging cells (or host cells) are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells for packaging adenovirus and ψ2 cells or PA317 cells for packaging retrovirus. Viral vectors used in gene therapy are typically generated by producing cell lines that package nucleic acid vectors into viral particles. The vector typically contains the minimum viral sequences required for packaging and subsequent integration into the host, with other viral sequences replaced by expression cassettes for the polynucleotides to be expressed. The missing viral functions are typically provided indirectly by the packaging cell line. For example, AAV vectors used for gene therapy typically only have the ITR sequences from the AAV genome required for packaging and integration into the host genome. Viral DNA is packaged in a cell line that contains a helper plasmid encoding other AAV genes (i.e., rep and cap) but lacking ITR sequences. The cell line can also be infected with adenovirus as a helper. The helper virus promotes the replication of the AAV vector and the expression of the AAV genes from the helper plasmid. Due to the lack of ITR sequences, the helper plasmid is not packaged in large quantities. Adenovirus contamination can be reduced by, for example, heat treatment, to which adenovirus is more sensitive than AAV. Additional methods for delivering nucleic acids to cells are known to those skilled in the art, such as those in US 2003 / 0087817 published on May 8, 2003, PCT application No. WO 2016 / 205764 published on December 22, 2016, and PCT application No. WO 2018 / 071868 published on April 19, 2018.

[0360] In various embodiments, base editor constructs can be engineered for delivery in one or more rAAV vectors. The rAAV associated with any method and composition provided herein can be any serotype, including any derivative or pseudotype (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 2 / 1, 2 / 5, 2 / 8, 2 / 9, 3 / 1, 3 / 5, 3 / 8 or 3 / 9). rAAV may include a genetic payload to be delivered to a cell (i.e., a recombinant nucleic acid vector expressing a transgenic gene of interest, such as a full base editor carried into a cell by rAAV). rAAV can be chimeric.

[0361] As used herein, the serotype of rAAV refers to the serotype of the capsid protein of the recombinant virus. Non-limiting examples of derivatives and pseudotypes include rAAV2 / 1, rAAV2 / 5, rAAV2 / 8, rAAV2 / 9, AAV2-AAV3 hybrids, AAVrh.10, AAVrh.74, AAVhu.14, AAV3a / 3b, AAVrh32.33, AAV-HSC15, AAV-HSC17, AAVhu.37, AAVrh.8, CHt-P6, AAV2 .5, AAV6.2, AAV2i8, AAV-HSC15 / 17, AAVM41, AAV9.45, AAV6 (Y445F / Y731F), AAV2.5T, AAV-HAE1 / 2, AAV clone 32 / 83, AAVShH10, AAV2 (Y->F), AAV8 (Y733F), AAV2.15, AAV2.4, AAVM41, and AAVr3.45. A non-limiting example of derivatives and pseudotypes having a chimeric VP1 protein is rAAV2 / 5-1VP1u, which has the genome of AAV2, the capsid backbone of AAV5, and the VP1u of AAV1. Other non-limiting examples of derivatives and pseudotypes having a chimeric VP1 protein are rAAV2 / 5-8VP1u, rAAV2 / 9-1VP1u, and rAAV2 / 9-8VP1u. In some embodiments, the capsid of the disclosed rAAV particles is AAV8 (serotype 8). In some embodiments, the capsid of the disclosed rAAV particles is AAV8 (serotype 9). In some embodiments, the capsid is serotype 2, 6, PHP.B, or PHP.eB.

[0362] Provided herein are compositions comprising a plurality of any of the disclosed rAAV particles. Exemplary compositions contain any of the disclosed rAAV8 particles, rAAV9 particles, rAAV2 particles, rAAV6 particles, rAAVPHP.B particles, and rAAVPHP.eB particles.

[0363] In some embodiments, methods are disclosed for administering a composition of rAAV8 particles or rAAV9 particles by intravenous administration. In other embodiments, tissues that are not well transduced by intravenous AAV9 injection can be transduced by other existing AAV variants, such as AAV4 transduction of the lungs, or by different delivery routes, such as AAV9 transduction of kidney cells by retroureteral infusion.

[0364] AAV derivatives / pseudotypes and methods of generating such derivatives / pseudotypes are known in the art (see, e.g., Mol. Ther. 2012 Apr;20(4):699-708. doi: 10.1038 / mt.2011.287. Epub 2012 Jan 24. The AAV vector toolkit: poised at the clinical crossroads. Asokan A1, Schaffer DV, Samulski RJ.). Methods for producing and using pseudotyped rAAV vectors are known in the art (see, e.g., Duan et al., J. Virol., 75:7662-7671, 2001; Halbert et al., J. Virol., 74:1524-1532, 2000; Zolotukhin et al., Methods, 28:158-167, 2002; and Auricchio et al., Hum. Molec. Genet., 10:3075-3081, 2001).

[0365] In some aspects, the present disclosure provides compositions containing a plurality of any of the disclosed rAAV particles. In some aspects, the present disclosure provides host cells containing a plurality of any of the disclosed rAAV particles. In some embodiments, the host cell is a mammalian cell, such as a human cell. In other embodiments, the host cell is a yeast cell, a plant cell, or a bacterial cell.

[0366] Methods for delivering any of the disclosed rAAV particles and compositions, and host cells comprising rAAV particles, to target cells or target tissues are known in the art. In some embodiments, any of the disclosed rAAV particles, host cells, or compositions are delivered to a subject, such as a mammalian subject. In some embodiments, the rAAV particles are delivered to a human subject.

[0367] In some embodiments, the disclosed rAAV particles and compositions are administered to the subject with a single injection (e.g., a single systemic injection). In some embodiments, the disclosed rAAV particles and compositions are administered to the subject with multiple injections. Known rAAV particles transduce target tissues within a few days, but typically allow three to four weeks to complete transduction, genomic integration, and clearance from cells. Therefore, in some aspects, any one of the disclosed rAAV particles or compositions is administered to the subject for a period of three weeks. In some aspects, any one of the disclosed rAAV particles or compositions is administered to the subject for a period of three to four weeks.

[0368] In some embodiments, any of the disclosed rAAV particles or compositions is administered at about 10 mg / kg of subject's body weight. 15 , about 10 14 , about 10 13 , about 10 12 , about 10 11 or less than about 10 11 A therapeutically effective amount of 10 vector genomes (vg) is administered to a subject or target tissue. In some embodiments, rAAV particles are administered at a dose of 10 15 to 10 14 , 10 14 to 10 13 , 10 13 to 10 12 , 10 12 to 10 11 or 10 12 to 10 11 In some embodiments, rAAV particles are administered at a dose of 10 vg per kg. 14 to 10 11 In some embodiments, any of the disclosed rAAV particles or compositions are administered to a target tissue of a subject at a dose lower than conventional doses for dual AAV particle delivery, such as described in PCT Publication No. WO 2020 / 236982, published on November 26, 2020, and Levy, JM, et al. Nat Biomed Eng 4, 97-110 (2020).

[0369] In some aspects, the present disclosure provides a single AAV vector for delivering a base editor to a target tissue (e.g., liver, neurons, heart, muscle, or eye tissue). In some embodiments, any of the disclosed rAAV particles or compositions are administered to the liver (liver) tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the cardiac tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the neuronal tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the muscle or neuromuscular tissue of a subject. In some embodiments, any of the disclosed rAAV particles or compositions are administered to the eye tissue of a subject.

[0370] In some embodiments, the disclosed rAAV particles provide transduction of target tissues to achieve expression and translation of a payload or transgene (e.g., a base editor according to the present disclosure), for a duration sufficient to install the desired mutation in the genome of the target cell. In some embodiments, the desired mutation is an A to G mutation. In some embodiments, the desired mutation is a C to T mutation. In some embodiments, the disclosed rAAV particles provide sufficient expression and translation of the base editor transgene for a duration sufficient to install the desired (on-target) mutation in the genome, with a tolerable degree of off-target effects (e.g., bystander editing). In some embodiments, the disclosed rAAV particles provide sufficient expression and translation of the base editor transgene for a duration sufficient to install the desired mutation in the genome, without obvious off-target editing. In some embodiments, the disclosed rAAV particles provide sufficient expression and translation of the base editor transgene for a duration sufficient to install the desired mutation in the genome, without obvious bystander editing.

[0371] Suitable routes of administration of the disclosed compositions of rAAV particles include, but are not limited to, topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, intragingival, intradental, intracochlear, transtympanic, intraorgan, epidural, intracapsular, intramuscular, intravenous, systemic, intravascular, intraosseous, periocular, intratumoral, intracerebral, parenteral, and intraventricular administration. In some embodiments, the route of administration is systemic (intravenous). In some embodiments, the pharmaceutical compositions described herein are administered topically to the site of disease.

[0372] In some aspects, provided herein are pharmaceutical compositions comprising any disclosed compositions and pharmaceutically acceptable carriers. In some embodiments, the pharmaceutical composition is formulated for delivery to a subject, for example, for base editing in the genome. In some embodiments, the disclosed composition is formulated as a composition suitable for intravenous or subcutaneous administration to a subject (e.g., human) according to conventional procedures. In some embodiments, provided herein are compositions formulated for delivery to a subject, for example, to a human subject, so as to achieve targeted genome modification within the subject. In some embodiments, cells are obtained from a subject and contacted with any pharmaceutical composition provided herein. In some embodiments, optionally after the desired genome modification has been achieved or detected in the cell, cells removed from the subject and in vitro contact with the pharmaceutical composition are reintroduced into the subject. Subjects contemplated for administration of the pharmaceutical composition include, but are not limited to, humans and / or other primates; mammals, domesticated animals, pets, and commercially relevant mammals, such as cattle, pigs, horses, sheep, cats, dogs, mice, and / or rats; and / or birds, including commercially relevant birds, such as chickens, ducks, geese, and / or turkeys.

[0373] The formulations of the pharmaceutical compositions described herein can be prepared by any method known in pharmacology or developed later. Typically, such preparation methods comprise the steps of combining the active ingredient with an excipient and / or one or more other auxiliary ingredients, and then, if necessary and / or desired, shaping and / or packaging the product into desired single or multiple dose units. The pharmaceutical formulations may additionally contain pharmaceutically acceptable excipients, as used herein, including any and all solvents, dispersion media, diluents or other liquid vehicles, dispersing or suspending aids, surfactants, isotonic agents, thickeners or emulsifiers, preservatives, solid binders, lubricants, etc., to suit the desired particular dosage form.

[0374] In some embodiments, the pharmaceutical composition of rAAV particles for administration by injection is a solution in a sterile isotonic aqueous buffer. If necessary, the drug may also include a solubilizer and a local anesthetic, such as lidocaine, to relieve pain at the injection site. Typically, the ingredients can be provided separately or mixed together in unit dosage form, for example, as a dry lyophilized powder or anhydrous concentrate in a sealed container such as an ampoule or sachet indicating the amount of active agent. When the drug is administered by infusion, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. In the case of administration of the pharmaceutical composition by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed before administration.

[0375] As used herein, the term "pharmaceutically acceptable carrier" refers to a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., a lubricant, magnesium talc, calcium or zinc stearate, or stearic acid), or solvent encapsulating material, involved in carrying or transporting a compound from one part of the body (e.g., a delivery site) to another part (e.g., an organ, tissue, or part of the body). A pharmaceutically acceptable carrier is "acceptable" in the sense that it is compatible with the other ingredients of the formulation and does not damage the tissues of the subject (e.g., physiological compatibility, sterility, physiological pH, etc.). Some examples of materials that can be used as pharmaceutically acceptable carriers include: (1) sugars such as lactose, glucose, and sucrose; (2) starches such as corn starch and potato starch; (3) cellulose and its derivatives such as sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants such as magnesium stearate, sodium lauryl sulfate, and talc; (8) excipients such as cocoa butter and suppository waxes; (9) oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (10) glycols such as propylene glycol; (11) polyols Alcohols, such as glycerol, sorbitol, mannitol, and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffers, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer's solution; (19) ethanol; (20) pH buffered solutions; (21) polyesters, polycarbonates, and / or polyanhydrides; (22) fillers, such as polypeptides and amino acids; (23) serum components, such as serum albumin, HDL, and LDL; (22) C2-C12 alcohols, such as ethanol; and (23) other nontoxic compatible substances used in pharmaceutical formulations. Wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, fragrances, preservatives, and antioxidants may also be present in the formulation. Terms such as "excipient," "carrier," "pharmaceutically acceptable carrier," and the like are used interchangeably herein.

[0376] The rAAV particle pharmaceutical compositions of the present disclosure can be administered or packaged, for example, as unit doses. When used with respect to the pharmaceutical compositions of the present disclosure, the term "unit dose" refers to physically discrete units suitable as unitary dosages for subjects, each unit containing a predetermined quantity of active material calculated to produce the desired therapeutic effect in association with the required diluent (i.e., carrier or vehicle).

[0377] rAAV vector sequences

[0378] In some aspects, the present disclosure provides the following rAAV vector nucleic acid sequences. In some embodiments, the disclosed vectors comprise a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 98%, at least 99% or at least 99.5% sequence identity to any one of SEQ ID NOs: 100-102. In certain embodiments, the disclosed vectors comprise a nucleic acid sequence of any one of SEQ ID NOs: 100-102. The components of each sequence and the length of ITR-to-ITR are indicated.

[0379] In some embodiments, any of the vectors described herein may comprise a nucleic acid sequence that differs from any of SEQ ID NOs: 100-102 by 1-5, 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, or more than 50 nucleotides. These differences may comprise nucleotides that have been inserted, deleted, or substituted relative to any of SEQ ID NOs: 100-102. In some embodiments, the disclosed vectors contain a fragment that shares about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 300, about 400, about 500, or more than 500 consecutive nucleotides with any of SEQ ID NOs: 100-102.

[0380] Single AAV SaABE8e

[0381] ITR-EFS promoter - SaABE8e (start codon-BPNLS-TadA-SaCas9 D10A-BPNLS-stop codon Code) - bH Poly A- sgRNA (protospacer, in bold) - U6 -ITR( Restriction sites for cloning ) [4804 bp is the ITR-to-ITR length]

[0382] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT GCGGCCTCTAGAATTCGCTAGCTAGGTCTTGAAAGGAGTGGGAATTGGCTCCGGTGCCCGTCAGTGGGCAGAGCGCACATCGCCCACAGTCCCCGAGAAGTTGGGGGGAGGGGTCGGCAATTGATCCGGTGCCTAGAGAAGGTGGCGCGGGGTAAACTGGGAAAGTGATGTCGTGTACTGGCTCCGCCTTTTTCCCGAGGGTGGGGGAGAACCGTATATAAGTGCAGTAGTCGCCGTGAACGTTCTTTTTCGCAACGGGTTTGCCGCCAGAACACAGGACCGGTGCCACC ATGAAACGGACAGCCGACGGAAGCGAGT TCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGAC CCTGGCCAAGAGGGCACGGGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGC GAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCC TGGTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGC CATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATG AACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCC TGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGAGCTCCATCAACTCCGG AGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGAGAGCGCAACACCTGAAAGCAGCGGGGGC AGCAGCGGGGGGTCAGGGAAGCGAAATTACATTCTGGGGCTGGCCATTGGCATTACATCAGTGGGCTATGGCATCA TTGACTACGAGACAAGGGACGTGATCGACGCCGGCGTGAGACTGTTCAAGGAGGCCAACGTGGAGAACAATGAGGG CCGGAGATCCAAGAGGGGAGCAAGGCGCCTGAAGCGGAGAAGGCGCCACAGAATCCAGAGAGTGAAGAAGCTGCTG TTCGATTACAACCTGCTGACCGACCACTCCGAGCTGTCTGGCATCAATCCTTATGAGGCCAGAGTGAAGGGCCTGT CCCAGAAGCTGTCTGAGGAGGAGTTTAGCGCCGCCCTGCTGCACCTGGCAAAGAGGAGAGGCGTGCACAACGTGAA TGAGGTGGAGGAGGACACCGGCAACGAGCTGTCCACAAAGGAGCAGATCAGCCGCAATTCCAAGGCCCTGGAGGAG AAGTATGTGGCCGAGCTGCAGCTGGAGCGGCTGAAGAAGGATGGCGAGGTGAGGGGCTCCATCAATCGCTTCAAGA CCTCTGACTACGTGAAGGAGGCCAAGCAGCTGCTGAAGGTGCAGAAGGCCTACCACCAGCTGGATCAGTCCTTTAT CGATACATATATCGACCTGCTGGAGACAAGGCGCACATACTATGAGGGACCAGGAGAGGGCTCTCCCTTCGGCTGG AAGGACATCAAGGAGTGGTACGAGATGCTGATGGGCCACTGCACCTATTTTCCAGAGGAGCTGAGAAGCGTGAAGT ACGCCTATAACGCCGATCTGTACAACGCCCTGAATGACCTGAACAACCTGGTCATCACCAGGGATGAGAACGAGAA GCTGGAGTACTATGAGAAGTTCCAGATCATCGAGAACGTGTTCAAGCAGAAGAAGAAGCCTACACTGAAGCAGATC GCCAAGGAGATCCTGGTGAACGAGGAGGACATCAAGGGCTACCGCGTGACCTCCACAGGCAAGCCAGAGTTCACCA ATCTGAAGGTGTATCACGATATCAAGGACATCACAGCCCGGAAGGAGATCATCGAGAACGCCGAGCTGCTGGATCA GATCGCCAAGATCCTGACCATCTATCAGAGCTCCGAGGACATCCAGGAGGAGCTGACCAACCTGAATAGCGAGCTG ACACAGGAGGAGATCGAGCAGATCAGCAATCTGAAGGGCTACACCGGCACACACAACCTGAGCCTGAAGGCCATCA ATCTGATCCTGGATGAGCTGTGGCACACAAACGACAATCAGATCGCCATCTTTAACCGGCTGAAGCTGGTGCCAAA GAAGGTGGACCTGTCCCAGCAGAAGGAGATCCCAACCACACTGGTGGACGATTTCATCCTGTCTCCCGTGGTGAAG CGGAGCTTCATCCAGAGCATCAAAGTGATCAACGCCATCATCAAGAAGTACGGCCTGCCCAATGATATCATCATCG AGCTGGCCAGGGAGAAGAACTCCAAGGACGCCCAGAAGATGATCAATGAGATGCAGAAGAGGAACCGCCAGACCAA TGAGCGGATCGAGGAGATCATCAGAACCACAGGCAAGGAGAACGCCAAGTACCTGATCGAGAAGATCAAGCTGCAC GATATGCAGGAGGGCAAGTGTCTGTATTCTCTGGAGGCCATCCCTCTGGAGGACCTGCTGAACAATCCATTCAACT ACGAGGTGGATCACATCATCCCCCGGAGCGTGAGCTTCGACAATTCTTTTAACAATAAGGTGCTGGTGAAGCAGGA GGAGAACAGCAAGAAGGGCAATAGGACCCCTTTCCAGTACCTGTCTAGCTCCGATTCTAAGATCAGCTACGAGACA TTCAAGAAGCACATCCTGAATCTGGCCAAGGGCAAGGGCCGCATCAGCAAGACCAAGAAGGAGTACCTGCTGGAGG AGCGGGACATCAACAGATTCTCCGTGCAGAAGGACTTCATCAACCGGAATCTGGTGGACACCAGATACGCCACACG CGGCCTGATGAATCTGCTGCGGTCTTATTTCAGAGTGAACAATCTGGATGTGAAGGTGAAGAGCATCAACGGCGGC TTCACCTCCTTTCTGCGGAGAAAGTGGAAGTTTAAGAAGGAGCGCAACAAGGGCTATAAGCACCACGCCGAGGATG CCCTGATCATCGCCAATGCCGACTTCATCTTTAAGGAGTGGAAGAAGCTGGACAAGGCCAAGAAAGTGATGGAGAA CCAGATGTTCGAGGAGAAGCAGGCCGAGAGCATGCCCGAGATCGAGACAGAGCAGGAGTACAAGGAGATTTTCATC ACACCTCACCAGATCAAGCACATCAAGGACTTCAAGGACTACAAGTATTCTCACAGGGTGGATAAGAAGCCCAACC GCGAGCTGATCAATGACACCCTGTATAGCACACGGAAGGACGATAAGGGCAATACCCTGATCGTGAACAATCTGAA CGGCCTGTACGACAAGGATAATGACAAGCTGAAGAAGCTGATCAACAAGTCTCCCGAGAAGCTGCTGATGTACCAC CACGATCCTCAGACATATCAGAAGCTGAAGCTGATCATGGAGCAGTACGGCGACGAGAAGAACCCACTGTATAAGT ACTATGAGGAGACAGGCAACTACCTGACAAAGTATAGCAAGAAGGATAATGGCCCCGTGATCAAGAAGATCAAGTA CTATGGCAACAAGCTGAATGCCCACCTGGACATCACCGACGATTACCCTAACTCTCGCAATAAGGTGGTGAAGCTG AGCCTGAAGCCATACCGGTTCGACGTGTACCTGGACAACGGCGTGTATAAGTTTGTGACAGTGAAGAATCTGGATG TGATCAAGAAGGAGAACTACTATGAGGTGAACAGCAAGTGCTACGAGGAGGCCAAGAAGCTGAAGAAGATCAGCAA CCAGGCCGAGTTCATCGCCTCTTTTTACAACAATGACCTGATCAAGATCAATGGCGAGCTGTATAGAGTGATCGGC GTGAACAATGATCTGCTGAACAGAATCGAAGTGAATATGATCGACATCACCTACAGGGAGTATCTGGAGAACATGA ATGATAAGAGGCCCCCTCGCATCATCAAGACCATCGCCTCTAAGACACAGAGCATCAAGAAGTACAGCACAGACAT CCTGGGGAACCTGTATGAAGTCAAGAGCAAGAAACATCCTCAGATTATCAAGAAAGCCTCTGGCGGCTCAAAAAGA ACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAATAGATCTCGACTGTGCCTTCTAGTTGCC AGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATA AAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAG GGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTCGAGCGGCCCAAGCTTAAA AAAATCTCGCCAACAAGTTGACGAGATAAACACGGCATTTTGCCTTGTTTTAGTAGATTCTGTAATTTTCATTACA GAGTACTAAAACCGCCGTTGCTCCAAGGTATGGCGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAATC GAAATACTTTCAAGTTACGGTAAGCATATGATAGTCCATTTAAAACATAATTTTAAAACTGCAAACTACCCAAGA AATTATTACTTTCTACGTCACGTATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAATTCTAATTATCTCTC TAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCTTCCTGCCCGACCTTGCGGCC GC CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT(SEQ ID NO: 100)

[0383] Single AAV SaKKH-ABE8e

[0384] ITR-EFS promoter - SaKKHABE8e (start codon-BPNLS-TadA-SaKKHCas9 D10A-BPNLS- stop codon) - bGH poly A - sgRNA (protospacer, in bold) - U6- ITR( The gray sequence contains Restriction sites for cloning )[4804 bp is the ITR-to-ITR length]

[0385] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT GCGGCCTCTA GAATTCGCTAGCTAGGTCTTGAAAGGAGTGGGAATTGGCTCCGGTGCCCGTCAGTGGGCAGAGCGCACATCGCCCACAGTCCCCGAGAAGTTGGGGGGAGGGGTCGGCAATTGATCCGGTGCCTAGAGAAGGTGGCGCGGGGTAAACTGGGAAAGTGATGTCGTGTACTGGCTCCGCCTTTTTCCCGAGGGTGGGGGAGAACCGTATATAAGTGCAGTAGTCGCCGTGAACGTTCTTTTTCGCAACGGGTTTGCCGCCAGAACACAGGACCGGTGCCACC ATGAAACGGACAGCCGACGGAAGCGAGT TCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGAC CCTGGCCAAGAGGGCACGGGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGC GAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCC TGGTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGC CATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATG AACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCC TGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGAGCTCCATCAACTCCGG AGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGAGAGCGCAACACCTGAAAGCAGCGGGGGC AGCAGCGGGGGGTCAGGGAAGCGAAATTACATTCTGGGGCTGGCCATTGGCATTACATCAGTGGGCTATGGCATCA TTGACTACGAGACAAGGGACGTGATCGACGCCGGCGTGAGACTGTTCAAGGAGGCCAACGTGGAGAACAATGAGGG CCGGAGATCCAAGAGGGGAGCAAGGCGCCTGAAGCGGAGAAGGCGCCACAGAATCCAGAGAGTGAAGAAGCTGCTG TTCGATTACAACCTGCTGACCGACCACTCCGAGCTGTCTGGCATCAATCCTTATGAGGCCAGAGTGAAGGGCCTGT CCCAGAAGCTGTCTGAGGAGGAGTTTAGCGCCGCCCTGCTGCACCTGGCAAAGAGGAGAGGCGTGCACAACGTGAA TGAGGTGGAGGAGGACACCGGCAACGAGCTGTCCACAAAGGAGCAGATCAGCCGCAATTCCAAGGCCCTGGAGGAG AAGTATGTGGCCGAGCTGCAGCTGGAGCGGCTGAAGAAGGATGGCGAGGTGAGGGGCTCCATCAATCGCTTCAAGA CCTCTGACTACGTGAAGGAGGCCAAGCAGCTGCTGAAGGTGCAGAAGGCCTACCACCAGCTGGATCAGTCCTTTAT CGATACATATATCGACCTGCTGGAGACAAGGCGCACATACTATGAGGGACCAGGAGAGGGCTCTCCCTTCGGCTGG AAGGACATCAAGGAGTGGTACGAGATGCTGATGGGCCACTGCACCTATTTTCCAGAGGAGCTGAGAAGCGTGAAGT ACGCCTATAACGCCGATCTGTACAACGCCCTGAATGACCTGAACAACCTGGTCATCACCAGGGATGAGAACGAGAA GCTGGAGTACTATGAGAAGTTCCAGATCATCGAGAACGTGTTCAAGCAGAAGAAGAAGCCTACACTGAAGCAGATC GCCAAGGAGATCCTGGTGAACGAGGAGGACATCAAGGGCTACCGCGTGACCTCCACAGGCAAGCCAGAGTTCACCA ATCTGAAGGTGTATCACGATATCAAGGACATCACAGCCCGGAAGGAGATCATCGAGAACGCCGAGCTGCTGGATCA GATCGCCAAGATCCTGACCATCTATCAGAGCTCCGAGGACATCCAGGAGGAGCTGACCAACCTGAATAGCGAGCTG ACACAGGAGGAGATCGAGCAGATCAGCAATCTGAAGGGCTACACCGGCACACACAACCTGAGCCTGAAGGCCATCA ATCTGATCCTGGATGAGCTGTGGCACACAAACGACAATCAGATCGCCATCTTTAACCGGCTGAAGCTGGTGCCAAA GAAGGTGGACCTGTCCCAGCAGAAGGAGATCCCAACCACACTGGTGGACGATTTCATCCTGTCTCCCGTGGTGAAG CGGAGCTTCATCCAGAGCATCAAAGTGATCAACGCCATCATCAAGAAGTACGGCCTGCCCAATGATATCATCATCG AGCTGGCCAGGGAGAAGAACTCCAAGGACGCCCAGAAGATGATCAATGAGATGCAGAAGAGGAACCGCCAGACCAA TGAGCGGATCGAGGAGATCATCAGAACCACAGGCAAGGAGAACGCCAAGTACCTGATCGAGAAGATCAAGCTGCAC GATATGCAGGAGGGCAAGTGTCTGTATTCTCTGGAGGCCATCCCTCTGGAGGACCTGCTGAACAATCCATTCAACT ACGAGGTGGATCACATCATCCCCCGGAGCGTGAGCTTCGACAATTCTTTTAACAATAAGGTGCTGGTGAAGCAGGA GGAGAACAGCAAGAAGGGCAATAGGACCCCTTTCCAGTACCTGTCTAGCTCCGATTCTAAGATCAGCTACGAGACA TTCAAGAAGCACATCCTGAATCTGGCCAAGGGCAAGGGCCGCATCAGCAAGACCAAGAAGGAGTACCTGCTGGAGG AGCGGGACATCAACAGATTCTCCGTGCAGAAGGACTTCATCAACCGGAATCTGGTGGACACCAGATACGCCACACG CGGCCTGATGAATCTGCTGCGGTCTTATTTCAGAGTGAACAATCTGGATGTGAAGGTGAAGAGCATCAACGGCGGC TTCACCTCCTTTCTGCGGAGAAAGTGGAAGTTTAAGAAGGAGCGCAACAAGGGCTATAAGCACCACGCCGAGGATG CCCTGATCATCGCCAATGCCGACTTCATCTTTAAGGAGTGGAAGAAGCTGGACAAGGCCAAGAAAGTGATGGAGAA CCAGATGTTCGAGGAGAAGCAGGCCGAGAGCATGCCCGAGATCGAGACAGAGCAGGAGTACAAGGAGATTTTCATC ACACCTCACCAGATCAAGCACATCAAGGACTTCAAGGACTACAAGTATTCTCACAGGGTGGATAAGAAGCCCAACC GCAAGCTGATCAATGACACCCTGTATAGCACACGGAAGGACGATAAGGGCAATACCCTGATCGTGAACAATCTGAA CGGCCTGTACGACAAGGATAATGACAAGCTGAAGAAGCTGATCAACAAGTCTCCCGAGAAGCTGCTGATGTACCAC CACGATCCTCAGACATATCAGAAGCTGAAGCTGATCATGGAGCAGTACGGCGACGAGAAGAACCCACTGTATAAGT ACTATGAGGAGACAGGCAACTACCTGACAAAGTATAGCAAGAAGGATAATGGCCCCGTGATCAAGAAGATCAAGTA CTATGGCAACAAGCTGAATGCCCACCTGGACATCACCGACGATTACCCTAACTCTCGCAATAAGGTGGTGAAGCTG AGCCTGAAGCCATACCGGTTCGACGTGTACCTGGACAACGGCGTGTATAAGTTTGTGACAGTGAAGAATCTGGATG TGATCAAGAAGGAGAACTACTATGAGGTGAACAGCAAGTGCTACGAGGAGGCCAAGAAGCTGAAGAAGATCAGCAA CCAGGCCGAGTTCATCGCCTCTTTTTACAAGAATGACCTGATCAAGATCAATGGCGAGCTGTATAGAGTGATCGGC GTGAACAATGATCTGCTGAACAGAATCGAAGTGAATATGATCGACATCACCTACAGGGAGTATCTGGAGAACATGA ATGATAAGAGGCCCCCTCATATCATCAAGACCATCGCCTCTAAGACACAGAGCATCAAGAAGTACAGCACAGACAT CCTGGGGAACCTGTATGAAGTCAAGAGCAAGAAACATCCTCAGATTATCAAGAAAGGCTCTGGCGGCTCAAAAAGA ACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAATAGATCTCGACTGTGCCTTCTAGTTGCC AGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATA AAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAG GGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTCGAGCGGCCCAAGCTTAAA AAAATCTCGCCAACAAGTTGACGAGATAAACACGGCATTTTGCCTTGTTTTAGTAGATTCTGTAATTTTCATTACA GAGTACTAAAACCGCCGTTGCTCCAAGGTATGGCGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAATC GAAATACTTTCAAGTTACGGTAAGCATATGATAGTCCATTTAAAACATAATTTTAAAACTGCAAACTACCCAAGA AATTATTACTTTCTACGTCACGTATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAATTCTAATTATCTCTC TAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCTTCCTGCCCGACCTTGCGGCC GC CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT(SEQ ID NO: 101)

[0386] Single AAV SauriABE8e

[0387] ITR-EFS Promoter-SauriABE8e (start codon-BPNLS-TadA-SauriCas9 D10A-BPNLS- stop codon) - bGH poly A-sgRNA (protospacer, in bold) - U6 -ITR (The gray sequence contains Restriction sites for cloning) [4828 bp is the ITR-to-ITR length]

[0388] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCC GGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCTGCGGCCTCTAGAAT TCGCTAGCTAGGTCTTGAAAGGAGTGGGAATTGGCTCCGGTGCCCGTCAGTGGGCAGAGCGCACATCGCCCACAGT CCCCGAGAAGTTGGGGGGGAGGGGTCGGCAATTGATCCGGTGCCTAGAGAAGGTGGCGGGGTAAACTGGGAAAGT GATGTCGTGTACTGGCTCCGCCTTTTTCCCGAGGGTGGGGGAGAACCGTATATAAGTGCAGTAGTCGCCGTGAACG TTCTTTTTCGCAACGGGTTTGCCGCCAGAACACAGGACCGGTGCCACCATGAAACGGACAGCCGACGGAAGCGAGT TCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGAC CCTGGCCAAGAGGGCACGGGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGC GAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCC TGGTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGC CATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATG AACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCC TGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGAGCTCCATCAACTCCGG AGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGAGAGCGCAACACCTGAAAGCAGCGGGGGC AGCAGCGGGGGGTCAATGCAGGAGAACCAGCAGAAGCAGAACTACATCCTGGGCCTGGCCATCGGAATCACCAGCG TCGGCTACGGACTGATCGATAGCAAGACAAGAGAAGTGATCGACGCCGGCGTTAGACTCTTTCCAGAAGCTGATAG CGAGAACAACTCCAACCGCAGAAGCAAGCGGGGCGCCAGACGGTTAAAACGGAGAAGAATCCACCGGCTGAACCGG GTCAAAGACCTGCTCGCTGATTACCAGATGATCGATCTTAACAATGTTCCTAAGAGCACCGACCCCTACACCATCA GAGTGAAGGGCCTCCGGGAGCCTCTGACAAAGAAGAATTCGCCATCGCCCTCCTGCATATCGCTAAGAGAAGAGG CCTGCACAACATCAGTGTGTCCATGGGCGACGAAGAGCAGGACAATGAACTGAGCACCAAGCAGCAGCTGCAAAAG AATGCCCAGCAACTGCAGGACAAGTATGTGTGCGAACTGCAGTTAGAACGGCTGACCAACATCAACAAGGTCAGAG GCGAGAAGAAACAGATTTAAGACAGAGGACTTTGTGAAAGAAGTGAAACAGCTGTGCGAAACCCAGAGACAGTACCA CAACATCGACGACCAATTCATCCAGCAGTACATCGACCTGGTGTCTACAAGACGGGAGTACTTCGAGGGCCCGGC AACGGCTCTCCATACGGCTGGGACGGCGACCTGCTGAAGTGGTACGAGAAGCTGATGGGCAGATGCACCTATTTC CCGAAGAACTGAGGTCCGTGAAGTACGCCTACAGCGCCGACCTCTTCAACGCCCTGAACGACCTGAACAACCTCGT TGTGACCAGGGATGACAATCCAAAGCTTGAGTACTACGAGAAGTACCACATTATTGAGAACGTGTTCAAGCAAAAG AAGAATCCCACACTCAAACAAATCGCCAAAGAGATCGGCGTGCCAAGATTACGACATCCGGGGCTATAGAATCACAA AGAGCGGCAAACCTCAGTTCACCTCTTTTAAGCTGTATCACGACCTGAAGAACATCTTCGAGCAGGCCAAATACCT GGAAGATGTGGAAATGCTGGACGAGATCGCCAAGATCCTGACCATCTACCAGGATGAGATTAGCATCAAGAAAGCC CTGGACCAGCTGCCCGAACTGCTGACAGAGAGCGAGAAATCTCAGATCGCACAGCTCACCCGGCTATACAGGCACCC ACAGACTGAGCCTGAAGTGCATCCACATTGTGATCGACGAGCTGTGGGAGAGCCCGAGAACCAGATGGAAATCTT TACCAGACTGAATCTGAAACCTAAGAAGGTGGAAATGAGCGAGATCGACAGCATACCCACCACCCTGGTCGACGAG TTCATCCTCTCACCTGTGGTGAAGCGGGCCTTCATCCAGAGCATCAAGGTAATCAACGCAGTGATCAATCGGTTCG GCCTGCCAGAGGACATCATCATCGAGCTGGCCAGAGAAAAGAATAGCAAGGATCGGAGAAAGTTCATTAACAAGCT GCAGAAACAAAATGAGGCCACAAGAAAGAAAATCGAACAGCTGCTGGCCAAGTACGGCAACACCAATGCCAAGTAC ATGATCGAGAAGATCAAGCTGCACGACATGCAGGAGGGCAAGTGCCTGTACAGCCTGGAGGCTATTCCTCTGGAAG ACCTGCTGAGCAACCCGACACACTACGAAGTTGACCACATTATCCCCAGATCTGTGAGCTTTGACAACAGCCTGAA CAACAAAGTGCTGGTGAAACAAAGCGAAAACAGCAAGAAGGGCAATCGCACCCCTTACCAGTACCTGAGCAGCAAC GAGTCTAAGATTAGCTACAACCAGTTTAAGCAGCACATCCTGAACCTGAGCAAGGCCAAGGACAGAATCAGCAAGA AAAAAAGAGATATGCTGCTGGAAGAGAGAGATATCAACAAGTTCGAAGTGCAGAAGGAATTCATTAACCGGAACCT GGTGGATACACGGTACGCCACCAGAGAACTGTCTAACCTGCTGAAGACCTACTTCAGCACCCATGACTACGCCGTG AAGGTGAAGACCATCAACGGCGGCTTCACTAACCACCTGAGGAAGGTGTGGGATTTCAAGAAGCACAGAAACCACG GCTACAAGCACCACGCCGAAGATGCCCTGGTGATCGCCAACGCCGACTTCCTGTTTAAGACACATAAGGCCCTGCG GAGAACCGATAAGATCCTGGAACAACCTGGCCTGGAAGTGAATGATACAACCGTGAAAGTGGACACCGAGGAAAAA TACCAGGAGCTGTTCGAGACACCTAAGCAAGTGAAGAACATCAAGCAGTTCCGGGACTTCAAGTACAGCCACCGAG TGGACAAGAAGCCTAACCGGCAGCTTATCAACGACACACTGTACTCCACCAGAGAGATTGATGGCGAAACCTACGT GGTGCAGACCCTTAAGGATCTGTACGCCAAGGACAACGAGAAAGTGAAGAAGCTGTTCACCGAAAGACCTCAGAAG ATCCTGATGTACCAGCACGACCCTAAGACCTTCGAGAAACTGATGACAATCCTGAACCAGTACGCTGAGGCCAAGA ACCCTCTGGCTGCTTATTACGAGGACAAAGGCGAGTACGTGACCAAGTACGCCAAGAAAGGCAATGGACCTGCCAT CCACAAGATCAAGTATATCGATAAGAAGCTTGGATCTTACCTGGATGTTAGCAACAAGTATCCTGAGACACAGAAC AAGCTTGTGAAGCTGTCCCTGAAGAGCTTTAGATTCGACATCTACAAGTGTGAACAGGGCTACAAGATGGTGTCCA TCGGATACCTGGACGTGCTGAAGAAAGATAACTACTACTACATCCCTAAGGACAAGTACGAGGCCGAGAAGCAGAA AAAGAAGATCAAGGAATCTGATCTTTTTGTGGGCAGCTTCTACTACAACGACCTCATCATGTACGAGGATGAACTG TTCAGAGTGATAGGAGTGAACAGCGACATCAACAATCTGGTTGAGCTAAACATGGTCGACATTACCTACAAGGACT TCTGCGAGGTGAACAACGTGACAGGCGAGAAAAGAATCAAAAAGACTATCGGCAAGCGCGTGGTCCTGATCGAGAA GTACACCACAGATATTCTAGGCAACCTGTACAAGACTCCCCTGCCTAAGAAGCCCCAGCTTATCTTCAAGCGGGGA GAACTGTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAATAGA TCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGC CACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGG GGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTA TGGCTCGAGCGGCCCAAGCTTAAAAAAATCTCGCCAACAAGTTGACGAGATAAACACGGCATTTTGCCTTGTTTTA GTAGATTCTGTAATTTTCATTACAGAGTACTAAAACCGTTGCTCCAAGGTATGGGTGCGGTGTTTCGTCCTTTCCA CAAGATATATAAAGCCAAGAAATCGAAATACTTTCAAGTTACGGTAAGCATATGATAGCCATTTTAAAACATAAT TTTAAAACTGCAAACTACCCAAGAAATTATTACTTTTCTACGTCACGTATTTTGTACTAATATCTTTGTGTTTTACAG TCAAATTAATTCTAATTATCTCTCTAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCATGGGAAATAGGCC CTCTTCCTGCCCGACCTTGCGGCCGCCTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT (SEQ ID NO: 102)

[0389] Methods for editing target nucleic acid molecules

[0390] In some aspects, provided herein is a method for contacting any disclosed AAV-encoded base editor with a nucleic acid molecule, e.g., a nucleic acid molecule (e.g., DNA) comprising a target sequence. In some embodiments, the nucleic acid molecule comprises DNA, e.g., single-stranded DNA or double-stranded DNA. The target sequence of the nucleic acid molecule may comprise a target nuclear base pair containing adenine (A). The target sequence of the nucleic acid molecule may comprise a target nuclear base pair containing cytosine (C). The target sequence may be a genomic sequence, e.g., a human genomic sequence. The target sequence may comprise a sequence associated with a disease or disorder, e.g., a target sequence with a point mutation. Target sequences with point mutations may be associated with cardiovascular disease.

[0391] In some embodiments, the target nucleotide sequence is a DNA sequence in a genome, for example, a eukaryotic genome. In certain embodiments, the target nucleotide sequence is in a mammalian (e.g., human) genome. In certain embodiments, the target nucleotide sequence is in the human genome. In other embodiments, the target nucleotide sequence is in the genome of rodents such as mice or rats. In other embodiments, the target nucleotide sequence is in the genome of domestic animals such as horses, cats, dogs or rabbits. In some embodiments, the target nucleotide sequence is in the genome of research animals. In some embodiments, the target nucleotide sequence is in the genome of a genetically engineered non-human subject. In some embodiments, the target nucleotide sequence is in a plant genome. In some embodiments, the target nucleotide sequence is in the genome of a microorganism such as a bacterium.

[0392] In some embodiments, the disclosed AAV-encoded base editors exhibit low off-target effects, such as low off-target editing frequencies. In some embodiments, the disclosed AAV-encoded base editors exhibit low off-target editing frequencies while exhibiting high on-target editing efficiency. In some embodiments, the use of TadA-8e deaminase or TadA-8e (V106W) deaminase in any disclosed AAV-encoded adenine base editor may exhibit an off-target editing frequency of 0.32% or less while maintaining an on-target editing efficiency of about 80% or more. See PCT Publication No. WO 2021 / 158921, published on August 12, 2021.

[0393] For one or more base editors evaluated, the disclosed base editors can provide (or produce) an on-target editing efficiency of greater than 50% or greater than 60% (e.g., greater than 70%, greater than 75%, greater than 80%, or greater than 85%) at the target nucleobase pair. Any disclosed editing method can produce an on-target editing efficiency of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, or at least about 85%.

[0394] In some embodiments, the disclosed BEs and editing methods comprising the step of contacting a cell comprising a target DNA sequence with any disclosed BE result in an actual or average off-target DNA editing frequency of about 2.0% or less, 1.75% or less, 1.5% or less, 1.2% or less, 1% or less, 0.9% or less, 0.8% or less, 0.75% or less, 0.7% or less, 0.65% or less, or 0.6% or less. These off-target editing frequencies can be obtained in sequences having any level of sequence identity to the target sequence. As used herein with reference to off-target DNA editing frequencies, the modifier "average" refers to the average of all editing events detected at sites other than a given target nucleobase pair (e.g., as detected by high-throughput sequencing).

[0395] In various embodiments, the disclosed editing methods result in an on-target DNA base editing efficiency of at least about 35%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 98% or 99% at the target nucleobase pair. The contacting step may result in a DNA base editing efficiency of at least about 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74% or 75%. In particular, the contacting step results in an on-target base editing efficiency greater than 75%. The contacting step may result in a DNA base editing efficiency between 60% and 85%. In certain embodiments, a base editing efficiency greater than 85% may be achieved.

[0396] These editing efficiencies can be achieved following administration of any disclosed AAV-encoded base editor in any target tissue, e.g., liver tissue, cardiac tissue, muscle tissue, neuronal tissue, or ocular tissue.

[0397] After administration of any of the disclosed rAAV vectors or particles to cardiac tissue or muscle tissue (e.g., skeletal muscle tissue), an editing efficiency of at least about 20%, at least about 22%, at least about 24%, at least about 27%, at least about 30%, at least about 33%, or at least about 36% can be achieved. These editing efficiencies represent a 2 to 2.5-fold increase compared to the editing efficiencies reported for dual AAV vectors in cardiac and muscle tissue. After administration of any of the disclosed rAAV vectors or particles to cardiac tissue or muscle tissue (e.g., cardiac tissue), an indel rate of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less can be achieved in cardiac cells or muscle cells.

[0398] In various embodiments, the disclosed editing methods result in a ratio of on-target edits to off-target edits of about 25: 1, 50: 1, 65: 1, 75: 1, 80: 1, 85: 1, 90: 1, 95: 1, 100: 1, 110: 1, 125: 1, or greater than 125: 1. In various embodiments, the disclosed editing methods result in a ratio of on-target edits to off-target edits of about 150: 1, 200: 1, 300: 1, 400: 1, 500: 1, 600: 1, 700: 1, 800: 1, 900: 1, 1000: 1, 1100: 1, 1200: 1, 1250: 1, 1275: 1, 1300: 1, 1325: 1, 1350: 1, 1400: 1, 1500: 1, or greater than 1500: 1. As used herein, the ratio of on-target editing: off-target editing is equivalent to the ratio of sequencing reads reflecting on-target deamination relative to the deamination of a known or predicted off-target site or a candidate off-target site. Candidate off-target sites can be identified, so the ratio of on-target editing: off-target editing can be measured using an experimental assay or a computational algorithm (e.g., Cas-OFFinder). For example, candidate off-target sites can be identified using an experimental assay (such as EndoV-Seq, GUIDE-Seq, or CIRCLE-Seq). In some embodiments, the ratio of on-target editing: off-target editing depends on the use of EndoV-Seq.

[0399] In some embodiments, the disclosed editing methods result in and the disclosed base editors generate minimal bystander edits (i.e., synonymous off-target point mutations at nucleobases near the target base and that do not change the outcome of the intended editing method). In some embodiments, the disclosed editing methods result in less than 10, less than 9, less than 8, less than 7, less than 6, less than 5, less than 4, less than 3, less than 2, less than 1, or zero non-silent bystander edits. For example, an editing method using the disclosed single AAV-encoded SaKKH-ABE8e editor in liver tissue can result in few (i.e., minimal) non-silent bystander edits.

[0400] Some aspects of the present disclosure are based on the following recognition: Any adenine base editor provided herein is capable of modifying specific DNA bases without generating a significant proportion of indels. As used herein, "indel" refers to the insertion or deletion of a nucleotide base in a DNA substrate. Such insertions or deletions can lead to frameshift mutations in gene coding regions. In some embodiments, it is desirable to generate an adenine base editor that effectively modifies (e.g., mutates or deaminates) specific nucleotides in DNA without generating a large number of insertions or deletions (i.e., indels) in nucleic acids (while having a lower RNA editing effect than existing adenine base editors).

[0401] In some embodiments, the disclosed editing methods using the disclosed BEs can result in less than 20%, 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1.5%, 1%, 0.5%, 0.2%, or 0.1% indel formation in a nucleic acid (e.g., DNA) comprising a target sequence. In some embodiments, the disclosed editing methods result in an indel rate of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less. See Figure 11A 、 11B and 15.

[0402] In some embodiments, the disclosed editing methods result in a base edit:indel ratio of at least about 5:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, or greater than 15:1.

[0403] Some aspects of the present disclosure are based on the following understanding: Any base editor provided herein can effectively generate expected mutations in DNA (e.g., DNA in the subject's genome) without generating a significant number of unexpected mutations, such as unexpected point mutations. In some embodiments, the expected mutation is a mutation generated by a specific base editor bound to gRNA, which is specifically designed to generate expected mutations (e.g., deamination). In some embodiments, the expected mutation is a mutation associated with a disease or disorder such as sickle cell disease. In some embodiments, the expected mutation is a point mutation of adenine (A) to guanine (G) associated with a disease or disorder. In some embodiments, the expected mutation is a point mutation of thymine (T) to cytosine (C) associated with a disease or disorder. In some embodiments, the expected mutation is a point mutation of adenine (A) to guanine (G) in a gene coding region. In some embodiments, the expected mutation is a point mutation of thymine (T) to cytosine (C) in a gene coding region.

[0404] In some embodiments, the expected mutation is a deamination that generates a stop codon (e.g., a premature stop codon) in the coding region of a gene. In some embodiments, the expected mutation is a mutation that eliminates a stop codon. In some embodiments, the expected mutation eliminates a stop codon comprising the nucleic acid sequence 5'-TAG-3', 5'-TAA-3', or 5'-TGA-3'.

[0405] In some embodiments, the expected mutation is a deamination that changes the regulatory sequence of a gene (e.g., a gene promoter or a gene repressor). In some embodiments, the expected mutation is a deamination introduced into a gene promoter. In specific embodiments, the deamination introduced into a gene promoter causes a decrease in the transcription of a gene operably linked to the gene promoter. In other embodiments, the deamination causes an increase in the transcription of a gene operably linked to the gene promoter.

[0406] In some embodiments, the expected mutation is a deamination that changes the splicing of a genetic sequence or gene. Therefore, in some embodiments, the expected deamination results in the introduction of a splice site in a gene. In other embodiments, the expected deamination results in the removal of a splice site. In some embodiments, the expected deamination results in the introduction of a stop codon in a gene. In other embodiments, the expected deamination results in the removal of a stop codon.

[0407] In some embodiments, any of the base editors provided herein are capable of generating a ratio of intended mutations to unintended mutations (e.g., intended point mutations: unintended point mutations) greater than 1:1. In some embodiments, any base editor provided herein is capable of generating a ratio of intended mutations to unintended mutations (e.g., intended point mutations: unintended point mutations) of at least 1.5:1, at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 150:1, at least 200:1, at least 250:1, at least 500:1, or at least 1000:1 or greater. It should be understood that the features of the base editors described in this and following sections of the disclosure can be applied to any base editor or method of using a base editor provided herein.

[0408] The disclosed AAV-encoded base editors disclosed herein have reduced and / or low RNA editing effects. In some embodiments, the base editors are evolved or engineered to have reduced RNA editing effects. As used herein, the term "RNA editing effect" refers to the modification (e.g., deamination) of nucleotides introduced into cellular RNA (e.g., messenger RNA (mRNA)). An important goal of DNA base editing efficiency is to modify (e.g., deaminize) specific nucleotides in DNA without introducing modifications of similar nucleotides in RNA. When the detected mutation is introduced into an RNA molecule at a frequency of 0.3% or less, the RNA editing effect is "low" or "reduced."

[0409] The present disclosure further provides methods of administering the disclosed base editors, wherein the methods produce reduced and / or low RNA editing effects. The present disclosure further provides base editors (e.g., disclosed ABEs) that induce (or produce, provide, or cause) low and / or undetectable RNA editing effects (see Figure 17 ). In some embodiments, the base editor provides an average adenosine (A) to inosine (I) (A- to -I) editing frequency of 0.3% or less in the cellular mRNA transcript. In some embodiments, the base editor provides an actual and / or consistent editing frequency of an average adenosine (A) to inosine (I) (A- to -I) of about 0.3% or less in the RNA. The base editor can provide an actual or average A- to -I editing frequency of about 0.5% or less, 0.4% or less, 0.35% or less, 0.25% or less, 0.2% or less, 0.15% or less, 0.12% or less, 0.1% or less, 0.08% or less, or 0.075% or less in the RNA.

[0410] Guide sequence (e.g., guide RNA)

[0411] The present disclosure further provides guide RNAs for use according to the disclosed editing methods. The present disclosure provides guide RNAs designed to recognize target sequences. Such gRNAs can be designed to have guide sequences (or "spacers") that are complementary to the original spacers within the target sequence.

[0412] Also provided is a guide RNA for use with one or more disclosed base editors, for example, in the disclosed method for editing nucleic acid molecules.Such gRNA can be designed to have a guide sequence having complementarity with the original spacer within the target sequence to be edited, and having a backbone sequence that specifically interacts with the napDNAbp domain of any disclosed base editor (e.g., the Cas9 nickase domain of the disclosed base editor). The guide RNA according to the disclosed editing method can be complementary to any original spacer sequence (SEQ ID NO: 430-565) listed in Table 1.

[0413] In various embodiments, the base editor can be compounded, bound or otherwise associated with one or more guide sequences (e.g., via any type of covalent or non-covalent bond). The guide sequence becomes associated or bound to the base editor and guides its positioning to a specific target sequence that is complementary to the guide sequence or a portion thereof. The specific design embodiment of the guide sequence will depend on the nucleotide sequence of the genomic target sequence (i.e., the desired site to be edited) and the type of napDNAbp present in the base editor (e.g., the type of Cas9 protein), as well as other factors, such as PAM sequence position, percentage G / C content in the target sequence, degree of microhomology, secondary structure, etc.

[0414] Typically, a guide sequence is any polynucleotide sequence that has sufficient complementarity to hybridize with the target sequence and instruct the sequence-specific binding of a napDNA bp (e.g., Cas9 or Cas9 variant) to the target sequence. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the degree of complementarity between the guide sequence and its corresponding target sequence is about or greater than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or higher. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, an algorithm based on the Burrows-Wheeler transformation (e.g., a Burrows-Wheeler aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).

[0415] In some embodiments, the guide sequence is about or greater than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. In some embodiments, the guide RNA is 15-300, 25-300, 50-300, 25-250, 25-200, 15-200, or 15-100 nucleotides in length. In some embodiments, the guide RNA is about 25 to 200 nucleotides in length. In some embodiments, the guide RNA is about 15 to 200 or 15 to 100 nucleotides in length.

[0416] The ability of the guide sequence to guide the sequence-specific binding of the base editor to the target sequence can be evaluated by any suitable assay.For example, the components of the base editor (including the guide sequence to be tested) can be provided to a host cell with a corresponding target sequence, for example, by transfection with a vector encoding the components of the base editor disclosed herein, followed by evaluation of preferential cutting within the target sequence. Similarly, the target sequence, the components of the base editor (including the guide sequence to be tested and a control guide sequence different from the test guide sequence) can be provided, and the binding or cleavage rate at the target sequence between the test and control guide sequence reactions can be compared to evaluate the cutting of the target polynucleotide sequence in situ. Other assays are possible, and those skilled in the art will appreciate it.

[0417] The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of the cell. Exemplary target sequences include those that are unique in the target genome.

[0418] In some embodiments, the guide sequence is selected to reduce the degree of secondary structure within the guide sequence. Secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. An example of such an algorithm is mFold, as described by Zuker & Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example folding algorithm is RNAfold, an online web server developed at the Institute of Theoretical Chemistry, University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., AR Gruber et al., 2008, Cell 106 (1): 23-24; and PA Carr & GM Church, 2009, Nature Biotechnology 27 (12): 1151-62). Additional algorithms can be found in Chuai, G. et al., DeepCRISPR: optimized CRISPRguide RNA design by deep learning, Genome Biol. 19:80 (2018), and U.S. Application Serial No. 61 / 836,080 and U.S. Patent No. 8,871,445, issued October 28, 2014, the entire contents of each of which are incorporated herein by reference.

[0419] The guide sequence of the gRNA is linked to a tracr mate (also referred to as a "backbone") sequence, which in turn hybridizes to the tracr sequence. The tracr mate sequence includes any sequence that has sufficient complementarity to the tracr sequence to promote one or more of the following: (1) excision of the guide sequence flanked by the tracr mate sequence in a cell containing the corresponding tracr sequence; and (2) formation of a complex at the target sequence, wherein the complex comprises the tracr mate sequence hybridized to the tracr sequence. Generally, the degree of complementarity refers to the optimal alignment of the tracr mate sequence and the tracr sequence along the length of the shorter of the two sequences. The optimal alignment can be determined by any suitable alignment algorithm and can further take into account secondary structure, such as self-complementarity within the tracr sequence or the tracr mate sequence. In some embodiments, when optimally aligned, the degree of complementarity between the tracr sequence and the tracr mate sequence along the length of the shorter of the two sequences is about or greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more. In some embodiments, the length of the tracr sequence is about or greater than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more nucleotides. In some embodiments, the tracr sequence and the tracr mate sequence are contained within a single transcript such that hybridization between the two produces a transcript having a secondary structure (e.g., a hairpin). The preferred loop-forming sequence for the hairpin structure is four nucleotides in length and most preferably has the sequence GAAA. However, longer or shorter loop sequences can be used, and alternative sequences can also be used. The sequence preferably includes a nucleotide triplet (e.g., AAA) and an additional nucleotide (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAG. In embodiments of the present invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins. In certain embodiments, the transcript has two, three, four, or five hairpins. In further embodiments of the present invention, the transcript has a maximum of five hairpins. In some embodiments, the individual transcripts further comprise a transcription termination sequence; preferably, this is a poly-T sequence, eg, six T nucleotides.

[0420] In some embodiments, the guide RNA used according to the disclosed editing method comprises a synthetic single guide RNA (sgRNA) containing modified ribonucleotides. In some embodiments, the guide RNA contains modifications, such as 2'-O-methylated nucleotides and phosphorothioate bonds. In some embodiments, the guide RNA contains 2'-O-methyl modifications in the first three and last three nucleotides and contains a phosphorothioate bond between the first three and last three nucleotides. Exemplary modified synthetic sgRNAs are disclosed in Hendel A. et al., Nat. Biotechnol. 33, 985-989 (2015), which is incorporated herein by reference.

[0421] In some embodiments, the guide RNA used according to the disclosed editing method includes a backbone structure recognized by Neisseria meningitidis Cas9 protein or domain (e.g., Nme2Cas9 domain). The backbone structure (or scaffold) recognized by the Nme2Cas9 protein can include the following provided sequence: 5'-[guide sequence]-gttgtagctccctttctcatttcggaaacgaaatgagaaccgttgctacaataaggccgtctgaaaagatgtgccgcaacgctctgccccttaaagcttctgctttaaggggcatcgttta-3' (SEQ ID NO: 719). The scaffold sequence is recognized by NmeCas9, Nme1Cas9, Nme2Cas9, and Nme3Cas9 proteins. Exemplary guide RNAs for editing with base editors containing Nme2Cas9 domains and variants thereof are described in Edraki et al., Molecular Cell 73, 714-726, which are incorporated herein by reference.

[0422] In other embodiments, the guide RNA used according to the disclosed editing method comprises a backbone structure recognized by a Streptococcus pyogenes Cas9 protein or domain (e.g., the SpCas9 domain of the disclosed base editor). The backbone structure recognized by the SpCas9 protein may comprise the following sequence: 5'-[guide sequence]-guuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuu-3' (SEQ ID NO: 339), wherein the guide sequence comprises a sequence complementary to the original spacer of the target sequence. See, U.S. Publication No. 2015 / 0166981, published on June 18, 2015, the disclosure of which is incorporated herein by reference. The guide sequence is typically 20 nucleotides long.

[0423] In other embodiments, the guide RNA used according to the disclosed editing method comprises a backbone structure recognized by a Campylobacter jejuni Cas9 protein or domain (e.g., a CjCas9 domain of a disclosed base editor). The backbone structure recognized by the CjCas9 protein can comprise the following sequence: 5'-[guide sequence]-gttttagtccctgaaaagggactaaaataaagagtttgcgggactctgcggggttacaatcccctaaaaccgcttttttt-3' (SEQ ID NO: 340), wherein the guide sequence comprises a sequence complementary to the protospacer of the target sequence.

[0424] In other embodiments, the guide RNA used in accordance with the disclosed editing methods comprises a backbone structure recognized by the Staphylococcus aureus Cas9 protein. The backbone structure recognized by the SaCas9 protein can comprise the following sequence: 5'-[guide sequence]-guuuuaguacucuguaaugaaaauuacagaaucuacuaaaacaaggcaaaaugccguguuuaucucgucaacuuguuggcgagauuuuuuu-3' (SEQ ID NO: 78). This is also the backbone structure recognized by SaKKH-Cas9 and SaCas9 ortholog SauriCas9.

[0425] The guide RNA used according to the disclosed editing methods can comprise a backbone structure as listed in Table 2 (SEQ ID NOs: 566-571).

[0426] Based on the present disclosure, the sequences of suitable guide RNAs for targeting the disclosed BEs to specific genomic target sites will be apparent to those skilled in the art. Such suitable guide RNA sequences typically comprise a guide sequence complementary to a nucleic acid sequence within 50 nucleotides upstream or downstream of the target nucleobase pair to be edited. Some exemplary guide RNA sequences suitable for targeting any of the provided BEs to specific target sequences are provided herein. Additional guide sequences are well known in the art and can be used with the base editors described herein. Additional exemplary guide sequences are disclosed in, for example, Jinek M., et al., Science 337:816-821 (2012); Mali P, Esvelt KM & Church GM (2013) Cas9 as a versatile tool for biological engineering, Nature Methods, 10,957-963; Li JF et al., (2013) Multiplex and homologous recombination-mediated genome editing in Arabidopsis and Nicotiana benthamiana using guide RNA andCas9, Nature Biotechnology, 31, 688-691; Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system, Nature Biotechnology 31, 227-229 (2013); Cong L et al., (2013) Multiplex genome engineering using CRIPSR / Cas systems, Science, 339, 819-823; Cho SW et al., (2013) Targeted genome engineering in human cells with the Cas9 RNA-guided endonuclease, NatureBiotechnology, 31, 230-232; Jinek, M. et al., RNA-programmed genome editing in human cells, eLife 2, e00471 (2013); Dicarlo, J.E. et al., Genomeengineering in Saccharomyces cerevisiae using CRISPR-Cas systems. NucleicAcid Res. (2013); Briner AE et al., (2014) Guide RNA functional modules direct Cas9 activity and orthogonality, Mol Cell, 56, 333-339, the entire contents of each of which are incorporated herein by reference.

[0427] Methods for generating Cas variants and base editors

[0428] The present invention, in various aspects, further relates to methods for preparing the disclosed improved base editors through various operating modes, including but not limited to, codon optimization to achieve higher expression levels in cells, and the use of nuclear localization sequences (NLS), preferably at least two NLSs (e.g., two bipartite NLSs) to increase the localization of the expressed base editors in the cell nucleus.

[0429] Preparation of base editors for increased expression in cells

[0430] Base editors contemplated herein can include modifications that result in increased expression, such as through codon optimization.

[0431] In some embodiments, the base editor (or its components) is codon optimized for expression in a specific cell (e.g., a eukaryotic cell). Eukaryotic cells can be those of a specific organism or those derived from a specific organism, such as mammals, including but not limited to humans, mice, rats, rabbits, dogs, or non-human primates. Generally, codon optimization refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest, by replacing at least one codon of a native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with a codon that is more frequently or most frequently used in the genes of the host cell, while maintaining the native amino acid sequence. Various species show specific preferences for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is generally associated with the translation efficiency of messenger RNA (mRNA), which in turn is thought to depend on factors such as the characteristics of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The advantage of the selected tRNA in the cell generally reflects the most commonly used codons in peptide synthesis. Thus, based on codon optimization, genes can be tailored for optimal gene expression in a given organism. Codon usage tables are readily available, for example, at the "Codon Usage Database," and these tables can be adjusted in a variety of ways. See, Y., et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28:292 (2000). Computer algorithms for codon optimization of specific sequences for expression in specific host cells are also available, such as GeneForge (Aptagen; Jacobus, Pa.). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more or all codons) in the sequence encoding the CRISPR enzyme correspond to the most frequently used codons for a specific amino acid.

[0432] Nuclear localization sequence and additional base editor components

[0433] In some embodiments, the base editors provided herein further include one or more nuclear targeting sequences, for example, nuclear localization sequences (NLS). In some embodiments, the NLS includes an amino acid sequence that promotes the import of proteins comprising NLS into the nucleus (for example, by nuclear transport). In some embodiments, any base editors provided herein further include one or more nuclear localization sequences (NLS). In certain embodiments, any base editor includes two NLSs. In some embodiments, one or more of the NLSs are dichotomous NLSs ("bpNLSs"). In certain embodiments, the disclosed base editors include two dichotomous NLSs. In some embodiments, the disclosed base editors include more than two dichotomous NLSs.

[0434] In some embodiments, the NLS is fused to the N-terminus of the base editor. In some embodiments, the NLS is fused to the C-terminus of the base editor. In some embodiments, the NLS is fused to the C-terminus of napDNAbp. In some embodiments, the NLS is fused to the N-terminus of adenosine deaminase. In some embodiments, the NLS is fused to the C-terminus of adenosine deaminase. In some embodiments, the NLS is fused to the base editor via one or more linkers. In some embodiments, the NLS is fused to the base editor without a linker.

[0435] In some embodiments, the NLS comprises the amino acid sequence of any of the NLS sequences provided or referenced herein. In some embodiments, the NLS comprises the amino acid sequence set forth in SEQ ID NO: 408 or SEQ ID NO: 409. Additional nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, the contents of which are incorporated herein by reference for disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 408), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 409), KRTADGSEFESPKKKRKV (SEQ ID NO: 410), or KRTADGSEFEPKKKRKV (SEQ ID NO: 411). In other embodiments, the NLS comprises the amino acid sequence: NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 482), PAAKRVKLD (SEQ ID NO: 483), RQRRNELKRSF (SEQ ID NO: 484), or NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 485).

[0436] In some embodiments, the base editor comprises a bpNLS. The bpNLS may comprise an amino acid sequence selected from the group consisting of KRTADGSEFEPKKKRKV (SEQ ID NO: 398), KRPAATKKAGQAKKKK (SEQ ID NO: 344), KKTELQTTNAENKTKKL (SEQ ID NO: 345), KRGINDRNFWRGENGRKTR (SEQ ID NO: 346), and RKSGKIAAIVVKRPRK (SEQ ID NO: 347). In certain embodiments, the bpNLS comprises the amino acid sequence shown in SEQ ID NO: 344 or 398.

[0437] In some embodiments, the base editors provided herein do not include a joint. In some embodiments, a joint is present between one or more domains or proteins (e.g., deaminase, napDNAbp, and / or NLS). In some embodiments, the "]-[" used in the general framework above indicates the presence of an optional joint.

[0438] In some embodiments, the general architecture of an exemplary base editor having a first adenosine deaminase, a second adenosine deaminase, and a napDNAbp domain comprises any of the following structures, wherein NLS is a nuclear localization sequence (e.g., any NLS provided herein), NH2 is the N-terminus of the base editor, and COOH is the C-terminus of the base editor.

[0439] An exemplary base editor comprising a deaminase, a napDNAbp domain, and an NLS (e.g., any NLS provided herein) can have the following architecture:

[0440] NH2-[deaminase domain]-[napDNAbp domain]-[NLS]-COOH;

[0441] NH2-[napDNAbp domain]-[deaminase domain]-[NLS]-COOH;

[0442] NH2-[NLS]-[deaminase domain]-[napDNAbp domain]-COOH; or

[0443] NH2-[NLS]-[napDNAbp domain]-[deaminase domain]-COOH.

[0444] In certain embodiments, the disclosed base editors comprise the following ABE architecture, wherein TadA-8e is an adenosine deaminase domain:

[0445] NH2-[bpNLS]-[TadA-8e]-[napDNAbp domain]-[bpNLS]-COOH;

[0446] NH2-[bpNLS]-[napDNAbp domain]-[TadA-8e]-[bpNLS]-COOH;

[0447] NH2-[bpNLS]-[TadA-8e]-[napDNAbp domain]-[bpNLS]-COOH; or

[0448] NH2-[bpNLS]-[napDNAbp domain]-[TadA-8e]-[bpNLS]-COOH.

[0449] Exemplary base editors comprising a cytosine deaminase, a napDNAbp domain, a UGI domain, and an NLS (e.g., any NLS provided herein) can have the following architecture: NH2-[napDNAbp domain]-[cytosine deaminase domain]-[UGI domain]-[bpNLS]-COOH; or NH2-[bpNLS]-[napDNAbp domain]-[cytosine deaminase domain]-[UGI domain]-COOH; NH2-[cytosine deaminase domain]-[napDNAbp domain]-[UGI domain]-[bpNLS]-COOH; or NH2-[bpNLS]-[cytosine deaminase domain]-[napDNAbp domain]-[UGI domain]-COOH.

[0450] Representative nuclear localization signal is a peptide sequence that guides protein to the nucleus of the cell expressing the sequence. Nuclear localization signal is mainly alkaline, can be positioned at almost any position in the protein amino acid sequence, usually comprises four amino acids (Autieri & Agrawal, (1998) J.Biol.Chem.273: 14731-37, incorporated herein by reference) to eight amino acid short sequences, and is usually rich in lysine and arginine residues (Magin et al., (2000) Virology 274: 11-16, incorporated herein by reference). Nuclear localization signal usually comprises proline residues. A variety of nuclear localization signals have been identified, and they have been used to realize the transport of biomolecules from cytoplasm to nucleus. See, e.g., Tinland et al., (1992) Proc. Natl. Acad. Sci. USA 89:7442-46; Moede et al., (1999) FEBS Lett. 461:229-34, which are incorporated herein by reference. It is currently believed that translocation involves nucleoporins.

[0451] Most NLSs can be divided into three major categories: (i) monopartite NLSs, such as the SV40 large T antigen NLS (PKKKRKV (SEQ ID NO: 408); (ii) bipartite motifs consisting of two basic domains separated by a variable number of spacer amino acids, such as the Xenopus laevis nucleoplasmin NLS (KRXXXXXXXXXXKKKL (SEQ ID NO: 486)); and (iii) atypical sequences, such as M9 of the hnRNP Al protein, influenza virus nucleoprotein NLS, and yeast Gal4 protein NLS (Dingwall and Laskey, Trends Biochem Sci. 1991 Dec;16(12):478-81).

[0452] Nuclear localization signals appear at different points in the amino acid sequence of a protein. NLS has been identified at the N-terminus, C-terminus, and central region of a protein. Therefore, this specification provides a base editor that can be modified with one or more NLSs at the C-terminus, N-terminus, and internal regions of a base editor. Residues of longer sequences that do not function as constituent NLS residues should be selected so as not to interfere with the nuclear localization signal itself, for example, in tension or space. Therefore, although there are no strict restrictions on the composition of sequences comprising NLS, in practice, such sequences may be functionally limited in length and composition.

[0453] The present disclosure contemplates any suitable manner of modifying a fusion protein (or base editor) to include one or more NLSs. In one aspect, the base editor can be engineered to express a fusion protein that is translated and fused to one or more NLSs at its N-terminus or its C-terminus (or both), i.e., to form a fusion protein-NLS fusion construct. In other embodiments, the nucleotide sequence encoding the fusion protein can be genetically modified to incorporate a reading frame encoding one or more NLSs in the internal region of the encoded fusion protein. In addition, NLS can include various amino acid linkers or spacers encoded between the fusion protein and the N-terminus, C-terminus, or internally attached NLS amino acid sequence. Therefore, the present disclosure also provides nucleotide constructs, vectors, and host cells for expressing a base editor comprising a fusion protein and one or more NLSs.

[0454] The base editors described herein may also include a nuclear localization signal, which is connected to the fusion protein via one or more joints (e.g., polymers, amino acids, polysaccharides, chemicals, or nucleic acid joint elements). In certain embodiments, an XTEN joint is used, as shown in SEQ ID NO: 412, to connect the NLS to the fusion protein. The joints within the intended scope of the present disclosure are not intended to have any limitations and may be any suitable type of molecule (e.g., polymers, amino acids, polysaccharides, nucleic acids, lipids, or any synthetic chemical joint domains), and are connected to the fusion protein via any suitable strategy that can achieve the formation of a bond (e.g., covalent bond, hydrogen bond) between the fusion protein and one or more NLSs.

[0455] The base editors described herein may also include one or more additional elements. In certain embodiments, the additional elements may include effectors of base repair, such as inhibitors of base repair.

[0456] In some embodiments, the base editors described herein may comprise one or more heterologous protein domains (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more domains in addition to the base editor components). The base editor may comprise any additional protein sequence, and optionally a linker sequence between any two domains. Other exemplary features that may be present are localization sequences, such as cytoplasmic localization sequences, export sequences, such as nuclear export sequences or other localization sequences, and sequence tags.

[0457] Examples of heterologous protein domains that can be fused to a base editor or its components (e.g., napDNAbp domains, nucleotide modification domains, or NLS domains) include, but are not limited to, epitope tags and reporter gene sequences. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins, including blue fluorescent protein (BFP). The base editor can be fused to a gene sequence encoding a protein or protein fragment that binds to a DNA molecule or binds to other cellular molecules, including but not limited to maltose binding protein (MBP), S-tag, LexA DNA binding domain (DBD) fusion, Gal4 DNA binding domain fusion, and herpes simplex virus (HSV) BP16 protein fusion. Additional domains that can form part of a base editor are described in U.S. Patent Publication No. 2011 / 0059502, published on March 10, 2011, which is incorporated herein by reference in its entirety.

[0458] In aspects of the present disclosure, reporter gene can be introduced into the cell to encode the gene product used as a marker, by which the change or modification of gene product expression is measured, reporter gene includes but is not limited to glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP) and autofluorescent protein, including blue fluorescent protein (BFP). In certain embodiments of the present disclosure, the gene product is luciferase. In further embodiments of the present disclosure, the expression of the gene product is reduced.

[0459] Other exemplary features that may be present are tags that can be used for the dissolution, purification or detection of base editors. Suitable protein tags provided herein include but are not limited to biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin-tags, FLAG tags, hemagglutinin (HA) tags, bgh poly A tags, polyhistidine tags and also referred to as histidine tags or His tags, maltose binding protein (MBP) tags, nus tags, glutathione-S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag 1, Softag 3), strep- tags, biotin ligase tags, FlAsH tags, V5 tags and SBP tags. Other suitable sequences will be apparent to those skilled in the art. In some embodiments, the base editor includes one or more His tags.

[0460] connector

[0461] In certain embodiments, a linker can be used to connect any peptide or peptide domain or domain of a base editor (e.g., a napDNAbp domain covalently linked to an adenosine deaminase domain, which is covalently linked to an NLS domain). The base editors described herein may comprise a linker having a length of 32 amino acids.

[0462] The joint can be as simple as a covalent bond, or it can be a polymer joint with a length of many atoms. In certain embodiments, the joint is a polypeptide or based on amino acids. In other embodiments, the joint is not peptide-like. In certain embodiments, the joint is a covalent bond (e.g., carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In certain embodiments, the joint is a carbon-nitrogen bond connected by an amide. In certain embodiments, the joint is a cyclic or non-cyclic, replaced or unreplaced, branched or non-branched aliphatic or heteroaliphatic joint. In certain embodiments, the joint is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the joint comprises a monomer, dimer or polymer of aminoalkanoic acid. In certain embodiments, the joint comprises an aminoalkanoic acid (e.g., glycine, acetate, alanine, β-alanine, 3-aminopropionic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.). In certain embodiments, the joint comprises a monomer, dimer or polymer of aminocaproic acid (Ahx). In certain embodiments, the joint is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the joint comprises a polyethylene glycol moiety (PEG). In other embodiments, the joint comprises an amino acid. In certain embodiments, the joint comprises a peptide. In certain embodiments, the joint comprises an aromatic or heteroaromatic moiety. In certain embodiments, the joint is based on a benzene ring. The joint can include a functionalized portion to promote the attachment of nucleophiles (e.g., thiol, amino) to the joint from the peptide. Any electrophilic reagent can be used as a part of a joint. Exemplary electrophilic reagents include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0463] In some embodiments, the length of the linker is 5-100 amino acids, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-110, 110-120, 120-130, 130-140, 140-150 or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the length of the linker is 32 amino acids in length. In an exemplary embodiment, the linker comprises the 32 amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 412), also referred to as an XTEN linker. In some embodiments, the linker comprises the 9 amino acid sequence SGGSGGSGGS (SEQ ID NO: 413). In some embodiments, the linker comprises the 4 amino acid sequence SGGS (SEQ ID NO: 414).

[0464] In some embodiments, the linker comprises the amino acid sequence (GGGGS) n (SEQ ID NO: 415), (G) n (SEQ ID NO: 416), (EAAAK) n (SEQ ID NO: 417), (GGS) n (SEQ ID NO: 418), (SGGS) n (SEQ ID NO: 419), (XP) n (SEQ ID NO: 420) or any combination thereof, wherein n is independently an integer between 1 and 30, and wherein X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS) n (SEQ ID NO: 421), wherein n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 422).

[0465] In some embodiments, the linker comprises SGSETPGTSESATPES (SEQ ID NO: 422) and SGGS (SEQ ID NO: 414). In some embodiments, the linker comprises SGGSSGSETPGTSESATPESSGGS (SEQ ID NO: 423). In some embodiments, the linker comprises SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 412). In some embodiments, the linker comprises GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (SEQ ID NO: 424). In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES (SEQ ID NO: 425). In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS (SEQ ID NO: 426). In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 427). In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS (SEQ ID NO: 428). It will be appreciated that any linker provided herein can be used to connect a first adenosine deaminase and a second adenosine deaminase; an adenosine deaminase domain (comprising, e.g., a first and / or second adenosine deaminase) and a napDNAbp; a napDNAbp and an NLS; or an adenosine deaminase domain and an NLS.

[0466] In some embodiments, any base editor provided herein comprises an adenosine deaminase and a napDNAbp fused to each other via a joint. In some embodiments, any base editor provided herein comprises a first adenosine deaminase and a second adenosine deaminase fused to each other via a joint. In some embodiments, any base editor provided herein comprises an NLS, which can be fused to an adenosine deaminase (e.g., a first and / or second adenosine deaminase) and a nucleic acid programmable DNA binding protein (napDNAbp). Various linker lengths and flexibilities can be employed between the adenosine deaminase (e.g., engineered ecTadA) and the napDNAbp (e.g., Cas9 domain), and / or between the first adenosine deaminase and the second adenosine deaminase (e.g., ranging from very flexible linkers of the form of SEQ ID NOs: 119, 121-124 (see, e.g., Guilinger JP, Thompson DB, Liu DR. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82; herein incorporated by reference in its entirety) to (XP) n (SEQ ID NO: 420)) to achieve the optimal length for deaminase activity for a particular application. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises the amino acid sequence (GGS) n (SEQ ID NO: 421) motif, wherein n is 1, 3 or 7. In some embodiments, the adenosine deaminase and napDNAbp and / or the first adenosine deaminase and the second adenosine deaminase of any base editor provided herein are fused via a linker comprising an amino acid sequence selected from SEQ ID NO: 119-132. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence (SGGS) 2-SGSETPGTSESATPES-(SGGS) 2 (SEQ ID NO: 412), which may also be referred to as (SGGS) 2-XTEN-(SGGS) 2 (SEQ ID NO: 429). In some embodiments, the linker comprises an amino acid sequence, wherein n is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker is 92 amino acids in length.

[0467] The above description is non-limiting for preparing base editors with increased expression, thereby improving editing efficiency.

[0468] Treatment methods and uses

[0469] Other aspects of the present disclosure provide methods for delivering a base editor into a cell to form a complete and functional Cas9 protein or a core base editor. For example, in some embodiments, a cell is contacted with a composition as described herein (e.g., a composition comprising a nucleotide sequence encoding a base editor or an AAV particle containing a nucleic acid vector comprising such a nucleotide sequence). In some embodiments, contact results in the delivery of such a nucleotide sequence into a cell, wherein the N-terminal portion of the Cas9 protein or the core base editor and the C-terminal portion of the Cas9 protein or the core base editor are expressed in the cell and connected to form a complete Cas9 protein or a complete core base editor.

[0470] It should be understood that any rAAV particles, nucleic acid molecules or compositions provided herein can be stably or transiently introduced into cells in any suitable manner. In some embodiments, the disclosed protein can be transfected into cells. In some embodiments, cells can be transduced or transfected by nucleic acid molecules. For example, cells can be transduced (e.g., with protein-encoding viruses) or transfected (e.g., with protein-encoding plasmids) with rAAV particles of protein-encoding nucleic acid molecules or viral genomes containing one or more nucleic acid molecules. This transduction can be a stable or transient transduction. In some embodiments, cells expressing proteins or containing proteins can be transduced or transfected with one or more guide RNA sequences, such as in the delivery of base editors. In some embodiments, protein-expressing plasmids can be introduced into cells by electroporation, transient transfection (e.g., lipofection) and stable genomic integration (e.g., nucleofection and piggybac) and viral transduction or other methods known to those skilled in the art. ...

Claims

1. A nucleic acid molecule comprising: (i) 5' inverted terminal repeat (ITR); (ii) a first nucleic acid fragment comprising a sequence encoding a base editor operably linked to a first promoter, wherein the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain; and a polyadenylation (poly A) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a second promoter; and (iv) 3'ITR, wherein the length between the 5' ITR and the 3' ITR is less than about 4.90 kb.

2. The nucleic acid molecule of claim 1, wherein the first nucleic acid segment does not encode an intein.

3. The nucleic acid molecule of claim 1 or 2, wherein the nucleic acid molecule does not comprise a post-transcriptional response element.

4. The nucleic acid molecule of any one of claims 1-3, wherein the second promoter is a U6 promoter.

5. The nucleic acid molecule of any one of claims 1-4, wherein the first promoter has a length of less than 325 nucleotides, less than 315 nucleotides, less than 300 nucleotides, less than 285 nucleotides, less than 270 nucleotides, less than 265 nucleotides, or less than 250 nucleotides.

6. The nucleic acid molecule of any one of claims 1-5, wherein the first promoter has a length of about 280 nucleotides.

7. The nucleic acid molecule of any one of claims 1-5, wherein the first promoter is an EF-1 alpha short (EFS) promoter, a MeCP2 promoter, a P3 promoter, or a U1A promoter.

8. The nucleic acid molecule of any one of claims 1-7, wherein the first promoter is the EF-1 alpha short (EFS) promoter.

9. The nucleic acid molecule of any one of claims 1-8, wherein the first promoter is a tissue-specific promoter.

10. The nucleic acid molecule of claim 9, wherein the first promoter is a cardiac tissue-specific promoter, a muscle tissue-specific promoter, or a neuronal tissue-specific promoter.

11. The nucleic acid molecule of any one of claims 1-10, wherein the base editor comprises (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain; and (ii) a deaminase domain.

12. The nucleic acid molecule of any one of claims 1-11, wherein the napDNAbp domain is a Cas9 domain.

13. The nucleic acid molecule of any one of claims 1-12, wherein the napDNAbp domain is selected from the group consisting of a Staphylococcus aureus Cas9 (SaCas9) domain, a Neisseria meningitidis 2 Cas9 (Nme2Cas9) domain, a Campylobacter jejuni Cas9 (CjCas9) domain, a Staphylococcus auris (SauriCas9) domain, and variants thereof.

14. The nucleic acid molecule of any one of claims 1-13, wherein the napDNAbp domain has nickase activity.

15. The nucleic acid molecule of any one of claims 1-14, wherein the napDNAbp domain is a compact variant of Streptococcus pyogenes Cas9 (SpCas9), Cpf1, CasX, CasY, C2c1, C2c2, C2c3, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Cas3, or CasΦ.

16. The nucleic acid molecule of any one of claims 1-14, wherein the napDNAbp domain is a SaCas9 domain, a SaCas9 nickase domain, a SaKKH domain, or a SaKKH nickase domain.

17. The nucleic acid molecule of any one of claims 1-14 and 16, wherein the napDNAbp domain is a SaKKH nickase domain.

18. The nucleic acid molecule of any one of claims 1-14 and 16, wherein the napDNAbp domain is a CjCas9 nickase domain.

19. The nucleic acid molecule of any one of claims 1-18, wherein the deaminase domain comprises an adenosine deaminase.

20. The nucleic acid molecule of any one of claims 1-19, wherein the deaminase domain consists of an adenosine deaminase monomer.

21. The nucleic acid molecule of any one of claims 1-20, wherein the deaminase domain is a TadA-8e, TadA-8e(V106W), TadA9, TadA20, or TadA7.10 deaminase.

22. The nucleic acid molecule of any one of claims 1-21, wherein the deaminase domain is a TadA-8e deaminase domain.

23. The nucleic acid molecule of any one of claims 1-22, wherein the base editor further comprises one or more nuclear localization sequences (NL).

24. The nucleic acid molecule of any one of claims 1-23, wherein the base editor further comprises a bipartite NLS (bpNLS).

25. The nucleic acid molecule of claim 24, wherein the bipartite nuclear localization signal comprises an amino acid sequence selected from the group consisting of KRTADGSEFEPKKKRKV (SEQ ID NO: 398), KRPAATKKAGQAKKKK (SEQ ID NO: 344), KKTELQTTNAENKTKKL (SEQ ID NO: 345), KRGINDRNFWRGENGRKTR (SEQ ID NO: 346), and RKSGKIAAIVVKRPRK (SEQ ID NO: 347).

26. The composition of claim 24 or 25, wherein the bipartite nuclear localization signal comprises the amino acid sequence shown in SEQ ID NO: 344 or 398.

27. The nucleic acid molecule of any one of claims 1-26, wherein the base editor is ABE8e, ABE8e(V106W), ABE9, ABE20, ABE7.10, or a variant thereof.

28. The nucleic acid molecule of any one of claims 1-27, wherein the base editor is SaKKH-ABE8e, SauriCas9-ABE8e, CjCas9-ABE8e, Nme2Cas9-ABE8e, or SaCas9-ABE8e.

29. The nucleic acid molecule of any one of claims 1-28, wherein the base editor comprises the following structure: NH2-[adenosine deaminase]-[napDNAbp domain]-COOH; or NH2-[napDNAbp domain]-[adenosine deaminase]-COOH, wherein each "]-[" in the structure indicates the presence of an optional linker sequence.

30. The nucleic acid molecule of any one of claims 1-18 and 23-26, wherein the deaminase domain comprises a cytosine deaminase.

31. The nucleic acid molecule of any one of claims 1-18, 23-26, and 30, wherein the deaminase domain is FERNY or evolved FERNY (evoFERNY).

32. The nucleic acid molecule of claims 1-18, 23-26, 30, and 31, wherein the base editor is BE3.9, FERNY-BE3.9, or a variant thereof.

33. The nucleic acid molecule of any one of claims 1-18, 23-26, and 30-32, wherein the base editor further comprises a uracil glycosylase inhibitor (UGI) domain.

34. The nucleic acid molecule of any one of claims 1-18 and 30-33, wherein the base editor comprises the following structure: NH2-[cytidine deaminase domain]-[napDNAbp domain]-[UGI]-COOH; NH2-[cytidine deaminase domain]-[UGI]-[napDNAbp domain]-COOH; NH2-[napDNAbp domain]-[UGI]-[cytidine deaminase domain]-COOH; NH2-[napDNAbp domain]-[cytidine deaminase domain]-[UGI]-COOH; NH2-[UGI]-[cytidine deaminase domain]-[napDNAbp domain]-COOH; or NH2-[UGI]-[napDNAbp domain]-[cytidine deaminase domain]-COOH, Each instance of "]-[" in the structures described herein indicates the presence of an optional linker sequence.

35. The nucleic acid molecule of any one of claims 1-18, 23-26, and 30-33, wherein the base editor comprises the amino acid sequence of SEQ ID NO: 21 or 22.

36. The nucleic acid molecule of any one of claims 1-29, wherein the base editor comprises the amino acid sequence of any one of SEQ ID NOs: 171-172 and 181-183.

37. The nucleic acid molecule of any one of claims 1-29 and 36, wherein the base editor comprises the amino acid sequence of SEQ ID NO:

171.

38. The nucleic acid molecule of any one of claims 1-37, wherein the poly A signal is a bovine growth hormone (bGH) signal, a human growth hormone (hGH) signal, or an SV40 signal.

39. The nucleic acid molecule of any one of claims 1-38, wherein the poly A signal is a bovine growth hormone (bGH) signal.

40. The nucleic acid molecule of any one of claims 1-39, wherein the first nucleic acid segment further comprises a minute virus of mice (MVM) intron.

41. The nucleic acid molecule of any one of claims 1-40, wherein the length between the 5' ITR and the 3' ITR is between 4.7 kb and 4.9 kb.

42. The nucleic acid molecule of any one of claims 1-41, wherein the length between the 5' ITR and the 3' ITR is about 4.65 kb, about 4.70 kb, about 4.725 kb, about 4.75 kb, about 4.80 kb, about 4.825 kb, about 4.85 kb, or about 4.90 kb.

43. The nucleic acid molecule of any one of claims 1-42, wherein the length between the 5' ITR and the 3' ITR is about 4.80 kb.

44. The nucleic acid molecule of any one of claims 1-43, wherein the nucleic acid molecule consists essentially of: (i) the 5' inverted terminal repeat (ITR); (ii) a sequence encoding a base editor operably linked to a first promoter; (iii) the polyadenylation (poly A) signal; (iv) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a second promoter, wherein the transcription direction of the second nucleic acid fragment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) the 3'ITR, wherein the length between the 5' ITR and the 3' ITR is less than about 4.90 kb.

45. The nucleic acid molecule of any one of claims 1-44, wherein the nucleic acid molecule is single-stranded or self-complementary.

46. The nucleic acid molecule of any one of claims 1-29 and 36-45, wherein the nucleic acid molecule comprises the sequence of any one of SEQ ID NOs: 100-102.

47. The nucleic acid molecule of any one of claims 1-29 and 36-46, wherein the nucleic acid molecule comprises the sequence of SEQ ID NO:

100.

48. The nucleic acid molecule of any one of claims 1-47, wherein the direction of transcription of the second nucleic acid segment is reversed relative to the direction of transcription of the first nucleic acid segment.

49. The nucleic acid molecule of any one of claims 1-29 and 36-48, wherein the nucleic acid molecule comprises from 5' to 3': (i) 5' inverted terminal repeat (ITR); (ii) a first nucleic acid fragment comprising a sequence encoding a SaKKH-ABE8e, SauriCas9-ABE8e, CjCas9-ABE8e base editor, or Nme2Cas9 base editor operably linked to a first promoter selected from the group consisting of EFS, MeCP2, P3, and U1A promoters; and a bGH polyadenylation (poly A) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the transcription direction of the second nucleic acid fragment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) 3'ITR, wherein the length between the 5' ITR and the 3' ITR is less than about 4.90 kb.

50. The nucleic acid molecule vector of claim 49, wherein the first promoter is EFS.

51. The nucleic acid molecule vector of claim 49 or 50, wherein the base editor is SaKKH-ABE8e.

52. The nucleic acid molecule of any one of claims 1-18, 23-26, 30-35, and 38-43, wherein the nucleic acid molecule comprises, from 5' to 3': (i) 5' inverted terminal repeat (ITR); (ii) a first nucleic acid fragment comprising a sequence encoding a CjCas9-FERNY-BE3.9 or CjCas9-evoFERNY-BE3.9 base editor operably linked to a first promoter selected from the group consisting of EFS, MeCP2, P3, and U1A promoters; and a bGH polyadenylation (poly A) signal; (iii) a second nucleic acid fragment encoding a guide RNA (gRNA) operably linked to a U6 promoter, wherein the transcription direction of the second nucleic acid fragment is reverse relative to the transcription direction of the nucleic acid molecule; and (iv) 3'ITR, wherein the length between the 5' ITR and the 3' ITR is less than about 4.90 kb.

53. The nucleic acid molecule vector of claim 52, wherein the first promoter is EFS.

54. The nucleic acid molecule vector of claim 52 or 53, wherein the base editor is CjCas9-evoFERNY-BE3.

9.

55. An AAV nucleic acid molecule comprising, in 5' to 3' order: (i) 5' inverted terminal repeat (ITR); (ii) a first nucleic acid fragment comprising a transgene operably linked to a first promoter, wherein the first promoter has a length of less than 300 nucleotides; and a transcription terminator that does not contain a post-transcriptional response element; (iii) a second nucleic acid fragment operably linked to a second promoter, wherein the transcription direction of the second nucleic acid fragment is reverse relative to the transcription direction of the first nucleic acid fragment; and (iv) 3'ITR, wherein the length between the 5' ITR and the 3' ITR is less than about 4.90 kb.

56. The AAV nucleic acid molecule of claim 55, wherein the first nucleic acid segment encodes a base editor and the second nucleic acid segment encodes a gRNA.

57. The AAV nucleic acid molecule of claim 56, wherein the base editor contains a napDNAbp domain, which is a compact Cas9 protein.

58. The AAV nucleic acid molecule of claim 57, wherein the napDNAbp domain is selected from the group consisting of a Staphylococcus aureus Cas9 (SaCas9) domain, a Neisseria meningitidis 2 Cas9 (Nme2Cas9) domain, a Campylobacter jejuni Cas9 (CjCas9) domain, a Staphylococcus auris (SauriCas9) domain, and variants thereof.

59. The nucleic acid molecule of claim 55, wherein the napDNAbp domain is a compact variant of Streptococcus pyogenes Cas9 (SpCas9), Cpf1, CasX, CasY, C2c1, C2c2, C2c3, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Cas3, or CasΦ.

60. The nucleic acid molecule of any one of claims 55-59, wherein the napDNAbp domain has nickase activity.

61. A recombinant AAV (rAAV) particle comprising the nucleic acid molecule of any one of claims 1-60 encapsidated in a capsid.

62. The rAAV particle of claim 61, wherein the capsid is of serotype 2, 6, 8, 9, PHP.B, or PHP.eB.

63. The rAAV particle of claim 61 or 62, wherein the capsid is of serotype 8 or serotype 9.

64. The rAAV particle of any one of claims 61-63, wherein the rAAV particle is a rAAV8-Sauri-ABE8e, rAAV9-Sauri-ABE8e, rAAV8-SaKKH-ABE8e, rAAV9-SaKKH-ABE8e, rAAV8-CjCas9-ABE8e, rAAV9-CjCas9-ABE8e, rAAV8-Nme2Cas9-ABE8e, or rAAV9-Nme2Cas9-ABE8e particle.

65. The rAAV particle of any one of claims 61-64, wherein the rAAV particle is a rAAV8-SaKKH-ABE8e particle.

66. A composition comprising a plurality of the rAAV particles of any one of claims 61-65.

67. A composition comprising the nucleic acid molecule of any one of claims 1-60.

68. A pharmaceutical composition comprising the composition of claim 66 or 67 and a pharmaceutically acceptable carrier.

69. A cell comprising the nucleic acid molecule of any one of claims 1-60, the rAAV particle of any one of claims 61-65, or the composition of any one of claims 66-68.

70. The cell of claim 69, wherein the cell is a bacterial cell.

71. The cell of claim 69, wherein the cell is a eukaryotic cell.

72. The cell of claim 69 or 71, wherein the cell is a yeast cell, a plant cell, or a mammalian cell.

73. The cell of claim 69, wherein the cell is a human cell.

74. A kit comprising the nucleic acid molecule of any one of claims 1-60, the rAAV particle of any one of claims 61-65, or the composition of any one of claims 66-68, and instructions for delivery to a cell.

75. A method for editing a target nucleic acid molecule, the method comprising contacting a cell comprising the target nucleic acid molecule with the nucleic acid molecule of any one of claims 1-60, the rAAV particle of any one of claims 61-65, or the composition of any one of claims 66-68.

76. The method of claim 75, wherein the contacting is performed in vitro.

77. The method of claim 75, wherein the contacting is performed in vivo.

78. The method of claim 75, wherein the contacting is performed in a subject.

79. The method of claim 78, wherein the subject has been diagnosed with a disease or condition.

80. The method of any one of claims 75-79, wherein the guide RNA comprises a guide sequence of at least 10 contiguous nucleotides that is complementary to a target sequence in the target nucleic acid molecule.

81. The method of any one of claims 75-80, wherein the guide RNA is about 25 to 200 nucleotides in length.

82. The method of any one of claims 75-81, wherein the target sequence comprises a point mutation associated with a disease or disorder.

83. The method of claim 82, wherein the point mutation comprises a T→C point mutation associated with the disease or condition.

84. The method of claim 82, wherein the point mutation comprises an A→G point mutation associated with the disease or condition.

85. The method of any one of claims 82-84, wherein the step of editing the target nucleic acid results in correction of the point mutation.

86. The method of any one of claims 75-85, wherein the contacting step results in the introduction of a splice site.

87. The method of any one of claims 75-85, wherein the contacting step results in removal of a splice site.

88. The method of any one of claims 75-85, wherein said contacting step results in the removal of a stop codon.

89. The method of any one of claims 75-85, wherein said contacting step results in the introduction of a stop codon.

90. The method of claim 89, wherein the stop codon comprises the nucleic acid sequence 5'-TAG-3', 5'-TAA-3', or 5'-TGA-3'.

91. The method of any one of claims 75-90, wherein the target sequence is within a PCSK9 gene or an ANGPTL3 gene.

92. The method of any one of claims 75-91, wherein the contacting step results in an editing efficiency of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, or at least about 85%.

93. The method of any one of claims 75-92, wherein the contacting step results in a base edit:indel ratio of at least about 5:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, or greater than about 15:

1.

94. The method of any one of claims 75-93, wherein the contacting step results in an indel rate of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less.

95. The method of any one of claims 75-94, wherein the contacting step results in minimal bystander editing.

96. A method comprising contacting a cell with the rAAV particle of any one of claims 61-65 or the composition of any one of claims 66-68, wherein the contacting results in delivery of the nucleic acid molecule into the cell.

97. The method of claim 96, wherein the cell is a human cell.

98. The method of claim 96 or 97, wherein the cell is a cardiac cell or a muscle cell.

99. The method of claim 98, wherein the contacting step results in an editing efficiency of at least about 20%, at least about 22%, at least about 24%, at least about 27%, at least about 30%, at least about 33%, or at least about 36% in the heart cells or muscle cells.

100. The method of any one of claims 96-99, wherein the contacting step results in an indel rate of 2.5% or less, 2.0% or less, 1.5% or less, 1.2% or less, or 1.0% or less in the heart cells or muscle cells.

101. A method comprising administering to a subject in need thereof a therapeutically effective amount of the rAAV particle of any one of claims 61-65, the pharmaceutical composition of claim 68, or the cell of any one of claims 69-73.

102. The method of claim 101, wherein the subject is a human.

103. The method of claim 101 or 102, wherein the subject has a disease or condition.

104. The method of any one of claims 101-103, wherein the rAAV particles are administered at a dose of about 10 g / kg body weight of the subject. 15 , about 10 14 , about 10 13 , about 10 12 , about 10 11 or less than about 10 11 A therapeutically effective amount of one vector genome (vg) is administered.

105. The method of any one of claims 101-104, wherein the rAAV particles are administered at a dose of about 8 x 10 10 , 4x10 10 , 1x10 10 or less than about 10 10 A therapeutically effective amount of vg is administered.

106. The method of any one of claims 101-105, wherein the rAAV particles are administered to cardiac tissue of the subject.

107. The method of any one of claims 101-106, wherein the rAAV particles are administered to skeletal muscle tissue of the subject.

108. The method of any one of claims 101-107, wherein the rAAV particles are administered to neuronal tissue of the subject.

109. Use of the nucleic acid molecule of any one of claims 1-60, the rAAV particle of any one of claims 61-65, the composition of any one of claims 66-68, or the cell of any one of claims 69-73 in the preparation of a medicament.

110. A method for preparing the rAAV particles of any one of claims 61-65.

Citation Information

Patent Citations

  • CAS9 proteins including ligand-dependent inteins

    US10077453B2

  • Adenosine nucleobase editors and uses thereof

    US10113163B2

  • Nucleobase editors and uses thereof

    US10167457B2

  • Process and device for inerting an aircraft fuel tank

    US20020158167A1

  • Regulation of endogenous gene expression in cells using zinc finger proteins

    US20030087817A1

Cited By

  • Media portrait generation method and system based on natural language processing

    CN121278103A