Rett Syndrome Treatment

The use of an adenine base editor to correct MECP2 gene mutations in Rett syndrome addresses gene dosage and repeated administration challenges, providing permanent functional MECP2 protein expression for a substantial patient group.

JP2025530183APending Publication Date: 2025-09-11THE UNIV COURT OF THE UNIV OF EDINBURGH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025514212
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-08
Filing Date
2023-09-07
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Current therapeutic strategies for Rett syndrome, particularly those targeting C-terminal deletions in the MECP2 gene, face challenges such as gene dosage issues, the need for repeated administration, and inefficiency in correcting mutations in postmitotic neurons like those found in the brain.

Method used

Employing a base editing construct, specifically an adenine base editor (ABE), to edit the stop codon in mutant MECP2 genes, allowing translational readthrough and maintaining functional MECP2 protein levels without overexpression, using a single guide RNA (sgRNA) to target the adenine base editor to the TGA stop codon.

Benefits of technology

Achieves permanent correction of MECP2 mutations, avoiding gene dosage issues and neurological disorders, and is applicable to a significant portion of Rett syndrome patients with C-terminal deletions, ensuring functional MECP2 protein expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025530183000001_ABST
    Figure 2025530183000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to compositions for use in treating a class of Rett syndrome mutations, i.e., C-terminal deletions, comprising a base editor that alters a stop codon in a mutant MECP2 gene. This alteration does not return the gene to its wild-type (WT) form, but rather re-establishes normal levels of a version of the MeCP2 protein that is functionally equivalent to the wild-type version.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Field of Disclosure The present disclosure relates to compositions for use in treating a class of Rett syndrome mutations, i.e., C-terminal deletions, comprising a base editor that alters a stop codon in a mutant MECP2 gene. This alteration does not return the gene to its wild-type (WT) form, but rather re-establishes normal levels of a version of the MeCP2 protein that is functionally equivalent to the wild-type version. [Background technology]

[0002] Background to the disclosure Rett syndrome (RTT) is a severe neurological disorder caused primarily by mutations in the X-linked gene MECP2. Loss of MeCP2 function has the most severe impact in the nervous system and minimal phenotypic consequences in other tissues (Ross et al. 2016). Due to the association of RTT with intellectual disability in other disorders, the MECP2 gene is frequently screened for mutations in clinical cases of developmental delay. This has resulted in the identification of numerous amino acid changes as variants that cause RTT or relatively benign or neutral variants that do not cause RTT (Krishnaraj et al. 2017, and Karczewski, et al. 2020). The RTT-like symptoms in Mecp2-null mice can be rescued by restoring MeCP2 expression, suggesting that the disorder may be curable. Most RTT-causing mutations disrupt two domains required for normal MeCP2 function, but some mutations affect the C-terminal region of these domains. These mutations cause RTT by significantly reducing MeCP2 protein levels (Guy et al. 2018). Interestingly, mice lacking this C-terminal region, which express normal levels of MeCP2 protein, do not have RTT-like symptoms, indicating that this region is not required for normal MeCP2 function. Summary of the Invention

[0003] Disclosure Overview A heterogeneous group of frameshift C-terminal deletions (CTDs) accounts for approximately 10% of mutations causing RTT. They occur within the deletion-prone region of MECP2, approximately 1110–1210 (numbered based on the e2 isoform) (Bebbington et al., 2010). We found that the outcome of deletions in this region varies depending on the resulting reading frame. In-frame deletions that remove portions of DNA in this region but still maintain the WT reading frame are present in large-scale sequencing databases (which exclude individuals with severe childhood disease), indicating that they are nonpathogenic, neutral variants (see Figure 3). The same is true for frameshift deletions that shift frame (n+1) and terminate with the amino acid sequence -Ser-Pro-Arg-Thr-Stop. However, additional RTT-causing frameshift deletions that appear in RettBASE (a database listing mutations found in RTT patients) shift to frame (n+2) and terminate with -Pro-Pro-Stop (see Figure 2).

[0004] The present disclosure is based on the inventors' initial hypothesis that for RTT subjects with deletions in exon 4 of MECP2 resulting in a frame (n+2) shift, the presence of two prolines before the stop codon disrupts translation termination, resulting in loss of MeCP2 protein and mRNA in a process similar to nonsense-mediated mRNA decay. Based on this, the inventors predicted that changing the stop codon after the two prolines to one encoding an amino acid would lead to further translation to the next stop codon, potentially preventing the loss of MeCP2 protein that is pathogenic in this class of mutations. Accordingly, in a first aspect, the disclosure provides a base editing construct for use in a method of treating Rett syndrome in a subject, wherein the subject comprises a mutant MECP2 gene comprising a translational frameshift (n+2) and a C-terminal deletion that results in expression of a truncated MECP2 gene terminating in -Pro-Pro-Stop, wherein the construct is capable of editing the stop codon and allowing translational readthrough.

[0005] As described above, the inventors observed from sequence information obtained from subjects with Rett syndrome that a specific C-terminal deletion, approximately from about 1110 to 1210 (numbered based on the e2 isoform), results in a +2 frameshift and the production of a truncated MECP2 protein terminating in -Pro-Pro-Stop (PPX). The encoding nucleic acid sequence is CCCCCCTGA. To avoid translation termination at the TGA stop codon, the present invention describes editing the adenine base in the TGA stop codon to another base, e.g., guanine, optionally cytosine or thymine, to yield a codon other than the stop codon. For example, changing the adenine base in the TGA stop codon to guanine results in a TGG (tryptophan) codon. As described in further detail herein, the inventors generated mice with a knock-in mutation that accurately models this A-to-G change in the CTD1 mouse model previously described in Guy et al. (2010). In contrast to CTD1 mice, this novel model does not display an RTT-like phenotype and exhibits a 100% survival rate up to one year, demonstrating that this approach may be used to develop treatments for this class of RTT mutations.

[0006] According to the present invention, one suitable base editing method involves using a single guide RNA (sgRNA) to target an adenine base editor (ABE) to the target adenine of the TGA stop codon. Due to the nature of the local DNA sequence, it may be appropriate to use an ABE that recognizes a non-canonical PAM sequence (see, e.g., Walton et al., 2020). Even more active PAM variant ABEs are described, for example, in Richter (2020), Walton (2020), Gaudelli (2020), and Chen (2022), the entire disclosures of which are incorporated herein by reference. The inventors observed that guide sequences are present in (almost) all CTD RTT- alleles. Therefore, treatment should be applicable to this entire class of mutations. The edited gene does not produce wild-type protein but encodes a truncated form that retains the function of native MeCP2 in the studies exemplified herein. The present disclosure offers several advantages in relation to other therapeutic strategies for treating RTT:

[0007] Avoiding gene dosage issues One of the main therapeutic strategies currently being explored for treating RTT is gene therapy, involving viral delivery of exogenous MECP2. However, a major challenge facing this form of gene therapy is that overexpression of MeCP2 leads to other neurological disorders, such as MECP2 duplication syndrome. Therefore, it is difficult to determine a therapeutic dose that avoids MeCP2 overexpression. By editing endogenous MECP2 using a base editor such as ABE, the gene remains under the control of its endogenous regulatory elements, thus avoiding any issues with gene dosage. Permanent Fix Other therapeutic approaches to RTT include RNA editing or the use of readthrough compounds for nonsense mutations. The limitation of these strategies is that they require repeated administration to maintain the corrected mRNA level. On the other hand, base editing techniques, such as using ABE, result in permanent gene correction and are proven therapeutic once a sufficient number of cells are modified.

[0008] Requires only a single DNA edit (does not rely on double-strand DNA breaks (DSBs), HDR, and codelivery of repair templates) Ideally, gene editing would simply revert the mutation back to WT. In theory, this could be achieved using homology-directed repair (HDR) of Cas9-induced double-strand breaks (DSBs) by supplying an exogenous DNA repair template. In this case, the strategy would be highly useful for ex vivo approaches, where edited cells can first be expanded in culture and then returned to the body, or when targeting a population of dividing cells where edited cells have a selective advantage over unedited cells. Unfortunately, this approach is not suitable for RTT because the target population is postmitotic neurons, where HDR levels are low and there is no means of selection. This strategy avoids this problem by relying solely on a single step of editing endogenous sequences rather than editing / inserting novel sequences. DSBs introduced as part of the HDR repair process pose a higher risk of undesired mutational events than base editing.

[0009] One sgRNA (or one set of sgRNAs) can be used for all CTD mutations RTT C-terminal deletions are a heterogeneous set of mutations, with many individual deletions occurring singly or in small numbers. This strategy targets problematic sequences that appear to be present across the entire class, thus making treatment applicable to approximately 10% of RTT patients, comparable to the most commonly detected missense mutations. Functional domains must not be affected Two important functional domains, MBD (methylated DNA binding domain; amino acids 78-162) and NID (NCoR1 / 2 interaction domain; amino acids 301-309), are unaffected by CTD mutations. ExAC and GnomAD data indicate that the C-terminal deletion-prone region is highly tolerant of missense mutations and in-frame deletions. The discovery and widespread implementation of the CRISPR / Cas system has dramatically expanded the genome engineering toolbox and revolutionized the future prospects of basic biological research and medicine. The recent development of an adenine base editor by fusing a deaminase domain to Cas9 enables guide RNA (gRNA)-targeted single-nucleotide deamination to convert A:T base pairs to G:C using the adenine base editor within a specific target window. Base editing has been widely demonstrated with high efficiency in various species, including human zygotes.

[0010] Various engineered base editors have been developed with improved DNA editing efficiency. See, for example, U.S. Patent Publication No. 2018 / 0073012, published March 15, 2018, U.S. Patent Publication No. 2017 / 0121693, published May 4, 2017, International Publication Nos. WO2017 / 070633, published April 27, 2017, and U.S. Patent Publication No. 2015 / 0166980, published June 18, 2015, U.S. Patent No. 9,840,699, issued December 12, 2017, and U.S. Patent No. 10,077,453, issued September 18, 2018, each of which is incorporated herein in its entirety. A base editor (BE) can be a fusion of a Cas ("CRISPR-associated") domain and a nucleobase (or "base")-modifying domain (e.g., a naturally occurring or evolved deaminase, e.g., an adenosine deaminase domain). In some cases, a base editor can also include a protein or domain that affects cellular DNA repair processes, increasing the efficiency and / or stability of the resulting single nucleotide changes. Base editors generally contain a catalytically impaired Cas9 domain fused to a nucleobase-modifying domain. The Cas9 domain directs the nucleobase-modifying domain to directly convert one base to another at a target site programmed by a guide RNA. To date, two classes of base editors have been developed: cytosine base editors (CBEs), which convert C*G to T·A, and adenine base editors (ABEs), which convert A·T to G*C. The present disclosure is directed to the use of ABEs.

[0011] ABEs are particularly useful for studying and correcting pathogenic alleles because, in principle, approximately half of pathogenic point mutations can be corrected by converting an A·T base pair to a G*C base pair (see, e.g., Gaudelli et al. 2017 and Richter et al. 2020, the entire disclosures of which are directed to readers skilled in the art and incorporated herein by reference). Many ABEs reported to date contain a single polypeptide chain containing a heterodimer of wild-type Escherichia coli (E. coli) TadA monomer (ecTadA or TadA), which plays a structural role in base editing, and a laboratory-evolved E. coli TadA monomer, TadA7.10 (also referred to herein as "TadA*"), which catalyzes deoxyadenosine deamination, along with the Cas9 (D10A) nickase. Wild-type E. coli TadA acts as a homodimer to deaminate adenosine located in the tRNA anticodon loop, generating inosine (I). Early ABE variants required heterodimeric TadA containing an N-terminal wild-type TadA monomer for maximal activity, but subsequent studies showed that later ABE variants with and without the wild-type TadA monomer were equally active.

[0012] Guide RNA-dependent off-target base editing has been reduced by strategies including engineering mutations in the Cas9 component of the base editor that increase DNA specificity, adding 5' guanosine nucleotides to the sgRNA, or delivering the base editor as a ribonucleoprotein complex (RNP). Guide RNA-independent off-target editing can result from Cas9-independent binding of the deaminase domain of the base editor to C or A bases. An ABE for use in accordance with the present disclosure can comprise a fusion protein comprising a nucleic acid DNA-binding protein (or napDNAbp) domain and an adenosine deaminase domain. The napDNAbp domain can comprise a Cas9 protein or a variant thereof, such as a Cas9 nickase or a Cas9 nickase with altered PAM specificity. The adenosine deaminase domain can comprise one or more adenosine deaminases. In certain teachings, the adenosine deaminase domain comprises a dimer of a first and a second adenosine deaminase. The dimer can be a heterodimer comprising a first adenosine deaminase that is different from the second adenosine deaminase. The first adenosine deaminase can be located N-terminal to the second adenosine deaminase. In various embodiments, the one or more adenosine deaminases are connected by a linker (e.g., a peptide linker).

[0013] Suitable ABEs may be able to preserve DNA editing efficiency, and in some embodiments, demonstrate improved DNA editing efficiency relative to existing adenine base editors, such as ABE7.10 (see, e.g., WO2020214842, the entire disclosure of which is incorporated herein by reference). WO2020214842 describes ABEs that exhibit reduced off-target editing effects while retaining high on-target editing efficiency, as well as more recent ABEs described in Richter (2020), Walton (2020), Gaudelli (2020), and Chen (2022). A suitable ABE may be compatible with a variety of Cas homologs, including small-sized, circularly permuted, and evolved Cas homologs. The present disclosure includes the use of compositions comprising an ABE, e.g., an ABE with reduced off-target effects, e.g., reduced RNA editing effects, e.g., a fusion protein comprising an nCas9 domain and an adenosine deaminase domain (e.g., a heterodimer of first and second adenosine deaminases), and one or more guide RNAs, e.g., single guide RNAs ("sgRNAs").

[0014] The present disclosure further teaches nucleic acid molecules encoding and / or expressing adenine base editors and their adenosine deaminase domains as described herein, as well as expression vectors or constructs for expressing the adenine base editors and gRNAs (e.g., sgRNAs) described herein, cells (e.g., host cells) comprising the nucleic acid molecules and expression vectors and one or more gRNAs (e.g., sgRNAs), and compositions for delivering and / or administering the nucleic acid molecules for expression in cells of RTT subjects having the described deletion in exon 4 of MECP2 resulting in a shift to frame (n+2). Nucleic acid sequences can be codon-optimized for expression in human cells using techniques well known to those of skill in the art. The present specification further teaches a complex comprising an adenine base editor described herein and a gRNA (e.g., sgRNA) bound to the Cas9 domain of a fusion protein. The guide RNA can be 50-150, e.g., 90-100, or 80-110 nucleotides in length and comprise a sequence of at least 10, at least 15, or at least 20 or 25 contiguous nucleotides that are complementary to the target nucleotide sequence. Typically, the sequence is no longer than 30 or 35 nucleotides in length.

[0015] The present disclosure further includes kits for expressing and / or transducing cells (e.g., host cells) with expression constructs encoding fusion proteins and gRNAs. There are further taught kits for administering the expressed fusion proteins and expressed gRNA (e.g., sgRNA) molecules to host cells. The present disclosure further teaches host cells that stably or transiently express the fusion proteins and gRNA (e.g., sgRNA) or complexes thereof. Because the present disclosure is directed to the treatment of RTT with specific C-terminal deletions that result in truncated MECP2 sequences terminating in a frameshift (n+2) and -Pro-Pro-Stop, the present teachings may further include first screening RTT subjects for such C-terminal frameshift (n+2) deletion mutations to identify subjects suitable for treatment according to the present invention. The skilled reader will be well aware of suitable screening techniques that can be used, including nucleic acid sequencing of the subject's MDCP2 gene sequence, PCR and other amplification techniques, hybridization techniques using suitable probes, etc.

[0016] Also taught are methods for editing a mutant MECP2 gene, e.g., a single nucleic acid base within a mutant MECP2 gene, using an adenine base editor described herein. Such methods include transducing (e.g., by transfection) into cells multiple complexes, each comprising a fusion protein (e.g., a fusion protein comprising a Cas9 nickase (nCas9) domain and an adenosine deaminase domain) and a gRNA (e.g., sgRNA) molecule. In certain embodiments, the methods include transfecting one or more nucleic acid constructs (e.g., a plasmid, phagemid, or viral vector) that each (or together) encode components of the fusion protein and gRNA (e.g., sgRNA) molecule complex. In other teachings, the methods disclosed herein can include introducing into cells complexes comprising the cloned fusion protein and gRNA (e.g., sgRNA) molecule that are expressed outside of the cells. In accordance with the present invention, delivery to neurons in the brain using a viral vector, such as an adeno-associated viral (AAV) vector (e.g., AAV9), is most suitable. However, a nucleic acid encoding a suitable base editor may be too large to fit into and be expressed by a single viral vector. Thus, in one teaching, as described herein, a base editor may be encoded in two or more separate portions, which can be reconstituted into a full-length molecule by protein splicing. This is facilitated by adding split intein sequences adjacent to the sequences to be spliced ​​(see, e.g., Chen (2020)).

[0017] In some embodiments, methods are provided for treating RTT due to a mutant MECP2 gene as described herein using the disclosed base editors. The methods described herein can include treating a subject having or at risk of developing RTT, comprising administering to the subject (in an effective amount) a fusion protein as described herein, a conjugate as described herein, a polynucleotide as described herein, a vector as described herein, or a pharmaceutical composition as described herein. DETAILED DESCRIPTION OF THE INVENTION

[0018] definition As used herein and in the claims, the singular forms "a," "an," and "the" include singular and plural references unless the context clearly dictates otherwise. Thus, for example, reference to "a substance" includes the singular substance and a plurality of such substances. As used herein, the term "adenosine deaminase domain" refers to a domain within a fusion protein that includes one or more adenosine deaminases. For example, the adenosine deaminase domain can include a heterodimer of a first adenosine deaminase and a second deaminase domain connected by a linker, or a single engineered adenosine deaminase domain. "Base editing" refers to genome editing techniques that involve converting a specific nucleic acid base to another at a targeted genomic locus. In certain embodiments, this can be achieved without the need for a double-stranded DNA break (DSB) or single-strand break (i.e., nicking). Other genome editing techniques, including CRISPR-based systems, begin with the introduction of a DSB at a locus of interest. Cellular DNA repair enzymes then repair the break, typically resulting in a random insertion or deletion of a base (indel) at the site of the DSB. However, when the desired goal is to introduce or correct a point mutation at a targeted locus rather than stochastically disrupt an entire gene, these genome editing techniques are unsuitable because they have low correction rates (e.g., typically 0.1%-5%) and the primary genome editing product is an indel. To increase the efficiency of gene correction without simultaneously introducing random indels, the CRISPR / Cas9 system has been modified to directly convert one DNA base to another without forming a DSB. See Komor, AC, et al, Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016), the entire disclosure of which is incorporated herein by reference.

[0019] "Base editing construct" refers to a system comprising a suitable enzyme and, optionally, a nucleic acid capable of binding to a target nucleic acid, for use in base editing. "Adenine base editors" (or "ABEs"). This type of editor converts A:T Watson-Crick nucleobase pairs to G:C Watson-Crick nucleobase pairs. Because the corresponding Watson-Crick paired bases are also exchanged as a result of the conversion, this category of base editors can also be called thymine base editors (or "TBEs"). The term "base editor" (or "BE"), as used herein, refers to a substance, including polypeptides, that can make modifications to bases (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA). In some embodiments, a base editor is capable of deaminating bases within a nucleic acid, e.g., a base within a DNA molecule. In the case of an adenine base editor, the base editor can deaminate adenine (A) in DNA. Such a base editor can include a nucleic acid-programmable DNA-binding protein (napDNAbp) fused to an adenosine deaminase. Some base editors include CRISPR-mediated fusion proteins utilized in the base editing methods described herein. In some embodiments, the base editor includes a nuclease-inactive Cas9 (dCas9) fused to a deaminase that binds to nucleic acids in a guide RNA-programmed manner via R-loop formation but does not cleave the nucleic acid. For example, the dCas9 domain of the fusion protein can include D10A and H840A mutations (which enable Cas9 to cleave only one strand of a nucleic acid duplex), as described in WO2017 / 070632, which is incorporated herein by reference in its entirety. Cas9, the DNA cleavage domain of Streptococcus pyogenes (S. pyogenes), includes two subdomains: an HNH nuclease subdomain and a RuvCl subdomain. The HNH subdomain cleaves the strand complementary to the gRNA (the "targeted strand," or the strand where editing or deamination occurs), while the RuvCl subdomain cleaves the non-complementary strand containing the PAM sequence (the "non-edited strand"). The RuvCl mutant D10A generates a nick in the targeted strand, while the HNH mutant H840A generates a nick in the non-edited strand (see Jinek et al., Science, 337:816-821 (2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)).

[0020] The term "base editor" encompasses CRISPR-mediated fusion proteins utilized in the multiplexed base editing methods described herein, as well as any base editors known or described in the art at the time of this filing or developed in the future. See Rees & Liu, Base editing: precision chemistry on the genome and transcriptome of living cells, Nat. Rev. Genet. 2018;19(12):770-788, and U.S. Patent Publication Nos. 2018 / 0073012, 2017 / 0121693, WO2017 / 070633, WO2015 / 0166980, WO2017 / 070633, WO2018 / 027078, WO2019 / 079347, WO2019 / 226593, U.S. Patent Publication Nos. Reference is made to Patent Publication No. 2015 / 0166980, U.S. Patent No. 10,077,453, International Publication No. WO2019 / 023680, International Publication No. WO2018 / 0176009, International Publication No. WO2020051360, International Publication No. WO2020102659, International Publication No. WO202086908, and International Publication No. WO2020214842, the disclosures of which are incorporated herein by reference in their entireties.

[0021] The term "Cas9" or "Cas9 nuclease" or "Cas9 domain" refers to CRISPR-associated protein 9 or a variant thereof, and encompasses any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any naturally occurring or engineered variant of Cas9. The term Cas9 therefore extends to compact Cas9 variants that have been developed, such as Nme2 variants (see Erdaki et al., 2019). The term Cas9 is not meant to be particularly limiting and may also be referred to as "Cas9 or a variant thereof." Exemplary Cas9 proteins are described herein and in the art. This disclosure is not limited with respect to the particular Cas9 used in the CRISPR-mediated fusion proteins utilized in this disclosure.

[0022] In some embodiments, proteins comprising Cas9 or a fragment thereof are referred to as "Cas9 variants." Cas9 variants share homology to Cas9 or a fragment thereof. Cas9 variants include functional fragments of Cas9. For example, Cas9 variants are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9. In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9.

[0023] In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9. As used herein, the term "dCas9" refers to nuclease-inactive or nuclease-dead Cas9 or a variant thereof, and includes any naturally occurring dCas9 from any organism, any naturally occurring dCas9 equivalent or functional fragment thereof, any dCas9 homolog, ortholog, or paralog from any organism, and any naturally occurring or engineered variant of dCas9. The term dCas9 is not meant to be particularly limiting and may also be referred to as "dCas9 or a variant thereof." Exemplary dCas9 proteins and methods for making dCas9 proteins are further described herein and / or in the art and are incorporated herein by reference. Any suitable mutation that inactivates both Cas9 endonucleases can be used to form dCas9, for example, the D10A and H840A mutations in the wild-type Streptococcus pyogenes Cas9 amino acid sequence or the D10A and N580A mutations in the wild-type Staphylococcus aureus Cas9 amino acid sequence.

[0024] As used herein, the term "nCas9" or "Cas9 nickase" refers to Cas9 or a variant thereof that cleaves or nicks only one strand of the target cleavage site, thereby introducing a nick into a double-stranded DNA molecule rather than creating a double-stranded break. This can be achieved by introducing appropriate mutations into wild-type Cas9 that inactivate one of Cas9's two endonuclease activities. Any suitable mutation that inactivates one Cas9 endonuclease activity while leaving the other intact, e.g., either the D10A or H840A mutation in the wild-type Streptococcus pyogenes Cas9 amino acid sequence, or the D10A mutation in the wild-type Staphylococcus aureus Cas9 amino acid sequence, is contemplated and can be used to form nCas9.

[0025] "CRISPR" refers to a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent fragments of a previous infection by a virus that has invaded a prokaryote. These fragments of DNA are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with an array of CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNAs, they effectively constitute the prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of the pre-crRNA requires a transcoded small RNA (tracrRNA), endogenous ribonuclease 3 (me), and the Cas9 protein. The tracrRNA acts as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then cleaves linear or circular nucleic acid targets that are complementary to the RNA with an endonuclease. Specifically, the target strand that is not complementary to the crRNA is first cleaved with an endonuclease and then trimmed 3'-5' with an exonuclease. In nature, DNA binding and cleavage typically require both a protein and an RNA. However, single-guide RNAs ("sgRNAs," or simply "gRNAs") can be engineered to incorporate both crRNA and tracrRNA embodiments into a single RNA species—the guide RNA. See, for example, Jinek M., et al., Science 337:816-821 (2012), the entire disclosure of which is incorporated herein by reference. Cas9 recognizes a short motif within the CRISPR repeat sequence (the PAM or protospacer adjacent motif) to help distinguish self from non-self.CRISPR biology and Cas9 nuclease sequence and structure are well known to those of skill in the art (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes," Ferretti JJ, et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E., et al., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Jinek M., et al., Science 337:816-821 (2012), the entire disclosures of each of which are incorporated herein by reference). Cas9 orthologs have been described in a variety of species, including, but not limited to, Streptococcus pyogenes, S. thermophiles, C. ulcerans, S. diphtheria, S. syrphidicola, P. intermedia, S. taiwanense, S. iniae, B. baltica, P. torquis, S. thermophilus, L. innocua, C. jejuni, and N. meningitidis.Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure, and include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire disclosure of which is incorporated herein by reference. Other relevant teachings of interest to the skilled reader, the entire disclosures of which are incorporated herein by reference, include Kleinstiver (2016), Slaymaker (2015), and Vakulskas (2018).

[0026] The term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In accordance with the present disclosure, the deaminase is an adenosine deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine to inosine in deoxyribonucleic acid (DNA) (thus converting the adenine base to a hypoxanthine base). The deaminases provided herein can be derived from any organism, such as bacteria. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase or deaminase domain does not occur in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase.

[0027] Adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be enzymes that convert adenosine (A) in DNA or RNA to inosine (I). Such adenosine deaminases can lead to the conversion of A:T to G:C base pairs. In some embodiments, the deaminase is a variant of a naturally occurring deaminase derived from an organism. In some embodiments, the deaminase does not occur in nature. For example, in some embodiments, the deaminase is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase.

[0028] In some embodiments, the adenosine deaminase is derived from a bacterium, such as E. coli, Staphylococcus aureus, S. typhimurium, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated E. coli TadA deaminase. For example, the truncated ecTadA may be missing one or more N-terminal amino acids relative to full-length ecTadA. In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the ecTadA deaminase does not include an N-terminal methionine. See U.S. Patent Publication No. 2018 / 0073012, the entire disclosure of which is incorporated herein by reference.

[0029] As used herein, the term "DNA-binding protein" or "DNA-binding protein domain" refers to any protein that localizes to and binds to a specific target DNA nucleotide sequence (e.g., a genomic locus). This term encompasses RNA-programmable proteins, which associate with (e.g., form a complex with) one or more nucleic acid molecules (i.e., including, for example, guide RNAs in the case of a Cas system) that direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., DNA sequence) that is complementary to the nucleic acid molecule(s) (or portion or region thereof) with which the protein is associated. Exemplary RNA programmable proteins include CRISPR-Cas9 proteins, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), as well as Cas9 equivalents, homologs, orthologs, or paralogs, including Casl2a (type V CRISPR-Cas system) (formerly known as Cpfl), C2cl (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR Cas9 equivalents can include Cas9 equivalents from any type of CRISPR system (e.g., Type II, V, VI), including C2c2 (C2c2-Cas system), GeoCas9, CjCas9, Casl2b, Casl2g, Casl2h, Casl2i, Casl3b, Casl3c, Casl3d, Casl4, Csn2, xCas9, SpCas9-NG, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, or Spy-macCas9. Additional Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the disclosure of which is incorporated herein by reference.

[0030] The term "DNA editing efficiency," as used herein, refers to the number or percentage of intended base pairs that are edited. For example, if a base editor edits 10% of the base pairs (e.g., in a cell or population of cells) that it is intended to target, the base editor can be described as 10% efficient. Some aspects of editing efficiency include the modification (e.g., deamination) of specific nucleotides within DNA without generating a large number or percentage of insertions or deletions (i.e., indels). Editing while generating less than 5% indels (as measured across the total target nucleotide substrates) is generally considered to be high editing efficiency. The generation of more than 20% indels is generally considered to be insufficient or low editing efficiency. Indel formation can be measured by techniques known in the art, including high-throughput screening of sequencing reads.

[0031] The term "off-target editing frequency," as used herein, refers to the number or percentage of unintended base pairs, e.g., DNA base pairs, that are edited. On-target and off-target editing frequencies can be measured by the methods and assays described herein, further taking into account techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves hybridization of a nucleic acid primer (e.g., a DNA primer) that has complementarity to a nucleic acid (e.g., DNA) region immediately upstream or downstream of a target sequence or off-target sequence of interest. Since the DNA target sequence and the Cas9-independent off-target sequence are known in advance in the methods disclosed herein, nucleic acid primers with sufficient complementarity to the target sequence of interest and the upstream or downstream region of the Cas9-independent off-target sequence can be designed using techniques known in the art, such as the PhusionU PCR kit (Fife Technologies), the Phusion HS II kit (Fife Technologies), and the Illumina MiSeq kit. Because many Cas9-dependent off-target sites share high sequence identity with the intended target site, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9-dependent off-target site can similarly be designed using techniques and kits known in the art. These kits utilize polymerase chain reaction (PCR) amplification to generate amplicons as intermediate products. Target and off-target sequences may include genomic loci, which further include protospacers and PAMs. Thus, the term "amplicon," as used herein, may refer to a nucleic acid molecule that constitutes a collection of genomic loci, protospacers, and PAMs. High-throughput sequencing techniques, as used herein, may further include Sanger sequencing and / or whole genome sequencing (WGS).

[0032] The terms "RNA editing activity," "RNA editing effect," and "RNA off-target editing," as used herein, refer to the introduction of modifications (e.g., deamination) into nucleotides in cellular RNA, such as messenger RNA (mRNA). A key objective of DNA base editing efficiency is the modification (e.g., deamination) of a specific nucleotide in DNA without introducing a similar nucleotide modification in RNA. RNA editing effect is "low" or "reduced" if the detected mutations are introduced into RNA molecules at a frequency of 0.3% or less. For reference, the ABEmax base editor introduces edits into RNA at a frequency of approximately 0.50%. RNA editing effect is "low" or "reduced" if mutations are detected at a scale of less than approximately 70,000 edits in the analyzed mRNA transcriptome. The number of RNA edits can be measured by techniques known in the art, including high-throughput screening of sequencing reads and RNA-seq. The effect of RNA editing on the function of proteins translated from edited mRNA transcripts can be predicted using the SIFT ("Sorting Intolerant from Tolerant") algorithm, which is based on predictions regarding amino acid sequence homology and physical properties.

[0033] The term "on-target editing," as used herein, refers to the introduction of an intended modification (e.g., deamination) into a nucleotide (e.g., adenine) in a target sequence using a base editor, such as those described herein. The term "off-target DNA editing," as used herein, refers to the introduction of an unintended modification (e.g., deamination) into a nucleotide (e.g., adenine) in a sequence outside the standard base editor binding window (i.e., from one protospacer position to another, typically 2-8 nucleotides in length). Off-target DNA editing can result from weak or nonspecific binding of the gRNA sequence to the target sequence. Exemplary teachings describing methods for reducing off-target editing, the entire disclosures of which are incorporated herein by reference, can be found in Grunwald (2019) and Rees (2019). The term "effective amount," as used herein, refers to an amount of a biologically active agent that is sufficient to elicit a desired biological response. For example, in some embodiments, an effective amount of a composition may refer to an amount of the composition that is sufficient to edit a nucleotide sequence, e.g., a target site in a genome. In some embodiments, an effective amount of a composition provided herein, e.g., a composition comprising a nuclease-inactive Cas9 domain, a deaminase domain, a gRNA, and optionally a growth factor and an anti-apoptotic factor, may refer to an amount of the composition that is sufficient to induce editing of a target site that is specifically bound and edited by the fusion protein. In some embodiments, an effective amount of a composition provided herein may refer to an amount of the composition that is sufficient to induce editing having the following characteristics: >50% product purity, <5% indels, and an editing window of 2-8 nucleotides. As will be understood by those of skill in the art, an effective amount of a substance, e.g., a composition or fusion protein-gRNA complex, may vary depending on various factors, such as the desired biological response, e.g., the particular allele, genome, or target site to be edited, the cell or tissue being treated, and the substance being used.

[0034] The term "evolved base editor" or "evolved base editor variant" refers to a base editor formed as a result of mutagenizing a reference or starting base editor. This term refers to embodiments in which a nucleic acid base-modifying domain is evolved or a separate domain is evolved. Mutagenizing a reference or starting base editor can include mutagenizing an adenosine deaminase. The amino acid sequence variation can include one or more mutated residues within the amino acid sequence of the reference base editor, for example, as a result of a change in the nucleotide sequence encoding the base editor that results in a codon change at any particular position in the coding sequence, a deletion of one or more amino acids (e.g., a truncated protein), an insertion of one or more amino acids, or any combination of the foregoing. The evolved base editor can include variants in one or more components or domains of the base editor (e.g., variants introduced into one or more adenosine deaminases).

[0035] The term "fusion protein," as used herein, refers to a hybrid polypeptide containing protein domains from at least two proteins. One protein can be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) portion of the protein, forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. The proteins can contain different domains, such as a nucleic acid-binding domain (e.g., the gRNA-binding domain of Cas9, which directs binding of the protein to a target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid-editing protein. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire disclosure of which is incorporated herein by reference.

[0036] The term "host cell," as used herein, refers to a cell capable of hosting, replicating, and introducing a nucleic acid or vector as discussed herein. In embodiments where the vector is a viral vector, a suitable host cell is one that can be infected by the viral vector, replicate it, and package it into viral particles that can infect fresh host cells. A cell can host a viral vector if it supports the expression of the viral vector's genes, replication of the viral genome, and / or production of viral particles. One criterion for determining whether a cell is a suitable host cell for a given viral vector is whether the cell can support the viral life cycle of the wild-type viral genome from which the viral vector is derived. For example, as provided in some embodiments described herein, if the viral vector is a modified M13 phage genome, a suitable host cell would be any cell capable of supporting the wild-type M13 phage life cycle. Suitable host cells for viral vectors useful in continuous evolution processes are well known to those of skill in the art, and the present disclosure is not limited in this respect. In some embodiments, the viral vector is a phage and the host cell is a bacterial cell. In some embodiments, the host cell is an E. coli cell. Suitable E. coli host strains will be apparent to those skilled in the art, but include, but are not limited to, New England Biolabs (NEB) Turbo, ToplOF', DH12S, ER2738, ER2267, and XL1-Blue MRF'. These strain names are art-recognized, and the genotypes of these strains are characterized. "Fresh," as used herein synonymously with the terms "uninfected" or "uninfected" in the context of host cells, refers to host cells that have not been infected with a viral vector containing a gene of interest as used in the continuous evolution processes provided herein. However, fresh host cells may be infected with a viral vector unrelated to the vector to be evolved, or with a vector of the same or similar type but that does not carry the gene of interest.

[0037] In some embodiments, the host cell is a prokaryotic cell, such as a bacterial cell. In some embodiments, the host cell is an E. coli cell. In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, an insect cell, or a mammalian cell. The type of host cell will, of course, vary depending on the viral vector used, and suitable host cell / vector combinations will be readily apparent to one of skill in the art. Because Rett Syndrome is a neurological disorder that affects the way the brain functions, the host cell can be a neurological cell, such as a neuron, e.g., an excitatory or inhibitory neuron, or a glial cell, e.g., an astrocyte and microglia.

[0038] The term "linker," as used herein, refers to a chemical group or molecule that links two molecules or domains, e.g., dCas9 and a deaminase. Typically, a linker is positioned between or adjacent to two groups, molecules, or other domains and is connected to each via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical domain. Chemical groups include, but are not limited to, disulfide, hydrazone, and azide domains. In some embodiments, the linker is 5 to 100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the linker is an XTEN linker. As used herein, the term "low toxicity" refers to the maintenance of viability of more than 60% in a population of cells after application of a base editing method or administration of a composition disclosed herein. The term may also refer to the prevention of apoptosis (cell death) in a population of more than 40% of cells. For example, a genome editing method that results in less than 30% (e.g., 25%, 20%, 15%, 10%, or 5%) cell death indicates low toxicity. Cytotoxicity can be assessed by a suitable staining assay, such as annexin V and propidium iodide staining assay followed by flow cytometry (e.g., FACS).

[0039] The term "mutation," as used herein, refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue; the deletion or insertion of one or more residues in a sequence; or the substitution of a residue in a genomic sequence in a subject to be modified. Mutations are typically described herein by identifying the original residue, followed by the position of the residue in the sequence, and identifying the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). Mutations may include various categories, such as single nucleotide polymorphisms, microduplication regions, indels, and inversions, and are in no way intended to be limiting. Mutations may include "loss-of-function" mutations, which are the result of a mutation that reduces or abolishes protein activity. Most loss-of-function mutations are recessive because, in heterozygotes, the second chromosomal copy carries a fully functional, unmutated version of the gene coding for the protein, and its presence compensates for the effect of the mutation. There are some exceptions where loss-of-function mutations are dominant; one example is haploinsufficiency, in which the organism cannot tolerate the approximately 50% reduction in protein activity experienced by heterozygotes. This explains several genetic diseases in humans, including Marfan syndrome, which is caused by mutations in the gene for a connective tissue protein called fibrillin.

[0040] In the context of this disclosure, "translational readthrough" is the result of editing the third base of a TGA stop codon so that the resulting codon no longer encodes a stop codon and translation can continue until the next stop codon is encountered. The terms "non-naturally occurring" and "engineered" are used interchangeably and indicate the involvement of the hand of man. These terms, when referring to a nucleic acid molecule or polypeptide (e.g., a deaminase), mean that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated and / or found in nature (e.g., an amino acid sequence not found in nature).

[0041] The term "nucleic acid," as used herein, refers to RNA and single- and / or double-stranded DNA. Nucleic acids can occur naturally, e.g., in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. Alternatively, a nucleic acid molecule can be a non-naturally occurring molecule, e.g., recombinant DNA or RNA, artificial chromosome, engineered genome or fragment thereof, or synthetic DNA, RNA, DNA / RNA hybrid, or contain non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, optionally purified, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, e.g., analogs with chemically modified bases or sugars and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise specified. In some embodiments, nucleic acids are selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, inosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); intervening bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0042] The term "backbone," as used herein to modify a guide RNA molecule, refers to the component of the guide RNA, also known as the crRNA / tracrRNA, that includes the core region. The backbone is separate from the guide sequence, or spacer, the region of the guide RNA that has complementarity to the protospacer of the nucleic acid molecule.

[0043] The term "nucleic acid programmable DNA binding protein (napDNAbp)" refers to any protein that can associate with (e.g., form a complex with) one or more nucleic acid molecules (i.e., which may be broadly referred to as "napDNAbp programming nucleic acid molecules," including, e.g., guide RNAs in the case of Cas systems), which directs or otherwise programs the protein to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) that are associated with the protein, thereby causing the protein to bind to the nucleotide sequence at the specific target site. The term napDNAbp encompasses CRISPR-Cas9 proteins, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), as well as Cas9 equivalents, homologs, orthologs, or paralogs, including Casl2a (type V CRISPR-Cas system) (formerly known as Cpfl), C2cl (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), GeoCas9, CjCas9, Cass The napDNAbp can comprise a Cas9 equivalent (e.g., type II, V, VI) from any type of CRISPR system, including Cas9-12b, Casl2g, Casl2h, Casl2i, Casl3b, Casl3c, Casl3d, Casl4, Csn2, xCas9, SpCas9-NG, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, or Spy-macCas9. The napDNAbp can be a Cas9 domain containing a nuclease-active Cas9 domain, a nuclease-inactive Cas9 (dCas9) domain, or a Cas9 nickase (nCas9) domain. Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353 (6299), the disclosure of which is incorporated herein by reference.However, nucleic acid programmable DNA-binding proteins (napDNAbp) that can be used in connection with the present disclosure are not limited to CRISPR-Cas systems. The claimed invention encompasses any such programmable protein, such as the Argonaute protein (NgAgo) from Natronobacterium gregoryi, which can be used for DNA-guided genome editing. The NgAgo-guided DNA system does not require a PAM sequence or guide RNA molecule, meaning that genome editing can be performed simply by expressing a generic NgAgo protein and introducing synthetic oligonucleotides into any genomic sequence. See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7):768-73, incorporated herein by reference.

[0044] In some embodiments, napDNAbp is an RNA-programmable nuclease, and when complexed with RNA, it is sometimes referred to as a nuclease:RNA complex. Typically, the bound RNA is referred to as a guide RNA (gRNA). A gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule is sometimes referred to as a single guide RNA (sgRNA), although "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology to a target nucleic acid (e.g., directs binding of the Cas9 (or equivalent) complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is homologous to tracrRNA, as depicted in Figure IE of Jinek et al., Science 337:816-821 (2012), the entire disclosures of which are incorporated herein by reference. Other examples of gRNAs (e.g., those comprising domain 2) can be found in U.S. Patent No. 9,340,799, entitled "mRNA-Sensing Switchable gRNA," and WO 2015 / 035136, entitled "Delivery System For Functional Nucleases," the entire disclosures of each of which are incorporated herein by reference. In some embodiments, a gRNA comprises two or more domains (1) and (2), and may be referred to as an "extended gRNA." For example, an extended gRNA may bind, e.g., to two or more Cas9 proteins and bind to a target nucleic acid at two or more distinct regions, as described herein. The gRNA comprises a nucleotide sequence complementary to a target site, mediates binding of a nuclease / RNA complex to said target site, and provides sequence specificity for the nuclease:RNA complex.In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, e.g., Cas9 (Csnl) from Streptococcus pyogenes (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes," Ferretti JJ et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E. et al., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Jinek M. et al., Science 337:816-821 (2012), the entire disclosures of each of which are incorporated herein by reference.

[0045] napDNAbp nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, and these proteins can, in principle, be targeted to any sequence specified by a guide RNA. Methods of using napDNAbp nucleases, e.g., Cas9, for site-specific cleavage (e.g., to modify genomes) are known in the art (e.g., Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature Biotechnology 31, 227-229 (2013); Jinek, M. et al. RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013); Jiang, W. et al. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013), the entire disclosures of each of which are incorporated herein by reference).

[0046] The term "napDNAbp programming nucleic acid molecule" or, equivalently, "guide sequence" refers to one or more nucleic acid molecules that associate with a napDNAbp protein and direct or otherwise program it to localize to a specific target nucleotide sequence (e.g., a genomic locus) that is complementary to one or more nucleic acid molecules (or a portion or region thereof) associated with the protein, thereby causing the napDNAbp protein to bind to the nucleotide sequence at the specific target site. A non-limiting example is a guide RNA for the Cas protein of a CRISPR-Cas genome editing system. The term "promoter" is art-recognized and refers to a nucleic acid molecule having a sequence that can be recognized by the cellular transcription machinery and initiate transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is active only in the presence of certain conditions. For example, a conditional promoter can be active only in the presence of a specific protein that connects proteins associated with regulatory elements in the promoter to the basal transcription machinery, or only in the absence of an inhibitory molecule. A subclass of conditionally active promoters are inducible promoters, which require the presence of a small molecule "inducer" for activity.

[0047] Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. A variety of constitutive, conditional, and inducible promoters are known to those of skill in the art, and those skilled in the art will be able to identify a variety of such promoters useful in practicing the present disclosure, which is not limited in this respect. In various embodiments, the present disclosure provides vectors having a suitable promoter for driving expression of a nucleic acid sequence encoding a fusion protein (or one or more individual components thereof). The term "recombinant," as used herein in the context of a protein or nucleic acid, refers to a protein or nucleic acid that does not occur in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0048] The term "subject," as used herein, refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject may be of either gender and at any stage of development. The term "target site" or "target nucleic acid" refers to the sequence within a nucleic acid molecule that is edited by a fusion protein (e.g., a dCas9-deaminase fusion protein provided herein). Target site also refers to the sequence within a nucleic acid molecule to which a complex of a fusion protein and a gRNA binds.

[0049] The terms "treatment," "treat," and "treating" refer to a clinical intervention intended to reverse, alleviate, delay the onset of, or inhibit the progression of a disease, disorder, or condition as described herein, or one or more symptoms thereof. As used herein, the terms "treatment," "treat," and "treating" refer to a clinical intervention intended to reverse, alleviate, delay the onset of, or inhibit the progression of a disease, disorder, or condition as described herein, or one or more symptoms thereof. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., taking into account a history of the symptoms and / or taking into account genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay recurrence.

[0050] As used herein, the term "variant" refers to a protein having characteristics that deviate from those occurring in nature while retaining at least one of its functions, i.e., binding, interaction, or enzymatic ability, and / or therapeutic properties. A "variant" is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the wild-type protein. For example, a variant of Cas9 can include a Cas9 having one or more changes in amino acid residues compared to the wild-type Cas9 amino acid sequence. As another example, a variant of a deaminase can include a deaminase having one or more changes in amino acid residues compared to the wild-type deaminase amino acid sequence, e.g., after reconstruction of the deaminase's ancestral sequence. These changes include substitution of different amino acid residues, truncations, covalent additions (e.g., of tags) and chemical modifications, including any other mutations. The term also encompasses fragments of the wild-type protein. The level or degree to which properties are retained may be reduced compared to the wild-type protein, but typically will be the same or similar in type. Generally, variants will be highly similar overall in many regions and identical to the amino acid sequences of the proteins described herein. Those skilled in the art will understand how to make and use variants that retain all, or at least some, of the functional capabilities or properties.

[0051] A variant protein can comprise, or alternatively consist of, an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of a wild-type protein or any of the proteins provided herein (e.g., Cas9 proteins, fusion proteins, and fusion proteins). Additional polypeptides provided in the present disclosure are encoded by polynucleotides that hybridize to the complement of a nucleic acid molecule encoding a protein, such as a Cas9 protein, under stringent hybridization conditions (e.g., hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45 degrees Celsius, followed by one or more washes in 0.2× SSC, 0.1% SDS at about 50-65 degrees Celsius), under highly stringent conditions (e.g., hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45 degrees Celsius, followed by one or more washes in 0.1× SSC, 0.2% SDS at about 68 degrees Celsius), or other stringent hybridization conditions known to those of skill in the art (e.g., Ausubel, FM et al, eds., 1989 Current Protocols in Molecular Biology , Green Publishing associates, Inc., and John Wiley & Sons Inc., New York, at pp. 6.3.1-6.3.6 and 2.10.3). By a polypeptide having an amino acid sequence that is at least, e.g., 95% "identical" to a query amino acid sequence, it is intended that the amino acid sequence of the subject polypeptide is identical to the query sequence, except that the subject polypeptide sequence may contain up to five amino acid changes per 100 amino acids of the query amino acid sequence. In other words, to obtain a polypeptide having an amino acid sequence that is at least 95% identical to the query amino acid sequence, up to 5% of the amino acid residues in the subject sequence may be inserted, deleted, or substituted with another amino acid. These changes in the reference sequence may occur at the amino- or carboxy-terminal positions of the reference amino acid sequence, or anywhere between those terminal positions, either individually among residues in the reference sequence or scattered in one or more contiguous groups within the reference sequence. As a practical matter, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to the amino acid sequence of a protein, e.g., a Cas9 protein, can be conventionally determined using known computer programs.

[0052] A preferred method for determining the best overall match between a query sequence (a sequence of the present disclosure) and a subject sequence, also called a global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)). In a sequence alignment, the query and subject sequences are both nucleotide sequences or both amino acid sequences. The result of the global sequence alignment is expressed as percent identity. Preferred parameters used in FASTDB amino acid alignment are: matrix=PAM 0, k-tuple=2, mismatch penalty=1, joining penalty=20, randomization group length=0, cutoff score=1, window size=sequence length, gap penalty=5, gap size penalty=0.05, window size=500 or the length of the subject amino acid sequence, whichever is shorter. If the subject sequence is shorter than the query sequence because of N- or C-terminal deletions, and not because of internal deletions, the results must be manually corrected.

[0053] This is because the FASTDB program does not take into account N- and C-terminal truncations of the subject sequence when calculating the global percent identity. For subject sequences truncated at the N- and C-termini, the percent identity to the query sequence is corrected by calculating, as a percentage of the total bases in the query sequence, the number of residues in the query sequence that are N- and C-terminal to the subject sequence that are not matched / aligned with the corresponding subject residues. Whether a residue is matched / aligned is determined by the results of the FASTDB sequence alignment. This percentage is then subtracted from the percent identity calculated by the FASTDB program above using the specified parameters to arrive at a final percent identity score. This final percent identity score is what is used for purposes of the present disclosure. Only residues at the N- and C-termini of the subject sequence that are not matched / aligned with the query sequence are considered for purposes of manually adjusting the percent identity score; that is, only query residue positions outside the farthest N- and C-terminal residues of the subject sequence. As used herein, the term "wild-type" is a term of the art understood by those skilled in the art and means the normal form of an organism, strain, gene or characteristic as it occurs in nature as distinguished from mutant or variant forms. The invention will now be further described by way of example and with reference to the figures shown below: [Brief explanation of the drawings]

[0054] [Figure 1]Figure 1 shows that CTD1 and CTD2 knock-in mice display unexpectedly distinct phenotypes. The CTD1 and CTD2 mouse alleles were designed to model the two most common RTT CTD patient mutations. (a) Male mice hemizygous for the CTD1 mutation display the expected RTT-like phenotype similar to Mecp2 null animals, whereas CTD2 mice appear indistinguishable from wild-type littermates. (b) CTD1 mice exhibit reduced survival (median age of death 20 weeks). CTD2 mice have a 100% survival rate at 1 year, similar to their WT littermates. (c) CTD2 mice express a truncated MeCP2 at WT levels (total brain protein from 6-week-old mice), but CTD1 mice have reduced levels of MeCP2 protein in the brain. (d) Due to differences in amino acid sequence between mouse and human, the mouse CTD2 allele encodes -SPX at the C-terminus rather than -PPX, which is seen in patients with the equivalent mutation. Both the mouse and patient alleles for CTD1 end in -PPX; [Figure 2] Figure 2 shows that human neutral variants and patient frameshift mutations appear to overlap. Regions prone to human MECP2 deletions are shown with neutral variants at the top and high-confidence Red mutations at the bottom (black lines indicate deleted nucleotides). Numbering of DNA and amino acid sequences is that of the E2 isoform. Patient mutations, CTD1 and CTD2, modeled in mice are shown; [Figure 3] Figure 3 shows missense mutations and in-frame deletions in the CTD region. Regions prone to human MECP2 deletions are shown using neutral variants from the ExAC and GnomAD databases: in-frame deletion variants are shown on top, and missense mutations are shown below. Green lines indicate deleted nucleotides. Letters below indicate missense amino acid changes, with those seen in more than 10 individuals shown in purple and those seen in more than 50 individuals shown in bold purple. DNA and amino acid sequence numbering is that of the E2 isoform; [Figure 4] Figure 4 shows that the pathogenicity of deletions correlates with the translation frame. Regions prone to deletion in human MECP2 are shown using amino acid sequences where frameshift deletions may occur. The C-terminal amino acid sequence of MeCP2 from the Lett mutation as well as a neutral variant with a C-terminus deletion are shown, with frameshifted amino acids in bold or italic typeface according to their frame. The stop codon X is underlined; [Figure 5] Figure 5 shows that CTD1 NS knock-in mice replace the TGA stop with a TGG tryptophan codon. (a) Partial DNA sequence of the mouse Mecp2 CTD1 allele, with the stop codon shown in red; the black arrow indicates the adenine base to be edited. In the CTD1 NS mouse allele, this A is changed to G. The resulting C-terminal amino acid sequence is shown. (b) Reduced levels of MeCP2 in the CTD1 mouse brain are restored to approximately WT levels in CTD1 NS mice, in which the stop codon is replaced by tryptophan (assayed in total brain protein from 6-week-old mice). (c) Quantification of the Western blot in (b). (d) Male mice hemizygous for the CTD1 mutation exhibit an RTT-like phenotype similar to Mecp2 null animals, and CTD1 NS mice are indistinguishable from wild-type littermates; [Figure 6] Figure 6 shows a cell culture system for testing base editing reagents. (a) Overview of the Flp-In T-Rex system showing the Mecp2 cDNA transgene under the control of a tetracycline-inducible promoter. (b) A single-copy Mecp2 cDNA transgene recapitulates mutant MeCP2 protein and mRNA levels seen from the knock-in mouse allele. (c) Scheme of work for transient transfection experiments to test ABE / guide RNA combinations; [Figure 7]Figure 7 shows A-to-G editing in cultured human cells expressing the MeCP2 CTD1 transgene. (a) Nucleotide sequence of the region surrounding the target adenine (A) in the CTD1 Mecp2 transgene. The target A is shown as position 0, and bystander A's within the guide RNA sequence are shown as positions 6 and 9. The guide RNA sequence is shown below, with its optimal editing window indicated by a bar. Guides 1 and 2 are accessible by SpG and SpRY ABE8e, while guide 5 is accessible only by SpRY ABE8e. (b) Quantification of adenine base editing at the target and bystander sites in the CTD1 Mecp2 transgene after transfection with different combinations of ABE and guide RNA expression constructs. High-throughput sequencing of PCR amplicons from triplicate transfections of each plasmid combination is shown. (c) Western blot of proteins from the ABE / guide combinations shown in (b). Levels of truncated MeCP2 protein are increased after successful editing. The edited CTD1 MeCP2 protein flows above the unedited protein due to the extended C-terminal tail. Endogenous full-length human MeCP2 from T-Rex cells is also present. Cells were harvested on day 6, 24 hours after CTD1 Mecp2 transgene induction. (d) Quantification of the data in panel (c). Total truncated MeCP2 levels (edited + unedited CTD1 MeCP2) were normalized to the histone H3 loading control. Expression levels are normalized to the average empty guide value (g0), which was set to 1. Individual data points and average values ​​are shown; [Figure 8]Figure 8 shows A-to-G editing using reconstituted split ABE expression constructs. (a) Diagram of the full-length SpRY ABE8e expression construct showing the three split sites in the SpCas9 component. Examples of split pairs of expression constructs are shown below. These constructs contain split intein sequences from the DnaB gene from the thermophilic bacterium Rhodothermus marinus. Addition of these sequences to the two parts of ABE8e results in protein splicing when the two polypeptides are co-expressed in the same cell, thus reconstituting the full-length ABE. (b) Quantification of A-to-G base editing after transfection of the split ABE pair and guide RNA expression construct into CTD1 Mecp2 Flp-In T-Rex cells. Editing at the target A and two bystander A, as in Figure 7, is shown. SpG and SpRY ABE8e variants were tested using three different split sites.

[0055] Referring to Figure 1, panels a-c of the figure are taken from Guy et al. (2018). Hum. Mol. Genet., 27, 2531-2545. All information regarding experimental details, materials, and methods used to generate the data is included in that publication. This data indicates that the C-terminus of the protein is not required for MeCP2 function (CTD2 mice) and that CTD1 mutations are pathogenic due to greatly reduced levels of MeCP2 protein. The ExAC and GnomAD sequencing databases (containing exome and genome sequence data from individuals in the general population not known to suffer from severe neurological disorders) were searched for deletions in the C-terminal region of MECP2. These deletions (see Figure 2) are depicted above the nucleotide and amino acid sequences of the region. Information was extracted from ExAC and GnomAD v2.1.1 and v3 in October 2019, and from GnomAD in May 2023. Three additional frameshift deletions were found in GnomAD v3: 1167del7, 1167del34, and 1183del4.

[0056] Below the sequence (see Figure 2), deletions found in the RettBASE database of RTT patient mutations are shown. To exclude non-pathogenic mutations in MECP2 (there are several of these in RettBASE), only cases with classic Rett syndrome and in which no mutations were present in either parent were studied. Figure 3 shows in-frame deletions and missense-neutral variants extracted from the ExAC and GnomAD databases in October 2019. This data emphasizes that amino acids in this region of MeCP2 are not critical for protein function. Three possible reading frames for the region prone to C-terminal deletions are shown in Figure 4: the normal reading frame in standard typeface, +1 in italics, and +2 in bold. Below, the C-terminal amino acid sequences of the frameshift deletions in Figure 2 are shown. The correct amino acids before the frameshift are shown in standard typeface, with those after the shifted frame formatted accordingly. It is clear that the pathogenic mutation shifts to the +2 frame, while the neutral variant shifts to the +1 frame. It can be seen that the C-terminal amino acid sequence -PPX is common to all RTT mutations.

[0057] To model the effect of editing a TGA stop codon to TGG (tryptophan) using ABE base editing (shown in Figure 5(a)), a knock-in mouse allele was generated in mouse ES cells. WT JU09 ES cells were transfected with SpCas9 plasmid pX330, in which the guide sequence GACCTGAGCCTGAGAGCTCTG was cloned into the BbsI site. The first G of the guide sequence is not present in the genomic sequence and was added to support transcription of the guide using a U6 promoter. Sequence: AGCAGTGCCTCCTCCCCACCTAAGAAGGAGCACCATCATCACCACCATCACTCAGAGTCCCCAAAGGCGCCCGTGCCACTGCTCCCACCCCATCAGCCCCCCTG G A single-stranded DNA repair template containing the sequence GCCTCAGGACTTGAGCAGCAGCATCTGCAAAGAAGAGAAGATGCCCCGAGGAGGCTCACTGGAAAGCGATGGCTGCCCCAAG was co-transfected with the Cas9 plasmid. The A to G change from the sequence seen in CTD1 mice is bold and underlined. ES cell clones were screened for the correct modification of the Mecp2 gene, and correctly targeted clones were used to generate mouse lines by injection into mouse blastocysts. The resulting chimeric founders were bred to establish the CTD1 NS line.

[0058] MeCP2 protein levels in the brains of CTD1 NS hemizygous male mice were measured by Western blotting of whole brain extracts (Fig. 5(b) (method as described in Guy et al., 2018). The CTD1 NS MeCP2 protein is slightly larger than CTD1, as expected due to the extended tail of missense amino acids, and is expressed at or above the levels of its WT littermates (quantified in Fig. 5(c)). CTD1 NS hemizygous males did not exhibit the RTT-like phenotype seen in the CTD1 line (and other RTT mouse models) and were indistinguishable from WT littermates when scored weekly as previously described (Guy et al., 2018 and references therein; Figure 5(d)). All CTD1 NS animals survived to 1 year (median survival of CTD1 male hemizygotes, 20 weeks). This mouse model indicates that, for example, using adenine base editing to change a stop codon in a CTD allele to tryptophan is likely to prove therapeutic for this class of patients.

[0059] A cell line model system was used to test the ABE plasmid and guide RNA sequence for adenine editing efficiency. Cell lines harboring a tetracycline-inducible Mecp2 cDNA transgene (Figure 6(a)) were generated using the Flp-In™ T-Rex™ system (ThermoFisher). WT and mutant Mecp2 cDNAs (mouse e1 isoform) were cloned into the vector pcDNA5 / FRT / TO, and the resulting plasmids were transfected into the Flp-In™ T-Rex™ 293 cell line along with the Flp plasmid pOG44. After selection with hygromycin B for recombination of the transgene into the Flp-In locus, pools of cells were characterized and used for further experiments. The cell line was induced to express the transgene using tetracycline, and Western blotting revealed that the differences in MeCP2 protein levels seen in the mouse knock-in model were recapitulated by varying induction periods and tetracycline concentrations (as described in Guy et al., 2018) (Figure 6(b)). cDNA was generated from total RNA preps derived from the induced cells and quantified by real-time qPCR. The differences in mRNA levels seen in the knock-in mouse model were also recapitulated in this cell line (Figure 6(b)). Figure 6(c) shows a typical transfection experiment. Cells are transfected with the ABE and guide constructs using Lipofectamine 2000 transfection reagent. After 24 hours, the transfection medium is replaced with standard growth medium, and the cells are then grown for an additional 4–5 days, after which cell pellets are harvested by trypsinization for further analysis. For protein or mRNA analysis, cells are induced with tetracycline (typically 0.5 μg / ml) for 3–24 hours. It is noteworthy that in these experiments, we did not select for transfected cells, but the transfection efficiency, as estimated by the expression of the GFP marker from the ABE plasmid, was high at approximately 80%. The editing efficiency is therefore lower than that when only transfected cells are considered.Figure 7 shows the results from a conventional transfection experiment to determine the editing efficiency of different ABE / guide RNA combinations.

[0060] The ABE constructs tested were SpG-ABE8 and SpRY-ABE8, fusions formed between the deaminase and nCas9 domains. These constructs consist of the adenine deaminase domain (ABE8e) as described by Richter et al. (2020) combined with the nCas9 SpG and SpRY variants described by Walton et al. (2020). Methods for generating the fusions are described in Kluesner et al. (2021); Nishida, K. et al. (2016); Komor, AC, 2016, and Gaudelli, NM et al. (2017). A guide sequence with 20 nucleotides homologous to the mouse Mecp2 sequence surrounding the stop codon to be edited was cloned into the plasmid pGuide (obtained from Addgene, plasmid number 64711, depository Dr. Kiran Musunuru). The guide homology sequence was: SpG guide 1: CCCCTGAGCCTCAGGACTTG (AGC), A to be edited into guide position 7 SpG guide 2: CTGAGCCTCAGGACTTGAGC (AGC), A to be edited into guide position 4 SpRY guide 5: CCTGAGCCTCAGGACTTGAG (CAG), A to be edited into guide position 5.

[0061] Guides 1 and 2 can be used by SpG- and SpRY-ABE8e for their NGN PAM sequences (shown in brackets after the guide sequences). Guide 5 can only be used for NAN PAM with SpRY-ABE8e. SpG-Cas9 recognizes NGN PAM, while SpRY-ABE8e has looser PAM specificity, recognizing NRN where R=A or G. As a non-targeting control, a pGuide plasmid without a guide sequence inserted was used in combination with ABE, which was designated "empty guide" or "g0."

[0062] For genomic DNA analysis, the region surrounding the edited base was amplified by PCR, and the amplicons were either Sanger sequenced for bulk PCR or used for high-throughput sequencing using the Illumina Mi-Seq platform. For bulk sequencing, sequencing traces were analyzed using the web tool EditR (Kluesner et al., 2018). Western blotting of protein extracts from cell pellets collected after tetracycline induction was used to measure post-editing protein levels (Figure 7(c) and (d)). The total intensity of the CTD1 MeCP2 bands (edited and unedited) in each lane was normalized to the histone H3 loading control band. These values ​​were then further normalized using the average value of the blank guide sample, which was set to 1 (Figure 7(d)). Each lane / value represents an independently transfected well of cells.

[0063] Figures 7(b)-(d) show all combinations of ABE8e and guide plasmids. Five days after transfection, cells were induced with tetracycline for 4 hours, then each dish was trypsinized, divided into three aliquots, pelleted, snap-frozen, and used to prepare genomic DNA, protein, or total RNA. Each data point represents an independently transfected dish of cells. Analysis of genomic DNA for editing efficiency revealed that guides 2 and 5 also had significant A-to-G edits at bystander positions 6 and 9 within the guide sequence. These edits do not affect the amino acid sequence of the CTD1 allele, as they are silent mutations here, but they would create missense edits in the WT allele. While data on missense-neutral variants indicate that such changes are unlikely to be deleterious to MeCP2 function, we decided to proceed with guide 1 (and SpG-ABE8e) because this combination produced the best and cleanest edits. Guide 1 was predicted to have the lowest risk of bystander editing, as having an on-target edit at position 7 of the guide means that the two downstream As are further away from the editing window than with Guide 1 (on-target As at position 4) or Guide 5 (position 5).

[0064] We also split the ABE8e construct into two parts and fitted them into two AAV vectors (Figure 8). The use of a split intein sequence in the construct means that the two halves are joined by protein splicing when expressed in the same cell to generate a functional full-length protein. AAV is an optimal vector for delivery to the brain (in our case, initially in mice), but has smaller size limitations than a single ABE expression construct. The use of two AAV vectors and splitting the nucleic acid into two parts, as achieved above, is described, for example, in Chen (2020). A pair of split constructs made from both SpG and SpRY ABE8e was tested in combination with SpG guide 1 and compared with the equivalent full-length ABE8e. Editing efficiency using the split ABE8e was comparable to that achieved with the full-length construct. The split site between amino acids 573 and 574 of SpCas9 was chosen for further study due to the nearly equal size of the N- and C-terminal constructs. References JPEG2025530183000002.jpg176162

Claims

1. 1. A base editing construct for editing a mutant MECP2 gene, comprising a C-terminal deletion that results in a translational frameshift (n+2) and expression of a truncated MECP2 gene that terminates at -Pro-Pro-Stop, wherein the construct is capable of editing the stop codon and allowing translational readthrough.

2. 10. The base editing construct of claim 1 for use in a method of treating Rett syndrome in a subject, wherein the subject comprises a mutant MECP2 gene comprising a translational frameshift (n+2) and a C-terminal deletion that results in expression of a truncated MECP2 gene that terminates at -Pro-Pro-Stop, and the construct is capable of editing the stop codon and allowing translational readthrough.

3. 3. The base editing construct of claim 1 or 2, comprising a base editor that edits the adenine base in the TGA stop codon to another base, such as guanine, optionally to cytosine or thymine.

4. 4. The base editing construct of Claim 3, comprising a single guide RNA (sgRNA) that targets an adenine base editor (ABE) to the targeted adenine of the TGA stop codon.

5. The base editing construct of claim 4, wherein the ABE is ABE8 and derivatives thereof, including fusions comprising AB8e and SpG-ABE8 and SpRY-ABE8.

6. The base editing construct of claim 4 or 5, wherein the sgRNA is 80 to 150, for example 90 to 100 nucleotides in length and comprises a sequence of at least 10, at least 15, or at least 20 or 25 contiguous nucleotides that are complementary to a target nucleotide sequence comprising a nucleic acid encoding -Pro-Pro-Stop.

7. sgRNA having the sequence CCCCTGAGCCTCAGGACTTG (AGC); CTGAGCCTCAGGACTTGAGC (AGC), or CCTGAGCCTCAGGACTTGAG (CAG) The base editing construct of claim 4 or 5, comprising:

8. A nucleic acid construct encoding the sgRNA and ABE according to any one of claims 4 to 7.

9. 9. The nucleic acid construct of claim 8, wherein the sgRNA and the ABE are provided by separate constructs.

10. An expression vector comprising the nucleic acid construct of claim 8 or 9.

11. The expression vector of claim 10, which is a plasmid, a phagemid and / or a viral vector.

12. The expression vector of claim 11 , comprising an adeno-associated virus (AAV) vector.

13. 13. The expression vector of claim 12, wherein the nucleic acid encoding the ABE is split and expressed by two or more AAV vectors.

14. A kit for expressing and / or transducing into a cell, for example, a host cell, comprising a base-editing construct according to claims 1 to 7, a nucleic acid construct according to claims 8 to 9, or an expression vector according to claims 10 to 13.

15. A base-editing construct described in claims 1 to 7, a nucleic acid construct described in claims 8 to 9, an expression vector described in claims 10 to 13, or a kit described in claim 14, for use in treating Rett syndrome.

16. A host cell that stably or transiently expresses the base editing construct described in claims 1 to 7, the nucleic acid construct described in claims 8 to 9, or the expression vector described in claims 10 to 13.