Rett syndrome therapy

By editing the stop codon in the MECP2 gene using an adenine base editor, the problem of loss of MeCP2 protein and mRNA in Rett syndrome was solved, and the normal expression of MeCP2 protein and the reduction of disease symptoms were achieved.

CN120225674APending Publication Date: 2025-06-27THE UNIV COURT OF THE UNIV OF EDINBURGH +1
View PDF 19 Cites 0 Cited by

Patent Information

Application Number
CN202380077365.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-08
Filing Date
2023-09-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In Rett Syndrome, the translational codon shift and stop codon problems caused by the loss of the C-terminal deletion of the MECP2 gene lead to the loss of MeCP2 protein and mRNA, which in turn causes disease.

Method used

The TGG codon was generated by editing the adenine base of the TGA stop codon in the MECP2 gene into guanine using the adenine base editor (ABE), thereby preventing translation termination and ensuring the expression of the MeCP2 protein.

Benefits of technology

This method can effectively prevent the loss of MeCP2 protein, restore normal MeCP2 protein levels, and thus alleviate or eliminate the symptoms of Retel syndrome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120225674A_ABST
    Figure CN120225674A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a composition for the treatment of a class of Rett Syndrome mutations (i.e., C-terminal deletion) comprising a base editor for altering a termination codon in the mutant MECP2 gene. This alteration does not return the gene to its wild type (WT) form, but reestablishes a normal level form that is functionally equivalent to the wild type form of the MeCP2 protein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to compositions for treating a class of Rett syndrome mutations (i.e., C-terminal deletions), which comprise base editors for altering stop codons in the mutant MECP2 gene. This alteration does not return the gene to its wild-type (WT) form, but rather re-establishes a normal level of a form that is functionally equivalent to the wild-type form of the MeCP2 protein. Background Art

[0002] Rett Syndrome (RTT) is a severe neurological disorder caused primarily by mutations in the X-linked gene MECP2. Loss of MeCP2 function has the most profound effects in the nervous system and minimal phenotypic consequences in other tissues (Ross et al., 2016). Due to its association with intellectual disability in RTT and other disorders, the MECP2 gene is frequently screened for mutations in clinical cases of developmental delay. This has defined many amino acid alterations as relatively benign or neutral variants that either cause or do not cause RTT (Krishnaraj et al., 2017, and Karczewski et al., 2020).

[0003] By restoring MeCP2 expression, RTT-like symptoms in Mecp2-null mice can be rescued, indicating that the disorder may be curable. Most mutations that cause RTT disrupt two domains required for normal MeCP2 function; however, there are mutations that affect the C-terminal region of these domains. These mutations cause RTT by significantly reducing MeCP2 protein levels (Guy et al., 2018). Interestingly, mice lacking this C-terminal region express normal levels of the MeCP2 protein and do not have RTT-like symptoms, indicating that this region is not necessary for normal MeCP2 function. Summary of the Invention

[0004] A heterogeneous group of frameshift C-terminal deletions (CTDs) account for ~10% of mutations causing RTT. They occur within a deletion-prone region of MECP2, approximately c.1110-1210 (based on e2 isoform numbering) (Bebbington et al., 2010). The inventors found that the outcome of deletions in this region depends on the resulting reading frame. In-frame deletions remove a portion of DNA in this region but still maintain the WT reading frame and are present in large-scale sequencing databases that exclude individuals with severe pediatric diseases, indicating that they are non-pathogenic neutral variants (see Figure 3)。 The same is true for frameshift deletions that shift into frame (n+1) and end with the amino acid sequence -Ser-Pro-Arg-Thr-Stop. However, other frameshift deletions that cause RTT appear in RettBASE (a database listing mutations found in RTT patients), shifting into frame (n+2) and ending with -Pro-Pro-Stop (see Figure 2 ).

[0005] The present disclosure is based on the initial hypothesis of the inventors that for RTT subjects, there is a deletion in exon 4 of MECP2, which results in a shift into frame (n+2), and the presence of two prolines that disrupt translation termination before the stop codon, leading to the loss of MeCP2 protein and mRNA in a process similar to nonsense-mediated decay. Based on this, the inventors predicted that changing the stop codon after the two prolines to a codon encoding an amino acid, resulting in further translation to the next stop codon, could prevent the loss of MeCP2 protein, the loss of which is pathogenic in this class of mutations.

[0006] Thus, in a first aspect, the present disclosure provides a base editing construct for a method of treating Rett syndrome in a subject, wherein the subject comprises a mutant MECP2 gene, the mutant MECP2 gene comprises a C-terminal deletion, the C-terminal deletion results in a translational frameshift (n+2) and the expression of a truncated MECP2 gene ending with -Pro-Pro-Stop, and the construct is capable of editing the stop codon to allow translation through.

[0007] As described above, the inventors observed from sequence information obtained from subjects with Rett syndrome that certain C-terminal deletions, approximately c.1110-1210 (based on e2 isoform numbering), result in a +2 frameshift and the production of a truncated MECP2 protein ending with -Pro-Pro-Stop (PPX). The nucleic acid sequence encoding it is CCCCCCTGA. To avoid translation termination at the TGA stop codon, the present invention describes base editing the adenine base in the TGA stop codon to another base, such as guanine, or optionally cytosine, or thymine, to produce a codon other than a stop codon. For example, changing the adenine base in the TGA stop codon to guanine produces the TGG (tryptophan) codon. As described in further detail herein, the inventors have generated mice with a knock-in mutation that precisely models this A-to-G change in the CTD1 mouse model previously described in Guy et al. (2108). Compared to the CTD1 mice, this new model did not show an RTT-like phenotype and 100% survived up to one year, indicating that this method can be used to develop therapies for this class of RTT mutations.

[0008] According to the present invention, a suitable base editing method involves using a single-guide RNA (sgRNA) to target an adenine base editor (ABE) to the target adenine of the TGA stop codon. Due to the nature of the local DNA sequence, it may be suitable to use an ABE that recognizes a non-canonical PAM sequence (see, for example, Walton et al., 2020). More active PAM-variant ABEs are described, for example, in Richter (2020), Walton (2020), Gaudelli (2020), and Chen (2022), the entire contents of which are incorporated herein by reference. The inventors have observed that the guide sequence is present in (almost) all CTD RTT-alleles. Thus, the therapy should be applicable to this entire class of mutations. The edited gene will not produce the wild-type protein but will encode a truncation that retains the native MeCP2 function in the work exemplified herein.

[0009] The present disclosure provides a number of advantages related to other treatment strategies for treating RTT:

[0010] Avoid the problem of gene dosage

[0011] One of the main treatment strategies currently being explored for treating RTT is gene therapy, which involves the viral delivery of exogenous MECP2. However, the major challenge faced by this form of gene therapy is that overexpression of MeCP2 leads to other neurological disorders, such as MECP2 duplication syndrome. Thus, it is difficult to determine a dose that is therapeutic while avoiding MeCP2 overexpression. By editing endogenous MECP2 using a base editor (such as ABE), the gene remains under the control of its endogenous regulatory elements, thus avoiding any issues of gene dosage.

[0012] Permanent correction

[0013] Other treatment methods for RTT include RNA editing methods or the use of read-through compounds for nonsense mutations. The limitation of these strategies is that they will require repeated dosing in order to maintain the level of corrected mRNA. On the other hand, base editing technologies such as the use of ABE will result in the permanent correction of the gene and will prove therapeutic once a sufficient number of cells have been modified.

[0014] Only one DNA editing is required (independent of the co-delivery of double-stranded DNA break (DSB), HDR and repair template)

[0015] Ideally, gene editing would simply revert the mutation back to WT, which in theory could be accomplished by homology-directed repair (HDR) of Cas9-induced double-strand breaks (DSBs) by providing an exogenous DNA repair template. In this case, the strategy would be very useful for ex vivo approaches where the edited cells could first be expanded in culture before returning them to the body, or if a dividing cell population in which the edited cells have a selective advantage over unedited cells is being targeted. Unfortunately, this method is not applicable to RTT because the target population is post-mitotic neurons in which HDR levels are low and there is no means of selection. This strategy avoids this problem by relying on a single step that edits only the endogenous sequence rather than editing / inserting a new sequence. DSBs introduced as part of the HDR repair process have a higher risk of unwanted mutagenic events than base editing.

[0016] One sgRNA (or a set of sgRNAs) can be used for all CTD mutations

[0017] The RTT C-terminal deletions are a heterogeneous group of mutations, many of which are seen only once or a small number of times for individual deletions. This strategy targets problem sequences that are thought to be present across the class, making the therapy applicable to approximately 10% of RTT patients, comparable to the most frequently detected missense mutations.

[0018] The functional domain should not be affected

[0019] Two key functional domains, the MBD (methylated DNA binding domain; a.a. 78 - 162) and the NID (NCoR1 / 2 interaction domain; a.a. 301 - 309) are not affected by CTD mutations. ExAC and GnomAD data show that the C-terminal deletion-prone region is highly tolerant of missense mutations and in-frame deletions.

[0020] The discovery and widespread implementation of the CRISPR / Cas system have significantly expanded the genomic engineering toolbox and have revolutionized the future prospects of basic biological research and medicine. The recent development of adenine base editors by fusing a deaminase domain to Cas9 enables guide RNA (gRNA)-targeted single nucleotide deamination to convert A:T base pairs to G:C within a specific target window using adenine base editors. The high efficiency of base editing in a range of species, including human zygotes, has been widely demonstrated.

[0021] Various engineered base editors with improved DNA editing efficiency have been developed. For example, see U.S. Patent Publication No. 2018 / 0073012, published on March 15, 2018; U.S. Patent Publication No. 2017 / 0121693, published on May 4, 2017; International Publication No. WO2017 / 070633, published on April 27, 2017; and U.S. Patent Publication No. 2015 / 0166980, published on June 18, 2015; U.S. Patent No. 9,840,699, issued on December 12, 2017; and U.S. Patent No. 10,077,453, issued on September 18, 2018, each of which is incorporated herein by reference in its entirety. A base editor (BE) can be a fusion of a Cas (“CRISPR-associated”) domain and a nucleobase (or “base”) modification domain (e.g., a native or evolved deaminase, such as an adenosine deaminase domain). In some cases, the base editor may also include a protein or domain that affects the cellular DNA repair process to increase the efficiency and / or stability of the resulting single nucleotide change.

[0022] Base editors generally contain a catalytically impaired Cas9 domain fused to a nucleobase modification domain. The Cas9 domain directs the nucleobase modification domain to directly convert one base to another at a guide RNA-programmed target site. Two classes of base editors have been developed to date: cytosine base editors (CBEs) that convert C*G to T·A and adenine base editors (ABEs) that convert A·T to G*C. This disclosure relates to the use of ABEs.

[0023] ABE (see, for example, Gaudelli et al., 2017 and Richter et al., 2020, to which the person skilled in the art is directed and the entire contents of which are incorporated herein by reference) is particularly useful for the study and correction of pathogenic alleles because in principle almost half of the pathogenic point mutations can be corrected by converting A·T base pairs into G·C base pairs. Many ABEs reported to date include a single polypeptide chain and Cas9(D10A) nickase, and the single polypeptide chain contains a heterodimer of a wild-type Escherichia coli (E. coli) TadA monomer (ecTadA or TadA) that plays a structural role during base editing and a laboratory-evolved E. coli TadA monomer TadA7.10 (also referred to herein as "TadA*") that catalyzes deoxyadenosine deamination. The wild-type E. coli TadA acts as a homodimer to deaminate adenosine located in the tRNA anticodon loop, producing inosine (I). Although early ABE variants required a heterodimeric TadA containing an N-terminal wild-type TadA monomer for maximal activity, subsequent work showed that subsequent ABE variants had comparable activity with and without the wild-type TadA monomer.

[0024] Guide RNA-dependent off-target base editing has been reduced by strategies including installing mutations that increase DNA specificity into the Cas9 component of the base editor, adding a 5' guanosine nucleotide to the sgRNA, or delivering the base editor as a ribonucleoprotein (RNP) complex. Guide RNA-independent off-target editing can occur in a Cas9-independent manner from the binding of the deaminase domain of the base editor to a C or A base.

[0025] The ABEs used according to the present disclosure can comprise a fusion protein comprising a nucleic acid DNA-binding protein (or napDNAbp) domain and an adenosine deaminase domain. The napDNAbp domain can comprise a Cas9 protein or a variant thereof, such as a Cas9 nickase or a Cas9 nickase with altered PAM specificity. The adenosine deaminase domain can comprise one or more adenosine deaminases. In certain teachings, the adenosine deaminase domain comprises a dimer of a first and a second adenosine deaminase. The dimer can be a heterodimer that comprises a first adenosine deaminase that is different from the second adenosine deaminase. The first adenosine deaminase can be located at the N-terminus of the second adenosine deaminase. In different embodiments, the one or more adenosine deaminases are linked by a linker (e.g., a peptide linker).

[0026] Suitable ABEs may be able to maintain DNA editing efficiency and, in some embodiments, exhibit improved DNA editing efficiency relative to existing adenine base editors (e.g., ABE7.10), see, e.g., WO2020214842, the entire content of which is incorporated herein. WO2020214842 describes ABEs that exhibit reduced off-target editing effects while maintaining high on-target editing efficiency, as well as updated ABEs described in Richter (2020), Walton (2020), Gaudelli (2020), and Chen (2022).

[0027] Suitable ABEs can be compatible with a variety of Cas homologs, including small-sized, circularly permuted, and evolved Cas homologs.

[0028] The present disclosure includes the use of a composition comprising an ABE (e.g., an ABE having reduced off-target effects (e.g., reduced RNA editing effects), e.g., a fusion protein comprising an nCas9 domain and an adenosine deaminase domain (e.g., a heterodimer of a first and a second adenosine deaminase)) and one or more guide RNAs (e.g., a single guide RNA (“sgRNA”)).

[0029] The present disclosure further teaches nucleic acid molecules encoding and / or expressing the adenine base editors and their adenosine deaminase domains as described herein, as well as expression vectors or constructs for expressing the adenine base editors and gRNAs (e.g., sgRNAs) as described herein, cells (e.g., host cells) comprising the nucleic acid molecules and expression vectors and one or more gRNAs (e.g., sgRNAs), and compositions for delivering and / or administering nucleic acid molecules for expression in cells of an RTT subject having the deletion in exon 4 of MECP2, which deletion results in a frame shift to frame (n+2). The nucleic acid sequences can be codon-optimized for expression in human cells using techniques well known to those skilled in the art.

[0030] The present specification further teaches a complex comprising the adenine base editor described herein and a gRNA (e.g., sgRNA) that binds to the Cas9 domain of the fusion protein. The length of the guide RNA can be 50-150, e.g., 90-100 or 80-110 nucleotides, and comprises a sequence of at least 10, at least 15, or at least 20 or 25 consecutive nucleotides that is complementary to the target nucleotide sequence. Generally, the length of the sequence does not exceed 30 or 35 nucleotides.

[0031] The present disclosure further includes kits for expressing and / or transducing cells (such as host cells) having an expression construct encoding a fusion protein and a gRNA. Kits for administering the expressed fusion protein and the expressed gRNA (such as sgRNA) molecules to host cells are further taught. The present disclosure further teaches host cells that stably or transiently express the fusion protein and the gRNA (such as sgRNA) or their complexes.

[0032] Since the present disclosure relates to the treatment of RTT, which is caused by a specific C-terminal deletion that results in a frameshift (n+2) and a truncated MECP2 sequence ending with -Pro-Pro-Stop, this teaching further includes first screening RTT subjects for such C-terminal frameshift (n+2) deletion mutations in order to identify subjects suitable for treatment according to the present invention. Suitable screening techniques that are well known to those skilled in the art include nucleic acid sequencing of the MDCP2 gene sequence of the subject, PCR and other amplification techniques, hybridization techniques using suitable probes, and the like.

[0033] Methods for editing the mutant MECP2 gene (such as a single nucleobase within the mutant MECP2 gene) as described herein are also taught. Such methods involve transducing (such as by transfection) cells with multiple complexes, each complex comprising a fusion protein (such as a fusion protein comprising a Cas9 nickase (nCas9) domain and an adenosine deaminase domain) and a gRNA (such as sgRNA) molecule. In certain embodiments, the method involves transfection of one or more nucleic acid constructs (such as plasmids, phagemids, or viral vectors), each of which encodes (or together encode) the components of a complex of a fusion protein and a gRNA (such as sgRNA) molecule. In other teachings, the methods disclosed herein can involve introducing into cells a complex comprising a fusion protein and a gRNA (such as sgRNA) molecule that has been expressed and cloned outside of these cells. According to the present invention, delivery to neurons in the brain using a viral vector such as an adeno-associated virus (AAV) vector (such as AAV9) is most suitable. However, the nucleic acid encoding a suitable base editor may be too large to fit into and be expressed by a single viral vector. Thus, in one teaching, as described herein, the base editor can be encoded in two or more separate parts, which can be reconstituted into a full-length molecule by protein splicing. This is facilitated by adding split intein sequences adjacent to the sequences to be joined (see, for example, Chen (2020)).

[0034] In some embodiments, methods are provided for treating Rett syndrome (RTT) using the disclosed base editors, where the RTT is caused by a mutant MECP2 gene as described herein. The methods described herein can include treating a subject having or at risk of developing RTT, which includes administering (in an effective amount) to the subject a fusion protein, a complex, a polynucleotide, a vector, or a pharmaceutical composition as described herein.

[0035] Definitions

[0036] As used herein and in the claims, the singular forms "a", "an", and "the" include singular and plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an agent" includes a single agent and a plurality of such agents.

[0037] As used herein, the term "adenosine deaminase domain" refers to a domain within a fusion protein that includes one or more adenosine deaminases. For example, the adenosine deaminase domain can include a heterodimer of a first adenosine deaminase and a second deaminase domain linked by a linker, or a single engineered adenosine deaminase domain.

[0038] "Base editing" refers to a genome editing technique that involves converting a specific nucleic acid base at a targeted genomic locus to another nucleic acid base. In certain embodiments, this can be achieved without the need for a double-stranded DNA break (DSB) or a single-strand break (i.e., a nick). Other genome editing techniques, including CRISPR-based systems, begin with the introduction of a DSB at the locus of interest. Subsequently, cellular DNA repair enzymes repair the break, typically resulting in random insertions or deletions (indels) of bases at the site of the DSB. However, when it is desired to introduce or correct a point mutation at a target locus rather than randomly disrupt an entire gene, these genome editing techniques are not suitable because the correction rate is low (e.g., typically 0.1% to 5%) and the major genome editing products are indels. To increase the efficiency of gene correction without introducing random indels simultaneously, the CRISPR / Cas9 system has been modified to directly convert one DNA base to another without DSB formation. See Komor, A.C. et al., Programmable editing of a target base in genomic DNA without double-bridged DNA cleavage. Nature 533, 420-424 (2016), which is incorporated herein by reference in its entirety.

[0039] "Base editing construct" refers to a system for base editing, which system comprises a suitable enzyme and optionally a nucleic acid capable of binding to a target nucleic acid.

[0040] "Adenine base editor" (or "ABE"). Such an editor converts an A:T Watson-Crick nucleobase pair into a G:C Watson-Crick nucleobase pair. Since the corresponding Watson-Crick paired bases are also swapped due to the conversion, base editors of this category can also be referred to as thymine base editors (or "TBE").

[0041] As used herein, the term "base editor" (or "BE") refers to a reagent comprising a polypeptide that is capable of modifying a base (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA). In some embodiments, the base editor is capable of deaminating a base within a nucleic acid (e.g., a base within a DNA molecule).

[0042] In the case of an adenine base editor, the base editor is capable of deaminating adenine (A) in DNA. Such a base editor can comprise a nucleic acid programmable DNA binding protein (napDNAbp) fused to an adenosine deaminase. Some base editors include CRISPR-mediated fusion proteins used in the base editing methods described herein. In some embodiments, the base editor comprises a nuclease-inactive Cas9 (dCas9) fused to a deaminase, which nuclease-inactive Cas9 binds to the nucleic acid in a guide RNA-programmed manner through the formation of an R-loop but does not cleave the nucleic acid. For example, the dCas9 domain of the fusion protein can comprise D10A and H840A mutations (which enable Cas9 to cleave only one strand of the nucleic acid duplex), as described in WO2017 / 070632 and the entire content of which is incorporated herein by reference. The DNA cleavage domain of Streptococcus pyogenes Cas9 comprises two subdomains, the HNH nuclease subdomain and the RuvCl subdomain. The HNH subdomain cleaves the strand complementary to the gRNA ("target strand", or the strand in which editing or deamination occurs), while the RuvCl subdomain cleaves the non-complementary strand containing the PAM sequence ("unedited strand"). The RuvCl mutant D10A creates a nick in the target strand, while the HNH mutant H840A creates a nick in the unedited strand (see Jinek et al., Science, 337:816-821 (2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)).

[0043] The term "base editor" encompasses CRISPR-mediated fusion proteins used in the multiplex base editing methods described herein and any base editor known or described in the art at the time of filing of the present application or developed in the future. See Rees & Liu, Base editing: precision chemistry on the genome and transcriptome of living cells, Nat. Rev. Genet. 2018; 19(12):770-788; and U.S. Patent Publication No. 2018 / 0073012; U.S. Patent Publication No. 2017 / 0121693; International Publication No. WO2017 / 070633; U.S. Patent Publication No. 2015 / 0166980; International Publication No. WO2017 / 070633; International Publication No. WO2018 / 027078; International Publication No. WO2019 / 079347; International Publication No. WO2019 / 226593; U.S. Patent Publication No. 2015 / 0166980; U.S. Patent No. 10,077,453; International Publication No. WO2019 / 023680; International Publication No. WO2018 / 0176009; International Publication No. WO2020051360; International Publication No. WO2020102659; International Publication No. WO202086908; and International Publication No. WO2020214842, the contents of each of which are incorporated herein by reference in their entirety.

[0044] The term "Cas9" or "Cas9 nuclease" or "Cas9 domain" refers to CRISPR-associated protein 9 or variants thereof, and includes any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or fragment thereof, any Cas9 homolog, ortholog or paralog from any organism, and any variant of naturally occurring or engineered Cas9. Thus, the term Cas9 extends to compact Cas9 variants that have been developed, such as the Nme2 variant (see Erdaki et al., 2019).

[0045] The term Cas9 is not meant to be particularly limiting and may be referred to as "Cas9 or variants thereof". Exemplary Cas9 proteins are described herein and are also described in the art. The present disclosure is not limited to a particular Cas9 that is used in the CRISPR-mediated fusion proteins used in the present disclosure.

[0046] In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". A Cas9 variant is homologous to Cas9 or a fragment thereof. Cas9 variants include functional fragments of Cas9. For example, a Cas9 variant has at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity or at least about 99.9% identity with wild-type Cas9. In some embodiments, compared to wild-type Cas9, a Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes. In some embodiments, a Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA-binding domain or a DNA-cleavage domain) such that the fragment has at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity or at least about 99.9% identity with the corresponding fragment of wild-type Cas9.

[0047] In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% of the amino acid length of the corresponding wild-type Cas9.

[0048] As used herein, the term "dCas9" refers to nuclease-inactive Cas9 or dead nuclease Cas9 or variants thereof, and includes any naturally occurring dCas9 from any organism, any naturally occurring dCas9 equivalent or functional fragment thereof, any dCas9 homolog, ortholog or paralog from any organism, and any variant of naturally occurring or engineered dCas9. The term dCas9 is not meant to be particularly limiting and may be referred to as "dCas9 or variants thereof". Exemplary dCas9 proteins and methods for preparing dCas9 proteins are further described herein and / or described in the art and incorporated herein by reference. Any suitable mutations that inactivate both Cas9 endonucleases, such as the D10A and H840A mutations in the wild-type Streptococcus pyogenes Cas9 amino acid sequence, or the D10A and N580A mutations in the wild-type Staphylococcus aureus Cas9 amino acid sequence, can be used to form dCas9.

[0049] As used herein, the term "nCas9" or "Cas9 nickase" refers to Cas9 or variants thereof that only cleaves or only nicks one of the strands at the target cleavage site, thereby introducing a nick rather than a double-strand break in a double-stranded DNA molecule. This can be achieved by introducing appropriate mutations in wild-type Cas9 to inactivate one of the two endonuclease activities of Cas9. Any suitable mutation that inactivates one Cas9 endonuclease activity but leaves the other intact is contemplated, such as one of the D10A or H840A mutations in the wild-type Streptococcus pyogenes Cas9 amino acid sequence, or the D10A mutation in the wild-type Staphylococcus aureus Cas9 amino acid sequence, which can be used to form nCas9.

[0050] "CRISPR" is a family of DNA sequences in bacteria and archaea (i.e., CRISPR clusters), which represent snippets of excised fragments of previous infections by viruses that have invaded prokaryotes. The excised fragments of DNA are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with a series of CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNAs, they effectively constitute a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNAs (crRNAs). In some types of CRISPR systems (e.g., type II CRISPR systems), the correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), an endogenous ribonuclease 3 (rne), and the Cas9 protein. The tracrRNA serves as a guide for the ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA cleaves linear or circular nucleic acid targets complementary to the RNA in an endonucleotide manner. Specifically, the target strand that is not complementary to the crRNA is first cleaved in an endonucleotide manner and then trimmed in an exonuclease 3'-5' manner. In nature, DNA binding and cleavage generally require a protein and two RNAs. However, a single-guide RNA ("sgRNA" or simply "gRNA") can be engineered to incorporate the functions of both crRNA and tracrRNA into a single RNA species - the guide RNA. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes short motifs (PAM or protospacer adjacent motif) in the CRISPR repeat sequences to help distinguish self from non-self.CRISPR biology and Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an MI strain of Streptococcus pyogenes.” Ferretti J.J. et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA enucleogenate in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), which are incorporated herein by reference in their entireties). Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes, Streptococcus thermophiles, Corynebacterium ulcerans, Streptococcus diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Bellia baltica, Psychroflexus torquis, Streptococcus thermophilus, Listeria innocua, Campylobacter jejuni, and Neisseria meningitidis. Based on the present disclosure, additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art, and such Cas9 nucleases and sequences include the Cas9 sequences and loci from organisms disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-C as immunity systems” (2013) RNA Biology 10:5, 726-737, the entire content of which is incorporated herein by reference.Other relevant teachings that are instructive to those skilled in the art and all content incorporated herein by reference include Kleinstiver (2016), Slaymaker (2015), and Vakulskas (2018).

[0051] The term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. According to the present disclosure, the deaminase is adenosine deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA) to inosine (and thus the adenine base is converted to a hypoxanthine base). The deaminases provided herein can be from any organism, such as bacteria. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring deaminase.

[0052] The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be enzymes that convert adenosine (A) in DNA or RNA to inosine (I). Such adenosine deaminases can result in the conversion of A:T to G:C base pairs. In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism. In some embodiments, the deaminase does not exist in nature. For example, in some embodiments, the deaminase has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring deaminase.

[0053] In some embodiments, the adenosine deaminase is derived from bacteria such as Escherichia coli (E. coli), Staphylococcus aureus (S. aureus), Salmonella typhimurium (S. typhi), Shewanella putrefaciens (S. putrefaciens), Haemophilus influenzae (H. influenzae), or Caulobacter crescentus (C. crescentus). In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is an Escherichia coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated Escherichia coli TadA deaminase. For example, relative to full-length ecTadA, the truncated ecTadA may lack one or more N-terminal amino acids. In some embodiments, relative to full-length ecTadA, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues. In some embodiments, relative to full-length ecTadA, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues. In some embodiments, the ecTadA deaminase does not contain an N-terminal methionine. Reference is made to U.S. Patent Publication No. 2018 / 0073012, which is incorporated herein by reference in its entirety.

[0054] As used herein, the term "DNA-binding protein" or "DNA-binding protein domain" refers to any protein that localizes to and binds to a specific target DNA nucleotide sequence (e.g., a locus of the genome). The term includes RNA-programmable proteins that associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., which include, for example, guide RNAs in the case of the Cas system), where the nucleic acid molecules direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a DNA sequence) that is complementary to one or more nucleic acid molecules (or portions or regions thereof) that are associated with the protein. Exemplary RNA-programmable proteins are CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and may include Cas9 equivalents from any type of CRISPR system (e.g., type II, type V, type VI), including Cas12a (type V CRISPR-Cas system) (originally called Cpf1), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), GeoCas9, CjCas9, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, or Spy-macCas9. Other Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016;353(6299), the contents of which are incorporated herein by reference.

[0055] As used herein, the term "DNA editing efficiency" refers to the number or proportion of the intended base pairs that are edited. For example, if a base editor edits 10% of the base pairs it is designed to target (e.g., within a cell or within a population of cells), the base editor can be described as 10% efficient. Some aspects of editing efficiency include the modification of specific nucleotides within DNA (e.g., deamination) without generating a large number or large percentage of insertions or deletions (i.e., indels). It is generally accepted that when less than 5% indels are generated (as measured on the total target nucleotide substrate), the editing is of high editing efficiency. Generation of more than 20% indels is generally considered poor or low editing efficiency. Indel formation can be measured by techniques known in the art, including high-throughput screening of sequencing reads.

[0056] As used herein, the term "off-target editing frequency" refers to the number or proportion of unintended base pairs (e.g., edited DNA base pairs). On-target and off-target editing frequencies can be measured by the methods and assays described herein, further in view of techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves the hybridization of nucleic acid primers (e.g., DNA primers) to nucleic acid (e.g., DNA) regions that are complementary to regions only upstream or downstream of the target sequence or off-target sequence of interest. Since the DNA target sequence and Cas9-independent off-target sequences are a priori known in the methods disclosed herein, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the target sequence or Cas9-independent off-target sequences of interest can be designed using techniques known in the art, such as the PhusionU PCR Kit (Fife Technologies), the Phusion HS II Kit (Fife Technologies), and the Illumina MiSeq Kit. Since many Cas9-dependent off-target sites have high sequence identity with the target site of interest, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9-dependent off-target sites can also be designed using techniques and kits known in the art. These kits utilize polymerase chain reaction (PCR) amplification, which generates amplicons as intermediates. The target sequence and off-target sequences can comprise genomic loci, which further comprise protospacers and PAMs.

[0057] Thus, as used herein, the term "amplicon" can refer to a nucleic acid molecule that constitutes an aggregate of a genomic locus, a protospacer, and a PAM. The high-throughput sequencing techniques used herein can further include Sanger sequencing and / or whole genome sequencing (WGS).

[0058] As used herein, the terms "RNA editing activity", "RNA editing effect", and "RNA off-target editing" refer to the introduction of a modification (e.g., deamination) into a nucleotide within a cellular RNA (e.g., messenger RNA (mRNA)). An important goal of DNA base editing efficiency is the modification (e.g., deamination) of a specific nucleotide within DNA without introducing a modification of a similar nucleotide within RNA. When a detected mutation is introduced into an RNA molecule at a frequency of 0.3% or less, the RNA editing effect is "low" or "reduced". For reference, the ABEmax base editor introduces editing into RNA at a frequency of approximately 0.50%. When a mutation is detected at a magnitude of less than approximately 70,000 edits within the analyzed mRNA transcriptome, the RNA editing effect is "low" or "reduced". The number of RNA edits can be measured by techniques known in the art, including sequencing reads and high-throughput screening of RNA-seq. The effect of RNA editing on the function of a protein translated from an edited mRNA transcript can be predicted by using the SIFT ("Sorting Intolerant from Tolerant") algorithm, which is based on the prediction of amino acid sequence homology and physical properties.

[0059] As used herein, the term "on-target editing" refers to the introduction of an intended modification (e.g., deamination) into a nucleotide (e.g., adenine) within a target sequence using a base editor described herein. As used herein, the term "off-target DNA editing" refers to the introduction of an unintended modification (e.g., deamination) into a nucleotide (e.g., adenine) within a sequence outside of the canonical base editor binding window (i.e., from one protospacer position to another, typically 2 to 8 nucleotides in length). Off-target DNA editing can be produced by weak or non-specific binding of the gRNA sequence to the target sequence. Exemplary teachings describing ways to reduce off-target editing (the entire contents of which are incorporated herein by reference) can be found in Grunwald (2019) and Rees (2019).

[0060] As used herein, the term "effective amount" refers to the amount of a bioactive agent sufficient to elicit a desired biological response. For example, in some embodiments, an effective amount of a composition can refer to the amount of the composition sufficient to edit a target site of a nucleotide sequence (e.g., a genome). In some embodiments, an effective amount of a composition provided herein (e.g., a composition comprising a nuclease-inactive Cas9 domain, a deaminase domain, a gRNA, and optionally a growth factor and an anti-apoptotic factor) can refer to the amount of the composition sufficient to induce editing of a target site specifically bound and edited by the fusion protein. In some embodiments, an effective amount of a composition provided herein can refer to the amount of the composition sufficient to induce editing having the following characteristics: >50% product purity; <5% insertions / deletions; and an editing window of 2-8 nucleotides. As will be understood by those skilled in the art, the effective amount of an agent (e.g., a composition or a fusion protein-gRNA complex) can vary depending on different factors, e.g., depending on the desired biological response, e.g., depending on the particular allele, genome, or target site to be edited, depending on the cell or tissue being targeted, and depending on the agent being used.

[0061] The term "evolved base editor" or "evolved base editor variant" refers to a base editor formed as a result of mutagenesis treatment of a reference base editor or a starting-point base editor. The term refers to embodiments in which the nucleobase modification domain has evolved or individual domains have evolved. Mutagenesis treatment of a reference base editor or a starting-point base editor can include mutagenesis treatment of adenosine deaminase. Amino acid sequence variation can include one or more mutant residues within the amino acid sequence of the reference base editor, e.g., as a result of a change in the nucleotide sequence encoding the base editor, which results in a codon change at any particular position in the coding sequence, a deletion of one or more amino acids (e.g., a truncated protein), an insertion of one or more amino acids, or any combination of the foregoing. An evolved base editor can include variants in one or more components or domains of the base editor (e.g., variants of one or more adenosine deaminases introduced).

[0062] As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains from at least two proteins. One protein can be located in the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) portion of the fusion protein, thereby forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein", respectively. The proteins can contain different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the protein to bind to a target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. Any protein provided herein can be prepared by any method known in the art. For example, the proteins provided herein can be prepared by recombinant protein expression and purification, which is particularly applicable to fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire content of which is incorporated herein by reference.

[0063] As used herein, the term "host cell" refers to a cell that can host, replicate, and transfer a nucleic acid or vector as discussed herein. In embodiments where the vector is a viral vector, a suitable host cell is a cell that can be infected by the viral vector, can replicate the viral vector, and can package the viral vector into viral particles capable of infecting fresh host cells. A cell can host a viral vector if it supports the expression of the viral vector gene, the replication of the viral genome, and / or the production of viral particles. One criterion for determining whether a cell is a suitable host cell for a given viral vector is to determine whether the cell can support the viral life cycle of the wild-type viral genome from which the viral vector is derived. For example, if the viral vector is a modified M13 phage genome, as provided in some embodiments herein, then a suitable host cell will be any cell that can support the wild-type M13 phage life cycle. Suitable host cells for viral vectors that can be used in continuous evolution processes are well known to those of skill in the art, and the present disclosure is not limited in this regard. In some embodiments, the viral vector is a phage and the host cell is a bacterial cell. In some embodiments, the host cell is an Escherichia coli cell.

[0064] Suitable E. coli host strains will be apparent to those skilled in the art and include, but are not limited to, New England Biolabs (NEB) Turbo, Top10F', DH12S, ER2738, ER2267, and XL1-Blue MRF'. These strain names are well recognized, and the genotypes of these strains have been fully characterized. As used herein, the term "fresh" in the context of host cells may be interchangeable with the terms "non-infected" or "uninfected" and refers to host cells that have not been infected with a viral vector that contains the gene of interest used in the continuous evolution process provided herein. However, fresh host cells may have been infected with a viral vector that is unrelated to the vector to be evolved or with a vector of the same or similar type that does not carry the gene of interest.

[0065] In some embodiments, the host cell is a prokaryotic cell, such as a bacterial cell. In some embodiments, the host cell is an E. coli cell. In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, an insect cell, or a mammalian cell. Of course, the type of host cell will depend on the viral vector used, and suitable host cell / vector combinations will be apparent to those skilled in the art.

[0066] Since Rett syndrome is a neurological disorder that affects the way the brain functions, the host cell can be a nerve cell, such as a neuron, for example, an excitatory or inhibitory neuron, or a glial cell, such as an astrocyte and a microglia.

[0067] As used herein, the term "linker" refers to a chemical group or molecule that connects two molecules or domains (such as dCas9 and a deaminase). Generally, the linker is located between or on both sides of two groups, molecules, or other domains and is covalently linked to each group, molecule, or other domain to connect the two. In some embodiments, the linker is an amino acid or multiple amino acids (such as a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical domain. Chemical groups include, but are not limited to, disulfide, hydrazone, and azide domains. In some embodiments, the linker has a length of 5-100 amino acids, such as a length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids. Longer or shorter linkers are also contemplated. In some embodiments, the linker is an XTEN linker.

[0068] As used herein, the term "low toxicity" means that after applying the base editing methods disclosed herein or administering the compositions disclosed herein, a cell population maintains a viability of higher than 60%. The term can also refer to preventing apoptosis (cell death) in more than 40% of the cell population. For example, a genome editing method that results in less than 30% (such as 25%, 20%, 15%, 10%, or 5%) cell death exhibits low toxicity. Cytotoxicity can be evaluated by appropriate staining assays (such as annexin V and propidium iodide staining assays) and subsequent flow cytometry (such as FACS).

[0069] As used herein, the term "mutation" refers to the replacement of a residue within a sequence (such as a nucleic acid or amino acid sequence) with another residue; the deletion or insertion of one or more residues within the sequence; or the replacement of a residue within the genomic sequence of a subject to be corrected. Mutations are generally described herein by identifying the original residue, then the position of the residue within the sequence, and the identity of the newly replaced residue. The different methods for preparing amino acid replacements (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). Mutations can include multiple categories, such as single nucleotide polymorphisms, microduplication regions, insertions / deletions, and inversions, and are not meant to be limiting in any way. Mutations can include "loss-of-function" mutations, which are the result of mutations that reduce or eliminate protein activity. Most loss-of-function mutations are recessive because in heterozygotes, the second chromosome copy carries an unmutated version of the gene encoding a fully functional protein, the presence of which compensates for the effects of the mutation. There are some exceptions where loss-of-function mutations are dominant, one example being haploinsufficiency, where the organism cannot tolerate a ~50% reduction in protein activity suffered by the heterozygote. This is the explanation for several human genetic diseases, including Marfan syndrome, which is caused by a gene mutation in a connective tissue protein called fibrillin.

[0070] In the context of the present disclosure, "readthrough" is the result of editing the third base of a TGA stop codon such that the resulting codon no longer encodes a stop codon and translation can continue until the next stop codon appears.

[0071] The terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of the human hand. When referring to a nucleic acid molecule or a polypeptide (such as a deaminase), these terms mean that the nucleic acid molecule or polypeptide is at least substantially free of at least one other component that is naturally associated with them in nature and / or as found in nature (such as an amino acid sequence not found in nature).

[0072] As used herein, the term "nucleic acid" refers to RNA and single-stranded and / or double-stranded DNA. Nucleic acids, for example, in the context of a genome, can be naturally occurring transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cosmids, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or those comprising non-naturally occurring nucleotides or nucleosides. Additionally, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as those having analogs different from the phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified, chemically synthesized, etc. In appropriate cases, such as in the case of chemically synthesized molecules, nucleic acids can contain nucleoside analogs, such as those having chemically modified bases or sugars and backbone modifications. Unless otherwise specified, nucleic acid sequences are in the 5' to 3' direction. In some embodiments, the nucleic acid is or comprises natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, inosinedenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoramidite linkages).

[0073] As used herein to modify guide RNA molecules, the term "scaffold" refers to the component of the guide RNA that contains the core region, also known as the crRNA / tracrRNA. The scaffold is separate from the guide sequence or the spacer region (the region of the guide RNA that is complementary to the protospacer of the nucleic acid molecule).

[0074] The term "nucleic acid programmable DNA binding protein (napDNAbp)" refers to any protein that can associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., which can be broadly referred to as "napDNAbp-programming nucleic acid molecules" and include, for example, guide RNAs in the case of the Cas system), where the nucleic acid molecules direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a locus of a genome) that is complementary to one or more nucleic acid molecules (or portions or regions thereof) associated with the protein, such that the protein binds to the nucleotide sequence at the specific target site. The term napDNAbp includes CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and can include Cas9 equivalents from any type of CRISPR system (e.g., type II, V, VI), including Cas12a (type V CRISPR-Cas system) (previously called Cpf1), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), GeoCas9, CjCas9, Cas12b, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, or Spy-macCas9. A napDNAbp can be a Cas9 domain that includes a nuclease-active Cas9 domain, a nuclease-inactive Cas9 (dCas9) domain, or a Cas9 nickase (nCas9) domain. Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016;353(6299), the content of which is incorporated herein by reference. However, the nucleic acid programmable DNA binding proteins (napDNAbp) that can be used in connection with the present disclosure are not limited to CRISPR-Cas systems. The claimed invention includes any such programmable protein, such as the Argonaute protein from Natronobacterium gregoryi (NgAgo), which can also be used for DNA-guided genome editing.The NgAgo-guide DNA system does not require a PAM sequence or guide RNA molecules, meaning that genome editing can be simply carried out by the expression of the general NgAgo protein and the introduction of synthetic oligonucleotides at any genomic sequence. See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7):768-73, which is incorporated herein by reference.

[0075] In some embodiments, napDNAbp is an RNA-programmable nuclease that, when in complex with RNA, can be referred to as a nuclease:RNA complex. Generally, one or more bound RNAs are referred to as guide RNAs (gRNAs). The gRNA can exist as a complex of two or more RNAs, or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), although "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Generally, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and guides binding of the Cas9 (or equivalent) complex to the target); and (2) a domain that binds the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is homologous to tracrRNA as described in FIG. 1E of Jinek et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Patent No. 9,340,799, titled "mRNA-Sensing Switchable gRNAs" and WO 2015 / 035136, titled "Delivery System For Functional Nucleases", the entire content of each of which is incorporated herein by reference. In some embodiments, the gRNA contains two or more domains (1) and (2) and can be referred to as an "extended gRNA". For example, an extended gRNA will, for example, bind two or more Cas9 proteins and bind the target nucleic acid at two or more different regions, as described herein. The gRNA contains a nucleotide sequence complementary to the target site that mediates binding of the nuclease / RNA complex to the target site, providing sequence specificity of the nuclease:RNA complex.In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 (Csnl) from Streptococcus pyogenes (see, e.g., “Complete genome sequence of an MI strain of Streptococcus pyogenes.” Ferretti J.J. et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III”. Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), the entire content of each of which is incorporated herein by reference).

[0076] The napDNAbp nuclease (e.g., Cas9) uses RNA:DNA hybridization to target DNA cleavage sites and, in principle, these proteins can be targeted to any sequence specified by a guide RNA. Methods for site-specific cleavage (e.g., modifying the genome) using a napDNAbp nuclease such as Cas9 are known in the art (see, e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, W.Y. et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature Biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); DiCarlo, J.E. et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).

[0077] The term “napDNAbp-programmed nucleic acid molecule” or equivalently “guide sequence” refers to one or more nucleic acid molecules that associate with a napDNAbp protein and direct or otherwise program the napDNAbp protein to localize to a specific target nucleotide sequence (e.g., a locus of the genome) that is complementary to one or more nucleic acid molecules (or portions or regions thereof) that associate with the protein, thereby enabling the napDNAbp protein to bind to the nucleotide sequence at the specific target site. Non-limiting examples are the guide RNAs of the Cas proteins of the CRISPR-Cas genome editing system.

[0078] The term "promoter" is well recognized in the art and refers to a nucleic acid molecule having a sequence recognized by the cellular transcriptional machinery and capable of initiating transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular environment, or conditionally active, meaning that the promoter is active only in the presence of specific conditions. For example, a conditional promoter can be active only in the presence of a specific protein that links a protein associated with a regulatory element in the promoter to the basal transcriptional machinery, or only in the absence of an inhibitory molecule. A subclass of conditionally active promoters is inducible promoters, which require the presence of a small molecule "inducer" for activity.

[0079] Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. A variety of constitutive, conditional, and inducible promoters are well known to those skilled in the art, and those skilled in the art will be able to identify a variety of such promoters useful in practicing the present disclosure, which is not limited in this regard. In different embodiments, the present disclosure provides vectors having a suitable promoter for driving the expression of a nucleic acid sequence encoding a fusion protein (or one or more individual components thereof).

[0080] In the context of a protein or nucleic acid, the term "recombinant" as used herein refers to a protein or nucleic acid that is not found in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0081] As used herein, the term "subject" refers to an individual organism, such as an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, such as a genetically engineered non-human subject. The subject can be of either sex and at any stage of development.

[0082] The term "target site" or "target nucleic acid" refers to a sequence within a nucleic acid molecule that is edited by a fusion protein (e.g., the dCas9-deaminase fusion protein provided herein). The target site further refers to a sequence within a nucleic acid molecule that binds to a complex of a fusion protein and a gRNA.

[0083] The terms "treatment", "treat" and "treating" refer to a clinical intervention, as described herein, that is intended to reverse, alleviate, delay the onset of, or inhibit the progression of a disease, disorder or condition, or one or more symptoms thereof. As used herein, the terms "treatment", "treat" and "treating" refer to a clinical intervention, as described herein, that is intended to reverse, mitigate, delay the onset of, or inhibit the progression of a disease, disorder or condition, or one or more symptoms thereof. In some embodiments, treatment can be administered after one or more symptoms have developed and / or after the disease has been diagnosed. In other embodiments, treatment can be administered in the absence of symptoms, for example, to prevent or delay the onset of symptoms or to inhibit the onset or progression of the disease. For example, treatment can be administered to a susceptible individual prior to the onset of symptoms (e.g., in view of a symptom history and / or in view of genetic or other susceptibility factors). Treatment can also continue after symptoms have resolved, for example, to prevent or delay their recurrence.

[0084] As used herein, the term "variant" refers to a protein having properties that deviate from those of the naturally occurring protein, which retains at least one function, namely, binding, interaction or enzymatic ability and / or its therapeutic properties. A "variant" has at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to the wild-type protein. For example, a variant of Cas9 can comprise Cas9 having one or more amino acid residue changes compared to the wild-type Cas9 amino acid sequence. As another example, a variant of a deaminase can comprise a deaminase having one or more amino acid residue changes compared to the wild-type deaminase amino acid sequence, such as after reconstruction of the ancestral sequence of the deaminase. These changes include chemical modifications, including substitutions of different amino acid residues, truncations, covalent additions (such as tags), and any other mutations. The term also includes fragments of the wild-type protein.

[0085] The level or degree of the retained property may be reduced relative to the wild-type protein, but is generally the same or similar in kind. Generally, variants are overall very similar and are identical to the amino acid sequence of the proteins described herein in many regions. Those skilled in the art will recognize how to make and use variants that maintain all or at least some of the functional capabilities or properties.

[0086] Variant proteins can comprise or alternatively consist of: an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical to the amino acid sequence of, for example, a wild-type protein or any protein provided herein (such as a Cas9 protein, a fusion protein, and a fusion protein protein), with the proviso that the amino acid sequence of the fusion protein protein is not an amino acid sequence that is 100% identical to the amino acid sequence of a Cas9 protein. Additional polypeptides provided in the present disclosure are encoded by polynucleotides that hybridize under stringent hybridization conditions (such as hybridizing to filter-bound DNA in 6x sodium chloride / sodium citrate (SSC) at approximately 45 degrees Celsius, followed by washing once or more with 0.2x SSC, 0.1% SDS, at approximately 50 - 65 degrees Celsius), under highly stringent conditions (such as hybridizing to filter-bound DNA in 6x sodium chloride / sodium citrate (SSC) at approximately 45 degrees Celsius, then washing once or more with 0.1x SSC, 0.2% SDS, at approximately 68 degrees Celsius), or under other stringent hybridization conditions known to those of skill in the art (see, for example, Ausubel, F.M. et al., eds., 1989 Current Protocol in Molecular Biology, Green publishing associates, Inc., and John Wiley & Sons Inc., New York, pp. 6.3.1 - 6.3.6 and 2.10.3) to the complement of a nucleic acid molecule encoding a protein such as a Cas9 protein.

[0087] By a polypeptide having an amino acid sequence that is at least, for example, 95% "identical" to a query amino acid sequence, it is intended that the amino acid sequence of the subject polypeptide is identical to the query sequence, except that the subject polypeptide sequence can include up to five amino acid alterations per 100 amino acids of the query amino acid sequence. In other words, to obtain a polypeptide having an amino acid sequence that is at least 95% identical to a query amino acid sequence, up to 5% of the amino acid residues in the subject sequence can be inserted, deleted, or substituted with another amino acid. These alterations of the reference sequence can occur at the amino - or carboxyl - terminal positions of the reference amino acid sequence or at any position between those terminal positions, singly scattered among the residues of the reference sequence or in one or more contiguous groups within the reference sequence.

[0088] In fact, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to the amino acid sequence of, for example, a protein such as a Cas9 protein can be routinely determined using known computer programs.

[0089] A preferred method for determining the best overall match between a query sequence (a sequence of the present disclosure) and a subject sequence (also known as global sequence alignment) can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6: 237-245 (1990)). In sequence alignment, both the query sequence and the subject sequence are either nucleotide sequences or amino acid sequences. The result of the global sequence alignment is expressed as a percentage of identity. The preferred parameters used in FASTDB amino acid alignment are: matrix = PAM 0 (Matrix = PAM 0), k-tuple = 2 (k-tuple = 2), mismatch penalty = 1 (Mismatch Penalty = 1), joining penalty = 20 (Joining Penalty = 20), randomization group length = 0 (Randomization Group Length = 0), cutoff score = 1 (Cutoff Score = 1), window size = sequence length (Window Size = sequence length), gap penalty = 5 (Gap Penalty = 5), gap size penalty = 0.05 (GapSize Penalty = 0.05), window size = 500 (Window Size = 500) or the length of the subject amino acid sequence, whichever is shorter. If the subject sequence is shorter than the query sequence due to N- or C-terminal deletions, rather than internal deletions, the results must be manually corrected.

[0090] This is because the FASTDB program does not consider N- and C-terminal truncations of the subject sequence when calculating the global percentage identity. For a subject sequence truncated at the N- and C-termini relative to the query sequence, the percentage identity is corrected by calculating the number of residues of the query sequence that are at the N- and C-termini of the subject sequence and that do not match / align with the corresponding subject residues as a percentage of the total bases of the query sequence. Whether a residue matches / aligns is determined by the result of the FASTDB sequence alignment. This percentage is then subtracted from the percentage identity to obtain the final percentage identity score, which is calculated using the specified parameters by the above FASTDB program. This final percentage identity score is for the purposes of the present disclosure. For the purpose of manually adjusting the percentage identity score, only the N- and C-terminal residues of the subject sequence that do not match / align with the query sequence are considered. That is, only the residue positions outside the outermost N- and C-terminal residues of the subject sequence are queried.

[0091] As used herein, the term "wild-type" is a term of the art understood by those skilled in the art and refers to the typical form of a naturally occurring organism, strain, gene, or trait that is distinct from mutant or variant forms.

[0092] The present invention will now be further described by way of example and with reference to the accompanying drawings, which show: BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 . CTD1 and CTD2 knock-in mice unexpectedly showed different phenotypes.

[0094] The CTD1 and CTD2 mouse alleles were designed to model the two most common RTT CTD patient mutations. (a) Male mice that were heterozygous for the CTD1 mutation showed the expected RTT-like phenotype, similar to Mecp2 knockout (Mecp2-null) animals, while CTD2 mice were indistinguishable from wild-type littermates. (b) CTD1 mice showed reduced survival (median age at death of 20 weeks). CTD2 mice had 100% survival at one year, like WT littermates. (c) CTD2 mice expressed WT levels of truncated MeCP2 (whole brain protein from 6-week-old mice), while CTD1 mice had reduced MeCP2 protein levels in the brain. (d) Due to differences in amino acid sequence between mice and humans, the mouse CTD2 allele encodes C-terminal -SPX, rather than -PPX found in patients with equivalent mutations. The alleles of CTD1 in both mice and patients have -PPX termini;

[0095] Figure 2 . Human neutral variants and patient frameshift mutations appear to overlap

[0096] The human MECP2 deletion-prone region is shown with neutral variants above and high-confidence Rett mutations below (black lines indicate deleted nucleotides). DNA and amino acid sequences are numbered for the E2 isoform. The patient mutations CTD1 and CTD2 modeled in mice are indicated;

[0097] Figure 3 . Missense mutations and in-frame deletions in the CTD region

[0098] The human MECP2 deletion-prone region is shown with neutral variants from the ExAC and GnomAD databases: in-frame deletion variants are shown above and missense mutations below. Green lines indicate deleted nucleotides. Letters below indicate missense amino acid changes, those found in more than 10 individuals are shown in purple, and those found in more than 50 individuals are shown in bold purple. DNA and amino acid sequences are numbered for the E2 isoform;

[0099] Figure 4 . The pathogenicity of deletions is related to the reading frame

[0100] The region commonly deleted in human MECP2 shows a possible amino acid sequence caused by a frameshift deletion. The C-terminal amino acid sequences of MeCP2 from Rett mutations and neutral variants with C-terminal deletions are shown, with the amino acids of their frameshifts in bold or italic font. The STOP codon X is underlined;

[0101] Figure 5 . The CTD1 NS knock-in mice replace the TGA stop codon with the TGG tryptophan codon

[0102] (a) The partial DNA sequence of the mouse Mecp2 CTD1 allele, with the stop codon shown in red and the adenine base to be edited indicated by a black arrow. This A is changed to G in the CTD1 NS mouse allele. The resulting C-terminal amino acid sequence is shown. (b) The reduced level of MeCP2 in the brains of CTD1 mice is restored to approximately WT levels in CTD1 NS mice, where the stop codon is replaced by tryptophan (measured in whole brain protein from 6-week-old mice). (c) Quantification of the western blot in (b). (d) Hemizygous male mice with the CTD1 mutation show an RTT-like phenotype similar to Mecp2 knockout (Mecp2-null) animals, while CTD1 NS mice are indistinguishable from wild-type littermates;

[0103] Figure 6 . A cell culture system for testing base editing reagents

[0104] (a) Overview of the Flp-In T-Rex system, showing the Mecp2 cDNA transgene under the control of a tetracycline-inducible promoter. (b) The single-copy Mecp2 cDNA transgene replicates the levels of mutant MeCP2 protein and mRNA seen from the knock-in mouse allele. (c) Workflow of a transient transfection experiment to test ABE / gRNA combinations;

[0105] Figure 7 . A-to-G editing in cultured human cells expressing the MeCP2 CTD1 transgene

[0106] (a) Nucleotide sequence of the region around the targeted adenine (A) in the CTD1 Mecp2 transgene. The targeted A is denoted as position 0 and the bystander A within the guide RNA sequence is at 6 and 9. The guide RNA sequences are shown below, with the optimal editing window indicated by bars. Guides 1 and 2 can be used by SpG and SpRY ABE8es, and guide 5 can only be used by SpRY ABE8e. (b) Quantification of adenine base editing at the targeted and bystander sites in the CTD1 Mecp2 transgene after transfection with different combinations of ABE and guide RNA expression constructs. High-throughput sequencing of PCR amplicons from triplicate transfections of each plasmid combination is shown. (c) Western blot of proteins from the ABE / guide combinations shown in (b). After successful editing, the level of the truncated MeCP2 protein increases. Due to the extended C-terminal tail, the edited CTD1 MeCP2 protein runs above the unedited protein. Endogenous full-length human MeCP2 from T-Rex cells is also present. Cells were harvested on day 6, 24 hours after induction of the CTD1 Mecp2 transgene. (d) Quantification of the data in (c). The total truncated MeCP2 level (edited + unedited CTD1 MeCP2) was normalized to the histone H3 loading control. Expression levels were normalized to the mean empty guide value (g0) which was set to 1. Individual data points and the mean are shown;

[0107] Figure 8 . A-to-G editing with a reconstituted split ABE expression construct

[0108] (a) Diagram of the full-length SpRY ABE8e expression construct, which shows 3 split sites in the SpCas9 component. Examples of pairs of split expression constructs are shown below. These constructs include split intein sequences from the DnaB gene of the thermophilic bacterium Rhodothermus marinus. When the two polypeptides are co-expressed in the same cell, addition of these sequences to the two parts of ABE8e results in protein splicing, thus reconstituting full-length ABE. (b) Quantification of A-to-G base editing after transfection of split ABE pairs and guide RNA expression constructs into CTD1 Mecp2 Flp-In T-Rex cells. As Figure 7 shown, editing occurs at the targeted A and two bystander As. SpG and SpRY ABE8e variants were tested using 3 different split sites.

[0109] Reference Figure 1 , panels a - c are taken from Guy et al. (2018). Hum. Mol. Genet., 27, 2531 - 2545. All information regarding the experimental details, materials, and methods used to generate the data is included in that publication.

[0110] This data shows that the function of MeCP2 does not require the C-terminus of the protein (CTD2 mice), and the CTD1 mutation is pathogenic due to a substantial reduction in the level of MeCP2 protein.

[0111] For deletions in the C-terminal region of MECP2, the ExAC and GnomAD sequencing databases (including exome and genome sequence data from individuals in the general population known not to have severe neurological diseases) were searched. These deletions (see Figure 2 ) are depicted above the nucleotide and amino acid sequences of this region. Information was extracted from ExAC and GnomAD v2.1.1 and v3 in October 2019 and from GnomAD in May 2023, when 3 additional frameshift deletions were found in GnomAD v3, namely 1167del7, 1167del34, and 1183del4.

[0112] The sequences (see Figure 2 ) of deletions found in the RettBASE database of mutations in RTT patients are shown below. To exclude non-pathogenic mutations in MECP2 (many of these are present in RettBASE), only cases with classical Rett syndrome and mutations not present in either parent were studied. Figure 3 In-frame deletions and missense neutral variants extracted from the ExAC and GnomAD databases in October 2019 are shown. This data emphasizes that the amino acids in this region of MeCP2 are not important for the function of the protein.

[0113] Three possible reading frames of the C-terminal deletion-prone region are shown in Figure 4 : the normal reading frame of the standard type, italic +1, and bold +2. The C-terminal amino acid sequences of the frameshift deletions in Figure 2 are shown below. The correct amino acids before the frameshift are shown in standard type, and the amino acids after the frameshift are formatted according to the frame they are shifted into. Clearly, pathogenic mutations shift into the +2 frame, while neutral ones shift into the +1 frame. It can be seen that the C-terminal amino acid sequence -PPX is common to all RTT mutations.

[0114] To simulate the effect of editing the TGA stop codon to TGG (tryptophan) using the ABE base editor (as in Figure 5(as shown in (a)), a knocked-in mouse allele was generated in mouse ES cells. WT JU09 ES cells were transfected with the SpCas9 plasmid pX330, and the guide sequence GACCTGAGCCTGAGAGCTCTG was cloned into the BbsI site. The initial G of the guide sequence does not exist in the genomic sequence and was added to assist in transcription of the guide using the U6 promoter. Co-transfection with the Cas9 plasmid was performed with the sequence: AGCAGCAGTGCCTCCTCCCCACCTAAGAAGGAGCACCATCATCACCACCATCACTCAGAGTCCCCAAAGGCGCCCGTGCCACTGCTCCCACCCCATCAGCCCCCCTG G single-stranded DNA repair template of GCCTCAGGACTTGAGCAGCAGCATCTGCAAAGAAGAGAAGATGCCCCGAGGAGGCTCACTGGAAAGCGATGGCTGCCCCAAG. The A-to-G change in the sequence found in CTD1 mice is shown in bold and underlined. ES cell clones were screened for correct modification of the Mecp2 gene, and a mouse line was generated using the correctly targeted clones by injection into mouse blastocysts. The resulting chimeric founders were bred to establish the CTD1 NS line.

[0115] MeCP2 protein levels in the brains of CTD1 NS hemizygous male mice were measured by Western blot of whole brain extracts, Figure 5 (b) (by the method described in Guy et al., 2018). As expected, due to the extended tail of the missense amino acids, the CTD1 NS MeCP2 protein is slightly larger than CTD1 and is expressed at levels equal to or higher than those in WT littermate mice (quantified in Figure 5 (c)).

[0116] CTD1 NS hemizygous males did not display the RTT-like phenotype seen in the CTD1 line (and other RTT mouse models), and when scored weekly as previously described, Figure 5 (d) were indistinguishable from WT littermate mice (Guy et al., 2018 and references therein). All CTD1 NS animals survived to one year (median survival of CTD1 male hemizygotes was 20 weeks).

[0117] This mouse model shows that changing the stop codon in the CTD allele to tryptophan using, for example, adenine base editing may prove therapeutic for this class of patients.

[0118] The adenine editing efficiency of ABE plasmids and guide RNA sequences was tested using a cell line model system. Using Flp-In TM T-Rex TMThe system (ThermoFisher) generated cell lines with tetracycline-inducible Mecp2 cDNA transgenes. Figure 6 (a). WT and mutant Mecp2 cDNAs (mouse e1 isoform) were cloned into the vector pcDNA5 / FRT / TO, and the resulting plasmids were co-transfected with the Flp plasmid pOG44 into Flp-In TM T-Rex TM 293 cell lines. After selection for integration of the transgene into the Flp-In locus with hygromycin B, cell banks were characterized and used for further experiments. The cell lines were induced with tetracycline to express the transgene, and differences in MeCP2 protein levels observed in the mouse knock-in model (as described in Guy et al., 2018) were recapitulated across a range of induction periods and tetracycline concentrations by western blotting. Figure 6 (b). cDNA was prepared from total RNA preps of induced cells and quantified by real-time qPCR. Differences in mRNA levels observed in the knock-in mouse model were also recapitulated in this cell system. Figure 6 (b). Shown in Figure 6 (c) are typical transfection experiments. Cells were transfected with ABE and guide constructs using Lipofectamine 2000 transfection reagent. After 24 h, the transfection medium was replaced with standard growth medium, and then the cells were grown for an additional 4 - 5 days, after which the cell pellet was harvested for further analysis by trypsinization. For protein or mRNA analysis, cells were induced with tetracycline (usually 0.5 μg / ml) for 3 - 24 h. It should be noted that there was no selection of transfected cells in these experiments, although the transfection efficiency estimated by expression of the GFP marker from the ABE plasmid was high, approximately 80%. Thus, the editing efficiency was lower if only transfected cells were considered. Figure 7 Results of typical transfection experiments are shown to determine the editing efficiency of different ABE / guide RNA combinations.

[0119] The ABE constructs tested were fusions SpG-ABE8 and SpRY-ABE8 formed between the deaminase and nCas9 domains. These constructs consisted of adenine deaminase domains as described in Richter et al. (2020) (ABE8e) combined with nCas9 SpG and SpRY variants as described in Walton et al. (2020). Methods for preparing the fusions were described by Kluesner et al. (2021); Nishida, K. et al. (2016); Komor, A.C., 2016 and Gaudelli N.M. et al. (2017).

[0120] A guide sequence of 20 nucleotides homologous to the mouse Mecp2 sequence was cloned into plasmid pGuide (obtained from Addgene, plasmid #64711, depositor Dr Kiran Musunuru), and the mouse Mecp2 sequence was around the stop codon to be edited. The guide homologous sequences are:

[0121] SpG guide 1: CCCCTGAGCCTCAGGACTTG (AGC), A is edited at guide position 7

[0122] SpG guide 2: CTGAGCCTCAGGACTTGAGC (AGC), A is edited at guide position 4

[0123] SpRY guide 5: CCTGAGCCTCAGGACTTGAG (CAG), A is edited at guide position 5

[0124] Due to their NGN PAM sequences (shown in parentheses after the guide sequences), guides 1 and 2 can be used with SpG- and SpRY-ABE8e. Due to NAN PAM, guide 5 can only be used with SpRY-ABE8e. SpG-Cas9 recognizes NGN PAM, and SpRY-ABE8e has a more relaxed PAM specificity recognizing NRN, where R = A or G.

[0125] As a non-targeting control, the pGuide plasmid without a guide sequence inserted was used in combination with ABE. This is called "empty guide" or "g0".

[0126] For genomic DNA analysis, the region around the edited base was amplified by PCR, and either bulk PCR was sequenced by Sanger sequencing or the amplicons were used for high-throughput sequencing using the Illumina Mi-Seq platform. For bulk sequencing, the sequencing traces were analyzed using the web tool EditR (Kluesner et al., 2018).

[0127] The edited protein levels were measured using Western blotting of protein extracts from cell pellets harvested after tetracycline induction, Figure 7 (c), (d). The total intensity of the CTD1 MeCP2 bands (edited and unedited) in each lane was normalized to the histone H3 loading control band. Then these values were further normalized, with the average of the empty guide samples set to 1, Figure 7 (d). Each lane / value represents an independently transfected well of cells.

[0128] Figure 7(b)-(d) show all combinations of ABE8e and the guide plasmids. Five days after transfection, the cells were induced with tetracycline for 4 hours, then each dish was trypsinized, and the cells were split into three aliquots. The pelleted cells were snap-frozen and used for the preparation of genomic DNA, protein, or total RNA. Each data point represents an independently transfected cell dish. When analyzing the editing efficiency of genomic DNA, it was found that guides 2 and 5 also had significant A-to-G editing at the bystander positions 6 and 9 within the guide sequences. Although these edits do not affect the amino acid sequence of the CTD1 allele as they are silent mutations here, they would introduce missense edits to the WT allele. Although the data on missense neutral variants suggest that this change is unlikely to be harmful to MeCP2 function, it was decided to proceed with guide 1 (and SpG-ABE8e) as this combination gave the highest and cleanest edits. Guide 1 was expected to have the lowest risk of bystander editing as the on-target edit at position 7 of the guide means that the two downstream As are further away from the editing window than in guide 1 (on-target A position 4) or guide 5 (position 5).

[0129] The inventors also split the ABE8e construct into two parts in order to package them into two AAV vectors ( Figure 8 ). Using an intein sequence split within the construct means that when they are expressed in the same cell, the two halves are joined by protein splicing to produce a functional full-length protein. AAV is the vector of choice for delivery to the brain (initially in mice in our case), but has a size limit less than that of a single ABE expression construct. As achieved above, the use of two AAV vectors and the splitting of the nucleic acid into two parts are described, for example, in Chen (2020). Split constructs made from both SpG and SpRY ABE8e were tested in combination with SpG guide 1 and compared with an equivalent full-length ABE8. The editing efficiency of the split ABE8e pairs was comparable to that achieved by the full-length constructs. Because the N-terminal and C-terminal constructs were of roughly equal size, the split site between amino acids 573 and 574 of SpCas9 was selected for further work.

[0130] References

[0131] 1. Ross et al., Human Molecular Genetics, Vol. 25, No. 20, October 15, 2016, pp. 4389 - 4404

[0132] 2. Krishnaraj et al., Human Mutation, Vol. 38, No. 8, 2017)

[0133] 3. Karczewski et al., Nature 581, 434 - 443 (2020)

[0134] 4. Bebbington et al., Journal of Medical Genetics 2010; 47:242-248

[0135] 5. Walton et al., Science, Vol. 368, No. 6488, pp290-296 (2020)

[0136] 6. Gaudelli et al., Nature Biotechnology Vol. 38, pp892-900 (2020)

[0137] 7. Richter et al., Nature Biotechnology Vol. 38, pp883-891 (2020)

[0138] 8. Chen et al.; 2022, bioRxiv (doi: https: / / doi.org / 10.1101 / 2022.08.12.503700 )

[0139] 9. Kleinstiver, B.P. et al., Nature Vol. 529, pp490-495 (2016)

[0140] 10. Slaymaker et al.; Science, Vol. 351, No. 6268, pp84-88, 2015

[0141] 11. Vakulskas et al., Nature Medicine Vol. 24, pp1216-1224 (2018)

[0142] 12. Chen et al Small Methods, Volume 4, Issue 9, (2020)

[0143] 13. Grünewald et al., Nat Biotechnol 37, 1041-1048 (2019)

[0144] 14. Rees et al., Science Advancesm, Vol. 5, No. 5, (2019)

[0145] 15. Kluesner et al (2021) Nature Communications, Vol. 12, Article number: 2437 (2021)

[0146] 16. Nishida, K. et al (2016) Science 16; 353(6305)

[0147] 17. Komor, A. C. et al., Nature 533, 420 - 424 (2016).

[0148] 18. Gaudelli, N. M. et al., Nature 551, 464 - 471 (2017).

[0149] 19. Erdaki, A. et al., Mol. Cell 73, 714 - 726 (2019).

Claims

1. A base editing construct for editing a mutant MECP2 gene, wherein the mutant MECP2 gene contains a C-terminal deletion that results in a translational frameshift (n+2) and the expression of a truncated MECP2 gene ending with -Pro-Pro-Stop, and wherein the construct is capable of editing the stop codon to allow translational readthrough.

2. The base editing construct according to claim 1, for use in a method of treating Rett syndrome in a subject, wherein the subject contains a mutant MECP2 gene, the mutant MECP2 gene contains a C-terminal deletion that results in a translational frameshift (n+2) and the expression of a truncated MECP2 gene ending with -Pro-Pro-Stop, and wherein the construct is capable of editing the stop codon to allow translational readthrough.

3. The base editing construct according to claim 1 or 2, which comprises a base editor that edits the adenine base in the TGA stop codon to another base, such as guanine or optionally cytosine or thymine.

4. The base editing construct according to claim 3, wherein the construct comprises a single guide RNA (sgRNA) to target the adenine base editor (ABE) to the target adenine in the TGA stop codon.

5. The base editing construct according to claim 4, wherein the ABE is ABE8 and its derivatives, the derivatives include AB8e and fusions, the fusions include SpG-ABE8 and SpRY-ABE8.

6. The base editing construct according to claim 4 or 5, wherein the length of the sgRNA is 80-150, such as 90-100 nucleotides, and contains a sequence of at least 10, at least 15, or at least 20 or 25 consecutive nucleotides that is complementary to the target nucleotide sequence containing the nucleic acid encoding -Pro-Pro-Stop.

7. The base editing construct according to any one of claims 4 or 5, wherein the sgRNA contains the sequence CCCCTGAGCCTCAGGACTTG (AGC); CTGAGCCTCAGGACTTGAGC (AGC), or CCTGAGCCTCAGGACTTGAG(CAG).

8. One or more nucleic acid constructs encoding the sgRNA and ABE according to any one of claims 4-7.

9. The plurality of nucleic acid constructs according to claim 8, wherein the sgRNA and ABE are provided by a plurality of separate constructs.

10. One or more expression vectors comprising one or more nucleic acid constructs according to claim 8 or 9.

11. The one or more expression vectors according to claim 10, wherein the one or more expression vectors are plasmids, phagemids, and / or one or more viral vectors.

12. The one or more expression vectors according to claim 11, wherein the expression vector contains an adeno-associated virus (AAV) vector.

13. One or more expression vectors according to claim 12, wherein the nucleic acid encoding the ABE is split and expressed by two or more AAV vectors.

14. A kit for expressing and / or transducing cells such as host cells, the kit comprising a base editing construct according to claims 1-7, a nucleic acid construct according to claims 8-9, or one or more expression vectors according to claims 10-13.

15. A base editing construct according to claims 1-7, a nucleic acid construct according to claims 8-9, one or more expression vectors according to claims 10-13, or a kit according to claim 14, which is used for treating Rett syndrome.

16. A host cell that stably or transiently expresses a base editing construct according to claims 1-7, a nucleic acid construct according to claims 8-9, or one or more expression vectors according to claims 10-13.

Citation Information

Patent Citations

  • CAS9 proteins including ligand-dependent inteins

    US10077453B2

  • Fusions of CAS9 domains and nucleic acid-editing domains

    US20150166980A1

  • Nucleobase editors and uses thereof

    US20170121693A1

  • Adenosine nucleobase editors and uses thereof

    US20180073012A1

  • Card-index marker.

    US770788A