Methods for editing single nucleotide polymorphisms using programmable base editor systems

CN112469824BActive Publication Date: 2026-09-04BEAM THERAPEUTICS INC
View PDF 24 Cites 0 Cited by

Patent Information

Application Number
CN201980046538.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-13
Filing Date
2019-05-11
Publication Date
2026-09-04
Estimated Expiration
2039-05-11

AI Technical Summary

Technical Problem

然而,迄今为止,所述策略已经取得了有限的成功

Benefits of technology

[0015]In various embodiments, the contact is in cells, eukaryotic cells, mammalian cells, or human cells. In various embodiments, the cells are in vivo or ex vivo. In various embodiments, the alteration is one or more of R106W, R168*, R133C, T158M, R255*, R270*, and R306C. In various embodiments, the A·T to G·C alteration at the RTT-associated SNP changes cysteine ​​to arginine, methionine to threonine, or the stop codon to arginine in the methyl CpG-binding protein 2 (MeCP2) polypeptide. In various embodiments, the RTT-associated SNP leads to the expression of the Mecp2 polypeptide, which contains arginine at amino acid positions 168, 133, 255, 270, or 306, or threonine at position 158. In various embodiments, the polynucleotide programmable DNA-binding domain is Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In various embodiments, the polynucleotide programmable DNA-binding domain comprises a modified SpCas9 with altered protospacer adjacent motif (PAM) specificity. In various embodiments, the altered PAM is specific to the nucleic acid sequence 5'-NGT-3'. In various embodiments, the modified SpCas9 comprises one or more amino acid substitutions, or corresponding amino acid substitutions, selected from L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, T1337R and L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D2V2LD1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q and T1337M. In various embodiments, the modified SpCas9 comprises one or more amino acid substitutions, or corresponding amino acid substitutions, of D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337, as well as L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, DK2L, D1332L, D1332V, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M. In various embodiments, the polynucleotide programmable DNA-binding domain is a nuclease-inactivating or nickase variant.In various embodiments, the nicking enzyme variant contains an amino acid substitution of D10A or a corresponding amino acid substitution thereof. In various embodiments, the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid (DNA). In various embodiments, the adenosine deaminase is a modified adenosine deaminase that does not exist in nature. In various embodiments, the adenosine deaminase is a TadA deaminase. In various embodiments, the TadA deaminase is TadA*7.10. In various embodiments, one or more guide RNAs comprise CRISPR RNA (crRNA) and a trans-encoded small RNA (tracrRNA), wherein the crRNA contains a nucleic acid sequence complementary to the Mecp2 nucleic acid sequence, which contains an RTT-related SNP. In various embodiments, the base editor is complexed with a single-stranded guide RNA (sgRNA), wherein the single-stranded guide RNA contains a nucleic acid sequence complementary to the Mecp2 nucleic acid sequence, which contains an RTT-related SNP. In various embodiments, the cell is a neuron. In various embodiments, the neuron expresses the Mecp2 polypeptide. In various implementation schemes, the cells are derived from subjects suffering from RTT.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112469824B_ABST
    Figure CN112469824B_ABST
Patent Text Reader

Abstract

The present application features compositions and methods for altering mutations associated with Rett syndrome (RTT). The present application provides compositions and methods using base editors that include a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain that binds a guide polynucleotide. The present application also provides a base editor system for editing nucleobases of a target nucleotide sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority and interest in U.S. Provisional Application No. 62 / 670,588, filed May 11, 2018; U.S. Provisional Application No. 62 / 780,838, filed December 17, 2018; and U.S. Provisional Application No. 62 / 817,986, filed March 13, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to base editors and methods of using said base editors, particularly base editors comprising a polynucleotide programmable nucleotide-binding domain and a nucleobase editing domain that guides the editing of the polynucleotide. This application also provides a base editor system for editing the nucleosides of a target nucleotide sequence. Background Technology

[0004] Rett syndrome (RTT or RETT) is caused by a heterogeneous group of mutations in the methyl CpG-binding protein 2 (Mecp2) gene that impair or eliminate the ability of the encoded protein to modify chromatin and transcriptional states in the central nervous system (CNS). When widespread throughout the CNS, gene therapy delivering functional Mecp2 or repairing endogenous Mecp2 mRNA transcripts using RNA editing are promising therapeutic interventions. However, both approaches must overcome significant challenges to achieve therapeutic efficacy. Mecp2 gene therapy must strictly control the dose of gene delivered per cell, otherwise there is a risk of mimicking the Mecp2 replication syndrome phenotype. RNA editing platforms cannot accurately correct the most prevalent Mecp2 mutations, which account for over 45% of RTT diagnoses, and cannot induce effective, unguided off-target editing.

[0005] The gene mutations in Mecp2 that cause Rett syndrome (RTT) are highly heterogeneous. Therefore, a popular therapeutic strategy is to deliver wild-type Mecp2 carried by recombinant adeno-associated virus (rAAV). Because this strategy is independent of the causal mutation in each individual, successful gene therapy would provide a treatment option for most of the RTT patient population. However, this strategy has achieved limited success to date. RTT patients are almost always heterozygous females with characteristic wild-type and mutant X-chromosome MeCP2 mosaic expression in the central nervous system (CNS) due to random X-chromosome inactivation. Therefore, rAAV delivery in neurons that already express wild-type MeCP2 and the expression of wild-type MeCP2 may partially mimic the phenotype of MeCP2 replication syndrome. Consistent with this, the high transduction efficiency of the central nervous system in RTT model mice results in approximately 2-fold higher MeCP2 expression compared to wild-type mice.

[0006] Therefore, there is a need for novel compositions and methods for treating Rett syndrome.

[0007] References

[0008] All publications, patents, and patent applications mentioned in this application are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application were expressly and independently cited. Unless otherwise stated, all publications, patents, and patent applications mentioned in this application are incorporated herein by reference in their entirety. Summary of the Invention

[0009] As described herein, this application is characterized by a composition and method for precisely correcting pathogenic amino acids using a programmable nucleobase editor. In particular, the composition and method of this application can be used to treat Rett syndrome (RTT). Therefore, this application provides a method for precisely correcting single nucleotide polymorphisms in the endogenous Mecp2 gene using an adenosine (A) base editor (ABE) to correct harmful mutations (e.g., R133C, T158M, R255*, R270*, R306C).

[0010] On one hand, this application provides a method for editing a MECP2 polynucleotide containing a single nucleotide polymorphism (SNP) associated with Rett syndrome (RTT), the method comprising contacting the MECP2 polynucleotide with a base editor, the base editor being complexed with one or more guiding polynucleotides, wherein the base editor includes a polynucleotide programmable DNA-binding domain and an adenosine deaminase domain, and wherein one or more guiding polynucleotides target the base editor to cause an A·T to G·C change in the RTT-associated SNP.

[0011] On the other hand, this application provides cells generated by introducing a base editor or a polynucleotide encoding the base editor into a cell or its progenitor cells, wherein the base editor includes a polynucleotide programmable DNA-binding domain and an adenosine deaminase domain; and one or more guiding polynucleotides targeting the base editor to cause an A·T to G·C change in RTT-related SNPs.

[0012] On the other hand, this application provides a method for treating RTT in a subject, the method comprising administering to the subject: a base editor or a polynucleotide encoding the base editor, wherein the base editor includes a polynucleotide programmable DNA-binding domain and an adenosine deaminase domain; and one or more guiding polynucleotides targeting the base editor to cause an A·T to G·C change in SNPs associated with RTT.

[0013] On the other hand, this application provides a base editor comprising: (i) a modified SpCas9 comprising an amino acid substitution of one or more of L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, T1337R and L1111, D1135L, S1136R, G1218S, E1219V, D1332A, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F and T1337M, or a corresponding amino acid substitution thereof; and (ii) an adenosine deaminase.

[0014] On the other hand, this application provides a base editor comprising: (i) a modified SpCas9 comprising one or more of the following amino acid substitutions, or corresponding amino acid substitutions: D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337, and L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337s and T1337M; and (ii) an adenosine deaminase.

[0015] In various embodiments, the contact is in cells, eukaryotic cells, mammalian cells, or human cells. In various embodiments, the cells are in vivo or ex vivo. In various embodiments, the alteration is one or more of R106W, R168*, R133C, T158M, R255*, R270*, and R306C. In various embodiments, the A·T to G·C alteration at the RTT-associated SNP changes cysteine ​​to arginine, methionine to threonine, or the stop codon to arginine in the methyl CpG-binding protein 2 (MeCP2) polypeptide. In various embodiments, the RTT-associated SNP leads to the expression of the Mecp2 polypeptide, which contains arginine at amino acid positions 168, 133, 255, 270, or 306, or threonine at position 158. In various embodiments, the polynucleotide programmable DNA-binding domain is Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In various embodiments, the polynucleotide programmable DNA-binding domain comprises a modified SpCas9 with altered protospacer adjacent motif (PAM) specificity. In various embodiments, the altered PAM is specific to the nucleic acid sequence 5'-NGT-3'. In various embodiments, the modified SpCas9 comprises one or more amino acid substitutions, or corresponding amino acid substitutions, selected from L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, T1337R and L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D2V2LD1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q and T1337M. In various embodiments, the modified SpCas9 comprises one or more amino acid substitutions, or corresponding amino acid substitutions, of D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337, as well as L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, DK2L, D1332L, D1332V, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M. In various embodiments, the polynucleotide programmable DNA-binding domain is a nuclease-inactivating or nickase variant.In various embodiments, the nicking enzyme variant contains an amino acid substitution of D10A or a corresponding amino acid substitution thereof. In various embodiments, the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid (DNA). In various embodiments, the adenosine deaminase is a modified adenosine deaminase that does not exist in nature. In various embodiments, the adenosine deaminase is a TadA deaminase. In various embodiments, the TadA deaminase is TadA*7.10. In various embodiments, one or more guide RNAs comprise CRISPR RNA (crRNA) and a trans-encoded small RNA (tracrRNA), wherein the crRNA contains a nucleic acid sequence complementary to the Mecp2 nucleic acid sequence, which contains an RTT-related SNP. In various embodiments, the base editor is complexed with a single-stranded guide RNA (sgRNA), wherein the single-stranded guide RNA contains a nucleic acid sequence complementary to the Mecp2 nucleic acid sequence, which contains an RTT-related SNP. In various embodiments, the cell is a neuron. In various embodiments, the neuron expresses the Mecp2 polypeptide. In various implementation schemes, the cells are derived from subjects suffering from RTT. Attached Figure Description

[0016] The features of this disclosure are specifically set forth in the appended claims. The features and advantages of this application can be better understood by referring to the following detailed description, which illustrates illustrative embodiments in which the principles of this disclosure and the accompanying drawings are applied:

[0017] Figure 1 This is a graph depicting the percentage of accurate corrections to the R255X RTT mutation using a base editor variant specific to NGT PAM.

[0018] Figure 2 This is a graph depicting the percentage of accurate corrections to the R255X RTT mutation using a base editor variant specific to NGT PAM.

[0019] Figure 3 This is a graph depicting the optimization of PAM variants with the amino acid substitution T1337L. It shows the precise percentage correction for the R255X RTT mutation using a base editor variant specific to NGT PAM.

[0020] Figure 4 This is a graph depicting the percentage of R255X RTT mutations precisely corrected by PAM base editor variants specific to NGT PAM, which are generated by shuffling mutations from other characterizing PAM variants. T1337Q is considered important for improving editing efficiency.

[0021] Figure 5 This is a graph depicting the changes in base editing efficiency when T1337 and D1332 are substituted with other amino acids. It shows the precise percentage correction for the R255X RTT mutation using a base editor variant specific to NGT PAM.

[0022] Figure 6 This is a graph depicting the importance of E1219V and R1335Q for base editing activities related to T1337. It shows the precise percentage correction for the R255X RTT mutation using a base editor variant specific to NGT PAM.

[0023] Figure 7 This is a graph depicting the percentage of accurate corrections to the R255X RTT mutation using a PAM variant base editor specific to the NGT PAM variants listed in Tables 8 and 9.

[0024] Figure 8 This is a graph depicting the percentage of accurate corrections to the R255X RTT mutation using a PAM variant base editor that is specific to the NGT PAM variants listed in Tables 8 and 9.

[0025] Figure 9 This is a graph depicting the percentage of R255X RTT mutations accurately corrected using a PAM variant base editor specific to the NGT PAM variants listed in Table 10. Detailed Implementation

[0026] This document describes a composition and method for providing a base editing system to precisely correct one or more mutations in the methylCpG-binding protein 2 (Mecp2) gene, which is causally associated with Rett syndrome (RTT), a progressive neurodevelopmental disorder, and its symptoms. RTT is an X-linked dominant disorder that primarily affects females, and in 96% of affected individuals, it is associated with mutations in the Mecp2 gene. It is characterized by seemingly normal early development followed by regression with loss of fine motor skills and effective communication, stereotyped movements, apraxia, or complete absence of gait. Other clinical features in affected individuals include abnormal acquired deceleration of head growth rate, paroxysmal breathing, gastrointestinal dysfunction, epilepsy, and scoliosis.

[0027] The most common RTT-causing mutation is the cytidine-to-thymidine (C→T) transition mutation, resulting in a C·G to T·A base pair substitution. This substitution can be reduced to a wild-type, non-pathogenic genomic sequence using an adenosine base editor (ABE) that catalyzes the A·T to G·C substitution. By extension, mutations causing high RTT are potential targets for conversion to wild-type sequences using ABEs without the risk of Mecp2 gene overexpression associated with gene therapy. Therefore, base editing from A·T to G·C DNA has the potential to precisely correct one or more of the most common RTT-causing mutations in the Mecp2 gene.

[0028] The descriptions and examples herein illustrate embodiments of this application in detail. It should be understood that this application is not limited to the specific embodiments described herein, and therefore variations are possible. Those skilled in the art will recognize that many variations and modifications are included within the scope of this application.

[0029] The chapter titles used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0030] While various features of this application may be described in the context of a single embodiment, features may also be provided independently or in any suitable combination. Conversely, for clarity, this application may be described herein in the context of individual embodiments, but it may also be implemented in a single embodiment.

[0031] definition

[0032] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art. The following references provide general definitions for many of the terms used in this application: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd edition, 1994); The Cambridge Dictionary of Science and Technology (Walker, ed., 1988); The Glossary of Genetics, 5th edition, R. Rieger et al., (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0033] In this application, the use of the singular includes the plural unless otherwise expressly stated. It must be noted that the singular forms “a,” “an,” and “the” used in the specification include plural objects unless the context clearly specifies otherwise. In this application, the use of “or” means “and / or” unless otherwise stated. Furthermore, the use of the term “comprising” and other forms (e.g., “include,” “includes,” and “included”) is not restrictive.

[0034] As used in this specification and claims, the terms "comprising" (and any form of inclusion, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "containing" (and any form of inclusion, such as "includes" and "include"), or "containing" (and any form of inclusion, such as "contains" and "contain") are inclusive or open-ended and do not exclude other unreferenced elements or method steps. It is contemplated that any embodiment discussed in this specification may be implemented with respect to any method or composition of this application, and vice versa. Furthermore, the compositions of this application may be used to implement the methods of this application.

[0035] The terms "about" or "approximately" refer to a specific value determined by those skilled in the art within an acceptable margin of error, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, "about" can mean within one or more standard deviations. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude of the numerical value, preferably within 5 times, more preferably within 2 times. Where a specific value is described in this application and claims, unless otherwise stated, it should be assumed that the term "about" means that the specific value is within an acceptable margin of error.

[0036] The ranges provided herein should be understood as abbreviations of all values ​​within the range. For example, the range 1 to 50 should be understood as including any number, combination of numbers, or subranges selected from groups of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0037] References to “some embodiments,” “one embodiment,” “another embodiment,” or “another embodiment” in the specification refer to specific features, structures, or characteristics described in connection with these embodiments, which are included in at least some embodiments but are not necessarily all embodiments of this application.

[0038] "Adenosine deaminase" refers to a polypeptide or fragment thereof capable of catalyzing the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolysis of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be derived from any organism, such as bacteria.

[0039] "Reagent" refers to any small molecule compound, antibody, nucleic acid molecule, or polypeptide or fragment thereof.

[0040] "Improvement" refers to reducing, inhibiting, weakening, decreasing, stopping, or stabilizing the development or progression of a disease.

[0041] "Change" refers to a change (increase or decrease) in the expression level or activity of a gene or polypeptide, as detected by methods known in the standard field (such as those described herein). As used herein, a change includes a 10% change in expression level, preferably a 25% change, more preferably a 40% change, and most preferably a 50% or greater change in expression level.

[0042] "Analogous molecules" are molecules that are not identical but have similar functions or structural features. For example, peptide analogs retain the biological activity of the corresponding naturally occurring peptide while possessing certain biochemical modifications that enhance the function of the analog compared to the naturally occurring peptide. These biochemical modifications can increase the analog's protease resistance, membrane permeability, or half-life without altering, for example, ligand binding. Analogs may include non-natural amino acids.

[0043] As used herein, “administration” means providing a patient or subject with one or more of the compositions described herein. Examples, but not limited to, administration of the compositions may be by intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection, or one or more such routes may be used. For example, parenteral administration may be by bolus injection or gradual infusion over time. Alternatively or concurrently, administration may be by oral route.

[0044] "Cytidine deaminase" refers to a polypeptide or fragment thereof capable of catalyzing a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1, derived from Petromyzon marinus (Petromyzon marinus cytosine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA), which is derived from mammals (e.g., humans, pigs, cattle, horses, monkeys, etc.) and APOBEC, are exemplary cytidine deaminases.

[0045] "Methyl CpG-binding protein 2 (Mecp2) protein" refers to a polypeptide or fragment thereof that has at least about 95% amino acid sequence identity with NCBI accession number NP_004983. In a particular embodiment, the Mecp2 protein contains one or more alterations relative to the following reference sequence. In a particular embodiment, the RTT-related Mecp2 protein contains one or more mutations selected from R106W, R168*, R133C, T158M, R255*, R270*, and R306C. Exemplary Mecp2 amino acid sequences are provided below.

[0046]

[0047] "Mecp2 polynucleotide" refers to a nucleic acid molecule encoding the Mecp2 protein or a fragment thereof. Exemplary Mecp2 polynucleotide sequences are provided below, which are available in NCBI accession number NM_004992. In a particular embodiment, the Mecp2 polynucleotide includes one or more alterations relative to the reference sequence below. In a particular embodiment, the RTT-related Mecp2 polynucleotide includes one or more mutations selected from 316C>T, 397C>T, 473C>T, 763C>T, 808C>T, and 916C>T.

[0048]

[0049]

[0050]

[0051]

[0052] "Base editor (BE)" or "nucleobase editor (NBE)" refers to the reagent that binds to a polynucleotide and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a polynucleotide-programmable nucleotide-binding domain that binds to a guiding polynucleotide (e.g., a guide RNA). In various embodiments, the reagent is a biomolecular complex comprising a protein domain with base-editing activity, i.e., capable of modifying bases (e.g., deoxyribonucleic acid) in a nucleic acid molecule (e.g., A, T, C, G, or U). In some embodiments, the polynucleotide-programmable DNA-binding domain is fused to or linked to a deaminase domain. In one embodiment, the reagent is a fusion protein comprising a domain with base-editing activity. In another embodiment, the protein domain with base-editing activity is linked to a guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to a deaminase). In some embodiments, the domain with base-editing activity enables deamination of bases within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating bases within a DNA molecule. In some embodiments, the base editor is capable of deaminating cytosine (C) or adenosine (A) within DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the adenosine deaminase evolved from TadA. In some embodiments, the polynucleotide programmable DNA-binding domain is a CRISPR-related (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is a catalytically dead Cas9 (dCas9) fused with a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused with a deaminase domain. In some embodiments, the base editor is fused with a base excision repair (BER) inhibitor. In some embodiments, the base excision repair inhibitor is a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the base excision repair inhibitor is an inosine base excision repair inhibitor. Detailed information about the base editor is described in International Patent Application No. PCT / 2017 / 045381 (International Patent Application No. 2018 / 027078) and Patent Application No. PCT / US2016 / 058344 (International Patent Application No. 2017 / 070632), the entire contents of which are incorporated herein by reference.See also Komor, AC et al. "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM et al. "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551, 464-471 (2017); Komor, AC et al. "Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity" Science Advances3:eaao4774 (2017) and Rees, HA et al. "Base editing: precision chemistry on the genome and transcriptome of living cells." Nat Rev Genet.2018Dec;19(12):770-788.doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.

[0053] For example, the cytidine base editor CBE used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8877 base pairs), (Addgene, Watertown, MA; Komor AC et al., 2017, Sci Adv., 30; 3(8): eaao4774. doi: 10.1126 / sciadv.aao4774), as shown below. It also includes a polynucleotide sequence having at least 95% or higher identity with the BE4 nucleic acid sequence.

[0054]

[0055]

[0056]

[0057]

[0058] BE4 amino acid sequence:

[0059]

[0060] For example, the cytidine base editor ABE used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8877 base pairs), (Addgene, Watertown, MA.; Gaudelli NM et al., Nature. 2017 Nov 23; 551(7681): 464-471. doi: 10.1038 / nature24644), as shown below. It also includes a polynucleotide sequence having at least 95% or higher identity with the ABE nucleic acid sequence.

[0061]

[0062] "Base editing activity" refers to the use of chemical alteration of bases within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, for example, converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, for example, converting A·T to G·C.

[0063] The term "base editor system" refers to a system for editing the nucleobases of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises (1) a polynucleotide-programmable nucleotide-binding domain and a deaminase domain for deaminating the nucleobase; and (2) a guide polynucleotide (e.g., guide RNA) to bind to the polynucleotide-programmable nucleotide-binding domain. In some embodiments, the base editor system comprises (1) a base editor (BE) comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain for deaminating the nucleobase; and (2) a guide RNA to bind to the polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE).

[0064] In some embodiments, a nuclease base editor system may include more than one base editing component. For example, a nuclease base editor system may include more than one deaminase. In some embodiments, a nuclease base editor system may include one or more cytidine deaminases and / or one or more adenosine deaminases. In some embodiments, single-stranded guide polynucleotides may be used to target different deaminases onto target nucleic acid sequences. In some embodiments, single-stranded guide polynucleotides may be used to target different deaminases onto target nucleic acid sequences.

[0065] The nucleobase components and polynucleotide-programmable nucleotide-binding components of a base editor system can bind covalently or nonvalently to each other. For example, in some embodiments, a deaminase domain can be targeted to a target nucleotide sequence via a polynucleotide-programmable nucleotide-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be fused to or linked to a deaminase domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can target a deaminase domain to a target nucleotide sequence through nonvalent interaction with or association with the deaminase domain. For example, in some embodiments, the nucleobase editing component, such as a ribonucleic acid editing component, may include additional heterologous portions or domains that can interact with, associate with, or form complexes with the additional heterologous portions or domains that are part of the polynucleotide-programmable nucleotide-binding domain. In some embodiments, the additional heterologous portion may be able to bind to, interact with, associate with, or form complexes with a polypeptide. In some embodiments, the additional heterologous portion may be able to bind to, interact with, associate with, or form complexes with a polynucleotide. In some embodiments, the additional heterologous portion may be able to bind to a guide polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a peptide linker. In some embodiments, the additional heterologous portion may be capable of binding to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 capsid protein domain, a PP7 capsid protein domain, an SfMuCom capsid protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.

[0066] The base editor system may further include a guiding polynucleotide component. It should be understood that components of the base editor system can be associated with each other through covalent bonds, non-covalent interactions, or any combination of their association and interactions. In some embodiments, a deaminase domain may be a guiding polynucleotide targeting a target nucleotide sequence. For example, in some embodiments, the nucleobase editing component of the base editor system may be a deaminase component. The deaminase component may include additional heterologous portions or domains (e.g., polynucleotide-binding domains, such as RNA or DNA-binding proteins) that can interact with, associate with, or form complexes with the portion or segment of the guiding polynucleotide (e.g., a polynucleotide motif). In some embodiments, the additional heterologous portions or domains (e.g., polynucleotide-binding domains, such as RNA or DNA-binding proteins) may be fused to or linked to the deaminase domain. In some embodiments, the additional heterologous portions may be able to bind to, interact with, associate with, or form complexes with polypeptides. In some embodiments, the additional heterologous portions may be able to bind to, interact with, associate with, or form complexes with polynucleotides. In some embodiments, the additional heterologous portion may be able to bind to a guiding polynucleotide. In some embodiments, the additional heterologous portion may be able to bind to a peptide linker. In some embodiments, the additional heterologous portion may be able to bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 capsid protein domain, a PP7 capsid protein domain, an SfMuCom capsid protein domain, an asosome α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.

[0067] In some embodiments, the base editor system may further include an inhibitor of the base excision repair (BER) component. It should be understood that the components of the base editor system can be associated with each other through covalent bonds, non-covalent interactions, or any combination of their association and interactions. The inhibitor of the BER component may include a base excision repair inhibitor. In some embodiments, the base excision repair inhibitor may be a uracil DNA glycosyltransferase inhibitor (UGI). In some embodiments, the base excision repair inhibitor may be an inosine base excision repair inhibitor. In some embodiments, the base excision repair inhibitor may target the target nucleotide sequence via a polynucleotide-programmable nucleotide binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain may be fused to or linked to the base excision repair inhibitor. In some embodiments, the polynucleotide-programmable nucleotide binding domain may be fused to or linked to a deaminase domain and the base excision repair inhibitor. In some embodiments, the polynucleotide-programmable nucleotide binding domain may target the base excision repair inhibitor to the target nucleotide sequence through non-covalent interaction with or association with the base excision repair inhibitor. For example, in some embodiments, the nuclear base editing component (e.g., a deaminase component) of the base editor system may include additional heterologous portions or domains (e.g., polynucleotide binding domains, such as RNA or DNA binding proteins) capable of interacting, associating, or forming complexes with a portion or segment of the guiding polynucleotide (e.g., a polynucleotide motif). In some embodiments, a base excision repair inhibitor may be a guiding polynucleotide targeting a target nucleotide sequence. For example, in some embodiments, the base excision repair inhibitor may include additional heterologous portions or domains (e.g., polynucleotide binding domains, such as RNA or DNA binding proteins) capable of interacting, associating, or forming complexes with a portion or segment of the guiding polynucleotide (e.g., a polynucleotide motif). In some embodiments, the additional heterologous portion or domain of the guiding polynucleotide (e.g., a polynucleotide binding domain, such as an RNA or DNA binding protein) may be fused to or linked to the base excision repair inhibitor. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming complexes with the polynucleotide. In some embodiments, the additional heterologous portion may be able to bind to a guiding polynucleotide. In some embodiments, the additional heterologous portion may be able to bind to a peptide linker. In some embodiments, the additional heterologous portion may be able to bind to a polynucleotide linker. The additional heterologous portion may be a protein domain.In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 capsid protein domain, a PP7 capsid protein domain, an SfMuCom capsid protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.

[0068] The term "Cas9" or "Cas9 domain" refers to a protein containing Cas9 or a fragment thereof (e.g., a protein containing the active, inactive, or partially active DNA-cutting domain of Cas9, and / or the gRNA-binding domain of Cas9). Cas9 nucleases are sometimes also called casnl nucleases or CRISPR (regularly spaced clusters of short palindromic repeats)-related nucleases. An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), whose amino acid sequence is shown below.

[0069]

[0070] The term "conserved amino acid substitution" or "conserved mutation" refers to the substitution of one amino acid for another that shares a common property. One functional approach to defining shared properties among individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins from homologous organisms (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on such analysis, amino acid groups can be defined where amino acids within a group preferentially exchange with each other and are therefore most similar in their effects on the overall protein structure (Schulz, GE and Schirmer, RH, ibid.). Non-restrictive examples of conserved mutations include amino acid substitutions, such as lysine replacing arginine and vice versa, thus maintaining a positive charge; glutamic acid replacing aspartic acid and vice versa, thus maintaining a negative charge; threonine replacing serine to maintain a free -OH group; and aminoamide replacing asparagine to maintain a free -NH2 group.

[0071] As may be used interchangeably herein, the terms "coding sequence" or "protein-coding sequence" refer to a segment of polynucleotides that encodes a protein. The start codon of this region or sequence is located near the 5' end, and the stop codon is located near the 3' end. Coding sequences may also be referred to as opening reading frames.

[0072] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolysis of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolysis of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolysis of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolysis of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolysis of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, adenosine deaminase catalyzes the hydrolysis of adenosine in deoxyribonucleic acid (DNA) to deaminase. The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be derived from any organism, such as bacteria. In some embodiments, the adenosine deaminase is derived from bacteria such as *Escherichia coli*, *Staphylococcus aureus*, *Streptococcus typhi*, *Streptococcus putrefactive*, *Haemophilus influenzae*, or *Clostridium crescentis*. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is inherently absent. For example, in some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase with at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity. For example, the deaminase domain is described in International Patent Application No. PCT / 2017 / 045381 (WO No. 2018 / 027078) and International Patent Application No. PCT / US2016 / 058344 (WO No. 2017 / 070632), which are incorporated herein by reference in their entirety.See also Komor, AC et al., "Programmable base editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM et al., "Programmable base editing of A·T to G·C ingenomic DNA without DNA cleavage" Nature 551,464-471(2017); Komor, AC et al., "Improved base excision repair inhibition and bacteriophage Mu Gam proteinyields C:G-to-T:A base editors with higher efficiency and product purity" Science Advances 3:eaao4774(2017) and Rees, HA et al., "Base editing: precisionchemistry on the genome and transcriptome of living cells." Nat RevGenet.2018Dec;19(12):770-788.doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.

[0073] "Detectable label" refers to a composition that, when linked to a molecule of interest, enables it to be detected by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. Useful labels include, for example, radioactive isotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., those commonly used in ELISA), biotin, digoxigenin, or haptens.

[0074] A "disease" is any condition or disorder that impairs or interferes with the normal function of cells, tissues, or organs. Examples of diseases include Rett syndrome.

[0075] "Effective amount" refers to the amount of the reagent or active compound described herein that provides relief relative to an untreated patient or an individual without disease. The effective amount of the active compound used to implement this application to treat a disease varies depending on the route of administration, age, weight, and the overall health status of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. This dosage is referred to as the "effective" amount. In one embodiment, the effective amount is sufficient to introduce the base editor of this application that alters the target gene (e.g., Mecp2) in cells (e.g., in vitro or in vivo cells). In one embodiment, the effective amount is the amount of base editor required to achieve a therapeutic effect (e.g., reduction or control of Rett syndrome or its symptoms or symptom). This therapeutic effect is insufficient to alter Mecp2 in all cells of a subject, tissue, or organ, but only sufficient to alter Mecp2 in about 1%, 5%, 10%, 25%, 50%, 75%, or more cells present in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to relieve one or more symptoms of Rett syndrome.

[0076] A “fragment” refers to a portion of a polypeptide or nucleic acid molecule. The portion preferably comprises at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of a reference nucleic acid molecule or polypeptide. A fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.

[0077] "Hybridization" refers to hydrogen bonds between complementary nucleobases, which can be Watson-Crick, Hoogsteen, or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that pair up by forming hydrogen bonds.

[0078] The terms "base repair inhibitor," "base repair inhibitor," "IBR," or their grammatical equivalents refer to proteins capable of inhibiting the activity of nucleic acid repair enzymes, such as base excision repair enzymes. In some embodiments, an IBR is an inosine base excision repair inhibitor. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, a base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, an IBR is an inhibitor of Endo V or hAAG. In some embodiments, an IBR is a catalytically inactivated EndoV or a catalytically inactivated hAAG. In some embodiments, a base repair inhibitor is a catalytically inactivated EndoV or a catalytically inactivated hAAG. In some embodiments, a base repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI refers to a protein capable of inhibiting uracil-DNA glycosylase base excision repair enzymes. In some embodiments, the UGI domain comprises wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI protein provided herein comprises a fragment of UGI and a protein homologous to UGI or a UGI fragment. In some embodiments, the base repair inhibitor is an inosine base excision repair inhibitor. In some embodiments, the base repair inhibitor is a “catalytically inactivated inosine-specific nuclease” or a “dead inosine-specific nuclease.” Without wishing to be bound by any particular theory, a catalytically inactivated inosine glycosylation enzyme (e.g., alkyladenine glycosylation enzyme (AAG)) can bind inosine but cannot create a debasement site or remove inosine, thereby spatially preventing the newly formed inosine moiety from being subjected to DNA damage / repair mechanisms. In some embodiments, a catalytically inactivated inosine-specific nuclease is capable of binding inosine in nucleic acids but does not cleave the nucleic acids. Non-limiting exemplary catalytically inactivated inosine-specific nucleases include, for example, catalytically inactivated alkyladenine glycosylation enzymes (AAG nucleases) from humans and, for example, catalytically inactivated endonuclease V (EndoV nucleases) from *E. coli*. In some implementations, the catalytically inactivated AAG nuclease contains an E125Q mutation or a corresponding mutation in another AAG nuclease.

[0079] The terms "isolated," "purified," or "biologically pure" refer to material free from various degrees of different components, which are typically present in the natural state of the material. "Isolated" indicates the degree of isolation from the original source or surrounding environment. "Purified" indicates a degree of separation beyond isolation. A "purified" or "biologically pure" protein is sufficiently free of other substances such that any impurities do not materially affect the biological properties of the protein or cause other adverse consequences. That is, if the nucleic acid or peptide of this application is substantially free of cellular material, viral material, or culture medium when produced using recombinant DNA technology, or substantially free of chemical precursors or other chemicals during chemical synthesis, it is purified. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can mean that a nucleic acid or protein essentially forms a single band in an electrophoretic gel. For proteins that can be modified, such as phosphorylation or glycosylation, different modifications can produce different isolated proteins, which can be purified separately.

[0080] "Separated polynucleotide" refers to a nucleic acid (e.g., DNA) that does not contain genes located flanking a gene in the naturally occurring genome of an organism from which the nucleic acid molecule of this application is derived. Therefore, the term includes, for example, recombinant DNA incorporated into a vector; DNA incorporated into an autonomously replicating plasmid or virus; or genomic DNA incorporated into a prokaryotic or eukaryotic organism; or DNA existing as an independent molecule separate from other sequences (e.g., cDNA or a genomic or cDNA fragment produced by PCR or restriction endonuclease digestion). Additionally, the term includes RNA molecules transcribed from DNA molecules, and recombinant DNA as part of a heterozygous gene encoding an additional polypeptide sequence.

[0081] "Isolated polypeptide" refers to the polypeptide of this application that has been separated from its naturally occurring accompanying components. Typically, the polypeptide is isolated when it contains at least 60% by weight of proteins and naturally occurring organic molecules naturally bound to it. Preferably, the formulation is at least 75%, more preferably at least 90%, and most preferably at least 99% by weight of the polypeptide of this application. The isolated polypeptide of this application can be obtained, for example, by extraction from a natural source, by expressing a recombinant nucleic acid encoding such polypeptide, or by chemically synthesizing a protein, and its purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis.

[0082] As used herein, the term "connector" can refer to a covalent connector (e.g., a covalent bond), a non-covalent connector, a chemical group, or a linker connecting two molecules or portions, such as two components of a protein complex or ribonucleic acid complex, or two domains of a fusion protein, such as a polynucleotide-programmable DNA-binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase or cytidine deaminase). Connectors can link different components or different portions of a base editor system. For example, in some embodiments, a connector can link a guiding polynucleotide-binding domain of a polynucleotide-programmable nucleotide-binding domain and a catalytic domain of a deaminase. In some embodiments, a connector can link a CRISPR peptide and a deaminase. In some embodiments, a connector can link Cas9 and a deaminase. In some embodiments, a connector can link dCas9 and a deaminase. In some embodiments, a connector can link nCas9 and a deaminase. In some embodiments, a connector can link a guiding polynucleotide and a deaminase. In some embodiments, a connector can link a deamination component and a polynucleotide-programmable nucleotide-binding component of a base editor system. In some embodiments, the adapter may connect the RNA-binding portion of the deamination component to the polynucleotide-programmable nucleotide-binding component of the base editor system. In some embodiments, the adapter may connect the RNA-binding portion of the deamination component to the RNA-binding portion of the polynucleotide-programmable nucleotide-binding component of the base editor system. The adapter may be located between or on either side of two groups, molecules, or other portions, and may connect to each via covalent or non-covalent interactions, thereby linking them. In some embodiments, the adapter may be an organic molecule, group, polymer, or chemical portion. In some embodiments, the adapter may be a polynucleotide. In some embodiments, the adapter may be a DNA adapter. In some embodiments, the adapter may be an RNA adapter. In some embodiments, the adapter may include an aptamer capable of binding a ligand. In some embodiments, the ligand may be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the adapter may include an aptamer derived from a riboswitch. The riboswitch of the derived aptamer can be selected from theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or pre-queosine1 (PreQ1) riboswitch. In some embodiments, the linker may comprise an aptamer that binds to a peptide or protein domain, such as a peptide ligand.In some embodiments, the peptide ligand may be a K homology (KH) domain, an MS2 capsid protein domain, a PP7 capsid protein domain, an SfMuCom capsid protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the peptide ligand may be part of a base editor system component. For example, the nucleobase editing component may include a deaminase domain and an RNA recognition motif.

[0083] In some embodiments, the linker may be an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker length may be about 5 to 100 amino acids, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. In some embodiments, the linker length may be about 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, 400 to 450, or 450 to 500 amino acids. Longer or shorter linkers may also be considered.

[0084] In some embodiments, the linker connects the gRNA-binding domain of an RNA-programmable nuclease, including the Cas9 nuclease domain, and the catalytic domain of a nucleic acid editing protein (e.g., cytidine or adenosine deaminase). In some embodiments, the linker connects dCas9 and the nucleic acid editing protein. For example, the linker is located between or on the sides of two groups, molecules, or other parts and is covalently linked to each, thereby connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker length is from 5 to 200 amino acids, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids. Longer or shorter linkers are also possible. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES, which may also be referred to as the XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises (SGGS).n (GGGS) n (GGGGS) n (G) n (EAAAK) n (GGS) n SGSETPGTSESATPES or (XP) n A motif, or any combination of the following, where n is independently an integer between 1 and 30, and X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises multiple proline residues and amino acids of length 5 to 21, 5 to 14, 5 to 9, or 5 to 7, such as PAPAP, PAPAPA, PAPAPAP, PAPAPAPA, P(AP)4, P(AP)7, P(AP) 10 This type of proline-rich connector is also known as a "rigid" connector.

[0085] In some embodiments, the domain of the base editor is fused via a linker comprising the following amino acid sequence:

[0086] SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the base editor domain is fused with a linker containing the amino acid sequence SGSETPGTSESATPES, also known as an XTEN linker. In some embodiments, the linker is 24 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGSSPGSETPGTSESATPESSGGS SGGS. In some embodiments, the linker is 92 amino acids long. In some embodiments, the linker contains the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAP GTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0087] As used herein, the term “mutation” refers to the substitution of a residue in a sequence of, for example, nucleic acids or amino acid sequences, by another residue, or the deletion or insertion of one or more residues in the sequence. Mutations are generally described herein by identifying the original residue, and subsequently the position of the original residue in the sequence and the identity of the newly substituted residue. Various methods for performing amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (2012)). In some embodiments, currently disclosed base editors can efficiently generate “desired mutations,” such as point mutations, in nucleic acids (e.g., nucleic acids within the genome of a subject) without generating a large number of undesired mutations, such as accidental point mutations. In some embodiments, the desired mutation is generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) that is specifically designed to bind to a guiding polynucleotide (e.g., gRNA) for generating the desired mutation. Typically, mutations that arise or are identified in a sequence (e.g., the amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain mutations. Those skilled in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence.

[0088] As used herein, the terms “nucleic acid” and “nucleic acid molecule” refer to compounds comprising nucleobases and an acidic moiety (e.g., a nucleoside, nucleotide, or polymer of nucleotides). Typically, polymerized nucleic acids, such as nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other by phosphodiester bonds. In some embodiments, “nucleic acid” refers to a single nucleic acid residue (e.g., a nucleotide and / or a nucleoside). In some embodiments, “nucleic acid” refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms “oligonucleotide,” “polynucleotide,” and “polynucleic acid” are used interchangeably to refer to polymers of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, “nucleic acid” encompasses RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can be naturally occurring, for example, in the context of genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, viscera, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or including non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified, chemically synthesized, etc. Where appropriate, for example, in the case of chemically synthesized molecules, nucleic acids may contain nucleoside analogs, such as those with chemically modified bases or sugars and backbone modifications. Unless otherwise specified, nucleic acid sequences are shown in a 5' to 3' orientation. In some embodiments, the nucleic acid is or comprises a natural nucleoside (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); or a nucleoside analogue (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5...). -methylcytosine, 2-aminoadenosine, 7-deazonoadenosine, 7-deazonoguanosine, 8-oxoadenosine, 8-oxoguanosine, O6-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); intercalating groups; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., thiophosphates and 5'-N-phosphoramide bonds).

[0089] The terms “nuclear localization sequence,” “nuclear localization signal,” or “NLS” refer to an amino acid sequence that facilitates protein delivery into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in International Patent Application No. PCT / EP2000 / 011690, filed November 23, 2000, by Plank et al., and in International Patent Application No. WO 2001 / 038547, published May 31, 2001, the contents of which are incorporated herein by reference due to their exemplary nuclear localization sequences. In other embodiments, the NLS is, for example, an optimized NLS as described by Koblan et al., Nature Biotech, 2018, doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises an amino acid sequence

[0090] KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0091] In this document, the terms “nucleobase,” “nitrogenous base,” or “base” are used interchangeably to refer to nitrogen-containing biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to form base pairs and stack on top of each other directly results in long-chain helical structures, such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases—adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U)—are referred to as primary or canonical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA may also contain other modified (non-primary) bases. Non-limiting exemplary modifications of nucleobases may include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydromethylcytosine. Hypoxanthine and xanthine can be produced in the presence of mutagens, and both can be produced through deamination (replacing an amino group with a carbonyl group). Hypoxanthine can be modified by adenine. Xanthine can be modified by guanine. Uracil can be produced by deamination of cytosine. A "nucleoside" consists of a nucleobase and a pentose sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a pentose sugar (ribose or deoxyribose), and at least one phosphate group.

[0092] The terms "nucleic acid programmable DNA-binding protein" or "napDNAbp" are used interchangeably with "polynucleotide programmable nucleotide-binding domain" to refer to a protein (e.g., DNA or RNA) that associates with a nucleic acid, such as a guide nucleic acid, directing the napDNAbp to a specific nucleic acid sequence. For example, a Cas9 protein can bind to a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as an active nuclease Cas9, a Cas9 nickase (nCas9), or an inactive nuclease Cas9 (dCas9). Examples of nucleic acid programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Other nucleic acid programmable DNA-binding proteins are also within the scope of this application, although they may not be specifically listed herein. See, for example, Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct; 1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4; 363(6422):88-91. doi:10.1126 / science.aav7271, the full text of each of these references is incorporated herein by reference.

[0093] As used herein, the terms "nucleobase editing domain" or "nucleobase editing protein" refer to proteins or enzymes that catalyze nucleobase modifications in RNA or DNA, such as the conversion of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymine nucleoside), the deamination of hypoxanthine (or inosine) by adenine (or adenosine), and the addition and insertion of non-template nucleotides. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., cytidine deaminase, cytosine deaminase, adenine deaminase, or adenosine deaminase). In some embodiments, the nucleobase editing domain may be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain may be an engineered or evolved nucleobase editing domain derived from a naturally occurring nucleobase editing domain. Nucleobase editing domains can originate from any organism, such as bacteria, humans, chimpanzees, gorillas, monkeys, cattle, dogs, rats, or mice. For example, nucleobase editing proteins are described in International Patent Application No. PCT / 2017 / 045381 (International Patent Application No. 2018 / 027078) and International Patent Application No. PCT / US2016 / 058344 (International Patent Application No. 2017 / 070632), each of which is incorporated herein by reference in its entirety. See also Komor, AC et al., “Programmable editing of a target basein genomic DNA without double-stranded DNA cleavage”, Nature 533, 420-424 (2016); Gaudelli, NM et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage”, Nature 551, 464-471 (2017); and Komor, AC et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity”, Science Advances 3:eaao4774 (2017), the contents of which are incorporated herein by reference.

[0094] As used in this article, “obtain” and “obtain reagent” include reagents that are synthesized, purchased or otherwise obtained.

[0095] As used herein, "patient" or "subject" means a mammalian subject or individual diagnosed with, having, developing, or suspected of having or developing a disease or condition. In some embodiments, the term "patient" refers to a mammalian subject with a higher-than-average probability of having a disease or condition. Exemplary patients may be humans, non-human primates, cats, dogs, pigs, cattle, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other animals that may benefit from the therapies disclosed herein. Exemplary human patients may be male and / or female.

[0096] "Patients in need" or "subjects in need" in this article refers to patients who have been diagnosed with or are suspected of having a disease or condition, such as, but not limited to, Rett syndrome (RTT).

[0097] The terms "pathogenic mutation," "pathogenic variation," "disease-causing mutation," "disease-causing variation," "harmful mutation," or "susceptibility mutation" refer to genetic variations or mutations that increase an individual's susceptibility to or predisposition to a particular disease or condition. In some embodiments, a pathogenic mutation comprises at least one wild-type amino acid replaced by at least one pathogenic amino acid in a protein encoded by a gene.

[0098] The term "non-conservative mutation" refers to the substitution of amino acids between different groups, such as lysine replacing tryptophan or phenylalanine replacing serine. In this case, non-conservative amino acid substitutions are preferably not substituted, interfere with the functional variant, or inhibit its function. Non-conservative amino acid substitutions can enhance the biological activity of the functional variant, thereby increasing the biological activity of the functional variant compared to the wild-type protein.

[0099] The terms “protein,” “peptide,” “polypeptide,” and their grammatical equivalents are used interchangeably herein to refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The term refers to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids in length. A protein, peptide, or polypeptide can refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified, for example, by adding chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. A protein, peptide, or polypeptide can also be a single molecule or a multi-molecule complex. A protein, peptide, or polypeptide can simply be a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof. As used herein, the term “fusion protein” refers to a hybrid polypeptide containing protein domains from at least two different proteins. A protein may be located on the N-terminal (N-terminal) portion or the C-terminal (C-terminal) portion of a fusion protein, thereby forming an N-terminal fusion protein or a C-terminal fusion protein, respectively. The protein may contain various domains, such as nucleic acid-binding domains (e.g., the gRNA-binding domain of Cas9 that guides protein binding to a target site) and nucleic acid-cleaving domains or catalytic domains of nucleic acids in acid-editing proteins. In some embodiments, the protein comprises a protein moiety, such as an amino acid sequence constituting the nucleic acid-binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In some embodiments, the protein is complexed with or associated with a nucleic acid, such as RNA or DNA. Any protein provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known, including those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (2012)), the entire contents of which are incorporated herein by reference.

[0100] The polypeptides and proteins disclosed herein (including their functional portions and functional variants) may contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, leucine, α-aminodecanoic acid, homoserine, S-acetaminomethylcysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, and dihydroindole-2-hydroxylamine. -Carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyllysine, N',N'-dibenzyllysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornene)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. Peptides and proteins can be associated with post-translational modifications of one or more amino acids in the peptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acylation (including acetylation and formylation), glycosylation (including N-linking and O-linking), amidation, hydroxylation, alkylation (including methylation and ethylation), ubiquitination, addition of pyrrolidone carboxylic acid, disulfide bond formation, sulfation, myristylation, palmitoylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation, and iodination.

[0101] The term "polynucleotide-programmable nucleotide-binding domain" refers to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide polynucleotide (e.g., guide RNA) that directs the polynucleotide-programmable DNA-binding domain to a specific nucleic acid sequence. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable RNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein can bind to guide RNA, which directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a Cas9 domain, such as an active nuclease Cas9, a Cas9 nickase (nCas9), or a nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpf1, Cas12b / C2c1 and Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h and Cas12i.Non-restricted examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, and Csc. Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein, type VI Cas effector protein, CARF, DinG, their homologs or their modified or engineered forms. Other nucleic acid-programmable DNA-binding proteins are also within the scope of this application, although they are not specifically listed herein.

[0102] As used herein in the context of proteins or nucleic acids, the term "recombinant" refers to a protein or nucleic acid that is not found in nature but is a product of human engineering. For example, in some embodiments, recombinant protein or nucleic acid molecules comprise an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0103] "Reduction" refers to a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0104] "Reference" refers to a standard or control condition. In one implementation, the reference is wild-type or healthy cells.

[0105] A "reference sequence" is a defined sequence used as the basis for sequence comparisons. The reference sequence can be a subset or the entirety of the specified sequence. For example, it can be a fragment of a full-length cDNA or gene sequence, or the complete cDNA or gene sequence. For peptides, the reference peptide sequence is typically at least about 16 amino acids in length, preferably at least about 20 amino acids, more preferably at least about 25 amino acids, and even more preferably about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the reference nucleic acid sequence is typically at least about 50 nucleotides in length, preferably at least about 60 nucleotides, more preferably at least about 75 nucleotides, and even more preferably about 100 nucleotides or about 300 nucleotides, or any integer between or near these values.

[0106] The terms “RNA-programmable nuclease” and “RNA-directed nuclease” are used in conjunction with one or more RNAs that are not cleaving targets (e.g., binding or associated). In some embodiments, when an RNA-programmable nuclease is complexed with RNA, it may be referred to as a nuclease:RNA complex. Typically, the bound RNA is referred to as guide RNA (gRNA). gRNA can exist as a complex of two or more RNAs or as a single-stranded RNA molecule. gRNA existing as a single-stranded RNA molecule may be referred to as single-stranded guide RNA (sgRNA), although “gRNA” is used interchangeably to refer to guide RNA existing as a single-molecule, bimolecule, or multimolecule complex. Typically, gRNA existing as a single-stranded RNA contains two domains: (1) a domain homologous to the target nucleic acid (e.g., and directing the Cas9 complex to bind to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided by Jinek et al. in Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNA (e.g., including such domain 2) can be found in the entire contents of U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled “Switchable Cas9 Nucleases and Uses,” and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled “Delivery System For Functional Nucleases,” which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more of domains (1) and (2) and may be referred to as an “extended gRNA.” For example, as described herein, an extended gRNA will, for example, bind two or more Cas9 proteins and bind target nucleic acids in two or more distinct regions. gRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site, providing sequence specificity of the nuclease:RNA complex.In some implementations, the RNA-programmable nuclease is a Cas9 endonuclease (of a CRISPR-related system), for example, Cas9 (Csnl) from Streptococcus pyogenes (see, for example, “Complete genome sequence of an Ml strain of Streptococcus pyogenes.”, Ferretti JJ., McShan WM., Ajdic DJ., Savic DJ., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN., Kenton S., Lai HS., Lin SP., Qian Y., JiaHG., Najar FZ., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW., Roe BA., McLaughlin RE., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNAmaturation by trans-encoded small RNA and host factor RNase”. III.", DeltchevaE., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471: 602 -607(2011).

[0107] The term "single nucleotide polymorphism (SNP)" is a variation in a single nucleotide at a specific location in the genome, where each variation is present in a population at a certain level (e.g., >1%). For example, at a specific base position in the human genome, the C nucleotide may be present in most individuals, but in a minority of individuals, that position is occupied by A. This means that there is an SNP at that specific position, and the two possible nucleotide variations, C or A, are considered alleles at that position. SNPs are the basis for differences in disease susceptibility. The severity of disease and how the body responds to treatment are also manifestations of genetic variation. SNPs can fall within the coding region of a gene, the non-coding region of a gene, or the intergenetic region (the region between genes). In some implementations, due to the degeneracy of the genetic code, SNPs within the coding sequence do not necessarily change the amino acid sequence of the resulting protein. There are two types of SNPs in coding regions: synonymous and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs change the amino acid sequence of the protein. There are two types of non-synonymous SNPs: missense and nonsense. SNPs not located in protein-coding regions can still affect gene splicing, transcription factor binding, messenger RNA degradation, or non-coding RNA sequences. Gene expression affected by this type of SNP is called an eSNP (expressed SNP), which can occur upstream or downstream of the gene. Single nucleotide variants (SNVs) are variations of a single nucleotide, with no frequency limit, and can occur in somatic cells. Somatic single nucleotide variants (e.g., those caused by cancer) can also be called single nucleotide variants.

[0108] "Specific binding" refers to nucleic acid molecules, peptides or complexes thereof (e.g., nucleic acid programmable DNA binding domains and guiding nucleic acids), compounds or molecules that recognize and bind to peptides and / or nucleic acids. This application describes acidic molecules that, while not specifically recognizing and binding to samples, such as other molecules in biological samples, do not actually recognize and bind to them.

[0109] Nucleic acid molecules that can be used in the methods of this application include any nucleic acid molecule encoding the polypeptide or fragment thereof of this application. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but will generally show substantial identity. Polynucleotides that have “substantially identical” to the endogenous sequence are generally capable of hybridizing with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules that can be used in the methods of this application include any nucleic acid molecule encoding the polypeptide or fragment thereof of this application. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but will generally show substantial identity. Polynucleotides that have “substantially identical” to the endogenous sequence are generally capable of hybridizing with at least one strand of a double-stranded nucleic acid molecule. “Hybridization” means the formation of a pair of double-stranded molecules between complementary polynucleotide sequences (e.g., the gene described herein) or portions thereof under various stringent conditions. (See, for example, Wahl, GM and SLBerger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).

[0110] For example, stringent salt concentrations are typically less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, and more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvents such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions will typically include at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Variations in other parameters, such as hybridization time, detergent concentration (e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of vector DNA, are well known to those skilled in the art. Various stringencies are achieved by combining these various conditions as needed. In a preferred embodiment, hybridization will occur at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In a more preferred embodiment, hybridization is performed at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA. In the most preferred embodiment, hybridization occurs at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. The useful variations under these conditions will be apparent to those skilled in the art.

[0111] For most applications, the stringency of the post-hybridization washing step will vary. Washing stringency conditions can be defined by salt concentration and temperature. As mentioned above, washing stringency can be increased by decreasing the salt concentration or by increasing the temperature. For example, the stringent salt concentration for the washing step will preferably be less than about 30 mM NaCl and 3 mM trisodium citrate, most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. The stringent temperature conditions for the washing step typically include at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In a preferred embodiment, the washing step will occur at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step will be carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In an even more preferred embodiment, the washing step will be carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Other variations of these conditions will be apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and have been described, for example, by Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biolog, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0112] "Subject" refers to mammals, including but not limited to human or non-human mammals such as cows, horses, dogs, sheep or cats.

[0113] "Substantially identical" means identical to a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein). Preferably, such a sequence is at least 60%, more preferably 80% or 85%, and even more preferably 90%, 95%, or even 99% identical to the sequence used for comparison at the amino acid or nucleic acid level.

[0114] Sequence identity is typically measured using sequence analysis software (e.g., the sequence analysis software package, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs from the Genetics Computing Group at the Biotechnology Center, University of Wisconsin (1710 University Avenue, Madison, Wis. 53705)). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conserved substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine, lysine, arginine; and phenylalanine, tyrosine. In an exemplary method for determining the degree of identity, the BLAST program can be used, where e -3 and e -100 The probability scores between them represent closely related sequences.

[0115] For example, COBALT is used for the following parameters:

[0116] a) Alignment parameters: Vacancy penalty -11, -1 and vacancy penalty -5, -1

[0117] b) CDD parameters: Use RPS BLAST; Blast E value 0.003; search for conservative columns and recalculate, then...

[0118] c) Query clustering parameters: Use query clustering; character length 4; maximum cluster distance 0.8; regular letters.

[0119] For example, the EMBOSS Needle can use the following parameters:

[0120] a) Matrix: BLOSUM62;

[0121] b) Open slots: 10;

[0122] c) Expand vacancy: 0.5;

[0123] d) Output format: in pairs;

[0124] e) Final open shot penalty: False;

[0125] f) Final open position: 10; and

[0126] g) Final vacancy expansion: 0.5.

[0127] The term "target site" refers to a sequence within a nucleic acid molecule that has been modified by a nucleobase editor. In one embodiment, the target site is deaminated by a deaminase or a fusion protein containing a deaminase (e.g., cytidine or adenine deaminase).

[0128] Because RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins can, in principle, target any sequence specified by the guiding RNA. Methods for site-specific cleavage (e.g., genome modification) using RNA-programmable nucleases (e.g., Cas9) are known in the art (see, for example, Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems., Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9., Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system., Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems., Nucleic acids). research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the full contents of each are incorporated herein by reference.

[0129] As used herein, the term "treatment" means to reduce or improve the associated condition and / or symptoms or to achieve the desired pharmacological and / or physiological effect. It should be understood that, although not excluded, treating a disease or condition does not require the complete elimination of the associated disease, condition, or symptoms. In some embodiments, the effect is therapeutic, i.e., but not limited to, the effect partially or completely reduces, diminishes, abrogates, abates, alleviates, reduces, decreases the intensity of the disease, or cures the disease and / or unpleasant symptoms attributed to the disease. In some embodiments, the effect is preventative, i.e., the effect protects against or prevents the occurrence or recurrence of the disease or condition. For this purpose, the method disclosed in this application includes administering a therapeutically effective amount of the composition described herein.

[0130] "Uracil glycosylation enzyme inhibitor" refers to an agent that inhibits the uracil excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to host uracil-DNA glycosylation enzymes and prevents the removal of uracil residues from DNA.

[0131] The enumeration of chemical groups in any definition of a variable herein includes defining the variable as any single group or combination of the listed groups. The description of embodiments of variables or aspects herein includes the embodiments as any single embodiment or in combination with any other embodiment or part thereof.

[0132] Any composition or method provided herein may be combined with one or more of any other composition or method provided herein.

[0133] DNA editing has become a viable method for altering disease states by correcting pathogenic mutations at the genetic level. Until recently, all DNA editing platforms functioned by inducing DNA double-strand breaks (DSBs) at designated genomic sites and relying on endogenous DNA repair pathways to determine the product outcome in a semi-random manner, resulting in complex populations of genetic products. While precise, user-defined repair outcomes can be achieved through homology-directed repair (HDR) pathways, several challenges have prevented efficient repair using HDR in treatment-related cell types. In practice, it is inefficient compared to competing, error-prone non-homologous end joining pathways. Furthermore, HDR is strictly limited to the G1 and S phases of the cell cycle, thus preventing precise repair of DSBs in post-mitotic cells. As a result, it has proven difficult or impossible to efficiently alter genomic sequences in these populations in a user-defined, programmable manner.

[0134] Nucleotide base editor

[0135] This document discloses a base editor or nucleobase editor for editing, modifying, or altering target nucleotide sequences of polynucleotides. The description herein refers to a nucleobase editor or base editor comprising a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain. When bound to a guiding polynucleotide (e.g., gRNA), the polynucleotide programmable nucleotide binding domain can specifically bind to the target polynucleotide sequence (i.e., through complementary base pairing between the bases of the binding guiding nucleic acid and the bases of the target polynucleotide sequence), thereby positioning the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0136] Multinucleotide programmable nucleotide binding domain

[0137] The term "polynucleotide-programmable nucleotide-binding domain" or "nucleic acid-programmable DNA-binding protein (napDNAbp)" refers to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide polynucleotide (e.g., guide RNA), which directs the polynucleotide-programmable nucleotide-binding domain to a specific nucleic acid sequence. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable RNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is the Cas9 protein. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is the Cpf1 protein.

[0138] CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposons, and binding plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the preceding mobile element, and a target nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA requires a reverse-coding small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA acts as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA performs endonuclease cleavage of linear or circular dsDNA targets complementary to the spacer. Target strands not complementary to crRNA are first endonucleated, then exonucleated 3'-5'. In practice, DNA binding and cleavage typically require both a protein and two RNAs. However, single-stranded guide RNAs (“sgRNA” or simply “gRNA”) can be engineered to integrate various aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 identifies short motifs (PAM or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish itself from non-self motifs.

[0139] Cas9 field of nucleobase editor

[0140] The Cas9 nuclease sequence and structure are well known to those skilled in the art (e.g., see “Completegenome sequence of an Ml strain of Streptococcus pyogenes,” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Deltcheva E., Chylinski. K., Sharma C.M., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.”, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), each incorporated herein by reference. Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus.Based on the disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include those from Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9families of type IICRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.

[0141] In some aspects, the nucleic acid programmable DNA-binding protein (napDNAbp) is a Cas9 domain. This document provides non-limiting exemplary Cas9 domains. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 cleavage enzyme. In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any amino acid sequence as described herein. In some embodiments, the amino acid sequence comprised of the Cas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any amino acid sequence listed herein. In some implementations, the Cas9 domain contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences listed herein. In some embodiments, the Cas9 domain comprises at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any amino acid sequence listed herein.

[0142] In some implementations, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain; that is, the Cas9 is a cleavage enzyme, referred to as the “nCas9” protein (as “cleavage enzyme” Cas9). The nuclease-inactivated Cas9 protein can be interchangeably referred to as the “dCas9” protein (for nuclease-dead Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains are known (e.g., see Jinek et al., Science., 337: 816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene”). expression."(2013)Cell, 28; 152(5): 1173-83, the entire contents of each of which are incorporated herein by reference. For example, the DNA cleavage domain of Cas9 is known to consist of two subdomains, namely the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science, 337: 816-821(2012); Qi et al., Cell., 28; 152(5): 1173). -83 (2013)). In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) a gRNA-binding domain of Cas9; or (2) a DNA-cutting domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant is homologous to Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to wild-type Cas9.In some implementations, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid variations compared to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA-binding domain or a DNA-cutting domain) such that the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to a corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a corresponding fragment of wild-type Cas9.

[0143] In some embodiments, the fragment is at least 100 amino acids long. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids long. In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows):

[0144]

[0145]

[0146] In some implementations, wild-type Cas9 corresponds to or contains the following nucleotide and / or amino acid sequences:

[0147]

[0148]

[0149] In some implementations, wild-type Cas9 corresponds to Cas9 of Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (nucleotide sequence below); and Uniprot reference sequence: Q99ZW2 (amino acid sequence below):

[0150]

[0151]

[0152]

[0153] In some implementations, the CRISPR protein-derived domain of the base editor may include those from *Corynebacterium ulcerans* (NCBI Refs: NC_015683.1, NC_017317.1); *Corynebacterium diphtheria* (NCBI Refs: NC_016782.1, NC_016786.1); *Spiroplasma syrphidicola* (NCBI Ref: NC_021284.1); *Prevotella intermedia* (NCBI Ref: NC_017861.1); *Spiroplasma taiwanense* (NCBI Ref: NC_021846.1); and *Streptococcus iniae* (NCBI Ref: NC_021846.1). Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_002342100.1); all or part of Cas9 of Streptococcus pyogenes or Staphylococcus aureus.

[0154] In some embodiments, dCas9 corresponds to or contains, partially or entirely, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain contains the D10A and H840A mutations or corresponding mutations in another Cas9. In some embodiments, dCas9 contains the amino acid sequence of dCas9 (D10A and H840A):

[0155]

[0156] In some implementations, the Cas9 domain contains the D10A mutation, while the residue at position 840 remains histidine at the corresponding position in the amino acid sequence provided above or in any amino acid sequence provided herein.

[0157] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided, such as Cas9 (dCas9) that leads to nuclease inactivation. For example, such mutations include other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical. In some embodiments, the provided dCas9 variants have the following amino acid sequences, which are shorter or longer than about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more.

[0158] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of the Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not comprise the full-length Cas9 sequence, but only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and other suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.

[0159] The Cas9 protein can associate with a guide RNA, which directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is the Cas9 domain, such as active Cas9, Cas9 nickase (nCas9), or inactive Cas9 (dCas9). Examples of nucleic acid-programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3.

[0160] Nuclease-inactivated Cas9 proteins can be interchangeably referred to as “dCas9” proteins (for nuclease-dead Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with inactivated DNA-cutting domains are known (e.g., see Jinek et al., Science. 337: 816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene”). The entire contents of each of these expressions are incorporated herein by reference. (2013) Cell., 28; 152(5): 1173-83. For example, the DNA cleavage domain of Cas9 is known to consist of two subdomains, namely the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science., 337: 816-821 (2012); Qi et al., Cell., 28; 152(5): 1173-83 (2013)).

[0161] In some embodiments, the Cas9 domain is a Cas9 cleavage enzyme. The Cas9 cleavage enzyme can be a Cas9 protein capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 cleavage enzyme cleaves the target strand of the double-stranded nucleic acid molecule, meaning that the Cas9 cleavage enzyme cleaves the strand that is base-paired (complementary) to the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 cleavage enzyme includes a D10A mutation and has a histidine residue at position 840. In some embodiments, the Cas9 cleavage enzyme cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule, meaning that the Cas9 cleavage enzyme cleaves the strand that is not base-paired to the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 cleavage enzyme contains an H840A mutation and has an aspartic acid residue at position 10 or a corresponding mutation. In some embodiments, the Cas9 nicking enzyme contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas9 nicking enzyme provided herein. Other suitable Cas9 nicking enzymes will be apparent to those skilled in the art based on this application and the knowledge of the art, and are within the scope of this application.

[0162] In some embodiments, the Cas9 domain is a nuclease-free Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-free dCas9 domain contains the D10X and H840X mutations of the amino acid sequences listed herein, or the corresponding mutations in any amino acid sequences provided herein, where X is any amino acid change. In some embodiments, the nuclease-free dCas9 domain contains the D10A and H840A mutations of the amino acid sequences listed herein, or the corresponding mutations in any amino acid sequences provided herein. As an example, the nuclease-free Cas9 domain contains the amino acid sequence listed in the cloning vector pPlatTET-gRNA2 (accession number BAV54124):

[0163]

[0164] It should be understood that other Cas9 proteins, including their variants and homologs (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), include their variants and homologs. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-dead Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.

[0165] Exemplary catalytic deactivation of Cas9 (dCas9):

[0166]

[0167] Exemplary catalytic Cas9 nickase (nCas9):

[0168]

[0169] Exemplary catalytically active Cas9:

[0170]

[0171] In some embodiments, Cas9 refers to Cas9 derived from archaea (e.g., nanoarchaea), which constitute domains and kingdoms of single-celled prokaryotic microorganisms. In some embodiments, the programmable nucleotide-binding protein may be a CasX or CasY protein, which has been described, for example, in Burstein et al., “New CRISPR-Cas systems from uncultivated microbes.”, Cell Res, February 21, 2017, doi: 10.1038 / cr.2017.21, the entire contents of which are incorporated herein by reference. Many CRISPR-Cas systems have been identified using genome-resolved metagenomics, including Cas9, which was first reported in the Archaea domain of life. This differentiated Cas9 protein was found in the little-studied nanoarchaea and is part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered, which are the most compact systems discovered to date. In some embodiments, in the base editor system described herein, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor system described herein, Cas9 is replaced by CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid-programmable DNA-binding proteins (napDNAbp) and are within the scope of this application.

[0172] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a CasX or CasY protein. In some embodiments, napDNAbp is a CasX protein. In some embodiments, napDNAbp is a CasY protein. In some embodiments, napDNAbp contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to that of a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein is a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein contains an amino acid sequence that has at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any CasX or CasY protein described herein. It should be understood that, according to this application, CasX and CasY from other bacterial species may also be used.

[0173] An example of CasX((uniprot.org / uniprot / F0NN87);

[0174] The amino acid sequence of the CAX protein associated with uniprot.org / uniprot / F0NH53)tr|F0NN87|F0NN87_SULIHCRISPR is as follows: OS = Icelandic Sulfur Leaf Bacterium (strain HVE10 / 4) GN = SiH_0402 PE = 4SV = 1.

[0175] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLE VEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0176] >tr|F0NH53|F0NH53_SULIR CRISPR-related protein, Casx OS=Icelandic sulfur leaf (strain REY15A)GN=SiRe_0771PE=4SV=1) amino acid sequence is as follows:

[0177] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAK VSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0178] δ-Proteobacterium CasX

[0179] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQ KWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA

[0180] An exemplary amino acid sequence of CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR-associated protein CasY [uncultured Saccharomyces spp.]) is as follows:

[0181]

[0182] It should be understood that polynucleotide-programmable nucleotide-binding domains can also include nucleic acid-programmable proteins that bind to RNA. For example, a polynucleotide-programmable nucleotide-binding domain can bind to a nucleic acid that guides the polynucleotide-programmable nucleotide-binding domain to RNA. Other nucleic acid-programmable DNA-binding proteins are also within the scope of this application, although they are not specifically listed herein.

[0183] The Cas proteins that can be used in this document include classes 1 and 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, and C... mr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h and Cas12i, CARF, DinG, their homologs or modified forms. Unmodified CRISPR enzymes can possess DNA cleavage activity, such as Cas9, which has two functional endonuclease domains: RuvC and HNH. CRISPR enzymes can be directed to cleave one or both strands at a target sequence, such as within the target sequence and / or within its complementary sequence. For example, a CRISPR enzyme can direct the cleavage of one or both strands at approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 base pairs from the first or last nucleotide of the target sequence.

[0184] Vectors encoding CRISPR enzymes mutated relative to the corresponding wild-type enzymes can be used, causing the mutated CRISPR enzymes to lack the ability to cleave one or both strands of a target polynucleotide containing the target sequence. Cas9 can refer to a polypeptide having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can also refer to a polypeptide having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to the wild-type or modified form of the Cas9 protein, which may include amino acid changes such as deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof.

[0185] In some implementations, the methods described herein can utilize engineered Cas proteins. The guide RNA (gRNA) is a short, synthetic RNA formed by Cas binding to the desired scaffold sequence and a user-defined... It consists of nucleotide spacer regions that define the genomic target to be modified. Therefore, technicians can alter the genomic target of the Cas protein, which is partially determined by the specificity of the gRNA targeting sequence to the genomic target compared to the rest of the genome.

[0186] The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. Upon target binding, Cas9 undergoes a second conformational change that positions the nuclease domains to cleave the opposite strand of the target DNA. The final result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired via one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the less efficient but high-fidelity homology-guided repair (HDR) pathway.

[0187] The “efficiency” of non-homologous end joining (NHEJ) and / or homologous guided repair (HDR) can be calculated using any convenient method. For example, in some cases, efficiency can be expressed as a percentage of successful HDR. For instance, a surveyor nuclease assay can be used to generate the cleavage product, and the percentage can be calculated as the ratio of product to plasmid. For example, as a result of successful HDR, a surveyor nuclease can be used to directly cleave DNA containing the newly integrated restriction sequence. The more plasmid cleaved, the higher the HDR percentage (higher HDR efficiency). As an illustrative example, the HDR fraction (percentage) can be calculated using the following equation [(cleavage product) / (platinum plus cleavage product)] (e.g., (b+c) / (a+b+c), where “a” is the band intensity of the DNA plasmid, and “b” and “c” are the cleavage products).

[0188] In some cases, efficiency can be expressed as the success rate of NHEJ. For example, the T7 endonuclease I assay can be used to generate cleavage products, and the product-to-recipient ratio can be used to calculate the percentage of NHEJ. The T7 endonuclease cleaves mismatched heteroduplex DNA caused by hybridization of wild-type and mutant DNA strands (NHEJ produces small random insertions or deletions (indels) at the original break site). A higher degree of cleavage indicates a higher percentage of NHEJ (higher NHEJ efficiency). As an illustrative example, the NHEJ score (percentage) can be calculated using the following equation: (1-(1-(b(c+c) / (a+b+c))1 / 2)×100, where “a” is the band intensity of the DNA plasmid, and “b” and “c” are the cleavage products (Ran et al., Cell, Sep 12, 2013; 154(6): 1380-9; Ran et al., Nat Protoc. Nov 2013; 8(11): 2281–2308).

[0189] The NHEJ repair pathway is the most active repair mechanism, frequently inducing small nucleotide insertions or deletions (indels) at DSB sites. The stochastic nature of NHEJ-mediated DSB repair has significant practical implications because cellular populations expressing Cas9 and gRNA or guide polynucleotides can lead to a wide variety of mutations. In most cases, NHEJ produces small insertions in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations, leading to premature stop codons within the target gene's open reading frame (ORF). The ideal end result is a loss-of-function mutation within the target gene.

[0190] Although NHEJ-mediated DSB repair often disrupts open reading frames of genes, homology-guided repair (HDR) can be used to produce specific nucleotide changes, ranging from single nucleotide changes to large insertions, such as the addition of fluorophores or tags.

[0191] To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered to the cell type of interest along with gRNA and Cas9 or Cas9 nickase. The repair template can contain the desired edit along with other homologous sequences immediately upstream and downstream of the target (called left and right homologous arms). The length of each homologous arm can vary depending on the size of the introduced alteration, with larger insertions requiring longer homologous arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and an exogenous repair template, HDR efficiency is typically low (<10% of modified alleles). Since HDR occurs at the S and G2 phases of the cell cycle, its efficiency can be improved by synchronizing the cell. Chemical or genetic repressor genes involved in NHEJ can also increase HDR frequency.

[0192] In some implementations, Cas9 is a modified Cas9. A given gRNA target sequence may have partially homologous other sites throughout the genome. These sites are called off-target sites and need to be considered when designing gRNAs. Besides optimizing gRNA design, CRISPR specificity can be improved by modifying Cas9. Cas9 generates double-strand breaks (DSBs) through the combined activity of its two nuclease domains, RuvC and HNH. The Cas9 nickase is a D10A mutant of SpCas9 that retains one nuclease domain and generates a DNA nick instead of a DSB. Nickase systems can also be combined with HDR-mediated gene editing for specific gene editing purposes.

[0193] In some cases, Cas9 is a variant of the Cas9 protein. When compared to the amino acid sequence of the wild-type Cas9 protein, the variant Cas9 peptide differs by one amino acid (e.g., by deletion, insertion, substitution, or fusion). In some cases, the variant Cas9 peptide has amino acid changes that reduce the nuclease activity of the Cas9 peptide (e.g., deletion, insertion, or substitution). For example, in some cases, the variant Cas9 peptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the corresponding wild-type Cas9 protein activity. In some cases, the variant Cas9 protein has no substantial nuclease activity. When the subject Cas9 protein is a variant Cas9 protein without substantial nuclease activity, it can be referred to as "dCas9".

[0194] In some cases, variant Cas9 proteins exhibit reduced nuclease activity. For example, variant Cas9 proteins exhibit less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of endonuclease activity. Wild-type Cas9 proteins, such as wild-type Cas9 proteins.

[0195] In some cases, variant Cas9 proteins can cleave the complementary strand of the guide target sequence, but have a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. For example, variant Cas9 proteins may have mutations (amino acid substitutions) that reduce the function of the RuvC domain. As a non-limiting example, in some embodiments, variant Cas9 proteins have D10A (aspartic acid of glutamic acid at amino acid position 10), thus enabling them to cleave the complementary strand of the double-stranded guide target sequence, but reducing their ability to cleave the non-complementary strand of the double-stranded guide target sequence (therefore, when variant Cas9 proteins cleave double-stranded target nucleic acids, it results in single-strand breaks (SSBs) instead of double-strand breaks (DSBs)) (see, for example, Jinek et al., Science. 2012 Aug. 17; 337(6096): 816-21).

[0196] In some cases, variant Cas9 proteins can cleave the non-complementary strand of a double-stranded guide target sequence, but have a reduced ability to cleave the complementary strand of the guide target sequence. For example, variant Cas9 proteins may have mutations (amino acid substitutions) that reduce the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, variant Cas9 proteins have the H840A mutation (histidine to alanine at amino acid position 840), and thus can cleave the non-complementary strand of the guide target sequence, but have a reduced ability to cleave the complementary strand of the guide target sequence (therefore, when the variant Cas9 protein cleaves the double-stranded guide target sequence, an SSB will be produced instead of a DSB). Such Cas9 proteins have a reduced ability to cleave guide target sequences (e.g., single-stranded guide target sequences), but retain the ability to bind to guide target sequences (e.g., single-stranded guide target sequences).

[0197] In some cases, the ability of variant Cas9 proteins to cleave both the complementary and non-complementary strands of double-stranded target DNA is reduced. As a non-limiting example, in some cases, variant Cas9 proteins possess both the D10A and H840A mutations, resulting in a reduced ability of the polypeptide to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0198] As another non-limiting example, in some cases, variant Cas9 proteins have mutations such as W476A and W1126A, resulting in a reduced ability of the peptide to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0199] As another non-limiting example, in some cases, variant Cas9 proteins possess mutations such as P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A, resulting in a reduced ability of the peptide to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0200] As another non-limiting example, in some cases, the variant Cas9 protein has mutations of H840A, W476A, and W1126A, resulting in a reduced ability of the peptide to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein carries mutations of H840A, D10A, W476A, and W1126A, resulting in a reduced ability of the peptide to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 has restored the catalytic His residue at position 840 in the Cas9 HNH domain (A840H).

[0201] As another non-limiting example, in some cases, variant Cas9 proteins carry mutations of H840A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A, resulting in a reduced ability of the peptide to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, variant Cas9 proteins have mutations of D10A, H840A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A, resulting in a reduced ability of the peptide to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the mutated Cas9 protein carries mutations such as W476A and W1126A, or mutations such as P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A, the mutated Cas9 protein cannot effectively bind to the PAM sequence. Therefore, in some such cases, when this variant Cas9 protein is used in a binding method, the method does not require the PAM sequence. In other words, in some cases, when this Cas9 variant protein is used in a binding method, the method may include a guide RNA, but the method can be performed in the absence of the PAM sequence (therefore, the binding specificity is provided by the targeting fragment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., partially inactivating one or another nuclease). As non-restrictive examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). Similarly, mutations other than alanine substitution are also acceptable.

[0202] In some implementations, variant Cas9 proteins with reduced catalytic activity (e.g., when the Cas9 protein has mutations such as D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, such as D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A) can still bind to the target DNA in a site-specific manner (because as long as the Cas9 variant protein retains the ability to interact with the guide RNA, it can still be guided by the guide RNA to the target DNA sequence).

[0203] Alternatives to Cas9 for Streptococcus pyogenes can include RNA-directed endonucleases from the Cpf1 family, which exhibit cleavage activity in mammalian cells. CRISPR (CRISPR / Cpf1) from Prevotella and Francisella 1 is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-directed endonuclease in the class II CRISPR / Cas system. This acquired immune mechanism has been found in Prevotella and Francisella. The Cpf1 gene is associated with a CRISPR locus and encodes an endonuclease that uses guide RNA to locate and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some limitations of the CRISPR / Cas9 system. Unlike Cas9 nucleases, Cpf1-mediated DNA cleavage results in double-strand breaks with short 3' overhangs. The staggered cutting pattern of Cpf1 opens up possibilities for directional gene transfer, similar to traditional restriction endonuclease cloning, which can improve gene editing efficiency. Like the aforementioned Cas9 variants and orthogonal homologs, Cpf1 can also expand the number of sites in CRISPR-targeted AT-rich regions or NGG PAM sites in the AT genome lacking SpCas9 support. The Cpf1 locus contains a hybrid α / β domain, followed by a helical region, a RuvC-II, and a zinc finger domain (RuvC-I). The Cpf1 protein possesses a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, Cpf1 lacks an HNH endonuclease domain, and its N-terminus lacks the α-helical recognition leaflet of Cas9. The Cpf1 CRISPR-Cas domain structure reveals that Cpf1 is functionally unique and is classified as a class 2, type V CRISPR system. The Cas1, Cas2, and Cas4 proteins encoded by the Cpf1 locus are more similar to the type II system compared to type I and type III. Functional Cpf1 does not require trans-activation CRISPR RNA (tracrRNA), thus only CRISPR (crRNA) is needed. This is advantageous for genome editing because Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (approximately half the size of Cas9). Compared to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex cleaves target DNA or RNA by recognizing a protospacer adjacent to the 5'-YTN-3' region. Upon PAM identification, Cpf1 introduces a sticky end-like DNA double-strand break with a 4 or 5 nucleotide overhang.

[0204] Some aspects of this application provide fusion proteins comprising domains that act as nucleic acid-programmable DNA-binding proteins, which can be used to guide proteins (e.g., base editors) to specific nucleic acid (e.g., DNA or RNA) sequences. In certain embodiments, the fusion protein comprises a nucleic acid-programmable DNA-binding protein domain and a deaminase domain. DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. An example of a programmable polynucleotide-binding protein with PAM specificity different from Cas9 is clusters of regularly spaced short palindromic repeats from *Prevotella* and *Francisella* 1 (Cpf1). Similar to Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate strong DNA interference with characteristics different from Cas9. Cpf1 is an RNA-guided endonuclease lacking tracrRNA, which utilizes a T-rich motif from a neighboring protospacer (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA via staggered double-strand breaks. Of the 16 Cpf1 family proteins, two enzymes from *Acidococcus pyogenes* and *Helicobacter pylori* have shown potent genome editing activity in human cells. Cpf1 proteins are known in the art and have been previously described, for example, in Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell. (165) 2016, pp. 949-962; the entire contents of which are incorporated herein by reference.

[0205] Also useful in the compositions and methods of this application is a nuclease-free Cpf1 (dCpf1) variant, which can be used as a multinucleotide-binding protein domain for guiding nucleotide sequence programmability. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9, but lacks the HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition leaf of Cas9. As shown in Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference), the RuvC-like domain of Cpf1 is responsible for DNA strand cleavage and inactivation of the RuvC-like domain, thus inactivating Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in Novida Cpf1 inactivate Cpf1 nuclease activity. In some embodiments, the dCpf1 of this application includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that any mutation, such as substitution mutations, deletions, or insertions that inactivate the RuvC domain of Cpf1, can be used according to this application.

[0206] In some embodiments, the nucleic acid-programmable nucleotide-binding protein of any fusion protein provided herein may be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 cleavage enzyme (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactivated Cpf1 (dCpf1). In some embodiments, Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with the Cpf1 sequence disclosed herein. In some embodiments, dCpf1 comprises an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity with the Cpf1 sequence disclosed herein, and comprises a mutation corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, or E1006A / D1255A. It should be understood that Cpf1 from other bacterial species may also be used according to this application.

[0207] The amino acid sequence of wild-type *Francisella novicida* Cpf1 is as follows. D917, E1006, and D1255 are in bold and underlined.

[0208] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0209] The following is the amino acid sequence of *Neofrancella* Cpf1 D917A. (A917, E1006, and D1255 are indicated in bold and underlined).

[0210] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0211] The following is the amino acid sequence of Neofrancella Cpf1 E1006A (D917, A1006 and D1255 are marked in bold and underline).

[0212] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0213] The following is the amino acid sequence of Francisella Cpf1 D1255A (the mutation positions of D917, E1006 and A1255 are shown in bold with underline).

[0214] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0215] The amino acid sequence of *Neofrancella* Cpf1 D917A / E1006A is as follows. (A917, A1006, and D1255 are indicated in bold and underlined).

[0216] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0217] The following is the amino acid sequence of *Neofrancella* Cpf1 D917A / D1255A. (A917, E1006, and A1255 are indicated in bold and underlined).

[0218] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0219] The amino acid sequence of *NeoFrancis* Cpf1 E1006A / D1255A is as follows (D917, A1006, and A1255 are indicated in bold and underline).

[0220] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0221] The following is the amino acid sequence of *Neofrancella* Cpf1 D917A / E1006A / D1255A. (A917, A1006, and A1255 are indicated in bold and underlined).

[0222] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0223] In some implementations, one of the Cas9 domains present in the fusion protein can be replaced by a guide nucleotide sequence programmable DNA-binding protein domain that does not require a PAM sequence.

[0224] In some implementations, a nucleic acid programmable DNA-binding protein (napDNAbp) is a single effector of a microbial CRISPR-Cas system. Single effectors of microbial CRISPR-Cas systems include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Generally, microbial CRISPR-Cas systems are classified into Class 1 and Class 2 systems. Class 1 systems have multi-subunit effector complexes, while Class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are Class 2 effectors. In addition to Cas9 and Cpf1, Shmakov et al., “Discovery and Functional Characterization of DiverseClass 2 CRISPR Cas Systems,” Mol. Cell, November 5, 2015; 60(3): 385-397, also describe three different Class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3), the entire contents of which are incorporated herein by reference. The effectors Cas12b / C2c1 and Cas12c / C2c3 of the two systems contain RuvC-like endonuclease domains associated with Cpf1. The third system contains an effector with two predicate HEPN RNase domains. Unlike the CRISPR RNA produced by Cas12b / C2c1, the production of mature CRISPR RNA is independent of tracrRNA. Cas12b / C2c1 depends on DNA cleavage by both CRISPR RNA and tracrRNA.

[0225] Crystal structures of *Bifidobacterium cas12b / C2c1* (AacC2c1) complexed with chimeric single-molecule guide RNA (sgRNA) have been reported. See, for example, Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism,” *Mol. Cell*, January 19, 2017; 65(2): 310-322, the entire contents of which are incorporated herein by reference. Crystal structures of *Bifidobacterium cas2c1* bound to target DNA as a ternary complex have also been reported. See, for example, Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1CRISPR-Cas endonuclease,” *Cell*, December 15, 2016; 167(7): 1814-1828, the entire contents of which are incorporated herein by reference. The catalytically active conformation of AacC2c1, containing both target and non-target DNA strands, was independently captured and placed in a single RuvC catalytic pocket. Cas12b / C2c1-mediated cleavage resulted in the target DNA being misaligned by seven nucleotides. Structural comparisons of the Cas12b / C2c1 ternary complex with previously identified Cas9 and Cpf1 counterparts demonstrate the diversity of mechanisms employed by the CRISPR-Cas9 system.

[0226] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, napDNAbp is a Cas12b / C2c1 protein. In some embodiments, napDNAbp is a Cas12c / C2c3 protein. In some embodiments, napDNAbp contains at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as naturally occurring Cas12b / C2c1 or Cas12c / C2c3 proteins. In some embodiments, napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp contains at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence identical to any napDNAbp sequence provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used according to this application.

[0227] A Cas12b / C2c1((uniprot.org / uniprot / T0D7A2#2)

[0228] sp|T0D7A2|C2C1_ALIAG CRISPR-related endonuclease C2c1 OS=Acidophilus cyclophosphamide (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1 SV=1) The amino acid sequence is as follows:

[0229]

[0230] BhCas12b (Histidine Bacillus) NCBI reference sequence: WP_095142515

[0231]

[0232] In some implementations, Cas12b is BvCas12B, a variant of BhCas12b, and includes the following variations relative to BhCas12B: S893R, K846R, and E837G. BvCas12b (Bacillus V3-13) NCBI reference sequence: WP_101661451.1

[0233]

[0234] In some embodiments, the Cas9 domain is a Cas9 domain derived from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is an active nuclease SaCas9, an inactive nuclease SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 domain contains an N579A mutation or a corresponding mutation in any amino acid sequence provided herein.

[0235] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain may bind to nucleic acid sequences having atypical PAM sequences. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain may bind to nucleic acid sequences having NNGRRT or NNNRRT PAM sequences. In some embodiments, the SaCas9 domain contains one or more of the E781X, N967X, and R1014X mutations or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain contains one or more of the E781K, N967K, and R1014H mutations, or one or more corresponding mutations, in any amino acid sequence provided herein. In some embodiments, the SaCas9 domain contains the E781K, N967K, or R1014H mutation or corresponding mutations in any amino acid sequence provided herein.

[0236] In some implementations, the variant Cas protein may be SpCas9, SpCas9-VRQR, SpCas9-VRER, xCas9(sp), SaCas9, SaCas9-KKH, SpCas9-MQKSER, SpCas9-LRKIQK, or SpCas9-LRVSQL.

[0237] An example amino acid sequence of SaCas9 is as follows:

[0238] KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQ KKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRKLVPKKVDLSQQK EIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE NSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGY KHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKL KKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVI KKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG

[0239] In the sequence, residue N579, which is underlined and bolded, can... through Mutation (e.g., mutation to A579) to produce SaCas9 nickase.

[0240] An exemplary amino acid sequence of SaCas9n is as follows:

[0241] KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQ KKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRKLVPKKVDLSQQK EIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE ASKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGY KHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKL KKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVI KKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG

[0242] In the sequence, residue N579, which is underlined and bolded, can... through Mutation (e.g., mutation to A579) to produce SaCas9 nickase.

[0243] An exemplary amino acid sequence of SaKKH Cas9 is as follows:

[0244] KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQ KKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRKLVPKKVDLSQQK EIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEITEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNR K LINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFY KNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPP H IIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG.

[0245] The residue A579 above is underlined and bolded, which can be mutated from N579 to produce the SaCas9 nickase. The residues K781, K967, and H1014 above, which can be mutated from E781, N967, and R1014 to produce SaKKH Cas9, are underlined and italicized.

[0246] The polynucleotide-programmable nucleotide-binding domain of the base editor may itself contain one or more domains. For example, the polynucleotide-programmable nucleotide-binding domain may contain one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide-binding domain may contain an endonuclease or an exonuclease. In this document, the term "exonuclease" refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from their free ends, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) internal regions within nucleic acids (e.g., DNA or RNA). In some embodiments, the endonuclease can cleave a single strand of a double-stranded nucleic acid. In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be a ribonuclease.

[0247] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can cleave zero, one, or both strands of the target polynucleotide. In some cases, the polynucleotide programmable nucleotide binding domain may include a nicking enzyme domain. Hereinafter, the term "nicking enzyme" refers to a polynucleotide programmable nucleotide binding domain containing a nuclease domain capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nicking enzyme can be derived from the fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide binding domain by introducing one or more mutations into the active polynucleotide programmable nucleotide binding domain. For example, in the case where the polynucleotide programmable nucleotide binding domain includes a Cas9-derived nicking enzyme domain, the Cas9-derived nicking enzyme domain may contain a D10A mutation and a histidine residue at position 840. In this case, residue H840 retains catalytic activity and can therefore cleave a single strand of the nucleic acid duplex. In another example, the Cas9-derived nicking enzyme domain may contain an H840A mutation, while the amino acid residue at position 10 remains D. In some implementations, the nicking enzyme may be derived from the fully catalytically active (e.g., native) form of the polynucleotide. A programmable nucleotide-binding domain is achieved by removing all or part of the nuclease domains that are not required for the nicking enzyme activity. For example, in cases where the programmable nucleotide-binding domain of the polynucleotide includes a nicking enzyme domain derived from Cas9, the Cas9-derived nicking enzyme domain may contain all or part of the RuvC domain or the HNH domain.

[0248] A base editor comprising a polynucleotide programmable nucleotide-binding domain includes a nuclease domain, which is capable of generating single-strand DNA breaks (nicks) on a specific polynucleotide target sequence (e.g., determined by the complementary sequence of the binding guide nucleic acid). In some embodiments, the strand of the target polynucleotide sequence of the nucleic acid duplex cleaved by a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) is the unedited strand (i.e., the fragment cleaved by the base editor is the opposite of the strand containing the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) can cleave the strand targeted to be edited in the DNA molecule. In this case, the non-target strand is not cleaved.

[0249] This document also provides a base editor comprising a polynucleotide programmable nucleotide-binding domain that is catalytically dead (i.e., unable to cleave the target polynucleotide sequence). In this document, the terms "catalytic death" and "nuclease death" are used interchangeably to refer to a polynucleotide programmable nucleotide-binding domain having one or more mutations and / or deletions that prevent it from cleaving a nucleic acid chain. In some embodiments, the catalytically dead polynucleotide programmable nucleotide-binding domain base editor may lack nuclease activity due to specific point mutations in one or more nuclease domains. For example, in the case where the base editor comprises a Cas9 domain, Cas9 may contain both the D10A and H840A mutations. Such mutations inactivate both nuclease domains, resulting in a loss of nuclease activity. In other embodiments, the catalytically dead polynucleotide programmable nucleotide-binding domain may comprise one or more deletions of all or part of a catalytic domain (e.g., a RuvC1 and / or HNH domain). In a further embodiment, the catalytically dead polynucleotide programmable nucleotide-binding domain comprises a point mutation (e.g., D10A or H840A) and all or part of a nuclease domain deletion.

[0250] This paper also considers mutations in polynucleotide programmable nucleotide-binding domains that can generate catalytic death from previous functional forms of the polynucleotide programmable nucleotide-binding domain. For example, in the case of catalytic death Cas9 (“dCas9”), variants with mutations other than D10A and H840A are provided, resulting in Cas9 with nuclease inactivation. Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain).

[0251] Based on this application and the knowledge of the art, other suitable nuclease-inactivating dCas9 domains will be apparent to those skilled in the art and are within the scope of this application. Such additional exemplary suitable nuclease-inactivating Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, for example, Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology., 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference). In some embodiments, the amino acid sequence contained in the dCas9 domain has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any dCas9 domain provided herein. In some embodiments, the Cas9 domain contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences listed herein. In some embodiments, the Cas9 domain comprises, compared to any amino acid sequence listed herein, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues.

[0252] Non-limiting examples of polynucleotide-programmable nucleotide-binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, large-scale nucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some cases, base editors contain polynucleotide-programmable nucleotide-binding domains that comprise a natural or modified protein or a portion thereof, which, through binding to a guiding nucleic acid, is capable of binding nucleic acid sequences (i.e., clusters of regularly spaced short palindromic repeats) mediated by nucleic acid modifications during the CRISPR process. Such proteins are referred to herein as “CRISPR proteins.” Therefore, this document discloses base editors that comprise polynucleotide-programmable nucleotide-binding domains comprising all or a portion of a CRISPR protein (i.e., base editors comprising all or a portion of a CRISPR protein as a domain, also referred to as “CRISPR protein-derived domains” of base editors). The domains of CRISPR proteins incorporated into base editors can be modified compared to wild-type or natural versions of CRISPR proteins. For example, as described below, domains derived from CRISPR proteins may contain one or more mutations, insertions, deletions, rearrangements, and / or recombinations, relative to wild-type or natural forms of CRISPR proteins.

[0253] In some embodiments, the CRISPR-derived domain incorporated into the base editor is a nuclease (e.g., deoxyribonuclease or ribonuclease) capable of binding the target polynucleotide when bound to a guiding nucleic acid. In some embodiments, the CRISPR-derived domain incorporated into the base editor is a cleavage enzyme capable of binding the target polynucleotide when bound to a guiding nucleic acid. In some embodiments, the CRISPR-derived domain incorporated into the base editor is a catalytic death domain capable of binding the target polynucleotide when bound to a guiding nucleic acid. In some embodiments, the target polynucleotide binding to the CRISPR-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide binding to the CRISPR-derived domain of the base editor is RNA.

[0254] In some implementations, the CRISPR protein-derived domain of the base editor may include those from *Corynebacterium ulcerans* (NCBI Refs: NC_015683.1, NC_017317.1); *Corynebacterium diphtheria* (NCBI Refs: NC_016782.1, NC_016786.1); *Spiroplasma syrphidicola* (NCBI Ref: NC_021284.1); *Prevotella intermedia* (NCBI Ref: NC_017861.1); *Spiroplasma taiwanense* (NCBI Ref: NC_021846.1); and *Streptococcus iniae* (NCBI Ref: NC_021846.1). Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_002342100.1); all or part of Cas9 of Streptococcus pyogenes or Staphylococcus aureus.

[0255] In some embodiments, the Cas9-derived domain of the base editor is a Cas9 domain derived from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is an active SaCas9 nuclease, an inactive SaCas9 nuclease (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 domain contains the N579X mutation. In some embodiments, the SaCas9 domain contains the N579A mutation. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to nucleic acid sequences with atypical PAM sequences. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to nucleic acid sequences with the NNGRRT PAM sequence. In some embodiments, the SaCas9 domain contains one or more of the E781X, N967X, and R1014X mutations.

[0256] Base editors may include domains derived entirely or partially from Cas9, specifically high-fidelity Cas9. In some embodiments, the high-fidelity Cas9 domain of the base editor is an engineered Cas9 domain containing one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA relative to the corresponding wild-type Cas9 domain. High-fidelity Cas9 domains with reduced electrostatic interaction with the sugar-phosphate backbone of DNA may have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., the wild-type Cas9 domain) contains one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some implementations, the Cas9 domain contains one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, or more.

[0257] In some implementations, the modified Cas9 is a high-fidelity Cas9 enzyme. In some implementations, the high-fidelity Cas9 enzyme is SpCas9 (K855A), eSpCas9 (1.1), SpCas9-HF1, or a highly precise Cas9 variant (HypaCas9). The modified Cas9 eSpCas9 (1.1) contains an alanine substitution that weakens the interaction between the HNH / RuvC groove and the non-target DNA strand, thereby preventing strand segregation and cleavage at off-target sites. Similarly, SpCas9-HF1 reduces off-target editing through an alanine substitution that disrupts the interaction between Cas9 and the DNA phosphate backbone. HypaCas9 contains mutations in the REC3 domain (SpCas9N692A / M694A / Q695A / H698A) that increase Cas9 proofreading and target recognition. All three high-fidelity enzymes produce fewer off-target edits compared to wild-type Cas9. Exemplary high-fidelity Cas9s are provided below.

[0258] High-fidelity Cas9 domain mutations compared to Cas9 are shown in bold and underlined.

[0259] MDKKYSIGL A IGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEK YPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFK SNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYK FIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMT AFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYALFDDKVMKQLKRRRYTGWG A LSRKLINGIRDKQSGKTILDFLKSDGFANRNFM A LIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETR A ITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVA KVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILDANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD

[0260] guidance

[0261] As used herein, the term "(one or more) guide polynucleotides" refers to a polynucleotide that is specific to a target sequence and can form a complex with a polynucleotide programmable nucleotide binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is guide RNA. As used herein, the term "guide RNA (gRNA)" and its grammatical equivalents can refer to RNA that is specific to target DNA and can form a complex with a Cas protein. The RNA / Cas complex helps to "guide" the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA endonuclease cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to the crRNA is first endonucleated and then exonucleated 3'-5'. In practice, DNA binding and cleavage typically require both a protein and two RNAs. However, single-stranded guide RNAs ("sgRNAs" or simply "gNRAs") can be engineered to integrate various aspects of both crRNA and tracrRNA into a single RNA species. For example, see Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 identifies short motifs (PAM or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish between itself and non-self. The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, for example, “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti, JJ et al., Natl. Acad. Sci. USA 98: 4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and hostfactor RNase III.” Deltcheva E. et al., Nature 471: 602-607 (2011); and Jinek M. et al., “Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.”, Science 337: 816-821 (2012), the entire contents of which are incorporated herein by reference). Cas9 orthologs have been described in a variety of species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus.Based on the disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include those from Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type IICRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a cleavage enzyme.

[0262] In some embodiments, the guiding polynucleotide is at least one single-stranded guide RNA (“sgRNA” or “gNRA”). In some embodiments, the guiding polynucleotide is at least one tracrRNA. In some embodiments, the guiding polynucleotide does not require a PAM sequence to guide the polynucleotide-programmable DNA-binding domain (e.g., Cas9 or Cpf1) to the target nucleotide sequence.

[0263] The multinucleotide programmable nucleotide-binding domains (e.g., CRISPR-derived domains) of the base editor disclosed herein can recognize target polynucleotide sequences by binding to a guide polynucleotide. The guide polynucleotide (e.g., gRNA) is typically single-stranded and can be programmed to specifically bind to the target sequence site of the polynucleotide (i.e., through complementary base pairing), thereby guiding the nucleic acid of the base editor target sequence bound to the guide polynucleotide. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some cases, the guide polynucleotide contains a natural nucleotide (e.g., adenosine). In some cases, the guide polynucleotide contains a non-natural (or non-natural) nucleotide (e.g., peptide nucleic acid or nucleotide analogue). In some cases, the target region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The length of the target region of the guide nucleic acid can be between 10 and 30 nucleotides, or between 15 and 25 nucleotides, or between 15 and 20 nucleotides.

[0264] In some embodiments, the guiding polynucleotide comprises two or more separate polynucleotides that can interact with each other, for example, through complementary base pairing (e.g., double-stranded guiding polynucleotide). For example, the guiding polynucleotide may comprise CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA). Alternatively, the guiding polynucleotide may comprise one or more trans-activating CRISPR RNAs (tracrRNA).

[0265] In type II CRISPR systems, targeting nucleic acids via a CRISPR protein (e.g., Cas9) typically requires complementary base pairing between a first RNA molecule (crRNA) containing a sequence that recognizes the target sequence and a second RNA molecule (trRNA) containing a repetitive sequence that forms a scaffold region that stabilizes the guide RNA-CRISPR protein complex. Such a double-stranded guide RNA system can be used to guide polynucleotides to direct the base editor disclosed herein to the target polynucleotide sequence.

[0266] In some embodiments, the base editor provided herein utilizes a single-stranded guide polynucleotide (e.g., gRNA). In some embodiments, the base editor provided herein utilizes a double-stranded guide polynucleotide (e.g., double-stranded gRNA). In some embodiments, the base editor provided herein utilizes one or more guide polynucleotides (e.g., multiple gRNAs). In some embodiments, a single-stranded guide polynucleotide is used in the different base editors described herein. For example, a single-stranded guide polynucleotide can be used in a cytidine base editor and an adenosine base editor.

[0267] In other embodiments, the guiding polynucleotide may contain both the polynucleotide targeting portion and the scaffold portion of the nucleic acid within a single molecule (i.e., a single-molecule guiding nucleic acid). For example, a single-molecule guiding polynucleotide may be a single-stranded guiding RNA (sgRNA or gRNA). In this document, the term "guiding polynucleotide sequence" encompasses any single-molecule, bimolecule, or multimolecule nucleic acid capable of interacting with and guiding a target polynucleotide sequence to that sequence.

[0268] Typically, a guiding polynucleotide (e.g., a crRNA / trRNA complex or gRNA) comprises a "polynucleotide targeting fragment" and a "protein-binding fragment," the latter containing a sequence capable of recognizing and binding to a target polynucleotide sequence. The guiding polynucleotide is stabilized within a polynucleotide programmable nucleotide-binding domain component of a base editor. In some embodiments, the guiding polynucleotide's polynucleotide targeting fragment recognizes and binds to a DNA polynucleotide, thereby facilitating base editing in DNA. In other cases, the guiding polynucleotide's polynucleotide targeting fragment recognizes and binds to an RNA polynucleotide, thereby facilitating base editing in RNA. In this document, a "fragment" refers to a portion or region of a molecule, such as a continuous extension of nucleotides in a guiding polynucleotide; a fragment may also refer to a region / part of a complex, such that a fragment may contain a region greater than or equal to 1. For example, in the case where the guiding polynucleotide comprises multiple nucleic acid molecules, its protein-binding fragment may include all or part of multiple individual molecules, for example, hybridizing along complementary regions. A binding fragment of RNA targeting DNA consisting of two separate molecules may comprise (i) 40 to 75 base pairs of the first RNA molecule, with a length of 100 base pairs; and (ii) 10 to 25 base pairs of the second RNA molecule, with a length of 50 base pairs. The definition of a “segment,” unless otherwise explicitly defined in a specific context, is not limited to a specific number of total base pairs, nor is it limited to any specific number of base pairs derived from any of the following. For a given RNA molecule, it is not limited to a specific number of separate molecules in the complex, and may include regions of RNA molecules of any total length, and may include regions complementary to other molecules.

[0269] Guide RNA or guide polynucleotides can contain two or more RNAs, such as CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). Guide RNA or guide polynucleotides can sometimes contain single-stranded RNA, or be a single-stranded guide RNA (sgRNA) formed by the fusion of crRNA and a portion of tracrRNA (e.g., the functional portion). Guide RNA or guide polynucleotides can also be double-stranded RNAs containing both crRNA and tracrRNA. Furthermore, crRNA can hybridize with target DNA.

[0270] As described above, the guide RNA or guide polynucleotide can be an expression product. For example, the DNA encoding the guide RNA can be a vector containing a sequence encoding the guide RNA. The guide RNA or guide polynucleotide can be transferred into cells by transfecting cells with isolated guide RNA or plasmid DNA containing a sequence encoding the guide RNA and a promoter. The guide RNA or guide polynucleotide can also be transferred into cells in other ways, such as using virus-mediated gene delivery.

[0271] Guide RNA or guide polynucleotides can be isolated. For example, guide RNA can be transfected into cells or organisms as isolated RNA. Guide RNA can be prepared by in vitro transcription using any in vitro transcription system known in the art. Guide RNA can be transferred into cells as isolated RNA rather than as a plasmid containing the coding sequence of the guide RNA.

[0272] Guide RNAs or guide polynucleotides may contain three regions: a first region at the 5' end, which may be complementary to a target site in the chromosomal sequence; a second inner region, which may form a stem-loop structure; and a third 3' region, which may be single-stranded. The first region of each guide RNA may also be different, allowing each guide RNA to direct the fusion protein to a specific target site. Furthermore, the second and third regions of each guide RNA may be identical across all guide RNAs.

[0273] The first region of the guide RNA or guide polynucleotide can be complementary to the sequence of the target site in the chromosomal sequence, allowing the first region of the guide RNA to pair with the bases of the target site. In some cases, the first region of the guide RNA may contain about 10 to 25 nucleotides or about 10 to 25 nucleotides (i.e., 10 nucleotides to nucleotides; or about 10 nucleotides to about 25 nucleotides; or 10 nucleotides to about 25 nucleotides; or 10 nucleotides to about 25 nucleotides; or about 10 nucleotides to about 25 nucleotides; or about 10 nucleotides to 25 nucleotides) or more. For example, the base-pairing region between the first region of the guide RNA and the target site in the chromosomal sequence may be or may be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more nucleotides. Sometimes, the length of the first region of the guide RNA may be or may be about 19, 20 or 21 nucleotides.

[0274] The guide RNA or guide polynucleotide may also contain a second region that forms a secondary structure. For example, a secondary structure formed by the guide RNA may contain a stem (or hairpin) and a loop. The lengths of the loop and the stem can vary. For example, the length of the loop may be in the range of about 3 to 10 nucleotides or about 3 to 10 nucleotides, and the length of the stem may be in the range of about 6 to 20 base pairs or about 6 to 20 base pairs. The stem may contain one or more protrusions of 1 to 10 or about 10 nucleotides. The total length of the second region may be in the range of about 16 to 60 nucleotides or about 16 to 60 nucleotides. For example, the length of the loop may be or may be about 4 nucleotides, and the length of the stem may be or may be about 12 base pairs.

[0275] The guide RNA or guide polynucleotide may also include a third region at its 3' end, which can be essentially single-stranded. For example, the third region is sometimes not complementary to any chromosomal sequence in the target cell, and sometimes not complementary to the rest of the guide RNA. Furthermore, the length of the third region can vary. The length of the third region can be greater than or greater than about 4 nucleotides. For example, the length of the third region can range from about 5 to 60 nucleotides.

[0276] Guide RNA or guide polynucleotides can target any exon or intron of a gene target. In some cases, the guide RNA can target exon 1 or exon 2 of a gene; the guide RNA can target exon 3 or 4 of a gene. The composition may contain multiple guide RNAs that all target the same exon, or in some cases, it may contain multiple guide RNAs that can target different exons. Both exons and introns of the gene can be targeted.

[0277] Guide RNA or guide polynucleotides can target nucleic acid sequences of approximately 20 nucleotides or more. Target nucleic acids can be less than or equal to approximately 20 nucleotides. The length of the target nucleic acid can be at least or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 30 nucleotides, or between 1 and 100 nucleotides. Target nucleic acids can be at most or at most about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or any value between 1 and 100 nucleotides in length. The target nucleic acid sequence can be or can be the 20 bases of the first 5' nucleotide of the PAM. Guide RNA can target nucleic acid sequences. The target nucleic acid can be at least or at least about 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 70, 1 to 80, 1 to 90 or 1 to 100 nucleotides.

[0278] Guide polynucleotides, such as guide RNA, can refer to nucleic acids that can hybridize with another nucleic acid, such as target nucleic acids or protospacers in the cell genome. Guide polynucleotides can be RNA. Guide polynucleotides can be DNA. Guide polynucleotides can be programmed or designed to bind to nucleic acid sequences specifically at a specific site. Guide polynucleotides can contain a polynucleotide chain and can be called single-stranded guide polynucleotides. Guide polynucleotides can contain two polynucleotide chains and can be called double-stranded guide polynucleotides. Guide RNA can be introduced into cells or embryos as RNA molecules. For example, RNA molecules can be transcribed in vitro and / or can be chemically synthesized. Guide RNA can be synthesized from DNA molecules (e.g., The gene fragment is transcribed into RNA. The guide RNA can then be introduced into cells or embryos as an RNA molecule. Guide RNA can also be introduced into cells or embryos in the form of a non-RNA nucleic acid molecule, such as a DNA molecule. For example, the DNA encoding the guide RNA can be operatively ligated to a promoter control sequence to express the guide RNA in target cells or embryos. The RNA coding sequence can be operatively ligated to a promoter sequence recognized by RNA polymerase III (Pol III). Plasmid vectors that can be used to express guide RNA include, but are not limited to, the px330 vector and the px333 vector. In some cases, the plasmid vector (e.g., the px333 vector) may contain at least two DNA sequences encoding guide RNA.

[0279] Methods for selecting, designing, and validating guide polynucleotides, such as guide RNAs and target sequences, are described herein and are known to those skilled in the art. For example, to minimize the potential impact of mass contamination on deaminase domains within debasase domains (e.g., AID domains), the number of deamination residues can be unintentionally targeted (e.g., off-target C residues may be minimized on ssDNA within the target nucleic acid locus). Additionally, software tools can be used to optimize gRNAs corresponding to target nucleic acid sequences, for example, to minimize off-target activity across the entire genome. For example, for each possible target domain selection using *Streptococcus pyogenes* Cas9, all off-target sequences (preceding the selected PAM, e.g., NAG or NGG) containing up to a certain number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) mismatched base pairs can be identified across the entire genome. First regions of gRNAs complementary to the target site can be identified, and all first regions (e.g., crRNAs) can be ranked according to their total predicted off-target score; the highest-ranked region indicates the most and least targeted activity. Candidate target gRNAs can be functionally evaluated using methods known in the art and / or described herein.

[0280] As a non-limiting example, DNA sequence search algorithms can be used to identify target DNA hybridization sequences in the crRNA of the guide RNA used with Cas9. Custom gRNA design software based on the public tool cas-offinder can be used for gRNA design, as described in Bae S., Park J., and Kim J.-S. Cas-OFFinder: A fast and general algorithm for searching for potential off-target sites of Cas9 RNA-guided nucleases, Bioinformatics 30, 1473-1475 (2014). The software scores the guide after calculating the tendency of the guide to deviate from the target genome-wide. Typically, for guides ranging from 17 to 24 nucleotides in length, matches ranging from perfect matches to 7 non-matches are considered. Once non-target sites are determined through calculation, a total score is calculated for each guide and summarized into a tabular output interface using a web version. In addition to identifying potential target sites adjacent to PAM sequences, the software also identifies all PAM adjacent sequences that differ from the selected target site by 1, 2, 3, or more nucleotides. The genomic DNA sequence of the target nucleic acid sequence, such as the target gene, can be obtained, and repetitive components can be screened using publicly available tools (such as the RepeatMasker program). RepeatMasker searches for repetitive elements and low-complexity regions in the input DNA sequence. The output is a detailed annotation of the repetitions present in the given query sequence.

[0281] Following identification, the first regions of the guide RNA, such as crRNAs, are classified based on their distance from the target site, their orthogonality, and the presence of 5' nucleotides that closely match the associated PAM sequence (e.g., based on the identification of closely matched 5'G sequences in the human genome containing the following: associated PAM, such as NGG PAM for Streptococcus pyogenes, NNGRRT or NNGRRV PAM for Staphylococcus aureus). As used herein, orthogonality refers to the number of sequences in the human genome containing the minimum number of mismatches with the target sequence. “High level of orthogonality” or “good orthogonality” can, for example, refer to a 20-mer target domain that has no identical sequences in the human genome other than the intended target, nor any sequences containing one or two mismatches at the target. Target domains with good orthogonality can be selected to minimize off-target DNA cleavage.

[0282] In some embodiments, the reporter system can be used to detect base editing activity and test candidate guide polynucleotides. In some embodiments, the reporter system may include a reporter gene-based assay, wherein base editing activity results in the expression of the reporter gene. For example, the reporter system may include a reporter gene containing an inactivated start codon, such as a mutation from 3'-TAC-5' to 3'-CAC-5' on the template strand. Upon successful deamination of the target C, the corresponding mRNA will be transcribed to 5'-AUG-3' instead of 5'-GUG-3', thereby enabling the translation of the reporter gene. Suitable reporter genes will be apparent to those skilled in the art. Non-limiting examples of reporter genes include genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secretory alkaline phosphatase (SEAP), or any other gene that is apparent to those skilled in the art to be expressed. The reporter system can be used to test many different gRNAs, for example, to determine which residue of the target DNA sequence the corresponding deaminase will target. It is also possible to test sgRNAs targeting non-template strands to evaluate specific base-editing proteins (e.g., Cas9 deaminase fusion proteins). In some embodiments, the gRNA can be engineered such that the start codon of the mutation does not pair with the gRNA bases. The guiding polynucleotide may comprise standard ribonucleotides, modified ribonucleotides (e.g., pseudouridine), ribonucleotide isomers, and / or ribonucleotide analogs. In some embodiments, the guiding polynucleotide may comprise at least one detectable label. The detectable label may be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tags, or suitable fluorescent dyes), a detection label (e.g., biotin, isohydroxydigitoxin, etc.), quantum dots, or gold particles.

[0283] The guide polynucleotide can be chemically synthesized, enzymatically synthesized, or a combination thereof. For example, guide RNA can be synthesized using a solid-phase synthesis method based on standard phosphorylamide. Alternatively, guide RNA can be synthesized in vitro by operatively linking DNA encoding guide RNA to a promoter control sequence recognized by phage RNA polymerase. Examples of suitable phage promoter sequences include T7, T3, SP6 promoter sequences, or variants thereof. In embodiments where the guide RNA comprises two separate molecules (e.g., crRNA and tracrRNA), crRNA can be chemically synthesized and tracrRNA can be enzymatically synthesized.

[0284] In some embodiments, the base editor system may contain multiple guide polynucleotides, such as gRNAs. For example, the gRNAs may target one or more target loci (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs) contained in the base editor system. The multiple gRNA sequences may be arranged in tandem and are preferably separated by direct repeats.

[0285] DNA sequences encoding guide RNA or guide polynucleotides can also be part of a vector. Furthermore, vectors can contain other expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selection marker sequences (e.g., GFP or antibiotic resistance genes, such as puromycin), origin of replication, etc. DNA molecules encoding guide RNA can also be linear. DNA molecules encoding guide RNA or guide polynucleotides can also be circular.

[0286] In some implementations, one or more components of the base editor system may be encoded by DNA sequences. Such DNA sequences can be introduced into expression systems, such as cDNA. They can be a single unit, together, or separately. For example, DNA sequences encoding a polynucleotide programmable nucleotide-binding domain and a guide RNA can be introduced into cells; each DNA sequence can be part of a separate molecule (e.g., a vector containing a sequence encoding a polynucleotide programmable nucleotide-binding domain and another containing a sequence encoding a guide RNA), or both can be part of the same molecule (e.g., a vector containing a sequence encoding (and regulating) a polynucleotide programmable nucleotide-binding domain and a sequence encoding (and regulating) the guide RNA).

[0287] Directing polynucleotides may contain one or more modifications to provide nucleic acids with novel or enhanced characteristics. Directing polynucleotides may contain nucleic acid affinity tags. Directing polynucleotides may contain synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.

[0288] In some cases, the gRNA or guide polynucleotide may contain modifications. Modifications can be made at any position on the gRNA or guide polynucleotide. More than one modification can be made to a single-stranded gRNA or guide polynucleotide. After modification, the gRNA or guide polynucleotide can undergo quality control. In some cases, quality control may include PAGE, HPLC, MS, or any combination thereof.

[0289] Modifications to gRNA or guide polynucleotides can be substitution, insertion, deletion, chemical modification, physical modification, stabilization, purification, or any combination thereof.

[0290] It can also be modified by 5'-adenosine, 5'-guanosine triphosphate cap, 5'-N7-methylguanosine triphosphate cap, 5'-triphosphate cap, 3'-phosphate ester, 3'-thiophosphate ester, 5'-phosphate ester, 5'-thiophosphate ester, cis-thymine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, d spacer, PC spacer, r spacer, 18-spacer, 9,3'-3' modification, 5'-5' modification, base-free, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesterol TEG, desulfurized biotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, double biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3'-DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methylribonucleoside analog, sugar-modified analog, wobbling / universal base, fluorescent dye labeling, 2'-fluoroRNA, 2'-O-methylRNA, phosphonate methyl ester, phosphodiester deoxyribonucleic acid, phosphodiester RNA, thiophosphate DNA, thiophosphate RNA, UNA, pseudouridine 5'-triphosphate, 5'-methylcytidine 5'-triphosphate, or any combination thereof modifying gRNA or guide polynucleotides.

[0291] In some cases, the modification is permanent. In others, it is temporary. In some cases, multiple modifications are made to the gRNA or guide polynucleotide. Modifications to gRNA or guide polynucleotides can alter the physicochemical properties of the nucleotide, such as its conformation, polarity, hydrophobicity, chemical reactivity, base pairing interactions, or any combination thereof.

[0292] Modifications can also involve phosphate thioester (PS) substitutes. In some cases, native phosphodiester bonds may be readily degraded by cellular nucleases, and modifying the internucleotide bonds with phosphate thioester (PS) bond substitutes makes them more stable against hydrolysis by cellular degradation. Modifications can increase the stability of gRNA or directing polynucleotides. Modifications can also enhance biological activity. In some cases, phosphate thioester-enhanced RNA gRNAs can inhibit RNase A, RNase T1, fetal bovine serum nuclease, or any combination thereof. These properties allow PS-RNA gRNAs to be used in applications where exposure to nucleases is highly likely, either in vivo or in vitro. For example, introducing a phosphate thioester (PS) bond between the last 3-5 nucleotides at the 5' or 6' end of the gRNA can inhibit exonuclease degradation. In some cases, phosphate thioester bonds can be added throughout the gRNA to reduce endonuclease attack.

[0293] Adjacent motifs of the original spacers

[0294] The term "protospacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the protospacer).

[0295] The protospacer adjacent motif (PAM), or PAM-like motif, refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). The PAM sequence is crucial for target binding, but the exact sequence depends on the type of Cas protein.

[0296] The base editors provided herein can contain CRISPR protein-derived domains capable of binding nucleotide sequences containing typical or atypical protospacer adjacent motif (PAM) sequences. The PAM site is a nucleotide sequence adjacent to the target polynucleotide sequence. Some aspects of this application provide base editors containing all or part of CRISPR proteins with different PAM specificities. For example, a typical Cas9 protein, such as Cas9 of Streptococcus pyogenes (spCas9), requires a typical NGGPAM sequence to bind to a specific nucleic acid region, where "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. The PAM can be CRISPR protein-specific and can differ between different base editors containing different CRISPR protein-derived domains. The PAM can be 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. The length of the PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. Typically, the length of the PAM is between 2 and 6 nucleotides.

[0297] In some embodiments, the Cas9 domain is the Cas9 domain of *Streptococcus pyogenes* (SpCas9). In some embodiments, the SpCas9 domain is an active nuclease SpCas9, an inactive nuclease SpCas9 (SpCas9d), or a SpCas9 cleavage enzyme (SpCas9n). In some embodiments, SpCas9 contains a D9X mutation or a corresponding mutation in any amino acid sequence provided herein, where X is any amino acid other than D. In some embodiments, SpCas9 includes a D9A mutation or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to nucleic acid sequences having atypical PAM sequences. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to nucleic acid sequences having NGG, NGA, or NGCG PAM sequences. In some embodiments, the SpCas9 domain contains one or more of the D1135X, R1335X, and T1336X mutations or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain contains one or more of the D1135E, R1335Q, and T1336R mutations or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain contains the D1135E, R1335Q, and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain contains one or more of the D1135X, R1335X, and T1336X mutations or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain contains one or more of the D1135V, R1335Q, and T1336R mutations or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain contains the D1135V, R1335Q, and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain contains one or more of the D1135X, G1217X, R1335X, and T1336X mutations, or corresponding mutations, in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain contains one or more of the D1135V, G1217R, R1335Q, and T1336R mutations, or corresponding mutations, in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain contains the D1135V, G1217R, R1335Q, and T1336R mutations, or corresponding mutations, in any amino acid sequence provided herein.

[0298] In some embodiments, the Cas9 domain of any fusion protein provided herein comprises at least 60%, at least 65%, at least 70%, at least 75%, or at least 80% of the amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with the Cas9 peptide described herein. In some embodiments, the Cas9 domain of any fusion protein provided herein comprises the amino acid sequence of any Cas9 peptide described herein. In some embodiments, the Cas9 domain of any fusion protein provided herein consists of the amino acid sequence of any Cas9 peptide described herein.

[0299] The sequence of an exemplary SpCas9 protein that can bind to the PAM sequence is shown below.

[0300] Exemplary SpCas9

[0301]

[0302] Exemplary SpCas9n

[0303]

[0304] Example SpEQR Cas9

[0305] E SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRIDLSQLGGD

[0306] The residues E1135, Q1335, and R1337 above are marked with underline and bold, which can be mutated from D1135, R1335, and T1337 to produce SpEQR Cas9.

[0307] Example SpVQR Cas9

[0308] V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRIDLSQLGGD

[0309] The residues V1135, Q1335, and R1337 are indicated by underline and bold, which can be mutated from D1135, R1335, and T1337 to produce SpVQR Cas9.

[0310] Example SpVRER Cas9

[0311] V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASA R ELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK E Y R STKEVLDATLIHQSITGLYETRIDLSQLGGD.

[0312] Exemplary SpVRQR Cas9

[0313] V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASA R ELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRIDLSQLGGD.

[0314] The residues V1135, R1218, Q1335, and R1337 above are indicated by underline and bold, which can be mutated from D1135, G1218, R1335, and T1337 to produce SpVRQR Cas9.

[0315] In some embodiments, the Cas9 domain is a recombinant Cas9 domain. In some embodiments, the recombinant Cas9 domain is a SpyMacCas9 domain. In some embodiments, the SpyMacCas9 domain is a nuclease-active SpyMacCas9, a nuclease-inactive SpyMacCas9 (SpyMacCas9d), or a SpyMacCas9 nickase (SpyMacCas9n). In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind nucleic acid sequences with atypical PAM sequences. In some embodiments, the SpyMacCas9 domain, SpCas9d domain, or SpCas9n domain can bind nucleic acid sequences with NAA PAM sequences.

[0316] Example SpyMacCas9

[0317]

[0318] DYLQNHNQQFDVLFNEIISFSKKCKLGKEHIQKIENVYSNKKNSASIEELAESFIKLLGFTQLGATSPFNFLGVKLNQKQYKGKKDYILPCTEGTLIRQSITGLYETRVDLSKIGED.

[0319] High-fidelity Cas9 domain

[0320] Some aspects of this application provide high-fidelity Cas9 domains. In some embodiments, the high-fidelity Cas9 domain is an engineered Cas9 domain comprising one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA relative to the corresponding wild-type Cas9 domain. High-fidelity Cas9 domains with reduced electrostatic interaction with the sugar-phosphate backbone of DNA may have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., the wild-type Cas9 domain) comprises one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain contains one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.

[0321] In some embodiments, any Cas9 fusion protein provided herein contains one or more of the N497X, R661X, Q695X, and / or Q926X mutations or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, any Cas9 fusion protein provided herein contains one or more of the N497A, R661A, Q695A, and / or Q926A mutations or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the Cas9 domain contains the D10A mutation or corresponding mutation in any amino acid sequence provided herein. High-fidelity Cas9 domains are known in the art and will be apparent to those skilled in the art. For example, Kleinstiver, BP et al., “High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects.”, Nature 529, 490-495 (2016); and Slaymaker, IM et al., “Rationally engineered Cas9 nucleases with improved specificity.”, Science 351, 84-88 (2015); the entire contents of each are incorporated herein by reference.

[0322] High-fidelity Cas9 domain mutations compared to Cas9 are shown in bold and underlined.

[0323] MDKKYSIGL AIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLNLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKIKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMT A FDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAHLFDDKVMKQLKRRRYTGWG A LSRKLINGIRDKQSGKTILDFLKSDGFANRNFM A LIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETR AITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRLKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD.

[0324] In some cases, variant Cas9 proteins carry mutations such as H840A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1218A, which reduces the peptide's ability to cleave target DNA or RNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, variant Cas9 proteins carry mutations such as D10A, H840A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1218A, which reduces the peptide's ability to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the mutated Cas9 protein carries mutations such as W476A and W1126A, or mutations such as P475A, W476A, N477A, ​​D1125A, W1126A, and D1218A, the mutated Cas9 protein cannot effectively bind to the PAM sequence. Therefore, in some such cases, when using this variant Cas9 protein in a binding method, the method does not require the PAM sequence. In other words, in some cases, when using this Cas9 variant protein in a binding method, the method may include a guide RNA, but the method can be performed in the absence of the PAM sequence (the binding specificity is thus provided by the targeting fragment of the guide RNA). Other residues can be mutated to achieve the aforementioned effect (i.e., partially inactivating one or more nucleases). As non-restrictive examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). Similarly, mutations other than alanine substitution are also acceptable.

[0325] In some embodiments, the CRISPR protein-derived domain of the base editor may contain all or a portion of the Cas9 protein having a standard PAM sequence (NGG). In other embodiments, the Cas9-derived domain of the base editor may employ an atypical PAM sequence. Such sequences have been described in the art and will be apparent to those skilled in the art. For example, a Cas9 domain binding an atypical PAM sequence has been described in Kleinstiver, BP, et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities,” Nature 523, 481-485 (2015). Kleinstiver, BP, et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition,” Nature Biotechnology 33, 1293-1298 (2015); the entire contents of each are incorporated herein by reference.

[0326] In some instances, the PAM recognized by the CRISPR protein-derived domain of the base editor disclosed herein can be provided to the cell on a single oligonucleotide encoding an insert (e.g., an AAV insert). In this case, providing the PAM on a single oligonucleotide can cleave target sequences that would otherwise be uncut because there is no adjacent PAM on the same polynucleotide as the target sequence.

[0327] In one implementation, *Streptococcus pyogenes* Cas9 (SpCas9) can be used as a CRISPR endonuclease for genome engineering. However, other methods can be used. In some cases, different endonucleases can be used to target certain genomic targets. In some cases, synthetic SpCas9-derived variants with non-NGG PAM sequences can be used. Additionally, other Cas9 orthologs from various species have been identified, and these “non-SpCas9s” can bind multiple PAM sequences, which can also be used in this application. For example, the relatively large size of SpCas9 (approximately 4kb coding sequence) can lead to inefficient expression of plasmids carrying SpCas9 cDNA in cells. Conversely, the coding sequence of *Staphylococcus aureus* Cas9 (SaCas9) is approximately 1,000 bases shorter than SpCas9, potentially enabling efficient expression in cells. Similar to SpCas9, the SaCas9 endonuclease is capable of modifying target genes in mammalian cells both in vitro and in vivo. In some cases, Cas proteins can target different PAM sequences. In some cases, the target gene may be adjacent to a Cas9 PAM (e.g., 5'-NGG). In other cases, other Cas9 orthologs may have different PAM requirements. For example, other PAMs include Streptococcus thermophilus (5'-NNAGAA for CRISPR1 and 5'-NGGNG for CRISPR3) and Neisseria meningitidis (5'-NNNNGATT).

[0328] In some implementations, for *Streptococcus pyogenes*, the target gene sequence may be preceding the 5'-NGG PAM (i.e., 5' to), and the 20-nt guide RNA sequence may pair with the opposite strand bases. This mediates Cas9 cleavage adjacent to the PAM. In some cases, the adjacent fragment may be or may be approximately 3 base pairs upstream of the PAM. In some cases, the adjacent fragment may be or may be approximately 10 base pairs upstream of the PAM. In some cases, the adjacent fragment may be or may be approximately 0-20 base pairs upstream of the PAM. For example, adjacent cleavages may be immediately adjacent to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 base pairs upstream of the PAM. Adjacent fragments may also be 1 to 30 base pairs downstream of the PAM.

[0329] Fusion proteins containing nuclear localization sequences (NLS) In some embodiments, the fusion proteins provided herein further include one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, such as nuclear localization sequences (NLS). In one embodiment, a binary NLS is used. In some embodiments, the NLS contains an amino acid sequence that facilitates the introduction of the NLS-containing protein into the cell nucleus (e.g., via nuclear transport). In some embodiments, any fusion protein provided herein also includes a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of a Cas9 domain. In some embodiments, the NLS is fused to the C-terminus of an nCas9 domain or a dCas9 domain. In some embodiments, the NLS is fused to the N-terminus of a deaminase. In some embodiments, the NLS is fused to the C-terminus of a deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to a fusion protein without linkers. In some embodiments, the NLS comprises the amino acid sequence of any NLS sequence provided or referenced herein. Other nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, the contents of which are incorporated herein by reference as exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequences PKKKRKVEGADKRTADGSEFESPKKKRKV, KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRKPKKRKKVK, or KMLKNRLCFLYQLY. In some embodiments, the NLS is present in or side-attached to a linker, such as the linker described herein. In some embodiments, the N-terminal or C-terminal NLS is a bifurcated NLS. The dichotomous NLS consists of two basic amino acid clusters separated by a relatively short spacer sequence (hence, a dichotomous NLS has 2 parts, unlike the monochotomous NLS). The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK, is a prototype of the ubiquitous dichotomous signal: two basic amino acid clusters separated by a spacer of approximately 10 amino acids. An exemplary dichotomous NLS sequence is as follows: PKKKRKVEGADKRTADGSEFES PKKKRKV.

[0330] In some embodiments, the fusion protein of this application does not contain a linker sequence. In some embodiments, one or more domains or linker sequences between proteins are present.

[0331] It should be understood that the fusion protein of this application may include one or more additional features. For example, in some embodiments, the fusion protein may include an inhibitor, a cytoplasmic localization sequence, an export sequence (e.g., a nuclear export sequence or other localization sequence), and a sequence tag that can be used for solubilization, purification, or detection of the fusion protein. Suitable protein tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, multihistidine tags (also known as histidine tags or His-tags), maltose-binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softags (e.g., Softag 1, Softag 3), streptococcal tags, biotin ligase tags, FLASH tags, V5 tags, and SBP tags. Other suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein includes one or more His tags.

[0332] Vectors encoding CRISPR enzymes containing one or more nuclear localization sequences (NLS) can be used. For example, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 NLSs can be used. CRISPR enzymes may contain NLSs at or near the ammunition terminus, or combinations of approximately or more than approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 NLSs at or near the carboxyl terminus (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxyl terminus). When more than one NLS is present, each can be selected independently, such that a single NLS may exist in more than one copy and / or in one or more copies together with one or more other NLSs.

[0333] The CRISPR enzyme used in the method may contain approximately 6 NLS. When the amino acid closest to the NLS is within approximately 50 amino acids in the polypeptide chain from the N-terminus or C-terminus, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, or 50 amino acids, the NLS is considered to be close to the N-terminus or C-terminus.

[0334] The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include, but are not limited to, NGG, NGA, NGC, NGN, NGT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NGRRR(N), TTTV, TYCV, TYCV, TATV, NNNGTAT, NNAGAAW, or NAAAAC. Y is pyrimidine; N is any nucleotide base; W is A or T.

[0335] Cas9 domain with reduced exclusivity

[0336] Typically, Cas9 proteins (such as the Cas9 of Streptococcus pyogenes (spCas9)) require a typical NGG PAM sequence to bind to a specific nucleic acid region, where the "N" in "NGG" stands for adenosine (A), thymidine (T), or cytosine (C), and the "G" stands for guanosine. This can limit the ability to edit desired bases in the genome. In some embodiments, the base-editing fusion proteins described herein may need to be placed in precise locations, such as regions containing target bases upstream of a PAM. See, for example, Komor, AC et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. Therefore, in some embodiments, any fusion protein described herein may contain a Cas9 domain capable of binding nucleotide sequences that do not contain a typical (e.g., NGG) PAM sequence. Cas9 domains that bind atypical PAM sequences have been described in the art and will be apparent to those skilled in the art. For example, the Cas9 domains that bind atypical PAM sequences have been described in “Engineered CRISPR-Cas9 nucleases with altered PAM specificities”, published by Kleinstiver, BP et al., Nature 523, 481-485 (2015). The full contents of the following articles are incorporated herein by reference: Kleinstiver, BP et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition”, Nature Biotechnology 33, 1293-1298 (2015); Nishimasu, H. et al., “Engineered CRISPR-Cas9 nuclease with expanded targeting space”, Science, 2018 Sep 21; 361(6408): 1259-1262; Chatterjee, P. et al., “Minimal PAM specificity of a highly similar SpCas9 ortholog”, Sci Adv. 2018 Oct 24; 4(10): eaau0766.doi:10.1126 / sciadv.aau0766.Several PAM variants are described in Table 1 below.

[0337] Table 1. Cas9 protein and corresponding PAM sequence

[0338]

[0339]

[0340] Nucleobase editing domain

[0341] This article describes a base editor comprising a fusion protein including a polynucleotide programmable nucleotide-binding domain and a nucleoside base editing domain (e.g., a deaminase domain). The base editor can be programmed to edit one or more bases in a target polynucleotide sequence by interacting with a guide polynucleotide capable of recognizing a target sequence. Once the target sequence is identified, the base editor is anchored to the polynucleotide to be edited, and then the deaminase domain component of the base editor can edit the target base.

[0342] In some embodiments, the nucleobase editing domain is a deaminase domain. In some cases, the deaminase domain may be cytosine deaminase or cytidine deaminase. In some embodiments, the terms "cytosine deaminase" and "cytidine deaminase" are used interchangeably. In some cases, the deaminase domain may be adenine deaminase or adenine deaminase. In some embodiments, the terms "adenine deaminase" and "adenine deaminase" are used interchangeably. Details of the nucleobase editing protein are described in International Patent Application No. PCT / 2017 / 045381 (WO No. 2018 / 027078) and International Patent Application No. PCT / US2016 / 058344 (WO No. 2017 / 070632), the entire contents of which are incorporated herein by reference. See also Komor, AC et al., “Programmable editing of a target base ingenomic DNA without double-stranded DNA cleavage”, Nature 533, 420-424 (2016); Gaudelli, NM et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage”, Nature 551, 464-471 (2017); and Komor, AC et al., “Improved baseexcision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity”, Science Advances 3:eaao4774 (2017), the full text of which is incorporated herein by reference.

[0343] C to T editing

[0344] In some embodiments, the base editor disclosed herein includes a fusion protein comprising a cytidine deaminase capable of deaminating a target cytidine (C) base of a polynucleotide to produce uridine (U), which has the base-pairing property of thymine. In some embodiments, for example in the case where the polynucleotide is double-stranded (e.g., DNA), the uridine base can then be substituted by a thymidine base (e.g., via a cellular repair mechanism) to produce a C:G to T:A transition. In other embodiments, the base editor cannot substitute U for T when deaminating C to U in a nucleic acid.

[0345] The deamination of target C in a polynucleotide to produce U is a non-limiting example of the type of base editing that can be performed by the base editor described herein. In another example, a base editor containing a cytidine deaminase domain can mediate the conversion of a cytosine (C) base to a guanine (G) base. For example, the U of the polynucleotide produced by deamination of cytosine from the cytosine in the cytosine deaminase domain of the base editor can be removed from the polynucleotide via a base excision repair mechanism (e.g., via a uracil DNA glycosylase (UDG) domain), producing a baseless site. The nucleobase opposite the baseless site can then be replaced with another base, such as C, via, for example, a transfer damage polymerase (e.g., via a base repair mechanism). Although the nucleobase opposite the baseless site is usually replaced with C, other substitutions (e.g., A, G, or T) can also occur.

[0346] Therefore, in some embodiments, the base editor described herein includes a deamination domain (e.g., a cytidine deaminase domain) capable of deaminating a target C to U in a polynucleotide. Furthermore, as described below, the base editor may include other domains that facilitate the conversion of U resulting from deamination to T or G. For example, a base editor containing a cytidine deaminase domain may also include a uracil glycosylase inhibitor (UGI) domain to mediate the substitution of U for T, thereby completing the C-to-T base editing event. In another instance, the base editor may incorporate a translesion polymerase to enhance the efficiency of C-to-G base editing, since the translesion polymerase can promote the incorporation of C opposite a base-free site (i.e., resulting in the incorporation of G at a base-free site, completing the C-to-G base editing activity).

[0347] A base editor containing a cytidine deaminase as a domain can deaminate the target C in any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. Typically, cytidine deaminase catalyzes the C nucleobase located in the context of a single-stranded portion of a polynucleotide. In some embodiments, the entire polynucleotide containing the target C can be single-stranded. For example, a cytidine deaminase incorporated into a base editor can deaminate the target C in a single-stranded RNA polynucleotide. In other embodiments, a base editor containing a cytidine deaminase domain can act on double-stranded polynucleotides, but the target C can be located in a portion of the polynucleotide that is in a single-stranded state during the deamination reaction. For example, in an embodiment where the NAGPB domain contains a Cas9 domain, several nucleotides can become unpaired during the formation of the Cas9-gRNA-target DNA complex, resulting in the formation of a Cas9 “R-loop complex.” These unpaired nucleotides can form bubbles of single-stranded DNA, which can then be used as a substrate for a single-stranded specific nucleotide deaminase, such as cytidine deaminase.

[0348] In some embodiments, the cytidine deaminase of the base editor may comprise all or a portion of the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C-to-U editing enzymes. Like proteins, the N-terminal domain of APOBEC is the catalytic domain, while the C-terminal domain is the pseudo-catalytic domain. More specifically, the catalytic domain is the zinc-dependent cytidine deaminase domain and is important for cytidine deamination. APOBEC family members include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminases. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of the APOBEC1 deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC2 deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3 deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3A deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3B deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3C deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3D deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3E deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3F deaminase. In some embodiments, the deaminase incorporated into the base editor includes all or a portion of APOBEC3G deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of APOBEC3H deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of APOBEC4 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of activation-induced deaminase (AID). In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of cytidine deaminase 1 (CDA1). It should be understood that the base editor may contain deaminases from any suitable organism (e.g., human or rat). In some embodiments, the deaminase domain of the base editor is derived from human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.In some embodiments, the deaminase domain of the base editor is derived from rats (e.g., rat APOBEC1). In some embodiments, the deaminase domain of the base editor is human APOBEC1. In some embodiments, the deaminase domain of the base editor is pmCDA1.

[0349] The base and amino acid sequences of PmCDA1 and the CDS of human AID are shown below.

[0350] >tr|A5H718|A5H718_PETMA Cytosine Deaminase OS=Gymnogalactiae OX=7757PE=2 SV=1:

[0351] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSG

[0352] TERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLK

[0353] IWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKT

[0354] LKRAEKRRSELSIMIQVKILHTTKSPAV

[0355] >EF094822.1 Isolation of PmCDA.21 cytosine deaminase mRNA and intact CDS from sea lamprey.

[0356] TGACACGACACAGCCGTGTATATGAGGAAGGGTAGCTGGATGGGGGGGGGGGGAATACGTTCAGAGAGGACATTAGCGAGCGTCTTGTTGGTGGCCTTGAGTCTAGACACCTGCAGACATGACCGACGCTGAGTACGTGAGAATCCATGAGAAGTTGGACATCTACACGTTTAAGAAACAGTTTTTCAACAACAAAAAATCCGTGTCGCATAGATGCTACGTTCTCTTTGAATTAAAACGACGGGGTGAACGTAGAGCGTGTTTTTGGGGCTATGCTGTGAATAAACCACAGAGCGGGACAGAACGTGGAATTCACGCCGAAATCTTTAGCATTAGAAAAGTCGAAGAATACCTGCGCGACAACCCCGGACAATTCACGATAAATTGGTACTCATCCTGGAGTCCTTGTGCAGATTGCGCTGAAAAGATCTTAGAATGGTATAACCAGGAGCTGCGGGGGAACGGCCACACTTTGAAAATCTGGGCTTGCAAACTCTATTACGAGAAAAATGCGAGGAATCAAATTGGGCTGTGGAACCTCAGAGATAACGGGGTTGGGTTGAATGTAATGGTAAGTGAACACTACCAATGTTGCAGGAAAATATTCATCCAATCGTCGCACAATCAATTGAATGAGAATAGATGGCTTGAGAAGACTTTGAAGCGAGCTGAAAAACGACGGAGCGAGTTGTCCATTATGATTCAGGTAAAAATACTCCACACCACTAAGAGTCCTGCTGTTTAAGAGGCTATGCGGATGGTTTTC

[0357] >tr|Q6QJ80|Q6QJ80_Human activation-induced cytidine deaminase OS=Homo sapiens OX=9606 GN=AICDA PE=2 SV=1:

[0358] MDSLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKAPV

[0359] >NG_011588.1:5001-15681 Human activation-induced cytidine deaminase (AICDA) on chromosome 12, reference sequence (LRG_17)

[0360]

[0361] The following provides other exemplary deaminases that can be fused with Cas9 according to aspects of this application. It should be understood that in some embodiments, the active domain of the corresponding sequence may be used, such as a domain without a localization signal (nuclear localization sequence, no nucleus output signal, cytoplasmic localization signal).

[0362] Human AID:

[0363]

[0364] (Bottom line: nuclear localization sequence; Double bottom line: nuclear output signal)

[0365] Mouse AID:

[0366]

[0367] (Bottom line: nuclear localization sequence; Double bottom line: nuclear output signal)

[0368] Canine AID:

[0369]

[0370] (Bottom line: nuclear localization sequence; Double bottom line: nuclear output signal)

[0371] Bovine AID:

[0372]

[0373] (Bottom line: nuclear localization sequence; Double bottom line: nuclear output signal)

[0374] Rat AID:

[0375]

[0376]

[0377] (Bottom line: nuclear localization sequence; Double bottom line: nuclear output signal)

[0378] mouse APOBEC-3

[0379]

[0380] (Italicized: Nucleic acid editing domain)

[0381] Rat APOBEC-3:

[0382]

[0383] (Italicized: Nucleic acid editing domain)

[0384] Rhesus monkey APOBEC-3G:

[0385]

[0386] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization signal) Chimpanzee APOBEC-3G:

[0387]

[0388]

[0389] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization signal) Green Monkey APOBEC-3G:

[0390]

[0391] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization signal) Human APOBEC-3G:

[0392]

[0393] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization signal) Human APOBEC-3F:

[0394]

[0395] (Italicized: Nucleic acid editing domain)

[0396] Human APOBEC-3B:

[0397] MNPQIRNPMERMYRDTFYDNFENEPILYGRSYTWLCYEVKIKRGRSNLLWDTGVFRGQVYFKPQYHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLSEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVTIMDYEEFAYCWENFVYNEGQQFMPWYKFDENYAFLHRTLKEILRY LMDPDTFTFNFNNDPLVLRRRQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQNQGN

[0398] (Italicized: Nucleic acid editing domain)

[0399] Rat APOBEC-3B:

[0400] MQPQGLGPNAGMGPVCLGCSHRRPYSPIRNPLKKLYQQTFYFHFKNVRYAWGRKNNFLCYEVNGMDCALPVPLRQGVFRKQGHIHAELCFIYWFHDKVLRVLSPMEEFKVTWYMSWSPCSKCAEQVARFLAAHRNLSLAIFSSRLYYYLRNPNYQQKLCRLIQEGVHVAAMDLPEFKKCWNKFVDNDGQPFRPWMRLRINFSFYDCKLQEIFSRMNLLREDVFYLQFNNSHRVKPVQNRYYRRKSYLCYQLERANGQEPLKGYLLYKKGEQHVEILFLEKMRSMELSQVRITCYLTWSPCPNCARQLAAFKKDHPDLILRIYTSRLYFWRKKFQKGLCTLWRSGIHVDVMDLPQFADCWTNFVNPQRPFRPWNELEKNSWRIQRRLRRIKESWGL

[0401] Bovine APOBEC-3B:

[0402] DGWEVAFRSGTVLKAGVLGVSMTEGWAGSGHPGQGACVWTPGTRNTMNLLREVLFKQQFGNQPRVPAPYYRRKTYLCYQLKQRNDLTLDRGCFRNKKQRHAERFIDKINSLDLNPSQSYKIICYITWSPCPNCANELVNFITRNNHLKLEIFASRLYFHWIKSFKMGLQDLQNAGISVAVMTHTEFEDCWEQFVDNQSRPFQPWDKLEQYSASIRRRLQRILTAPI

[0403] Chimpanzee APOBEC-3B:

[0404] MNPQIRNPMEWMYQRTFYYNFENEPILYGRSYTWLCYEVKIRRGHSNLLWDTGVFRGQMYSQPEHHAEMCFLSWFCGNQLSAYKCFQITWFVSWTPCPDCVAKLAKFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYNEGQPFMPWYKFDDNYAFLHRTLKEIIRHLMDPDTFTFNFNNDPLVLRRHQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGQVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQVRASSLCMVPHRPPPPPQSPGPCLPLCSEPPLGSLLPTGRPAPSLPFLLTASFSFPPPASLPPLPSLSLSPGHLPVPSFHSLTSCSIQPPCSSRIRETEGWASVSKEGRDLG

[0405] Human APOBEC-3C:

[0406] MNPQIRNPMKAMYPGTFYFQFKNLWEANDRNETWLCFTVEGIKRRSVVSWKTGVFRNQVDSETHCHAERCFLSWFCDDILSPNTKYQVTWYTSWSPCPDCAGEVAEFLARHSNVNLTIFTARLYYFQYPCYQEGLRSLSQEGVAVEIMDYEDFKYCWENFVYNDNEPFKPWKGLKTNFRLLKRRLRESLQ

[0407] (Italic: Nucleic acid editing domain)

[0408] Gorilla APOBEC-3C3C

[0409] MNPQIRNPMKAMYPGTFYFQFKNLWEANDRNETWLCFTVEGIKRRSVVSWKTGVFRNQVDSETHCHAERCFLSWECDDILSPNTNYQVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLYYFQDTDYQEGLRSLSQEGVAVKIMDYKDFKYCWENFVYNDDEPFKPWKGLKYNFRFLKRRLQEILE

[0410] Human APOBEC-3A:

[0411] MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN

[0412] (Italic: nucleic acid editing domain)

[0413] Cynomolgus macaque APOBEC-3A:

[0414] MDGSPASRPRHLMDPNTFTFNFNNDLSVRGRHQTYLCYEVERLDNGTWVPMDERRGFLCNKAKNVPCGDYGCHVELRFLCEVPSWQLDPAQTYRVTWFISWSPCFRRGCAGQVRVFLQENKHVRLRIFAARIYDYDPLYQEALRTLRDAGAQVSIMTYEEFKHCWDTFVDRQGRPFQPWDGLDEHSQALSGRLRAILQNQGN

[0415] (Italic: nucleic acid editing domain)

[0416] Bovine APOBEC-3A:

[0417] MDEYTFTENFNNQGWPSKTYLCYEMERLDGDATIPLDEYKGFVRNKGLDQPEKPCHAELYFLGKIHSWNLDRNQHYRLTCFISWSPCYDCAQKLTTFLKENHHISLHILASRIYTHNRFGCHQSGLCELQAAGARITIMTFEDFKHCWETFVDHKGKPFQPWEGLNVKSQALCTELQAILKTQQN

[0418] (Italicized: Nucleic acid editing domain)

[0419] Human APOBEC-3H:

[0420]

[0421]

[0422] (Italicized: Nucleic acid editing domain)

[0423] Rhesus monkey APOBEC-3H:

[0424] MALLTAKTFSLQFNNKRRVNKPYYPRKALLCYQLTPQNGSTPTRGHLKNKKKDHAEIRFINKIKSMGLDETQCYQVTCYLTWSPCPSCAGELVDFIKAHRHLNLR IFASRLYYHWRPNYQEGLLLLLCGSQVPVEVMGLPEFTDCWENFVDHKEPPSFNPSEKLEELDKNSQAIKRRLERIKSRSVDVLENGLRSLQLGPVTPSSSIRNSR

[0425] Human APOBEC-3D:

[0426]

[0427] (Italicized: Nucleic acid editing domain)

[0428] Human APOBEC-1:

[0429] MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR

[0430] Mouse APOBEC-1:

[0431] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNTRCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYPPSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK

[0432] Rat APOBEC-1:

[0433] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK

[0434] Human APOBEC-2:

[0435] MAQKEEAAVATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKGGQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEEPEIQAALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK

[0436] Mouse APOBEC-2:

[0437] MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEVQSKGGQAQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0438] Rat APOBEC-2:

[0439] MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYLWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0440] Bovine APOBEC-2:

[0441] MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEPEIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0442] Petromyzon marinus CDA1 (pmCDAl):

[0443] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSFMIQVKILHTTKSPAV

[0444] Human APOBEC3G D316R D317R:

[0445] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQ

[0446] VYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDP

[0447] KVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKFNYDEFQHCWSKFVYSQ

[0448] RELFEPWNNLPKYYILLHFMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVER

[0449] MHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTC

[0450] FTSWSPCFSCAQEMAKFISKKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFT YSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN

[0451] Human APOBEC3G Chain A:

[0452] MDPPTFTFNFNNEPWWGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISFTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQ

[0453] Human APOBEC3G Chain A D120R D121R:

[0454] MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQ

[0455] Some aspects of this application are based on the understanding that modulating the catalytic activity of the deaminase domain of any fusion protein described herein, for example by influencing the synthetic ability of the fusion protein (e.g., a base editor) through point mutations in the deaminase domain. For example, mutations that reduce, but do not eliminate, the catalytic activity of the deaminase domain in a base-edited fusion protein can decrease the likelihood that the deaminase domain will catalyze the deamination of residues adjacent to the target residue, thereby narrowing the deamination window. The ability to narrow the deamination window can prevent unwanted deamination of residues adjacent to a specific target residue, which can reduce or prevent off-target effects.

[0456] For example, in some embodiments, the APOBEC deaminase incorporated into the base editor may contain one or more mutations selected from rAPOBEC1, namely H121X, H122X, R126X, R126X, R118X, W90X, W90X, and R132X, or one or more corresponding mutations in another APOBEC deaminase, wherein X is any amino acid. In some embodiments, the APOBEC deaminase incorporated into the base editor may contain one or more mutations selected from rAPOBEC1, namely H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E, or one or more corresponding mutations in another APOBEC deaminase.

[0457] In some embodiments, the APOBEC deaminase incorporated into the base editor may contain one or more mutations selected from hAPOBEC3G, specifically D316X, D317X, R320X, R320X, R313X, W285X, W285X, R326X, or one or more corresponding mutations. Another APOBEC deaminase contains a mutation where X is any amino acid. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing one or more mutations selected from hAPOBEC3G, specifically D316R, D317R, R320A, R320E, R313A, W285A, W285Y, R326E, or one or more corresponding mutations. In another APOBEC deaminase...

[0458] In some embodiments, the APOBEC deaminase incorporated into the base editor may contain the H121R and H122R mutations of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may contain an APOBEC deaminase containing the R126A mutation of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may contain an APOBEC deaminase containing the R126E mutation of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may contain an APOBEC deaminase containing the R118A mutation of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W90A mutation of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W90Y mutation of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the R132E mutation of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing both the W90Y and R126E mutations of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the R126E and R132E mutations of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W90Y and R132E mutations of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W90Y, R126E, and R132E mutations of rAPOBEC1, or one or more corresponding mutations of another APOBEC deaminase.

[0459] In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the D316R and D317R mutations of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R320A mutation of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the R320E mutation of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the R313A mutation of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W285A mutation of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W285Y mutation of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the R326E mutation of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing both the W285Y and R320E mutations of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the R320E and R326E mutations of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W285Y and R326E mutations of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise an APOBEC deaminase containing the W285Y, R320E, and R326E mutations of hAPOBEC3G, or one or more corresponding mutations of another APOBEC deaminase.

[0460] Many modified cytidine deaminases are commercially available, including but not limited to SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, which are available from Addgene (plasmids 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, and 85177).

[0461] The following provides other exemplary deaminases that can be fused with Cas9 according to aspects of this application. It should be understood that in some embodiments, active domains of the corresponding sequences may be used, such as domains without localization signals (nuclear localization sequences, no nuclear output signals, cytoplasmic localization signals).

[0462] Detailed information about C-to-T nucleobase editing proteins is described in International Patent Application No. PCT / US2016 / 058344 (International Patent Application No. WO2017 / 070632) and Komor, AC et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage”, Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference.

[0463] A to G editing

[0464] In some embodiments, the base editor described herein may include a deaminase domain comprising adenosine deaminase. This adenosine deaminase domain of the base editor facilitates the editing of adenine (A) nucleoside bases to guanine (G) nucleoside bases by deaminating A to form inosine (I), exhibiting the base-pairing properties of G. Adenosine deaminase enables the deamination (i.e., removal of the amino group) of adenine from deoxyadenosine residues in deoxyribonucleic acid (DNA).

[0465] In some embodiments, the nucleobase editors provided herein can be prepared by fusing one or more protein domains together to create a fusion protein. In some embodiments, the fusion proteins provided herein contain one or more features that improve the base-editing activity (e.g., efficiency, selectivity, and specificity) of the fusion protein. For example, the fusion proteins provided herein may contain a Cas9 domain with reduced nuclease activity. In some embodiments, the fusion proteins provided herein may have a Cas9 domain without nuclease activity (dCas9) or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, referred to as Cas9 cleavage enzyme (nCas9). Without being bound by any particular theory, the presence of catalytic residues (e.g., H840) maintains the activity of Cas9 in cleaving the unedited (e.g., undeamination) strand containing a T opposite to target A. “Residues” of the catalytic residues of Cas9 (e.g., D10 to A10) prevent cleavage of the edited strand containing the target A residue. These Cas9 variants can generate single-strand DNA breaks (gaps) at specific locations based on target sequences defined by gRNA, leading to the repair of the unedited strand and ultimately resulting in an unedited daughter strand. In some embodiments, the A-to-G base editor also includes an inosine base excision repair inhibitor, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactivated inosine-specific nuclease. Without being bound by any particular theory, the UGI domain or the catalytically inactivated inosine-specific nuclease can inhibit or prevent base excision repair of deaminoadenosine residues (e.g., inosine), which can enhance the activity or efficiency of the base editor.

[0466] A base editor containing adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In some embodiments, the base editor containing adenosine deaminase can deaminate target A of a polynucleotide containing RNA. For example, the base editor may contain an adenosine deaminase domain capable of deaminating target A of RNA polynucleotides and / or DNA-RNA hybrid polynucleotides. In one embodiment, the adenosine deaminase incorporated into the base editor comprises all or part of an adenosine deaminase that acts on RNA (ADAR, e.g., ADAR1 or ADAR2). In another embodiment, the adenosine deaminase incorporated into the base editor comprises all or part of an adenosine deaminase that acts on tRNA (ADAT). A base editor containing an adenosine deaminase domain can also deaminate the A nucleobase of a DNA polynucleotide. In one embodiment, the adenosine deaminase domain of the base editor comprises all or part of ADAT, which contains one or more mutations that allow ADAT to deaminate target A in DNA. For example, the base editor may contain all or part of ADAT from Escherichia coli (EcTadA), which contains one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I157F, or another mutant adenosine deaminase.

[0467] Adenosine deaminases can be derived from any suitable organism. In some embodiments, adenosine deaminases are derived from prokaryotes. In some embodiments, adenosine deaminases are derived from bacteria. In some embodiments, adenosine deaminases are derived from *Escherichia coli*, *Staphylococcus aureus*, *Salmonella typhi*, *Shewanella putrefactive*, *Haemophilus influenzae*, *Crescentia*, or *Bacillus subtilis*. In some embodiments, adenosine deaminases are derived from *Escherichia coli*. In some embodiments, adenine deaminases are naturally occurring adenosine deaminases comprising one or more mutations corresponding to any mutations provided herein (e.g., mutations in ecTadA). The corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determination of homologous residues. Mutations in any naturally occurring adenosine deaminase (e.g., homologous to ecTadA) corresponding to any mutation described herein (e.g., any mutation identified in ecTadA) can be generated accordingly.

[0468] TadA

[0469] In a particular embodiment, TadA is any of the TadA described in International Patent Application No. PCT / US2017 / 045381 (International Patent Application No. WO2018 / 027078), which is incorporated herein by reference in its entirety.

[0470] In one embodiment, the fusion protein of this application comprises wild-type TadA linked to TadA7.10, which is linked to a Cas9 nickase. In a particular embodiment, the fusion protein comprises a single TadA7.10 domain (e.g., provided in monomeric form). In other embodiments, the ABE7.10 editor comprises TadA7.10 and TadA (wt) capable of forming heterodimers. The relevant sequences are as follows:

[0471] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (referred to as "Ta").

[0472] TadA7.10:

[0473] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD

[0474] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any amino acid sequence listed in any adenosine deaminase provided herein. It should be understood that the adenosine deaminases provided herein may include one or more mutations (e.g., any mutations provided herein). This application provides any deaminase domain having a certain percentage of identity, plus any mutations described herein or combinations thereof. In some embodiments, compared to a reference sequence or any adenosine deaminase provided herein, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations. In some embodiments, compared to any amino acid sequence known in the art or described herein, adenosine deaminase comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues.

[0475] In some embodiments, the TadA deaminase is a whole-celled Escherichia coli TadA deaminase. For example, in some embodiments, the adenosine deaminase comprises the following amino acid sequence:

[0476] MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEI KAQKKAQSSTD.

[0477] However, it should be understood that other adenosine deaminases that can be used in this application are obvious to those skilled in the art and are within the scope of this application. For example, the adenosine deaminase can be a homolog of an adenosine deaminase that acts on tRNA (ADAT). Without limitation, the amino acid sequence of an exemplary ADAT homolog includes the following:

[0478] Staphylococcus aureus TadA:

[0479] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGS LMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN Bacillus subtilis TadA:

[0480] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML

[0481] VIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCS GTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIK ALKKADRAEGAGPAV

[0482] Shewanella putrefaciens (S. putrefaciens) TadA:

[0483] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE

[0484] TadA from Haemophilus influenzae F3031 (H.influenzae):

[0485] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFG

[0486] ASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK

[0487] TadA from Caulobacter crescentus (C.crescentus):

[0488] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADD

[0489] PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI TadA from Geobacter sulfurreducens (G.sulphurredensens): MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP

[0490] In some embodiments, the adenosine deaminase contains a D108X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X represents any amino acid other than the corresponding amino acid in the wild type. In some embodiments, the adenosine deaminase contains a D108G, D108N, D108V, D108A, or D108Y mutation, or a corresponding mutation in another adenosine deaminase. However, it should be understood that additional deaminases can be similarly compared to identify homologous amino acid residues that can be mutated as provided herein.

[0491] In some embodiments, the adenosine deaminase contains an A106X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an A106V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0492] In some embodiments, the adenosine deaminase contains an E155X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein the presence of X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an E155D, E155G, or E155V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0493] In some embodiments, the adenosine deaminase contains a D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains D147Y, a mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0494] It should be understood that any mutations provided herein (e.g., based on the TadA reference sequence) can be introduced into other adenosine deaminases (such as *Escherichia coli* TadA (ecTadA)), *Staphylococcus aureus* TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). It will be apparent to those skilled in the art that additional deaminases can be similarly compared to identify homologous amino acid residues that can be mutated as provided herein. Therefore, any mutation identified in the TadA reference sequence can be made in other adenosine deaminases having homologous amino acid residues (e.g., ecTada). It should also be understood that any mutations provided herein can be made alone or in any combination in the TadA reference sequence or in another adenosine deaminase. For example, an adenosine deaminase may contain mutations of D108N, A106V, E155V, and / or D147Y in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains the following mutation sets (mutation sets separated by semicolons) in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V and E55V; D108N, A106V and D147Y; D108N, E55V and D147Y; A106V, E55V and D147Y; and D108N, A106V, E55V and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be carried out in the adenosine deaminase (e.g., wild-type TadA or ecTadA).

[0495] In some embodiments, the adenosine deaminase comprises H8X, T17X, ​​L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X and / or K157X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X represents an amino acid other than the corresponding amino acid in any wild-type adenosine deaminase. In some embodiments, the adenosine deaminase includes one or more corresponding mutations in another adenosine deaminase, such as H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, 1951, V102A, F104L, A106V, R107C, R107H, R107P, D108G or D108N, D108V or D108A or D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D and / or K157R mutations in the TadA reference sequence.

[0496] In some embodiments, the adenosine deaminase contains one or more H8X, D108X, and / or N127X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein X represents the presence of any amino acid. In some embodiments, the adenosine deaminase contains one or more H8Y, D108N, and / or N127S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.

[0497] In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, or 6 mutations relative to the TadA reference sequence, specifically one or more mutations in H8X, D108X, N127X, D147X, R152X, and Q154X, or a corresponding mutation in another adenosine deaminase, wherein X represents any amino acid present in the wild-type adenosine deaminase other than the corresponding amino acid. In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, 6, 7, or 8 mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X represents any amino acid present in the wild-type adenosine deaminase other than the corresponding amino acid. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, or 5 mutations selected from H8X, D108X, N127X, E155X, and T166X in the TadA reference sequence, or a corresponding mutation in another adenosine. The deaminase, where X represents any amino acid present in the wild-type adenosine deaminase other than the corresponding amino acid.

[0498] In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, or 6 mutations relative to the TadA reference sequence, said mutations being selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H, or corresponding mutations of one or more other adenosine deaminases. In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, 6, 7, or 8 mutations in the TadA reference sequence, said mutations being selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H, or corresponding mutations of another adenosine deaminase. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, or 5 mutations selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, or 6 mutations selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q in the TadA reference sequence, or one or more corresponding mutations, or a mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, 6, 7, or 8 mutations selected from H8Y, R126W, L68Q, D108N, N127S, D147Y, and E155V in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, or 5 mutations selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0499] In some embodiments, the adenosine deaminase contains one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains the D108N, D108G, or D108V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains the A106V and D108N mutations in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains the R107C and D108N mutations in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains the H8Y, D108N, N127S, D147Y, and Q154H mutations in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of H8Y, R24W, D108N, N127S, D147Y, and E155V in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of D108N, D147Y, and E155V in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of H8Y, D108N, and N127S in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of A106V, D108N, D147Y, and E155V in the TadA reference sequence, or corresponding mutations in another adenosine deaminase (e.g., ecTadA).

[0500] In some embodiments, the adenosine deaminase contains one or more of the following mutations in the TadA reference sequence: a, S2X, H8X, I49X, L84X, H123X, N127X, I156X, and / or K160X, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains one or more of the following mutations in the TadA reference sequence: S2A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S, or one or more corresponding mutations in another adenosine deaminase.

[0501] In some embodiments, the adenosine deaminase comprises an L84X mutant adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an L84F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0502] In some embodiments, the adenosine deaminase contains an H123X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an H123Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0503] In some embodiments, the adenosine deaminase contains an I157X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an I157F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0504] In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, 6, or 7 mutations in the TadA reference sequence, said mutations being selected from the group consisting of L84X, A106X, D108X, H123X, D147X, E155X, and I156X, or a corresponding mutation in another adenosine deaminase, wherein X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, or 6 mutations, said mutations being selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, or 5 mutations selected from the group consisting of H8X, A106X, D108X, N127X, and K160X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X represents any amino acid other than the corresponding amino acid present in the wild-type adenosine deaminase.

[0505] In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, 6, or 7 mutations in the TadA reference sequence, said mutations being selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F, or a corresponding mutation of another adenosine deaminase. In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, 5, or 6 mutations, said mutations being selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V in the TadA reference sequence.

[0506] In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, or 5 mutations selected from H8Y, A106T, D108N, N127S, and K160S in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.

[0507] In some embodiments, the adenosine deaminase comprises one or more of the following mutations in the TadA reference sequence: E25X, R26X, R107X, A142X, and / or A143X, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more mutations in the TadA reference sequence, namely E25M, E25D, E25A, E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R07K, R107A, R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more mutations corresponding to the TadA reference sequence described herein, or one or more corresponding mutations in another adenosine deaminase.

[0508] In some embodiments, the adenosine deaminase contains an E25X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an E25M, E25D, E25A, E25R, E25V, E25S, or E25Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0509] In some embodiments, the adenosine deaminase contains an R26X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an R26G, R26N, R26Q, R26C, R26L, or R26K mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0510] In some embodiments, the adenosine deaminase contains an R107X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an R107P, R07K, R107A, R107N, R107W, R107H, or R107S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0511] In some embodiments, the adenosine deaminase contains the A142X mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains the A142N, A142D, A142G mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase.

[0512] In some embodiments, the adenosine deaminase comprises an A143X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q and / or A143R mutations in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0513] In some embodiments, the adenosine deaminase comprises one or more of the following mutations in the TadA reference sequence: H36X, N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X, and / or K161X, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations in the TadA reference sequence: H36L, N37T, N37S, P48T, P48L, I49V, R51H, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T, or one or more corresponding mutations in another adenosine deaminase.

[0514] In some embodiments, the adenosine deaminase contains an H36X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an H36L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0515] In some embodiments, the adenosine deaminase contains an N37X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an N37T or N37S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0516] In some embodiments, the adenosine deaminase contains a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains a P48T or P48L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0517] In some embodiments, the adenosine deaminase contains an R51X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an R51H or R51L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0518] In some embodiments, the adenosine deaminase contains an S146X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an S146R or S146C mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0519] In some embodiments, the adenosine deaminase contains a K157X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains a K157N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0520] In some embodiments, the adenosine deaminase contains a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains a P48S, P48T, or P48A mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0521] In some embodiments, the adenosine deaminase contains an A142X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine. In some embodiments, the adenosine deaminase contains an A142N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0522] In some embodiments, the adenosine deaminase contains a W23X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains a W23R or W23L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0523] In some embodiments, the adenosine deaminase contains an R152X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an R152P or R52H mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0524] In one embodiment, the adenosine deaminase may comprise mutants H36L, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F, and K157N. In some embodiments, the adenosine deaminase comprises the following combinations of mutations relative to the TadA reference sequence, wherein each mutation in the combination is separated by an underscore, and each combination of mutations is enclosed in parentheses:

[0525] (A106V_D108N),(R107C_D108N),

[0526] (H8Y_D108N_S127S_D147Y_Q154H),

[0527] (H8Y_R24W_D108N_N127S_D147Y_E155V), (D108N_D147Y_E155V),

[0528] (H8Y_D108N_S127S),(H8Y_D108N_N127S_D147Y_Q154H),

[0529] (A106V_D108N_D147Y_E155V),(D108Q_D147Y_E155V)

[0530] (D108M_D147Y_E155V),(D108L_D147Y_E155V),(D108K_D147Y_E155V),

[0531] (D108I_D147Y_E155V),

[0532] (D108F_D147Y_E155V),(A106V_D108N_D147Y),(A106V_D108M_D147Y_E155V),

[0533] (E59A_A106V_D108N_D147Y_E155V),(E59A catdead_A106V_D108N_D147Y_E155V),

[0534] (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y),

[0535] (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F),(D103A_D104N),

[0536] (G22P_D103A_D104N),(G22P_D103A_D104N_S138A),(D103A_D104N_S138A),

[0537] (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F),

[0538] (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F),

[0539] (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155V_I156F),

[0540] (R26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F),

[0541] (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F),

[0542] (R26C_L84F_A106V_R107H_D108N_H123Y_A142N_D147Y_E155V_I156F),

[0543] (L84F_A106V_D108N_H123Y_A142N_A143L_D147Y_E155V_I156F),

[0544] (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F),

[0545] (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I156F),

[0546] (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F),

[0547] (A106V_D108N_A142N_D147Y_E155V),

[0548] (R26G_A106V_D108N_A142N_D147Y_E155V),

[0549] (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V),

[0550] (R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V),

[0551] (E25D_R26G_A106V_D108N_A142N_D147Y_E155V),

[0552] (A106V_R107K_D108N_A142N_D147Y_E155V),

[0553] (A106V_D108N_A142N_A143G_D147Y_E155V),

[0554] (A106V_D108N_A142N_A143L_D147Y_E155V),

[0555] (H36L_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N),

[0556] (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F),

[0557] (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T),

[0558] (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F),

[0559] (N72S_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F),

[0560] (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F),

[0561] (H36L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N),

[0562] (H36L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F),

[0563] (L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T),

[0564] (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0565] (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N),

[0566] (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E),

[0567] (H36L_G67V_L84F_A106V_D108N_H123Y_S146T_D147Y_E155V_I156F),

[0568] (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F),

[0569] (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L),

[0570] (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0571] (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F),

[0572] (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0573] (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0574] (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L),

[0575] (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F),

[0576] (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N),

[0577] (N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T),

[0578] (L84F_A106V_D108N_D147Y_E155V_I156F),

[0579] (R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K161T),

[0580] (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K161T),

[0581] (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E_K161T),

[0582] (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E),(R74Q

[0583] L84F_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0584] (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0585] (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0586] (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0587] (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F),

[0588] (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F),

[0589] (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F),

[0590] (P48S_A142N),

[0591] (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N),(P48T_I49V_A142N),

[0592] (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N),

[0593] (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F),

[0594] (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N),

[0595] (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N),

[0596] (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N),

[0597] (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N),

[0598] (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F_K157N),

[0599] (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N),

[0600] (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N),

[0601] (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T),

[0602] (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152H_E155V_I156F_K157N),

[0603] (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N),

[0604] (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N),

[0605] (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_E155V_I156F_K157N),

[0606] (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_R152P_E155V_I156F_K157N),

[0607] (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T),

[0608] (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N),

[0609] (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_R152P_E155V_I156F_K157N).

[0610] In some embodiments, the fusion proteins provided herein include one or more features that enhance the base-editing activity of the fusion protein. For example, any fusion protein provided herein may include a Cas9 domain with reduced nuclease activity. In some embodiments, any fusion protein provided herein may have a Cas9 domain without nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, referred to as Cas9 cleavage enzyme (nCas9).

[0611] In some embodiments, the adenosine deaminase contains a D108X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild type. In some embodiments, the adenosine deaminase contains a D108G, D108N, D108V, D108A, or D108Y mutation, or a corresponding mutation in another adenosine deaminase.

[0612] In some embodiments, the adenosine deaminase contains an A106X, E155X, or D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an E155D, E155G, or E155V mutation. In some embodiments, the adenosine deaminase contains D147Y.

[0613] It should be understood that any mutations provided herein (e.g., based on the TadA reference sequence) can be introduced into other adenosine deaminases (such as *E. coli* TadA (ecTadA)), *S. aureus* TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). Any mutations identified relative to the TadA reference sequence can be performed in other adenosine deaminases having homologous amino acid residues. It should also be understood that any mutations provided herein, relative to the TadA reference sequence or another adenosine deaminase, can be performed alone or in any combination.

[0614] For example, adenosine deaminases may contain D108N, A106V, E155V and / or D147Y mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains the following mutation sets in the TadA reference sequence (mutation sets are separated by semicolons), or the corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V and E55V; D108N, A106V and D147Y; D108N, E55V and D147Y; A106V, E55V and D147Y; and D108N, A106V, E55V and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be performed in the adenosine deaminase (e.g., the TadA reference sequence or ecTadA).

[0615] In some embodiments, the adenosine deaminase comprises H8X, T17X, ​​L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X and / or K157X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X represents an amino acid other than the corresponding amino acid in any wild-type adenosine deaminase. In some embodiments, the adenosine deaminase includes one or more corresponding mutations in another adenosine deaminase, such as H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, 1951, V102A, F104L, A106V, R107C, R107H, R107P, D108G or D108N, D108V or D108A or D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D and / or K157R mutations in the TadA reference sequence. In some embodiments, the adenosine deaminase contains one or more H8X, D108X, and / or N127X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein X represents the presence of any amino acid. In some embodiments, the adenosine deaminase contains one or more H8Y, D108N, and / or N127S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.

[0616] In some embodiments, the adenosine deaminase comprises one or more of the following mutations in the TadA reference sequence: H8X, R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X, K161X, Q163X and / or T166X, or one or more corresponding mutations in another adenosine deaminase, wherein X represents any amino acid other than the corresponding amino acid present in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase includes one or more of the following mutations in the TadA reference sequence: H8Y, R26W, M61I, L68Q, M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G or E155V or E155D, K161Q, Q163H and / or T166P, or one or more corresponding mutations in another adenosine deaminase.

[0617] In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, or 6 mutations selected from one or more mutations in the TadA reference sequence H8X, D108X, N127X, D147X, R152X, and Q154X, or a corresponding mutation in another adenosine deaminase, wherein X represents any amino acid present in the wild-type adenosine deaminase other than the corresponding amino acid. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, 6, 7, or 8 mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase, wherein X represents any amino acid present in the wild-type adenosine deaminase other than the corresponding amino acid. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, or 5 mutations selected from H8X, D108X, N127X, E155X, and T166X in the TadA reference sequence, or a corresponding mutation in another adenosine. The deaminase, where X represents any amino acid present in the wild-type adenosine deaminase other than the corresponding amino acid.

[0618] In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, or 6 mutations selected from H8X, A106X, D108X, or one or more mutations of another adenosine deaminase, wherein X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, 6, 7, or 8 mutations selected from H8X, R126X, L68X, D108X, N127X, D147X, and E155X, or corresponding mutations or mutations of another adenosine deaminase, wherein X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains 1, 2, 3, 4, or 5 mutations selected from H8X, D108X, A109X, N127X, and E155X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X represents any amino acid other than the corresponding amino acid present in the wild-type adenosine deaminase.

[0619] In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, or 6 mutations selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H in the TadA reference sequence, or corresponding mutations of one or more other adenosine deaminases. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, 6, 7, or 8 mutations in the TadA reference sequence, selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H, or corresponding mutations of another adenosine deaminase. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, or 5 mutations selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, or 6 mutations selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q in the TadA reference sequence, or one or more corresponding mutations, or a mutation in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, 6, 7, or 8 mutations selected from H8Y, R126W, L68Q, D108N, N127S, D147Y, and E155V in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises 1, 2, 3, 4, or 5 mutations selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0620] Any mutations provided herein, and any other mutations (e.g., based on the ecTadA amino acid sequence), can be introduced into any other adenosine deaminase. Any mutations provided herein can be performed alone or in any combination of the TadA reference sequence or another adenosine deaminase.

[0621] Detailed information about A-to-G nucleobase editing proteins is described in the international PCT application PCT / 2017 / 045381 (WO2018 / 027078) and in “Programmable base editing of A·T to G·C in genomic DNA without DNA cutting” by Gaudelli, N.Maud et al., Nature, 551, 464-471 (2017), the entire contents of which are incorporated herein by reference.

[0622] Cytidine deaminase

[0623] In one embodiment, the fusion protein of this application comprises a cytidine deaminase. In some embodiments, the cytidine deaminase provided herein is capable of deaminating cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the cytosine deaminase provided herein is capable of deaminating cytosine in DNA. Cytidine deaminases can be derived from any suitable organism. In some embodiments, the cytidine deaminase is a naturally occurring cytidine deaminase containing one or more mutations corresponding to any mutation provided herein. Those skilled in the art will be able to identify the corresponding residues in any homologous protein, for example, by sequence alignment and determination of homologous residues. Therefore, those skilled in the art will be able to generate mutations in any naturally occurring cytidine deaminase corresponding to any mutation described herein. In some embodiments, the cytidine deaminase is derived from prokaryotes. In some embodiments, the cytidine deaminase is derived from bacteria. In some embodiments, the cytidine deaminase is derived from mammals (e.g., humans).

[0624] In some embodiments, the cytidine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the cytidine deaminase amino acid sequences shown herein. It should be understood that the cytidine deaminases provided herein may include one or more mutations (e.g.,...

Claims

1. A base editor for complexing with one or more guide polynucleotides in the preparation of a base for editing single nucleotide polymorphisms (SNPs) associated with Rett syndrome (RTT). MECP2 The use of polynucleotide drugs, the editing by making the... MECP2 The polynucleotide contacts the base editor that guides the complexation of one or more polynucleotides, wherein the base editor comprises: i) A SpCas9 polynucleotide programmable DNA binding domain, wherein the SpCas9 polynucleotide programmable DNA binding domain has a protospacer adjacent motif specific to the nucleic acid sequence 5'-NGT-3', and wherein the SpCas9 polynucleotide programmable DNA binding domain is composed of the following SpCas9 amino acid sequence: DKK YSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGET AEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTE ELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLL FKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQG DSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPE NIVIEMA RENQTTQK GQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRL SDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERG GLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAH DAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRP LIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQT GGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD and an amino acid change combination consisting of one of the following combinations of amino acid changes with reference to the SpCas9 amino acid sequence, wherein the first amino acid of the SpCas9 amino acid sequence is at position 2: L1111R, D1135L, G1218R, E1219F, A1322R, R1335V and T1337R; L1111R, D1135V, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218S, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219V, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, D1332A, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335Q and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V and T1337L; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337I; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337V; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337F; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337M; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337S; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337N; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337R; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337H; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337Q; D1135L, S1136R, G1218K, E1219I, A1322R, R1335Q, and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; L1111R, D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135V, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218R, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335V and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337V; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R and R1335Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V and T1337Q; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337T; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337R; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337Q; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218S, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218A, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218V, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337F; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337Y; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337F; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337Y; D1135G, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337L; D1135I, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136A, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136W, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136H, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136K, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218K, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218Q, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218T, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218N, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219I, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219A, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219N, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219Q, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219G, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219S, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219T, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335N, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335S and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335T and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335F, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Y, and T1337L; or D1135L, S1136R, G1218S, E1219V, N1286Q, A1322R, I1331F, R1335Q, and T1337L; and ii) Adenosine deaminase domain; and One or more of the described guiding polynucleotides target the base editor to cause an A•T to G•C change in an RTT-related SNP, wherein the RTT-related SNP is described in the... MECP2 The c.763C>T mutation in polynucleotides.

2. The use according to claim 1, wherein the contact is within a cell.

3. The use according to claim 2, wherein the cell is a eukaryotic cell.

4. The use according to claim 2, wherein the cell is a mammalian cell.

5. The use according to claim 2, wherein the cell is a human cell.

6. The use according to claim 2, wherein the cell is in vivo.

7. The use according to claim 2, wherein the cell is in vitro.

8. The use according to any one of claims 1 to 7, wherein, The SNP caused an amino acid change in R255X in the Mecp2 protein.

9. The use according to any one of claims 1 to 7, wherein the A•T to G•C change at the SNP associated with RTT will change the stop codon in the methylCpG binding protein 2 polypeptide to arginine.

10. The use according to any one of claims 1 to 7, wherein the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid.

11. The use according to claim 10, wherein the adenosine deaminase is a TadA deaminase.

12. The use according to claim 11, wherein the TadA deaminase is TadA * 7.

10.

13. The use according to any one of claims 1 to 7, wherein the one or more guiding polynucleotides comprise CRISPR RNA (crRNA) and trans-encoded small RNA, wherein the crRNA contains... Mecp2 Nucleic acid sequences complementary to each other, the Mecp2 The nucleic acid sequence is contained in the Mecp2 The c.763C>T mutation in polynucleotides.

14. The use according to any one of claims 1 to 7, wherein the base editor is complexed with a single-stranded guide RNA comprising a nucleic acid sequence complementary to the Mecp2 nucleic acid sequence, wherein the Mecp2 nucleic acid sequence is contained in the... Mecp2 The c.763C>T mutation in polynucleotides.

15. A cell generated by introducing a base editor or a polynucleotide encoding the base editor into a cell or its progenitor cell, wherein the base editor comprises: i) A SpCas9 polynucleotide programmable DNA binding domain, wherein the SpCas9 polynucleotide programmable DNA binding domain has a protospacer adjacent motif specific to the nucleic acid sequence 5'-NGT-3', and wherein the SpCas9 polynucleotide programmable DNA binding domain is composed of the following SpCas9 amino acid sequence: DKK YSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGET AEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTE ELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLL FKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQG DSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPE NIVIEMA RENQTTQK GQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRL SDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERG GLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAH DAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRP LIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQT GGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD and an amino acid change combination consisting of one of the following combinations of amino acid changes with reference to the SpCas9 amino acid sequence, wherein the first amino acid of the SpCas9 amino acid sequence is at position 2: L1111R, D1135L, G1218R, E1219F, A1322R, R1335V and T1337R; L1111R, D1135V, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218S, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219V, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, D1332A, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335Q and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V and T1337L; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337I; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337V; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337F; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337M; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337S; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337N; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337R; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337H; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337Q; D1135L, S1136R, G1218K, E1219I, A1322R, R1335Q, and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; L1111R, D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135V, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218R, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335V and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337V; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R and R1335Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V and T1337Q; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337T; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337R; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337Q; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218S, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218A, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218V, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337F; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337Y; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337F; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337Y; D1135G, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337L; D1135I, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136A, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136W, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136H, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136K, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218K, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218Q, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218T, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218N, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219I, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219A, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219N, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219Q, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219G, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219S, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219T, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335N, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335S and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335T and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335F, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Y, and T1337L; or D1135L, S1136R, G1218S, E1219V, N1286Q, A1322R, I1331F, R1335Q, and T1337L; and ii) Adenosine deaminase domain; and One or more guiding polynucleotides; The cell or its progenitor cell contains a single nucleotide polymorphism (SNP) associated with Rett syndrome (RTT), wherein one or more of the guide polynucleotides target a base editor to cause an A•T to G•C change in the RTT-associated SNP, and wherein the RTT-associated SNP is in MECP2 The c.763C>T mutation in polynucleotides.

16. The cell of claim 15, wherein the cell is a neuron.

17. The cell of claim 16, wherein the neuron expresses the Mecp2 polypeptide.

18. The cell according to any one of claims 15 to 17, wherein the cell is derived from a subject suffering from RTT.

19. The cell according to any one of claims 15 to 17, wherein the cell is a mammalian cell.

20. The cell according to any one of claims 15 to 17, wherein the cell is a human cell.

21. The cell according to any one of claims 15 to 17, wherein the SNP causes an amino acid alteration of R255X in the Mecp2 protein.

22. The cell according to any one of claims 15 to 17, wherein the A•T to G•C change at the SNP associated with RTT changes the stop codon in the methyl CpG binding protein 2 polypeptide to arginine.

23. The cell according to any one of claims 15 to 17, wherein the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid.

24. The cell of claim 23, wherein the adenosine deaminase is a TadA deaminase.

25. The cell of claim 24, wherein the TadA deaminase is TadA* 7.

10.

26. The cell according to any one of claims 15 to 17, wherein the one or more guiding polynucleotides comprise CRISPR RNA (crRNA) and trans-encoded small RNA, wherein the crRNA contains... Mecp2 Nucleic acid sequences complementary to each other, wherein Mecp2 The nucleic acid sequence is contained in the Mecp2 The c.763C>T mutation in polynucleotides.

27. The cell according to any one of claims 15 to 17, wherein the base editor is complexed with a single-stranded guide RNA, the single-stranded guide RNA comprising... Mecp2 Nucleic acid sequences complementary to each other, wherein Mecp2 The nucleic acid sequence is contained in the Mecp2 The c.763C>T mutation in polynucleotides.

28. Use of the cells according to claim 15 in the preparation of a medicament for treating RTT in a subject.

29. The use of a base editor system for preparing a medicament for treating Rett syndrome (RTT) in a subject, wherein the subject contains a single nucleotide polymorphism (SNP) associated with RTT, and wherein the base editor system comprises: A base editor or a polynucleotide encoding the base editor, wherein the base editor comprises: i) A SpCas9 polynucleotide programmable DNA binding domain, wherein the SpCas9 polynucleotide programmable DNA binding domain has a protospacer adjacent motif specific to the nucleic acid sequence 5'-NGT-3', and wherein the SpCas9 polynucleotide programmable DNA binding domain is composed of the following SpCas9 amino acid sequence: DKK YSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGET AEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTE ELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLL FKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQG DSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPE NIVIEMA RENQTTQK GQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRL SDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERG GLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAH DAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRP LIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQT GGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD and an amino acid change combination consisting of one of the following combinations of amino acid changes with reference to the SpCas9 amino acid sequence, wherein the first amino acid of the SpCas9 amino acid sequence is at position 2: L1111R, D1135L, G1218R, E1219F, A1322R, R1335V and T1337R; L1111R, D1135V, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218S, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219V, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, D1332A, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335Q and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V and T1337L; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337I; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337V; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337F; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337M; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337S; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337N; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337R; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337H; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337Q; D1135L, S1136R, G1218K, E1219I, A1322R, R1335Q, and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; L1111R, D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135V, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218R, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335V and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337V; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R and R1335Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V and T1337Q; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337T; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337R; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337Q; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218S, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218A, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218V, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337F; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337Y; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337F; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337Y; D1135G, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337L; D1135I, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136A, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136W, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136H, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136K, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218K, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218Q, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218T, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218N, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219I, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219A, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219N, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219Q, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219G, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219S, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219T, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335N, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335S and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335T and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335F, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Y, and T1337L; or D1135L, S1136R, G1218S, E1219V, N1286Q, A1322R, I1331F, R1335Q, and T1337L; and ii) Adenosine deaminase domain; and One or more targeted base editor-guided polynucleotides to cause an A•T to G•C change in an RTT-related SNP, wherein said RTT-related SNP is in MECP2 The c.763C>T mutation in polynucleotides.

30. The use according to claim 29, wherein the subject is a mammal.

31. The use according to claim 30, wherein the mammal is a human.

32. The use according to any one of claims 29 to 31, wherein the SNP causes an amino acid change in R255X in the Mecp2 protein.

33. The use according to any one of claims 29 to 31, wherein the A•T to G•C change at the SNP associated with RTT changes the stop codon in the methyl CpG binding protein 2 polypeptide to arginine.

34. The use according to any one of claims 29 to 31, wherein the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid.

35. The use according to claim 34, wherein the adenosine deaminase is a TadA deaminase.

36. The use according to claim 35, wherein the TadA deaminase is TadA * 7.

10.

37. The use according to any one of claims 29 to 31, wherein the one or more guiding polynucleotides comprise CRISPR RNA (crRNA) and trans-encoded small RNA, wherein the crRNA comprises a nucleic acid sequence complementary to the Mecp2 nucleic acid sequence, wherein the Mecp2 nucleic acid sequence is contained in the MECP2 The c.763C>T mutation in polynucleotides.

38. The use according to any one of claims 29 to 31, wherein the base editor is complexed with a single-stranded guide RNA, wherein the single-stranded guide RNA comprises a nucleic acid sequence complementary to the Mecp2 nucleic acid sequence, wherein the Mecp2 nucleic acid sequence is contained in the MECP2 The c.763C>T mutation in polynucleotides.

39. A base editor, comprising: (i) Modified SpCas9 having a protospacer adjacent motif specific to the nucleic acid sequence 5'-NGT-3', and consisting of the following SpCas9 amino acid sequence: DKK YSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGET AEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTE ELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLL FKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQG DSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPE NIVIEMA RENQTTQK GQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRL SDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERG GLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAH DAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRP LIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQT GGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD and an amino acid alteration combination consisting of one of the following combinations of amino acid alterations to the SpCas9 amino acid sequence, wherein the first amino acid of the SpCas9 amino acid sequence is at position 2: L1111R, D1135L, G1218R, E1219F, A1322R, R1335V and T1337R; L1111R, D1135V, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218S, E1219F, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219V, A1322R, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, D1332A, R1335V, and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335Q and T1337R; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V and T1337L; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337I; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337V; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337F; L1111R, D1135V, G1218R, E1219F, A1322R, R1335V, and T1337M; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337Q; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, and R1335Q; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337S; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337N; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337R; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337H; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337Q; D1135L, S1136R, G1218K, E1219I, A1322R, R1335Q, and T1337K; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; L1111R, D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135V, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218R, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332A, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335V and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337V; D1135L, S1136R, G1218S, E1219V, A1322R, D1332S, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332T, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332V, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332L, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332K, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, D1332R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R and R1335Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337T; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337R; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V and T1337Q; D1135L, S1136R, G1218R, E1219F, A1322R, R1335V, and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337T; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337R; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337Q; D1135L, S1136R, G1218S, E1219L, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337T; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337R; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337Q; D1135L, S1136R, G1218S, E1219F, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218S, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218A, E1219F, A1322R, R1335G, and T1337C; D1135L, S1136R, G1218V, E1219H, A1322R, R1335L and T1337N; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337F; D1135L, S1136R, G1218S, E1219L, A1322R, R1335A, and T1337Y; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337W; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337F; D1135L, S1136R, G1218S, E1219I, A1322R, R1335A and T1337Y; D1135G, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337L; D1135I, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136A, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136W, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136H, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136K, G1218S, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218K, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218Q, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218T, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218N, E1219V, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219I, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219A, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219N, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219Q, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219G, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219L, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219S, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219T, A1322R, R1335Q and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335L and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335I, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335N, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335S and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335T and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335F, and T1337L; D1135L, S1136R, G1218S, E1219V, A1322R, R1335Y, and T1337L; or D1135L, S1136R, G1218S, E1219V, N1286Q, A1322R, I1331F, R1335Q, and T1337L; and (ii) Adenosine deaminase.

Citation Information

Patent Citations

  • Towel rings

    CN3315821D

  • cell phone

    CN3329834D

  • Dehydrated liposomes

    US4880635A

  • Antineoplastic agent-entrapping liposomes

    US4906477A

  • Paucilamellar lipid vesicles

    US4911928A