Base editing methods and compositions for editing PRNP in the treatment of prion disease
Base editing compositions targeting the PRNP gene through CRISPR nucleases and deaminases address the limitations of ASO therapies by reducing PrP expression and misfolding, providing an effective treatment for prion disease across various genetic backgrounds.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-04-02
AI Technical Summary
Current therapies for prion disease, such as antisense oligonucleotide (ASO)-mediated approaches, face limitations in potency, biodistribution, and long-term tolerability for central nervous system targeting, necessitating the development of new therapeutics to halt or delay prion disease progression.
Base editing methods using compositions comprising CRISPR nucleases and deaminases, guided by specific RNAs, are employed to introduce mutations in the PRNP gene, reducing or eliminating PrP expression by blocking transcription, translation, or forming truncated variants that lack misfolding properties, thereby treating prion disease.
The base editing strategies effectively reduce PrP levels and prevent misfolding, offering a broader therapeutic benefit to a wider population, including those without known PRNP mutations, by installing premature stop codons or disrupting transcription factor binding, thus slowing disease progression.
Smart Images

Figure US2025046856_02042026_PF_FP_ABST
Abstract
Description
BASE EDITING METHODS AND COMPOSITIONS FOR EDITING PRNP IN THE TREATMENT OF PRION DISEASE RELATED APPLICATIONS
[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application, U.S.S.N. 63 / 700,235, filed September 27, 2024, and U.S. Provisional Application U.S.S.N. 63 / 718,534, filed November 8, 2024, both of which are incorporated herein by reference. GOVERNMENT SUPPORT
[0002] This invention was made with government support under Grant No. NS132315 awarded by the National Institutes of Health. The government has certain rights in the invention. REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0003] The contents of the electronic sequence listing (B119570203WO00-SEQ- BDP.xml; Size: 886,137 bytes; and Date of Creation: September 16, 2025) is herein incorporated by reference in its entirety. BACKGROUND OF INVENTION
[0004] Prion disease is a fatal neurodegenerative disease caused by the misfolding of prion protein encoded by the PRNP gene. The normal functioning prion protein is sometimes referred to a “PrP” or “PrPc” whereas the misfolded, disease-associated form is sometimes referred to as “PrPSc.” Subtypes of prion disease include Creutzfeldt-Jakob Disease (CJD), Gerstmann-Straussler-Scheinker disease (GSS), and fatal familial insomnia (FFI)1.
[0005] Misfolded PrP (or PrPSc) causes prion disease via a toxic gain of function1. In sporadic prion disease, which accounts for 85% of cases, an apparently spontaneous misfolding event of PrP initiates disease. In genetic cases (which account for ~15% of the cases), protein-coding mutations in the PRNP gene that encodes PrP increases the risk of protein misfolding3. In less than 1% of human cases, prion disease can be acquired by infection with exogenous pathogenic PrP from select organisms, such as cattle or other humans4. Prion disease occurs in many mammals5; inoculation of mice with pathogenic prion isolates triggers disease progression, which faithfully recapitulates human disease pathology and serves as a preclinical animal model for prion disease6.
[0006] While there is currently no cure for the disease, depleting PrP in the brain is an established strategy to prevent or stall templated misfolding of PrP. PrP has a signaling function related to myelin maintenance on peripheral nerves7, but reduction or elimination of PrP appears to be compatible with healthy life7,8. Heterozygous PRNP knockout mice show enhanced resistance to prion disease, and homozygous PRNP knockout mice are completely resistant to prion disease9. Antisense oligonucleotide (ASO)-mediated whole-brain PrP lowering is correlated with slowing of disease progression in prion infected mice2. While no homozygous PRNP-null humans have been observed, heterozygous null humans appear healthy, and occur at a frequency of ~1 / 18,000 healthy individuals, which suggests that homozygotes may be too rare to observe with currently available sequencing data8. And while no approved therapies are currently available for human use, a PrP-lowering antisense oligonucleotide (ASO)2is currently in a Phase I clinical trial.
[0007] Unfortunately, while ASO-mediated knockdown of PrP holds promise as potential disease-modifying therapy for prion disease, central nervous system (CNS)-targeting of ASOs currently suffer from limited potency and biodistribution to deep brain regions10, the need for repeated intrathecal dosing, and an unknown long-term tolerability profile. Accordingly, new therapeutics to halt and / or delay prion disease progression are needed. SUMMARY OF INVENTION
[0008] This present disclosure relates to the discovery that base editing may be an effective approach to treating prion disease. In various aspects and embodiments, the present disclosure provides (a) compositions that include base editors comprising napDNAbps (e.g. CRISPR nucleases) and a deaminase (e.g., an adenosine deaminase or a cytidine deaminase), guide RNAs that specifically target various regions of the PRNP gene (e.g., the promoter region, the promoter-proximal region, or the protein coding region), nucleic acid molecules encoding said base editors or their napDNAbp and / or deaminase components, complexes comprising a base editor complexed with a PRNP-targeting guide RNA and / or nucleic acid molecules encoding such complexes, vectors, delivery systems, cells, pharmaceutical compositions, and / or kits, as well as (b) methods of using and / or administering these various compositions to introduce or otherwise install one or more mutations or edits in the PRNP gene—such as in the promoter region, promoter-proximal region, or codon region—which have the effect of reducing and / or eliminating the expression and / or activity of PrP in a cell. In various aspect, the reduced and / or eliminated expression and / or activity of PrP in a cell can result from (i) reduced and / or blocked PRNP transcription, (ii) reduced and / or blocktranslation of the PRNP gene (e.g. by mutational elimination of the start codon), or (iii) formation of a truncated variant of PrP that lacks or otherwise has reduced misfolding properties observed for PrP. Such compositions, methods, and editing strategies as disclosed herein can be used to treat prior disease.
[0009] Exemplary methods of base editing that are contemplated herein include: (1) a method of installing one or more base edits in the promoter region of the PRNP gene by contacting the PRNP promoter region with a base editing composition disclosed herein thereby inhibiting, blocking, or otherwise reducing transcription of the PRNP gene, with concomitant reduction and / or elimination of PRNP mRNA, and with concomitant reduction and / or elimination of PrP expression levels; (2) a method of installing one or more base edits in a promoter-proximal region of the PRNP gene by contacting the PRNP promoter-proximal region with a base editing composition disclosed herein thereby inhibiting, blocking, or otherwise reducing transcription of the PRNP gene, with concomitant reduction and / or elimination of PRNP mRNA, and with concomitant reduction and / or elimination of PrP expression levels; (3) a method of installing one or more mutations in the start codon of the PRNP gene by contacting a target site comprising the start codon with a base editing composition disclosed herein thereby knocking down the translation of the PRNP gene, which consequently reduces and / or eliminates the production of PrP and consequently PrPSc formation; and (4) a method of installing one or more nonsense mutations in a codon (e.g. at the codons in the PRNP gene corresponding to amino acid positions R37, W57, W65, W81, or Q83 of PrP) by contacting a target codon with a base editing composition disclosed herein thereby converting the target codon to a pre-mature stop codon (i.e. nonsense mutation), which consequently results in the production of a truncated form or variant of PrP, and wherein the truncated variant of PrP lacks a PrPSc forming activity.
[0010] Each of these methods may be used to treat and / or minimize prion disease by minimizing and / or eliminating the production of PRNP mRNA, forming a non-disease causing PrP truncated variant, or altogether knocking out PrP expression.
[0011] In various aspects, the present disclosure provides base editing systems, compositions, and methods for installing a stop codon in a PRNP gene encoding a prion protein (PrP), e.g., to produce a non-pathogenic truncated PrP protein, to mutate the start codon of the PrP protein in the PRNP gene to prevent translation or to mutate the 273 bp promoter and promoter-proximal region of the PRNP gene to prevent transcription factorbinding and lower mRNA transcription thereby lowering PrP expression. The present disclosure also provides guide RNAs configured to direct the base editors to install a premature stop codon in the PRNP gene corresponding to an amino acid position in the PrP protein selected from the group consisting of R37, W57, W65, W81 and Q83. Insertion of a stop codon truncates the PrP protein and prevents it from catalyzing the misfolding of other mutant PrP proteins. Likewise, mutation of the start codon prevents translation from initiating. Likewise, mutation of the 273 bp promoter and promoter-proximal region would prevent transcription factor binding leading to lower mRNA transcription and reduced PrP protein expression. Also, provided herein are methods for treating or preventing prion disease, e.g., via truncating a PrP protein. The present disclosure also provides complexes comprising the base editor and guide RNAs disclosed herein, as well as vectors, cells, pharmaceutical compositions, and kits.
[0012] In various other aspects, the inventors have discovered an alternative approach to ASO therapy for depleting PrP levels in neural tissue. Instead of transiently inhibiting PrP protein production, which is the goal of ASO-based therapies, the inventors herein show that base editing may be used to either (1) install a premature stop codon in the PRNP gene encoding the PrP protein, to produce a non-pathogenic truncated PrP variant, (2) to mutate the start codon of the PrP protein in the PRNP gene to prevent translation, or (3) to mutate the 273 bp promoter and promoter-proximal region to prevent transcription factor binding and lower mRNA transcription thereby lowering PrP expression. Without wishing to be bound by any particular theory, it is generally believed that truncated PrP variants disclosed herein are too short to misfold into a pathogenic protein and / or to serve as a template to induce template-mediated misfolding of endogenous PrP proteins. Importantly, installation of an early stop codon in the PRNP gene, disruption of its start codon or disruption of transcription factor binding to the promoter or promoter-proximal region would be largely independent of the etiology or nature of the mutation, or even the existence of a mutation (e.g., PRNP mutations are known in the art to be heterogenous and about 85% of patients do not have a mutation), greatly expanding the population that would benefit from such a therapy. Additionally, depletion of PrP is thought to be more beneficial than restoration to non- pathogenic forms (e.g., via the use of gene editing to correct genomic mutations that cause prion disease), which can still be templated to misfold by pathogenic prion proteins.
[0013] Accordingly, some aspects of the present disclosure relate to base editing compositions for installing a stop codon at a target location in a PRNP gene encoding a prion protein (PrP), e.g., to produce a non-pathogenic truncated PrP protein. In someembodiments, the compositions comprise one or more nucleotide sequences encoding one or more portions of a base editor (e.g., deaminase domain, napDNAbp domain, NLS, UGI domain, etc.). In some embodiments, the base editing compositions produce R37X, W57X, W65X, W81X or Q83X mutations in a PrP protein encoded by the PRNP gene.
[0014] Because CBEs convert C-G base pairs into T-A base pairs they enable conversion of codons encoding for arginine (Arg), glutamine (Gln), and / or tryptophan (Trp) into stop codons (e.g., UAA, UAG, UGA) (see Figure 1). For example, the codon CGA encodes for the amino acid glutamine (Gln), which can be converted into the stop codon UGA, using base editing to mutate the C to a T. Analogously, ABEs may be used to mutate an ATG start codon to GTG or ACG, thus preventing translation of the PrP protein from occurring.
[0015] Those of skill in the art will understand that installation of an edit, e.g., using base editing, requires a guide RNA to position the base editor near a codon to be converted into a premature stop codon within an editing window of the base editor or to mutate the start codon of the PrP protein in the PRNP gene to prevent translation. Accordingly, in some embodiments, the compositions comprise a nucleotide sequence encoding a guide RNA configured to direct a base editor to install a premature stop codon at a target site in the PRNP gene. In some embodiments, the target site in the PRNP gene corresponds to any one of the following amino acid positions R37, W57, W65, W81, or Q83 in the PrP protein. Without wishing to be bound by any particular theory, it is known in the art that heterozygous PRNP R37X mutations have been observed in healthy humans,8and thus, in some embodiments, the target site in the PRNP gene corresponds to amino acid position R37 in the human PrP protein.
[0016] In some cases, the compositions disclosed herein comprise one or more nucleotide sequences encoding one or more portions of a split nucleobase editor. Split nucleobase editors are base editors (e.g., fusion proteins) that have been split into an N-terminal portion and a C-terminal portion. The N-terminal portion of the nucleobase editor is fused at its C- terminus to an intein-N, whereas an intein-C is fused to the N-terminus of a C-terminal portion of the nucleobase editor. Upon co-expression of the nucleotide sequences, the two sections of the nucleobase editor are recombined (e.g., ligated) via intein-mediated protein splicing. In some embodiments, the compositions further comprise a nucleotide sequence encoding a gRNA configured to direct the recombined nucleobase editor to a target site in the PRNP gene corresponding to any one of amino acid positions R37, W57, W65, W81, or Q83 in the PrP protein or to mutate an ATG start codon to GTG or ACG, thus preventingtranslation of the PrP protein from occurring. Such configurations are useful, for example, for delivering nucleobase editors using recombinant adeno-associated viruses (rAAV).
[0017] The present disclosure also provides compositions comprising one or more nucleotide sequences encoding one or more portions of an intact nucleobase editor (e.g., not a split nucleobase editor). In some embodiments, the compositions comprise a polynucleotide comprising one or more of the following: (1) a first nucleotide sequence encoding a guide RNA; (2) a second nucleotide sequence encoding a deaminase domain (e.g., TadCBEd or ABE8e or ABE8e(V106W)) of the nucleobase editor (3) a third nucleotide sequence encoding a nucleic acid programmable DNA binding protein (napDNAbp) domain (e.g., SauriCas9, enCjCas9, SpCas9, SaCas9) of the nucleobase editor; (4) a fourth nucleotide sequence encoding a terminator sequence (e.g., a polyT sequence); (5) a fifth nucleotide sequence encoding a UGI domain; (6) a sixth nucleotide sequence encoding an inverted terminal repeat (ITR) domain; or (7) a seventh nucleotide sequence encoding one or more miR target sites. In some embodiments, the one or more nucleotide sequences are operably linked to a promoter sequence, e.g., Cbh, EFS, hSYN, and / or U6. In some cases, the one or more nucleotide sequences are operably linked to a hSYN promoter.
[0018] In other aspects, the present disclosure provides compositions comprising rAAV particles for the targeted delivery of a base editor and a gRNA into a target cell (e.g., neurons). In some embodiments, the compositions comprise one or more rAAV particles comprising one or more of the base editing compositions disclosed herein. For instance, in some embodiments, the one or more recombinant rAAV particles comprises any one or more of the following: (1) a first nucleotide sequence encoding a gRNA configured to direct the recombined nucleobase editor to a target site in the PRNP gene corresponding to any one of amino acid positions R37, W57, W65, W81, or Q83 in the PrP protein or to mutate an ATG start codon to GTG or ACG, thus preventing translation of the PrP protein from occurring, (2) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N, and (3) a third nucleotide sequence encoding an intein- C fused to the N-terminus of a C-terminal portion of the split nucleobase editor.
[0019] In other embodiments, however, compositions disclosed herein comprise one or more rAAV particles comprising one or more of the following: (1) a first nucleotide sequence encoding a gRNA configured to direct the recombined nucleobase editor to a target site in the PRNP gene corresponding to any one of amino acid positions R37, W57, W65, W81, or Q83 in the PrP protein or to mutate an ATG start codon to GTG or ACG, thus preventing translation of the PrP protein from occurring, (2) a second nucleotide sequence encoding adeaminase domain of the nucleobase editor; (3) a third nucleotide sequence encoding a napDNAbp domain of the nucleobase editor; (4) a fourth nucleotide sequence encoding a UGI domain; (5) a fifth nucleotide sequence encoding a terminator sequence; (6) a sixth nucleotide sequence encoding an ITR domain; and (7) a seventh nucleotide sequence encoding an miR target site.
[0020] In some embodiments, compositions comprising rAAV particles are administered in high doses (e.g., 1x1014vg / kg). Those of skill in the art will understand that high doses of rAAVs are typically needed to overcome low base editing efficiencies of current base editor systems and the high tropism of rAAV particles for multiple cell types (e.g., non-specific delivery of the rAAV particle to a non-target tissue and / or cell). Accordingly, certain aspects of the present disclosure relate to base editing compositions with improved base editing efficiencies and on-target specificity, relative to existing base editing systems currently known in the art.
[0021] In particular, the inventors of the present disclosure have now discovered that base editing compositions comprising one or more of an evolved deaminase domain (e.g., evoFERNY, evoAPOBEC, and TadCBEd), a guide RNA with flip-and-extend scaffold (F+E), a neuronal promoter (e.g., hSYN) and / or microRNA (miR) binding sites (e.g., miR- 122, miR-183, etc.), or any combination thereof, are capable of reducing base editor expression in non-targeted tissues (e.g., tissues other than neuronal tissue), resulting in about 60% reduction in PrP levels at the target tissue (e.g., brain tissue) with about 7-fold lower total viral titers (e.g., relative to base editors lacking the above recited components). Thus, in some embodiments, the compositions disclosed herein comprise one or more of the following: (1) a first nucleotide sequence encoding a guide RNA with a flip-and-extend scaffold; (2) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N; (3) a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of a split nucleobase editor; (4) a fourth nucleotide sequence encoding a terminator sequence (e.g., a polyT sequence); (5) a fifth nucleotide sequence encoding a UGI domain; (6) a sixth nucleotide sequence encoding an inverted terminal repeat (ITR) domain; or (7) a seventh nucleotide sequence encoding one or more miR target sites. In some embodiments, the one or more nucleotide sequences is operably linked to a promoter, e.g., Cbh, hSYN, and / or U6. In some cases, the one or more nucleotide sequences is operably linked to a hSYN promoter.
[0022] In some embodiments, the compositions comprise one or more polynucleotides comprising one or more of the aforementioned nucleotide sequences. For example, in someembodiments, the one or more polynucleotides comprises one or more of the following: (1) a first nucleotide sequence encoding a guide RNA with a flip-and-extend scaffold; (2) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N; (3) a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of a split nucleobase editor; (4) a fourth nucleotide sequence encoding a terminator sequence (e.g., a polyT sequence); (5) a fifth nucleotide sequence encoding a UGI domain; (6) a sixth nucleotide sequence encoding an inverted terminal repeat (ITR) domain; or (7) a seventh nucleotide sequence encoding one or more miR target sites. In some embodiments, the one or more nucleotide sequences is operably linked to a promoter, e.g., Cbh, hSYN, and / or U6. In some cases, the one or more nucleotide sequences is operably linked to a hSYN promoter.
[0023] In some embodiments, an N-terminal portion of the split nucleobase editor comprises an evolved deaminase. In other embodiments, a C-terminal portion of the split nucleobase editor comprises the evolved deaminase. The evolved deaminase may be an evolved adenosine deaminase (e.g., ABE8e) or an evolved cytidine deaminase (e.g., evoFERNY, evoAPOBEC, and TadCBEd). In some embodiments, one or more polynucleotides comprises one or more nucleotide sequences encoding one or more portions of a nucleobase editor (or split nucleobase editor). In some embodiments, a vector comprises one or more polynucleotide disclosed herein. In some embodiments, the vector is an AAV vector. In some embodiments, the AAV vectors comprises an ITR domain. The ITR domain, according to some embodiments, has a stereotype selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, and / or AAV11). In some embodiments, the one or more nucleotide sequences encode one or more miR target sites. In some cases, the miR target site comprises miR-122 and / or miR-138, although other miR sites are also contemplated herein.
[0024] As mentioned above, in some embodiments, the compositions disclosed herein comprise nucleotide sequences encoding one or more miR target sites used to de-target the base editing composition to non-targeted tissue (e.g., non-neuronal tissue). Any suitable number of target sites (e.g., between 1 and 6) and any suitable number of miRs (e.g., between 1 and 6) may be used to de-target the base editing compositions. In some embodiments, the miR target sites may be encoded within a 3′-untranslated region (UTR) of a nucleobase editor. Without wishing to be bound by any particular theory, it is believed that miRs are small non-coding RNAs of approximately 22 nucleotides in length. They bind to the 3′-UTR of target mRNA based on sequence complementary and result in target mRNA degradation orsuppression in translation. Because many miRs are tissue specific, they can be used to de- target tissues non-specifically infected with one or more rAAV particles (e.g., encoding the base editing compositions disclosed herein). For example, base editing compositions comprising one or more rAAV particles comprising a one or more nucleotide sequences encoding a nucleobase editor and one or more nucleotide sequence encoding miR-122 target sites will not be expressed in the liver cells infected with the rAAV particles because miR- 122 is abundant in the liver, but would be expressed in neuronal cells because miR-122 is not present in neuronal tissue.
[0025] The present disclosure also provides guide RNA sequences comprising a flip-and- extend (F+E) scaffold. Without wishing to be bound by any particular theory, gRNAs with F+E scaffolds are more stable and enhance assembly with the napDNAbp (e.g., Cas9n). The “F” in F+E stands for “flip” and corresponds to an A-U base pair that is flipped in order to remove a putative Pol-III terminator sequence (e.g., four (4) consecutive U’s) from the gRNA stem-loop region. It is believed that such edits avoid premature termination of U6 Pol-III transcription, which increases the concentration of gRNA produced. The “E” in F+E stands for “extend” and corresponds to extension of the Cas9n-binding hairpin structure (e.g., by an additional 5 bp). In some embodiments, the gRNA comprises a flip scaffold only; and in other embodiments, the gRNA comprises an extension scaffold only.
[0026] Additional aspects of the disclosure relate to complexes comprising any one of the base editors and guide RNAs disclosed herein, or vectors comprising a nucleic acid sequence encoding any one of the base editors and guide RNAs of the complexes disclosed herein and / or a nucleic acid sequence encoding any one of the guide RNAs disclosed herein. The disclosure further provides for (1) vectors comprising a nucleic acid sequence encoding the base editors or any portion of the base editors (e.g., split nucleobase editors), (2) cells comprising the any one of the compositions, guide RNAs, complexes, polynucleotides, and / or vectors disclosed herein, (3) pharmaceutical compositions comprising any one of the compositions, guide RNAs, complexes, polynucleotides, vectors, or cells disclosed herein, or (4) kits comprising any one of the compositions, guide RNAs, complexes, polynucleotides, vectors, cells, or pharmaceutical compositions disclosed herein.
[0027] The present disclosure further provides for one or more methods, for example, of using the one or more compositions, guide RNAs, complexes, polynucleotides, vectors, cells, pharmaceutical compositions, and / or kits disclosed herein or the treatment of a disease or disorder associated with a PRNP gene. For example, in some embodiments, the methods relate to producing a truncated PrP protein variant by contacting a cell, or nucleic acidencoding a PRNP gene, with any one of the compositions, guide RNAs, complexes, polynucleotides, vectors, cells, and / or kits disclosed herein. In some embodiments, the method results in recombination of a N-terminal portion of a split nucleobase editor and a C- terminal portion of the nucleobase editor to yield a recombined nucleobase editor. The contacting also produces a guide RNA that directs the recombined nucleobase editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 of a PrP protein encoded by the PRNP gene (e.g., producing a truncated PrP variant), to mutate the start codon of the PrP protein in the PRNP gene to prevent translation or to mutate the 273 bp promoter and promoter-proximal region to prevent transcription factor binding and mRNA transcription therby lowering PrP expression.
[0028] In some embodiments, the methods produce an N-terminally truncated PrP variant. In other embodiments, the methods produce a C-terminally truncated PrP variant. Yet in other embodiments, still, the methods mutate a start codon to prevent initiation of translation.
[0029] Other methods are directed to reducing the templated misfolding of a PrP protein by pathogenic mutants in a subject. Again, as discussed above, it is believed that because truncated PrP proteins, e.g., generated using split nucleobase editors, do not mediate PrP protein misfolding via pathogenic mutants, that truncated variants may be used to reduce templated misfolding included by pathogenic mutants. Such methods would enable one to slow the disease progression of a subject having, or suspected of having, a prion disease or prion-associated disease.
[0030] Other methods disclosed herein are directed to administering to a subject in need thereof (e.g., having or suspected of developing prion disease) a therapeutically effective amount of any one of the compositions, guide RNAs, complexes, polynucleotides, vectors, and / or cells disclosed herein. Methods for treating or preventing prion disease are also contemplated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] FIGs. 1A-1G illustrates the development of base editing strategies to install a stop codon in a PRNP locus. FIG. 1A shows a schematic of the CBE-mediated stop codon installation as a strategy to knockdown cellular prion protein (PrP). The PRNP locus consists of N-terminal (dark blue, amino acids 1-144) and C-terminal (light blue, amino acids 145- 253) domains. The signal peptide (grey, amino acids 1-22), octapeptide repeat (OPR) region (dashed box, amino acids 51-90), and GPI signal (grey, amino acids 230-253) arehighlighted. CBE may convert CAG (Gln), CAA (Gln), CGA (Arg), or TGG (Trp) codons to a stop codon. sgRNA spacers that install stop codons in PRNP evaluated in this study are shown as half-arrows. Truncated PrP no longer templates fibril formation. FIG. 1B illustrates the frequency of the desired stop codon installation or indel formation from candidate sgRNAs using BE4max via plasmid transfection of HEK293T cells. FIG. 1C shows the editing efficiency at bystander positions with BE4max and PRNP R37X sgRNA via plasmid transfection of HEK293T cells. Throughout the application, nomenclature “R37X” refers to a substitution of R at residue 37 of the wildtype PrP protein with “X,” which can be any alternative amino acid in place of R. This nomenclature extends to other mutations as well. Three silent mutations (G35G, S36S, and Y38Y) are possible due to bystander editing from cytosine base editing. FIG. 1D shows a schematic of dual-AAV PHP.eB BE3.9max with PRNP R37X sgRNA. The N-terminal AAV encodes a Cbh promoter, APOBEC deaminase domain, and amino acids 1-572 of SpCas9 fused to NpuN intein. The C-terminal AAV encodes a Cbh promoter, NpuC intein, amino acids 573-1367 of SpCas9, and one copy of the UGI domain. Both AAVs contain a U6 Pol III cassette expressing the PRNP R37-targeting sgRNA. FIG. 1E shows the experimental design for initial assessment of the effect of PRNP base editing on PrP levels. 5-8 week-old Tg25109 mice were treated retro-orbitally with dual- AAV PHP.eB BE3.9max for installation of PRNP R37X at a dose of 1x1014vg / kg. Brain was harvested 100-days post-injection to assess editing efficiency via HTS and PrP protein reduction via ELISA. FIG. 1F shows the editing efficiency of R37X installation in untreated (n=3) and dual-AAV PHP.eB BE3.9max-treated mice (n=3). FIG. 1G, PrP levels in dual- AAV PHP.eB BE3.9max-treated mice (n=3) in the bulk brain hemisphere normalized to those of untreated mice (n=3). Dots represent individual biological replicates (n=3) and error bars.
[0032] FIGs. 2A-2F illustrates that in vivo base editing provides protection from pathogenic human prion challenge. FIG. 2A shows the design of the human pathogenic prion challenge study. Tg25109 mice were divided into two cohorts: a human prion isolate inoculation group and an uninoculated control group. Among the human prion isolate inoculation group, n=13 received dual-AAV PHP.eB BE3.9max PRNP R37X treatment and n=11 received dual-AAV PHP.eB BE3.9max Dnmt1 control treatment. Among the uninoculated control group, n=5 received dual-AAV PHP.eB BE3.9max PRNP R37X treatment and n=1 remained untreated. Mice were treated with AAV at total dose of 1x1014vg / kg at age of 6-9 weeks. One week after AAV treatment, mice were inoculated with either E200K or sCJD prion isolates. After prion inoculation, mice were monitored for weight loss,nest building behavior, and lifespan. Study endpoint was 600-day post prion isolate inoculation (92-95 weeks of age). The uninoculated control group was euthanized to harvest brain hemispheres for analysis via HTS and PrP ELISA. FIG. 2B shows Kaplan-Meier curve of Tg25109 mice inoculated with either the E200K (purple) or sCJD MM1 pathogenic human prion isolate (red). Mice were treated with dual-AAV PHP.eB BE3.9max encoding sgRNAs programmed to install either PRNP R37X or the control Dnmt1 A8T edit. Median survival from each treatment condition is marked. FIG. 2C shows body weight (lines represent mean and shaded areas represent 95% CI) for all timepoints with ≥ 2 animals surviving, and FIG. 2D shows nest building score of Tg25109 mice in the human prion challenge study, fitted to the LOESS model. FIG. 2E shows the frequency of the desired R37X edit and indels, and FIG. 2F shows PrP protein level in the bulk brain hemisphere of mice from the uninoculated control group treated with dual-AAV PHP.eB BE3.9max with PRNP R37X sgRNA (n=5), and in untreated mice from the uninoculated control group (n=1, marked as a white circle with a black dot) or from additional untreated adult Tg25109 mice (n=7, marked as white circles). Dots represent individual biological replicates and error bars represent 95% CI. Significance was calculated by two-tailed Student’s t test. **, P ≤ 0.01; ***, P ≤ 0.001; ****, P ≤ 0.0001. or bars represent 95% CI.
[0033] FIGs. 3A-3G illustrates optimization of CBE strategy for improved potency. FIG. 3A. Frequency of the desired R37X edit and the bystander edits in HEK293T cells after plasmid transfection of the specified CBEs with the PRNP R37X sgRNA. The editors are a fusion of the deaminase domain denoted on the x-axis with SpCas9 (D10A) nickase and two UGI domains. FIG. 3B. Frequency of the desired R37X edit in HEK293T cells after plasmid transfection of TadCBEd and PRNP R37X sgRNAs without or with the specified scaffold modifications. ‘sgRNA’ denotes an sgRNA with the canonical scaffold; ‘F-sgRNA’ denotes an sgRNA with a U-to-A flip scaffold; ‘F+E-sgRNA’ denotes an sgRNA with both a U-to-A flip and a 5-bp stem extension in the scaffold. FIG. 3C. Schematic of dual-AAV PHP.eB BE3.9max or TadCBEd with PRNP R37X sgRNA. The N-terminal AAV (4.9 kb or 4.7 kb, including ITRs) encodes a Cbh promoter, APOBEC or TadCBEd deaminase domain, amino acid 1-572 of SpCas9 fused to NpuN intein. The C-terminal AAV (4.9 kb) encodes a Cbh promoter, NpuC intein, amino acid 573-1367 of SpCas9, and one copy of the UGI domain. Both AAVs contain a U6 Pol III cassette expressing the PRNP R37-targeting sgRNA. 5-8 week-old Tg25109 mice were treated retro-orbitally with AAVs at a dose of 1.5x1013vg / kg. FIG. 3D. Frequency of the desired R37X edit, and FIG. 3E. PrP protein level in the bulk brain hemisphere of Tg25109 mice untreated (n=12), or treated with dual-AAV PHP.eB packagingBE3.9max with PRNP R37X sgRNA (n=8), TadCBEd with PRNP R37X sgRNA (n=8), or TadCBEd with PRNP R37X F+E-sgRNA (n=12) at a total dose of 1.5x1013vg / kg and harvested 35 days post-injection. FIG. 3F. Comparison of the frequency of the desired R37X edit in the bulk brain hemisphere of Tg25109 mice, 35 days (n=8 for each condition) or 100 days (n=6 for each condition) after treatment with either dual-AAV PHP.eB BE3.9max or TadCBEd encoding PRNP sgRNA. FIG. 3G Scatter plot of the frequency of the R37X edit (x-axis) and PrP protein level (y-axis) in the bulk brain hemisphere of Tg25109 mice. Mice were treated with dual-AAV PHP.eB BE3.9max or TadCBEd packaging PRNP R37X sgRNA, harvested 35 days (n=8 for each condition) or 100 days (n=6 for each condition) after treatment for analysis. Linear regression is shown with solid line and 95% CI is shown with dotted lines. Dots represent individual biological replicates (n=3 unless noted otherwise) and error bars represent 95% CI. Significance was calculated by two-tailed Student’s t test. ns, P > 0.05; *, P ≤ 0.05; **, P ≤ 0.01; ****, P ≤ 0.0001.
[0034] FIGs. 4A-4D shows the off-target analysis of the R37X base editing strategy in the human and mouse genome. FIG. 4A shows the percentage of C•G-to-T•A substitution in BE-AAV treated samples above background (untreated samples) at top 100 CIRCLE-seq nominated off-target sites in the human genome (GRCh37). Genomic DNA was extracted from HEK293T cells untreated (n=3) or after 3 days following transfection of plasmids encoding TadCBEd and the PRNP R37X sgRNA (n=3). Off-target editing with P > 0.01 compared to untreated control are labeled with hollow circles, and those with P ≤ 0.01 are labeled with solid circles. Each dot represents mean of three biological replicates. FIGs. 4B and 4C show the percentage of C•G-to-T•A substitution in BE-AAV treated samples above background (untreated samples) at top 100 CIRCLE-seq nominated off-target sites in the mouse genome (GRCm38). Genomic DNAs was extracted from bulk brain hemisphere of Tg25109 mice untreated (n=6), or 35 days (FIG. 4B) and 100 days (FIG. 4C) after treatment with dual-AAV PHP.eB BE3.9max with PRNP R37X sgRNA (n=6), or dual-AAV PHP.eB TadCBEd with PRNP R37X sgRNA (n=6) at a total dose of 1.5x1013vg / kg. Each dot represents mean of six biological replicates. FIG. 4D shows the percentage of C•G-to-T•A substitution in BE-AAV treated samples above background at top 100 CIRCLE-seq nominated off-target sites in the mouse genome (GRCm38). Genomic DNA was extracted from bulk brain hemispheres of Tg25109 mice untreated (n=5), or 600 days after treatment with dual-AAV PHP.eB BE3.9max with PRNP R37X sgRNA (n=5) at a total dose of 1x1014vg / kg. Each dot represents mean of five biological replicates. In all panels, significance wascalculated by one-tailed Student’s t test. Plots showing individual data points and error bars are provided in FIGS. 11-14.
[0035] FIGs. 5A-5F illustrates the engineering of tissue-specific expression of dual-AAV TadCBEd. FIG. 5A shows the frequency of the desired R37X edit, and FIG. 5B shows the PrP protein level in the bulk brain hemisphere of Tg25109 mice treated with dual-AAV PHP.eB TadCBEd encoding PRNP R37X F+E-sgRNA, with Cbh (n=12), hSYN (n=11) or EFS (n=6) promoter driving the expression of the base editor. Data for the “Cbh” condition are equivalent to the “TadCBEd + PRNP R37X F+E-sgRNA” condition shown in FIG. 3D and FIG. 3E, and are replotted for comparison. FIG. 5C shows the frequency of the desired R37X edit in the liver of mice untreated or treated with dual-AAV PHP.eB TadCBEd and PRNP R37X F+E-sgRNA with Cbh promoter (Cbh; n=6), hSYN promoter (hSYN; n=6), hSYN promoter plus miR-183 target site incorporation (hSYN+mi-R183; n=6), hSYN promoter plus miR-122 target site incorporation (hSYN+miR-122; n=5), or hSYN promoter plus miR-183 and miR-122 target site incorporation (hSYN+miR-183+miR-122; n=5). FIG. 5D shows a schematic of dual-AAV PHP.eB TadCBEd with PRNP R37X F+E-sgRNA with hSYN promoter driving the expression of the base editor and miR target site incorporation. ‘miR-183’ construct contains 4 copies of miR-183 target sites; ‘miR-122’ construct contains 3 copies of miR-122 target sites; ‘miR-183+miR-122’ construct contains 3 copies each of miR-183 and miR-122 target sites. FIG. 5E shows the frequency of the desired R37X edit. FIG.5F, PrP protein level in the bulk brain hemisphere of Tg25109 mice harvested 35 days after treatment with dual-AAV PHP.eB TadCBEd with PRNP R37X F+E-sgRNA, with or without the specified miR target site incorporation. Data for the “hSYN” condition are equivalent to the “hSYN” condition shown in FIGs 5A and 5B and are replotted for comparison. Dots represent individual biological replicates and error bars represent 95% CI. Significance was calculated by two-tailed Student’s t test. ns, P > 0.05; *, P ≤ 0.05; ****, P ≤ 0.0001.
[0036] FIGs. 6A-6C shows the in vitro validation of PrP reduction with BE4max- mediated PRNP R37X installation. FIG. 6A shows representative flow cytometry gating plot showing the fluorescence signal in HEK293T cells treated with BE4max, Cas9 nuclease or dead (deaminase-inactive) BE4max editor, with either PRNP R37X sgRNA (blue) or non- targeting sgRNA (red). Six days after plasmid transfection, cells were incubated with a fluorescently conjugated 6D11 antibody that binds PrP and were analyzed by flow cytometry. FIG. 6B shows the mean fluorescence intensity (MFI) of the HEK293T cells after staining with 6D11 antibodies. MFI are plotted after normalizing the MFI of cells treated with PRNPR37X sgRNAs to MFI of cells treated with non-targeting sgRNAs. FIG. 6C shows the frequency of the PRNP R37X edit and indels in HEK293T cells after plasmid transfection with BE4max, Cas9 nuclease or dead BE4max, with PRNP R37-targeting sgRNA. Dots represent individual biological replicates (n=3), and error bars represent 95% CI
[0037] FIGs. 7A-7B shows the optimization of in vivo delivery of the exemplary base editing system. FIG. 7A shows the experimental design for assessment of in vivo administration route. 4-week-old Tg66 mice were treated with dual-AAV PHP.eB BE3.9max for installation of PRNP R37X by systemic administration through retro-orbital injection or direct CNS administration through intracerebroventricular (ICV) stereotaxic injection. Mice treated with retro-orbital injection received a total of 4x1012vg / kg AAV (2x1012vg / kg each for N- and C-terminal AAV) (n=2). Mice treated with ICV injection received a total of 5.5x1011vg / kg AAV (2.5x1011vg / kg each for N- and C-terminal AAV and 5.0x1010vg / kg for AAV encoding EGFP fused to nuclear membrane-localized Klarsicht / ANC-1 / Syne-1 homology (KASH) domain (EGFP:KASH)) (n=2). Mouse brains were harvested five weeks after treatment. GFP-positive nuclei were sorted from ICV-injected samples via FACS to enrich for AAV-transduced cells. Genomic DNA extracted from the brain tissues were analyzed for editing efficiency. FIG. 7B shows the frequency of the desired R37X edit in the bulk brain hemisphere of mice untreated (n=2) or treated with dual-AAV PHP.eB BE3.9max by retro-orbital injection (n=2), or by ICV injection (n=2). Dots represent individual biological replicates and error bars represent standard error of mean (s.e.m).
[0038] FIGs. 8A-8F illustrates the development of a single-AAV cytosine base editor (CBE) delivery strategy. FIG. 8A shows a schematic of single-AAV compatible CBE- mediated stop codon installation strategy and frequency of the desired stop codon installation in HEK293T cells after plasmid transfection of size-minimized CBEs and corresponding sgRNAs. Four size-minimized Cas9 domain (enCjCas9, evoCjCas9, SauriCas9, and eNme2- Cas9) were fused to TadCBEd and a UGI domain to generate candidate size-minimized CBEs. (see Example 1, Note 1). FIG. 8B and 8C show the frequency of the desired Q91X edit (FIG. 8B) and R37X edit (FIG. 8C) in HEK293T cells after plasmid transfection of the base editors and the sgRNAs with the specified modifications. Positions of the poly-U stretch in the enCjCas9 sgRNA scaffold (FIG. 8B) and SauriCas9 sgRNA scaffold (FIG. 8C) are highlighted. “No flip” refers to sgRNA with canonical scaffold; UnF refers to sgRNA with the Un-An position flipped to An-Un (see Supplementary Note 2). FIG. 8D shows a schematic of SauriCas9-TadCBEd, composed of N-terminal NLS (N-NLS), TadCBEd domain, linker 1, SauiCas9 domain, linker 2, UGI domain, and C-terminal NLS (C-NLS).FIG. 8D also shows the frequency of the desired R37X edit in HEK293T cells after plasmid transfection of the SauriCas9-TadCBEd with the specified modification and the PRNP R37X F-guide RNA (see Example 1, Note 3). FIG. 8E shows a schematic of single-AAV PHP.eB SauriCas9-TadCBEd with PRNP R37X F-sgRNA and single-AAV PHP.eB enCjCas9- TadCBEd with PRNP Q91X F-sgRNA. AAVs (5.1 kb and 5.0 kb, respectively, including ITRs) encode an EFS promoter, TadCBEd deaminase domain, SauriCas9 or enCjCas9 domain, a UGI domain, and a U6 Pol III cassette expressing either the PRNP R37-targeting sgRNA or PRNP Q91-targeting sgRNA. FIG. 8F, shows the frequency of the specified stop codon installation in the bulk brain hemisphere of Tg25109 mice, 35 days after treatment with either single-AAV PHP.eB SauriCas9-TadCBEd PRNP R37X F-sgRNA (n=6) or single-AAV PHP.eB enCjCas9-TadCBEd PRNP Q91X F-sgRNA (n=5). AAVs were retro- orbitally administered to 5-8 weeks-old Tg25109 mice at a total dose of 1.5x1013vg / kg. Dots represent individual biological replicates (n=3 unless noted otherwise) and error bars represent 95% CI.
[0039] FIGs. 9A-9J shows the development of an adenosine base editor (ABE)-mediated start codon disruption strategy for truncated PrP reduction. FIG. 9A illustrates a schematic of dual- or single-AAV compatible ABE-mediated start codon disruption strategies and provides the frequency of the desired M1V edit and the bystander edits in HEK293T cells after plasmid transfection of SaCas9-ABE8e with A1 sgRNA, SauriCas9-ABE8e with A2 or A3 sgRNA, and SpCas9-ABE8e with A4 or A5 sgRNA. Spacer sequences of sgRNAs are highlighted, and possible bystander edit positions are marked with arrows. FIG. 9B shows schematics of AAV constructs for dual and single AAV constructs. For dual constructs comprising a dual-AAV PHP.eB SpCas9-ABE8e(V106W) with A5 PRNP M1V F+E- sgRNA, the N-terminal AAV (4.7 kb, including ITRs) encodes a Cbh promoter, ABE8e(V106W) deaminase domain, amino acid 1-572 of SpCas9 fused to NpuN intein, and the C-terminal AAV (4.6 kb, including ITRs) encodes a Cbh promoter, NpuC intein, amino acid 573-1367 of SpCas9. Both AAVs contain a U6 Pol III cassette expressing the PRNP M1-targeting A5 F+E-sgRNA. The single-AAV PHP.eB SauriCas9-ABE8e with A3 PRNP M1V F-guide RNA (5.1 kb, including ITRs) encodes an EFS promoter, ABE8e deaminase domain, SauriCas9, and a U6 Pol III cassette expressing the PRNP M1-targeting A3 F- sgRNA. FIG. 9C shows the frequency of the desired M1V edit, and FIG. 9D shows PrP protein levels in the bulk brain hemisphere of Tg25109 mice either untreated (n=6) or treated with (1) dual-AAV PHP.eB SpCas9-ABE8e(V106W) with A5 PRNP M1V F+E-sgRNA (n=5), or (2) single-AAV PHP.eB SauriCas9-ABE8e with A3 PRNP M1V F-sgRNA (n=6).FIG.9E shows the frequency of the desired M1V edit, and FIG. 9F shows PrP protein level in the bulk brain hemisphere of Tg25109 mice treated with single-AAV PHP.eB SauriCas9- ABE8e with A3 M1V F-sgRNA at 5.0x1012vg / kg (n=6), 1.5x1013vg / kg (n=6) or 4.5x1013vg / kg (n=6) total dose. Data for the “1.5x1013vg / kg” condition is equivalent to the “SauriCas9-ABE8e” condition in FIGs. 9D and 9E. FIG. 9G shows the frequency of the desired M1V edit, and FIG. 9H shows PrP protein level in the bulk brain hemisphere of Tg25109 mice treated with single-AAV PHP.eB SauriCas9-ABE8e with A3 M1V F-sgRNA with either EFS promoter (n=6) or pCALM1 promoter (n=4) driving the expression of the SauriCas9-ABE8e. Data for the “EFS” condition are equivalent to the “SauriCas9-ABE8e” condition in FIGs. 9D and 9E. FIG. 9I shows a representative allele frequency table showing editing outcomes after treatment with a single-AAV PHP.eB SauriCas9-ABE8e with A3 PRNP M1V F-sgRNA. Incidences where bystander edits occur without on-target editing are shown within a box. FIG. 9J shows the total frequency of bystander editing that leads to N3S, N3D, and N3G mutation in genomic DNA harvested from brain hemisphere of Tg25109 mice treated with single-AAV PHP.eB SauriCas9-ABE8e with A3 PRNP M1V F-sgRNA, harvested 35 days post-treatment (n=4). Dots represent individual biological replicates (n=3 unless noted otherwise) and error bars represent 95% CI
[0040] FIGs. 10A-10C show the characterization of tissue-specific expression of dual- AAV PHP.eB TadCBEd with PRNP R37X F+E-guide RNA strategy. FIG 10A shows the viral genome concentration normalized to the diploid genome (vg / dg) in the liver of Tg25109 mice untreated (n=6), or treated with dual-AAV PHP.eB TadCBEd PRNP R37X F+E-sgRNA with the Cbh promoter (Cbh; n=6), hSYN promoter (hSYN; n=6), hSYN promoter with miR- 183 target sites incorporation (hSYN+miR-183; n=6), hSYN promoter with miR-122 target sites incorporation (hSYN+miR-122; n=5), or hSYN promoter with miR-183 and miR-122 target sites incorporation (hSYN+miR-183+miR-122; n=5). ddPCR was performed with probes specific for Cas9 N- and C-terminus, and Actb housekeeping gene. FIGS 10B and 10C show cargo transgene transcripts normalized to Gusb transcript in the liver (FIG. 10B) and DRG (FIG.10C) of mice treated with dual-AAV PHP.eB TadCBEd PRNP R37X F+E- sgRNA, with Cbh promoter (Cbh; n=6), hSYN promoter (hSYN; n=6), or hSYN promoter with miR-183 and miR-122 target site incorporation (hSYN+miR183+miR122; n=4). ddPCR was performed with probes specific for Cas9 N- and C-terminus, and Gusb transcript. Dots represent individual biological replicates and error bars represent 95% CI.
[0041] FIG. 11 shows the off-target editing in HEK293T cells after plasmid transfection of TadCBEd and PRNP R37X sgRNA. The percentage of C•G-to-T•A substitution at the top100 CIRCLE-seq nominated off-target sites in the human genome (GRCh37) is shown; see FIG. 5A. Genomic DNA was extracted from HEK293T cells untreated (n=3) or after 3 days following transfection of plasmids encoding TadCBEd and the PRNP R37X sgRNA (n=3). Dots represent individual biological replicates and error bars represent 95% CI.
[0042] FIG. 12 shows off-target editing in mouse brain tissues harvested 35 days after treatment with dual-AAV PHP.eB BE3.9max and dual-AAV PHP.eB TadCBEd encoding PRNP R37X sgRNA. The percentage of C•G-to-T•A substitution at the top 100 CIRCLE- seq-nominated off-target sites in the mouse genome (GRCm38) is shown; see FIG. 5B. Genomic DNA was extracted from bulk brain hemispheres of Tg25109 mice untreated (n=6), or 35 days after treatment with dual-AAV PHP.eB BE3.9max with PRNP R37X sgRNA (n=6), or 35 days after treatment with dual-AAV PHP.eB TadCBEd with PRNP R37X sgRNA (n=6) at a total dose of 1.5x1013vg / kg. Dots represent individual biological replicates and error bars represent 95% CI.
[0043] FIG. 13 shows the off-target editing in mouse brain tissues harvested 100 days after treatment with dual-AAV PHP.eB BE3.9max and dual-AAV PHP.eB TadCBEd encoding PRNP R37X sgRNA. The percentage of C•G-to-T•A substitution at the top 100 CIRCLE-seq-nominated off-target sites in the mouse genome (GRCm38) is shown; see FIG. 5C. Genomic DNA was extracted from bulk brain hemispheres of Tg25109 mice untreated (n=6), or 100 days after treatment with dual-AAV PHP.eB BE3.9max with PRNP R37X sgRNA (n=6), or 100 days after treatment with dual-AAV PHP.eB TadCBEd with PRNP R37X sgRNA (n=6) at a total dose of 1.5x1013vg / kg. Dots represent individual biological replicates, and error bars represent 95% CI.
[0044] FIG. 14 shows the off-target editing in mouse brain tissues harvested 600 days after treatment with dual-AAV PHP.eB BE3.9max encoding PRNP R37X sgRNA. The percentage of C•G-to-T•A substitution at the top 100 CIRCLE-seq-nominated off-target sites in the mouse genome (GRCm38) is shown; see FIG. 5D. Genomic DNA was extracted from bulk brain hemispheres of Tg25109 mice untreated (n=5), or 600 days after treatment with dual-AAV PHP.eB BE3.9max with PRNP R37X sgRNA (n=5) at a total dose of 1x1014vg / kg. Dots represent individual biological replicates, and error bars represent 95% CI.
[0045] FIG. 15 shows a representative flow cytometry gating used for the analysis of PrP levels in HEK293T cells. FIG. 15A and 15B show representative flow cytometry gating plot for untreated HEK293T cells either without (FIG. 15A) or with (FIG. 15B) incubation using a fluorescently conjugated 6D11 antibody that binds PrP. FIG. 15C shows a representative flow cytometry gating plot for BE4max-treated HEK293T cells incubated with 6D11antibodies. Individual cells were gated based on forward scatter-height (FSC-H) and side scatter-height (SSC-H) ratios, as well as on forward scatter-area (FSC-A) and forward scatter-height (FSC-H) ratios. Cells were gated based on the intensity of the allophycocyanin (APC) signal.
[0046] FIG. 16 shows the percent editing of PRNP at position R37X by various base editors to lower PrP expression.
[0047] FIG. 17 shows a schematic of the proposed target sites within the PRNP gene to reduce PrP expression. The contemplated target sites include, but are not limited to, disruption of the PRNP promoter region, disruption of the M1V start codon, generation of a stop codon at R37X, and disruption of the promoter or promoter-proximal region upstream of exon 1 to prevent transcription factor recruitment and lower mRNA transcription thereby lowering PrP expression.
[0048] FIG. 18 shows a map of the promoter (tailed region) with transcription start site (TSS) and gRNA target sites. The sequences of gRNAs shown are shown in Table 8.
[0049] FIG. 19 shows a plot of the percentage of base editing achieved for each gRNA tested as a function of the PRNP promoter position.
[0050] FIG. 20 shows the editing efficiency of all bases in the PRNP promoter for each gRNA tested (see Table 8). The results from this study allow the activity of individual gRNAs targeting the promoter to be tested to determine which are the most effective for editing.
[0051] FIGs. 21A and 21B describe the base editing of the PRNP promoter with TadCBEd to lower PRNP mRNA levels. FIG. 21A shows the bulk editing results of the PRNP promoter for various gRNAs tested. FIG. 21B shows the fold change in PRNP mRNA expression following bulk editing with TadCBEd base editors, as shown in FIG. 21A.
[0052] FIG. 22A-22C describe the base editing of the PRNP promoter with TadCBEd to lower PrP expression. FIG. 22A shows a western blot of PrP protein expression following gene editing with gRNA306, gRNA307, gRNA310, gRNA310, gRNA311, gRNA312, gRNA315, and gRNA-R37X. HSP90 was included as a loading control. FIG. 22B shows the fold change in PrP expression following base editing the same gRNAs. FIG. 22C shows the relative percentage of C>T edits versus deletions installed by each gRNA.
[0053] FIGs. 23A and 23B. FIG. 23A shows a schematic of the proposed promoter- proximal region upstream of exon 1 within the PRNP gene to reduce PrP expression. FIG. 23B shows the editing efficiency of the promoter-proximal region using various base editors (e.g., spCas9 or SauriCas9) when combined with gRNA 310.
[0054] FIG. 24 describes the base editing of the PRNP promoter-proximal region site using various base editors in combination with gRNA316, gRNA317, or gRNA 318 to lower PrP expression.
[0055] FIG. 25A compares the editing efficiency of edits made at the R37X position, the promoter region, or the promoter-proximal region of the PRNP gene using various base editors. FIG. 25B compares the band intensity from western blot analysis of the resulting PrP protein following base editing at either the R37X position, the promoter or the promoter- proximal region of the PRNP gene. A lower band intensity correlates with lower PrP expression. DEFINITIONS
[0056] As used herein and in the claims, the singular forms “a,” “an,” and “the” include the singular and the plural reference unless the context clearly indicates otherwise. Thus, for example, a reference to “an agent” includes a single agent and a plurality of such agents.
[0057] An “adeno-associated virus” or “AAV” is a virus which infects humans and some other primate species. The wild-type AAV genome is a single-stranded deoxyribonucleic acid (ssDNA), either positive- or negative-sensed. The genome comprises two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs): rep and cap between the ITRs. The rep ORF comprises four overlapping genes encoding Rep proteins required for the AAV life cycle. The cap ORF comprises overlapping genes encoding capsid proteins: VP1, VP2 and VP3, which interact together to form the viral capsid. VP1, VP2 and VP3 are translated from one mRNA transcript, which can be spliced in two different manners: either a longer or shorter intron can be excised resulting in the formation of two isoforms of mRNAs: a ~2.3 kb- and a ~2.6 kb-long mRNA isoform. The capsid forms a supramolecular assembly of approximately 60 individual capsid protein subunits into a non-enveloped, T-1 icosahedral lattice capable of protecting the AAV genome. The mature capsid is composed of VP1, VP2, and VP3 (molecular masses of approximately 87, 73, and 62 kDa respectively) in a ratio of about 1:1:10.
[0058] rAAV particles may comprise a nucleic acid vector (e.g., a recombinant genome), which may comprise at a minimum: (a) one or more heterologous nucleic acid regions comprising a sequence encoding a protein or polypeptide of interest (e.g., a split Cas9 or split nucleobase) or an RNA of interest (e.g., a gRNA), or one or more nucleic acid regions comprising a sequence encoding a Rep protein; and (b) one or more regions comprising inverted terminal repeat (ITR) sequences (e.g., wild-type ITR sequences or engineered ITRsequences) flanking the one or more nucleic acid regions (e.g., heterologous nucleic acid regions). In some embodiments, the nucleic acid vector is between 4 kb and 5 kb in size (e.g., 4.2 to 4.7 kb in size). In some embodiments, the nucleic acid vector further comprises a region encoding a Rep protein. In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, a double-stranded nucleic acid vector may be, for example, a self-complimentary vector that contains a region of the nucleic acid vector that is complementary to another region of the nucleic acid vector, initiating the formation of the double-strandedness of the nucleic acid vector.
[0059] As used herein, the term “adenosine deaminase” or “adenosine deaminase domain” refers to a protein or enzyme that catalyzes a deamination reaction of an adenosine (or adenine). The terms are used interchangeably. In certain embodiments, the disclosure provides nucleobase editor fusion proteins comprising one or more adenosine deaminase domains. For instance, an adenosine deaminase domain may comprise a heterodimer of a first adenosine deaminase and a second deaminase domain, connected by a linker. Adenosine deaminases (e.g., engineered adenosine deaminases or evolved adenosine deaminases) provided herein may be enzymes that convert adenine (A) to inosine (I) in DNA or RNA. Such adenosine deaminase can lead to an A:T to G:C base pair conversion. In some embodiments, the deaminase is a variant of a naturally-occurring deaminase from an organism. In some embodiments, the deaminase does not occur in nature. For example, in some embodiments, the deaminase is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring deaminase.
[0060] In some embodiments, the adenosine deaminase is derived from a bacterium, such as, E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is an E. coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated E. coli TadA deaminase. For example, the truncated ecTadA may be missing one or more N-terminal amino acids relative to a full-length ecTadA. In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5 ,6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full length ecTadA. In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5 ,6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full lengthecTadA. In some embodiments, the ecTadA deaminase does not comprise an N-terminal methionine. Reference is made to U.S. Patent Publication No. 2018 / 0073012, published March 15, 2018, which is incorporated herein by reference.
[0061] In genetics, the “antisense” strand of a segment within double-stranded DNA is the template strand, and which is considered to run in the 3' to 5' orientation. By contrast, the “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. In the case of a DNA segment that encodes a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription, and eventually undergoes (typically, not always) translation into a protein. The antisense strand is thus responsible for the RNA that is later translated to protein, while the sense strand possesses a nearly identical makeup to that of the mRNA. Note that for each segment of dsDNA, there will possibly be two sets of sense and antisense, depending on which direction one reads (since sense and antisense is relative to perspective). It is ultimately the gene product, or mRNA, that dictates which strand of one segment of dsDNA is referred to as sense or antisense.
[0062] “Base editing” refers to genome editing technology that involves the conversion of a specific nucleic acid base into another at a targeted genomic locus. In certain embodiments, this can be achieved without requiring double-stranded DNA breaks (DSB), or single stranded breaks (i.e., nicking). To date, other genome editing techniques, including CRISPR-based systems, begin with the introduction of a DSB at a locus of interest. Subsequently, cellular DNA repair enzymes mend the break, commonly resulting in random insertions or deletions (indels) of bases at the site of the DSB. However, when the introduction or correction of a point mutation at a target locus is desired rather than stochastic disruption of the entire gene, these genome editing techniques are unsuitable, as correction rates are low (e.g., typically 0.1% to 5%), with the major genome editing products being indels. In order to increase the efficiency of gene correction without simultaneously introducing random indels, the present inventors previously modified the CRISPR / Cas9 system to directly convert one DNA base into another without DSB formation. See, Komor, A.C., et al., Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016), the entire contents of which is incorporated by reference herein.
[0063] The terms “base editor (BE)” and “nucleobase editor,” which are used interchangeably herein, refer to an agent comprising a polypeptide that is capable of making amodification to a base (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA) that converts one base to another (e.g., A to G, A to C, A to T, C to T, C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G). In some embodiments, the nucleobase editor is capable of deaminating a base within a nucleic acid such as a base within a DNA molecule. In the case of an adenine nucleobase editor, the nucleobase editor is capable of deaminating an adenine (A) in DNA. Such nucleobase editors may include a nucleic acid programmable DNA binding protein (napDNAbp) fused to an adenosine deaminase. Some nucleobase editors include CRISPR-mediated fusion proteins that are utilized in the base editing methods described herein. In some embodiments, the nucleobase editor comprises a nuclease-inactive Cas9 (dCas9) fused to a deaminase which binds a nucleic acid in a guide RNA-programmed manner via the formation of an R-loop, but does not cleave the nucleic acid. For example, the dCas9 domain of the fusion protein may include a D10A and a H840A mutation (which renders Cas9 capable of cleaving only one strand of a nucleic acid duplex), as described in PCT / US2016 / 058344, which published as WO 2017 / 070632 on April 27, 2017, and is incorporated herein by reference in its entirety. The DNA cleavage domain of S. pyogenes Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA (the “targeted strand”, or the strand in which editing or deamination occurs), whereas the RuvC1 subdomain cleaves the non-complementary strand containing the PAM sequence (the “non-edited strand”). The RuvC1 mutant D10A generates a nick in the targeted strand, while the HNH mutant H840A generates a nick on the non-edited strand (see Jinek et al., Science, 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83 (2013)).
[0064] In some embodiments, a nucleobase editor is a macromolecule or macromolecular complex that results primarily (e.g., more than 80%, more than 85%, more than 90%, more than 95%, more than 99%, more than 99.9%, or 100%) in the conversion of a nucleobase in a polynucleic acid sequence into another nucleobase (i.e., a transition or transversion) using a combination of 1) a nucleotide-, nucleoside-, or nucleobase-modifying enzyme and 2) a nucleic acid binding protein that can be programmed to bind to a specific nucleic acid sequence.
[0065] In some embodiments, the nucleobase editor comprises a DNA binding domain (e.g., a programmable DNA binding domain such as a dCas9 or nCas9) that directs it to a target sequence. In some embodiments, the nucleobase editor comprises a nucleobase modification domain fused to a programmable DNA binding domain (e.g., a dCas9 or nCas9). The terms “nucleobase modifying enzyme” and “nucleobase modification domain,”which are used interchangeably herein, refer to an enzyme that can modify a nucleobase and convert one nucleobase to another (e.g., a deaminase such as a cytidine deaminase or a adenosine deaminase). The nucleobase modifying enzyme of the nucleobase editor may target cytosine (C) bases in a nucleic acid sequence and convert the C to thymine (T) base. In some embodiments, C to T editing is carried out by a deaminase, e.g., a cytidine deaminase. In some embodiments, A to G editing is carried out by a deaminase, e.g., an adenosine deaminase. Nucleobase editors that can carry out other types of base conversions (e.g., C to G) are also contemplated.
[0066] A “split nucleobase editor” refers to a nucleobase editor that is provided as an N- terminal portion (also referred to as a N-terminal half) and a C-terminal portion (also referred to as a C-terminal half) encoded by two separate nucleic acids. The polypeptides corresponding to the N-terminal portion and the C-terminal portion of the nucleobase editor may be combined to form a complete nucleobase editor. In some embodiments, for a nucleobase editor that comprises a dCas9 or nCas9, the “split” is located in the dCas9 or nCas9 domain, at positions as described herein in the split Cas9. Accordingly, in some embodiments, the N-terminal portion of the nucleobase editor contains the N-terminal portion of the split Cas9, and the C-terminal portion of the nucleobase editor contains the C-terminal portion of the split Cas9. Similarly, intein-N or intein-C may be fused to the N-terminal portion or the C-terminal portion of the nucleobase editor, respectively, for the joining of the N- and C-terminal portions of the nucleobase editor to form a complete nucleobase editor.
[0067] In some embodiments, a nucleobase editor converts a C to a T. In some embodiments, the nucleobase editor comprises a cytosine deaminase. A “cytosine deaminase”, or “cytidine deaminase,” refers to an enzyme that catalyzes the chemical reaction “cytosine + H2O → uracil + NH3” or “5-methyl-cytosine + H2O → thymine + NH3.” As it may be apparent from the reaction formula, such chemical reactions result in a C to U / T nucleobase change. In the context of a gene, such a nucleotide change, or mutation, may in turn lead to an amino acid change in the protein, which may affect the protein’s function, e.g., loss-of-function or gain-of-function. In some embodiments, the C to T nucleobase editor comprises a dCas9 or nCas9 fused to a cytidine deaminase. In some embodiments, the cytidine deaminase domain is fused to the N-terminus of the dCas9 or nCas9. In some embodiments, the nucleobase editor further comprises a domain that inhibits uracil glycosylase, and / or a nuclear localization signal. Such nucleobase editors have been described in the art, e.g., in Rees & Liu, Nat Rev Genet. 2018;19(12):770-788 and Koblan et al., Nat Biotechnol. 2018;36(9):843-846; as well as. U.S. Patent Publication No.2018 / 0073012, published March 15, 2018, which issued as U.S. Patent No. 10,113,163; on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published May 4, 2017, which issued as U.S. Patent No. 10,167,457 on January 1, 2019; PCT Publication No. WO 2017 / 070633, published April 27, 2017; U.S. Patent Publication No. 2015 / 0166980, published June 18, 2015; U.S. Patent No. 9,840,699, issued December 12, 2017; U.S. Patent No. 10,077,453, issued September 18, 2018; PCT Publication No. WO 2019 / 023680, published January 31, 2019; PCT Publication No. WO 2018 / 0176009, published September 27, 2018, PCT Application No PCT / US2019 / 033848, filed May 23, 2019, PCT Application No. PCT / US2019 / 47996, filed August 23, 2019; PCT Application No. PCT / US2019 / 049793, filed September 5, 2019; International Patent Application No. PCT / US2020 / 028568, filed April 17, 2020; PCT Application No. PCT / US2019 / 61685, filed November 15, 2019; PCT Application No. PCT / US2019 / 57956, filed October 24, 2019; PCT Publication No. PCT / US2019 / 58678, filed October 29, 2019, the contents of each of which are incorporated herein by reference in their entireties.
[0068] In some embodiments, a nucleobase editor converts an A to a G. In some embodiments, the nucleobase editor comprises an adenosine deaminase. An “adenosine deaminase” is an enzyme involved in purine metabolism. It is needed for the breakdown of adenosine from food and for the turnover of nucleic acids in tissues. Its primary function in humans is the development and maintenance of the immune system. An adenosine deaminase catalyzes hydrolytic deamination of adenosine (forming inosine, which base pairs as G) in the context of DNA. There are no known natural adenosine deaminases that act on DNA. Instead, known adenosine deaminase enzymes only act on RNA (tRNA or mRNA). Evolved deoxyadenosine deaminase enzymes that accept DNA substrates and deaminate dA to deoxyinosine have been described, e.g., in PCT Application PCT / US2017 / 045381, filed August 3, 2017, which published as WO 2018 / 027078, PCT Application No. PCT / US2019 / 033848, which published as WO 2019 / 226953, PCT Application No PCT / US2019 / 033848, filed May 23, 2019, and PCT Patent Application No. PCT / US2020 / 028568, filed April 17, 2020; each of which is herein incorporated by reference by reference.
[0069] Exemplary adenosine and cytidine nucleobase editors are also described in Rees & Liu, Base editing: precision chemistry on the genome and transcriptome of living cells, Nat. Rev. Genet. 2018;19(12):770-788; as well as U.S. Patent Publication No. 2018 / 0073012, published March 15, 2018, which issued as U.S. Patent No. 10,113,163, on October 30, 2018; U.S. Patent Publication No. 2017 / 0121693, published May 4, 2017, which issued as U.S.Patent No. 10,167,457 on January 1, 2019; PCT Publication No. WO 2017 / 070633, published April 27, 2017; U.S. Patent Publication No. 2015 / 0166980, published June 18, 2015; U.S. Patent No. 9,840,699, issued December 12, 2017; and U.S. Patent No. 10,077,453, issued September 18, 2018, the contents of each of which are incorporated herein by reference in their entireties.
[0070] The term “Cas9” or “Cas9 nuclease” refers to an RNA-guided nuclease comprising a Cas9 domain, or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or the gRNA binding domain of Cas9). A “Cas9 domain” as used herein, is a protein fragment comprising an active or inactive cleavage domain of Cas9 and / or the gRNA binding domain of Cas9. A “Cas9 protein” is a full length Cas9 protein. A Cas9 nuclease is also referred to sometimes as a casn1 nuclease or a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and a Cas9 domain. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA target complementary to the spacer. The target strand not complementary to crRNA is first cut endonucleolytically, then trimmed 3′-5′ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gNRA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the entire contents of which are hereby incorporated by reference. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self versus non-self. Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and hostfactor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816- 821(2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, a Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.
[0071] As used herein, the term “nCas9” or “Cas9 nickase” refers to a Cas9 or a variant thereof, which cleaves or nicks only one of the strands of a target cut site thereby introducing a nick in a double strand DNA molecule rather than creating a double strand break. This can be achieved by introducing appropriate mutations in a wild-type Cas9 which inactivates one of the two endonuclease activities of the Cas9. Any suitable mutation which inactivates one Cas9 endonuclease activity but leaves the other intact is contemplated, such as one of D10A or H840A mutations in the wild-type S. pyogenes Cas9 amino acid sequence, or a D10A mutation in the wild-type S. aureus Cas9 amino acid sequence, may be used to form the nCas9.
[0072] The term “cDNA” refers to a strand of DNA copied from an RNA template. cDNA is complementary to the RNA template.
[0073] As used herein, a “cytosine deaminase” encoded by the CDA gene is an enzyme that catalyzes the removal of an amine group from cytidine (i.e., the base cytosine when attached to a ribose ring) to uridine (C to U) and deoxycytidine to deoxyuridine (C to U). A non-limiting example of a cytosine deaminase is APOBEC1 (“apolipoprotein B mRNA editing enzyme, catalytic polypeptide 1”). Another example is AID (“activation-induced cytosine deaminase”). Under standard Watson-Crick hydrogen bond pairing, a cytosine base hydrogen bonds to a guanine base. When cytidine is converted to uridine (or deoxycytidine is converted to deoxyuridine), the uridine (or the uracil base of uridine) undergoes hydrogen bond pairing with the base adenine. Thus, a conversion of “C” to uridine (“U”) by cytosine deaminase will cause the insertion of “A” instead of a “G” during cellular repair and / orreplication processes. Since the adenine “A” pairs with thymine “T”, the cytosine deaminase in coordination with DNA replication causes the conversion of a C·G pairing to a T·A pairing in the double-stranded DNA molecule.
[0074] “CRISPR” is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of prior infections by a virus that have invaded the prokaryote. The snippets of DNA are used by the prokaryotic cell to detect and destroy DNA from subsequent attacks by similar viruses and effectively compose, along with an array of CRISPR-associated proteins (including Cas9 and homologs thereof) and CRISPR-associated RNA, a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA target complementary to the RNA. Specifically, the target strand not complementary to crRNA is first cut endonucleolytically, then trimmed 3´-5′ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gRNA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species – the guide RNA. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the entire contents of which is hereby incorporated by reference. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self versus non-self. CRISPR biology, as well as Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species,including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.
[0075] The term “deaminase” or “deaminase domain” refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine (or adenine) deaminase, which catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA) to inosine. In other embodiments, the deaminase is a cytidine (or cytosine) deaminase, which catalyzes the hydrolytic deamination of cytidine or cytosine.
[0076] The deaminases provided herein may be from any organism, such as a bacterium. In some embodiments, the deaminase or deaminase domain is a variant of a naturally- occurring deaminase from an organism. In some embodiments, the deaminase or deaminase domain does not occur in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring deaminase.
[0077] As used herein, the term “DNA binding protein” or “DNA binding protein domain” refers to any protein that localizes to and binds a specific target DNA nucleotide sequence (e.g. a gene locus of a genome). This term embraces RNA-programmable proteins, which associate (e.g. form a complex) with one or more nucleic acid molecules (i.e., which includes, for example, guide RNA in the case of Cas systems) that direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., DNA sequence) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) associated with the protein. Exemplary RNA-programmable proteins are CRISPR- Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g. engineered or modified), and may include a Cas9 equivalent from any type of CRISPR system (e.g. type II, V, VI), including Cpf1 (a type-V CRISPR-Cas systems) (now known as Cas12a), C2c1 (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system), C2c3 (a type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14,Argonaute, and nCas9. Further Cas-equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), the contents of which are incorporated herein by reference.
[0078] The term “DNA editing efficiency,” as used herein, refers to the number or proportion of intended base pairs that are edited. For example, if a nucleobase editor edits 10% of the base pairs that it is intended to target (e.g., within a cell or within a population of cells), then the nucleobase editor can be described as being 10% efficient. Some aspects of editing efficiency embrace the modification (e.g. deamination) of a specific nucleotide within DNA, without generating a large number or percentage of insertions or deletions (i.e., indels). It is generally accepted that editing while generating less than 5% indels (as measured over total target nucleotide substrates) is high editing efficiency. The generation of more than 20% indels is generally accepted as poor or low editing efficiency. Indel formation may be measured by techniques known in the art, including high-throughput screening of sequencing reads.
[0079] The term “off-target editing frequency,” as used herein, refers to the number or proportion of unintended base pairs, e.g. DNA base pairs, that are edited. On-target and off- target editing frequencies may be measured by the methods and assays described herein, further in view of techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves the hybridization of nucleic acid primers (e.g., DNA primers) with complementarity to nucleic acid (e.g., DNA) regions just upstream or downstream of the target sequence or off-target sequence of interest. Because the DNA target sequence and the Cas9-independent off-target sequences are known a priori in the methods disclosed herein, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the target sequence and Cas9-independent off-target sequences of interest may be designed using techniques known in the art, such as the PhusionU PCR kit (Life Technologies), Phusion HS II kit (Life Technologies), and Illumina MiSeq kit. Since many of the Cas9-dependent off-target sites have high sequence identity to the target site of interest, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9-dependent off-target site may likewise be designed using techniques and kits known in the art. These kits make use of polymerase chain reaction (PCR) amplification, which produces amplicons as intermediate products. The target and off- target sequences may comprise genomic loci that further comprise protospacers and PAMs. Accordingly, the term “amplicons,” as used herein, may refer to nucleic acid molecules that constitute the aggregates of genomic loci, protospacers and PAMs. High-throughputsequencing techniques used herein may further include Sanger sequencing and Illumina- based next-generation genome sequencing (NGS).
[0080] The term “on-target editing,” as used herein, refers to the introduction of intended modifications (e.g., deamination’s) to nucleotides (e.g., adenine) in a target sequence, such as using the nucleobase editors described herein. The term “off-target DNA editing,” as used herein, refers to the introduction of unintended modifications (e.g. deamination’s) to nucleotides (e.g. adenine) in a sequence outside the canonical nucleobase editor binding window (i.e., from one protospacer position to another, typically 2 to 8 nucleotides long). Off-target DNA editing can result from weak or non-specific binding of the gRNA sequence to the target sequence.
[0081] As used herein, the terms “upstream” and “downstream” are terms of relativity that define the linear position of at least two elements located in a nucleic acid molecule (whether single or double-stranded) that is orientated in a 5ʹ-to-3ʹ direction. In particular, a first element is upstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 5ʹ to the second element. For example, a SNP is upstream of a Cas9-induced nick site if the SNP is on the 5ʹ side of the nick site. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 3ʹ to the second element. For example, a SNP is downstream of a Cas9-induced nick site if the SNP is on the 3 ʹ side of the nick site. The nucleic acid molecule can be a DNA (double or single stranded). RNA (double or single stranded), or a hybrid of DNA and RNA. The analysis is the same for single strand nucleic acid molecule and a double strand molecule since the terms upstream and downstream are in reference to only a single strand of a nucleic acid molecule, except that one needs to select which strand of the double stranded molecule is being considered. Often, the strand of a double stranded DNA which can be used to determine the positional relativity of at least two elements is the “sense” or “coding” strand. In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5ʹ to 3ʹ, and which is complementary to the antisense strand of DNA, or template strand, which runs from 3ʹ to 5ʹ. Thus, as an example, a SNP nucleobase is “downstream” of a promoter sequence in a genomic DNA (which is double-stranded) if the SNP nucleobase is on the 3ʹ side of the promoter on the sense or coding strand.
[0082] The term “effective amount,” as used herein, refers to an amount of a biologically active agent that is sufficient to elicit a desired biological response. For example, in some embodiments, an effective amount of a nucleobase editor may refer to the amount of theeditor that is sufficient to edit a target site nucleotide sequence, e.g., a genome. In some embodiments, an effective amount of a nucleobase editor provided herein, e.g., of a fusion protein comprising a nickase Cas9 domain and a guide RNA may refer to the amount of the fusion protein that is sufficient to induce editing of a target site specifically bound and edited by the fusion protein. As will be appreciated by the skilled artisan, the effective amount of an agent, e.g., a fusion protein, a nuclease, a hybrid protein, a protein dimer, a complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide, may vary depending on various factors as, for example, on the desired biological response, e.g., on the specific allele, genome, or target site to be edited, on the cell or tissue being targeted, and on the agent being used.
[0083] The term “fusion protein” as used herein refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C- terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic domain of a nucleic-acid editing protein. Another example includes a Cas9 or equivalent thereof fused to an adenosine deaminase. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.
[0084] Two proteins or protein domains are considered to be “fused” when a peptide bond is formed linking the two proteins or two protein domains. In some embodiments, a linker (e.g., a peptide linker) is present between the two proteins or two protein domains. The term “linker,” as used herein, refers to a chemical group or a molecule linking two molecules or moieties, e.g., two domains of a fusion protein, such as, for example, a nuclease-inactive Cas9 domain and a nucleic acid editing domain (e.g., a deaminase domain). Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In someembodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.
[0085] The term “guide nucleic acid” or “napDNAbp-programming nucleic acid molecule” or equivalently “guide sequence” refers the one or more nucleic acid molecules which associate with and direct or otherwise program a napDNAbp protein to localize to a specific target nucleotide sequence (e.g., a gene locus of a genome) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) associated with the protein, thereby causing the napDNAbp protein to bind to the nucleotide sequence at the specific target site. A non-limiting example is a guide RNA of a Cas protein of a CRISPR- Cas genome editing system. Chemically, guide nucleic acids can be all RNA, all DNA, or a chimeric of RNA and DNA. The guide nucleic acids may also include nucleotide analogs. Guide nucleic acids can be expressed as transcription products or can be synthesized.
[0086] As used herein, a “guide RNA” can refer to a synthetic fusion of the endogenous bacterial crRNA and tracrRNA that provides both targeting specificity and a scaffold and / or binding ability for Cas9 nuclease to a target DNA. This synthetic fusion does not exist in nature and is also commonly referred to as an sgRNA. However, the term, guide RNA, also embraces equivalent guide nucleic acid molecules that associate with Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and which otherwise program the Cas9 equivalent to localize to a specific target nucleotide sequence. The Cas9 equivalents may include other napDNAbps from any type of CRISPR system (e.g., type II, V, VI), including Cpf1 (a type-V CRISPR-Cas systems) (now known as Cas12a), C2c1 (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system) and C2c3 (a type V CRISPR-Cas system). Further Cas- equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences are and structures of guide RNAs are provided herein. In addition, methods for designing appropriate guide RNA sequences are provided herein.
[0087] A guide RNA is a particular type of guide nucleic acid which is mostly commonly associated with a Cas protein of a CRISPR-Cas9 and which associates with Cas9, directing the Cas9 protein to a specific sequence in a DNA molecule that includes complementarity tothe protospacer sequence for the guide RNA. Functionally, guide RNAs associate with Cas9, directing (or programming) the Cas9 protein to a specific sequence in a DNA molecule that includes a sequence complementary to the protospacer sequence for the guide RNA. A gRNA is a component of the CRISPR / Cas system. Typically, a guide RNA comprises a fusion of a CRISPR-targeting RNA (crRNA) and a trans-activation crRNA (tracrRNA), providing both targeting specificity and scaffolding / binding ability for Cas9 nuclease. A “crRNA” is a bacterial RNA that confers target specificity and requires tracrRNA to bind to Cas9. A “tracrRNA” is a bacterial RNA that links the crRNA to the Cas9 nuclease and typically can bind any crRNA. The sequence specificity of a Cas DNA-binding protein is determined by gRNAs, which have nucleotide base-pairing complementarity to target DNA sequences. The native gRNA comprises a 20 nucleotide (nt) spacer sequence which specifies the DNA sequence to be targeted, and is immediately followed by an 80 nt scaffold sequence, which associates the gRNA with Cas9. In some embodiments, a spacer of the present disclosure has a length of 15 to 100 nucleotides, or more. For example, a spacer may have a length of 15 to 90, 15 to 85, 15 to 80, 15 to 75, 15 to 70, 15 to 65, 15 to 60, 15 to 55, 15 to 50, 15 to 45, 15 to 40, 15 to 35, 15 to 30, or 15 to 20 nucleotides. In some embodiments, the spacer is 20 nucleotides long. For example, the spacer may be 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. At least a portion of the target DNA sequence is complementary to the spacer of the gRNA. For Cas9 to successfully bind to the DNA target sequence, a region of the target sequence is complementary to the spacer of the gRNA sequence and is immediately followed by the correct protospacer adjacent motif (PAM) sequence (e.g., NGG for Cas9 and TTN, TTTN, or YTN for Cpf1). In some embodiments, a spacer is 100% complementary to its target sequence. In some embodiments, the spacer sequence is less than 100% complementary to its target sequence and is, thus, considered to be partially complementary to its target sequence. For example, a targeting sequence may be 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% complementary to its target sequence. In some embodiments, the spacer of template DNA or target DNA may differ from a complementary region of a gRNA by 1, 2, 3, 4 or 5 nucleotides.
[0088] In some embodiments, the guide RNA is about 15-120 nucleotides long and comprises a sequence of at least 10 contiguous nucleotides that is complementary to a target sequence. In some embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 97, 98, 99, 100,101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 nucleotides long. In some embodiments, the guide RNA comprises a sequence of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more contiguous nucleotides that is complementary to a target sequence. Sequence complementarity refers to distinct interactions between adenine and thymine (DNA) or uracil (RNA), and between guanine and cytosine.
[0089] As used herein, a “spacer sequence” is the sequence of the guide RNA (~20 nts in length) which has the same sequence (with the exception of uridine bases in place of thymine bases) as the protospacer of the PAM strand of the target (DNA) sequence, and which is complementary to the target strand (or non-PAM strand) of the target sequence.
[0090] As used herein, the “target sequence” refers to the ~20 nucleotides in the target DNA sequence that have complementarity to the protospacer sequence in the PAM strand. The target sequence is the sequence that anneals to or is targeted by the spacer sequence of the guide RNA. The spacer sequence of the guide RNA and the protospacer have the same sequence (except the spacer sequence is RNA, and the protospacer is DNA).
[0091] As used herein, the terms “guide RNA core,” “guide RNA scaffold sequence,” and “backbone sequence,” which are used interchangeably, refer to the region (or sequence) within the gRNA that is responsible for Cas9 binding. It does not include the 20 bp spacer sequence that is used to guide Cas9 to target DNA. This region also known as the crRNA / tracrRNA. The guide RNA backbone sequence is separate from the guide sequence, or spacer, region of the guide RNA, which has complementarity to a protospacer of a nucleic acid molecule.
[0092] As used herein, the term “protospacer” refers to the sequence (e.g., a ~20 bp sequence) in DNA adjacent to the PAM (protospacer adjacent motif) sequence which shares the same sequence as the spacer sequence of the guide RNA, and which is complementary to the target sequence of the non-PAM strand. The spacer sequence of the guide RNA anneals to the target sequence located on the non-PAM strand. In order for Cas9 to function it also requires a specific protospacer adjacent motif (PAM) that varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes a PAM sequence of NGG that is found directly downstream of the protospacer sequence in the genomic DNA, on the non-target strand. The skilled person will appreciate that the literature in the state of the art sometimes refers to the “protospacer” as the ~20-nt target-specific guide sequence on the guide RNA itself, rather than referring to it as a “spacer” (and that the protospacer (DNA) and the spacer (RNA) have the same sequence). Thus, the term “protospacer” as used herein may be used interchangeably with the term“spacer.” The context of the description surrounding the appearance of either “protospacer” or “spacer” will help inform the reader as to whether the term is refence to the gRNA or the DNA sequence. Both usages of these terms are acceptable since the state of the art uses both terms in each of these ways.
[0093] A “protospacer adjacent motif” (PAM) is typically a sequence of nucleotides located adjacent to (e.g., within 10, 9, 8, 7, 6, 5, 4, 3, 3, or 1 nucleotide(s) of a target sequence). A PAM sequence is “immediately adjacent to” a target sequence if the PAM sequence is contiguous with the target sequence (that is, if there are no nucleotides located between the PAM sequence and the target sequence). In some embodiments, a PAM sequence is a wild-type PAM sequence. Examples of PAM sequences include, without limitation, NGG, NGR, NNGRR(T / N), NNNNGATT, NNAGAAW, NGGAG, NAAAAC, AWG, and CC. In some embodiments, a PAM sequence is obtained from Streptococcus pyogenes (e.g., NGG or NGR). In some embodiments, a PAM sequence is obtained from Staphylococcus aureus (e.g., NNGRR(T / N)). In some embodiments, a PAM sequence is obtained from Neisseria meningitidis (e.g., NNNNGATT). In some embodiments, a PAM sequence is obtained from Streptococcus thermophilus (e.g., NNAGAAW or NGGAG). In some embodiments, a PAM sequence is obtained from Treponema denticola (e.g., NAAAAC). In some embodiments, a PAM sequence is obtained from Escherichia coli (e.g., AWG). In some embodiments, a PAM sequence is obtained from Pseudomonas aeruginosa (e.g., CC). Other PAM sequences are contemplated. A PAM sequence is typically located downstream (i.e., 3′) from the target sequence, although in some embodiments a PAM sequence may be located upstream (i.e., 5′) from the target sequence.
[0094] The term “host cell,” as used herein, refers to a cell that can host, replicate, and transfer a phage vector useful for a continuous evolution process as provided herein. In embodiments where the vector is a viral vector, a suitable host cell is a cell that may be infected by the viral vector, can replicate it, and can package it into viral particles that can infect fresh host cells. A cell can host a viral vector if it supports expression of genes of viral vector, replication of the viral genome, and / or the generation of viral particles. One criterion to determine whether a cell is a suitable host cell for a given viral vector is to determine whether the cell can support the viral life cycle of a wild-type viral genome that the viral vector is derived from. In some embodiments, the viral vector is a phage and the host cell is a bacterial cell. In some embodiments, the host cell is an E. coli cell. Suitable E. coli host strains will be apparent to those of skill in the art, and include, but are not limited to, New England Biolabs (NEB) Turbo, Top10F’, DH12S, ER2738, ER2267, and XL1-Blue MRF’.These strain names are art recognized and the genotype of these strains has been well characterized. It should be understood that the above strains are exemplary only and that the invention is not limited in this respect. The term “fresh,” as used herein interchangeably with the terms “non-infected” or “uninfected” in the context of host cells, refers to a host cell that has not been infected by a viral vector comprising a gene of interest as used in a continuous evolution process provided herein. A fresh host cell can, however, have been infected by a viral vector unrelated to the vector to be evolved or by a vector of the same or a similar type but not carrying the gene of interest.
[0095] In some embodiments, the host cell is a prokaryotic cell, for example, a bacterial cell. In some embodiments, the host cell is an E. coli cell. In some embodiments, the host cell is a eukaryotic cell, for example, a yeast cell, a plant cell, an insect cell, or a mammalian cell. In some embodiments, the cell is a human cell. The type of host cell, will, of course, depend on the viral vector employed, and suitable host cell / viral vector combinations will be readily apparent to those of skill in the art.
[0096] An “intein” is a segment of a protein that is able to excise itself and join the remaining portions (the exteins) with a peptide bond in a process known as protein splicing. Inteins are also referred to as “protein introns.” The process of an intein excising itself and joining the remaining portions of the protein is herein termed “protein splicing” or “intein- mediated protein splicing.” In some embodiments, an intein of a precursor protein (an intein containing protein prior to intein-mediated protein splicing) comes from two genes. Such intein is referred to herein as a split intein. For example, in cyanobacteria, DnaE, the catalytic subunit α of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene is herein referred as “intein-N.” The intein encoded by the dnaE-c gene is herein referred as “intein-C.”
[0097] Other intein systems may also be used. For example, a synthetic intein based on the dnaE intein, the Cfa-N and Cfa-C intein pair, has been described (e.g., in Stevens et al., J Am Chem Soc. 2016 Feb 24;138(7):2162-5, incorporated herein by reference). As another example, a synthetic intein based on the dnaE intein, the Nostoc punctiforme (Npu) intein pair, has been described (see Zettler, J., Schutz, V. & Mootz, H. D., The naturally split Npu DnaE intein exhibits an extraordinarily high rate in the protein trans-splicing reaction. FEBS letters 583, 909-914 (2009), incorporated herein by reference). Non-limiting examples of intein pairs that may be used in accordance with the present disclosure include: Cfa DnaE intein, Npu DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyXintein, Rma DnaB intein and Cne Prp8 intein (e.g., as described in US Patent 8,394,604, incorporated herein by reference).
[0098] Exemplary nucleotide and amino acid sequences of inteins are provided below, as SEQ ID NOs: 1-8. In some embodiments, the inteins used in accordance with the disclosed napDNAbp domains (e.g., Cas9 domains) comprise the Npu intein-N comprising the amino acid sequence of SEQ ID NO: 2 and the Npu intein-C comprising the amino acid sequence of SEQ ID NO: 4. In some embodiments, the inteins used in accordance with the disclosed nucleobase editors comprise the Npu intein-N comprising the amino acid sequence of SEQ ID NO: 2 and the Npu intein-C comprising the amino acid sequence of SEQ ID NO: 4. In some embodiments, the inteins used in accordance with the disclosed constructs encoding any of the disclosed napDNAbp domains (e.g., a Cas9 domain) comprise the Npu intein-N DNA comprising the nucleotide sequence of SEQ ID NO: 1 and the Npu intein-C DNA comprising the nucleotide sequence of SEQ ID NO: 3. In some embodiments, the inteins used in accordance with the disclosed constructs encoding any of the disclosed nucleobase editors comprise the Npu intein-N DNA comprising the nucleotide sequence of SEQ ID NO: 1 and the Npu intein-C DNA comprising the nucleotide sequence of SEQ ID NO: 3.
[0099] In some embodiments, the intein-N comprises an amino acid sequence that is at least 90%, 95%, 98%, or 99% identical to the amino acid of SEQ ID NOs: 2 or 6. In some embodiments, the intein-N comprises an amino acid sequence that differs from the amino acid of SEQ ID NOs: 2 or 6 by 1, 2, 3, 4, 5, 6, or 7 amino acids. In some embodiments, the intein-N comprises the nucleotide sequence of SEQ ID NOs: 1 or 5. In some embodiments, the intein-N used in accordance with the disclosed constructs comprises a nucleotide sequence that is at least 90%, 95%, 98%, or 99% identical to the nucleotide sequence of SEQ ID NOs: 1 or 5. In some embodiments, the intein-N used in accordance with the disclosed constructs comprises a nucleotide sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 10-15 nucleotides from the nucleotide sequence of SEQ ID NOs: 1 or 5.
[0100] In some embodiments, the intein-C comprises an amino acid sequence that is at least 90%, 95%, 98%, or 99% identical to the amino acid of SEQ ID NOs: 4 or 8. In some embodiments, the intein-C comprises an amino acid sequence that differs from the amino acid of SEQ ID NOs: 4 or 8 by 1, 2, 3, 4, or 5 amino acids. In some embodiments, the intein- C comprises the nucleotide sequence of SEQ ID NOs: 3 or 7. In some embodiments, the intein-C used in accordance with the disclosed constructs comprises a nucleotide sequence that is at least 90%, 95%, 98%, or 99% identical to the nucleotide sequence of SEQ ID NOs: 3 or 7. In some embodiments, the intein-C used in accordance with the disclosed constructscomprises a nucleotide sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from the nucleotide sequence of SEQ ID NOs: 3 or 7.
[0101] In particular embodiments, the intein-N comprises the amino acid sequence as set forth in SEQ ID NO: 2. In some embodiments, the intein-C comprises the amino acid sequence as set forth in SEQ ID NO: 4.
[0102] DnaE Intein-N DNA:
[0103] TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTG CCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGAT AACAATGGTAAATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGC AGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGG ACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGA GCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT (SEQ ID NO: 1)
[0104] Npu DnaE N-terminal Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEY CLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN (SEQ ID NO: 2)
[0105] DnaE Intein-C DNA:
[0106] ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTAT GATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTT CTAAT (SEQ ID NO: 3)
[0107] Npu DnaE C-terminal Protein:
[0108] MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN (SEQ ID NO: 4)
[0109] Cfa-N DNA:
[0110] TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGC CTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACA AGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACA AGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGA TCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGA GCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA (SEQ ID NO: 5)
[0111] Cfa-N Protein:
[0112] CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRG EQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP (SEQ ID NO: 6)
[0113] Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAA GTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAG TGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC (SEQ ID NO: 7)
[0114] Cfa-C Protein:
[0115] MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLV ASN (SEQ ID NO: 8)
[0116] Intein-N and intein-C may be fused to the N-terminal portion of the split Cas9 and the C-terminal portion of the split Cas9, respectively, for the joining of the N-terminal portion of the split Cas9 and the C-terminal portion of the split Cas9. For example, in some embodiments, an intein-N is fused to the C-terminus of the N-terminal portion of the split Cas9, i.e., to form a structure of N-[N-terminal portion of the split Cas9]-[intein-N]-C. In some embodiments, an intein-C is fused to the N-terminus of the C-terminal portion of the split Cas9, i.e., to form a structure of N-[intein-C]-[C-terminal portion of the split Cas9]-C. The mechanism of intein-mediated protein splicing for joining the proteins the inteins are fused to (e.g., split Cas9) is known in the art, e.g., as described in Shah et al., Chem Sci. 2014; 5(1):446–461, incorporated herein by reference.
[0117] The term “mutation,” as used herein, refers to a substitution of a residue within a sequence, e.g., a nucleic acid or amino acid sequence, with another residue; a deletion or insertion of one or more residues within a sequence; or a substitution of a residue within a sequence of a genome in a subject to be corrected. Mutations are typically described herein by identifying the original residue followed by the position of the residue within the sequence and by the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). Mutations can include avariety of categories, such as single base polymorphisms, microduplication regions, indel, and inversions, and is not meant to be limiting in any way. Mutations can include “loss-of- function” mutations which are mutations that reduce or abolish a protein activity. Most loss- of-function mutations are recessive, because in a heterozygote the second chromosome copy carries an unmutated version of the gene coding for a fully functional protein whose presence compensates for the effect of the mutation. There are some exceptions where a loss-of- function mutation is dominant, one example being haploinsufficiency, where the organism is unable to tolerate the approximately 50% reduction in protein activity suffered by the heterozygote. This is the explanation for a few genetic diseases in humans, including Marfan syndrome, which results from a mutation in the gene for the connective tissue protein called fibrillin. Mutations also embrace “gain-of-function” mutations, which is one which confers an abnormal activity on a protein or cell that is otherwise not present in a normal condition. Many gain-of-function mutations are in regulatory sequences rather than in coding regions, and can therefore have a number of consequences. Because of their nature, gain-of-function mutations are usually dominant. Many loss-of-function mutations are recessive, such as autosomal recessive.
[0118] The term “napDNAbp” which stand for “nucleic acid programmable DNA binding protein” refers to any protein that may associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., which may broadly be referred to as a “napDNAbp- programming nucleic acid molecule” and includes, for example, guide RNA in the case of Cas systems) which direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a gene locus of a genome) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) associated with the protein, thereby causing the protein to bind to the nucleotide sequence at the specific target site. This term napDNAbp embraces CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and may include a Cas9 equivalent from any type of CRISPR system (e.g., type II, V, VI), including Cpf1 (a type-V CRISPR-Cas systems) (now known as Cas12a), C2c1 (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system), C2c3 (a type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, Cas13d, Cas14, Argonaute, and nCas9. Further Cas-equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353 (6299), the contents of which are incorporated herein by reference. However, the nucleic acid programmable DNA bindingprotein (napDNAbp) that may be used in connection with this invention are not limited to CRISPR-Cas systems. The invention embraces any such programmable protein, such as the Argonaute protein from Natronobacterium gregoryi (NgAgo) which may also be used for DNA-guided genome editing. NgAgo-guide DNA system does not require a PAM sequence or guide RNA molecules, which means genome editing can be performed simply by the expression of generic NgAgo protein and introduction of synthetic oligonucleotides on any genomic sequence. See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7):768-73, which is incorporated herein by reference.
[0119] In some embodiments, the napDNAbp is a RNA-programmable nuclease, when in a complex with an RNA, may be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) is referred to as a guide RNA (gRNA). gRNAs can exist as a complex of two or more RNAs, or as a single RNA molecule. gRNAs that exist as a single RNA molecule may be referred to as single-guide RNAs (sgRNAs), though “gRNA” is used interchangeably to refer to guide RNAs that exist as either single molecules or as a complex of two or more molecules. Typically, gRNAs that exist as single RNA species comprise two domains: (1) a domain that shares homology to a target nucleic acid (e.g., and directs binding of a Cas9 (or equivalent) complex to the target); and (2) a domain that binds a Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as a tracrRNA, and comprises a stem-loop structure. For example, in some embodiments, domain (2) is homologous to a tracrRNA as depicted in Figure 1E of Jinek et al., Science 337:816-821(2012), the entire contents of which is incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Patent No. 9,340,799, entitled “mRNA-Sensing Switchable gRNAs,” and International Patent Application No. PCT / US2014 / 054247, filed September 6, 2013, published as WO 2015 / 035136 and entitled “Delivery System For Functional Nucleases,” the entire contents of each are herein incorporated by reference. In some embodiments, a gRNA comprises two or more of domains (1) and (2), and may be referred to as an “extended gRNA.” For example, an extended gRNA will, e.g., bind two or more Cas9 proteins and bind a target nucleic acid at two or more distinct regions, as described herein. The gRNA comprises a nucleotide sequence that complements a target site, which mediates binding of the nuclease / RNA complex to said target site, providing the sequence specificity of the nuclease:RNA complex. In some embodiments, the RNA- programmable nuclease is the (CRISPR-associated system) Cas9 endonuclease, for example Cas9 (Csn1) from Streptococcus pyogenes (see, e.g., “Complete genome sequence of an M1strain of Streptococcus pyogenes.” Ferretti J.J. et al.., Proc. Natl. Acad. Sci. U.S.A. 98:4658- 4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816- 821(2012), the entire contents of each of which are incorporated herein by reference.
[0120] The napDNAbp nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins are able to be targeted, in principle, to any sequence specified by the guide RNA. Methods of using napDNAbp nucleases, such as Cas9, for site- specific cleavage (e.g., to modify a genome) are known in the art (see e.g., Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, W.Y. et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature Biotechnology 31, 227-229 (2013); Jinek, M. et al. RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, J.E. et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic Acid Res. (2013); Jiang, W. et al. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature Biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).
[0121] A “uracil glycosylase inhibitor (UGI)” refers to a protein that inhibits the activity of uracil-DNA glycosylase. Suitable UGI proteins for use in accordance with the present disclosure include, for example, those published in Wang et al., J. Biol. Chem. 264:1163- 1171(1989); Lundquist et al., J. Biol. Chem. 272:21408-21419(1997); Ravishankar et al., Nucleic Acids Res. 26:4880-4887(1998); and Putnam et al., J. Mol. Biol. 287:331-346(1999), each of which is incorporated herein by reference. Non-limiting, exemplary proteins that may be used as a UGI of the present disclosure and their respective sequences are provided below. In some embodiments, the UGI is a variant of a naturally-occurring deaminase from an organism, and the variants do not occur in nature. For example, in some embodiments, the UGI is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring UGI from an organism or any UGIs provided herein (e.g., a UGI comprising the amino acid sequence of any one of SEQ ID NOs: 17-20). In some embodiments, the UGI comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 30%, no more than 25%, no more than 20%, no more than 15%, no more than 10%, no more than 5%, no more than 1% longer or shorter)than any of the UGIs provided herein. In some embodiments, the UGI comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 20 amino acids, no more than 15 amino acids, no more than 10 amino acids, no more than 5 amino acids, no more than 2 amino acids longer or shorter) than any of the UGIs provided herein.
[0122] A “nuclear localization signal” or “NLS” refers to as an amino acid sequence that “tags” a protein for import into the cell nucleus by nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface. One or more NLS may be added to the N- or C-terminus of a protein, or internally (e.g., between two protein domains). For example, one or more NLS may be added to the N- or C-terminus of a nucleobase editor, or between the Cas9 and the deaminase in a nucleobase editor. In some embodiments, 1, 2, 3, 4, 5, or more NLS may be added. Nuclear localization sequences are known in the art and would be apparent to the skilled artisan. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, filed November 23, 2000, the contents of which are incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In some embodiments, a NLS comprises a bipartite nuclear localization signal comprising an amino acid sequence selected from the group consisting of KRTADGSEFEPKKKRKV (SEQ ID NO: 9), KRPAATKKAGQAKKKK (SEQ ID NO: 10), KKTELQTTNAENKTKKL (SEQ ID NO: 11), KRGINDRNFWRGENGRKTR(SEQ ID NO: 12), RKSGKIAAIVVKRPRK(SEQ ID NO: 13), PKKKRKV (SEQ ID NO: 14) or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 15). In some embodiments, a linker is inserted between the Cas9 and the deaminase. In certain embodiments, the NLS comprises the amino acid sequence of SEQ ID NO: 9-15. In some embodiments, the NLS comprises the amino acid sequence of SEQ ID NO: 192-197.
[0123] An NLS can be classified as monopartite or bipartite. A non-limiting example of a monopartite NLS is the sequence PKKKRKV (SEQ ID NO: 14) in the SV40 Large T- antigen. A “bipartite” NLS typically contains two clusters of basic amino acids, separated by a spacer of about 10 amino acids. One non-limiting example of a bipartite NLS is the NLS of nucleoplasmin, KRPAATKKAGQAKKKK (spacer underlined) (SEQ ID NO: 10). In some embodiments, the NLS used in accordance with the present disclosure is the NLS of nucleoplasmin comprising the amino acid sequence of KRPAATKKAGQAKKKK (SEQ ID NO: 10). Other bipartite NLSs that may be used in accordance with the present disclosure include, without limitation: SV40 bipartite NLS (KRTADGSEFESPKKKRKV (SEQ ID NO: 192), e.g., as described in Hodel et al., J Biol Chem. 2001 Jan 12;276(2):1317-25,incorporated herein by reference); Kanadaptin bipartite NLS (KKTELQTTNAENKTKKL (SEQ ID NO: 11), e.g., as described in Hubner et al., Biochem J. 2002 Jan 15;361(Pt 2):287- 96, incorporated herein by reference); influenza A nucleoprotein bipartite NLS (KRGINDRNFWRGENGRKTR (SEQ ID NO: 12), e.g., as described in Ketha et al., BMC Cell Biology. 2008;9:22, incorporated herein by reference); and ZO-2 bipartite NLS (RKSGKIAAIVVKRPRK (SEQ ID NO: 13), e.g., as described in Quiros et al., Nusrat A, ed. Molecular Biology of the Cell. 2013;24(16):2528-2543, incorporated herein by reference).
[0124] The nucleotide sequence encoding an NLS is “operably linked” to the nucleotide sequence encoding a protein to which the NLS is fused (e.g., a Cas9 or a nucleobase editor) when two coding sequences are “in-frame with each other” and are translated as a single polypeptide fusing two sequences.
[0125] Nucleic acids of the present disclosure may include one or more genetic elements. A “genetic element” refers to a particular nucleotide sequence that has a role in nucleic acid expression (e.g., promoter, enhancer, terminator) or encodes a discrete product of an engineered nucleic acid (e.g., a nucleotide sequence encoding a guide RNA, a protein and / or an RNA interference molecule).
[0126] A “promoter” refers to a control region of a nucleic acid sequence at which initiation and rate of transcription of the remainder of a nucleic acid sequence are controlled. A promoter may also contain sub-regions at which regulatory proteins and molecules may bind, such as RNA polymerase and other transcription factors. Promoters may be constitutive, inducible, activatable, repressible, tissue-specific, or any combination thereof. A promoter drives expression or drives transcription of the nucleic acid sequence that it regulates. Herein, a promoter is considered to be “operably linked” when it is in a correct functional location and orientation in relation to a nucleic acid sequence it regulates to control (“drive”) transcriptional initiation and / or expression of that sequence.
[0127] A promoter may be one naturally associated with a gene or sequence, as may be obtained by isolating the 5′ non-coding sequences located upstream of the coding segment of a given gene or sequence. Such a promoter is referred to as an “endogenous promoter.” In some embodiments, a coding nucleic acid sequence may be positioned under the control of a recombinant or heterologous promoter, which refers to a promoter that is not normally associated with the encoded sequence in its natural environment. Such promoters may include promoters of other genes; promoters isolated from any other cell; and synthetic promoters or enhancers that are not “naturally occurring” such as, for example, those that contain different elements of different transcriptional regulatory regions and / or mutations thatalter expression through methods of genetic engineering that are known in the art. In addition to producing nucleic acid sequences of promoters and enhancers synthetically, sequences may be produced using recombinant cloning and / or nucleic acid amplification technology, including polymerase chain reaction (PCR).
[0128] In some embodiments, promoters used in accordance with the present disclosure are “inducible promoters,” which are promoters that are characterized by regulating (e.g., initiating or activating) transcriptional activity when in the presence of, influenced by or contacted by an inducer signal. An inducer signal may be endogenous or a normally exogenous condition (e.g., light), compound (e.g., chemical or non-chemical compound) or protein that contacts an inducible promoter in such a way as to be active in regulating transcriptional activity from the inducible promoter. Thus, a “signal that regulates transcription” of a nucleic acid refers to an inducer signal that acts on an inducible promoter. A signal that regulates transcription may activate or inactivate transcription, depending on the regulatory system used. Activation of transcription may involve directly acting on a promoter to drive transcription or indirectly acting on a promoter by inactivation a repressor that is preventing the promoter from driving transcription. Conversely, deactivation of transcription may involve directly acting on a promoter to prevent transcription or indirectly acting on a promoter by activating a repressor that then acts on the promoter.
[0129] In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. In the case of a DNA segment that encodes a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription, and eventually undergoes (typically, not always) translation into a protein. The antisense strand is thus responsible for the RNA that is later translated to protein, while the sense strand possesses a nearly identical makeup to that of the mRNA. Note that for each segment of dsDNA, there will possibly be two sets of sense and antisense, depending on which direction one reads (since sense and antisense is relative to perspective). It is ultimately the gene product, or mRNA, that dictates which strand of one segment of dsDNA is referred to as sense or antisense.
[0130] The term “subject,” as used herein, refers to an individual organism, for example, an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, a goat, a cow, a cat, or a dog. In some embodiments, the subject is a vertebrate, anamphibian, a reptile, a fish, an insect, a fly, or a nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject may be of either sex and at any stage of development.
[0131] A “subject in need thereof” refers to an individual who has a disease, a sign and / or symptom of a disease, or a predisposition toward a disease, with the purpose to cure, heal, alleviate, relieve, alter, remedy, ameliorate, improve, or affect the disease, the symptom of the disease, or the predisposition toward the disease. In some embodiments, the subject is a mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is human. In some embodiments, the mammal is a rodent. In some embodiments, the rodent is a mouse. In some embodiments, the rodent is a rat. In some embodiments, the mammal is a companion animal. A “companion animal” refers to pets and other domestic animals. Non-limiting examples of companion animals include dogs and cats; livestock, such as horses, cattle, pigs, sheep, goats, and chickens; and other animals, such as mice, rats, guinea pigs, and hamsters.
[0132] The term “target site” refers to a sequence within a nucleic acid molecule that is edited by a base editor (BE) or nucleobase editor disclosed herein. The term “target site,” in the context of a single strand, also can refer to the “target strand” which anneals or binds to the spacer sequence of the guide RNA. The target site can refer, in certain embodiments, to a segment of double-stranded DNA that includes the protospacer (i.e., the strand of the target site that has the same nucleotide sequence as the spacer sequence of the guide RNA) on the PAM-strand (or non-target strand) and target strand, which is complementary to the protospacer and the spacer alike, and which anneals to the spacer of the guide RNA, thereby targeting or programming a Cas9 nucleobase editor to target the target site.
[0133] A “transcriptional terminator” is a nucleic acid sequence that causes transcription to stop. A transcriptional terminator may be unidirectional or bidirectional. It is comprised of a DNA sequence involved in specific termination of an RNA transcript by an RNA polymerase. A transcriptional terminator sequence prevents transcriptional activation of downstream nucleic acid sequences by upstream promoters. A transcriptional terminator may be necessary in vivo to achieve desirable expression levels or to avoid transcription of certain sequences. A transcriptional terminator is considered to be “operably linked to” a nucleotide sequence when it is able to terminate the transcription of the sequence it is linked to.
[0134] The most commonly used type of terminator is a forward terminator. When placed downstream of a nucleic acid sequence that is usually transcribed, a forward transcriptionalterminator will cause transcription to abort. In some embodiments, bidirectional transcriptional terminators are provided, which usually cause transcription to terminate on both the forward and reverse strand. In some embodiments, reverse transcriptional terminators are provided, which usually terminate transcription on the reverse strand only.
[0135] In prokaryotic systems, terminators usually fall into two categories (1) rho- independent terminators and (2) rho-dependent terminators. Rho-independent terminators are generally composed of palindromic sequence that forms a stem loop rich in G-C base pairs followed by several T bases. Without wishing to be bound by theory, the conventional model of transcriptional termination is that the stem loop causes RNA polymerase to pause, and transcription of the poly-A tail causes the RNA:DNA duplex to unwind and dissociate from RNA polymerase.
[0136] In eukaryotic systems, the terminator region may comprise specific DNA sequences that permit site-specific cleavage of the new transcript so as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3′ end of the transcript. RNA molecules modified with this polyA tail appear to more stable and are translated more efficiently. Thus, in some embodiments involving eukaryotes, a terminator may comprise a signal for the cleavage of the RNA. In some embodiments, the terminator signal promotes polyadenylation of the message. The terminator and / or polyadenylation site elements may serve to enhance output nucleic acid levels and / or to minimize read through between nucleic acids.
[0137] Terminators for use in accordance with the present disclosure include any terminator of transcription described herein or known to one of ordinary skill in the art. Examples of terminators include, without limitation, the termination sequences of genes such as, for example, the bovine growth hormone terminator, and viral termination sequences, such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA, and arcA terminator. In some embodiments, the termination signal may be a sequence that cannot be transcribed or translated, such as those resulting from a sequence truncation.
[0138] A “Woodchuck Hepatitis Virus (WHP) Posttranscriptional Regulatory Element (WPRE)” is a DNA sequence that, when transcribed creates a tertiary structure enhancing expression. Commonly used in molecular biology to increase expression of genes delivered by viral vectors. WPRE is a tripartite regulatory element with gamma, alpha, and beta components.
[0139] The full WPRE sequence is 609 bp long:
[0140] GCTTATCGATAATCAACCTCTGGATTACAAAATTTGTGAAAGATTGAC TGGTATTCTTAACTATGTTGCTCCTTTTACGCTATGTGGATACGCTGCTTTAATGC CTTTGTATCATGCTATTGCTTCCCGTATGGCTTTCATTTTCTCCTCCTTGTATAAAT CCTGGTTGCTGTCTCTTTATGAGGAGTTGTGGCCCGTTGTCAGGCAACGTGGCGT GGTGTGCACTGTGTTTGCTGACGCAACCCCCACTGGTTGGGGCATTGCCACCACC TGTCAGCTCCTTTCCGGGACTTTCGCTTTCCCCCTCCCTATTGCCACGGCGGAACT CATCGCCGCCTGCCTTGCCCGCTGCTGGACAGGGGCTCGGCTGTTGGGCACTGAC AATTCCGTGGTGTTGTCGGGGAAATCATCGTCCTTTCCTTGGCTGCTCGCCTATGT TGCCACCTGGATTCTGCGCGGGACGTCCTTCTGCTACGTCCCTTCGGCCCTCAATC CAGCGGACCTTCCTTCCCGCGGCCTGCTGCCGGCTCTGCGGCCTCTTCCGCGTCTT CGCCTTCGCCCTCAGACGAGTCGGATCTCCCTTTGGGCCGCCTCCCCGCATCGAT ACCG (SEQ ID NO: 16).
[0141] The terms “nucleic acid” and “polynucleotide,” as used herein, refer to a compound comprising a nucleobase and an acidic moiety, e.g., a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides are linear molecules, in which adjacent nucleotides are linked to each other via a phosphodiester linkage. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g. nucleotides and / or nucleosides). In some embodiments, “nucleic acid” refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms “oligonucleotide” and “polynucleotide” can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, “nucleic acid” encompasses RNA as well as single and / or double-stranded DNA. Nucleic acids may be naturally occurring, for example, in the context of a genome, a transcript, an mRNA, tRNA, rRNA, siRNA, snRNA, a plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule may be a non-naturally occurring molecule, e.g., a recombinant DNA or RNA, an artificial chromosome, an engineered genome (e.g., an engineered viral vector), an engineered vector, or fragment thereof, or a synthetic DNA, RNA, or DNA / RNA hybrid, optionally including non-naturally occurring nucleotides or nucleosides. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can comprise nucleoside analogssuch as analogs having chemically modified bases or sugars, and backbone modifications. A nucleic acid sequence is presented in the 5′ to 3′ direction unless otherwise indicated. In some embodiments, a nucleic acid is or comprises natural nucleosides (e.g. adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, 5-methylcytidine, 2-aminoadenosine, C5- bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8- oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5′-N-phosphoramidite linkages).
[0142] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein, and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide may also be a single molecule or may be a multi-molecular complex. A protein, peptide, or polypeptide may be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. The term “fusion protein” as used herein refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy- terminal (C-terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy- terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic domain of a nucleic-acid editing protein. In some embodiments, a protein is in a complex with, or is in association with, a nucleic acid, e.g., RNA or DNA. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteinsprovided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), which are incorporated herein by reference.
[0143] The term “recombinant” as used herein in the context of proteins or nucleic acids refers to proteins or nucleic acids that do not occur in nature, but are the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that comprises at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations as compared to any naturally occurring sequence. The fusion proteins (e.g., nucleobase editors) described herein are made by recombinant technology. Recombinant technology is familiar to those skilled in the art.
[0144] The term “pharmaceutically-acceptable carrier” means a pharmaceutically- acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc magnesium, calcium or zinc stearate, or steric acid), or solvent encapsulating material, involved in carrying or transporting the compound from one site (e.g., the delivery site) of the body, to another site (e.g., organ, tissue or portion of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the tissue of the subject (e.g., physiologically compatible, sterile, physiologic pH, etc.).
[0145] “A therapeutically effective amount” as used herein refers to the amount of each therapeutic agent (e.g., nucleobase editor, rAAV) described in the present disclosure required to confer therapeutic effect on the subject, either alone or in combination with one or more other therapeutic agents. Effective amounts vary, as recognized by those skilled in the art, depending on the particular condition being treated, the severity of the condition, the individual subject parameters including age, physical condition, size, gender, and weight, the duration of the treatment, the nature of concurrent therapy (if any), the specific route of administration and like factors within the knowledge and expertise of the health practitioner. These factors are well known to those of ordinary skill in the art and can be addressed with no more than routine experimentation. It is generally preferred that a maximum dose of the individual components or combinations thereof be used, that is, the highest safe dose according to sound medical judgment. It will be understood by those of ordinary skill in theart, however, that a subject may insist upon a lower dose or tolerable dose for medical reasons, psychological reasons or for virtually any other reasons. Empirical considerations, such as the half-life, generally will contribute to the determination of the dosage. For example, therapeutic agents that are compatible with the human immune system, such as polypeptides comprising regions from humanized antibodies or fully human antibodies, may be used to prolong half-life of the polypeptide and to prevent the polypeptide being attacked by the host's immune system.
[0146] The terms “treatment,” “treat,” and “treating,” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. As used herein, the terms “treatment,” “treat,” and “treating” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay onset of a symptom or inhibit onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, for example, to prevent or delay their recurrence.
[0147] As used herein, the term “variant” refers to a protein having characteristics that deviate from what occurs in nature that retains at least one functional i.e. binding, interaction, or enzymatic ability and / or therapeutic property thereof. A “variant” is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the wild type protein. For instance, a variant of Cas9 may comprise a Cas9 that has one or more changes in amino acid residues as compared to a wild type Cas9 amino acid sequence. As another example, a variant of a deaminase may comprise a deaminase that has one or more changes in amino acid residues as compared to a wild type deaminase amino acid sequence, e.g. following ancestral sequence reconstruction of the deaminase. These changes include chemical modifications, including substitutions of different amino acid residues truncations, covalent additions (e.g. of a tag), and any other mutations. The term also encompasses circular permutants, mutants, truncations, or domains of a reference sequence,and which display the same or substantially the same functional activity or activities as the reference sequence. This term also embraces fragments of a wild type protein.
[0148] The level or degree of which the property is retained may be reduced relative to the wild type protein but is typically the same or similar in kind. Generally, variants are overall very similar, and in many regions, identical to the amino acid sequence of the protein described herein. A skilled artisan will appreciate how to make and use variants that maintain all, or at least some, of a functional ability or property.
[0149] The variant proteins may comprise, or alternatively consist of, an amino acid sequence which is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%, identical to, for example, the amino acid sequence of a wild-type protein, or any protein provided herein.
[0150] By a polypeptide having an amino acid sequence at least, for example, 95% “identical” to a query amino acid sequence, it is intended that the amino acid sequence of the subject polypeptide is identical to the query sequence except that the subject polypeptide sequence may include up to five amino acid alterations per each 100 amino acids of the query amino acid sequence. In other words, to obtain a polypeptide having an amino acid sequence at least 95% identical to a query amino acid sequence, up to 5% of the amino acid residues in the subject sequence may be inserted, deleted, or substituted with another amino acid. These alterations of the reference sequence may occur at the amino- or carboxy-terminal positions of the reference amino acid sequence or anywhere between those terminal positions, interspersed either individually among residues in the reference sequence or in one or more contiguous groups within the reference sequence.
[0151] As a practical matter, whether any particular polypeptide is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, for instance, the amino acid sequence of a protein such as a PrP protein, can be determined conventionally using known computer programs. A preferred method for determining the best overall match between a query sequence (a sequence of the present invention) and a subject sequence, also referred to as a global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)). In a sequence alignment the query and subject sequences are either both nucleotide sequences or both amino acid sequences. The result of said global sequence alignment is expressed as percent identity. Preferred parameters used in a FASTDB amino acid alignment are: Matrix=PAM 0, k-tuple=2, Mismatch Penalty=1, Joining Penalty=20, Randomization Group Length=0,Cutoff Score=1, Window Size=sequence length, Gap Penalty=5, Gap Size Penalty=0.05, Window Size=500 or the length of the subject amino acid sequence, whichever is shorter.
[0152] If the subject sequence is shorter than the query sequence due to N- or C-terminal deletions, not because of internal deletions, a manual correction must be made to the results. This is because the FASTDB program does not account for N- and C-terminal truncations of the subject sequence when calculating global percent identity. For subject sequences truncated at the N- and C-termini, relative to the query sequence, the percent identity is corrected by calculating the number of residues of the query sequence that are N- and C- terminal of the subject sequence, which are not matched / aligned with a corresponding subject residue, as a percent of the total bases of the query sequence. Whether a residue is matched / aligned is determined by results of the FASTDB sequence alignment. This percentage is then subtracted from the percent identity, calculated by the above FASTDB program using the specified parameters, to arrive at a final percent identity score. This final percent identity score is what is used for the purposes of the present invention. Only residues to the N- and C-termini of the subject sequence, which are not matched / aligned with the query sequence, are considered for the purposes of manually adjusting the percent identity score. That is, only query residue positions outside the farthest N- and C-terminal residues of the subject sequence.
[0153] The term “vector,” as used herein, refers to a nucleic acid that can be modified to encode a gene of interest and that is able to enter into a host cell and replicate within the host cell, and then transfer a replicated form of the vector into another host cell. Exemplary suitable vectors include viral vectors, such as AAV vectors or bacteriophages and filamentous phage, and conjugative plasmids. Additional suitable vectors will be apparent to those of skill in the art based on the instant disclosure.
[0154] As used herein the term “wild type” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene or characteristic as it occurs in nature as distinguished from mutant or variant forms. DETAILED DESCRIPTION
[0155] As disclosed herein, the inventors have found that base editing represents a strategy for the treatment of prion diseases. In certain aspects and embodiments, the disclosure provides: (a) compositions comprising base editors that include napDNAbps (e.g., CRISPR nucleases) and a deaminase (e.g., adenosine deaminase or cytidine deaminase), guide RNAs targeting one or more regions of the PRNP gene (e.g., promoter, promoter-proximal, or coding regions), nucleic acid molecules encoding the base editor, its napDNAbp and / or deaminase components, or a base editor–guide RNA complex; vectors, delivery vehicles, engineered cells, pharmaceutical formulations, and kits comprising the same; and (b) methods of using or administering such compositions to install one or more edits within the PRNP gene—e.g., in the promoter, promoter-proximal, or coding sequence—to reduce or abolish PrP expression and / or activity. In certain embodiments, the reduction of PrP activity may occur through (i) inhibition of PRNP transcription, (ii) disruption of translation (e.g., by editing the start codon), or (iii) generation of a truncated PrP variant that lacks misfolding and aggregation properties characteristic of PrPSc. These compositions and methods are useful for preventing, treating, or ameliorating prion disease.
[0156] Non-limiting examples of base editing approaches contemplated herein include:
[0157] Promoter Editing: Introducing one or more edits into the PRNP promoter by contacting the promoter with a base editing composition as disclosed, thereby inhibiting or reducing PRNP transcription and consequently lowering PRNP mRNA and PrP protein levels.
[0158] Promoter-Proximal Editing: Introducing one or more edits into the promoter- proximal region of PRNP with a base editing composition, similarly reducing PRNP transcription, PRNP mRNA, and PrP expression.
[0159] Start Codon Disruption: Editing the PRNP start codon using a base editing composition to abolish translation initiation, thereby reducing or eliminating PrP production and downstream PrPSc accumulation.
[0160] Nonsense Mutation Installation: Introducing one or more nonsense mutations in PRNP coding regions—e.g., at codons corresponding to amino acid residues R37, W57, W65, W81, or Q83—converting them into premature stop codons and leading to the expression of a truncated PrP protein that lacks pathogenic PrPSc-forming activity.
[0161] Each of these approaches may be employed individually or in combination to prevent, treat, or reduce the severity of prion disease by reducing and / or blocking PRNP mRNA synthesis, generating non-pathogenic truncated PrP species, or fully eliminating PrP expression.
[0162] Accordingly, in various aspects, the present disclosure provides base editing compositions, systems, and methods for installing a pre-mature stop codon in a PRNP gene encoding a prion protein (PrP), e.g., to produce a non-pathogenic truncated PrP protein, to mutate the start codon of the PrP protein in the PRNP gene to prevent translation or to mutate the 273 bp promoter and promoter-proximal region to prevent transcription factor bindingand lower mRNA transcription thereby lowering PrP expression. The present disclosure also provides guide RNAs configured to direct the base editors to install a premature stop codon in the PRNP gene corresponding to an amino acid position in the PrP protein, e.g., introducing stop codons a PRNP codons corresponding to amino acid positions R37, W57, W65, W81 and Q83 in PrP. Also provided herein are methods for treating or preventing prion disease, e.g., via truncating a PrP protein or decreasing mRNA and protein expression by promoter disruption. The present disclosure also provides complexes comprising the base editor (e.g., the fusion proteins that make up the base editors) and guide RNAs, as well as vectors, cells, pharmaceutical compositions, and kits.
[0163] Aspects of the present disclosure relate to compositions comprising split nucleobase editors for installing a stop codon in a PRNP gene encoding a prion protein (PrP), e.g., to produce a non-pathogenic truncated PrP protein, to mutate the start codon of the PrP protein in the PRNP gene to prevent translation or to mutate the 273 bp promoter and promoter-proximal region to prevent transcription factor binding and lower mRNA transcription thereby lowering PrP expression. This is necessary, for example, to overcome the packaging size limit when AAVs are used as the delivery vehicle. Split nucleobase editors (e.g., CBE or ABE) are divided into an N-terminal portion (or “half”) and a C- terminal half. Each base editor half is fused to half of a fast-splicing split-intein. Following co-infection by AAV particles expressing each base editor–split intein half, protein splicing in trans reconstitutes the full-length base editor. Unlike other approaches utilizing small molecules or sgRNA to bridge split Cas9, intein splicing removes all exogenous sequences and regenerates a native peptide bond at the split site, resulting in a single reconstituted protein (e.g., a protein that is identical in sequence to the unmodified nucleobase editor).
[0164] Split-intein CBEs and split-intein ABEs are disclosed that are packaged into dual AAV genomes to enable efficient base editing in somatic tissues of therapeutic relevance, including liver, heart, muscle, retina, and brain. The resulting AAVs were used to achieve base editing efficiencies at test loci for both CBEs and ABEs that, in each of these tissues, meets or exceeds therapeutically relevant editing thresholds for the treatment of human genetic diseases at AAV dosages that are known to be well-tolerated in humans. In particular, the disclosed AAV-nucleobase editor vectors achieved significant editing efficiencies to enable about a 60% reduction in PrP protein levels in the mouse brain. Accordingly, the invention provides split nucleobase editors and nucleic acids and vectors encoding same, as well as cells, compositions, methods, and kits, that utilize the disclosed split nucleobase editors, and vectors.
[0165] Aspects of the present disclosure relate to compositions encoding a split nucleobase editor and guide RNA. The compositions, according to some embodiments, comprise one or more nucleotide sequences. A description of the various compositions contemplated herein and the components encoded by the one or more nucleotide sequences are discussed below.
[0166] In some embodiments, the compositions comprise a first nucleotide sequence encoding a guide RNA and a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N. In some embodiments, the guide RNA is configured to direct a recombined nucleobase editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene or to mutate the start codon of the PrP protein in the PRNP gene to prevent translation. In some embodiments, the composition further comprises a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the split nucleobase editor.
[0167] In some embodiments, the compositions comprise a first nucleotide sequence encoding a guide RNA and a third nucleotide sequence encoding an intein-C fused to the N- terminus of a C-terminal portion of the split nucleobase editor. In some embodiments, the guide RNA is configured to direct a recombined nucleobase editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene or to mutate the start codon of the PrP protein in the PRNP gene to prevent translation. In some embodiments, the composition further comprises and a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N.
[0168] In some embodiments, the compositions comprise a first nucleotide sequence encoding a guide RNA, a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N, and a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the split nucleobase editor. In some embodiments, the guide RNA is configured to direct a recombined nucleobase editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene or to mutate the start codon of the PrP protein in the PRNP gene to prevent translation. In some cases, the N-terminal portion of the split nucleobase editor and the C-terminal portion of the split nucleobase editor may be joined to form a recombined nucleobase editor (e.g., following expression in a cell).
[0169] Aspects of the present disclosure relate to compositions encoding an intact nucleobase editor (e.g., not a split nucleobase editor) and guide RNA. In some embodiments, the compositions comprise one or more nucleotide sequences encoding one or more portionsof an intact nucleobase editor (e.g., not a split nucleobase editor). In some embodiments, the compositions comprise one or more polynucleotides comprising one or more of the following: (1) a first nucleotide sequence encoding a guide RNA; (2) a second nucleotide sequence encoding a deaminase domain (e.g., TadCBEd or ABE8e or ABE8e(V106W)) of the nucleobase editor (3) a third nucleotide sequence encoding a nucleic acid programmable DNA binding protein (napDNAbp) domain (e.g., SauriCas9, enCjCas9, SpCas9, SaCas9) of the nucleobase editor; (4) a fourth nucleotide sequence encoding a terminator sequence (e.g., a polyT sequence); (5) a fifth nucleotide sequence encoding a UGI domain; (6) a sixth nucleotide sequence encoding an inverted terminal repeat (ITR) domain; or (7) a seventh nucleotide sequence encoding one or more miR target sites. In some embodiments, the one or more nucleotide sequences is operably linked to a promoter sequence, e.g., Cbh, EFS, hSYN, and / or U6. In some cases, the one or more nucleotide sequences is operably linked to a hSYN promoter.
[0170] In some embodiments, compositions disclosed herein comprise a polynucleotide encoding one or more domains of an intact nucleobase editor. In some embodiments, the polynucleotide encodes for a deaminase domain, a napDNAbp domain, and a UGI domain of the nucleobase editor. In some embodiments, the polynucleotide further encodes one or more additional nucleotide sequences, for example, a terminator sequence, an ITR sequence, or a promoter sequence.
[0171] Other aspects of the disclosure relate to compositions comprising one or more recombinant adeno associated virus (rAAV) particles.
[0172] In some embodiments, the compositions disclosed herein comprises a first rAAV particle comprising a first nucleotide sequence encoding a guide RNA (gRNA) and a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N. In some embodiments, the guide RNA is configured to direct a recombined base editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene or to mutate the start codon of the PrP protein in the PRNP gene to prevent translation. Additionally, in some embodiments, the compositions further comprise a second recombinant adeno associated virus (rAAV) particle comprising a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C- terminal portion of the split nucleobase editor.
[0173] In other embodiments, the compositions disclosed herein comprise a second rAAV particle comprising a first nucleotide sequence encoding a guide RNA (gRNA) and a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminalportion of the split nucleobase editor. In some embodiments, the guide RNA is configured to direct a recombined base editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene or to mutate the start codon of the PrP protein in the PRNP gene to prevent translation. Additionally, in some embodiments, the compositions further comprise a first recombinant adeno associated virus (rAAV) particle comprising a second nucleotide sequence encoding a N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N.
[0174] In some embodiments, still, the compositions disclosed herein comprise a first recombinant adeno associated virus (rAAV) particle comprising a second nucleotide sequence encoding a N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N and a second recombinant adeno associated virus (rAAV) particle comprising a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the split nucleobase editor. The first rAAV particle and / or the second rAAV particle further comprises a first nucleotide sequence encoding a guide RNA configured to direct a recombined base editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene or to mutate the start codon of the PrP protein in the PRNP gene to prevent translation. As described above, and elsewhere herein, in some embodiments, the N-terminal portion of the split nucleobase editor and the C- terminal portion of the split nucleobase editor may be joined to form a recombined nucleobase editor.
[0175] In some embodiments, compositions disclosed herein comprises a first rAAV particle comprising a first nucleotide sequence encoding a guide RNA (gRNA) and a second nucleotide sequence encoding a complete nucleobase editor (e.g., fusion protein comprising napDNAbp domain and a deaminase domain). As will be discussed more in detail below, compositions comprising a single rAAV particle comprise size-minimized Cas9 domains (e.g., enCjCas9, evoCjCas9, SauriCas9, and eNme2-Cas9) fused to a deaminase (e.g., a cytidine deaminase such as TadCBEd or an adenine deaminase such as ABE8e).
[0176] In other embodiments, however, compositions disclosed herein comprise one or more rAAV particles comprising one or more polynucleotides comprising one or more of the following: (1) a first nucleotide sequence encoding a gRNA configured to direct the recombined nucleobase editor to a target site in the PRNP gene corresponding to any one of amino acid positions R37, W57, W65, W81, or Q83 in the PrP protein or to mutate an ATG start codon to GTG or ACG, thus preventing translation of the PrP protein from occurring, (2) a second nucleotide sequence encoding a deaminase domain of the nucleobase editor; (3)a third nucleotide sequence encoding a napDNAbp domain of the nucleobase editor; (4) a fourth nucleotide sequence encoding a UGI domain; (5) a fifth nucleotide sequence encoding a terminator sequence; (6) a sixth nucleotide sequence encoding an ITR domain; and (7) a seventh nucleotide sequence encoding an miR target site.
[0177] In some embodiments, any one of the compositions disclosed herein comprise one or more additional nucleotide sequences (e.g., a fourth, fifth, sixth, seventh, etc.).
[0178] In some cases, the one or more nucleotide sequences encode a fourth nucleotide sequence encoding a terminator sequence (e.g., poly T sequence).
[0179] In some embodiments, the one or more nucleotide sequences encode a uracil glycosylase inhibitor (UGI) domain. Amino acid sequences of exemplary UGI domains are provided below:
[0180] Bacillus phage PBS2 (Bacteriophage PBS2) Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLT SDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 17)
[0181] Erwinia tasmaniensis SSB (thermostable single-stranded DNA binding protein)
[0182] MASRGVNKVILVGNLGQDPEVRYMPNGGAVANITLATSESWRDKQTGET KEKTEWHRVVLFGKLAEVAGEYLRKGSQVYIEGALQTRKWTDQAGVEKYTTEVVV NVGGTMQMLGGRSQGGGASAGGQNGGSNNGWGQPQQPQGGNQFSGGAQQQARP QQQPQQNNAPANNEPPIDFDDDIP (SEQ ID NO: 18)
[0183] UdgX (binds to uracil in DNA but does not excise)
[0184] MAGAQDFVPHTADLAELAAAAGECRGCGLYRDATQAVFGAGGRSARIM MIGEQPGDKEDLAGLPFVGPAGRLLDRALEAADIDRDALYVTNAVKHFKFTRAAGG KRRIHKTPSRTEVVACRPWLIAEMTSVEPDVVVLLGATAAKALLGNDFRVTQHRGE VLHVDDVPGDPALVATVHPSSLLRGPKEERESAFAGLVDDLRVAADVRP (SEQ ID NO: 19)
[0185] UDG (catalytically inactive human UDG, binds to uracil in DNA but does not excise)
[0186] MIGQKTLYSFFSPSPARKRHAPSPEPAVQGTGVAGVPEESGDAAAIPAKK APAGQEEPGTPPSSPLSAEQLDRIQRNKAAALLRLAARNVPVGFGESWKKHLSGEFG KPYFIKLMGFVAEERKHYTVYPPPHQVFTWTQMCDIKDVKVVILGQEPYHGPNQAH GLCFSVQRPVPPPPSLENIYKELSTDIEDFVHPGHGDLSGWAKQGVLLLNAVLTVRA HQANSHKERGWEQFTDAVVSWLNQNSNGLVFLLWGSYAQKKGSAIDRKRHHVLQ TAHPSPLSVYRGFFGCRHFSKTNELLQKSGKKPIDWKEL (SEQ ID NO: 20)
[0187] In some embodiments, the UGI domain comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5% , or is 100% identical to the amino acid sequence of SEQ ID NOs: 17-20.
[0188] In some embodiments, the one or more nucleotide sequences encode an inverted terminal repeat (ITR) domain of an adenovirus (AAV). The AAV may be of any suitable serotype known to the skilled artisan, including for example, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV-DJ, AAV- DJ / 8, MyoAAV, AAV-MG, AAV-PHP.eB, AAVBI30, AAV-PHP.S, AAV2.7m8, AAV- QuadYF, and AAV2-retro. In some embodiments, the AAV serotype is AAV2 or AAV9.
[0189] In some embodiments, the one or more nucleotide sequences encode one or more microRNA (miR) targeting sites. The one or more miR targeting sites may be encoded in a 3′ untranslated region (UTR) of a split nucleobase editor, according to some embodiments. Additionally, or alternatively, the one or more miR targeting sites may be encoded in a 5′ UTR region of the split nucleobase editor. In some embodiments, the one or more miR target sites comprise miR124, miR204, miR181, miR1, miR206, miR126, miR134, miR133, miR208, miR302, miR338, miR219, miR9, miR218, miR7, miR128, miR125, miR138, miR132, miR212, miR137, miR31, miR127, miR143, miR346, miR708, miR10, miR192, miR194, miR215, miR216, miR122a, miR92a, miR483, miR130, miR292, miR217, miR193a, miR128a, miR150, miR181a, miR155, and / or miR142 target sites.
[0190] In some embodiments, the one or more miR target sites comprises an miR183 target site, an miR 122 target, or both. Other miR target sites are also contemplated herein. For example, in some embodiments, the miR target sites comprise miR124, miR204, and / or miR181 target sites. In some embodiments, an eye tissue comprises miR124, miR204, and / or miR181.
[0191] In some embodiments, the one or more miR target sites comprise miR1, miR206, miR126, miR134, miR133, miR208, miR302, miR338, miR219, miR124, miR9, miR218, miR7, and / or miR128 target sites. In some embodiments, a heart tissue comprises miR1, miR206, miR126, miR134, miR133, miR208, miR302, miR338, miR219, miR124, miR9, miR218, miR7, and / or miR128.
[0192] In some embodiments, the one or more miR target sites comprise miR125, miR138, miR132, miR212, miR137, miR31, miR127, miR143, miR346, and / or miR708 target sites. In some embodiments, a brain / nervous tissue comprises miR125, miR138, miR132, miR212, miR137, miR31, miR127, miR143, miR346, and / or miR708.
[0193] In some embodiments, the one or more miR target sites comprise miR10, miR192, miR204, miR194, miR215, and / or miR216 target sites. In some embodiments, a kidney tissue comprises target sites comprise miR10, miR192, miR204, miR194, miR215, and / or miR216.
[0194] In some embodiments, the one or more miR target sites comprise miR122a, miR192, miR92a, and / or miR483 target sites. In some embodiments, a liver tissue comprises miR122a, miR192, miR92a, and / or miR483.
[0195] In some embodiments, the one or more miR target site comprises miR126, and wherein the miR126 target sites. In some embodiments, a lung tissue comprises miR126.
[0196] In some embodiments, the one or more miR target sites comprise miR126, miR130, miR302, and / or miR292 target sites. In some embodiments, a hematopoietic and pluripotent stem cell comprise miR126, miR130, miR302, and / or miR292.
[0197] In some embodiments, one or more miR target sites comprise miR216 and / or miR217 target sites. In some embodiments, a pancreas tissue comprises miR216 and / or miR217.
[0198] In some embodiments, the one or more miR target sites comprise miR133, miR1, miR206, miR134, miR193a, and / or miR128a target sites. In some embodiments, a muscle tissue comprises miR133, miR1, miR206, miR134, miR193a, and / or miR128a.
[0199] In some embodiments, the one or more miR target sites comprise miR150, miR181a, miR155, and / or miR142 target sites. In some embodiments one or more tissues of an immune system comprise miR150, miR181a, miR155, and / or miR142.
[0200] In some embodiments, the seventh nucleotide sequence encodes one or more miR targets in a 3′ UTR of the nucleobase editor. However, the seventh nucleotide sequence may encode more than one miR target in the 3′ UTR of the nucleobase editor, according to some embodiments. For example, in some embodiments, the seventh nucleotide sequence encodes two or more miR target sites, three or more miR target sites, four or more miR target sites, five or more miR target sites, or six or more miR target sites in the 3’ untranslated region (UTR) of the nucleobase editor.
[0201] In some embodiments, the one or more nucleotide sequences is operably linked to a promoter. For example, in some embodiments, the first nucleotide sequence encoding the gRNA is operably linked to a promoter. Additionally, in some embodiments, the second nucleotide sequence encoding the N-terminal portion of the split nucleobase editor is operably linked to a promoter. Similar, in some embodiments, a third nucleotide sequenceencoding the C-terminal portion of the split nucleobase editor is operably linked to a promoter.
[0202] Any one of the nucleotide sequences disclosed herein may be operably linked to any suitable promoter known to the skilled artisan. In some embodiments, the promoter is selected from the group consisting of Cbh, hSYN, and U6. In some embodiments, one or more of the nucleotides sequences is operably linked to a tissue specific promoter (e.g., a neuronal tissue promoter). In some embodiments, the promoter is a neuronal tissue promoter (e.g., hSYN). In some embodiments, the promoter is hSYN.
[0203] Other aspects of the disclosure relate to methods of making the disclosed split nucleobase editors, as well as methods of using the split nucleobase editors or nucleic acid molecules encoding the nucleobase editors in applications including editing a nucleic acid molecule, e.g., a genome (e.g., PRNP gene). Such methods involve transducing (e.g., via transfection) cells with a plurality of complexes each comprising a portion of a split nucleobase editor (e.g., a nucleobase editor comprising a napDNAbp (e.g., nCas9) domain and a deaminase domain) and / or a gRNA molecule. In some embodiments, the nucleic acid constructs encoding the N-terminal and C-terminal portions of the split nucleobase editor are transfected separately from one another. In certain embodiments, the methods involve the transfection of nucleic acid constructs (e.g., plasmids) that each (or together) encode the components of a complex of split nucleobase editor and a gRNA molecule.
[0204] In certain embodiments of the disclosed methods of making the disclosed split nucleobase editors, one or more nucleic acid constructs that encode the split nucleobase editor is transfected into the cell separately from the plasmid that encodes the gRNA molecule. In certain embodiments, these components are encoded on a single construct and transfected together. In other embodiments, the methods disclosed herein involve the introduction into cells of one or more nucleic acid vectors encoding a split nucleobase editor and gRNA molecule that has been expressed and cloned outside of these cells. In some embodiments, these vectors are delivered as part of an rAAV vector.
[0205] It should be appreciated that any nucleobase editor, e.g., any of the nucleobase editors provided herein, may be introduced into the cell in any suitable way, either stably or transiently. In some embodiments, a nucleobase editor may be transfected into the cell. In some embodiments, the cell may be transduced or transfected with a nucleic acid construct that encodes a nucleobase editor. For example, a cell may be transduced (e.g., with a virus encoding a nucleobase editor), or transfected (e.g., with a plasmid encoding a nucleobase editor) with a nucleic acid that encodes a nucleobase editor, or the translated nucleobaseeditor. Such transduction may be a stable or transient transduction. In some embodiments, cells expressing a nucleobase editor or containing a nucleobase editor may be transduced or transfected with one or more gRNA molecules, for example, when the nucleobase editor comprises a Cas9 (e.g., nCas9) domain. In some embodiments, a plasmid expressing one or more portions of a nucleobase editor may be introduced into cells through electroporation, transient (e.g., lipofection) and stable genome integration (e.g., nucleofection and piggybac), viral transduction, or other methods known to those of skill in the art. In particular, embodiments, plasmids expressing one or more portions of any of the disclosed nucleobase editors may be delivered to cells through nucleofection.
[0206] In some aspects, the disclosed split nucleobase editors are delivered to the cell (or the subject) by use of recombinant AAV (rAAV) particles. In some embodiments, any of the disclosed split nucleobase editors is fused to split intein pairs that are packaged into two separate rAAV particles that, when co-delivered to a cell, reconstitute the functional editor protein. Accordingly, the disclosure provides dual rAAV vectors and dual rAAV vector particles that comprise expression constructs that encode two portions (or “two halves”) of any of the disclosed nucleobase editors, wherein the encoded nucleobase editor is divided between the two halves at a split site. In some embodiments, the disclosed rAAV vectors encoding the split nucleobase editors may comprise a nucleotide sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the sequences of SEQ ID NOs: 167-179, and 182-183.
[0207] In some embodiments, the disclosed rAAV vectors encode a single rAAV vector and rAAV vector particles that comprise an expression construct that encodes a single nucleobase editor. In some embodiments, the disclosed rAAV vectors encoding a single nucleobase editors may comprise a nucleotide sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the sequences of SEQ ID NOs: 180-181, and 184-185. Other aspects of the disclosure related to complexes comprising the split nucleobase editors (e.g., the fusion proteins that make up the base editors) and guide RNAs, as well as vectors, cells, pharmaceutical compositions, and kits; each of which is discussed in more detail below.
[0208] Other aspects of the disclosure relate to preventing the PRNP gene from being translated into protein (or reducing the amount that is translated into protein), such as for example, by disrupting the M1V start codon as described in An, M., et al. Nat Med. 2025 Jan14;31(4):1319–1328, which is hereby incorporated by reference in its entirety (See Example 1).
[0209] In other aspects still, the disclosure relates to preventing or reducing PRNP gene expression by base editing the PRNP promoter region. In some embodiments, any one of the base editing system disclosed herein may be used to edit the PRNP promoter region. In some embodiments, a gRNA is used to guide the base editor to the PRNP target gene. In some embodiments, the gRNA comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% identical to any one of the nucleic acid sequences of SEQ ID NOs: 21-38. In some embodiments, the gRNA comprises a nucleic acid sequence that is identical to any one of the nucleic acid sequences of SEQ ID NOs: 21-38. Table 8. Exemplary gRNAs for targeting a base editor to the promoter region of a PRNP gene
[0210] In additional aspects, the disclosure relates to preventing or reducing PRNP gene expression by base editing the promoter-proximal region site to prevent transcription factor binding and lower mRNA transcription therby lowering PrP expression. In some embodiments, any one of the base editing system disclosed herein may be used to edit the PRNP promoter-proximal region site. In some embodiments, a gRNA is used to guide the base editor to the PRNP target site (e.g., the promoter-proximal region site). In some embodiments, the gRNA comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% identical to any one of the nucleic acid sequences of SEQ ID NOs: 771-773. In some embodiments, the gRNA comprises a nucleic acid sequence that is identical to any one of the nucleic acid sequences of SEQ ID NOs: 771-773. Table 9. Exemplary gRNAs for targeting a base editor to the promoter-proximal region site of a PRNP genenapDNAbp domain
[0211] In some embodiments, any one of the nucleobase editors (or a recombined nucleobase editor) disclosed herein comprise a napDNAbp domain and a deaminase domain. In some embodiments, the napDNAbp domain is fused to the deaminase domain, although this is not a requirement of the invention. Any suitable napDNAbp domain known in the art may be used in the base editors described herein, such as those described in detail in United State Patent Application No: 17 / 797701 by David Liu, et al., filed on August 4, 2022, which is incorporated herein by reference in its entirety. For example, in various embodiments, the napDNAbp may be any Class 2 CRISPR-Cas system, including any type II, type V, or type VI CRISPR-Cas enzyme. Given the rapid development of CRISPR-Cas as a tool for genomeediting, there have been constant developments in the nomenclature used to describe and / or identify CRISPR-Cas enzymes, such as Cas9 and Cas9 orthologs. This application references CRISPR-Cas enzymes with nomenclature that may be old and / or new as described in United State Patent Application 63 / 136,194 (described elsewhere herein) or Makarova et al., The CRISPR Journal, Vol. 1, No. 5, 2018, which is incorporated herein by reference in its entirety.
[0212] Other napDNAbps are also possible in other embodiments. For example, in some embodiments, the napDNAbp comprises the canonical SpCas9, or any ortholog Cas9 protein, or any variant Cas9 protein —including any naturally occurring variant, mutant, or otherwise engineered version of Cas9 — that is known or that may be made or evolved through a directed evolutionary or otherwise mutagenic process. In various embodiments, the Cas9 or Cas9 variants have a nickase activity, i.e., only cleave one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variants have inactive nucleases, i.e., are “dead” Cas9 proteins. Other variant Cas9 proteins that may be used are those having a smaller molecular weight than the canonical SpCas9 (e.g., for easier delivery) or having modified or rearranged primary amino acid structure (e.g., the circular permutant formats).
[0213] In various embodiments described herein, the base editors comprise a napDNAbp, such as a Cas9 protein. These proteins are “programmable” by way of their becoming complexed with a guide RNA (or a pegRNA, as the case may be), which guides the Cas9 protein to a target site on the DNA which possess a sequence that is complementary to the spacer portion of the gRNA (or pegRNA) and also which possesses the required PAM sequence. However, in certain embodiment envisioned here, the napDNAbp may be substituted with a different type of programmable protein, such as a zinc finger nuclease or a transcription activator-like effector nuclease (TALEN). See U.S. Ser. No. 12 / 965,590; U.S. Ser. No. 13 / 426,991 (U.S. Pat. No. 8,450,471); U.S. Ser. No. 13 / 427,040 (U.S. Pat. No. 8,440,431); U.S. Ser. No. 13 / 427,137 (U.S. Pat. No. 8,440,432); and U.S. Ser. No. 13 / 738,381, all of which are incorporated by reference herein in their entirety. In addition, TALENS are described in WO 2015 / 027134, US 9,181,535, Boch et al., "Breaking the Code of DNA Binding Specificity of TAL-Type III Effectors", Science, vol. 326, pp. 1509-1512 (2009), Bogdanove et al., TAL Effectors: Customizable Proteins for DNA Targeting, Science, vol. 333, pp. 1843-1846 (2011), Cade et al., "Highly efficient generation of heritable zebrafish gene mutations using homo- and heterodimeric TALENs", Nucleic Acids Research, vol. 40, pp. 8001-8010 (2012), and Cermak et al., "Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting", Nucleic AcidsResearch, vol. 39, No. 17, e82 (2011), each of which are incorporated herein by reference. See also, for example, in Carroll et al., “Genome Engineering with Zinc-Finger Nucleases,” Genetics, Aug 2011, Vol. 188: 773-782; Durai et al., “Zinc finger nucleases: custom- designed molecular scissors for genome engineering of plant and mammalian cells,” Nucleic Acids Res, 2005, Vol. 33: 5978-90; and Gaj et al., “ZFN, TALEN, and CRISPR / Cas-based methods for genome engineering,” Trends Biotechnol. 2013, Vol.31: 397-405, each of which are incorporated herein by reference in their entireties.
[0214] In some embodiments, the napDNAbp domain comprises SpCas9, SpCas9 (D10A) nickase, enCjCas9, evoCjCas9, SauriCas9, eNme2-C Cas9, or SaCas9. Nucleotide sequences and / or amino acid sequences for exemplary napDNAbp domains are provided below: SpCas9 (D10A) nickase
[0215] GACAAGAAGTACAGCATCGGCCTGGCCATCGGCACCAACTCTGTGGGC TGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGGTGCTG GGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTC GACAGCGGCGAAACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAG ATACACCAGACGGAAGAACCGGATCTGCTATCTGCAAGAGATCTTCAGCAACGA GATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGT GGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGA GGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGT GGACAGCACCGACAAGGCCGACCTGCGGCTGATCTATCTGGCCCTGGCCCACAT GATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACCTGAACCCCGACAACAG CGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAG ACTGAGCAAGAGCAGACGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAA GAAGAATGGCCTGTTCGGAAACCTGATTGCCCTGAGCCTGGGCCTGACCCCCAA CTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGA CACCTACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGC CGACCTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATC CTGAGAGTGAACACCGAGATCACCAAGGCCCCCCTGAGCGCCTCTATGATCAAG AGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGCTCTCGTGCGGCAG CAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTAC GCCGGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGGACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGA GAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAACGGCAGCATCCCCCACCAG ATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTACCCAT TCCTGAAGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCT ACTACGTGGGCCCTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAA AGAGCGAGGAAACCATCACCCCCTGGAACTTCGAGGAAGTGGTGGACAAGGGC GCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAGAACCTGCCCA ACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATA ACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCC TGAGCGGCGAGCAGAAAAAGGCCATCGTGGACCTGCTGTTCAAGACCAACCGGA AAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAGAAAATCGAGTGCTTCG ACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACATA CCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAA CGAGGACATTCTGGAAGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGA GATGATCGAGGAACGGCTGAAAACCTATGCCCACCTGTTCGACGACAAAGTGAT GAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCCGGAAGC TGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGA AGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCC TGACCTTTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAGGGCGATAGCC TGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTAAGAAGGGCATCC TGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGC CCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGA CAGAAGAACAGCCGCGAGAGAATGAAGCGGATCGAAGAGGGCATCAAAGAGCT GGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACACCCAGCTGCAGAACG AGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGG AACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGA GCTTTCTGAAGGACGACTCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGA ACCGGGGCAAGAGCGACAACGTGCCCTCCGAAGAGGTCGTGAAGAAGATGAAG AACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTTCGAC AATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTC ATCAAGAGACAGCTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATC CTGGACTCCCGGATGAACACTAAGTACGACGAGAATGACAAGCTGATCCGGGAA GTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCGGAAGGATTTCC AGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAACGCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCG AGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAGATGATCGCCAAGA GCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTTCTTCTACAGCAACATCA TGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGGC CTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGG GATTTTGCCACCGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAA AAGACCGAGGTGCAGACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCCAAGAGG AACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAGTACGG CGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAA AAGGGCAAGTCCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATC ATGGAAAGAAGCAGCTTCGAGAAGAATCCCATCGACTTTCTGGAAGCCAAGGGC TACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAGTACTCCCTGTTC GAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAA GGGAAACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGC CACTATGAGAAGCTGAAGGGCTCCCCCGAGGATAATGAGCAGAAACAGCTGTTT GTGGAACAGCACAAGCACTACCTGGACGAGATCATCGAGCAGATCAGCGAGTTC TCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCTAC AACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTG TTTACCCTGACCAATCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCA TCGACCGGAAGAGGTACACCAGCACCAAAGAGGTGCTGGACGCCACCCTGATCC ACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTCAGCTGGGAG GTGAC (SEQ ID NO: 39) SauriCas9
[0216] CAGGAGAACCAGCAGAAGCAAAATTACATCCTGGGCCTGGCCATCGG CATCACCAGCGTGGGCTATGGCCTGATCGACAGCAAGACCAGAGAAGTGATTGA CGCCGGCGTGCGGCTATTCCCAGAGGCCGACTCTGAAAACAACAGCAATAGAAG ATCTAAGCGGGGCGCCCGGAGACTGAAAAGACGAAGAATCCACAGACTGAACA GAGTGAAAGACCTCCTGGCTGACTACCAGATGATCGACTTAAACAACGTGCCCA AGTCTACCGACCCCTACACCATCCGGGTGAAGGGACTGCGGGAACCTCTGACCA AGGAAGAGTTTGCCATCGCTCTGCTGCATATCGCCAAGAGAAGAGGCCTGCACA ACATCTCCGTGAGCATGGGCGATGAGGAACAGGACAACGAGCTGTCCACCAAGC AGCAGCTGCAGAAGAACGCTCAGCAGCTGCAGGACAAATACGTGTGCGAGCTGC AACTGGAAAGACTGACCAACATCAACAAGGTCAGAGGCGAGAAGAACCGGTTCAAGACCGAAGATTTCGTGAAGGAAGTGAAGCAGCTGTGCGAGACCCAGCGGCA GTACCACAACATCGACGATCAGTTCATCCAGCAGTACATCGACCTGGTGAGCAC CCGGAGAGAATACTTCGAGGGCCCTGGCAACGGCTCTCCATATGGCTGGGATGG AGATCTGCTGAAGTGGTATGAGAAGCTGATGGGCAGATGCACCTACTTCCCTGA GGAGCTGAGAAGCGTGAAGTACGCCTACAGCGCCGATCTGTTTAACGCCCTGAA CGATCTGAACAACCTGGTTGTGACCCGCGACGACAACCCTAAGCTGGAATACTA CGAGAAATACCACATTATCGAGAACGTGTTCAAGCAGAAGAAAAATCCAACTCT GAAGCAAATCGCCAAAGAGATCGGCGTGCAGGATTACGACATCAGAGGATACA GAATTACCAAGTCCGGTAAGCCTCAGTTCACCAGCTTCAAACTCTACCACGACCT GAAAAACATCTTTGAACAGGCCAAATACCTGGAAGATGTGGAGATGCTGGATGA GATAGCTAAAATCCTGACAATCTACCAAGACGAGATCAGCATCAAGAAAGCCCT GGACCAGCTGCCTGAGCTGCTGACCGAGAGCGAAAAAAGCCAGATCGCTCAGCT GACCGGCTACACCGGTACACATAGACTGTCTCTGAAGTGCATCCACATCGTGATC GACGAGCTGTGGGAGAGCCCCGAAAACCAGATGGAAATCTTCACCAGACTGAAC CTCAAGCCAAAGAAGGTGGAAATGAGCGAGATCGACAGCATCCCTACCACACTG GTGGATGAGTTCATCCTGAGCCCTGTGGTGAAGCGGGCCTTCATCCAGTCCATCA AGGTGATCAACGCTGTGATCAACAGATTCGGCCTGCCCGAGGACATCATTATCG AGCTGGCCAGAGAGAAGAACAGCAAGGACAGAAGGAAGTTCATCAACAAGCTG CAGAAACAGAATGAGGCCACCCGGAAAAAAATCGAGCAGCTGCTGGCCAAGTA CGGCAATACCAATGCCAAGTACATGATCGAAAAGATCAAGCTGCATGACATGCA GGAGGGCAAGTGTCTGTACAGCCTGGAAGCTATCCCCCTGGAAGACCTGCTGTCT AATCCTACACACTACGAGGTGGACCACATCATCCCTAGAAGCGTGTCCTTCGACA ACAGCCTGAACAACAAGGTTCTGGTGAAGCAAAGCGAGAACAGCAAGAAGGGC AATAGGACCCCTTACCAGTACCTGAGCAGCAACGAGTCCAAGATCTCTTACAAC CAGTTCAAGCAGCACATTCTGAACCTGTCTAAGGCCAAAGATAGAATCAGCAAG AAGAAACGAGATATGCTGCTGGAAGAACGGGACATCAACAAATTCGAGGTGCA GAAGGAATTCATCAACAGAAACCTTGTGGACACCCGGTACGCCACTCGGGAGCT GAGCAACCTGCTGAAGACCTACTTCAGCACACACGACTACGCCGTGAAAGTGAA GACCATCAACGGCGGCTTCACAAACCACCTGAGGAAGGTGTGGGACTTCAAGAA GCACCGGAACCACGGCTACAAGCACCACGCCGAGGATGCCCTGGTCATCGCCAA CGCCGACTTTCTGTTCAAAACCCACAAGGCCCTGAGACGGACAGACAAGATCCT GGAACAGCCTGGACTGGAAGTCAACGACACCACCGTGAAGGTGGACACAGAGG AGAAGTACCAGGAGTTATTCGAGACACCGAAACAAGTGAAGAACATCAAGCAGT TTAGAGATTTCAAGTATTCTCACAGAGTTGACAAGAAGCCCAACCGGCAGCTGATCAATGATACCCTGTACAGTACCAGAGAGATCGATGGCGAAACCTACGTGGTCC AAACACTGAAAGACCTGTACGCCAAGGACAATGAAAAGGTGAAAAAGCTCTTTA CAGAACGGCCTCAAAAGATACTGATGTACCAGCACGATCCTAAGACCTTTGAGA AACTGATGACCATTCTGAATCAGTACGCTGAGGCAAAGAATCCTCTGGCCGCTTA TTACGAGGATAAGGGCGAATACGTGACCAAGTACGCCAAGAAGGGCAACGGCC CTGCCATCCACAAGATCAAATACATCGACAAGAAACTGGGCAGCTACCTGGACG TGAGTAACAAATATCCTGAGACACAGAACAAGCTGGTGAAACTGTCTCTGAAGA GCTTTAGATTCGACATCTACAAATGTGAACAGGGCTACAAGATGGTGTCCATTGG CTACCTCGACGTACTGAAGAAGGACAACTACTACTACATCCCCAAAGATAAGTA CGAGGCCGAGAAGCAGAAAAAGAAGATCAAGGAAAGCGACCTCTTCGTGGGCA GCTTCTACTACAACGACCTGATCATGTACGAGGACGAACTCTTCCGGGTGATCGG AGTGAACTCCGATATCAACAACCTGGTTGAGCTGAATATGGTCGACATCACCTAC AAGGATTTCTGCGAGGTGAACAACGTGACAGGCGAGAAGAGAATCAAGAAAAC CATCGGAAAGAGAGTGGTGCTGATCGAGAAGTATACGACCGACATCCTGGGAAA TCTGTATAAAACGCCCCTGCCTAAGAAGCCCCAGCTCATTTTCAAGAGAGGCGA GCTG (SEQ ID NO: 40) enCjCas9
[0217] GCCCGCATCCTCGCTTTCGCAATCGGAATCTCTAGTATCGGATGGGCCT TCTCTGAAAACGACGAACTGAAAGACTGCGGCGTGAGAATCTTCACAAAGGTTG AAAACCCTAAAACAGGCGAGTCTTTAGCTCTGCCACGTAGGTTGGCCCGCTCCGC CCGAAAAAGGTATGCTCGGCGGAAGGCTCGCCTCAACCACTTGAAGCATTTGAT AGCTAATGAGTTCAAACTGAACTACGAAGATTACCAGTCCTTCGACGAGTCATTG GCAAAAGCCTACAAAGGCAGCCTTATCAGTCCTTATGAGTTGAGATTTCGCGCAC TCAACGAACTGCTTTCTAAGCAAGACTTTGCTAGGGTCATTCTGCACATCGCAAA ACGGCGAGGTTATGACGATATCAAGAACTCCGACGATAAAGAAAAGGGAGCCAT TCTCAAGGCGATCAAACAGAATGAGGAAAAATTGGCAAACTACCAGAGTGTGGG CGAGTATCTGTATAAAGAGTATTTCCAGAAGTTTAAGGAAAACAGCAAGGAGTT TACAAACGTCAGAAATAAAAAGGAGTCTTACGAGAGATGCATCGCGCAGTCATT CCTCAAAGATGAGCTGAAGCTGATATTTAAGAAGCAACGCGAATTTGGTTTCTCA TTCTCTAAGAAGTTCGAAGAGGAGGTTCTTTCCGTGGCGTTTTACAAGAGGGCGC TCAAAGACTTCTCCCACCTGGTTGGTAACTGTAGTTTCTTCACGGATGAGAAGCG AGCTCCCAAAAATTCTCCCCTGGCTTTCATGTTTGTTGCCCTGACTCGGATCATTA ACCTGCTGAACAACCTGAAAAATACTGAAGGGATCTTGTATACGAAGGACGACCTAAATGCACTCCTGAATGAAGTGCTCAAAAACGGAACTCTAACCTATAAACAGA CCAAGAAATTACTGGGGCTCTCTGACGACTACGAGTTCAAGGGCGAGAAGGGTA CTTATTTTATCGAATTCAAAAAGTATAAGGAGTTCATTAAAGCATTGGGGGAACA CAACCTCAGCCAGGACGATCTCAATGAAATTGCCAAGGACATCACGCTGATTAA AGACGAGATAAAACTGAAAAAGGCACTGGCCAAGTATGACCTCAACCAGAACC AGATCGACTCTCTGTCCAAGCTGGAGTTCAAAGACCACCTAAACATATCCTTCAA AGCCCTGAAACTGGTCACCCCTCTAATGCTCGAAGGAAAAAAATACGACGAGGC GTGTAATGAACTGAATCTTAAGGTGGCCATCAATGAGGATAAGAAGGACTTTCTT CCAGCCTTTAACGAGACATATTACAAAGACGAGGTCACAAACCCGGTTGTGCTG AGGGCCATAAAAGAGTATCGGAAGGTTCTGAATGCCCTCCTGAAGAAGTACGGC AAAGTGCACAAAATAAATATCGAATTGGCTAGGGAGGTGGGGAAGAACCATTCT CAGCGAGCAAAGATCGAGAAAGAGCAGAATGAGAACTACAAAGCCAAGAAAGA CGCCGAACTGGAGTGCGAAAAGCTGGGGCTTAAAATAAACAGTAAAAACATCCT GAAATTAAGATTGTTCAAAGAGCAAAAGGAGTTTTGCGCCTACTCAGGGGAAAA AATCAAAATATCAGACCTGCAGGACGAGAAAATGCTGGAGATCGACCATATCTA TCCGTATAGCAGGTCATTTGACGATTCCTACATGAACAAAGTGCTTGTGTTTACC AAACAGAACCAAGAAAAGCTGAACCAAACCCCCTTTGAGGCTTTCGGAAACGAC TCAGCCAAGTGGCAGAAAATCGAAGTCCTAGCCAAGAATCTGCCTACAAAAAAA CAAAAGAGGATTCTTGATAAGAACTATAAGGACAAGGAACAGAAAAACTTTAAA GACAGGAACCTGAATGACACGAGGTACATTGCGCGACTGGTTCTAAACTATACC AAAGACTACCTGGATTTCCTCCCTCTGAGCGACGACGAGAATACTAAACTGAAT GATACCCAGAAAGGCTCAAAGGTCCACGTTGAGGCTAAGTCCGGGATGCTGACT AGCGCCCTCCGCCACACGTGGGGCTTCAGCGCCAAAGATCGGAATAATCATCTTC ATCACGCTATTGATGCAGTAATCATAGCCTACGCTAACAACAGCATCGTGAAAG CCTTCTCCGATTTCAAGAAAGAACAGGAGTCTAATAGCGCCGAGTTGTACGCCA AGAAAATTTCCGAATTGGACTATAAAAATAAGAGAAAATTCTTCGAACCCTTCTC CGGGTTTCGCCAAAAGGTCTTAGATAAGATCGACGAGATTTTCGTTTCCAAGCCC GAAAGAAAAAAGCCTTCAGGGGCACTGCACGAAGAGACATTCCGCAAGGAAGA GGAATTTTACCAATCTTACGGTGGTAAAGAGGGAGTTCTGAAGGCTCTGGAGCTT GGGAAGATCCGCAAGGTAAACGGGAAAATCGTGAAAAACGGGGACATGTTCAG GGTGGATATCTTCAAGCACAAAAAGACCAACAAGTTCTACGCAGTACCCATCTA CACTATGGATTTCGCTTTAAAGGTTCTCCCAAATAAGGCGGTGGCTCGATCGAAG AAAGGAGAGATCAAGGACTGGATCTTAATGGATGAAAATTACGAGTTTTGCTTCT CGCTCTACAAAGATAGCCTGATTCTGATCCAGACAAAAAAGATGCAGGAACCAGAATTTGTTTATTATAACGCCTTCACGAGCAGTACAGTGTCCCTGATTGTGAGCAA GCATGATAACAAGTTCGAGACTCTGTCTAAGAATCAGAAAATCCTTTTCAAGAAC GCCAACGAGAAGGAGGTCATCGCAAAGTCAATTGGCATCCAAAACCTGAAGGTG TTCGAGAAATACATAGTGTCCGCACTCGGTGAAGTAACTAAAGCCGAATTTCGA CAGCGCGAGGATTTTAAGAAA (SEQ ID NO: 41) SaCas9
[0218] KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRG ARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALL HLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRG SINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGW KDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYY EKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITA RKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLK AINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSI KVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENA KYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLV KQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDIN RFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWK FKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIE TEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVN NLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYY EETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRF DVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNND LIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKY STDILGNLYEVKSKKHPQIIKKG (SEQ ID NO: 42)
[0219] Residue N579 of SEQ ID NO: SaCas9, which is underlined and in bold, may be mutated (e.g., to a A579) to yield a SaCas9 nickase. enCjCas9 MARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRL ARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSK QDFARVILHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQ KFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGILYTKD DLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHN LSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVT PLMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPVVLRAIKEYRK VLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGL KINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNKVL VFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNF KDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTS ALRHTWGFSAKDRNNHLHHAIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKIS ELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKPERKKPSGALHEETFRKEEEFYQSYG GKEGVLKALELGKIRKVNGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTMDFALKVL PNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKDMQEPEFVYYNAFTSST VSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKSIGIQNLKVFEKYIVSALGEVTKAE FRQREDFKK (SEQ ID NO: 43)
[0220] evoCjCas9 GCCCGCATCCTGGCCTTTGCCATCGGCATCAGCAGCATCGGATGGGCCTTCAGCG AGAATGACGAGCTGAAGGACTGCGGTGTGAGAATCTTTACAAAGGTCGAGAACC CCAAGACCGGCGAAAGCCTGGCGCTGCCAAGACGGCTGGCCCGGTCTGCTCGGA AGAGATACGCCCGGAGAAAAGCCAGACTGAATCACCTGAAACACCTGATCGCCA ACGAGTTCAAACTGAACTACGAGGACTACCAGAGCTTCGATGAGAGCTTGGCTA AGGCCTACAAAGGCAGCCTGATCAGCCCCTACGAGCTGAGATTCAGAGCTCTGA ATGAGCTCCTGAGCAAGCAGGACTTCGCTAGAGTGATCCTGCACATCGCCAAGC GAAGAGGCTACGACGACATCAAGAACTCTGATGACAAGGAGAAGGGCGCCATTC TGAAAGCCATCAAGCAGAACGAAGAGAAACTGGCCAATTACCAGAGCGTGGGC GAGTACCTGTACAAGGAGTACTTCCAGAAGTTCAAAGAAAATAGCAAAGAGTTC ACCAACGTGCGGAACAAGAAGGAGTCTTATGAAAGATGTATCGCCCAGAGCTTC CTGAAAGACGAGCTGAAACTGATCTTCAAGAAGCAAAGAGAGTTTGGCTTCAGC TTCAGCAAAAAATTTGAAGAAGAGGTGCTGTCTGTCGCCTTCTACAAGAGGGCTC TGAAGGACTTCAGCCACCTGGTGGGCAATTGCAGCTTTTTCACAGACGAAAAGC GGGCCCCTAAGAACAGCCCTCTGGCCTTCATGTTCGTGGCTCTGACCAGAATCAT CAACCTGCTGAACAACCTGAAAAATACCGAGGGCATCCTCTATACCAAGGACGA TCTGAACGCCCTGCTGAATGAGGTGCTCAAGAATGGCACCCTGACCTACAAGCA AACAAAAAAACTGCTGGGCCTGTCTGACGACTACGAGTTTAAAGGCGAGAAGGGCACCTATTTCATCGAATTTAAGAAGTACAAGGAATTCATCAAGGCACTGGGCGA ACACAACCTGTCTCAGGACGACCTGAACGAGATCGCCAAGGACATCACCCTGAT CAAGGATGAGATCAAGCTGAAAAAAGCTCTGGCCAAGTACGACCTGAATCAAAA CCAGATCGACAGCTTAAGCAAGCTGGAATTTAAGGATCACCTGAACATCTCCTTT AAGGCCCTGAAGCTGGTGACCCCACTGATGCTGGAAGGAAAGAAGTACGACGAA GCATGCAACGAGCTTAACCTGAAGGTTGCTATCAACGAGGATAAGAAGGATTTC CTGCCTGCCTTTAACGAGACATACTACAAGGATGAGGTGACCAACCCCGTGGTG CTGAGAGCTATCAAAGAGTACAGAAAGGTGCTGAACGCCCTGCTGAAGAAGTAC GGCAAGGTCCACAAGATCAATATTGAGCTGGCCCGGGAGGTTGGAAAGAACCAC TCTCAGAGAGCAAAGATCGAGAAGGAGCAAAACGAGAACTACAAAGCGAAGAA GGACGCCGAACTGGAGTGCGAGAAGCTTGGCCTGAAGATCAACTCTAAGAATAT CCTGAAACTCAGACTTTTCAAAGAACAGAAGGAATTCTGTGCCTACAGCGGCGA GAAGATCAAAATTTCTGACCTGCAGGATGAAAAGATGCTGGAGATCGACCACAT CTACCCTTACTCCAGAAGCTTCGACGACAGCTATATGAACAAAGTGCTGGTGTTC ACAAAGCAGAACCAGGAGAAGCTGAATCAGACCCCTTTCGAGGCCTTCGGCAAT GACTCCGCCAAGTGGCAGAAAATCGAGGTGCTGGCCAAAAACCTGCCAACCAAG AAACAGAAGAGAATCCTCGACAAGAACTACAAGGACAAGGAACAGAAGAACTT CAAGGATCGGAACCTGAACGACACCCGGTACATCGCCAGGCTGGTGTTAAATTA CACCAAGGACTACCTGGATTTCCTGCCCCTGAGCGACGACGAGAACACCAAGCT GAACGACACACAGAAGGGCAGCAAGGTGCACGTGGAAGCCAAGAGCGGCATGC TGACCAGCGCACTGCGCCACACTTGGGGCTTCAGCGCTAAGGACCGGAACAACC ACCTGCATCACGCCATCGATGCCGTGATCATAGCCTACGCCAACAACTCAATCGT GAAAGCTTTCAGTGACTTTAAGAAAGAACAGGAGAGCAACTCTGCCGAACTGTA CGCCAAGAAAATTAGCGAGCTGGACTACAAGAACAAACGGAAGTTCTTCGAACC TTTCTCAGGATTTAGACAGAAGGTGCTGGATAAGATCGATGAAATCTTCGTGTCC AAGCCCGAGAGAAAGAAGCCTAGCGGAGCCCTGCACAAGGAAACCTTCAGAAA GGAAGAAGAGTTCTACCAGTCTTATGGGGGCAAAGAAGGCGTGCTGAAGGCCCT GGAACTCGGCAAGATCAGAAAGGTGAAGGGAAAGATTGTGAAGAACGGCGACA TGTTCAGAGTGGACATCTTCAAGCACAAGAAGACCAACAAGTTCTATGCCGTGC CTATCTACACAATGGATTTCGCCCTGAAAGTGCTGCCTAACAAGGCTGTCGCCAG ATCTAAGAAGGGCGAGATCAAGGACTGGATCCTGATGGACGAAAACTACGAGTT CTGCTTCAGCCTGTACAAGGACAGCCTGATCCTGATCCAAACAAAGAAGATGCA GGAGCCTGAATTCGTGTACTACAACGCCTTCACAAGCAGCACCGTGAGCCTGATC GTGTCTAAACATGATAACAAGTTCGAAACCCTGTCCAAGAACCAAAAGATCCTGTTCAAGAACGCCAACGAGAAGGAAGTGATCGCCAAGGGCATCGGCATTCAGAAT CTGAAGGTGTTCGAGAAATATATCGTGTCCGCTCTGGGAGAGGTTACAAAGGCC GAGTTTCGGCAGAGAGAAGATTTTAAGAAG (SEQ ID NO: 44)
[0221] eNme2-C Cas9 GCAGCATTCAAGTCAAACCCAATCAATTACATCCTGGGACTGGCAATCGGAATC GCATCCGTGGGATGGGCTATGGTGGAGATCGACGAGGAGGGGAATCCTATCCGG CTGATCGATCTGGGCGTGAGAGTGTTTGAGAGGGCCGAGGTGCCAAAGACCGGC GATTCTCTGGCTATGGCCCGGAGACTGGCACGGAGCGTGAGGCGCCTGACACGG AGAAGGGCACACAGGCTGCTGAGGGCACGCCGGCTGCTGAAGAGAGAGGGCGT GCTGCAGGCAGCAGACTTCGATGAGAATGGCCTGATCACGAGCTTGCCAAACAC CCCCTGGCAGCTGAGAGCAGCCGCCCTGGACAGGAAGCTGACACCACTGGAGTG GTCTGCCGTGCTGCTGCACCTGATCAAGCACCGCGGCTACCTGAGCCAGCGGAA GAACGAGGGAGAGACAGCAGCCAAGGAGCTGGGCGCCCTGCTGAAGGGAGTGG CCAACAATGCCCACGCCCTGCAGACCGGCGATTTCAGGACACCTGCCGAGCTGG CCCTGAATAAGTTTGAGAAGGAGTCCGGCCACATCAGAAACCAGAGGGGCGACT ATAGCCACACCTTCTCCCGCAAGGATCTGCAGGCCGAGCTGATCCTGCTGTTCGA GAAGCAGAAGGAGTTTGGCAATCCACACGTGAGCGGAGGCCTGAAGGAGGGAA TCGAGACCCTGCTGATGACACAGAGGCCTGCCCTGTCCGGCGACGCAGTGCAGA AGATGCTGGGGCACTGCACCCTCGAGCCTACAGAGCCAAAGGCCGCCAAGAACA CCTACACAGCCGAGCGGTTTATCTGGCTGACAAAGCTGAACAATCTGAGAATCCT GGAGCAGGGATCCGAGAGGCCACTGACCGACACAGAGAGGTCCACCCTGATGG ATGAGCCTTACCGGAAGTCTAAACTGACATATGCCCAGGCCAGAAAGCTGCTGG GCCTGGAGGACACCGCCTTCTTTAAGGGCCTGAGATACGGCAAGGATAATGCCG AGGCCTCCACACTGATGGAGATGAAGGCCTATCACGCCATCTCTCGCGCCCTGGA GAAGGAGGGCCTGAAGGACAAGAAGTCCCCCCTGAACCTGAGCTCCGAGCTGCA GGATGAGATCGGCACCGCCTTCTCTCTGTTTAAGACCGACGAGGATATCACAGGC CGCCTGAAGGACAGGGTGCAGCCTGAGATCCTGGAGGCCCTGCTGAAGCACATC TCTTTCGATAAGTTTGTGCAGATCAGCCTGAAGGCCCTGAGAAGGATCGTGCCAC TGATGGAGCAGGGCAAGCGGTACGACGAGGCCTGCGCCGAGATCTACGGCGTTC ACTATGGCAAGAAGAACACAGAGGAGAAGATCTATCTGCCCCCTATCCCTGCCG ACGAGATCAGAAATCCTGTGGTGCTGAGGGCCCTGTCCCAGGCAAGAAAAGTGA TCAACGGAGTGGTGCGCCGGTACGGATCTCCAGCCCGGATCCACATCGAGACCG CCAGAGAAGTGGGCAAGAGCTTCAAGGACCGGAAGGAGATCGCGAAGAGACAGGAGGAGAATCGCAAGGATCGGGAGAAGGCCGCCGCCAAGTTTAGGGAGTACTTC CCTAACTTTGTGGGCGAGCCAAAGTCTAAGGACATCCTGAAGCTGCGCCTGTACG AGCAGCAGCACGGCAAGTGTCTGTATAGCGGCAAAGAGATCAATCTGGTGCGGC TGAACGAGAAGGGCTATGTGGAGATCGATCACGCCCTGCCTTTCTCCAGAACCTG GGACGATTCTTTTAACAATAAGGTGCTGGTGCTGGGCAGCGAGAACCAGAATAA GGGCAATCAGACACCATACGAGTATTTCAATGGCAAGGACAACTCCAGGGAGTG GCAGGAGTTCAAGGCCCGCGTGGAGACCTCTAGATTTCCCAGTAGCAAGAAGCA GCGGATCCTGCTGCAGAAGTTCGACGAGGATGGCTTTAAGGAGTGCAACCTGAA TGACACCAGATACGTGAACCGGTTCCTGTGCCAGTTTGTGGCCGATCACATCCTG CTGACCGGCAAGGGCAAGAGAAGGGTGGTCGCCTCTAATGGCCAGATCACAAAC CTGCTGAGGGGGTTTTGGAGACTGAGGAAGGTGCGGGCAGAGAATGACAGACAC CACGCACTGGATGCAGTGGTGGTGGCATGCAGCACCGTGGCAATGCAGCAGAAG ATCACAAGATTCGTGAGGTATAAGGAGATGAACGCCTTTGACGGCAAGACCGTC GATAAGGAGACAGGCAAGGTGCTGTACCAGAAGACCCACTTCCCCCAGCCTTGG GAGTTCTTTGCCCAGGAAGTTATGATCCGGGTGTTCGGCAAGCCAGACGGCAAG CCTGAGTTTGAGGAGGCCGATACCCCAGAGAAGCTGAGGACACTGCTGGCAGAG AAGCTGTCTAGCAGGCCAGAGGCAGTGCACGAGTACGTGACCCCGCTGTTCGTG TCCAGGGCACCCAATCGGAAGATGTCTGGCGCCCACAAGGACACACTGAGAAGC GCCAAGAGGTTTGTGAAGCACAACGAGAAGATCTCCGTGAAGAGAGTGTGGCTG ACCGAGATCAAGCTGGCCGATCTGGAGAACATGGTGAATTACAAGAACGGCAGG GAGATCGAGCTGTATGAGGCCCTGAAGGCAAGGCTGGAGGCCTACGGAGGAAAT GCCAAGCAGGCCTTCGACCCAAAGGATAACCCCTTTTATAAGAAGGGAGGACAG CTGGTGAAGGCCGTGCGGGTGGAGAAGACCCAGAAGAGCGGCGTGCTGCTGAAT AAGAAGAACGCCTACACAATCGCCGACAATGGTGATATGGTGAGAGTGGACGTG TTCTGTAAGGTGGATAAGAAGGGCAAGAATCAGTACTTTATCGTGCCTATCTATG CCTGGCAGGTGGCCGAGAACATCCTGCCAGACATCGATTGCAAGGGCTACAGAA TCGACGATAGCTATACATTCTGTTTTTCCCTGCACAAGTATGACCTGATCGCCTTC CAGAAGGATGAGAAGTCCAAGGTGGAGTTTGCCTACTATATCAATTGCGACTCCT CTAGCGGCGGGTTCTACCTGGCCTGGCACGATAAGGGCAGCAGGGAGCAGCGGT TTCGCATCTCCACCCAGAATCTGGCGCTGATCCAGAAGTATCAGGTGAACGAGCT GGGCAAGGAGATCAGGCCATGTCGGCTGAAGAAGCGCCCACCCGTGCGG (SEQ ID NO: 45).
[0222] S. pyogenes Cas9 wild type (NCBI Reference Sequence: NC_002737.2, Uniprot Reference Sequence: Q99ZW2)
[0223] MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGA LLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVE EDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFR GHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLE NLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKAL VRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPL ARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKED YFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFE DREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFL KSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKE HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDN KVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFR KDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDF ATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSP TVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLI IKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIH LFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 46)
[0224] S. pyogenes dCas9 (D10A and H840A)
[0225] MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGA LLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVE EDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFR GHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLE NLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKAL VRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPL ARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKED YFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFE DREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFL KSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKE HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDN KVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFR KDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDF ATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSP TVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLI IKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIH LFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (SEQ ID NO: 47)
[0226] S. pyogenes Cas9 Nickase (D10A)
[0227] MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGA LLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVE EDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFR GHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLE NLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKAL VRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPL ARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKED YFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFE DREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKE HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDN KVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFR KDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDF ATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSP TVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLI IKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIH LFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 48)
[0228] VRER-nCas9 (D10A / D1135V / G1218R / R1335E / T1337R) S. pyogenes Cas9 Nickase
[0229] MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGA LLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVE EDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFR GHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLE NLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKAL VRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPL ARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKED YFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFE DREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFL KSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKE HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDN KVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFR KDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDF ATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFVSP TVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLI IKLPKYSLFELENGRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIH LFTLTNLGAPAAFKYFDTTIDRKEYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 49)
[0230] VQR-nCas9 (D10A / D1135V / R1335Q / T1337R) S. pyogenes Cas9 Nickase
[0231] MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGA LLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVE EDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFR GHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLE NLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKAL VRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPL ARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKED YFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFE DREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFL KSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKE HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDN KVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFR KDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDF ATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFVSP TVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLI IKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIH LFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 50)
[0232] EQR-nCas9 (D10A / D1135E / R1335Q / T1337R) S. pyogenes Cas9 Nickase
[0233] MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGA LLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVE EDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFR GHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLE NLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKAL VRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPL ARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKED YFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFE DREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFL KSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKE HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDN KVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFR KDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDF ATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFESP TVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLI IKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIH LFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 51)
[0234] VRQR-nCas9 (D10A / D1135V / G1218R / R1335Q / T1337R) S. pyogenes Cas9 Nickase
[0235] MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGA LLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVE EDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFR GHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKAL VRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPL ARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKED YFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFE DREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFL KSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTV KVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKE HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDN KVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLS ELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFR KDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDF ATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFVSP TVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLI IKLPKYSLFELENGRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIH LFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 52)
[0236] SaKKH-nCas9 (D10A / E782K / N968K / R1015H) S. aureus Cas9 Nickase
[0237] MKRNYILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKR GARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAA LLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEV RGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFG WKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEY YEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDIT ARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSL KAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQ SIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKEN AKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVL VKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKW KFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEI ETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLIV NNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKY YEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYR FDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYKN DLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKK YSTDILGNLYEVKSKKHPQIIKKG (SEQ ID NO: 53)
[0238] Streptococcus thermophilus CRISPR1 Cas9 (St1Cas9) Nickase (D9A)
[0239] MSDLVLGLAIGIGSVGVGILNKVTGEIIHKNSRIFPAAQAENNLVRRTNRQ GRRLTRRKKHRRVRLNRLFEESGLITDFTKISINLNPYQLRVKGLTDELSNEELFIALK NMVKHRGISYLDDASDDGNSSIGDYAQIVKENSKQLETKTPGQIQLERYQTYGQLRG DFTVEKDGKKHRLINVFPTSAYRSEALRILQTQQEFNPQITDEFINRYLEILTGKRKYY HGPGNEKSRTDYGRYRTSGETLDNIFGILIGKCTFYPDEFRAAKASYTAQEFNLLNDL NNLTVPTETKKLSKEQKNQIINYVKNEKAMGPAKLFKYIAKLLSCDVADIKGYRIDK SGKAEIHTFEAYRKMKTLETLDIEQMDRETLDKLAYVLTLNTEREGIQEALEHEFAD GSFSQKQVDELVQFRKANSSIFGKGWHNFSVKLMMELIPELYETSEEQMTILTRLGK QKTTSSSNKTKYIDEKLLTEEIYNPVVAKSVRQAIKIVNAAIKEYGDFDNIVIEMARET NEDDEKKAIQKIQKANKDEKDAAMLKAANQYNGKAELPHSVFHGHKQLATKIRLW HQQGERCLYTGKTISIHDLINNSNQFEVDHILPLSITFDDSLANKVLVYATANQEKGQ RTPYQALDSMDDAWSFRELKAFVRESKTLSNKKKEYLLTEEDISKFDVRKKFIERNL VDTRYASRVVLNALQEHFRAHKIDTKVSVVRGQFTSQLRRHWGIEKTRDTYHHHAV DALIIAASSQLNLWKKQKNTLVSYSEDQLLDIETGELISDDEYKESVFKAPYQHFVDT LKSKEFEDSILFSYQVDSKFNRKISDATIYATRQAKVGKDKADETYVLGKIKDIYTQD GYDAFMKIYKKDKSKFLMYRHDPQTFEKVIEPILENYPNKQINEKGKEVPCNPFLKY KEEHGYIRKYSKKGNGPEIKSLKYYDSKLGNHIDITPKDSNNKVVLQSVSPWRADVY FNKTTGKYEILGLKYADLQFEKGTGTYKISQEKYNDIKKKEGVDSDSEFKFTLYKND LLLVKDTETKEQQLFRFLSRTMPKQKHYVELKPYDKQKFEGGEALIKVLGNVANSG QCKKGLGKSNISIYKVRTDVLGNQHIIKNEGDKPKLDF (SEQ ID NO: 54)
[0240] Streptococcus thermophilus CRISPR3Cas9 (St3Cas9) Nickase (D10A)
[0241] MTKPYSIGLAIGTNSVGWAVITDNYKVPSKKMKVLGNTSKKYIKKNLLG VLLFDSGITAEGRRLKRTARRRYTRRRNRILYLQEIFSTEMATLDDAFFQRLDDSFLVPDDKRDSKYPIFGNLVEEKVYHDEFPTIYHLRKYLADSTKKADLRLVYLALAHMIKY RGHFLIEGEFNSKNNDIQKNFQDFLDTYNAIFESDLSLENSKQLEEIVKDKISKLEKKD RILKLFPGEKNSGIFSEFLKLIVGNQADFRKCFNLDEKASLHFSKESYDEDLETLLGYI GDDYSDVFLKAKKLYDAILLSGFLTVTDNETEAPLSSAMIKRYNEHKEDLALLKEYI RNISLKTYNEVFKDDTKNGYAGYIDGKTNQEDFYVYLKNLLAEFEGADYFLEKIDRE DFLRKQRTFDNGSIPYQIHLQEMRAILDKQAKFYPFLAKNKERIEKILTFRIPYYVGPL ARGNSDFAWSIRKRNEKITPWNFEDVIDKESSAEAFINRMTSFDLYLPEEKVLPKHSL LYETFNVYNELTKVRFIAESMRDYQFLDSKQKKDIVRLYFKDKRKVTDKDIIEYLHAI YGYDGIELKGIEKQFNSSLSTYHDLLNIINDKEFLDDSSNEAIIEEIIHTLTIFEDREMIK QRLSKFENIFDKSVLKKLSRRHYTGWGKLSAKLINGIRDEKSGNTILDYLIDDGISNR NFMQLIHDDALSFKKKIQKAQIIGDEDKGNIKEVVKSLPGSPAIKKGILQSIKIVDELV KVMGGRKPESIVVEMARENQYTNQGKSNSQQRLKRLEKSLKELGSKILKENIPAKLS KIDNNALQNDRLYLYYLQNGKDMYTGDDLDIDRLSNYDIDHIIPQAFLKDNSIDNKV LVSSASNRGKSDDFPSLEVVKKRKTFWYQLLKSKLISQRKFDNLTKAERGGLLPEDK AGFIQRQLVETRQITKHVARLLDEKFNNKKDENNRAVRTVKIITLKSTLVSQFRKDFE LYKVREINDFHHAHDAYLNAVIASALLKKYPKLEPEFVYGDYPKYNSFRERKSATEK VYFYSNIMNIFKKSISLADGRVIERPLIEVNEETGESVWNKESDLATVRRVLSYPQVN VVKKVEEQNHGLDRGKPKGLFNANLSSKPKPNSNENLVGAKEYLDPKKYGGYAGIS NSFAVLVKGTIEKGAKKKITNVLEFQGISILDRINYRKDKLNFLLEKGYKDIELIIELPK YSLFELSDGSRRMLASILSTNNKRGEIHKGNQIFLSQKFVKLLYHAKRISNTINENHRK YVENHKKEFEELFYYILEFNENYVGAKKNGKLLNSAFQSWQNHSIDELCSSFIGPTGS ERKGLFELTSRGSAADFEFLGVKIPRYRDYTPSSLLKDATLIHQSVTGLYETRIDLAKL GEG (SEQ ID NO: 55)
[0242] S. aureus Cas9 wild type
[0243] MKRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKR GARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAA LLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEV RGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFG WKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEY YEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDIT ARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSL KAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQ SIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVL VKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDI NRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKW KFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEI ETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIV NNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKY YEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYR FDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNN DLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKK YSTDILGNLYEVKSKKHPQIIKKG (SEQ ID NO: 56)
[0244] S. aureus Cas9 Nickase (D10A)
[0245] MKRNYILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKR GARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAA LLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEV RGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFG WKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEY YEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDIT ARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSL KAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQ SIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKEN AKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVL VKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDI NRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKW KF1KKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMP EIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYK YYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPY RFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYN NDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIK KYSTDILGNLYEVKSKKHPQIIKKG (SEQ ID NO: 57)
[0246] Streptococcus thermophilus wild type CRISPR3 Cas9 (St3Cas9)
[0247] MTKPYSIGLDIGTNSVGWAVITDNYKVPSKKMKVLGNTSKKYIKKNLLG VLLFDSGITAEGRRLKRTARRRYTRRRNRILYLQEIFSTEMATLDDAFFQRLDDSFLV PDDKRDSKYPIFGNLVEEKVYHDEFPTIYHLRKYLADSTKKADLRLVYLALAHMIKY RGHFLIEGEFNSKNNDIQKNFQDFLDTYNAIFESDLSLENSKQLEEIVKDKISKLEKKD RILKLFPGEKNSGIFSEFLKLIVGNQADFRKCFNLDEKASLHFSKESYDEDLETLLGYI GDDYSDVFLKAKKLYDAILLSGFLTVTDNETEAPLSSAMIKRYNEHKEDLALLKEYI RNISLKTYNEVFKDDTKNGYAGYIDGKTNQEDFYVYLKNLLAEFEGADYFLEKIDRE DFLRKQRTFDNGSIPYQIHLQEMRAILDKQAKFYPFLAKNKERIEKILTFRIPYYVGPL ARGNSDFAWSIRKRNEKITPWNFEDVIDKESSAEAFINRMTSFDLYLPEEKVLPKHSL LYETFNVYNELTKVRFIAESMRDYQFLDSKQKKDIVRLYFKDKRKVTDKDIIEYLHAI YGYDGIELKGIEKQFNSSLSTYHDLLNIINDKEFLDDSSNEAIIEEIIHTLTIFEDREMIK QRLSKFENIFDKSVLKKLSRRHYTGWGKLSAKLINGIRDEKSGNTILDYLIDDGISNR NFMQLIHDDALSFKKKIQKAQIIGDEDKGNIKEVVKSLPGSPAIKKGILQSIKIVDELV KVMGGRKPESIVVEMARENQYTNQGKSNSQQRLKRLEKSLKELGSKILKENIPAKLS KIDNNALQNDRLYLYYLQNGKDMYTGDDLDIDRLSNYDIDHIIPQAFLKDNSIDNKV LVSSASNRGKSDDFPSLEVVKKRKTFWYQLLKSKLISQRKFDNLTKAERGGLLPEDK AGFIQRQLVETRQITKHVARLLDEKFNNKKDENNRAVRTVKIITLKSTLVSQFRKDFE LYKVREINDFHHAHDAYLNAVIASALLKKYPKLEPEFVYGDYPKYNSFRERKSATEK VYFYSNIMNIFKKSISLADGRVIERPLIEVNEETGESVWNKESDLATVRRVLSYPQVN VVKKVEEQNHGLDRGKPKGLFNANLSSKPKPNSNENLVGAKEYLDPKKYGGYAGIS NSFAVLVKGTIEKGAKKKITNVLEFQGISILDRINYRKDKLNFLLEKGYKDIELIIELPK YSLFELSDGSRRMLASILSTNNKRGEIHKGNQIFLSQKFVKLLYHAKRISNTINENHRK YVENHKKEFEELFYYILEFNENYVGAKKNGKLLNSAFQSWQNHSIDELCSSFIGPTGS ERKGLFELTSRGSAADFEFLGVKIPRYRDYTPSSLLKDATLIHQSVTGLYETRIDLAKL GEG (SEQ ID NO: 58)
[0248] Streptococcus thermophilus CRISPR1 Cas9 wild type (St1Cas9)
[0249] MSDLVLGLDIGIGSVGVGILNKVTGEIIHKNSRIFPAAQAENNLVRRTNRQ GRRLTRRKKHRRVRLNRLFEESGLITDFTKISINLNPYQLRVKGLTDELSNEELFIALK NMVKHRGISYLDDASDDGNSSIGDYAQIVKENSKQLETKTPGQIQLERYQTYGQLRG DFTVEKDGKKHRLINVFPTSAYRSEALRILQTQQEFNPQITDEFINRYLEILTGKRKYY HGPGNEKSRTDYGRYRTSGETLDNIFGILIGKCTFYPDEFRAAKASYTAQEFNLLNDL NNLTVPTETKKLSKEQKNQIINYVKNEKAMGPAKLFKYIAKLLSCDVADIKGYRIDK SGKAEIHTFEAYRKMKTLETLDIEQMDRETLDKLAYVLTLNTEREGIQEALEHEFADGSFSQKQVDELVQFRKANSSIFGKGWHNFSVKLMMELIPELYETSEEQMTILTRLGK QKTTSSSNKTKYIDEKLLTEEIYNPVVAKSVRQAIKIVNAAIKEYGDFDNIVIEMARET NEDDEKKAIQKIQKANKDEKDAAMLKAANQYNGKAELPHSVFHGHKQLATKIRLW HQQGERCLYTGKTISIHDLINNSNQFEVDHILPLSITFDDSLANKVLVYATANQEKGQ RTPYQALDSMDDAWSFRELKAFVRESKTLSNKKKEYLLTEEDISKFDVRKKFIERNL VDTRYASRVVLNALQEHFRAHKIDTKVSVVRGQFTSQLRRHWGIEKTRDTYHHHAV DALIIAASSQLNLWKKQKNTLVSYSEDQLLDIETGELISDDEYKESVFKAPYQHFVDT LKSKEFEDSILFSYQVDSKFNRKISDATIYATRQAKVGKDKADETYVLGKIKDIYTQD GYDAFMKIYKKDKSKFLMYRHDPQTFEKVIEPILENYPNKQINEKGKEVPCNPFLKY KEEHGYIRKYSKKGNGPEIKSLKYYDSKLGNHIDITPKDSNNKVVLQSVSPWRADVY FNKTTGKYEILGLKYADLQFEKGTGTYKISQEKYNDIKKKEGVDSDSEFKFTLYKND LLLVKDTETKEQQLFRFLSRTMPKQKHYVELKPYDKQKFEGGEALIKVLGNVANSG QCKKGLGKSNISIYKVRTDVLGNQHIIKNEGDKPKLDF (SEQ ID NO: 59)
[0250] CasX from Sulfolobus islandicus (strain REY15A)
[0251] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKN NEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEV FKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWV LTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSS VTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSN MRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGE LIRGEG (SEQ ID NO: 60)
[0252] CasY from Sulfolobus islandicus (strain REY15A)
[0253] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKN NEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEV FKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVL TRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSV TNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSN MRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGE LIRGEG (SEQ ID NO: 61)
[0254] In some embodiments, the napDNAbp domain comprises a nucleotide sequence or amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or is 100% identical to SEQ ID NO: 39-61 Deaminase domains
[0255] In some embodiments, the nucleobase editors (or a recombined nucleobase editor) disclosed herein comprise a deaminase domain. A deaminase domain may be a cytosine deaminase domain or an adenosine deaminase domain.
[0256] Nucleobase editors configured to convert a C to T, in some embodiments, comprise a cytosine deaminase. A “cytosine deaminase” refers to an enzyme that catalyzes the chemical reaction “cytosine + H2O → uracil + NH3” or “5-methyl-cytosine + H2O → thymine + NH3.” As it may be apparent from the reaction formula, such chemical reactions result in a C to U / T nucleobase change. In the context of a gene, such a nucleotide change, or mutation, may in turn lead to an amino acid change in the protein, which may affect the protein’s function, e.g., loss-of-function or gain-of-function. In some embodiments, the C to T base editor comprises a SpCas9, enCjCas9, evoCjCas9, SauriCas9, eNme2-C Cas9, or SaCas9 or a variant thereof (e.g., a nickase) provided herein fused to a cytosine deaminase. In some embodiments, the cytosine deaminase domain is fused to the N-terminus of the Cas9 variant.
[0257] Non-limiting examples of suitable cytosine deaminase domains are provided below, as SEQ ID NOs: 62-97.
[0258] Human AID
[0259] MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGY LRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNL SLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKA WEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 62)
[0260] Mouse AID
[0261] MDSLLMKQKKFLYHFKNVRWAKGRHETYLCYVVKRRDSATSCSLDFGH LRNKSGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVAEFLRWNPNL SLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIGIMTFKDYFYCWNTFVENRERTFKA WEGLHENSVRLTRQLRRILLPLYEVDDLRDAFRMLGF (SEQ ID NO: 63)
[0262] Dog AID
[0263] MDSLLMKQRKFLYHFKNVRWAKGRHETYLCYVVKRRDSATSFSLDFGH LRNKSGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGYPNLSLRIFAARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENREKTFKA WEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 64)
[0264] Bovine AID
[0265] MDSLLKKQRQFLYQFKNVRWAKGRHETYLCYVVKRRDSPTSFSLDFGHL RNKAGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGYPNLS LRIFTARLYFCDKERKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKA WEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 65)
[0266] Mouse APOBEC-3
[0267] MGPFCLGCSHRKCYSPIRNLISQETFKFHFKNLGYAKGRKDTFLCYEVTR KDCDSPVSLHHGVFKNKDNIHAEICFLYWFHDKVLKVLSPREEFKITWYMSWSPCFE CAEQIVRFLATHHNLSLDIFSSRLYNVQDPETQQNLCRLVQEGAQVAAMDLYEFKK CWKKFVDNGGRRFRPWKRLLTNFRYQDSKLQEILRPCYIPVPSSSSSTLSNICLTKGL PETRFCVEGRRMDPLSEEEFYSQFYNQRVKHLCYYHRMKPYLCYQLEQFNGQAPLK GCLLSEKGKQHAEILFLDKIRSMELSQVTITCYLTWSPCPNCAWQLAAFKRDRPDLIL HIYTSRLYFHWKRPFQKGLCSLWQSGILVDVMDLPQFTDCWTNFVNPKRPFWPWK GLEIISRRTQRRLRRIKESWGLQDLVNDFGNLQLGPPMS (SEQ ID NO: 66)
[0268] Rat APOBEC-3
[0269] MGPFCLGCSHRKCYSPIRNLISQETFKFHFKNLRYAIDRKDTFLCYEVTRK DCDSPVSLHHGVFKNKDNIHAEICFLYWFHDKVLKVLSPREEFKITWYMSWSPCFEC AEQVLRFLATHHNLSLDIFSSRLYNIRDPENQQNLCRLVQEGAQVAAMDLYEFKKC WKKFVDNGGRRFRPWKKLLTNFRYQDSKLQEILRPCYIPVPSSSSSTLSNICLTKGLP ETRFCVERRRVHLLSEEEFYSQFYNQRVKHLCYYHGVKPYLCYQLEQFNGQAPLKG CLLSEKGKQHAEILFLDKIRSMELSQVIITCYLTWSPCPNCAWQLAAFKRDRPDLILHI YTSRLYFHWKRPFQKGLCSLWQSGILVDVMDLPQFTDCWTNFVNPKRPFWPWKGL EIISRRTQRRLHRIKESWGLQDLVNDFGNLQLGPPMS (SEQ ID NO: 67)
[0270] Rhesus macaque APOBEC-3G
[0271] MVEPMDPRTFVSNFNNRPILSGLNTVWLCCEVKTKDPSGPPLDAKIFQGK VYSKAKYHPEMRFLRWFHKWRQLHHDQEYKVTWYVSWSPCTRCANSVATFLAKD PKVTLTIFVARLYYFWKPDYQQALRILCQKRGGPHATMKIMNYNEFQDCWNKFVD GRGKPFKPRNNLPKHYTLLQATLGELLRHLMDPGTFTSNFNNKPWVSGQHETYLCY KVERLHNDTWVPLNQHRGFLRNQAPNIHGFPKGRHAELCFLDLIPFWKLDGQQYRV TCFTSWSPCFSCAQEMAKFISNNEHVSLCIFAARIYDDQGRYQEGLRALHRDGAKIA MMNYSEFEYCWDTFVDRQGRPFQPWDGLDEHSQALSGRLRAI (SEQ ID NO: 68)
[0272] Chimpanzee APOBEC-3G
[0273] MKPHFRNPVERMYQDTFSDNFYNRPILSHRNTVWLCYEVKTKGPSRPPLD AKIFRGQVYSKLKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDVA TFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCW SKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTSNFNNELWVRGRHET YLCYEVERLHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLH QDYRVTCFTSWSPCFSCAQEMAKFISNNKHVSLCIFAARIYDDQGRCQEGLRTLAKA GAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLEEHSQALSGRLRAILQNQGN (SEQ ID NO: 69)
[0274] Green monkey APOBEC-3G
[0275] MNPQIRNMVEQMEPDIFVYYFNNRPILSGRNTVWLCYEVKTKDPSGPPLD ANIFQGKLYPEAKDHPEMKFLHWFRKWRQLHRDQEYEVTWYVSWSPCTRCANSVA TFLAEDPKVTLTIFVARLYYFWKPDYQQALRILCQERGGPHATMKIMNYNEFQHCW NEFVDGQGKPFKPRKNLPKHYTLLHATLGELLRHVMDPGTFTSNFNNKPWVSGQRE TYLCYKVERSHNDTWVLLNQHRGFLRNQAPDRHGFPKGRHAELCFLDLIPFWKLDD QQYRVTCFTSWSPCFSCAQKMAKFISNNKHVSLCIFAARIYDDQGRCQEGLRTLHRD GAKIAVMNYSEFEYCWDTFVDRQGRPFQPWDGLDEHSQALSGRLRAI (SEQ ID NO: 70)
[0276] Human APOBEC-3G
[0277] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLD AKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMA TFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCW SKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTFNFNNEPWVRGRHETY LCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLD QDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEA GAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN (SEQ ID NO: 71)
[0278] Human APOBEC-3F
[0279] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPRL DAKIFRGQVYSQPEHHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAE FLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVY SEGQPFMPWYKFDDNYAFLHRTLKEILRNPMEAMYPHIFYFHFKNLRKAYGRNESW LCFTMEVVKHHSPVSWKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLYYFWDTDYQEGLRSLSQEGASVEI MGYKDFKYCWENFVYNDDEPFKPWKGLKYNFLFLDSKLQEILE (SEQ ID NO: 72)
[0280] Human APOBEC-3B
[0281] MNPQIRNPMERMYRDTFYDNFENEPILYGRSYTWLCYEVKIKRGRSNLL WDTGVFRGQVYFKPQYHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKL AEFLSEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVTIMDYEEFAYCWENF VYNEGQQFMPWYKFDENYAFLHRTLKEILRYLMDPDTFTFNFNNDPLVLRRRQTYL CYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQ IYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRD AGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQNQGN (SEQ ID NO: 73)
[0282] Human APOBEC-3C
[0283] MNPQIRNPMKAMYPGTFYFQFKNLWEANDRNETWLCFTVEGIKRRSVVS WKTGVFRNQVDSETHCHAERCFLSWFCDDILSPNTKYQVTWYTSWSPCPDCAGEV AEFLARHSNVNLTIFTARLYYFQYPCYQEGLRSLSQEGVAVEIMDYEDFKYCWENFV YNDNEPFKPWKGLKTNFRLLKRRLRESLQ (SEQ ID NO: 74)
[0284] Human APOBEC-3A
[0285] MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMD QHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGC AGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCW DTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN (SEQ ID NO: 75)
[0286] Human APOBEC-3H
[0287] MALLTAETFRLQFNNKRRLRRPYYPRKALLCYQLTPQNGSTPTRGYFENK KKCHAEICFINEIKSMGLDETQCYQVTCYLTWSPCSSCAWELVDFIKAHDHLNLGIFA SRLYYHWCKPQQKGLRLLCGSQVPVEVMGFPKFADCWENFVDHEKPLSFNPYKML EELDKNSRAIKRRLERIKIPGVRAQGRYMDILCDAEV (SEQ ID NO: 76)
[0288] Human APOBEC-3D
[0289] MNPQIRNPMERMYRDTFYDNFENEPILYGRSYTWLCYEVKIKRGRSNLL WDTGVFRGPVLPKRQSNHRQEVYFRFENHAEMCFLSWFCGNRLPANRRFQITWFVS WNPCLPCVVKVTKFLAEHPNVTLTISAARLYYYRDRDWRWVLLRLHKAGARVKIM DYEDFAYCWENFVCNEGQPFMPWYKFDDNYASLHRTLKEILRNPMEAMYPHIFYFH FKNLLKACGRNESWLCFTMEVTKHHSAVFRKRGVFRNQVDPETHCHAERCFLSWF CDDILSPNTNYEVTWYTSWSPCPECAGEVAEFLARHSNVNLTIFTARLCYFWDTDYQEGLCSLSQEGASVKIMGYKDFVSCWKNFVYSDDEPFKPWKGLQTNFRLLKRRLREIL Q (SEQ ID NO: 77)
[0290] Human APOBEC-1
[0291] MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKI WRSSGKNTTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHP GVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGD EAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPP HILLATGLIHPSVAWR (SEQ ID NO: 78
[0292] Mouse APOBEC-1
[0293] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVW RHTSQNTSNHVEVNFLEKFTTERYFRPNTRCSITWFLSWSPCGECSRAITEFLSRHPY VTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYPPSNEAY WPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWA TGLK (SEQ ID NO: 79)
[0294] Rat APOBEC-1
[0295] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHV TLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 80)
[0296] Petromyzon marinus CDA1 (pmCDA1)
[0297] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACF WGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILE WYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCR KIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAV (SEQ ID NO: 81)
[0298] Evolved pmCDA1 (evoCDA1)
[0299] MTDAEYVRIHEKLDIYTFKKQFSNNKKSVSHRCYVLFELKRRGERRACF WGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILE WYNQELRGNGHTLKIWVCKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCR KIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMFQVKILHTTKSPAV (SEQ ID NO: 82)
[0300] Human APOBEC3G D316R_D317R
[0301] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLD AKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMA TFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKIMNYDEFQHCW SKFVYSQRELFEPWNNLPKYYILLHIMLGEILRHSMDPPTFTFNFNNEPWVRGRHETY LCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLD QDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEA GAKISIMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN (SEQ ID NO: 83)
[0302] Human APOBEC3G chain A
[0303] MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCN QAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISK NKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPF QPWDGLDEHSQDLSGRLRAILQ (SEQ ID NO: 84)
[0304] Human APOBEC3G chain A D120R_D121R
[0305] MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCN QAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISK NKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISIMTYSEFKHCWDTFVDHQGCPF QPWDGLDEHSQDLSGRLRAILQ (SEQ ID NO: 85)
[0306] evo APOBEC1
[0307] MSSKTGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPNV TLFIYIARLYHLANPRNRQGLRDLISSGVTIQIMTEQESGYCWHNFVNYSPSNESHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQSQLTSFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 86)
[0308] YE1
[0309] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSYSPCGECSRAITEFLSRYPHV TLFIYIARLYHHADPENRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 87)
[0310] YE2
[0311] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSYSPCGECSRAITEFLSRYPHV TLFIYIARLYHHADPRNRQGLEDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 88)
[0312] YEE
[0313] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSYSPCGECSRAITEFLSRYPHV TLFIYIARLYHHADPENRQGLEDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 89)
[0314] EE
[0315] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHV TLFIYIARLYHHADPENRQGLEDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 90)
[0316] R33A
[0317] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELAKETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHV TLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 91)
[0318] R33A+K34A
[0319] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELAAETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHV TLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 92)
[0320] AALN
[0321] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELAAETCLLYEINWGGRHSIW RHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHV TLFIYIARLYHLANPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATG LK (SEQ ID NO: 93)
[0322] FERNY
[0323] MFERNYDPRELRKETYLLYEIKWGKSGKLWRHWCQNNRTQHAEVYFLE NIFNARRFNPSTHCSITWYLSWSPCAECSQKIVDFLKEHPNVNLEIYVARLYYHEDER NRQGLRDLVNSGVTIRIMDLPDYNYCWKTFVSDQGGDEDYWPGHFAPWIKQYSLK L (SEQ ID NO: 94)
[0324] evo FERNY
[0325] MFERNYDPRELRKETYLLYEIKWGKSGKLWRHWCQNNRTQHAEVYFLE NIFNARRFNPSTHCSITWYLSWSPCAECSQKIVDFLKEHPNVNLEIYVARLYYPENER NRQGLRDLVNSGVTIRIMDLPDYNYCWKTFVSDQGGDEDYWPGHFAPWIKQYSLK L (SEQ ID NO: 95)
[0326] APOBEC deaminase
[0327] TCCTCAGAGACTGGGCCTGTCGCCGTCGATCCAACCCTGCGCCGCCGG ATTGAACCTCACGAGTTTGAAGTGTTCTTTGACCCCCGGGAGCTGAGAAAGGAG ACATGCCTGCTGTACGAGATCAACTGGGGAGGCAGGCACTCCATCTGGAGGCAC ACCTCTCAGAACACAAATAAGCACGTGGAGGTGAACTTCATCGAGAAGTTTACC ACAGAGCGGTACTTCTGCCCCAATACCAGATGTAGCATCACATGGTTTCTGAGCT GGTCCCCTTGCGGAGAGTGTAGCAGGGCCATCACCGAGTTCCTGTCCAGATATCC ACACGTGACACTGTTTATCTACATCGCCAGGCTGTATCACCACGCAGACCCAAGG AATAGGCAGGGCCTGCGCGATCTGATCAGCTCCGGCGTGACCATCCAGATCATG ACAGAGCAGGAGTCCGGCTACTGCTGGCGGAACTTCGTGAATTATTCTCCTAGCA ACGAGGCCCACTGGCCTAGGTACCCACACCTGTGGGTGCGCCTGTACGTGCTGG AGCTGTATTGCATCATCCTGGGCCTGCCCCCTTGTCTGAATATCCTGCGGAGAAA GCAGCCCCAGCTGACCTTCTTTACAATCGCCCTGCAGTCTTGTCACTATCAGAGG CTGCCACCCCACATCCTGTGGGCCACAGGCCTGAAG (SEQ ID NO: 96)
[0328] TadCBEd
[0329] AGTTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTG ACCCTGGCCAAGAGGGCACGGGATGAGAGGAAGGCGCCTGTGGGAGCCGTGCT GGTGCTGAACAATAGAGTGATCGGCGAGGGCTGGAACAGAGCCATCGGCCTGCA CGACCCAACAGCCCATGCCGAAATTATAGCCCTGAGACAGGGCGGCCTGGTCAT GCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTG ATGTGCGCCGGCGCCATGATCAACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGA GGAACTCAAAAAGAGGCGCCGCAGGCTCCCTGATGAACGTGCTGAACTACCCCG GAATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCG CCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAA GGCCCAGAGCTCCATCAAC (SEQ ID NO: 97)
[0330] In some embodiments, the cytidine deaminase comprises a nucleotide sequence or amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or is 100% identical to SEQ ID NO: 62-97.
[0331] In some embodiments, a nucleobase editor (or a recombined nucleobase editor) converts an A to G. In some embodiments, the base editor comprises an adenosine deaminase. An “adenosine deaminase” is an enzyme involved in purine metabolism. It is needed for the breakdown of adenosine from food and for the turnover of nucleic acids in tissues. Its primary function in humans is the development and maintenance of the immune system. An adenosine deaminase catalyzes hydrolytic deamination of adenosine (forming inosine, which base pairs as G) in the context of DNA. There are no known adenosine deaminases that act on DNA. Instead, known adenosine deaminase enzymes only act on RNA (tRNA or mRNA). Evolved deoxyadenosine deaminase enzymes that accept DNA substrates and deaminate dA to deoxyinosine for use in adenosine nucleobase editors have been described, e.g., in PCT Application PCT / US2017 / 045381, filed August 3, 2017, which published as WO 2018 / 027078, PCT Application No. PCT / US2019 / 033848, which published as WO 2019 / 226953, PCT Application No PCT / US2019 / 033848, filed May 23, 2019, and PCT Application No. PCT / US2020 / 028568, filed April 17, 2020; each of which is herein incorporated by reference. Non-limiting examples of evolved adenosine deaminases that accept DNA as substrates are provided below. In some embodiments, an adenosine deaminase comprises any of the following amino acid sequences, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or at least 99.9% identical to any of the following amino acid sequences (SEQ ID NOs: 98- 166):
[0332] ecTadA
[0333] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 98)
[0334] ecTadA (D108N)
[0335] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 99)
[0336] ecTadA (D108G)
[0337] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARGAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 100)
[0338] ecTadA (D108V)
[0339] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARVAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 101)
[0340] ecTadA (H8Y, D108N, N127S)
[0341] SEVEFSYEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARNAKTGAAGSLMDVLHHPGMSHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 102)
[0342] ecTadA (H8Y, D108N, N127S, E155D)
[0343] SEVEFSYEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARNAKTGAAGSLMDVLHHPGMSHRVEITEGILADECAALLSDFFRMRRQDIKAQKK AQSSTD (SEQ ID NO: 103)
[0344] ecTadA (H8Y, D108N, N127S, E155G)
[0345] SEVEFSYEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARNAKTGAAGSLMDVLHHPGMSHRVEITEGILADECAALLSDFFRMRRQGIKAQKK AQSSTD (SEQ ID NO: 104)
[0346] ecTadA (H8Y, D108N, N127S, E155V)
[0347] SEVEFSYEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARNAKTGAAGSLMDVLHHPGMSHRVEITEGILADECAALLSDFFRMRRQVIKAQKK AQSSTD (SEQ ID NO: 105)
[0348] ecTadA (A106V, D108N, D147Y, and E155V)
[0349] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSYFFRMRRQVIKAQKK AQSSTD (SEQ ID NO: 106)
[0350] ecTadA (S2A, I49F, A106V, D108N, D147Y, E155V)
[0351] AEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPF GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSYFFRMRRQVIKAQKK AQSSTD (SEQ ID NO: 107)
[0352] ecTadA (H8Y, A106T, D108N, N127S, K160S)
[0353] SEVEFSYEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG TRNAKTGAAGSLMDVLHHPGMSHRVEITEGILADECAALLSDFFRMRRQEIKAQSK AQSSTD (SEQ ID NO: 108)
[0354] ecTadA (R26G, L84F, A106V, R107H, D108N, H123Y, A142N, A143D, D147Y, E155V, I156F)
[0355] SEVEFSHEYWMRHALTLAKRAWDEGEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VHNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNDLLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 109)
[0356] ecTadA (E25G, R26G, L84F, A106V, R107H, D108N, H123Y, A142N, A143D, D147Y, E155V, I156F)
[0357] SEVEFSHEYWMRHALTLAKRAWDGGEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VHNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNDLLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 110)
[0358] ecTadA (E25D, R26G, L84F, A106V, R107K, D108N, H123Y, A142N, A143G, D147Y, E155V, I156F
[0359] SEVEFSHEYWMRHALTLAKRAWDDGEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VKNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNGLLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 111)
[0360] ecTadA (R26Q, L84F, A106V, D108N, H123Y, A142N, D147Y, E155V, I156F
[0361] SEVEFSHEYWMRHALTLAKRAWDEQEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNALLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 112)
[0362] ecTadA (E25M, R26G, L84F, A106V, R107P, D108N, H123Y, A142N, A143D, D147Y, E155V, I156F
[0363] SEVEFSHEYWMRHALTLAKRAWDMGEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VPNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNDLLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 113)
[0364] ecTadA (R26C, L84F, A106V, R107H, D108N, H123Y, A142N , D147Y, E155V, I156F)
[0365] SEVEFSHEYWMRHALTLAKRAWDECEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VHNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNALLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 114)
[0366] ecTadA (L84F, A106V , D108N, H123Y, A142N, A143L, D147Y, E155V, I156F)
[0367] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNLLLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 115)
[0368] ecTadA (R26G, L84F, A106V, D108N, H123Y, A142N , D147Y, E155V, I156F)
[0369] SEVEFSHEYWMRHALTLAKRAWDEGEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNALLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 116)
[0370] ecTadA (R51H, L84F, A106V, D108N, H123Y, D147Y, E155V, I156F, K157N)
[0371] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GHHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFNAQK KAQSSTD (SEQ ID NO: 117)
[0372] ecTadA (E25A, R26G, L84F, A106V, R107N, D108N, H123Y, A142N, A143E, D147Y, E155V, I156F)
[0373] SEVEFSHEYWMRHALTLAKRAWDAGEVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVNNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECNELLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 118)
[0374] ecTadA (N37T, P48T, L84F, A106V, D108N, H123Y, D147Y, E155V, I156F)
[0375] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHTNRVIGEGWNRTI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 119)
[0376] ecTadA (N37S, L84F, A106V, D108N, H123Y, D147Y, E155V, I156F)
[0377] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHSNRVIGEGWNRPIG RHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFKAQKK AQSSTD (SEQ ID NO: 120)
[0378] ecTadA (H36L, L84F, A106V, D108N, H123Y, D147Y, E155V, I156F)
[0379] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVLNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 121)
[0380] ecTadA (H36L, P48L, L84F, A106V, D108N, H123Y, D147Y, E155V, I156F)
[0381] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVLNNRVIGEGWNRLI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 122)
[0382] ecTadA (H36L, L84F, A106V, D108N, H123Y, D147Y, E155V, K57N, I156F)
[0383] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVLNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFNAQK KAQSSTD (SEQ ID NO: 123 )
[0384] ecTadA (H36L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F)
[0385] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVLNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 124)
[0386] ecTadA (L84F, A106V, D108N, H123Y, S146R, D147Y, E155V, I156F)
[0387] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLRYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 125)
[0388] ecTadA (N37S, R51H, L84F, A106V, D108N, H123Y, D147Y, E155V, I156F
[0389] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHSNRVIGEGWNRPIG HHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFKAQKK AQSSTD (SEQ ID NO: 126)
[0390] ecTadA (R51L, L84F, A106V, D108N, H123Y, D147Y, E155V, I156F, K157N
[0391] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFNAQK KAQSSTD (SEQ ID NO: 127)
[0392] saTadA (D108N)
[0393] GSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQ QPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADNP KGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN (SEQ ID NO: 128)
[0394] saTadA (D107A_D108N)
[0395] GSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQ QPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGAANP KGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN (SEQ ID NO: 129)
[0396] saTadA (G26P_D107A_D108N)
[0397] GSHMTNDIYFMTLAIEEAKKAAQLPEVPIGAIITKDDEVIARAHNLRETLQ QPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGAANP KGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN (SEQ ID NO: 130)
[0398] saTadA (G26P_D107A_D108N_S142A)
[0399] GSHMTNDIYFMTLAIEEAKKAAQLPEVPIGAIITKDDEVIARAHNLRETLQ QPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGAANP KGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACATLLTTFFKNLRANKKSTN (SEQ ID NO: 131)
[0400] saTadA (D107A_D108N_S142A)
[0401] GSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQ QPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGAANP KGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACATLLTTFFKNLRANKKSTN (SEQ ID NO: 132)
[0402] ecTadA (P48S)
[0403] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRSI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 133)
[0404] ecTadA (P48T)
[0405] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRTI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 134)
[0406] ecTadA (P48A)
[0407] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRAI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 135)
[0408] ecTadA (A142N)
[0409] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECNALLSDFFRMRRQEIKAQKK AQSSTD (SEQ ID NO: 136)
[0410] ecTadA (W23R)
[0411] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVHNNRVIGEGWNRPIG RHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGA RDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKA QSSTD (SEQ ID NO: 137)
[0412] ecTadA (W23L)
[0413] SEVEFSHEYWMRHALTLAKRALDEREVPVGAVLVHNNRVIGEGWNRPIG RHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGA RDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKA QSSTD (SEQ ID NO: 138)
[0414] ecTadA (R152P)
[0415] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMPRQEIKAQKK AQSSTD (SEQ ID NO: 139)
[0416] ecTadA (R152H)
[0417] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMHRQEIKAQKK AQSSTD (SEQ ID NO: 140)
[0418] ecTadA (L84F, A106V, D108N, H123Y, D147Y, E155V, I156F)
[0419] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLSYFFRMRRQVFKAQK KAQSSTD (SEQ ID NO: 141)
[0420] ecTadA (H36L, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F, K157N)
[0421] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVLNNRVIGEGWNRPI GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMRRQVFNAQK KAQSSTD (SEQ ID NO: 142)
[0422] ecTadA (H36L, P48S, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F , K157N)
[0423] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVLNNRVIGEGWNRSI GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMRRQVFNAQK KAQSSTD (SEQ ID NO: 143)
[0424] ecTadA (H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F , K157N)
[0425] SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVLNNRVIGEGWNRAI GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMRRQVFNAQK KAQSSTD (SEQ ID NO: 144)
[0426] ecTadA (W23L, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, R152P, E155V, I156F, K157N)
[0427] SEVEFSHEYWMRHALTLAKRALDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKK AQSSTD (SEQ ID NO: 145)
[0428] ecTadA (W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, R152P, E155V, I156F, K157N)
[0429] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKK AQSSTD (SEQ ID NO: 146)
[0430] Staphylococcus aureus TadA:
[0431] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETL QQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGAD DPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN (SEQ ID NO: 147)
[0432] Bacillus subtilis TadA:
[0433] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSI AHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPK GGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE (SEQ ID NO: 148)
[0434] Salmonella typhimurium (S. typhimurium) TadA:
[0435] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHN HRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGA MVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFR MRRQEIKALKKADRAEGAGPAV (SEQ ID NO: 149)
[0436] Shewanella putrefaciens (S. putrefaciens) TadA:
[0437] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHD PTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDE KTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQR AQQGIE (SEQ ID NO: 150)
[0438] Haemophilus influenzae F3031 (H. influenzae) TadA:
[0439] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEG WNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEK ALLKSLSDK (SEQ ID NO: 151)
[0440] Caulobacter crescentus (C. crescentus) TadA:
[0441] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAG NGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGR VVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI (SEQ ID NO: 152)
[0442] Geobacter sulfurreducens (G. sulfurreducens) TadA:
[0443] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGH NLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERV VFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKA KATPALFIDERKVPPEP (SEQ ID NO: 153)
[0444] Streptococcus pyogenes (S. pyogenes) TadA
[0445] MPYSLEEQTYFMQEALKEAEKSLQKAEIPIGCVIVKDGEIIGRGHNAREES NQAIMHAEIMAINEANAHEGNWRLLDTTLFVTIEPCVMCSGAIGLARIPHVIYGASN QKFGGADSLYQILTDERLNHRVQVERGLLAADCANIMQTFFRQGRERKKIAKHLIKE QSDPFD (SEQ ID NO: 154
[0446] TadA 7.10:
[0447] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKK AQSSTD (SEQ ID NO: 155)
[0448] TadA 7.10 (V106W) (E. coli)
[0449] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGW RNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKK AQSSTD (SEQ ID NO: 156)
[0450] TadA-8e (E. coli)
[0451] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKK AQSSIN (SEQ ID NO: 157)
[0452] TadA-8e(V106W) (E. coli)
[0453] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGW RNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKK AQSSIN (SEQ ID NO: 158)
[0454] Aquifex aeolicus (A. aeolicus) TadA
[0455] MGKEYFLKVALREAKRAFEKGEVPVGAIIVKEGEIISKAHNSVEELKDPTA HAEMLAIKEACRRLNTKYLEGCELYVTLEPCIMCSYALVLSRIEKVIFSALDKKHGG VVSVFNILDEPTLNHRVKWEYYPLEEASELLSEFFKKLRNNII (SEQ ID NO: 159)
[0456] Tad1
[0457] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIG LYDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKK AQSSIN (SEQ ID NO: 160)
[0458] Tad2
[0459] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKK AQSSIN (SEQ ID NO: 161)
[0460] Tad3
[0461] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYGLIDATLYVTFEPCVMCAGAIIHSRIGRVVFGVR NSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKA QSSIN (SEQ ID NO: 162)
[0462] Tad4
[0463] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIG LHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKK AQSSIN (SEQ ID NO: 163)
[0464] Tad6
[0465] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIG LYDPTAHAEIMALRQGGLVMQNYGLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGV RNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRQVFNAQKK AQSSIN (SEQ ID NO: 164)
[0466] Tad6-SR
[0467] SEVEFSHEYWMRHALTLAKRARDEGEVPVGAVLVLNNRVIGEGWNRAIG LYDPTAHAEIMALRQGGLVMQNYGLIDATLYSTFEPCVMCAGAMIHSRIGRVVFGV RNSKRGAAGSLMNVLNYPGMDHRVEITEGILADECAALLCDFYRMPRRVFNAQKK AQSSIN (SEQ ID NO: 165)
[0468] ABE8e MKRTADGSEFESPKKKRKVSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLN NRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGA MIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFY RMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGSDKKYSIGL AIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKR TARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDV DKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFG NLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNL SDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQS KNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPH QIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKV KYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVE DRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKP ENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLY YLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNV PSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQ ITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFY SNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKK TEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGK SKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDE IIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYF DTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSKRTADGSEFEPKK KRKV (SEQ ID NO: 166)
[0469] In some aspects, the fusion proteins of the present disclosure comprise cytidine base editors (CBEs) comprising a napDNAbp domain (e.g., any of the Cas14a1 variants provided herein) and a cytosine deaminase domain that enzymatically deaminates a cytosine nucleobase of a C:G nucleobase pair to a uracil. The uracil may be subsequently converted to a thymine (T) by the cell’s DNA repair and replication machinery. The mismatched guanine (G) on the opposite strand may subsequently be converted to an adenine (A) by the cell’s DNA repair and replication machinery. In this manner, a target C:G nucleobase pair is ultimately converted to a T:A nucleobase pair. Other cytosine deaminase domains besides those provided herein are known in the art, and a person of ordinary skill in the art would recognize which cytosine deaminase domains could be used in the fusion proteins of the present disclosure.
[0470] The CBE fusion proteins described herein may further comprise one or more nuclear localization signals (NLSs) and / or one or more uracil glycosylase inhibitor (UGI) domains. Thus, the base editor fusion proteins may comprise the structure: NH2-[first nuclear localization sequence]-[cytosine deaminase domain]-[napDNAbp domain]-[first UGI domain]-[second UGI domain]-[second nuclear localization sequence]-COOH, wherein each instance of “]-[” indicates the presence of an optional linker sequence. The CBE fusion proteins of the present disclosure may comprise modified (or evolved) cytosine deaminase domains, such as deaminase domains that recognize an expanded PAM sequence, haveimproved efficiency of deaminating 5′-GC targets, and / or make edits in a narrower target window.
[0471] In some aspects, the fusion proteins of the disclosure comprise an adenine base editor. Some aspects of the disclosure provide fusion proteins that comprise a nucleic acid programmable DNA binding protein (napDNAbp), and at least two adenosine deaminase domains. Without wishing to be bound by any particular theory, dimerization of adenosine deaminases (e.g., in cis or in trans) may improve the ability (e.g., efficiency) of the fusion protein to modify a nucleic acid base (for example, to deaminate adenine). In some embodiments, any of the fusion proteins may comprise 2, 3, 4, or 5 adenosine deaminase domains. In some embodiments, any of the fusion proteins provided herein comprises two adenosine deaminases. In some embodiments, any of the fusion proteins provided herein contain only two adenosine deaminases. In some embodiments, the adenosine deaminases are the same. In some embodiments, the adenosine deaminases are any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminases are different. Other adenosine deaminase domains besides those provided herein are known in the art, and a person of ordinary skill in the art would recognize which adenosine deaminase domains could be used in the fusion proteins of the present disclosure.
[0472] In some embodiments, the general architecture of exemplary fusion proteins with a first adenosine deaminase, a second adenosine deaminase, and a napDNAbp (e.g., any of the Cas14a1 variants provided herein) comprises any one of the following structures, where NLS is a nuclear localization sequence (e.g., any NLS provided herein), NH2 is the N- terminus of the fusion protein, and COOH is the C-terminus of the fusion protein: NH2-[first adenosine deaminase]-[second adenosine deaminase]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[napDNAbp]-[second adenosine deaminase]-COOH; NH2- [napDNAbp]-[first adenosine deaminase]-[second adenosine deaminase]-COOH; NH2- [second adenosine deaminase]-[first adenosine deaminase]-[napDNAbp]-COOH; NH2- [second adenosine deaminase]-[napDNAbp]-[first adenosine deaminase]-COOH; NH2- [napDNAbp]-[second adenosine deaminase]-[first adenosine deaminase]-COOH.
[0473] In some embodiments, the fusion proteins provided herein do not comprise a linker. In some embodiments, a linker is present between one or more of the domains or proteins (e.g., first adenosine deaminase, second adenosine deaminase, and / or napDNAbp). In some embodiments, the “]-[” used in the general architecture above indicates the presence of an optional linker. Exemplary fusion proteins comprising a first adenosine deaminase, a second adenosine deaminase, a napDNAbp, and an NLS are provided: NH2-[NLS]-[firstadenosine deaminase]-[second adenosine deaminase]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[NLS]-[second adenosine deaminase]-[napDNAbp]-COOH; NH2- [first adenosine deaminase]-[second adenosine deaminase]-[NLS]-[napDNAbp]-COOH; NH2-[first adenosine deaminase]-[second adenosine deaminase]-[napDNAbp]-[NLS]- COOH; NH2-[NLS]-[first adenosine deaminase]-[napDNAbp]-[second adenosine deaminase]-COOH; NH2-[first adenosine deaminase]-[NLS]-[napDNAbp]-[second adenosine deaminase]-COOH; NH2-[first adenosine deaminase]-[napDNAbp]-[NLS]- [second adenosine deaminase]-COOH; NH2-[first adenosine deaminase]-[napDNAbp]- [second adenosine deaminase]-[NLS]-COOH; NH2-[NLS]-[napDNAbp]-[first adenosine deaminase]-[second adenosine deaminase]-COOH; NH2-[napDNAbp]-[NLS]-[first adenosine deaminase]-[second adenosine deaminase]-COOH; NH2-[napDNAbp]-[first adenosine deaminase]-[NLS]-[second adenosine deaminase]-COOH; NH2-[napDNAbp]- [first adenosine deaminase]-[second adenosine deaminase]-[NLS]-COOH; NH2-[NLS]- [second adenosine deaminase]-[first adenosine deaminase]-[napDNAbp]-COOH; NH2- [second adenosine deaminase]-[NLS]-[first adenosine deaminase]-[napDNAbp]-COOH; NH2-[second adenosine deaminase]-[first adenosine deaminase]-[NLS]-[napDNAbp]- COOH; NH2-[second adenosine deaminase]-[first adenosine deaminase]-[napDNAbp]- [NLS]-COOH; NH2-[NLS]-[second adenosine deaminase]-[napDNAbp]-[first adenosine deaminase]-COOH; NH2-[second adenosine deaminase]-[NLS]-[napDNAbp]-[first adenosine deaminase]-COOH; NH2-[second adenosine deaminase]-[napDNAbp]-[NLS]- [first adenosine deaminase]-COOH; NH2-[second adenosine deaminase]-[napDNAbp]-[first adenosine deaminase]-[NLS]-COOH; NH2-[NLS]-[napDNAbp]-[second adenosine deaminase]-[first adenosine deaminase]-COOH; NH2-[napDNAbp]-[NLS]-[second adenosine deaminase]-[first adenosine deaminase]-COOH; NH2-[napDNAbp]-[second adenosine deaminase]-[NLS]-[first adenosine deaminase]-COOH; NH2-[napDNAbp]- [second adenosine deaminase]-[first adenosine deaminase]-[NLS]-COOH. Nucleobase editor Sequences
[0474] Sequence 1: Dual-AAV BE3.9max PRNP R37X sgRNA (N-terminus) ITR-Cbh promoter-NLS-APOBEC deaminase domain-linker-SpCas9 (amino acids 1-572)- NpuN-NLS-WPRE-bovine growth hormone(bGH)-derived poly(A)-PRNP R37X sgRNA (reverse complement)-human U6 promoter (reverse complement)-ITR CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCG ACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCC AACTCCATCACTAGGGGTTCCTGCGGCCTCTAGATCAGGGTACCCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTC AATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGG TAAACTGCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTA TTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTGTGCCCAGTACATGACCTT ATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATG GTCGAGGTGAGCCCCACGTTCTGCTTCACTCTCCCCATCTCCCCCCCCTCCCCACC CCCAATTTTGTATTTATTTATTTTTTAATTATTTTGTGCAGCGATGGGGGCGGGGG GGGGGGGGGGGCGCGCGCCAGGCGGGGCGGGGCGGGGCGAGGGGCGGGGCGG GGCGAGGCGGAGAGGTGCGGCGGCAGCCAATCAGAGCGGCGCGCTCCGAAAGT TTCCTTTTATGGCGAGGCGGCGGCGGCGGCGGCCCTATAAAAAGCGAAGCGCGC GGCGGGCGGGAGTCGCTGCGACGCTGCCTTCGCCCCGTGCCCCGCTCCGCCGCCG CCTCGCGCCGCCCGCCCCGGCTCTGACTGACCGCGTTACTCCCACAGGTGAGCGG GCGGGACGGCCCTTCTCCTCCGGGCTGTAATTAGCTGAGCAAGAGGTAAGGGTTT AAGGGATGGTTGGTTGGTGGGGTATTAATGTTTAATTACCTGGAGCACCTGCCTG AAATCACTTTTTTTCAGGTTGGACCGGTGCCACCATGAAACGGACAGCCGACGG AAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCCTCAGAGACTGGGC CTGTCGCCGTCGATCCAACCCTGCGCCGCCGGATTGAACCTCACGAGTTTG AAGTGTTCTTTGACCCCCGGGAGCTGAGAAAGGAGACATGCCTGCTGTACG AGATCAACTGGGGAGGCAGGCACTCCATCTGGAGGCACACCTCTCAGAACA CAAATAAGCACGTGGAGGTGAACTTCATCGAGAAGTTTACCACAGAGCGGT ACTTCTGCCCCAATACCAGATGTAGCATCACATGGTTTCTGAGCTGGTCCCC TTGCGGAGAGTGTAGCAGGGCCATCACCGAGTTCCTGTCCAGATATCCACA CGTGACACTGTTTATCTACATCGCCAGGCTGTATCACCACGCAGACCCAAGG AATAGGCAGGGCCTGCGCGATCTGATCAGCTCCGGCGTGACCATCCAGATC ATGACAGAGCAGGAGTCCGGCTACTGCTGGCGGAACTTCGTGAATTATTCT CCTAGCAACGAGGCCCACTGGCCTAGGTACCCACACCTGTGGGTGCGCCTG TACGTGCTGGAGCTGTATTGCATCATCCTGGGCCTGCCCCCTTGTCTGAATA TCCTGCGGAGAAAGCAGCCCCAGCTGACCTTCTTTACAATCGCCCTGCAGTC TTGTCACTATCAGAGGCTGCCACCCCACATCCTGTGGGCCACAGGCCTGAA GTCTGGAGGATCTAGCGGAGGATCCTCTGGCAGCGAGACACCAGGAACAAGCGA GTCAGCAACACCAGAGAGCAGTGGCGGCAGCAGCGGCGGCAGCGACAAGAAGTA CAGCATCGGCCTGGCCATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGA GTACAAGGTGCCCAGCAAGAAATTCAAGGTGCTGGGCAACACCGACCGGCACAGCAT CAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGAAACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGCTA TCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGA CTGGAAGAGTCCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTC GGCAACATCGTGGACGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTG AGAAAGAAACTGGTGGACAGCACCGACAAGGCCGACCTGCGGCTGATCTATCTGGCC CTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACCTGAACCCC GACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGT TCGAGGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCA GACTGAGCAAGAGCAGACGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAAGA AGAATGGCCTGTTCGGAAACCTGATTGCCCTGAGCCTGGGCCTGACCCCCAACTTCAA GAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGA CGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTTCT GGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACAC CGAGATCACCAAGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCA CCAGGACCTGACCCTGCTGAAAGCTCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAA AGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCCGGCTACATTGACGGCGGAGC CAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGGACGGCACC GAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTC GACAACGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGG CGGCAGGAAGATTTTTACCCATTCCTGAAGGACAACCGGGAAAAGATCGAGAAGATCC TGACCTTCCGCATCCCCTACTACGTGGGCCCTCTGGCCAGGGGAAACAGCAGATTCG CCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAACTTCGAGGAAGTGG TGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAGAA CCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGT GTATAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTC CTGAGCGGCGAGCAGAAAAAGGCCATCGTGGACCTGCTGTTCAAGACCAACCGGAAA GTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAGAAAATCGAGTGCCTGTCCTAC GAGACAGAGATCCTGACAGTGGAGTATGGCCTGCTGCCAATCGGCAAGATC GTGGAGAAGAGGATCGAGTGTACCGTGTACTCTGTGGATAACAATGGCAAC ATCTATACACAGCCCGTGGCACAGTGGCACGATAGGGGAGAGCAGGAGGTG TTCGAGTATTGCCTGGAGGACGGCAGCCTGATCAGGGCAACCAAGGACCAC AAGTTCATGACAGTGGATGGCCAGATGCTGCCCATCGACGAGATTTTCGAG CGGGAGCTGGACCTGATGAGAGTGGATAACCTGCCTAATAGCGGAGGCAGT AAAAGAACAGCAGACGGGAGTGAGTTTGAGCCCAAGAAAAAGAGAAAGGTGTAAGATCTGATAATCAACCTCTGGATTACAAAATTTGTGAAAGATTGACTGGTATTC TTAACTATGTTGCTCCTTTTACGCTATGTGGATACGCTGCTTTAATGCCTTTGTAT CATGCTATTGCTTCCCGTATGGCTTTCATTTTCTCCTCCTTGTATAAATCCTGGTTA GTTCTTGCCACGGCGGAACTCATCGCCGCCTGCCTTGCCCGCTGCTGGACAGGGG CTCGGCTGTTGGGCACTGACAATTCCGTGGTGCGACTGTGCCTTCTAGTTGCCAG CCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCC CACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTC ATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAG ACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTCGAGAAAAAAAGCAC CGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTC TAGCTCTAAAACTGCCCCGGGTATCGGCTGCCGGTGTTTCGTCCTTTCCACAAGA TATATAAAGCCAAGAAATCGAAATACTTTCAAGTTACGGTAAGCATATGATA GTCCATTTTAAAACATAATTTTAAAACTGCAAACTACCCAAGAAATTATTACT TTCTACGTCACGTATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAATT CTAATTATCTCTCTAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCA TGGGAAATAGGCCCTCTTCCTGCCCGACCTTGCGGCCGCAGGAACCCCTAGTG ATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGAC CAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAG CGCGCAG (SEQ ID NO: 167)
[0475] Sequence 2: Dual-AAV BE3.9max PRNP R37X sgRNA (C-terminus) ITR-Cbh promoter-NLS-NpuC-SpCas9 (amino acids 573-1367)-UGI-NLS-WPRE-bovine growth hormone(bGH)-derived poly(A)-PRNP R37X sgRNA (reverse complement)-human U6 promoter (reverse complement)-ITR CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCG ACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCC AACTCCATCACTAGGGGTTCCTGCGGCCTCTAGATCAGGGTACCCGTTACATAAC TTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTC AATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGG TAAACTGCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTA TTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTGTGCCCAGTACATGACCTT ATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATG GTCGAGGTGAGCCCCACGTTCTGCTTCACTCTCCCCATCTCCCCCCCCTCCCCACC CCCAATTTTGTATTTATTTATTTTTTAATTATTTTGTGCAGCGATGGGGGCGGGGG GGGGGGGGGGGCGCGCGCCAGGCGGGGCGGGGCGGGGCGAGGGGCGGGGCGG GGCGAGGCGGAGAGGTGCGGCGGCAGCCAATCAGAGCGGCGCGCTCCGAAAGT TTCCTTTTATGGCGAGGCGGCGGCGGCGGCGGCCCTATAAAAAGCGAAGCGCGC GGCGGGCGGGAGTCGCTGCGACGCTGCCTTCGCCCCGTGCCCCGCTCCGCCGCCG CCTCGCGCCGCCCGCCCCGGCTCTGACTGACCGCGTTACTCCCACAGGTGAGCGGGCGGGACGGCCCTTCTCCTCCGGGCTGTAATTAGCTGAGCAAGAGGTAAGGGTTT AAGGGATGGTTGGTTGGTGGGGTATTAATGTTTAATTACCTGGAGCACCTGCCTG AAATCACTTTTTTTCAGGTTGGACCGGTGCCACCATGAAACGGACAGCCGACGG AAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCATCAAGATTGCTACAC GGAAATACCTGGGAAAGCAGAACGTGTACGACATCGGCGTGGAGCGGGATC ACAACTTCGCCCTGAAGAATGGCTTTATCGCCAGCAATTGCTTCGACTCCGTG GAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACATACCACGATCTGC TGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTG AAAACCTATGCCCACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGAT ACACCGGCTGGGGCAGGCTGAGCCGGAAGCTGATCAACGGCATCCGGGACAAGCAG TCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCA TGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGT GTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCG CCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGA TGGGCCGGCACAAGCCCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCA CCCAGAAGGGACAGAAGAACAGCCGCGAGAGAATGAAGCGGATCGAAGAGGGCATCA AAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACACCCAGCTGCAGA ACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGG AACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTT TCTGAAGGACGACTCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGG CAAGAGCGACAACGTGCCCTCCGAAGAGGTCGTGAAGAAGATGAAGAACTACTGGCG GCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTTCGACAATCTGACCAAGGCC GAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAGCTGGTG GAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACT AAGTACGACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCA AGCTGGTGTCCGATTTCCGGAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAA CTACCACCACGCCCACGACGCCTACCTGAACGCCGTCGTGGGAACCGCCCTGATCAA AAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTG CGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTTC TTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGAT CCGGAAGCGGCCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAA GGGCCGGGATTTTGCCACCGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGT GAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCCAAGAG GAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAGTACGGCGG CTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGG CAAGTCCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGA AGCAGCTTCGAGAAGAATCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGA AAAAGGACCTGATCATCAAGCTGCCTAAGTACTCCCTGTTCGAGCTGGAAAACGGCCG GAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAAACGAACTGGCCCTGCC CTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGGCTCC CCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACG AGATCATCGAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCT GGACAAAGTGCTGTCCGCCTACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGC CGAGAATATCATCCACCTGTTTACCCTGACCAATCTGGGAGCCCCTGCCGCCTTCAAG TACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAAGAGGTGCTGGAC GCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCT CAGCTGGGAGGTGACAGCGGCGGGAGCGGCGGGAGCGGGGGGAGCACTAAT CTGAGCGACATCATTGAGAAGGAGACTGGGAAACAGCTGGTCATTCAGGAG TCCATCCTGATGCTGCCTGAGGAGGTGGAGGAAGTGATCGGCAACAAGCCAGAGTCTGACATCCTGGTGCACACCGCCTACGACGAGTCCACAGATGAGAAT GTGATGCTGCTGACCTCTGACGCCCCCGAGTATAAGCCTTGGGCCCTGGTC ATCCAGGATTCTAACGGCGAGAATAAGATCAAGATGCTGAGCGGAGGATCC AAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTA AGATCTGATAATCAACCTCTGGATTACAAAATTTGTGAAAGATTGACTGGTATTC TTAACTATGTTGCTCCTTTTACGCTATGTGGATACGCTGCTTTAATGCCTTTGTAT CATGCTATTGCTTCCCGTATGGCTTTCATTTTCTCCTCCTTGTATAAATCCTGGTTA GTTCTTGCCACGGCGGAACTCATCGCCGCCTGCCTTGCCCGCTGCTGGACAGGGG CTCGGCTGTTGGGCACTGACAATTCCGTGGTGCGACTGTGCCTTCTAGTTGCCAG CCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCC CACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTC ATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAG ACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTCGAGAAAAAAAGCAC CGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTC TAGCTCTAAAACTGCCCCGGGTATCGGCTGCCGGTGTTTCGTCCTTTCCACAAGA TATATAAAGCCAAGAAATCGAAATACTTTCAAGTTACGGTAAGCATATGATA GTCCATTTTAAAACATAATTTTAAAACTGCAAACTACCCAAGAAATTATTACT TTCTACGTCACGTATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAATT CTAATTATCTCTCTAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCA TGGGAAATAGGCCCTCTTCCTGCCCGACCTTGCGGCCGCAGGAACCCCTAGTG ATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGAC CAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAG CGCGCAG (SEQ ID NO: 168)
[0476] Sequence 3: Dual-AAV TadCBEd PRNP R37X sgRNA (N-terminus) ITR-Cbh promoter-NLS-TadCBEd deaminase domain-linker-SpCas9 (amino acids 1-572)- NpuN-NLS-WPRE-bovine growth hormone(bGH)-derived poly(A)-PRNP R37X sgRNA (reverse complement)-human U6 promoter (reverse complement)-ITR CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCG ACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCC AACTCCATCACTAGGGGTTCCTGCGGCCTCTAGATCAGGGTACCCGTTACATAAC TTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTC AATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGG TAAACTGCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTA TTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTGTGCCCAGTACATGACCTT ATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATG GTCGAGGTGAGCCCCACGTTCTGCTTCACTCTCCCCATCTCCCCCCCCTCCCCACC CCCAATTTTGTATTTATTTATTTTTTAATTATTTTGTGCAGCGATGGGGGCGGGGG GGGGGGGGGGGCGCGCGCCAGGCGGGGCGGGGCGGGGCGAGGGGCGGGGCGG GGCGAGGCGGAGAGGTGCGGCGGCAGCCAATCAGAGCGGCGCGCTCCGAAAGT TTCCTTTTATGGCGAGGCGGCGGCGGCGGCGGCCCTATAAAAAGCGAAGCGCGC GGCGGGCGGGAGTCGCTGCGACGCTGCCTTCGCCCCGTGCCCCGCTCCGCCGCCG CCTCGCGCCGCCCGCCCCGGCTCTGACTGACCGCGTTACTCCCACAGGTGAGCGG GCGGGACGGCCCTTCTCCTCCGGGCTGTAATTAGCTGAGCAAGAGGTAAGGGTTT AAGGGATGGTTGGTTGGTGGGGTATTAATGTTTAATTACCTGGAGCACCTGCCTG AAATCACTTTTTTTCAGGTTGGACCGGTGCCACCATGAAACGGACAGCCGACGG AAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCAGTTCTGAGGTGGAGT TTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGGCACGGGATGAGAGGAAGGCGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGA GTGATCGGCGAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCC CATGCCGAAATTATAGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTAC AGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGC GCCGGCGCCATGATCAACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGG AACTCAAAAAGAGGCGCCGCAGGCTCCCTGATGAACGTGCTGAACTACCCC GGAATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGT GCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTC AGAAGAAGGCCCAGAGCTCCATCAACTCTGGAGGATCTAGCGGAGGATCCTCT GGCAGCGAGACACCAGGAACAAGCGAGTCAGCAACACCAGAGAGCAGTGGCGG CAGCAGCGGCGGCAGCGACAAGAAGTACAGCATCGGCCTGGCCATCGGCACCAACT CTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGT TCGACAGCGGCGAAACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGA TACACCAGACGGAAGAACCGGATCTGCTATCTGCAAGAGATCTTCAGCAACGAGATGG CCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGG ATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGCCTACC ACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAA GGCCGACCTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCA CTTCCTGATCGAGGGCGACCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCAT CCAGCTGGTGCAGACCTACAACCAGCTGTTCGAGGAAAACCCCATCAACGCCAGCGG CGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGACGGCTGGAAAA TCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGC CCTGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGC CAAACTGCAGCTGAGCAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCA GATCGGCGACCAGTACGCCGACCTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCAT CCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCAAGGCCCCCCTGAGCGC CTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGCTCTC GTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACG GCTACGCCGGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAA GCCCATCCTGGAAAAGATGGACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGA GGACCTGCTGCGGAAGCAGCGGACCTTCGACAACGGCAGCATCCCCCACCAGATCCA CCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTACCCATTCCTGAAG GACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCA TCACCCCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCA TCGAGCGGATGACCAACTTCGATAAGAACCTGCCCAACGAGAAGGTGCTGCCCAAGC ACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGCTGACCAAAGTGAAATACGT GACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGCCATCGT GGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTA CTTCAAGAAAATCGAGTGCCTGTCCTACGAGACAGAGATCCTGACAGTGGAGT ATGGCCTGCTGCCAATCGGCAAGATCGTGGAGAAGAGGATCGAGTGTACCG TGTACTCTGTGGATAACAATGGCAACATCTATACACAGCCCGTGGCACAGTG GCACGATAGGGGAGAGCAGGAGGTGTTCGAGTATTGCCTGGAGGACGGCA GCCTGATCAGGGCAACCAAGGACCACAAGTTCATGACAGTGGATGGCCAGA TGCTGCCCATCGACGAGATTTTCGAGCGGGAGCTGGACCTGATGAGAGTGG ATAACCTGCCTAATAGCGGAGGCAGTAAAAGAACAGCAGACGGGAGTGAGTTT GAGCCCAAGAAAAAGAGAAAGGTGTAAGATCTGATAATCAACCTCTGGATTACA AAATTTGTGAAAGATTGACTGGTATTCTTAACTATGTTGCTCCTTTTACGCTATGT GGATACGCTGCTTTAATGCCTTTGTATCATGCTATTGCTTCCCGTATGGCTTTCATTTTCTCCTCCTTGTATAAATCCTGGTTAGTTCTTGCCACGGCGGAACTCATCGCCG CCTGCCTTGCCCGCTGCTGGACAGGGGCTCGGCTGTTGGGCACTGACAATTCCGT GGTGCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCC TTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAA TTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCA GGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGT GGGCTCTATGGCTCGAGAAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAA CGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAACTGCCCCGGGTATCGGCT GCCGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAATCGAAATAC TTTCAAGTTACGGTAAGCATATGATAGTCCATTTTAAAACATAATTTTAAAAC TGCAAACTACCCAAGAAATTATTACTTTCTACGTCACGTATTTTGTACTAATA TCTTTGTGTTTACAGTCAAATTAATTCTAATTATCTCTCTAACAGCCTTGTAT CGTATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCTTCCTGCCCGA CCTTGCGGCCGCAGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGC TCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGC CCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG (SEQ ID NO: 169)
[0477] Sequence 5: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (N-terminus) ITR-Cbh promoter-NLS-TadCBEd deaminase domain-linker-SpCas9 (amino acids 1-572)- NpuN-NLS-WPRE-bovine growth hormone(bGH)-derived poly(A)-PRNP R37X F+E- sgRNA (reverse complement)-human U6 promoter (reverse complement)-ITR CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCG ACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCC AACTCCATCACTAGGGGTTCCTGCGGCCTCTAGATCAGGGTACCCGTTACATAAC TTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTC AATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGG TAAACTGCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTA TTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTGTGCCCAGTACATGACCTT ATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATG GTCGAGGTGAGCCCCACGTTCTGCTTCACTCTCCCCATCTCCCCCCCCTCCCCACC CCCAATTTTGTATTTATTTATTTTTTAATTATTTTGTGCAGCGATGGGGGCGGGGG GGGGGGGGGGGCGCGCGCCAGGCGGGGCGGGGCGGGGCGAGGGGCGGGGCGG GGCGAGGCGGAGAGGTGCGGCGGCAGCCAATCAGAGCGGCGCGCTCCGAAAGT TTCCTTTTATGGCGAGGCGGCGGCGGCGGCGGCCCTATAAAAAGCGAAGCGCGC GGCGGGCGGGAGTCGCTGCGACGCTGCCTTCGCCCCGTGCCCCGCTCCGCCGCCG CCTCGCGCCGCCCGCCCCGGCTCTGACTGACCGCGTTACTCCCACAGGTGAGCGG GCGGGACGGCCCTTCTCCTCCGGGCTGTAATTAGCTGAGCAAGAGGTAAGGGTTT AAGGGATGGTTGGTTGGTGGGGTATTAATGTTTAATTACCTGGAGCACCTGCCTG AAATCACTTTTTTTCAGGTTGGACCGGTGCCACCATGAAACGGACAGCCGACGG AAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCAGTTCTGAGGTGGAGT TTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGGCAC GGGATGAGAGGAAGGCGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGA GTGATCGGCGAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCC CATGCCGAAATTATAGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTAC AGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGC GCCGGCGCCATGATCAACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGG AACTCAAAAAGAGGCGCCGCAGGCTCCCTGATGAACGTGCTGAACTACCCC GGAATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTC AGAAGAAGGCCCAGAGCTCCATCAACTCTGGAGGATCTAGCGGAGGATCCTCT GGCAGCGAGACACCAGGAACAAGCGAGTCAGCAACACCAGAGAGCAGTGGCGG CAGCAGCGGCGGCAGCGACAAGAAGTACAGCATCGGCCTGGCCATCGGCACCAACT CTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGT TCGACAGCGGCGAAACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGA TACACCAGACGGAAGAACCGGATCTGCTATCTGCAAGAGATCTTCAGCAACGAGATGG CCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGG ATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGCCTACC ACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAA GGCCGACCTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCA CTTCCTGATCGAGGGCGACCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCAT CCAGCTGGTGCAGACCTACAACCAGCTGTTCGAGGAAAACCCCATCAACGCCAGCGG CGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGACGGCTGGAAAA TCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGC CCTGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGC CAAACTGCAGCTGAGCAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCA GATCGGCGACCAGTACGCCGACCTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCAT CCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCAAGGCCCCCCTGAGCGC CTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGCTCTC GTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACG GCTACGCCGGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAA GCCCATCCTGGAAAAGATGGACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGA GGACCTGCTGCGGAAGCAGCGGACCTTCGACAACGGCAGCATCCCCCACCAGATCCA CCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTACCCATTCCTGAAG GACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCA TCACCCCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCA TCGAGCGGATGACCAACTTCGATAAGAACCTGCCCAACGAGAAGGTGCTGCCCAAGC ACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGCTGACCAAAGTGAAATACGT GACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGCCATCGT GGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTA CTTCAAGAAAATCGAGTGCCTGTCCTACGAGACAGAGATCCTGACAGTGGAGT ATGGCCTGCTGCCAATCGGCAAGATCGTGGAGAAGAGGATCGAGTGTACCG TGTACTCTGTGGATAACAATGGCAACATCTATACACAGCCCGTGGCACAGTG GCACGATAGGGGAGAGCAGGAGGTGTTCGAGTATTGCCTGGAGGACGGCA GCCTGATCAGGGCAACCAAGGACCACAAGTTCATGACAGTGGATGGCCAGA TGCTGCCCATCGACGAGATTTTCGAGCGGGAGCTGGACCTGATGAGAGTGG ATAACCTGCCTAATAGCGGAGGCAGTAAAAGAACAGCAGACGGGAGTGAGTTT GAGCCCAAGAAAAAGAGAAAGGTGTAAGATCTGATAATCAACCTCTGGATTACA AAATTTGTGAAAGATTGACTGGTATTCTTAACTATGTTGCTCCTTTTACGCTATGT GGATACGCTGCTTTAATGCCTTTGTATCATGCTATTGCTTCCCGTATGGCTTTCAT TTTCTCCTCCTTGTATAAATCCTGGTTAGTTCTTGCCACGGCGGAACTCATCGCCG CCTGCCTTGCCCGCTGCTGGACAGGGGCTCGGCTGTTGGGCACTGACAATTCCGT GGTGCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCC TTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAA TTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCA GGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGT GGGCTCTATGGCTCGAGAAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTAAACTTGCTATGCTGTTTCCAGCATAGCTCTTAAACTGCCCCG GGTATCGGCTGCCGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAA TCGAAATACTTTCAAGTTACGGTAAGCATATGATAGTCCATTTTAAAACATA ATTTTAAAACTGCAAACTACCCAAGAAATTATTACTTTCTACGTCACGTATTT TGTACTAATATCTTTGTGTTTACAGTCAAATTAATTCTAATTATCTCTCTAAC AGCCTTGTATCGTATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCT TCCTGCCCGACCTTGCGGCCGCAGGAACCCCTAGTGATGGAGTTGGCCACTCCC TCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCC CGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG (SEQ ID NO: 170)
[0478] Sequence 6: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (C-terminus) ITR-Cbh promoter-NLS-NpuC-SpCas9 (amino acids 573-1367)-UGI-NLS-WPRE-bovine growth hormone(bGH)-derived poly(A)-PRNP R37X F+E-sgRNA (reverse complement)- human U6 promoter (reverse complement)-ITR CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCG ACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCC AACTCCATCACTAGGGGTTCCTGCGGCCTCTAGATCAGGGTACCCGTTACATAAC TTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTC AATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGG TAAACTGCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTA TTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTGTGCCCAGTACATGACCTT ATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATG GTCGAGGTGAGCCCCACGTTCTGCTTCACTCTCCCCATCTCCCCCCCCTCCCCACC CCCAATTTTGTATTTATTTATTTTTTAATTATTTTGTGCAGCGATGGGGGCGGGGG GGGGGGGGGGGCGCGCGCCAGGCGGGGCGGGGCGGGGCGAGGGGCGGGGCGG GGCGAGGCGGAGAGGTGCGGCGGCAGCCAATCAGAGCGGCGCGCTCCGAAAGT TTCCTTTTATGGCGAGGCGGCGGCGGCGGCGGCCCTATAAAAAGCGAAGCGCGC GGCGGGCGGGAGTCGCTGCGACGCTGCCTTCGCCCCGTGCCCCGCTCCGCCGCCG CCTCGCGCCGCCCGCCCCGGCTCTGACTGACCGCGTTACTCCCACAGGTGAGCGG GCGGGACGGCCCTTCTCCTCCGGGCTGTAATTAGCTGAGCAAGAGGTAAGGGTTT AAGGGATGGTTGGTTGGTGGGGTATTAATGTTTAATTACCTGGAGCACCTGCCTG AAATCACTTTTTTTCAGGTTGGACCGGTGCCACCATGAAACGGACAGCCGACGG AAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCATCAAGATTGCTACAC GGAAATACCTGGGAAAGCAGAACGTGTACGACATCGGCGTGGAGCGGGATC ACAACTTCGCCCTGAAGAATGGCTTTATCGCCAGCAATTGCTTCGACTCCGTG GAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACATACCACGATCTGC TGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTG AAAACCTATGCCCACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGAT ACACCGGCTGGGGCAGGCTGAGCCGGAAGCTGATCAACGGCATCCGGGACAAGCAG TCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCA TGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGT GTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCG CCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGA TGGGCCGGCACAAGCCCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCA CCCAGAAGGGACAGAAGAACAGCCGCGAGAGAATGAAGCGGATCGAAGAGGGCATCA AAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACACCCAGCTGCAGA ACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGG AACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGACTCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGG CAAGAGCGACAACGTGCCCTCCGAAGAGGTCGTGAAGAAGATGAAGAACTACTGGCG GCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTTCGACAATCTGACCAAGGCC GAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAGCTGGTG GAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACT AAGTACGACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCA AGCTGGTGTCCGATTTCCGGAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAA CTACCACCACGCCCACGACGCCTACCTGAACGCCGTCGTGGGAACCGCCCTGATCAA AAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTG CGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTTC TTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGAT CCGGAAGCGGCCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAA GGGCCGGGATTTTGCCACCGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGT GAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCCAAGAG GAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAGTACGGCGG CTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGG CAAGTCCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGA AGCAGCTTCGAGAAGAATCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGA AAAAGGACCTGATCATCAAGCTGCCTAAGTACTCCCTGTTCGAGCTGGAAAACGGCCG GAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAAACGAACTGGCCCTGCC CTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGGCTCC CCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACG AGATCATCGAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCT GGACAAAGTGCTGTCCGCCTACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGC CGAGAATATCATCCACCTGTTTACCCTGACCAATCTGGGAGCCCCTGCCGCCTTCAAG TACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAAGAGGTGCTGGAC GCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCT CAGCTGGGAGGTGACAGCGGCGGGAGCGGCGGGAGCGGGGGGAGCACTAATCTG AGCGACATCATTGAGAAGGAGACTGGGAAACAGCTGGTCATTCAGGAGTCCATC CTGATGCTGCCTGAGGAGGTGGAGGAAGTGATCGGCAACAAGCCAGAGTCTGAC ATCCTGGTGCACACCGCCTACGACGAGTCCACAGATGAGAATGTGATGCTGCTG ACCTCTGACGCCCCCGAGTATAAGCCTTGGGCCCTGGTCATCCAGGATTCTAACG GCGAGAATAAGATCAAGATGCTGAGCGGAGGATCCAAAAGAACCGCCGACGGC AGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAAGATCTGATAATCAACCTC TGGATTACAAAATTTGTGAAAGATTGACTGGTATTCTTAACTATGTTGCTCCTTTT ACGCTATGTGGATACGCTGCTTTAATGCCTTTGTATCATGCTATTGCTTCCCGTAT GGCTTTCATTTTCTCCTCCTTGTATAAATCCTGGTTAGTTCTTGCCACGGCGGAAC TCATCGCCGCCTGCCTTGCCCGCTGCTGGACAGGGGCTCGGCTGTTGGGCACTGA CAATTCCGTGGTGCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTC CCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAA ATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGG GTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGG GATGCGGTGGGCTCTATGGCTCGAGAAAAAAAGCACCGACTCGGTGCCACTTTTTCA AGTTGATAACGGACTAGCCTTATTTAAACTTGCTATGCTGTTTCCAGCATAGCTCTTAAA CTGCCCCGGGTATCGGCTGCCGGTGTTTCGTCCTTTCCACAAGATATATAAAGC CAAGAAATCGAAATACTTTCAAGTTACGGTAAGCATATGATAGTCCATTTTA AAACATAATTTTAAAACTGCAAACTACCCAAGAAATTATTACTTTCTACGTCA CGTATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAATTCTAATTATCT CTCTAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCATGGGAAATAG GCCCTCTTCCTGCCCGACCTTGCGGCCGCAGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGC CCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG (SEQ ID NO: 171)
[0479] Sequence 7: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (N-terminus) with hSYN promoter ITR-hSYN promoter-NLS-TadCBEd deaminase domain-linker-SpCas9 (amino acids 1- 572)-NpuN-NLS-WPRE-bovine growth hormone(bGH)-derived poly(A)-PRNP R37X F+E- sgRNA (reverse complement)-human U6 promoter (reverse complement)-ITR CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCG ACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCC AACTCCATCACTAGGGGTTCCTGCGGCCTCTAGATCAGGGTACCAGTGCAAGTGG GTTTTAGGACCAGGATGAGGCGGGGTGGGGGTGCCTACCTGACGACCGACCCCG ACCCACTGGACAAGCACCCAACCCCCATTCCCCAAATTGCGCATCCCCTATCAGA GAGGGGGAGGGGAAACAGGATGCGGCGAGGCGCGTGCGCACTGCCAGCTTCAG CACCGCGGACAGTGCCTTCGCCCCCGCCTGGCGGCGCGCGCCACCGCCGCCTCA GCACTGAAGGCGCGCTGACGTCACTCGCCGGTCCCCCGCAAACTCCCCTTCCCGG CCACCTTGGTCGCGTCCGCGCCGCCGCCGGCCCAGCCGGACCGCACCACGCGAG GCGCGAGATAGGGGGGCACGGGCGCGACCATCTGCGCTGCGGCGCCGGCGACTC AGCGCTGCCTCAGTCTGCGGTGGGCAGCGGAGGAGTCGTGTCGTGCCTGAGAGC GCAGGACCGGTGCCACCATGAAACGGACAGCCGACGGAAGCGAGTTCGAGTCAC CAAAGAAGAAGCGGAAAGTCAGTTCTGAGGTGGAGTTTTCCCACGAGTACTG GATGAGACATGCCCTGACCCTGGCCAAGAGGGCACGGGATGAGAGGAAGG CGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCT GGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATAG CCCTGAGACAGGGCGGCCTGGTCATGCAGAACTACAGACTGATTGACGCCA CCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGCCATGATCA ACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCG CCGCAGGCTCCCTGATGAACGTGCTGAACTACCCCGGAATGAATCACCGCG TCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCG ATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAGA GCTCCATCAACTCTGGAGGATCTAGCGGAGGATCCTCTGGCAGCGAGACACCAG GAACAAGCGAGTCAGCAACACCAGAGAGCAGTGGCGGCAGCAGCGGCGGCAGC GACAAGAAGTACAGCATCGGCCTGGCCATCGGCACCAACTCTGTGGGCTGGGCCGTG ATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGGTGCTGGGCAACACCGAC CGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGAAACA GCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAAC CGGATCTGCTATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCT TCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGC ACCCCATCTTCGGCAACATCGTGGACGAGGTGGCCTACCACGAGAAGTACCCCACCAT CTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGACCTGCGGCTGAT CTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGA CCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTA CAACCAGCTGTTCGAGGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCAT CCTGTCTGCCAGACTGAGCAAGAGCAGACGGCTGGAAAATCTGATCGCCCAGCTGCC CGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCCTGAGCCTGGGCCTGAC CCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGC CGACCTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATCCTG AGAGTGAACACCGAGATCACCAAGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACG ACGAGCACCACCAGGACCTGACCCTGCTGAAAGCTCTCGTGCGGCAGCAGCTGCCTG AGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCCGGCTACATTGA CGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATG GACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCA GCGGACCTTCGACAACGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGC CATTCTGCGGCGGCAGGAAGATTTTTACCCATTCCTGAAGGACAACCGGGAAAAGATC GAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCCCTCTGGCCAGGGGAAAC AGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAACTTC GAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAAC TTCGATAAGAACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGT ACTTCACCGTGTATAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAA GCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGCCATCGTGGACCTGCTGTTCAAGAC CAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAGAAAATCGAGTG CCTGTCCTACGAGACAGAGATCCTGACAGTGGAGTATGGCCTGCTGCCAAT CGGCAAGATCGTGGAGAAGAGGATCGAGTGTACCGTGTACTCTGTGGATAA CAATGGCAACATCTATACACAGCCCGTGGCACAGTGGCACGATAGGGGAGA GCAGGAGGTGTTCGAGTATTGCCTGGAGGACGGCAGCCTGATCAGGGCAAC CAAGGACCACAAGTTCATGACAGTGGATGGCCAGATGCTGCCCATCGACGA GATTTTCGAGCGGGAGCTGGACCTGATGAGAGTGGATAACCTGCCTAATAG CGGAGGCAGTAAAAGAACAGCAGACGGGAGTGAGTTTGAGCCCAAGAAAAAGA GAAAGGTGTAAGATCTGATAATCAACCTCTGGATTACAAAATTTGTGAAAGATTG ACTGGTATTCTTAACTATGTTGCTCCTTTTACGCTATGTGGATACGCTGCTTTAAT GCCTTTGTATCATGCTATTGCTTCCCGTATGGCTTTCATTTTCTCCTCCTTGTATAA ATCCTGGTTAGTTCTTGCCACGGCGGAACTCATCGCCGCCTGCCTTGCCCGCTGC TGGACAGGGGCTCGGCTGTTGGGCACTGACAATTCCGTGGTGCGACTGTGCCTTC TAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAG GTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTG AGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAG GATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTCGAG AAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTAA ACTTGCTATGCTGTTTCCAGCATAGCTCTTAAACTGCCCCGGGTATCGGCTGCCGGTG TTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAATCGAAATACTTTCAAG TTACGGTAAGCATATGATAGTCCATTTTAAAACATAATTTTAAAACTGCAAA CTACCCAAGAAATTATTACTTTCTACGTCACGTATTTTGTACTAATATCTTTG TGTTTACAGTCAAATTAATTCTAATTATCTCTCTAACAGCCTTGTATCGTATA TGCAAATATGAAGGAATCATGGGAAATAGGCCCTCTTCCTGCCCGACCTTGC GGCCGCAGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTC GCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGC GGCCTCAGTGAGCGAGCGAGCGCGCAG (SEQ ID NO: 172)
[0480] Sequence 8: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (C-terminus) with hSYN promoter ITR-hSYN promoter-NLS-NpuC-SpCas9 (amino acids 573-1367)-UGI-NLS-WPRE-bovine growth hormone(bGH)-derived poly(A)-PRNP R37X F+E-sgRNA (reverse complement)- human U6 promoter (reverse complement)-ITRCTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCG ACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCC AACTCCATCACTAGGGGTTCCTGCGGCCTCTAGATCAGGGTACCAGTGCAAGTGG GTTTTAGGACCAGGATGAGGCGGGGTGGGGGTGCCTACCTGACGACCGACCCCG ACCCACTGGACAAGCACCCAACCCCCATTCCCCAAATTGCGCATCCCCTATCAGA GAGGGGGAGGGGAAACAGGATGCGGCGAGGCGCGTGCGCACTGCCAGCTTCAG CACCGCGGACAGTGCCTTCGCCCCCGCCTGGCGGCGCGCGCCACCGCCGCCTCA GCACTGAAGGCGCGCTGACGTCACTCGCCGGTCCCCCGCAAACTCCCCTTCCCGG CCACCTTGGTCGCGTCCGCGCCGCCGCCGGCCCAGCCGGACCGCACCACGCGAG GCGCGAGATAGGGGGGCACGGGCGCGACCATCTGCGCTGCGGCGCCGGCGACTC AGCGCTGCCTCAGTCTGCGGTGGGCAGCGGAGGAGTCGTGTCGTGCCTGAGAGC GCAGGACCGGTGCCACCATGAAACGGACAGCCGACGGAAGCGAGTTCGAGTCAC CAAAGAAGAAGCGGAAAGTCATCAAGATTGCTACACGGAAATACCTGGGAAA GCAGAACGTGTACGACATCGGCGTGGAGCGGGATCACAACTTCGCCCTGAA GAATGGCTTTATCGCCAGCAATTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAA GATCGGTTCAACGCCTCCCTGGGCACATACCACGATCTGCTGAAAATTATCAAGGACA AGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGAAGATATCGTGCTGACCCT GACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCCCACCTG TTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGG CTGAGCCGGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTG GATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACG ACAGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAGGGCGATA GCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTAAGAAGGGCATCC TGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCC GAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAG AACAGCCGCGAGAGAATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCA GATCCTGAAAGAACACCCCGTGGAAAACACCCAGCTGCAGAACGAGAAGCTGTACCTG TACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGAACTGGACATCAACCGG CTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGACTCCA TCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGC CCTCCGAAGAGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCA AGCTGATTACCCAGAGAAAGTTCGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGA GCGAACTGGATAAGGCCGGCTTCATCAAGAGACAGCTGGTGGAAACCCGGCAGATCA CAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACGACGAGAATGA CAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTC CGGAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACG ACGCCTACCTGAACGCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGA AAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAGATGATCGCCAA GAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTTCTTCTACAGCAACATCATG AACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGGCCTCTGA TCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCA CCGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGC AGACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGAT CGCCAGAAAGAAGGACTGGGACCCTAAGAAGTACGGCGGCTTCGACAGCCCCACCGT GGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGTCCAAGAAACTGAA GAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAAT CCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCA AGCTGCCTAAGTACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCT CTGCCGGCGAACTGCAGAAGGGAAACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGGCTCCCCCGAGGATAATGAGCA GAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATCGAGCAGATC AGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCG CCTACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCT GTTTACCCTGACCAATCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATC GACCGGAAGAGGTACACCAGCACCAAAGAGGTGCTGGACGCCACCCTGATCCACCAG AGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTCAGCTGGGAGGTGACAGC GGCGGGAGCGGCGGGAGCGGGGGGAGCACTAATCTGAGCGACATCATTGAGAA GGAGACTGGGAAACAGCTGGTCATTCAGGAGTCCATCCTGATGCTGCCTGAGGA GGTGGAGGAAGTGATCGGCAACAAGCCAGAGTCTGACATCCTGGTGCACACCGC CTACGACGAGTCCACAGATGAGAATGTGATGCTGCTGACCTCTGACGCCCCCGA GTATAAGCCTTGGGCCCTGGTCATCCAGGATTCTAACGGCGAGAATAAGATCAA GATGCTGAGCGGAGGATCCAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCA AGAAGAAGAGGAAAGTCTAAGATCTGATAATCAACCTCTGGATTACAAAATTTG TGAAAGATTGACTGGTATTCTTAACTATGTTGCTCCTTTTACGCTATGTGGATACG CTGCTTTAATGCCTTTGTATCATGCTATTGCTTCCCGTATGGCTTTCATTTTCTCCT CCTTGTATAAATCCTGGTTAGTTCTTGCCACGGCGGAACTCATCGCCGCCTGCCTT GCCCGCTGCTGGACAGGGGCTCGGCTGTTGGGCACTGACAATTCCGTGGTGCGAC TGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGAC CCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGC ATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAA GGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTAT GGCTCGAGAAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGC CTTATTTAAACTTGCTATGCTGTTTCCAGCATAGCTCTTAAACTGCCCCGGGTATCGGC TGCCGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAATCGAAATA CTTTCAAGTTACGGTAAGCATATGATAGTCCATTTTAAAACATAATTTTAAAA CTGCAAACTACCCAAGAAATTATTACTTTCTACGTCACGTATTTTGTACTAAT ATCTTTGTGTTTACAGTCAAATTAATTCTAATTATCTCTCTAACAGCCTTGTA TCGTATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCTTCCTGCCCG ...
Claims
CLAIMSWhat is claimed is:
1. A composition comprising:(i) a first nucleotide sequence encoding a guide RNA (gRNA); and(ii) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N, wherein the guide RNA is configured to direct a recombined nucleobase base editor to install a premature stop codon at one or more of positions R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene.
2. The composition of claim 1, further comprising a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the split nucleobase editor.
3. A composition comprising: a first nucleotide sequence encoding a guide RNA (gRNA); and a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C- terminal portion of a split nucleobase editor, wherein the guide RNA is configured to direct a recombined nucleobase base editor to install a premature stop codon at one or more of positions R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene.
4. The composition of claim 3, further comprising a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N.
5. A composition, comprising:(i) a first nucleotide sequence encoding a guide RNA (gRNA);(ii) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N and a guide RNA; and(iii) a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C- terminal portion of the split nucleobase editor and the guide RNA,wherein the guide RNA is configured to direct a recombined base editor to install a premature stop codon at one or more of positions R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene, and wherein the N-terminal portion of the split nucleobase editor and the C- terminal portion of the split nucleobase editor may be joined to form a recombined nucleobase editor.
6. The composition of any one of claims 1-4, wherein the N-terminal portion of the split nucleobase editor and the C-terminal portion of the split nucleobase editor may be joined to form the recombined nucleobase editor.
7. The composition of any one of claims 1-6, wherein the first nucleotide sequence encoding the gRNA is operably linked to a promoter.
8. The composition of any one of claims 1-6, wherein the second nucleotide sequence encoding the N-terminal portion of the split nucleobase editor is operably linked to a promoter.
9. The composition of any one of claims 1-6, wherein the third nucleotide sequence encoding the C-terminal portion of the split nucleobase editor is operably linked to a promoter.
10. The composition of any one of claims 7-9, wherein the promoter is selected from the group consisting of Cbh, hSYN, and U6.
11. The composition of any one of claims 1-10, further comprising a fourth nucleotide sequence encoding a terminator sequence.
12. The composition of any one of claims 1-11, further comprising a fifth nucleotide sequence encoding a uracil glycosylase inhibitor (UGI) domain.
13. The composition of claim 12, wherein the UGI comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5% , or is 100% identical to the amino acid sequence of any one of SEQ ID NOs: 17-20.
14. The composition of any one of claims 1-13, further comprising a sixth nucleotide sequence encoding an inverted terminal repeat (ITR) domain of an adeno associated virus (AAV).
15. The composition of claim 14, wherein the AAV has a serotype selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, and AAV11.
16. The composition of claim 15, wherein the AAV serotype is AAV2 or AAV9.
17. The composition of any one of claims 1-16, wherein the recombined nucleobase editor comprises a napDNAbp domain and a deaminase domain.
18. The composition of claim 17, wherein the napDNAbp domain comprises a nucleotide sequence or an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, or is 100% identical to any one of SEQ ID NOs: 39-45 or 46-61.
19. The composition of claim 17, wherein the deaminase domain comprises an adenosine deaminase.
20. The composition of claim 19, wherein the adenosine deaminase domain comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, least 99.5%, or is 100% identical to the amino acid sequence of SEQ ID NOs: 98-166.
21. The composition of claim 17, wherein the deaminase domain comprises a cytosine deaminase.
22. The composition of claim 21, wherein the cytosine deaminase comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, least 99.5%, or at least 100% identical to the amino acid sequence of any one of SEQ ID NO: 62-97.
23. The composition of any one of claims 1-22, wherein the guide RNA comprises a flip- and-extend scaffold.
24. The composition of any one of claims 1-22, wherein the gRNA comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, least 99.5%, or at least 100% identical to the nucleotide sequence of any one of SEQ ID NOs: 221, 222, or 225.
25. The composition of any one of claims 1-24, wherein the gRNA and the recombined nucleobase editor form a complex.
26. The composition of any one of claims 1-25, wherein the N-terminal portion of the split nucleobase editor comprises an amino acid sequence that is at least 80%, is at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5% or is 100% identical to a portion of any one of SEQ ID NOs: 47-61 that correspond to amino acids 1-573 or 1-637 of SEQ ID NO: 46.
27. The composition of any one of claims 1-26, wherein at least a portion of the C- terminal portion of the split nucleobase editor comprises an amino acid sequence that is at least 80%, is at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5% or is 100% identical to a portion of any one of SEQ ID NOs: 47-61 that correspond to amino acids 574-1368 or 638-1368 of SEQ ID NO: 46.
28. The composition of any one of claims 1-27, wherein the recombined nucleobase editor comprises a nucleotide sequence that is at least 80%, is at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5% or is 100% identical to the nucleotide sequence of SEQ ID NO: 186-191.
29. The compositions of any one of claims 1-28, further comprising a seventh nucleotide sequence encoding one or more microRNA (miRNA) target sites.
30. The composition of claim 29, wherein the microRNA target site is encoded within a 3' untranslated region (UTR) of the nucleobase editor.
31. The composition of claim 29 or 30, wherein the miR target site comprises miR124, miR204, miR181, miRl, miR206, miR126, miR134, miR133, miR208, miR302, miR338, miR219, miR9, miR218, miR7, miR128, miR125, miR138, miR132, miR212, miR137, miR31, miR127, miR143, miR346, miR708, miRlO, miR192, miR194, miR215, miR216, miR122a, miR92a, miR483, miR130, miR292, miR217, miR193a, miR128a, miR150, miR181a, miR155, and / or miR142 target sites.
32. The composition of any one of claims 29-31, wherein the miR target site comprises an miRl 83 target site, an miR122 target site, or both.
33. The composition of any one of claims 29-31, wherein the miR target sites comprise miR124, miR204, and / or miR181 target sites, and wherein miR124, miR204, and / or miR181 are present in eye tissue.
34. The composition of any one of claims 29-31 wherein the miR target sites comprise miRl, miR206, miR126, miR134, miR133, miR208, miR302, miR338, miR219, miR124, miR9, miR218, miR7, and / or miR128 target sites, and wherein one or more of miRl, miR206, miR126, miR134, miR133, miR208, miR302, miR338, miR219, miR124, miR9, miR218, miR7, and / or miR128 are present in heart tissue.
35. The composition of any one of claims 29-31, wherein the miR target sites comprise miR125, miR138, miR132, miR212, miR137, miR31, miR127, miR143, miR346, and / or miR708 target sites, and wherein one or more of miR125, miR138, miR132, miR212, miR137, miR31, miR127, miR143, miR346, and / or miR708are present in brain / nervous tissue.
36. The composition of any one of claims 29-31, wherein the miR target sites comprise miR10, miR192, miR204, miR194, miR215, and / or miR216 target sites, and wherein one or more of miR10, miR192, miR204, miR194, miR215, and / or miR216 are present in kidney tissue.
37. The composition of any one of claims 29-31, wherein the miR target sites comprise miR122a, miR192, miR92a, and / or miR483 target sites, and wherein one or more of miR122a, miR192, miR92a, and / or miR483 are present in liver tissue.
38. The composition of any one of claims 29-31, wherein the miR target site comprises miR126, and wherein the miR126 is present in lung tissue.
39. The composition of any one of claims 29-31, wherein the miR target sites comprise miR126, miR130, miR302, and / or miR292 target sites, and wherein one or more of miR126, miR130, miR302, and / or miR292 are present in hematopoietic and pluripotent stem cells.
40. The composition of any one of claims 29-31, wherein the miR target sites comprise miR216 and / or miR217 target sites, and wherein the miR216 and / or miR217 are present in pancreas tissue.
41. The composition of any one of claims 29-31, wherein the target sites comprise miR133, miR1, miR206, miR134, miR193a, and / or miR128a target sites, and wherein one or more of miR133, miR1, miR206, miR134, miR193a, and / or miR128a present in muscle tissue.
42. The composition of any one of claims 29-31, wherein the target sites comprise miR150, miR181a, miR155, and / or miR142 target sites, and wherein one or more of miR150, miR181a, miR155, and / or miR142 are present in organs and / or tissues of the immune system.
43. The composition of any one of claims 29-42, wherein the seventh nucleotide sequence encodes one or more miR target sites in the 3’ untranslated region (UTR) of the nucleobase editor.
44. The composition of any one of claims 29-42, wherein the seventh nucleotide sequence encodes two or more miR target sites in the 3’ untranslated region (UTR) of the nucleobase editor.
45. The composition of any one of claims 29-42, wherein the seventh nucleotide sequence encodes three or more miR target sites in the 3’ untranslated region (UTR) of the nucleobase editor.
46. The composition of any one of claims 29-42, wherein the seventh nucleotide sequence encodes four or more miR target sites in the 3’ untranslated region (UTR) of the nucleobase editor.
47. The composition of any one of claims 29-42, wherein the seventh nucleotide sequence encodes five or six miR target sites in the 3’ untranslated region (UTR) of the nucleobase editor.
48. A composition, comprising: a first recombinant adeno associated virus (rAAV) particle comprising: (i) a first nucleotide sequence encoding a guide RNA (gRNA) and (ii) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N, wherein the guide RNA is configured to direct a recombined base editor to install a premature stop codon at one or more of positions R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene.
49. The composition of claim 48, further comprising a second recombinant adeno associated virus (rAAV) particle comprising a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the split nucleobase editor.
50. A composition, comprising: a second recombinant adeno associated virus (rAAV) particle comprising: (i) a first nucleotide sequence encoding a guide RNA (gRNA); and (ii) a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C- terminal portion of the split nucleobase editor, wherein the guide RNA is configured to direct a recombined base editor to install a premature stop codon at one or more of positions R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene.
51. The composition of claim 50, further comprising a first recombinant adeno associated virus (rAAV) particle comprising a second nucleotide sequence encoding a N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N.
52. A composition, comprising: a first recombinant adeno associated virus (rAAV) particle comprising a second nucleotide sequence encoding a N-terminal portion of a split nucleobase editor fused at its C- terminus to an intein-N; and a second recombinant adeno associated virus (rAAV) particle comprising a third nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the split nucleobase editor, wherein the first recombinant adeno associated virus (rAAV) particle and / or the second recombinant adeno associated virus (rAAV) particle further comprises a first nucleotide sequence encoding a guide RNA, wherein the guide RNA is configured to direct a recombined base editor to install a premature stop codon at one or more of positions R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene, and wherein the N-terminal portion of the split nucleobase editor and the C-terminal portion of the split nucleobase editor may be joined to form a recombined nucleobase editor.
53. The compositions of any one of claims 48-52, further comprising a fourth nucleotide sequence encoding one or more microRNA (miRNA) target sites.
54. The composition of claim 53, wherein the microRNA target site is encoded within a 3′ untranslated region (UTR) of the nucleobase editor.
55. The composition of claim 53 or 54, wherein the miR target site comprises miR124, miR204, miR181, miR1, miR206, miR126, miR134, miR133, miR208, miR302, miR338, miR219, miR9, miR218, miR7, miR128, miR125, miR138, miR132, miR212, miR137, miR31, miR127, miR143, miR346, miR708, miR10, miR192, miR194, miR215, miR216, miR122a, miR92a, miR483, miR130, miR292, miR217, miR193a, miR128a, miR150, miR181a, miR155, and / or miR142 target sites.
56. A guide RNA (gRNA) comprising a flip-and-extend scaffold configured to direct a nucleobase editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 in a PrP protein encoded by a PRNP gene.
57. The gRNA of claim 56, wherein the PRNP gene encodes for a prion protein (PrP).
58. The gRNA of claim 56 or 57, wherein installing the premature stop codon in the PRNP gene produces a truncated PrP protein.
59. The composition of any one of claims 56-58, wherein the gRNA comprises a spacer sequence comprising a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, least 99.5%, or at least 100% identical to the nucleotide sequence of any one of SEQ ID NOs: 737, 740-744, or 749-752.
60. The composition of any one of claims 56-58, wherein the gRNA comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, least 99.5%, or at least 100% identical to the nucleotide sequence of any one of SEQ ID NOs: 221-222 or 225.
61. The composition of any one of claims 56-60, wherein the gRNA and the recombined nucleobase editor form a complex.
62. A complex comprising a base editor and the guide RNA of any one of claim 56-61.
63. A vector comprising a nucleic acid sequence encoding the base editor and guide RNA of the complex of claim 62 and / or the guide RNA of any one of claims 56-61.
64. A vector comprising a nucleic acid sequence encoding the base editor or a portion of the base editor of any one of claims 1-55.
65. A cell comprising the compositions of any one of claims 1-55, the guide RNA of any one of claims 56-61, complex of claim 62, or vector of claim 64.
66. A pharmaceutical composition comprising compositions of any one of claims 1-55, guide RNA of any one of claims 56-61, complex of claim 62, the vector of claim 64, or the cell of claim 65.
67. A kit comprising the compositions of any one of claims 1-55, guide RNA of any one of claims 56-61, complex of claim 62, the vector of claim 64, the cell of claim 65, or the pharmaceutical composition of claim 66.
68. A method of producing truncated PrP protein variants using base editing, the method comprising contacting a cell, or nucleic acid encoding the PRNP gene, with the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67, wherein the contacting results in recombination of a N-terminal portion of a split nucleobase editor and a C-terminal portion of the split nucleobase editor to yield a recombined nucleobase editor, wherein the guide RNA directs the recombined base editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 of a PrP protein encoded by the PRNP gene, and wherein installing the premature stop codon produces truncated PrP protein variants.
69. The method of claim 68, wherein the truncated PrP protein variant is not a pathogenic mutant.
70. The method of claim 68 or 69, wherein the truncated PrP protein variant at least partially inhibits the templated misfolding of native PrP by pathogenic PrP mutants.
71. The method of any one of claims 68-70, wherein the truncated PrP protein is a N- terminally truncated PrP protein.
72. The method of any one of claims 68-70, wherein the truncated PrP protein is a C- terminally truncated PrP protein.
73. The method of any one of claims 68-72, wherein the gRNA installs the premature stop codon at amino acid position R37 in the PrP protein.
74. A method of reducing the templated misfolding of a prion protein (PrP) by pathogenic mutants in a subject, the method comprising contacting a PRNP gene encoding the PrPprotein with the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67, wherein the guide RNA comprises a flip-and-extend scaffold configured to direct the base editor to install a premature stop codon at one or more at one or more of positions R37, W57, W65, W81, or Q83 in the PrP protein, wherein installing the premature stop codon produces truncated PrP protein variants, and wherein the truncated PrP protein variants are not pathogenic mutants.
75. A method comprising administering to a subject in need thereof a therapeutically effective amount of with the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67.
76. A method for treating or preventing prion disease, the method comprising administering to a subject in need thereof a therapeutically effective amount of the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67.
77. Use of the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67 in the manufacture of a medicament for the treatment of a disease or disorder associated with a PRNP gene.
78. The composition of any one of claims 1-55, further comprising one or more polynucleotides encoding one or more of the nucleotide sequences.
79. The composition of claim 78, wherein a first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 167 [corresponding to sequence 1, Dual-AAV BE3.9max PRNP R37X sgRNA (N-terminus)] and a second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 168 [corresponding to sequence 2, Dual-AAV BE3.9max PRNP R37X sgRNA (C-terminus)].
80. The composition of claim 78, wherein the first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 169 [corresponding to sequence 3, Dual-AAV TadCBEd PRNP R37X sgRNA (N-terminus)] and the second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 168 [corresponding to sequence 2, (C-terminus)].
81. The composition of claim 78, wherein the first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 170 [corresponding to Sequence 5: Dual-AAV TadCBEd PRNP R37X F+E- sgRNA (N-terminus)] and the second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 171 [Sequence 6: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (C-terminus)].
82. The composition of claim 78, wherein the first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 172 [corresponding to sequence 7: Dual-AAV TadCBEd PRNP R37X F+E- sgRNA (N-terminus) with hSYN promoter] and the second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 173 [corresponding to sequence 8: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (C-terminus) with hSYN promoter].
83. The composition of claim 78, wherein the first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 174 [corresponding to Sequence 9: Dual-AAV TadCBEd PRNP R37X F+E- sgRNA (N-terminus) with hSYN promoter + miR-183 target sites] and the second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 175 [corresponding to Sequence 10: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (C-terminus) with hSYN promoter + miR- 183 target sites].
84. The composition of claim 78, wherein the first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 176 [corresponding to Sequence 11: Dual-AAV TadCBEd PRNP R37X F+E- sgRNA (N-terminus) with hSYN promoter + miR-122 target sites] and the second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 177 [corresponding to Sequence 12: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (C-terminus) with hSYN promoter + miR- 122 target sites].
85. The composition of claim 78, wherein the first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 178 [corresponding to Sequence 13: Dual-AAV TadCBEd PRNP R37X F+E- sgRNA (N-terminus) with hSYN promoter + miR-183+ miR-122 target sites] and the second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 179 [corresponding to Sequence 14: Dual-AAV TadCBEd PRNP R37X F+E-sgRNA (C-terminus) with hSYN promoter + miR- 183 + miR-122 target sites].
86. The composition of claim 78, wherein the first polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 182 [corresponding to Sequence 17: Dual-AAV SpCas9-ABE8e(V106W) PRNP M1V F+E-sgRNA (N-terminus)] and the second polynucleotide comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, at least 99.5%, at least 99.8% or is 100% identical to the nucleic acid sequence of SEQ ID NO: 183 [corresponding to Sequence 18: Dual-AAV SpCas9-ABE8e(V106W) PRNP M1V F+E-sgRNA (C-terminus)].
87. Use of the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67 for producing truncated PrP protein variants using base editing comprising contacting a cell, or nucleic acid encoding the PRNP gene, with the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67, wherein the contacting results in recombination of a N-terminal portion of a split nucleobase editor and a C-terminal portion of the split nucleobase editor to yield a recombined nucleobase editor, wherein the guide RNA directs the recombined base editor to install a premature stop codon at position R37, W57, W65, W81, or Q83 of a PrP protein encoded by the PRNP gene, and wherein installing the premature stop codon produces truncated PrP protein variants.
88. Use of the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67 for reducing the templated misfolding of a prion protein (PrP) by pathogenic mutants in a subject comprising contacting a PRNP gene encoding the PrP protein with the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67,wherein the guide RNA comprises a flip-and-extend scaffold configured to direct the base editor to install a premature stop codon at one or more at one or more of positions R37, W57, W65, W81, or Q83 in the PrP protein, wherein installing the premature stop codon produces truncated PrP protein variants, and wherein the truncated PrP protein variants are not pathogenic mutants.
89. Use of the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67 for treating or preventing prion disease, the method comprising administering to a subject in need thereof a therapeutically effective amount of the composition of any one of claims 1-55, the guide RNA of any one of claims 56-61, the complex of claim 62, the vector of claim 64, the cell claim 65, the pharmaceutical composition of claim 66, and / or the kit of claim 67.
90. A composition comprising: (i) a first nucleotide sequence encoding a guide RNA (gRNA); and (ii) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N, wherein the guide RNA comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% identical to any one of the nucleic acid sequences of SEQ ID NOs: 21-38.
91. The composition of claim 90, wherein the gRNA is configured to direct a recombined nucleobase base editor to install a C>T or A>G base edit in a promoter region of a PRNP gene.
92. The composition of claim 91, wherein installation of the base edit reduces the transcription of the PRNP gene.
93. A composition comprising: (i) a first nucleotide sequence encoding a guide RNA (gRNA); and(ii) a second nucleotide sequence encoding an N-terminal portion of a split nucleobase editor fused at its C-terminus to an intein-N, wherein the guide RNA comprises a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, or at least 99.9% identical to any one of the nucleic acid sequences of SEQ ID NOs: A-Z.
94. The composition of claim 93, wherein the gRNA is configured to direct a recombined nucleobase base editor to install a C>T or A>G base edit in a promoter-proximal region of a PRNP gene.
95. The composition of claim 94, wherein installation of the base edit reduces the transcription of the PRNP gene.
Citation Information
Patent Citations
Transgenic animals secreting desired proteins into milk
EP0264166A1
CAS9 proteins including ligand-dependent inteins
US10077453B2
Adenosine nucleobase editors and uses thereof
US10113163B2
Nucleobase editors and uses thereof
US10167457B2
Multiple domain proteins
US20110059502A1