Nucleases and uses thereof
Programmable RNA-guided DNA endonucleases like MM762 address off-target and packaging issues in base editors, enabling high-fidelity and efficient in vivo genome editing for therapeutic applications, notably in Duchenne muscular dystrophy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-03-26
AI Technical Summary
Existing base editors face challenges in achieving high-fidelity, efficient, and safe in vivo genome editing due to off-target activity and packaging limitations, which hinder their therapeutic potential for diseases like Duchenne muscular dystrophy.
Development of programmable RNA-guided DNA endonucleases, such as MM762 and its derivatives, engineered for stringent target recognition and compatibility with single-AAV packaging, resulting in high-fidelity base editors like eMM762ABE and eMM762CBE, which reduce off-target events and enhance editing efficiency.
These editors achieve near-complete multiplex editing in human cells and restore dystrophin expression in a humanized mouse model of Duchenne muscular dystrophy, outperforming dual-AAV SpCas9-based editors, with minimal off-target activity and efficient in vivo delivery.
Smart Images

Figure PCTCN2025122889-FTAPPB-I100001 
Figure PCTCN2025122889-FTAPPB-I100002 
Figure PCTCN2025122889-FTAPPB-I100003
Abstract
Description
NUCLEASES AND USES THEREOF
[0001] REFERENCE TO RELATED APPLICATIONS
[0002] The instant application claims the priority to and the benefit of the filing date of PCT / CN2024 / 120148, filed on September 20, 2024, PCT / CN2024 / 121995, filed on September 27, 2024, and PCT / CN2025 / 095501, filed on May 16, 2025, the entire contents of which, including any drawings and sequence listing, are incorporated herein by reference.
[0003] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0004] The disclosure contains a Sequence Listing XML file which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on September 22, 2025, by software “WIPO Sequence” according to WIPO Standard ST. 26, is named SYP001PCT. xml, and is 277, 122 bytes in size. According to WIPO Standard ST. 26, symbol “t” is used to denote both T in DNA and U in RNA. Thus, in the instant sequence listing prepared according to ST. 26, wherever a sequence is an RNA, the T in the sequence shall be deemed as U.BACKGROUND
[0005] The precise installation of single-nucleotide changes without introducing double-strand breaks (DSBs) has been enabled by base editors, which couple a nucleobase deaminase to a programmable DNA-targeting scaffold, such as a Cas nickase1, 2. By directly converting C·G-to-T·A or A·T-to-G·C base pairs, these systems obviate the need for donor templates and cellular DSB repair pathways, thereby minimizing unwanted insertions, deletions, and chromosomal rearrangements3-7. This transformative capability has already enabled a wide array of applications, from functional genomics to potential therapies for monogenic diseases8-10.
[0006] Despite their therapeutic potential, base editors raise safety concerns due to off-target deamination8, 11. Two distinct classes of off-target events have been reported: Cas-independent deamination, in which the deaminase alone promiscuously modifies exposed single-stranded DNA or RNA12-16, and Cas-dependent off-target activity, in which the programmable nuclease mistargets genomic loci due to mismatch tolerance between the guide RNA and the target sequence17. Cas-independent off-target activity has been largely mitigated through protein engineering, notably via directed evolution of the deaminase domain14, 16, 18-21. In contrast, Cas‐dependent off‐target DNA editing remains a critical obstacle, with unintended conversions observed at sites bearing partial complementarity to the guide RNA22. High-fidelity SpCas9 variants23-29 and truncated single-guide RNA (sgRNA) 30, 31 have been shown to reduce SpCas9-dependent off-target activity; however, complete elimination remains challenging due to the intrinsic mismatch tolerance of sgRNA-DNA pairing across the genome32. Moreover, these strategies often result in a trade-off, with reduced on-target editing efficiency33-35.
[0007] In addition to concerns regarding editing specificity, achieving efficient delivery remains a major obstacle for the in vivo application of base editors. Base editors built on wild‐type SpCas9 achieve robust on‐target activity but exceed the ~4.7 kb packaging capacity of adeno‐associated virus (AAV) vectors, complicating in vivo delivery. Recent efforts to miniaturize base editors using IscB‐derived scaffolds have enabled single‐AAV delivery36-41, but this has come at the expense of reduced editing efficiency, restricted target-adjacent motif (TAM) compatibility and shortened spacer length requirements-factors that collectively limit their utility and elevate the risk of off-target activity42, 43.
[0008] Hence there remains an unmet need to identify or develop new programmable RNA-guided DNA endonucleases, nickases, and / or base editors to overcome one or more defects of the existing endonucleases, nickases, and / or base editors in the art.
[0009] Citation or identification of any document in the disclosure is not an admission that such a document is available as prior art to the disclosure. Each of the references mentioned or cited in the disclosure is incorporated by reference in its entirety.SUMMARY
[0010] In the disclosure, the applicant identified and developed new programmable RNA-guided DNA endonucleases, nickases, and base editors and other tools to meet the unmet need. In particular, the applicant established a new route to compact, high-fidelity base editing by leveraging the programmable RNA-guided DNA endonucleases. Through computational discovery, cryo-EM structure analysis, and systematic mutagenesis, the applicant nominated MM762 (SN007) -a 762-amino-acid programmable RNA-guided DNA endonuclease with intrinsically stringent target recognition-as a next-generation editing scaffold. By applying a multi-tiered engineering strategy, the applicant developed two base editors, eMM762ABE and eMM762CBE, that retain high on-target activity while markedly reducing off-target editing. Both editors are fully compatible with single-AAV packaging, enabling efficient in vivo delivery alongside their sgRNA expression cassettes. In human cell lines, eMM762ABE and eMM762CBE exhibited editing efficiencies comparable to SpCas9-based base editors, but reduced Cas-dependent off-target events by up to 404-fold. In primary human T cells, they supported near-complete multiplex editing of clinically relevant loci with minimal off-target activity. In a humanized mouse model of Duchenne muscular dystrophy, a single AAV encoding eMM762ABE restored dystrophin expression in over 90%of myofibers, outperforming dual-AAV SpCas9ABE8e. Collectively, these compact and high-fidelity editors provided a versatile platform for therapeutic genome editing, addressing long-standing challenges in delivery, efficiency, and safety.
[0011] Included in the disclosure is a collection of programmable RNA-guided DNA endonucleases, including SIMM nuclease ( “SN” for short) 001 to 007 (SN001-SN007) newly identified in the disclosure, and their derivatives and uses, meeting the unmet need in the art. Also included in the disclosure is a programmable guide nucleic acid (e.g., programmable guide RNA, also known as “guide RNA” or “gRNA” ) suitable for use to guide the corresponding programmable RNA-guided DNA endonucleases or derivatives thereof in the disclosure to a target DNA. Also included in the disclosure is a system or composition comprising the programmable RNA-guided DNA endonucleases or derivatives thereof in the disclosure and the corresponding programmable guide nucleic acid in the disclosure suitable for use to target (e.g., function on) a target DNA. Also included in the disclosure is a method of using or use of the system in the disclosure to target (e.g., function on) a target DNA.
[0012] Table 1
[0013] In an aspect, the disclosure provides a polypeptide comprising an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11.
[0014] In some embodiments, the polypeptide has at least one of endonuclease activity, nickase activity, and DNA binding property.
[0015] In some embodiments, the polypeptide comprises:
[0016] (a) an amino acid substitution at a position of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 that is corresponding to a position selected from the group consisting of D10, D386, H387, N410, H509, and D512 of SEQ ID NO: 13, for example, the positions of SEQ ID NO: 9 corresponding to the positions of D10, D386, H387, N410, H509, and D512 of SEQ ID NO: 13 are positions 10, D373, H374, N397, H496, and D499 of SEQ ID NO: 9, respectively, and optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ; or
[0017] (b) an amino acid substitution relative to any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 that is corresponding to an amino acid substitution selected from the group consisting of D10A, D386A, H387A, N410A, H509A, and D512A relative to SEQ ID NO: 13, for example, the amino acid substitutions relative to SEQ ID NO: 9 corresponding to the amino acid substitutions of D10A, D386A, H387A, N410A, H509A, and D512A relative to SEQ ID NO: 13 are D10A, D373A, H374A, N397A, H496A, and D499A relative to SEQ ID NO: 9, respectively. In some embodiments, the polypeptide comprises:
[0018] (a) an amino acid substitution at a position selected from the group consisting of D10, D386, H387, N410, H509, and D512 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ;
[0019] (b) an amino acid substitution at a position selected from the group consisting of 41, 42, 45, 46, 49, 50, 53, 56, 57, 58, 64, 66, 69, 72, 73, 80, 81, 82, 83, 84, 86, 87, 92, 95, 96, 102, 105, 106, 107, 108, 110, 112, 113, 114, 115, 116, 117, 118, 120, 121, 128, 135, 136, 138, 140, 141, 145, 153, 154, 157, 162, 163, 185, 187, 198, 200, 246, 250, 255, 268, 344, 349, 355, 368, 408, 412, 431, 432, 434, 443, 446, 449, 450, 479, 500, 523, 532, 535, 537, 549, 550, 553, 557, 560, 564, 575, 576, 577, 582, 583, 588, 594, 614, 618, 622, 626, 636, 640, 644, 645, 683, 684, 697, 698, 699, 710, 716, 718, 731, 732, 739, 740, 742, 743, 745, 747, 749, 751, 754, 758, and 759 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with Arginine (R) or Gly (G) ;
[0020] (c) an amino acid substitution at a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with P, T, L, A, or D; or
[0021] (d) a combination of any two or three of (a) , (b) , and (c) .
[0022] In some embodiments, the polypeptide comprises:
[0023] (a) an amino acid substitution at a position of D10 and / or D386 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ;
[0024] (b) an amino acid substitution at a position selected from the group consisting of 45, 46, 49, 120, 138, 200, 432, 549, 560, 582, 583, 640, 644, 710, and 747 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with Arginine (R) or Gly (G) ;
[0025] (c) an amino acid substitution at a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with P, T, L, or A; or
[0026] (d) a combination of any two or three of (a) , (b) , and (c) .
[0027] In some embodiments, the polypeptide comprises:
[0028] (a) an amino acid substitution of D10A and / or D386A relative to SEQ ID NO: 13;
[0029] (b) an amino acid substitution selected from the group consisting of A45R, G46R, E49R, D120R, Q138R, I200R, F432R, S549R, G560R, F582R, M583R, S640R, E644G, K710R, and T747R relative to SEQ ID NO: 13;
[0030] (c) a combinational substitution of (i) R391P, A393T, I396L, and G400A (hpHNHv22) or (ii) R391D, N397D, and G400A (hpHNHv8) , relative to SEQ ID NO: 13; or
[0031] (d) a combination of any two or three of (a) , (b) , and (c) .
[0032] In some embodiments, the polypeptide comprises:
[0033] (a) D10A or D386A;
[0034] (b) a combinational substitution selected from the group consisting of:
[0035] (1) A45R, G46R, E49R, and T747R;
[0036] (2) E49R, M583R, and T747R;
[0037] (3) E49R, F432R, S640R, and T747R;
[0038] (4) E49R, S549R, M583R, S640R, and T747R;
[0039] (5) E49R, F432R, M583R, and T747R;
[0040] (6) E49R, Q138R, M583R, S640R, and T747R;
[0041] (7) E49R, Q138R, F432R, and T747R;
[0042] (8) E49R, S549R, M583R, and T747R;
[0043] (9) E49R, F432R, M583R, and T747R;
[0044] (10) E49R, I200R, F432R, M583R, and T747R;
[0045] (11) G46R, I200R, F432R, F582R, and T747R;
[0046] (12) G46R, I200R, F582R, and T747R;
[0047] (13) G46R, D120R, I200R, F432R, F582R, and T747R;
[0048] (14) G46R, D120R, I200R, F432R, G560R, F582R, M583R, and T747R;
[0049] (15) G46R, D120R, I200R, F432R, G560R, F582R, and T747R;
[0050] (16) E49R, D120R, F432R, M583R, E644R, and T747R;
[0051] (17) E49R, F432R, G560R, M583R, and T747R;
[0052] (18) E49R, F432R, M583R, E644G, and T747R;
[0053] (19) E49R, F432R, M583R, K710R, and T747R; or
[0054] (20) E49R, Q138R, F432R, M583R, and T747R;
[0055] (c) R391P, A393T, I396L, and G400A (hpHNHv22) ; or
[0056] (d) a combination of any two or three of (a) , (b) , and (c) ,
[0057] relative to SEQ ID NO: 13.
[0058] In some embodiments, the polypeptide comprises (a) D10A; (b) E49R, I200R, F432R, M583R, and T747R; and (c) R391P, A393T, I396L, and G400A (hpHNHv22) , relative to SEQ ID NO: 13.
[0059] In some embodiments, the polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 21-24. In some embodiments, the polypeptide comprises:
[0060] (a) an amino acid substitution at a position selected from the group consisting of D10, D373, H374, N397, H496, and D499 of SEQ ID NO: 9; optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ;
[0061] (b) an amino acid substitution at a position selected from the group consisting of E44, D128, S546, Y568, M569, S623, and E667 of SEQ ID NO: 9; optionally, the amino acid substitution is an amino acid substitution with Arginine (R) ;
[0062] (c) an amino acid substitution at a position of SEQ ID NO: 9 that is corresponding to a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with P, T, L, A, or D; or
[0063] (d) a combination of any two or three of (a) , (b) , and (c) .
[0064] In some embodiments, the polypeptide comprises:
[0065] (a) an amino acid substitution of D10A and / or D373A relative to SEQ ID NO: 9;
[0066] (b) an amino acid substitution selected from the group consisting of E44R, D128R, S546R, Y568R, M569R, S623R, and E667R relative to SEQ ID NO: 9;
[0067] (c) an amino acid substitution relative to SEQ ID NO: 9 that is corresponding to a combinational substitution of (i) R391P, A393T, I396L, and G400A (hpHNHv22) or (ii) R391D, N397D, and G400A (hpHNHv8) relative to SEQ ID NO: 13; or
[0068] (d) a combination of any two or three of (a) , (b) , and (c) .
[0069] In some embodiments, the polypeptide comprises:
[0070] (a) D10A or D373A relative to SEQ ID NO: 9;
[0071] (b) a combinational substitution selected from the group consisting of:
[0072] (1) E44R+ Y568R E667R relative to SEQ ID NO: 9;
[0073] (2) E44R+S546R+Y568R+E667R relative to SEQ ID NO: 9;
[0074] (3) E44R+S546R+M569R+E667R relative to SEQ ID NO: 9;
[0075] (4) E44R+D128R+S546R+Y568R+E667R relative to SEQ ID NO: 9; or
[0076] (5) E44R+D128R+S546R+Y568R+S623R+E667R relative to SEQ ID NO: 9;
[0077] (c) an amino acid substitution relative to SEQ ID NO: 9 that is corresponding to a combinational substitution of R391P, A393T, I396L, and G400A (hpHNHv22) relative to SEQ ID NO: 13; or
[0078] (d) a combination of any two or three of (a) , (b) , and (c) .
[0079] In some embodiments, the polypeptide comprises (a) D10A and (b) E44R+D128R+S546R+Y568R+S623R+E667R, relative to SEQ ID NO: 9.
[0080] In some embodiments, the polypeptide comprises an amino acid sequence of SEQ ID NO: 71 (SN005-D10A+E44R+D128R+S546R+Y568R+S623R+E667R) (enSN005) .
[0081] In another aspect, the disclosure provides a fusion protein comprising the polypeptide of the disclosure and a functional domain.
[0082] In some embodiments, the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a deaminase or a catalytic domain thereof, an uracil glycosylase inhibitor (UGI) , an uracil glycosylase (UNG) , a methylpurine glycosylase (MPG) , a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease or a catalytic domain thereof (e.g., T5 exonuclease) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having RNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment thereof, and any combination thereof.
[0083] In some embodiments, the deaminase or catalytic domain thereof is an adenine deaminase or a catalytic domain thereof (e.g., tRNA adenosine deaminase (TadA) , such as, TadA8e, TadA8.17, TadA8.20, TadA9, TadA8e-V106W, TadA8EV106W+D108Q TadA-CDa, TadA-CDb, TadA-CDc, TadA-CDd, TadA-CDe, TadA-dual, TADAC-1.2, TADAC-1.14, TADAC-1.17, TADAC-1.19, TADAC-2.5, TADAC-2.6, TADAC-2.9, TADAC-2.19, TADAC-2.23, TadA8e-N46L, TadA8e-N46P, TadA* (8.17m) , TadA (8.8m) ) ; or the deaminase or catalytic domain thereof is a cytosine deaminase or a catalytic domain thereof (e.g., an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase, an activation induced deaminase (AID) , a cytidine deaminase 1 from Petromyzon marinus (pmCDA1) , DddA, or a functional variant thereof, e.g., APOBEC1 (rAPOBEC1) , APOBEC2, APOBEC3, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, hAPOBEC3-W104A, CBE6c-V106W) . In some embodiments, the fusion protein further comprises a DNA binding domain; optionally, the DNA binding domain comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 72 or 54 (Sto7d*) .
[0084] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 25, 49-53, 70, and 73; or the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 31 and 64-67.
[0085] In some embodiments, the fusion protein comprises the polypeptide and a reverse transcriptase or a catalytic domain thereof; optionally, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 36. In some embodiments, the fusion protein comprises the polypeptide, a DNMT3l domain, a DNMT3a domain, and a KRAB domain; optionally, the fusion protein comprises, consists essentially of, or consists of a sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to the sequence of SEQ ID NO: 40.
[0086] In an aspect, the disclosure provides a system comprising:
[0087] (1) the polypeptide of the disclosure or the fusion protein of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the polypeptide or the fusion protein, and
[0088] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0089] (i) a scaffold sequence capable of forming a complex with the polypeptide or the fusion protein; and
[0090] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.
[0091] In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48; or the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48; or the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0092] In some embodiments, the guide sequence is in a length of from about 19 to about 24 nucleotides; optionally, the guide sequence is about 20 nucleotides in length.
[0093] In some embodiments, the guide sequence is any one of SEQ ID NOs: 80-260.
[0094] In yet another aspect, the disclosure provides a polynucleotide comprising a sequence encoding the polypeptide of the disclosure or the fusion protein of the disclosure.
[0095] In yet another aspect, the disclosure provides a vector comprising the polynucleotide of the disclosure; optionally, the vector is a plasmid vector, a viral vector (e.g., a recombinant AAV (rAAV) vector, a recombinant lentivirus vector) , a ribonucleoprotein (RNP) , or a lipid nanoparticle (LNP) .
[0096] In yet another aspect, the disclosure provides a cell comprising the polypeptide of the disclosure, the fusion protein of the disclosure, the system of the disclosure, the polynucleotide of the disclosure, or the vector of the disclosure.
[0097] In yet another aspect, the disclosure provides a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.
[0098] In yet another aspect, the disclosure provides a guide RNA comprising a guide sequence of any one of SEQ ID NOs: 80-260 5’ to a scaffold sequence of SpCas9.
[0099] In yet another aspect, the disclosure provides a guide RNA comprising a guide sequence of any one of SEQ ID NOs: 80-260 5’ to a scaffold sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0100] The details of one or more embodiments of the disclosure are set forth in the description below. Other features or advantages of the disclosure will be apparent from the following drawings and detailed description of several embodiments, and also from the appended claims. It is understood that any aspect or embodiment of the disclosure can be combined with any other one or more aspects or embodiments of the disclosure, including aspects or embodiments only described in one sub-section, only in the examples, or only in the claims, to constitute another embodiment explicitly or implicitly disclosed herein unless otherwise indicated.BRIEF DESCRIPTION OF THE DRAWINGS
[0101] An understanding of the features and advantages of the disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure may be utilized, and the accompanying drawings of which:
[0102] Fig. 1 illustrates an exemplary target DNA, and an exemplary system comprising (1) an exemplary guide nucleic acid comprising a guide sequence and a scaffold sequence and (2) an exemplary napDNAbp (programmable RNA-guided DNA endonuclease or mutant thereof in the disclosure) .
[0103] Fig. 2 shows endonuclease activity of SN007 with various lengths of guide sequences as indicated by fluorescent signals.
[0104] Fig. 3 shows endonuclease activity of SN007 with various lengths of guide sequences as indicated by NGS sequencing.
[0105] Fig. 4 shows endonuclease activity and nickase activity of SN007 and mutants thereof vs. SpCas9 and nickases thereof, as well as negative control (NT) .
[0106] Fig. 5 shows nickase activity of SN007 mutants (with single mutation) vs. SpCas9 nickase (H840A) and SN007 nickase (D386A) .
[0107] Fig. 6 | Comparison of the specificities of diverse Cas and IscB, and identification of functional Cas. a, Phylogenetic tree of IscB and representative CRISPR-Cas systems. Systems selected for this study are marked with black dots, newly identified Cas with red dots and previously reported Cas with blue dots. b-f, Specificity of the IscB orthologue OgeuIscB (b) , type II-Aorthologues SpCas9 (c) and SlugCas9 (d) , type II-C orthologue Nme2Cas9 (e) , and type II-D orthologue MG34-16 (f) was assessed in human cells using a GFP activation reporter (Fig. 13a) . Editors were programmed with either a perfectly matched sgRNA (wt) or single‐mismatch sgRNAs (sm; Fig. 13b) , and efficiencies with mismatched guides were normalized to the wt sgRNA (set to 1) . Data are mean ± s.d. (n = 3) . g, Three-dimensional plot of protein size (x axis) , editing efficiency (y axis) , and specificity (z axis) for IscB and representative Cas. A golden cube in the upper middle corner marks the ideal region representing compact Cas scaffolds (<800 amino acids) that achieve both high editing efficiency (>60%) and high specificity (>90%) . h, GFP activation efficiencies mediated by Cas candidates, SN001 and SN003-SN007, as measured by flow cytometry. The most efficient Cas, SN007 (MM762) , is highlighted in bold. i, Specificity of MM762 assessed by mismatch tolerance assay. Data are mean ± s.d. (n = 3) .
[0108] Fig. 7 | Cryo-EM structure of the MM762-sgRNA-target DNA ternary complex. a, Domain architecture of MM762, highlighting the positions of the RuvC, BH, REC, HNH, WED, and PI domains. b, Cryo-EM density map (left) and corresponding atomic model (right) of the MM762-sgRNA-target DNA ternary complex. c, Close-up view of the PI and WED domains of MM762 engaging the 3′-TGG PAM and its reverse complementary CCA sequence. Key residues mediating specific recognition of the GG / CC dinucleotides (K626, N664, and K728) are in bold. d, Close-up view of the REC domain engaging the sgRNA scaffold, highlighting the residues and nucleotides that form stabilizing contacts shown as sticks. e, Zoomed-in view of the extensive contacts between MM762 residues and the sgRNA-DNA heteroduplex. f, Schematic of sgRNA. g, Detailed view of the sgRNA tertiary structure.
[0109] Fig. 8 | Directed engineering of MM762 to enhance efficiency for base editor development. a, Schematic of the MM762 protein engineering workflow. MM762 mutant library was generated from the wild-type MM762 (D386A) nickase by substituting non-positively charged residues (X) with arginine (R) , and variants exhibiting enhanced GFP activation were selected for further optimization. b, Single‐mutation MM762 variants were screened for enhanced activity. SpCas9 (H840A) (green) and wild-type MM762 (D386A) (light red) are indicated; variants selected for subsequent combinatorial engineering are highlighted with boxes color-coded according to their corresponding protein regions. Each dot denotes the activity of an individual variant. c, Combinatorial assembly of MM762 variants incorporating selected single amino-acid substitutions. Mutations included in each combination are indicated by colored squares corresponding to regions shown in Fig. 3b. Editing efficiencies of the resulting D10A nickase variants were measured using the GFP activation reporter; data represent mean ± s.d. (n = 3) . d, Schematic of the construction strategy for the development of eMM762ABEv1. e, Overview of PAM / TAM-matched target sites used for comparative analysis of editing efficiencies. f, Violin plot of A·T-to-G·C editing efficiencies for ABE8e variants derived from IminiIscB, eMM762v1, MG34-16, SpCas9, SlugCas9-NNG and eNme-C at endogenous HEK293T loci. Each point represents the mean of the highest editing efficiency observed per locus. Violin width indicates density; the central bar denotes the median, and the thinner bars mark the first and third quartiles. Plots are truncated at the observed minima and maxima. *P <0.05; ***P <0.001; ****P < 0.0001; ns = not significant, P ≥ 0.05; statistical significance determined by one-way ANOVA followed by Tukey’s multiple comparisons test.
[0110] Fig. 9 | Integrated engineering strategies to optimize the efficiency of MM762-based compact base editors. a, Schematic of Sto7d*fusions to eMM762ABEv1. The mutated Sto7d (Sto7d*, K12L) DNA-binding domain was appended to the N-terminus of the deaminase (nSto7d*) , to the C-terminus of eMM762v1 (cSto7d*) , or inserted between the deaminase and eMM762v1 (inSto7d*) . b, Comparison of A·T-to-G·C conversion efficiencies mediated by eMM762ABEv1, eMM762ABEv1-nSto7d* (eMM762ABEv2) , eMM762ABEv1-cSto7d*and eMM762ABEv1-inSto7d*at the DMD E45, E50 and E55 loci. c, Computational design of a high-performance HNH domain (hpHNH) in MM762 using Funclib. The MM762 structure is shown with the HNH and RuvC nuclease domains coloured light red and blue, respectively. The inset highlights the HNH regions targeted for computational mutagenesis and outlines the design strategy. d, Editing efficiencies of two hpHNH variants, engineered on the eMM762v11-ABE8e backbone, as measured by the BE-GFP activation reporter. Data are mean ± s.d. (n = 3) . e, Schematic of eMM762ABEv3~v5 development. eMM762ABEv3, derived from eMM762ABEv2, substitutes eMM762v1 with eMM762v2; eMM762ABEv4 introduces the hpHNHv22 mutation into eMM762ABEv3; and eMM762ABEv5 replaces the XTEN linker in eMM762ABEv4 with an optimized connector between the deaminase and eMM762v2. f, A·T-to-G·C editing observed at ATXN2, B2M, and TRAC loci in HEK293T cells. g, Violin plot of A·T-to-G·C editing efficiencies achieved by eMM762ABEv1, eMM762ABEv3, eMM762ABEv5, SpCas9ABE8e and HiFiABE8e across ten target sites in HEK293T cells. Each dot represents the mean of the highest editing efficiency observed at each locus across three independent biological replicates. h, Diagram of the MM762 sgRNA structure, with base pairs chosen for mutation to G-C highlighted. i, Editing efficiencies of sgRNA variants and wild-type sgRNA in combination with eMM762ABEv3 at the DMD E50, ATXN2 and CD52 loci in HEK293T cells. j, Schematic of the development of eMM762CBEv1~v4. k, Comparison of C·G-to-T·A conversion efficiencies induced by eMM762CBEv1~v4 and SpCas9CBE6c. l, Violin plot of efficiencies achieved by eMM762CBEv1, eMM762CBEv4, and SpCas9CBE6c across ten target loci in HEK293T cells. All heat maps depicting editing rates across the entire protospacer, averaged over three independent experiments. All violin width reflects the density of observations at each editing level; the central bar marks the median, the thinner bars denote the first and third quartiles, and the plot is truncated at the observed minimum and maximum values. **P <0.01; ***P <0.001; statistical significance determined by one-way ANOVA followed by Tukey’s multiple comparisons test.
[0111] Fig. 10 | Characterizing the off-target effects of engineered MM762 and eMM762ABE in mammalian cells. a, GUIDE-seq profiling of on-and off-target cleavage by eMM762v2 and SpCas9 at endogenous loci in HEK293T cells. GUIDE-seq read counts for each site (shown at right) reflect cleavage frequency; mismatches within the spacer or PAM are highlighted in color. b, Summary of GUIDE-seq analysis of eMM762v2 and SpCas9 at five endogenous sites. The percentage of off-target sequencing reads relative to total reads is shown for each nuclease. c, ON / OFF ratios of GUIDE-seq reads across five genomic sites, calculated as log2 [ (ON-target reads + 1) / (OFF-target reads + 1) ] . N = 5. d, On-target (ON) and off-target (OT) editing efficiencies of eMM762v2, SpCas9, HiFi Cas9 and eSpCas9 (1.1) proteins directed to VEGFA and FANCF loci in HEK293 cells. Data are mean ± s.d. (n = 3) . e, Specificity and efficiency scores summarized across seven tested endogenous loci for eMM762v2, SpCas9, HiFi Cas9 and eSpCas9 (1.1) . f, A·T-to-G·C conversion efficiencies at ABE-induced off-target sites, shown alongside the on-target site at the HEK293 site4 locus in HEK293T cells. g, Normalized off-target editing efficiencies of eMM762ABEv3, SpCas9ABE8e, and HiFiABE8e. Editing efficiencies at each off-target site (as shown in Fig. 5f) were normalized to those of eMM762ABEv3 (set as 1.0 for each site) . Each dot represents a single off-target site. h, Specificity and efficiency scores of eMM762ABEv3, SpCas9ABE8e, and HiFiABE8e. *P <0.05; **P <0.01; ***P <0.001; ****P < 0.0001; ns = not significant, P ≥ 0.05; statistical significance determined by one-way ANOVA followed by Tukey’s multiple comparisons test.
[0112] Fig. 11 | Assessment of multiplex editing efficiency and fidelity of eMM762ABE and eMM762CBE in primary human T cells. a, Schematic diagram of multiplex base editing in human T cells. Base editor mRNA and multiple sgRNAs targeting distinct genomic loci are co-delivered by electroporation to achieve simultaneous editing. Targeted loci are indicated alongside corresponding sgRNAs. b, A·T-to-G·C conversion efficiencies (left y axis) and indels frequencies (right y axis) induced by eMM762ABEv3 and SpCas9ABE8e in primary human T cells, with base editing rates shown at splice site positions. Data represent mean values from two independent donors (n = 2) . c, C·G-to-T·A editing efficiencies (left y axis) and indels frequencies (right y axis) induced by eMM762CBEv1 and SpCas9CBE6c in primary human T cells, with base editing rates shown at splice site positions. Data represent mean values from two independent donors (n = 2) . d, Protein reduction mediated by adenine base editors, as measured by flow cytometry at three target sites. e, Protein reduction achieved using cytosine base editors, as assessed by flow cytometry at three target loci. f, Off-target analysis of eMM762ABEv3 compared with SpCas9ABE8e at the B2M locus in primary human T cells. g, Comparative off-target profiling of eMM762CBEv1 and SpCas9CBE6c at the B2M locus in primary human T cells, assessed by targeted deep sequencing. Data represent mean values from two independent donors (n = 2) . ns = not significant, P ≥ 0.05; statistical significance determined by one-way ANOVA followed by Tukey’s multiple comparisons test.
[0113] Fig. 12 | Therapeutic applications of eMM762ABE in humanized mdx mice. a, A·T-to-G·C conversion efficiencies of eMM762ABEv3, SpCas9ABE8e and HiFiABE8e at the DMD E50 on-target site and 19 in silico–predicted off-target loci in HEK293T cells, as determined by targeted deep sequencing. Target site sequences are shown with the spacer region and PAM indicated by a red line, and splice site positions highlighted in bold red. Bars indicate editing efficiencies at the most frequently edited adenine for each locus, with the corresponding position (A#) shown in parentheses. b, Schematic of AAV vector designs for SpCas9ABE8e and eMM762ABE variants. eMM762ABE variants, together with their sgRNA expression cassette, were packaged into a single AAV vector. SpCas9ABE8e was split using a Rhodothermus marinus (Rma) intein, with the editor components and sgRNA delivered in two separate AAV particles. A muscle-specific Spc5.12 promoter was used to drive base editor expression. c, Overview of the in vivo intramuscular delivery of AAV–ABE constructs into the tibialis anterior muscle of the right leg of 3-week-old DMDΔmE5051, KIhE50 / Y mice. The left leg was injected with saline as a control. Black arrows indicate time points for tissue collection following injection. d, In vivo A·T-to-G·C editing efficiencies at position A7 of the DMD E50 locus mediated by eMM762ABEv3 and SpCas9ABE8e. e, Exon skipping efficiencies determined by deep sequencing of total mRNA extracted from muscle tissue. f, Quantification of dystrophin-positive (Dys+) fibers in cross-sections of tibialis anterior (TA) muscles. g, Immunofluorescence analysis comparing dystrophin (Dys+, green) expression restored by eMM762ABEv3 and SpCas9ABE8e systems; spectrin (purple) serves as a control. Scale bar, 100 μm. h, Western blot analysis of dystrophin expression in TA muscles 6 weeks post-injection with ABEs or saline; vinculin was used as a loading control. i, Quantification of dystrophin expression based on Western blot signal normalized to vinculin. Data are mean ± s.d. (n = 3) . ns = not significant, P ≥ 0.05; ****P < 0.0001; statistical significance determined by one-way ANOVA followed by Tukey’s multiple comparisons test.
[0114] Fig. 13 | evaluation of fidelity of programmable RNA-guided DNA endonucleases by using mismatch tolerance assay in human cells. a, Schematic of the GFP activation reporter assay used to assess dsDNA cleavage (endonuclease) activity. The reporter encodes a BFP–T2A–EGFxxFP cassette in which EGFP is split into two fragments-EGFx (CDS 1–561 bp) and xFP (CDS 112–720 bp) -separated by a short insertion carrying the target sequence flanked by the appropriate TAM or PAM. Cas-or IscB-triggered double-strand breaks at the insertion site initiate single-strand annealing (SSA) -mediated repair of EGFxxFP, restoring EGFP expression. b, Schematic illustrating the design of single‐mismatched guide RNAs against the indicated target sequence. c, Mismatch tolerance of the high‐fidelity SpCas9 variant HiFi Cas9 assessed using GFP activation reporter in human cells. Data are mean ± s.d. (n = 3) .
[0115] Fig. 14 | Comparison of Cas and IscB. Features of OgeuIscB and Cas proteins.
[0116] Fig. 15 | PAM analysis of different Cas proteins. PAM of Cas proteins determined by in vitro plasmid depletion assay.
[0117] Fig. 16 | Preliminary optimization of the MM762 sgRNA scaffold, and characterization of MM762. a, Structure of the MM762 sgRNA, in which the crRNA-tracrRNA duplex is designated as stem-loop 1 (SL1) and highlighted in red for truncation analysis. b, Editing efficiencies of truncated sgRNA variants co-transfected with MM762 in HEK293T cells, measured using a single-target GFP activation reporter. Data represent mean (n = 2) . c, Mismatch tolerance of the MM762 and SpCas9 assessed using a GFP activation reporter system with sgRNAs containing double mismatches at varying positions. Data are mean ± s.d. (n = 3) . d-e, Optimal spacer length for MM762, as determined by a GFP activation reporter (d) and by amplicon sequencing at an endogenous locus in HEK293T cells (e) . Data are mean ± s.d. (n = 3) . f, Diagram of the GFP‐based reporter system configured to measure nickase activity of the editors. g, Impact of defined mutations on the nickase editing efficiency of MM762. Single-target and dual-target GFP reporter assays were employed to evaluate DSB cleavage and nickase activities, respectively. Data are mean ± s.d. (n = 3) . h, Editing efficiencies of MM762ABE and MGABE-derived from MM762 and MG34-16, respectively-compared to SpCas9ABE8e at the DMD E50 locus. Data represent mean (n = 2) . i, Efficiency of MM762-based CBE at the DMD E50 locus. Data represent mean (n = 2) .
[0118] Fig. 17 | Purification and activity of MM762, and Cryo-EM data processing workflows for wtMM762 (D10A) -sgRNA-target DNA and eMM762v2 (D10A) -sgRNA-target DNA complexes. a, Size-exclusion chromatography profile of MM762 (left) and SDS–PAGE of peak fractions (right) , showing purified MM762 protein and denaturing gel of the corresponding sgRNA. b, Native PAGE analysis of MM762-sgRNA complex-mediated cleavage of a 708 bp linear dsDNA substrate containing a PAM and 20 bp target sequence (n =4) . c, Cryo-EM processing workflow for wtMM762 (D10A) -sgRNA-target DNA complex. Shown are a representative raw micrograph, selected 2D class averages, 3D reconstruction and refinement steps. A red arrow highlights a region with improved density in one 3D class. Local resolution estimation and FSC curve are shown. d, Cryo-EM processing workflow for eMM762v2 (D10A) -sgRNA-target DNA complex. Shown are raw micrograph, 2D class averages, 3D reconstructions and refinements, as well as local resolution map, FSC curve and angular distribution plot.
[0119] Fig. 18 | Characterization of base editing activity window. a, Schematic of the construction strategy. b~f, Base editing activity windows of lminiIscB-ABE8e (b) , eMM762ABEv1 (c) , MG34-16-ABE8e (d) , SlugCas9-NNG-ABE8e (e) and eNme2-C-ABE8e (f) , showing pooled A·T-to-G·C conversion efficiencies across all protospacer positions (the PAM-distal base of protospacer defined as position 1) . The dotted line marks the threshold for the editing window, set at 25%of the mean peak activity for each base editor. Positions not represented in any of the tested sites were shaded in grey. Each point represents the percentage conversion observed for an adenine at that position of one protospacer.
[0120] Fig. 19 | Fusion of Sto7d*to enhance eMM762ABEv1 activity. Editing efficiencies of eMM762ABEv1 and Sto7d*fusion variants were assessed at three endogenous loci in HEK293T cells (related to Fig. 4b) . Each dot denotes the mean of the highest efficiency observed per locus across three independent biological replicates. Data are mean ± s.d. (n = 3) .
[0121] Fig. 20 | Computational engineering for high-performance HNH domain of MM762. a, Sequences of the top 29 hpHNH designs ranked by total protein-energy scores. b, Schematic of the BE-GFP activation reporter assay used to quantify base editing activity of editor constructs. c, Editing efficiencies of hpHNH variants incorporated into eMM762v11-ABE8e, as measured by the BE-GFP activation reporter. Data represent mean (n = 2) .
[0122] Fig. 21 | Screening of combinatorial eMM762 variants and linker optimization for enhanced base editing. a, Schematic of the construction strategy. b, Editing efficiencies of ABE8e variants derived from eMM762v# (v1–v12) , measured using a BE–GFP activation reporter. MFI, mean fluorescence intensity. Data represent mean (n =2) . c, Cryo-EM structures of the eMM762v2-sgRNA-target DNA ternary complex (left) compared with the wild-type MM762-sgRNA-target DNA ternary complex (right) . Mutations in eMM762v2 are highlighted in red. d, Linker sequences evaluated. e, Schematic of constructs for linker testing. f, A·T-to-G·C editing efficiencies of base editors with different linkers at the DMD E50 A7 position in HEK293T cells, quantified by EditR. Data represent mean (n = 2) .
[0123] Fig. 22 | Comparison of eMM762ABE variants, SpCas9-ABE8e, and HiFiABE8e. a, Editing efficiencies of eMM762ABE variants compared with SpCas9-ABE8e and HiFiABE8e in HEK293T cells (heat maps shown in Fig. 3f) . Data are mean ± s.d. b, Heat maps show A·T-to-G·C editing efficiencies across the protospacer in HEK293T cells, averaged from three independent biological replicates. c, Indel frequencies induced by eMM762ABE variants, SpCas9ABE8e, and HiFiABE8e at the same genomic sites shown in Fig. 3g. Corresponding A·T-to-G·C editing efficiencies are presented in Fig. 4g.
[0124] Fig. 23 | Engineering sgRNA scaffold to increase activity. a, Sequence alignment of the wild-type sgRNA and sgRNA variants. Dots indicate positions identical to the wild-type sequence; hyphens denote nucleotide deletions; substituted bases are highlighted in colour. b, Schematic of the sgRNA scaffold engineering strategy. c, Editing efficiencies of truncated sgRNA variants co-transfected with eMM762v2 (D10A) nickase, as measured by a dual-target GFP activation reporter. The sgRNA SL1-Δ14 is shown in yellow, and its efficiency is indicated by the red dashed line. Data are mean ± s.d. (n = 3) . d, A·T-to-G·C editing efficiencies induced by G-C pair substitution sgRNA variants co-transfected with eMM762ABEv2 at the DMD E50 locus in HEK293T cells, showing editing at position A7. The red dashed line denotes the sgRNA SL1-Δ14 efficiency, and variants selected for further validation in Fig. 4h, i are highlighted in purple. Data represent mean (n = 2) . e, Violin plot summarizing the highest editing efficiencies presented in Fig. 4i.
[0125] Fig. 24 | Comparison of eMM762CBEv1, eMM762CBEv4, and SpCas9CBE6c. Heat maps show C·G-to-T·Aediting efficiencies at eight endogenous loci in HEK293T cells, averaged over three independent biological replicates.
[0126] Fig. 25 | Off-target assessment of engineered MM762 and eMM762ABE. a, In addition to the GUIDE-seq data shown in Fig. 5. b, On-target (ON) and off-target (OT) editing efficiencies of eMM762v2, SpCas9, HiFi Cas9 and eSpCas9 (1.1) proteins in HEK293 cells. Data are mean ± s.d. (n = 3) . c, A·T-to-G·C conversion rates at ABE-induced off-target sites in HEK293T cells. d, Schematic of the orthogonal R-loop assay used to evaluate Cas-independent off-target deamination by base editors. e, Cas-independent off-target activity of eMM762ABEv1, eMM762ABEv2 and SpCas9ABE8e at five R-loops generated by dSaCas9. N. D. = Not Detected. Data are mean ± s.d. (n = 3) .
[0127] Fig. 26 | Characterizing the fidelity of eMM762ABE and eMM762CBE in T-cells. a, Targeted deep sequencing–based analysis of off-target editing by eMM762ABEv3 and SpCas9ABE8e at the CD52 locus in primary human T cells. b, Off-target editing comparison between eMM762CBEv1 and SpCas9CBE6c at the TRAC locus in primary human T cells, evaluated by targeted deep sequencing. Data represent mean values from two independent donors (n = 2) .
[0128] Fig. 27 | Evaluating the efficiency and fidelity of eMM762ABE in mice. a, A·T-to-G·C conversion efficiencies of eMM762ABEv3, SpCas9ABE8e and HiFiABE8e at 11 in silico–predicted off-target loci of the on-target DMD E50 site in N2A cells, as determined by targeted deep sequencing. Data are mean ± s.d. (n = 3) . b, Schematic of the treatment strategy in DMDΔmE5051, KIhE50 / Y mice, in which mouse Dmd exons 50 and 51 are deleted and replaced with human DMD exon 50. eMM762ABE targets and disrupts the conserved adenine within the complementary sequence of the splice donor site (GU) , promoting programmable skipping of exon 50 and enabling restoration of dystrophin expression. c, Gel electrophoresis of RT-PCR products from treated DMDΔmE5051, KIhE50 / Y mice. d, In vivo A·T-to-G·C editing efficiencies at position A7 of the DMD E50 locus mediated by eMM762ABE variants and SpCas9ABE8e. e, Comparison of exon skipping efficiencies among eMM762ABEv1, eMM762ABEv2, eMM762ABEv3 and SpCas9ABE8e. f, Immunofluorescence analysis of dystrophin (Dys+, green) expression restored by eMM762ABEv1 and eMM762ABEv2; spectrin (purple) serves as a control. Scale bar, 100 μm. g, Quantification of dystrophin-positive (Dys+) fibers in tibialis anterior (TA) muscle cross-sections. h, Western blot analysis of dystrophin expression in TA muscles 6 weeks after injection with eMM762ABEv1, eMM762ABEv2 or saline; vinculin served as a loading control. i, Densitometric quantification of dystrophin expression from Western blots, normalized to vinculin levels.
[0129] Fig. 28 shows the fold-change of base editing efficiency of multiple ABEs containing MM745 mutants with one or more substitutions with R relative to the ABE (SEQ ID NO: 70) containing MM745-D10A (SEQ ID NO: 69) .
[0130] Fig. 29 shows the base editing efficiency of multiple ABEs containing MM745 mutants with one or more substitutions with R, and with or without DNA binding domain (SEQ ID NO: 72) , relative to the ABE (SEQ ID NO: 70) containing MM745-D10A (SEQ ID NO: 69) .
[0131] Fig. 30 shows the base editing efficiency of ABE (SEQ ID NO: 73) containing both DBD (SEQ ID NO: 72) and enSN005 nickase (SEQ ID NO: 71) at endogenous loci of PDCD1, B2M, or CD52.
[0132] Fig. 31 shows guide sequences and primers used in the examples.
[0133] Fig. 32 shows guide sequences and primers used in the examples.
[0134] Fig. 33 shows guide sequences and primers used in the examples.
[0135] Fig. 34 shows guide sequences and primers used in the examples.
[0136] Fig. 35 shows guide sequences and primers used in the examples.
[0137] Fig. 36 shows guide sequences and primers used in the examples.
[0138] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.DETAILED DESCRIPTION
[0139] The disclosure will be described with respect to particular embodiments, but the disclosure is not limited thereto in any respect. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which this disclosure belongs. Terms as set forth hereinafter are generally to be understood in their plain and ordinary meaning or common sense unless indicated otherwise.
[0140] Definition
[0141] The disclosure will be described with respect to particular embodiments, but the disclosure is not limited thereto in any respect. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which this disclosure belongs. Terms as set forth hereinafter are generally to be understood in their plain and ordinary meaning or common sense unless indicated otherwise.
[0142] Similar to programmable RNA-guided DNA endonucleases Cas9, Cas12, and IscB, the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure are capable of binding to a target DNA (e.g., a dsDNA) as guided by a guide nucleic acid (e.g., a guide RNA) comprising a guide sequence targeting the DNA. The programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure may be associated with the guide nucleic acid (e.g., a guide RNA) , which localizes / targets the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure to a target DNA that comprises a DNA strand (i.e., a target strand) that is reversely complementary to the guide nucleic acid, or a portion thereof (e.g., the guide sequence of a guide RNA) . In other words, the guide nucleic acid is “programed” to localize and bind the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure to the target DNA. Binding of the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure to the target DNA enables the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure or a construct comprising the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure to access to and function on the target DNA. For this purpose, the guide nucleic acid comprises a scaffold sequence responsible for (capable of) forming a complex with the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure, and a guide sequence that is intentionally designed to be responsible for (capable of) hybridizing to a target sequence of the target DNA, thereby guiding the complex comprising the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure and the guide nucleic acid to the target DNA such that the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure is indirectly bound to the target DNA. The ability of the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure to be bound to a target DNA by being guided by such a guide nucleic acid makes the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure a nucleic acid programmable DNA binding protein (napDNAbp) or nucleic acid programmable DNA binding domain (napDNAbd) , similar to Cas9, Cas12, and IscB.
[0143] Referring to Fig. 1, an exemplary dsDNA is depicted to comprise a 5’ to 3’ single DNA strand and a 3’ to 5’ single DNA strand, the 5’ to 3’ single DNA strand comprises an exemplary first deoxyribonucleotide dA, and the 3’ to 5’ single DNA strand comprises an exemplary second deoxyribonucleotide dT that base pairs with the dA.
[0144] An exemplary guide nucleic acid is depicted to comprise a guide sequence and a scaffold sequence. The guide sequence is designed according to base pairing principle to be capable of hybridizing to a part of the 3’ to 5’ single DNA strand, and so the guide sequence “targets” that part. And thus, the 3’ to 5’ single DNA strand is referred to as a “target strand (TS) ” of the dsDNA, while the opposite 5’ to 3’ single DNA strand is referred to as a “nontarget strand (NTS) ” of the dsDNA. That part of the target strand based on which the guide sequence is designed and to which the guide sequence is capable to hybridize is referred to as a “target sequence” , while the opposite part on the nontarget strand corresponding to that part is referred to as the “protospacer sequence” , which is typically 100%(fully) reversely complementary to the target sequence, if there is no intentional or unintentional mismatch. Generally, as is conventional in the art, a nucleic acid sequence (e.g., a DNA sequence) is written in 5’ to 3’ direction / orientation unless explicitly indicated otherwise.
[0145] For example, for a DNA sequence of ATGC, it is usually understood as 5’ -ATGC-3’ unless otherwise indicated. Its reverse sequence is 5’ -CGTA-3’ . Its fully complementary sequence is 5’ -TACG-3’ . Its fully reverse complementary sequence is 5’ -GCAT-3’ (3’ -TACG-5’ ) . Note that the fully complementary sequence usually does not have the ability to base-pair / hybridize with the original sequence.
[0146] Generally, the double-strand sequence of a dsDNA may be represented with the sequence of its 5’ to 3’ single DNA strand conventionally written in 5’ to 3’ direction / orientation unless otherwise indicated.
[0147] For example, for a dsDNA having a 5’ to 3’ single DNA strand of 5’ -ATGC-3’ and a 3’ to 5’ single DNA strand of 3’ -TACG-5’ as shown below, the dsDNA may be simply represented as 5’ -ATGC-3’ .
[0148] 5’-----ATGC -----3’
[0149] 3’-----TACG -----5’
[0150] It should be noted that either the 5’ to 3’ single DNA strand or the 3’ to 5’ single DNA strand of a dsDNA can be a nontarget strand from which a protospacer sequence is selected.
[0151] In the sense of base editing, the strand on which the target nucleotide (e.g., the deoxyribonucleotide dA in Fig. 1) to be edited is located is termed as an edited strand, and the opposite strand is termed as a non-edited strand. As used herein, the nontarget strand is the edited strand, and the target strand is the non-edited strand.
[0152] Typically for a gene, the 5’ to 3’ single DNA strand of the gene is sense strand, and the 3’ to 5’ single DNA strand of the gene is antisense strand. Either the sense strand or the antisense strand can be a nontarget strand from which a protospacer sequence is selected.
[0153] To hybridize to a dsDNA, such as, a dsDNA 5’ -ATGC-3’ , the guide sequence of a gRNA can be in one embodiment designed to have a sequence of 5’ -AUGC-3’ that is fully reversely complementary to the 3’ to 5’ strand of the dsDNA (3’ -TACG-5’ ) , which would be set forth in ATGC in the electric sequence listing and marked as an RNA sequence according to WIPO standard ST. 26; and in another embodiment, the guide sequence of a gRNA can be designed to have a sequence of 5’ -GCAU-3’ that is fully reversely complementary to the 5’ to 3’ strand of the dsDNA (5’ -ATGC-3’ ) , which would be set forth in GCAT in the electric sequence listing and marked as an RNA sequence according to WIPO standard ST. 26.
[0154] In the case that the guide sequence of a gRNA is fully reversely complementary to the target sequence of a target DNA and the target sequence of a target DNA is fully reversely complementary to the protospacer sequence of the target DNA, the guide sequence of a gRNA is identical to the protospacer sequence of a target DNA except for the difference between the U in the guide sequence due to its RNA nature and the corresponding T in the protospacer sequence due to its DNA nature. According to WIPO standard ST. 26, symbol “t” is used to denote both T in DNA and U in RNA (See “Table 1: List of nucleotides symbols” , the definition of symbol “t” is “thymine in DNA / uracil in RNA (t / u) ” ) . Thus, in the electronic sequence listing of the disclosure prepared according to WIPO standard ST.26, such a guide sequence of a gRNA could be set forth in the same sequence as a corresponding protospacer sequence of a target DNA in the same length. For convenience, a single SEQ ID NO entry in the electronic sequence listing can be used to denote both such a guide sequence of a gRNA and a protospacer sequence of a target DNA, despite whether the SEQ ID NO entry is marked as DNA or RNA in the electronic sequence listing. When a reference is made to such a SEQ ID NO entry that sets forth a protospacer / guide sequence, it refers to either a protospacer sequence that is a DNA sequence or a guide sequence of a gRNA that is an RNA sequence depending on the context, no matter whether it is marked as a DNA or an RNA in the electronic sequence listing. As used herein, if a DNA sequence, for example, 5’ -ATGC-3’ is transcribed to an RNA sequence, with each dT (deoxythymidine, or “T” for short) in the primary sequence of the DNA sequence replaced with a U (uridine) and each dA (deoxyadenosine, or “A” for short) , dG (deoxyguanosine, or “G” for short) , and dC (deoxycytidine, or “C” for short) replaced with A (adenosine) , G (guanosine) , and C (cytidine) , respectively, resulting in 5’ -AUGC-3’ , it is said in the disclosure that the DNA sequence “encodes” the RNA sequence.
[0155] As used herein, the term “polypeptide” and “protein” are used interchangeably to refer to a polymer of amino acids of any length. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass a polymer of amino acids that has been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component.
[0156] As used herein, the polypeptide of the disclosure may refer to any polypeptide in the disclosure, for example, a wild type polypeptide in the disclosure, a mutant of a reference polypeptide in the disclosure, and more specifically, any one of the programmable RNA-guided DNA endonucleases of the disclosure and any mutant thereof, such as, a nickase, a dead version (an endonuclease-deficient mutant) , or a high-efficiency mutant of any one of the programmable RNA-guided DNA endonucleases of the disclosure.
[0157] As used herein, the term “reference polypeptide” is used in the context of designing and developing a new polypeptide based on an original polypeptide (e.g., a wild-type polypeptide) . For example, the original polypeptide is mutated to generate the new polypeptide. In that case, the original polypeptide is a reference of the new polypeptide and termed as a reference polypeptide. The properties of a new polypeptide can be evaluated with the reference polypeptide as a reference from which the new polypeptide is derived. For example, one or more of the properties (e.g., endonuclease activity, nickase activity) of the new polypeptide can be compared with the reference polypeptide from which the new polypeptide is derived.
[0158] With respect to a polypeptide in the disclosure, the terms “variant” , “mutant” , and “engineered polypeptide” as used herein are used interchangeably to refer to a mutant of a reference polypeptide (e.g., a wild type polypeptide) generated by introducing an amino acid mutation into the reference polypeptide.
[0159] The amino acid sequence of a protein often starts with a most N-terminal Methionine (M) (i.e., at position 1) , since the coding sequence for the protein would require a 5’ end start codon ATG to initiate its transcription and translation, and the start codon ATG encodes amino acid Met. If a start codon ATG is separately present upstream of the coding sequence for a polypeptide with a most N-terminal Met at position 1, then the codon for the Met on the coding sequence may be omitted as needed and hence the Met of the polypeptide is deleted. Usually, such deletion of Met at position 1 would not affect the function of the polypeptide. In another view, the Met encoded by the upstream start codon ATG and the polypeptide lacking most N-terminal Met can together be regarded as the corresponding polypeptide having the N-terminal Met. In some embodiments, the polypeptide of the disclosure comprises a deletion of the most N-terminal Methionine (M) relative to a reference polypeptide. For convenience, when reference is made to any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 or a mutant thereof or any other polypeptide in the disclosure, if Met is present at position 1 of any of those referred polypeptides, the reference is also made to an N-terminal truncation of any of those referred polypeptides lacking the most N-terminal Methionine (M) . For convenience, when reference is made to any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 or a mutant thereof or any other polypeptide in the disclosure, if Met is not present at position 1 of any of those referred polypeptides and start codon ATG is essential for its expression, the reference is also made to the polypeptide with addition of an most N-terminal Methionine (M) at N-terminal of the amino acid at position 1 of the polypeptide.
[0160] As used herein, the description of “a mutant comprising an amino acid mutation (e.g., substitution) at a position that is corresponding to a given position (e.g., D10) or that is a given position (e.g., D10) of a given polypeptide (e.g., SN007) ” or similar description means that the mutant is a mutant of the given polypeptide (as a reference polypeptide) or another different polypeptide (as a reference polypeptide) and comprises an amino acid mutation at a position of the amino acid sequence of the mutant corresponding to the given position of the amino acid sequence of the given polypeptide.
[0161] The position of the mutation in the mutant may be the same as the given position of the given polypeptide, for example, when the mutant is a mutant of the given polypeptide and has exactly the same length as the given polypeptide.
[0162] The position of the amino acid mutation in the amino acid sequence of the mutant may be different from the given position of the given polypeptide, for example, when the mutant is a mutant of the given polypeptide but does not have exactly the same length as the given polypeptide, for example, when the mutant comprises a N-terminal truncation as compared with the given polypeptide and thus the first N-terminal amino acid of the mutant is not corresponding to the first N-terminal amino acid of the given polypeptide but to an internal amino acid within the given polypeptide, but the position of the mutation in the mutant can be determined by alignment of the mutant and the given polypeptide to identify the corresponding amino acids in the two sequences as understood by a skilled in the art. For example, if the mutant has a N-terminal truncation of 20 amino acids as compared with the given polypeptide, then the mutant comprising an amino acid mutation at a position corresponding to D10 of a given polypeptide means that the mutant comprises an amino acid mutation at position 41 of the mutant since position 41 in the mutant is corresponding to D10 in the given polypeptide as determined by alignment of the mutant and the given polypeptide. Alternatively, the mutant may be a mutant of another polypeptide (as a reference polypeptide) different from the given polypeptide but said another reference polypeptide is structurally similar to the given polypeptide and therefore the position of the mutation of the mutant can be determined according to the given position of the given polypeptide, for example, by sequence alignment of the mutant (or said another polypeptide) and the given polypeptide. For example, the mutant is mutated from a first reference polypeptide (e.g., SN001) comprises an amino acid mutation (e.g., substitution) at an XXX position of the first reference polypeptide corresponding to YYY position (e.g., D10) of a second reference polypeptide (e.g., SN007) , where the first and the second reference polypeptides are not the same but structurally similar, for example, the first and the second reference polypeptides are conservative at one or more amino acid residues. For example, SN001 and SN007 in the disclosure are two different polypeptides but structurally similar and conservative at D10 (numbered according to the sequence of SN007) . Therefore, with respect to a mutant of SN001 comprising an amino acid mutation at a position corresponding to D10 of SN007, the position of the mutation of the SN001 mutant can be determined to be position D10 of SN001 by sequence alignment of SN001 and SN007.
[0163] With respect to an amin acid residue of a polypeptide, by “conserved” or “conservative” it means that the amino acid residue is constant (not changed) across all indicated polypeptides. With respect to a motif of a polypeptide or a nucleic acid, by “conserved” or “conservative” it means that the motif is constant (not changed) across all indicated polypeptides or nucleic acids. As used herein, the term “motif” refers to a segment of a polypeptide or a nucleic acid, consisting of multiple (more than one) amino acids or nucleotides.
[0164] As used herein, the description of a mutant “comprising an amino acid mutation that is corresponding to a given amino acid mutation (e.g., substitution, such as, D10A) or that is a given amino acid mutation (e.g., substitution, such as, D10A) relative to a given polypeptide” means that the mutant is a mutant of the given polypeptide (as a reference polypeptide) or another different polypeptide (as a reference polypeptide) and comprises the same type of amino acid mutation (e.g., D-to-Asubstitution) as the given amino acid mutation at a position in the mutant corresponding to the position (e.g., D10) of the given amino acid mutation (e.g., D10A) numbered according to the given polypeptide. For example, a mutant comprising an amino acid substitution corresponding to D10A relative to a given polypeptide (e.g., SN007) refers to the fact that the given polypeptide comprises amino acid D (Asp) at position 61, and the mutant comprises amino acid A (Ala) at a position of a reference polypeptide from which the mutant is generated corresponding to the position D10 of the given polypeptide. The corresponding relationship of positions in two or more amino acid sequences as determined by sequence alignment is explained in the previous paragraphs.
[0165] As used herein, “amino acid mutation” includes addition, deletion, and / or substitution. Insertion is also a kind of addition that often occurs within a polypeptide. Truncation is also a kind of deletion. N-terminal or C-terminal truncation / deletion means the truncation / deletion occurs at the N-terminal or C-terminal of a polypeptide. Typically, C-terminal truncation / deletion refers to a deletion of one or more amino acids starting from the most C-terminal amino acid of a polypeptide to the N-terminal of the polypeptide. Typically, N-terminal truncation / deletion refers to a deletion of one or more amino acids starting from the most N-terminal amino acid of a polypeptide to the C-terminal of the polypeptide. Alternatively, in some embodiments, the most N-terminal amino acid of a polypeptide is Met (corresponding to the start codon ATG of a nucleic acid encoding the polypeptide) , and the N-terminal truncation / deletion starts from the second amino acid of the polypeptide downstream of (C’ to) (on the right side of) the Met to the C-terminal of the polypeptide. As a specific but non-limiting example, a N-terminal truncation of 20 amino acids of a polypeptide refers to a deletion of amino acids at positions 1-20 of the polypeptide, or in some embodiments, a deletion of amino acids at positions 2-21 of the polypeptide where the most N-terminal amino acid of the polypeptide is Met and retained after the deletion.
[0166] As used herein, a “conservative substitution” refers to a substitution of an amino acid made among amino acids within one of the following four groups:
[0167] (1) non-polar amino acids, including Glycine (Gly / G) , Alanine (Ala / A) , Valine (Val / V) , Cysteine (Cys / C) , Proline (Pro / P) , Leucine (Leu / L) , Isoleucine (Ile / I) , Methionine (Met / M) , Tryptophan (Trp / W) , and Phenylalanine (Phe / F) ;
[0168] (2) negatively charged amino acids, including Aspartic Acid (Asp / D) and Glutamic Acid (Glu / E) ;
[0169] (3) polar amino acids, including Serine (Ser / S) , Threonine (Thr / T) , Tyrosine (Tyr / Y) , Asparagine (Asn / N) , and Glutamine (Gln / Q) ; and
[0170] (4) positively charged amino acids, including Lysine (Lys / K) , Arginine (Arg / R) , and Histidine (His / H) .
[0171] As used herein, the terms “non-naturally occurring” and “engineered” are used interchangeably and refer to artificial participation. When these terms are used to describe a nucleic acid or a polypeptide, it is meant that the nucleic acid or polypeptide is at least substantially freed from at least one other component of its association in nature or as found in nature.
[0172] As used herein, the term “wild type” has the meaning commonly understood by those skilled in the art to mean a typical form of an organism, a strain, a gene, or a feature that distinguishes it from a mutant or variant when it exists in nature. It can be isolated from sources in nature and not intentionally modified.
[0173] As used herein, the term “endonuclease activity” is used interchangeably with “ (ds) DNA endonuclease activity” or “dsDNA cleavage activity” against both strands of a dsDNA. As used herein, the term “nickase activity” is used interchangeably with “ (ds) DNA nickase activity” or “ssDNA cleavage activity” against one strand of a dsDNA, and the substrate for such activity is a dsDNA and is not a ssDNA. As used herein, the term “endonuclease” is used interchangeably with “ (ds) DNA endonuclease” , and the term “nickase” is used interchangeably with “(ds) DNA nickase” . As used herein, the term “nick” is used interchangeably with “ssDNA cleavage” against one strand of a dsDNA, and the substrate for such activity is a dsDNA rather than a ssDNA. Unless otherwise indicated (for example, indicated for “off-target” ) , the term “endonuclease activity” in reference to the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure refers to guide sequence specific (on-target) endonuclease activity. Unless otherwise indicated, the term “nickase activity” in reference to the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure refers to guide sequence specific (on-target) nickase activity.
[0174] Although in the disclosure sometime reference is made to reduced nickase activity of a polypeptide compared to the nickase activity of a reference polypeptide, it does not mean to acknowledge that the reference polypeptide is a nickase. A reference polypeptide that has endonuclease activity may also show positive results in a nickase activity evaluation assay due to its inherent capability of cleaving one strand of a dsDNA, which however does not make it a nickase. Typically, nickase refers to an indicated polypeptide substantially lacking endonuclease activity and substantially having nickase activity.
[0175] As used herein, a “fusion protein” refers to a protein created through the joining (via a linker or not) of two or more originally separate proteins, or portions thereof.
[0176] As used herein, the derivatives of the programmable RNA-guided DNA endonucleases in the disclosure include, but are not limited to, the mutants (e.g., nickase) of the programmable RNA-guided DNA endonucleases and the fusion proteins (e.g., base editor, prime editor) comprising the programmable RNA-guided DNA endonucleases or mutants (e.g., nickase) thereof.
[0177] As used herein, the term “heterologous” in reference to polypeptide domains refers to the fact that the polypeptide domains do not naturally occur together (e.g., in the same polypeptide) . For example, in fusion proteins generated by the hand of man, a polypeptide domain from one polypeptide may be fused to a polypeptide domain from a different polypeptide. The two polypeptide domains would be considered “heterologous” with respect to each other, as they do not naturally occur together.
[0178] As used herein, the term “heterologous” in reference to nucleotide sequences refers to the fact that the nucleotide sequences do not naturally occur together (e.g., in the same polynucleotide) . For example, in a guide nucleic acid generated by the hand of man, a guide sequence intentionally designed to target a human gene locus may be fused to a scaffold sequence from a microorganism. The two nucleotide sequences would be considered “heterologous” with respect to each other, as they do not naturally occur together.
[0179] As used herein, the terms “nucleic acid” , “nucleic acid molecule” , or “polynucleotide” are used interchangeably. They refer to a polymer of deoxyribonucleotides or ribonucleotides or their mixtures of any length in either single-or double-stranded form, and, unless otherwise stated, encompass known analogs of natural nucleotides that can function in a similar manner as naturally occurring nucleotides. The terms encompass nucleic acid-like structures with synthetic backbones, as well as amplification products. DNAs and RNAs are both polynucleotides. The polymer may include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine) , nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O (6) -methylguanine, and 2-thiocytidine) , chemically modified bases, biologically modified bases (e.g., methylated bases) , intercalated bases, modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose) , or modified phosphate groups (e.g., phosphorothioates and 5′-N-phosphoramidite linkages) . The terms include both modified and unmodified.
[0180] As used herein, the terms “programmable nucleic acid” , “guide nucleic acid” , and “programmable guide nucleic acid” refer to a nucleic acid-based molecule capable of guiding a polypeptide (for example, the programmable RNA-guided DNA endonucleases or derivatives thereof of the disclosure) to a target nucleic acid, by comprising a scaffold sequence capable of forming a complex with the polypeptide and comprising a guide sequence capable of hybridizing to the target nucleic acid. The terms include, but are not limited to, RNA-based molecules, e.g., a guide RNA. As used herein, the terms “guide RNA (gRNA) ” and “RNA guide” are used interchangeably. As used in the disclosure, the term “guide sequence” is used interchangeably with the term “spacer sequence” or “spacer” . As used herein, the term “complex” refers to a grouping of two or more molecules, e.g., grouping of a polypeptide in the disclosure and a guide nucleic acid in the disclosure (via the scaffold sequence of the guide nucleic acid) . In some embodiments, the complex comprises a nucleic acid and a polypeptide interacting with (e.g., binding to, coming into contact with, adhering to) one another. As used herein, the term “complex” can refer to a grouping of a guide nucleic acid and a polypeptide. As used herein, the term “complex” can refer to a grouping of a guide nucleic acid, a polypeptide, and a target nucleic acid (e.g., a target DNA) .
[0181] With respect to a scaffold sequence in the disclosure, the terms “variant” and “mutant” are used interchangeably to refer to a mutant of a reference scaffold sequence (e.g., a scaffold sequence in Table 1) generated by introducing a nucleotide mutation into the reference scaffold sequence.
[0182] As described herein, the guide sequence is so designed to be capable of hybridizing to a target sequence of a target DNA. As used herein, the term “hybridize” , “hybridizing” , or “hybridization” refers to a reaction in which one or more polynucleotide sequences react to form a complex that is stabilized via hydrogen bonding between the bases of the one or more polynucleotide sequences. The hydrogen bonding may occur by Watson Crick base pairing, Hoogstein binding, or in any other sequence specific manner. As used herein, the hybridization of a guide sequence and a target sequence is so stabilized to permit a polypeptide that is complexed with a guide nucleic acid comprising the guide sequence to act (e.g., cleave, deaminize) at or near the target sequence or its complement.
[0183] For the purpose of hybridization, in some embodiments, the guide sequence is reversely complementary to a target sequence. As used herein, the term “reverse complementary” refers to the ability of nucleobases of a first polynucleotide sequence, such as a guide sequence, to base pair with nucleobases of a second polynucleotide sequence, such as a target sequence, by traditional Watson-Crick base-pairing. Two reverse complementary polynucleotide sequences are able to non-covalently bind under appropriate temperature and solution ionic strength conditions. In some embodiments, a first polynucleotide sequence (e.g., a guide sequence) comprises 100%(fully) reverse complementarity to a second nucleic acid (e.g., a target sequence) . In some embodiments, a first polynucleotide sequence (e.g., a guide sequence) is reverse complementary to a second polynucleotide sequence (e.g., a target sequence) if the first polynucleotide sequence comprises at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%complementarity to the second nucleic acid (i.e., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%of the nucleotides of the first polynucleotide sequence can base-pair with the nucleotides of the second polynucleotide sequence) . As used herein, the term “substantially complementary” refers to a first polynucleotide sequence (e.g., a guide sequence) that has a certain level of complementarity to a second polynucleotide sequence (e.g., a target sequence) (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%of the first polynucleotide sequence can base-pair with the second polynucleotide sequence, or at most 1, 2, 3, 4, or 5 contiguous or non-contiguous nucleotides of the first polynucleotide sequence mismatch the nucleotides of the second polynucleotide sequence) . In some embodiments, the level of complementarity is such that the first polynucleotide sequence (e.g., a guide sequence) can hybridize to the second polynucleotide sequence (e.g., a target sequence) with sufficient affinity to permit a polypeptide that is complexed with the first polynucleotide sequence or a nucleic acid comprising the first polynucleotide sequence to act (e.g., cleave, deaminize) on the target sequence or its complement. In some embodiments, a guide sequence that is substantially complementary to a target sequence has 100%or less than 100%complementarity to the target sequence. In some embodiments, a guide sequence that is substantially complementary to a target sequence has at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%complementarity to the target sequence, and / or has at most 1, 2, 3, 4, or 5 contiguous or non-contiguous nucleotide mismatches from the target sequence.
[0184] With respect to a system in the disclosure comprising a polypeptide (e.g., SN007 or a mutant thereof) in the disclosure and a guide nucleic acid in the disclosure, wild type system may be used to refer to a system comprising a wild type polypeptide in the disclosure (e.g., SN007 of SEQ ID NO: 13 in Table 1) and a guide nucleic acid comprising the scaffold sequence in the disclosure corresponding to the wild type polypeptide (e.g., the scaffold sequence of SEQ ID NO: 14 in Table 1) , and variant system may be used to refer to a system comprising a mutant of a wild type or reference polypeptide (e.g., a mutant of SN007, such as, a SN007 nickase of SEQ ID NO: 23 or 24) in the disclosure and / or a guide nucleic acid comprising a mutant of the scaffold sequence in the disclosure corresponding to the wild type or reference polypeptide (e.g., a mutant of the scaffold sequence of SEQ ID NO: 14 in Table 1) .
[0185] As used herein, the terms “protospacer adjacent motif (PAM) ” and “target adjacent motif (TAM) ” are used interchangeably and refer to a short nucleotide sequence (or a motif) immediately 3’ to a protospacer sequence on the nontarget strand of a target dsDNA recognizable by a polypeptide of the disclosure.
[0186] As used herein, the term “identity” or “sequence identity” refers to the overall relatedness between polymeric molecules, e.g., between nucleic acids (e.g., DNA and / or RNA) and / or between polypeptides. In some embodiments, polymeric molecules are considered to be “substantially identical” to one another if their sequences are at least 80%, 85%, 90%, 95%, or 99%identical. Calculation of the percent identity of two nucleic acids or polypeptides, for example, can be performed by aligning the two sequences for optimal comparison purpose (e.g., gaps can be introduced in one or both of a first and a second sequences for optimal alignment, and non-identical sequences can be disregarded for comparison purposes) . In certain embodiments, the length of a sequence aligned for comparison purpose is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100%of the length of a reference sequence. The nucleotides at corresponding positions are then compared. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. As is well known in the art, nucleic acids or polypeptides may be compared using any of a variety of algorithms, including those available in commercial computer programs such as BLASTN for nucleotide sequences and BLASTP, gapped BLAST, and PSI-BLAST for amino acid sequences. In some embodiments, the sequence identity is calculated by global alignment, for example, using the Needleman-Wunsch algorithm, for example, using an online tool at ebi. ac. uk / Tools / psa / emboss_needle / . In some embodiments, the sequence identity is calculated by local alignment, for example, using the Smith-Waterman algorithm, for example, using an online tool at ebi. ac. uk / Tools / psa / emboss_water / .
[0187] As used herein, the terms “upstream” and “downstream” refer to relative positions within a single nucleic acid (e.g., DNA) or within a single polypeptide. “Upstream” and “downstream” relate to the 5’ to 3’ direction of a single nucleic acid, respectively, in which transcription occurs, or N-to C-orientation of a single polypeptide, respectively, in which translation occurs. For a first sequence and a second sequence present on the same strand of a single nucleic acid written in 5’ to 3’ direction or a single polypeptide written in N-to-C orientation, the first sequence is upstream of the second sequence when the 3’ end or C-terminal of the first sequence is on the left side of the 5’ end or N-terminal of the second sequence, and the first sequence is downstream of the second sequence when the 5’ end or N-terminal of the first sequence is on the right side of the 3’ end or C-terminal of the second sequence. For example, a promoter is usually at the upstream of a coding sequence under the regulation of the promoter; and on the other hand, a coding sequence under the regulation of a promoter is usually at the downstream of the promoter.
[0188] As used herein, the term “regulatory element” refers to a DNA sequence that controls or impacts one or more aspects of transcription and / or expression and is intended to include promoters, enhancers, silencers, termination signals, internal ribosome entry sites (IRES) , and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences) . Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences) . Regulatory elements may also direct expression in a time-dependent manner, e.g., in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type specific.
[0189] As used herein, the term “operably linked” refers to a juxtaposition wherein the components described are in a relationship permitting them to function in their intended manner. A regulatory element “operably linked” to a functional element is associated in such a way that transcription, expression, and / or activity of the functional element is achieved under conditions compatible with the regulatory element. In some embodiments, “operably linked” regulatory elements are contiguous (e.g., covalently linked) with the functional elements of interest; in some embodiments, regulatory elements act in trans to or otherwise at a distance from the functional elements of interest.
[0190] As used herein, the term “cell” is understood to refer not only to a particular individual cell, but to the progeny or potential progeny of the cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term.
[0191] As used herein, the term “in vivo” means inside the body of an organism, and the terms “ex vivo” or “in vitro” means outside the body of an organism.
[0192] As used herein, the term “treat” , “treatment” , or “treating” is an approach for obtaining beneficial or desired results including clinical results. For purposes of the disclosure, the beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms resulting from a disease, diminishing the extent of a disease, stabilizing a disease (e.g., delaying the worsening of a disease) , delaying the spread (e.g., metastasis) of a disease, delaying the recurrence of a disease, reducing recurrence rate of a disease, delay or slowing the progression of a disease, ameliorating a disease state, providing a remission (partial or total) of a disease, decreasing the dose of one or more other medications required to treat a disease, delaying the progression of a disease, increasing the quality of life, and prolonging survival. Also encompassed by the term is a reduction of pathological consequence of a disease (such as cancer) . The methods of the disclosure contemplate any one or more of these aspects of treatment.
[0193] As used herein, the term “disease” includes the terms “disorder” and “condition” and is not limited to those specific diseases that have been medically or clinically defined.
[0194] As used herein, reference to “not” a value or parameter generally means and describes “other than” a value or parameter. For example, the method is not used to treat cancer of type X means the method may be used to treat cancer of types other than X.
[0195] As used herein, the singular forms “a” , “an” , and “the” include plural referents unless the context clearly dictates otherwise. That is, articles “a / an” and “the” are used herein to refer to one or more than one (i.e., at least one) grammatical object of the article. For example, “an element” means one element or more than one element, e.g., two elements.
[0196] As used herein, the term “and / or” in a phrase such as “A and / or B” is intended to mean either or both of the alternatives, including both A and B, A or B, A (alone) , and B (alone) . Likewise, the term “and / or” in a phrase such as “A, B, and / or C” is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone) ; B (alone) ; and C (alone) .
[0197] As used herein, when the term “about” is ahead of a serious of numbers (for example, about 1, 2, 3) , it is understood that each of the serious of numbers is modified by the term “about” (that is, about 1, about 2, about 3) . The term “about X-Y” used herein has the same meaning as “about X to about Y. ”
[0198] As used herein, a numerical range includes the end values of the range and each specific value within the range, for example, “16 to 100 nucleotides” includes 16 nucleotides and 100 nucleotides and each specific value between 16 and 100, e.g., 17, 23, 34, 52, 78.
[0199] It is understood that embodiments of the disclosure described herein include “consisting” and / or “consisting essentially of” embodiments.
[0200] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely” , “only” , and the like in connection with the recitation of claim elements, or use of a “negative” limitation.
[0201] Overview
[0202] The clinical translation of base editing hinges on the ability to balance efficiency, specificity, and delivery8, 85. In the disclosure, the applicant introduced several programmable RNA-guided DNA endonucleases and a class of compact, high-fidelity base editors-eMM762ABE and eMM762CBE-developed through extensive engineering based on MM762. These editors demonstrated a unique combination of high activity, markedly reduced off-target editing, and compatibility with single-AAV delivery, thereby addressing several of the most pressing limitations of current SpCas9-derived base editors.
[0203] The cryo-EM analysis in the disclosure revealed that MM762 forms an extensive interaction network with the sgRNA-DNA heteroduplex, exceeding the contact footprint of SpCas9. This architecture, together with the distinctive RNA-effector module (REM) described in MG34-150, enforces stringent guide-target pairing and extensive R-loop contacts, providing a structural basis for the intrinsic high fidelity. Partial exposure of sgRNA stem loops explains why scaffold truncation enhanced activity, while engineered substitutions introduced new electrostatic interactions that further stabilized sgRNA and target DNA binding. These structural insights not only rationalize the fidelity of MM762 but also guided its optimization into efficient base editors. Thus, unlike engineered high-fidelity SpCas9 variants33, 35, which often compromise on-target activity, MM762 and its derivatives retain efficient editing while preserving high specificity.
[0204] Through systematic mutagenesis, auxiliary domain fusion, computational redesign, and sgRNA scaffold refinement, the applicant substantially enhanced the efficiency of eMM762-derived base editors without compromising specificity. These engineering strategies yielded eMM762ABE variants that matched or surpassed SpCas9-based ABE8e while exhibiting higher specificity than HiFiABE8e. Notably, eMM762ABEv3 showed minimal nuclease-dependent and nuclease-independent off-target activity, as confirmed by GUIDE-seq, orthogonal R-loop assays and targeted deep sequencing.
[0205] The therapeutic relevance of eMM762ABE and eMM762CBE was demonstrated in both ex vivo and in vivo models. In primary human T cells, these editors enabled near-complete knockout of B2M, CD52, TRAC, and PD1 via multiplex editing, while dramatically reducing off-target activity by up to 404-fold compared to SpCas9-based counterparts and generating substantially fewer unintended indels than SpCas9ABE8e, highlighting their potential for safer and more precise genome editing in cell therapy applications. In vivo, eMM762ABEv3 delivered by a single AAV vector efficiently disrupted the DMD exon 50 splice donor site in a humanized DMD mouse model, restoring dystrophin expression in >90%of myofibers with higher specificity than dual-AAV SpCas9ABE8e. This compact system represents a clinically actionable alternative to larger editors that require split-intein systems or dual-AAV vector packaging85, 86.
[0206] In conclusion, the results in the disclosure established eMM762 as a structurally defined, compact, efficient, and high-fidelity nuclease scaffold. The resulting base editors, eMM762ABE and eMM762CBE, unite robust efficiency and exceptional specificity with a size compatible with single-AAV delivery, making them strong candidates for next-generation therapeutic genome editing.
[0207] The disclosure provides, in part, programmable RNA-guided DNA endonucleases including SN001-SN007, and derivatives (including mutants and fusion proteins) thereof, systems comprising the programmable RNA-guided DNA endonuclease or the derivative and a guide nucleic acid, e.g., a guide RNA, and uses thereof or methods of using the same.
[0208] Without wishing to be bound by theory, it is believed that the programmable RNA-guided DNA endonucleases of the disclosure contain a RuvC nuclease domain separated into three segments: RuvC I, II, and III domains, which is responsible for single-strand cleavage at the nontarget strand of a dsDNA, and a HNH nuclease domain, which is responsible for single-strand cleavage at the target strand of a dsDNA, together leading to double-strand cleavage of the dsDNA.
[0209] It is widely thought that Cas9 evolved from IscB, which contains a small effector protein (420-500 aa in length) and a large scaffold sequence of guide RNA. As IscB evolved, its protein domains have increased in size, while the gRNA has decreased. For example, OgeuIscB as a representative IscB has a size of 496 aa and a gRNA with a scaffold sequence having a size of 206 nt, whereas SpCas9 as a representative Cas9 has a size of 1368 aa and a gRNA with a scaffold sequence having a size of 76 nt. In contrast to IscB and Cas9, the programmable RNA-guided DNA endonucleases of the disclosure have a size of 735-771 amino acids longer than OgeuIscB (496 aa) and short than SpCas9 (1368 aa) , and the scaffold sequences of their gRNA have a size of 129-135 nt shorter than the scaffold sequence for OgeuIscB (206 nt) and longer than the scaffold sequence for SpCas9 (76 nt) . Therefore, without wishing to be bound by theory, it is believed that the programmable RNA-guided DNA endonucleases of the disclosure are evolutionarily intermediates between Cas9 and its ancestor IscB, yet not classified as either Cas9 or IscB.
[0210] The programmable RNA-guided DNA endonucleases of the disclosure have one or more of the structural or functional features, including but not limited to the followings, that Cas9 and IscB lack and hence distinguish them from Cas9 and IscB:
[0211] (1) significantly distinct length of amino acid sequence different from SpCas9 and OgeuIscB;
[0212] (2) significantly low sequence identity (about 20%) to SpCas9 and to OgeuIscB;
[0213] (3) enrichment in arginine and lysine content;
[0214] (4) containing RRXRR motif in REC domain;
[0215] (5) containing a Zn-finger within the HNH domain;
[0216] (6) containing a Zn-binding ribbon motif in the recognition domain;
[0217] (7) containing a Zn-binding ribbon motif in the vicinity of RuvC-II;
[0218] (8) recognizing PAM from both the major and minor grooves of a target dsDNA;
[0219] (9) from the major groove side of a target dsDNA, employing the amino acid of the polypeptide of the disclosure corresponding to K728 of SN007 to interact with the dsDNA; and
[0220] (10) from the minor groove side of a target dsDNA, employing the amino acids of the polypeptide of the disclosure corresponding to N664 and K662 of SN007 to interact with the dsDNA.
[0221] The programmable RNA-guided DNA endonucleases of the disclosure have one or more of the advantages, including but not limited to the followings, compared to Cas9 (e.g., SpCas9) , IscB (e.g., OgeuIscB) , and Cas12a (Cpf1) (e.g., FnCas12) :
[0222] (1) suitable size for efficient delivery, especially single AAV delivery;
[0223] (2) high on-target editing;
[0224] (3) low off-target editing;
[0225] (4) wide PAM recognition; and
[0226] (5) ability to be engineered to a nickase due to the presence of HNH domain.
[0227] Representative polypeptides
[0228] In an aspect, the disclosure provides a programmable RNA-guided DNA endonuclease as set forth in any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11. In some embodiments, the disclosure provides a wild type polypeptide as set forth in any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11. In some embodiments, the disclosure provides a reference polypeptide as set forth in any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11. In some embodiments, the disclosure provides a reference polypeptide that is any polypeptide in the disclosure.
[0229] In another aspect, the disclosure provides a polypeptide comprising an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11.
[0230] In some embodiments, the polypeptide is not any nuclease disclosed in Aliaga Goltsman, D. S., Alexander, L. M., Lin, JL. et al. Compact Cas9d and HEARO enzymes for genome editing discovered from uncultivated microbes. Nat Commun 13, 7602 (2022) . https: / / doi. org / 10.1038 / s41467-022-35257-7. In some embodiments, the polypeptide is not any one of MG33-1, MG33-2, MG33-3, MG33-4, MG33-5, MG33-6, MG33-7, MG33-8, MG33-9, MG33-10, MG33-11, MG33-12, MG33-13, MG33-14, MG33-15, MG33-16, MG33-17, MG33-18, MG33-19, MG33-20, MG33-21, MG33-22, MG33-23, MG33-24, MG33-26, MG33-25, MG33-27, MG33-28, MG33-29, MG33-30, MG33-31, MG33-32, MG33-33, MG33-34, MG34-1, MG34-3, MG34-5, MG34-7, MG34-8, MG34-9, MG34-10, MG34-12, MG34-13, MG34-14, MG34-15, MG34-16, MG34-17, MG34-19, MG34-20, MG34-21, MG34-22, MG34-23, MG34-24, MG102-1, MG102-2, MG102-3, MG102-4, MG102-5, MG102-6, MG102-7, MG102-8, MG102-9, MG102-10, MG102-11, MG102-12, MG102-13, MG102-14, MG102-15, MG102-16, MG102-17, MG102-18, MG102-19, MG102-20, MG102-21, MG102-22, MG102-23, MG102-24, MG102-25, MG102-27, MG102-28, MG102-29, MG102-30, MG102-31, MG102-32, MG102-33, MG102-26, MG102-35, MG102-36, MG102-37, MG102-38, MG102-39, MG102-40, MG102-41, MG102-42, MG102-43, MG102-44, MG102-45, MG102-46, MG102-47, MG102-48, MG143-1, MG144-1, MG144-2, MG144-3, MG144-4, MG145-1, MG35-1, MG35-2, MG35-3, MG35-7, MG35-8, MG35-9, MG35-10, MG35-11, MG35-12, MG35-13, MG35-14, MG35-15, MG35-16, MG35-17, MG35-18, MG35-19, MG35-20, MG35-21, MG35-22, MG35-23, MG35-24, MG35-25, MG35-30, MG35-31, MG35-32, MG35-33, MG35-34, MG35-35, MG35-36, MG35-37, MG35-38, MG35-39, MG35-40, MG35-41, MG35-42, MG35-43, MG35-44, MG35-45, MG35-46, MG35-47, MG35-48, MG35-49, MG35-50, MG35-51, MG35-52, MG35-53, MG35-54, MG35-55, MG35-56, MG35-57, MG35-58, MG35-59, MG35-60, MG35-61, MG35-62, MG35-63, MG35-64, MG35-65, MG35-66, MG35-67, MG35-68, MG35-69, MG35-70, MG35-71, MG35-72, MG35-73, MG35-74, MG35-75, MG35-76, MG35-77, MG35-78, MG35-79, MG35-80, MG35-81, MG35-82, MG35-83, MG35-84, MG35-85, MG35-86, MG35-87, MG35-88, MG35-89, MG35-90, MG35-91, MG35-92, MG35-93, MG35-94, MG35-95, MG35-96, MG35-97, MG35-98, MG35-99, MG35-100, MG35-101, MG35-102, MG35-103, MG35-104, MG35-105, MG35-106, MG35-107, MG35-108, MG35-109, MG35-110, MG35-111, MG35-112, MG35-113, MG35-114, MG35-115, MG35-116, MG35-117, MG35-118, MG35-119, MG35-120, MG35-121, MG35-122, MG35-123, MG35-124, MG35-125, MG35-126, MG35-127, MG35-128, MG35-129, MG35-130, MG35-131, MG35-132, MG35-133, MG35-134, MG35-135, MG35-136, MG35-137, MG35-138, MG35-139, MG35-140, MG35-141, MG35-142, MG35-143, MG35-144, MG35-146, MG35-147, MG35-148, MG35-149, MG35-150, MG35-151, MG35-152, MG35-153, MG35-154, MG35-155, MG35-156, MG35-157, MG35-158, MG35-159, MG35-160, MG35-161, MG35-162, MG35-163, MG35-164, MG35-165, MG35-166, MG35-168, MG35-170, MG35-171, MG35-172, MG35-175, MG35-212, MG35-213, MG35-214, MG35-223, MG35-301, MG35-323, MG35-385, MG35-387, MG35-4, MG35-5, MG35-6, MG35-145, MG35-419, MG35-420, MG35-421, MG35-176, MG35-177, MG35-178, MG35-179, MG35-180, MG35-181, MG35-183, MG35-184, MG35-185, MG35-186, MG35-187, MG35-188, MG35-189, MG35-190, MG35-191, MG35-192, MG35-193, MG35-194, MG35-195, MG35-196, MG35-197, MG35-198, MG35-199, MG35-200, MG35-201, MG35-202, MG35-203, MG35-204, MG35-205, MG35-206, MG35-207, MG35-208, MG35-209, MG35-210, MG35-211, MG35-215, MG35-216, MG35-217, MG35-218, MG35-219, MG35-220, MG35-221, MG35-222, MG35-224, MG35-225, MG35-226, MG35-227, MG35-228, MG35-229, MG35-230, MG35-231, MG35-232, MG35-233, MG35-234, MG35-235, MG35-236, MG35-237, MG35-238, MG35-239, MG35-240, MG35-241, MG35-242, MG35-243, MG35-244, MG35-245, MG35-246, MG35-247, MG35-248, MG35-249, MG35-250, MG35-251, MG35-252, MG35-253, MG35-254, MG35-255, MG35-256, MG35-257, MG35-258, MG35-259, MG35-260, MG35-261, MG35-262, MG35-263, MG35-264, MG35-265, MG35-266, MG35-267, MG35-268, MG35-269, MG35-270, MG35-271, MG35-272, MG35-273, MG35-274, MG35-275, MG35-276, MG35-277, MG35-278, MG35-279, MG35-280, MG35-281, MG35-282, MG35-283, MG35-284, MG35-285, MG35-286, MG35-287, MG35-288, MG35-289, MG35-290, MG35-291, MG35-292, MG35-293, MG35-294, MG35-295, MG35-296, MG35-297, MG35-298, MG35-299, MG35-300, MG35-302, MG35-303, MG35-304, MG35-305, MG35-307, MG35-308, MG35-309, MG35-310, MG35-311, MG35-312, MG35-313, MG35-314, MG35-315, MG35-316, MG35-317, MG35-318, MG35-319, MG35-320, MG35-321, MG35-322, MG35-324, MG35-325, MG35-326, MG35-327, MG35-328, MG35-329, MG35-330, MG35-331, MG35-332, MG35-333, MG35-334, MG35-335, MG35-336, MG35-337, MG35-338, MG35-339, MG35-340, MG35-341, MG35-342, MG35-343, MG35-344, MG35-345, MG35-346, MG35-347, MG35-348, MG35-349, MG35-350, MG35-351, MG35-352, MG35-353, MG35-354, MG35-355, MG35-356, MG35-357, MG35-358, MG35-359, MG35-360, MG35-361, MG35-362, MG35-363, MG35-364, MG35-365, MG35-366, MG35-367, MG35-368, MG35-369, MG35-370, MG35-371, MG35-372, MG35-373, MG35-374, MG35-375, MG35-376, MG35-377, MG35-378, MG35-379, MG35-384, MG35-386, MG35-388, MG35-389, MG35-390, MG35-391, MG35-392, MG35-393, MG35-394, MG35-395, MG35-396, MG35-397, MG35-398, MG35-399, MG35-400, MG35-401, MG35-402, MG35-403, MG35-404, MG35-405, MG35-406, MG35-408, MG35-409, MG35-410, MG35-411, MG35-412, MG35-413, MG35-414, MG35-415, MG35-416, MG35-417, MG35-418, MG35-182, MG35-306, MG35-380, MG35-381, MG35-382, MG35-383, MG35-407, MG35-422, MG35-423, MG35-424, MG35-425, MG35-426, MG35-427, MG35-428, MG35-429, MG35-430, MG35-431, MG35-432, MG35-433, MG35-434, MG35-435, MG35-436, MG35-437, MG35-438, MG35-439, MG35-440, MG35-441, MG35-442, MG35-443, MG35-444, MG35-445, MG35-446, MG35-447, MG35-448, MG35-449, MG35-450, MG35-451, MG35-452, MG35-453, MG35-454, MG35-455, MG35-456, MG35-457, MG35-458, MG35-459, MG35-460, MG35-461, MG35-462, MG35-463, MG35-464, MG35-465, MG35-466, MG35-467, MG35-468, MG35-469, MG35-470, MG35-471, MG35-472, MG35-473, MG35-474, MG35-475, MG35-476, MG35-477, MG35-478, MG35-479, MG35-480, MG35-481, MG35-482, MG35-483, MG35-484, MG35-485, MG35-486, MG35-487, MG35-488, MG35-489, MG35-490, MG35-491, MG35-492, MG35-493, MG35-494, MG35-495, MG35-496, MG35-497, MG35-498, MG35-499, MG35-500, MG35-501, MG35-502, MG35-503, MG35-504, MG35-505, MG35-506, MG35-507, MG35-508, MG35-509, MG35-510, MG35-511, MG35-512, and MG35-513.
[0231] In some embodiments, the polypeptide in the disclosure is nucleic acid programmable (programmable nucleic acid-guided) . In some embodiments, the polypeptide in the disclosure is RNA programmable (programmable RNA-guided) .
[0232] In some embodiments, the polypeptide in the disclosure is an endonuclease or has endonuclease activity. In some embodiments, the polypeptide in the disclosure is a nickase or has nickase activity. In some embodiments, the polypeptide in the disclosure is a DNA binding protein or has DNA binding property. In some embodiments, the polypeptide has at least one of endonuclease activity, nickase activity, and DNA binding property.
[0233] In some embodiments, the polypeptide in the disclosure comprises a REC domain split into two separate domains.
[0234] In some embodiments, the REC domain is split into two domains separated by a RRXRR motif.
[0235] Mutants
[0236] In some embodiments, the polypeptide in the disclosure is a mutant of a reference polypeptide (which could be any polypeptide in the disclosure, e.g., any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11) in the disclosure, i.e., comprising an amino acid mutation relative to (compared to) the reference polypeptide.
[0237] In some embodiments, the polypeptide has one or more of the properties, including but not limited to the followings, compared to the reference polypeptide:
[0238] (1) higher endonuclease activity;
[0239] (2) lower endonuclease activity;
[0240] (3) higher nickase activity;
[0241] (4) higher ability to bind to a target dsDNA;
[0242] (5) higher base editing efficiency when used in a base editor for base editing;
[0243] (6) higher prime editing efficiency when used in a prime editor for prime editing;
[0244] (7) higher epigenomic editing efficiency when used in an epigenomic editor for epigenomic editing;
[0245] (8) lower off-target endonuclease activity;
[0246] (9) lower off-target nickase activity;
[0247] (10) wider PAM recognition;
[0248] (11) higher on-target editing; and
[0249] (12) lower off-target editing.
[0250] In some embodiments, the polypeptide comprises an amino acid mutation relative to (compared to) the amino acid sequence of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11.
[0251] In some embodiments, the polypeptide comprising the amino acid mutation comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the amino acid sequence of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11.
[0252] In some embodiments, the amino acid mutation is an amino acid addition, deletion, or substitution.
[0253] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 714, 715, 716, 717, 718, 719, 720, 721, 722, 723, 724, 725, 726, 727, 728, 729, 730, 731, 732, 733, 734, 735, 736, 737, 738, 739, 740, 741, 742, 743, 744, 745, 746, and / or 747 of SEQ ID NO: 1.
[0254] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 714, 715, 716, 717, 718, 719, 720, 721, 722, 723, 724, 725, 726, 727, 728, 729, 730, 731, 732, 733, 734, 735, 736, 737, 738, 739, 740, 741, 742, 743, 744, 745, 746, and / or 747 of SEQ ID NO: 3.
[0255] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 714, 715, 716, 717, 718, 719, 720, 721, 722, 723, 724, 725, 726, 727, 728, 729, 730, 731, 732, 733, 734, 735, 736, 737, 738, 739, 740, 741, 742, 743, 744, 745, 746, 747, 748, 749, 750, 751, 752, 753, 754, 755, 756, 757, 758, 759, 760, 761, 762, 763, 764, 765, 766, 767, 768, 769, 770, and / or 771 of SEQ ID NO: 5.
[0256] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 714, 715, 716, 717, 718, 719, 720, 721, 722, 723, 724, 725, 726, 727, 728, 729, 730, 731, 732, 733, 734, and / or 735 of SEQ ID NO: 7.
[0257] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 714, 715, 716, 717, 718, 719, 720, 721, 722, 723, 724, 725, 726, 727, 728, 729, 730, 731, 732, 733, 734, 735, 736, 737, 738, 739, 740, 741, 742, 743, 744, and / or 745 of SEQ ID NO: 9.
[0258] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 714, 715, 716, 717, 718, 719, 720, 721, 722, 723, 724, 725, 726, 727, 728, 729, 730, 731, 732, 733, 734, 735, 736, 737, 738, 739, 740, 741, 742, 743, 744, 745, 746, 747, 748, 749, 750, 751, 752, 753, 754, 755, 756, 757, 758, 759, 760, and / or 761 of SEQ ID NO: 11.
[0259] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 714, 715, 716, 717, 718, 719, 720, 721, 722, 723, 724, 725, 726, 727, 728, 729, 730, 731, 732, 733, 734, 735, 736, 737, 738, 739, 740, 741, 742, 743, 744, 745, 746, 747, 748, 749, 750, 751, 752, 753, 754, 755, 756, 757, 758, 759, 760, 761, and / or 762 of SEQ ID NO: 13.
[0260] In some embodiments, the amino acid substitution is a conservative amino acid substitution or a non-conservative amino acid substitution.
[0261] In some embodiments, the amino acid substitution is an amino acid substitution with an amino acid residue that is different from the amino acid residue at the position of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11.
[0262] In some embodiments, the amino acid substitution is an amino acid substitution with
[0263] (1) a non-polar amino acid residue (such as, Glycine (Gly / G) , Alanine (Ala / A) , Valine (Val / V) , Cysteine (Cys / C) , Proline (Pro / P) , Leucine (Leu / L) , Isoleucine (Ile / I) , Methionine (Met / M) , Tryptophan (Trp / W) , Phenylalanine (Phe / F) ,
[0264] (2) a polar amino acid residue (such as, Serine (Ser / S) , Threonine (Thr / T) , Tyrosine (Tyr / Y) , Asparagine (Asn / N) , Glutamine (Gln / Q) ) ,
[0265] (3) a positively charged amino acid residue (such as, Lysine (Lys / K) , Arginine (Arg / R) , Histidine (His / H) ) , or
[0266] (4) a negatively charged amino acid residue (such as, Aspartic Acid (Asp / D) , Glutamic Acid (Glue / E) ) .
[0267] In some embodiments, the amino acid substitution is an amino acid substitution with a positively charged amino acid residue, such as, Arginine (R) .
[0268] In some embodiments, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) .
[0269] Mutations for nickase
[0270] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 that is corresponding to a position selected from the group consisting of D10, D386, H387, N410, H509, and / or D512 of SEQ ID NO: 13. For example, the positions of SEQ ID NO: 9 corresponding to the positions of D10, D386, H387, N410, H509, and D512 of SEQ ID NO: 13 are positions D10, D373, H374, N397, H496, and D499 of SEQ ID NO: 9. In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position selected from the group consisting of D10, D386, H387, N410, H509, and / or D512 of SEQ ID NO: 13. In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position selected from the group consisting of D10, D373, H374, N397, H496, and / or D499 of SEQ ID NO: 9, respectively.
[0271] In some embodiments, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) .
[0272] In some embodiments, the amino acid substitution is relative to any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 and is corresponding to an amino acid substitution selected from the group consisting of D10A, D386A, H387A, N410A, H509A, and / or D512A relative to SEQ ID NO: 13. For example, the amino acid substitutions relative to SEQ ID NO: 9 corresponding to the amino acid substitutions of D10A, D386A, H387A, N410A, H509A, and D512A relative to SEQ ID NO: 13 are D10A, D373A, H374A, N397A, H496A, and D499A relative to SEQ ID NO: 9, respectively. In some embodiments, the amino acid substitution is selected from the group consisting of D10A, D386A, H387A, N410A, H509A, and / or D512A relative to SEQ ID NO: 13. In some embodiments, the amino acid substitution is selected from the group consisting of D10A, D373A, H374A, N397A, H496A, and / or D499A relative to SEQ ID NO: 9.
[0273] In some embodiments, the amino acid substitution is relative to any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 and is corresponding to an amino acid substitution of D10A or D386A relative to SEQ ID NO: 13. For example, the amino acid substitutions relative to SEQ ID NO: 9 corresponding to the amino acid substitutions of D10A and D386A relative to SEQ ID NO: 13 are D10A and D373A relative to SEQ ID NO: 9, respectively. In some embodiments, the amino acid substitution is D10A or D386A relative to SEQ ID NO: 13. In some embodiments, the amino acid substitution is D10A or D373A relative to SEQ ID NO: 9.
[0274] In some embodiments, the polypeptide comprising the amino acid mutation is a nickase or has nickase activity. In some embodiments, the amino acid mutation leads to increased nickase activity, or the polypeptide comprising the amino acid mutation has increased nickase activity, compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0275] Mutations of SN007 (MM762)
[0276] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of D10, D386, H387, N410, H509, and / or D512 of SEQ ID NO: 13.
[0277] In some embodiments, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) .
[0278] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of D10A, D386A, H387A, N410A, H509A, and / or D512A relative to SEQ ID NO: 13.
[0279] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of D10A or D386A relative to SEQ ID NO: 13.
[0280] In some embodiments, the polypeptide comprising the amino acid mutation is a nickase or has nickase activity.
[0281] In some embodiments, the amino acid mutation leads to increased nickase activity, or the polypeptide comprising the amino acid mutation has increased nickase activity, compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0282] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 41, 42, 45, 46, 49, 50, 53, 56, 57, 58, 64, 66, 69, 72, 73, 80, 81, 82, 83, 84, 86, 87, 92, 95, 96, 102, 105, 106, 107, 108, 110, 112, 113, 114, 115, 116, 117, 118, 120, 121, 128, 135, 136, 138, 140, 141, 145, 153, 154, 157, 162, 163, 185, 187, 198, 200, 246, 250, 255, 268, 344, 349, 355, 368, 408, 412, 431, 432, 434, 443, 446, 449, 450, 479, 500, 523, 532, 535, 537, 549, 550, 553, 557, 560, 564, 575, 576, 577, 582, 583, 588, 594, 614, 618, 622, 626, 636, 640, 644, 645, 683, 684, 697, 698, 699, 710, 716, 718, 731, 732, 739, 740, 742, 743, 745, 747, 749, 751, 754, 758, and / or 759 of SEQ ID NO: 13. In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 45, 46, 49, 120, 138, 200, 432, 549, 560, 582, 583, 640, 644, 710, and / or 747 of SEQ ID NO: 13.
[0283] In some embodiments, the amino acid substitution is an amino acid substitution with a positively charged amino acid residue, such as, Arginine (R) .
[0284] In some embodiments, the amino acid substitution is an amino acid substitution with Gly (G) .
[0285] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of 41R, 42R, 45R, 46R, 49R, 50R, 53R, 56R, 57R, 58R, 64R, 66R, 69R, 72R, 73R, 80R, 81R, 82R, 83R, 84R, 86R, 87R, 92R, 95R, 96R, 102R, 105R, 106R, 107R, 108R, 110R, 112R, 113R, 114R, 115R, 116R, 117R, 118R, 120R, 121R, 128R, 135R, 136R, 138R, 140R, 141R, 145R, 153R, 154R, 157R, 162R, 163R, 185R, 187R, 198R, 200R, 246R, 250R, 255R, 268R, 344R, 349R, 355R, 368R, 408R, 412R, 431R, 432R, 434R, 443R, 446R, 449R, 450R, 479R, 500R, 523R, 532R, 535R, 537R, 549R, 550R, 553R, 557R, 560R, 564R, 575R, 576R, 577R, 582R, 583R, 588R, 594R, 614R, 618R, 622R, 626R, 636R, 640R, 644R, 644G, 645R, 683R, 684R, 697R, 698R, 699R, 710R, 716R, 718R, 731R, 732R, 739R, 740R, 742R, 743R, 745R, 747R, 749R, 751R, 754R, 758R, and / or 759R relative to SEQ ID NO: 13.
[0286] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of A45R, G46R, E49R, D120R, Q138R, I200R, F432R, S549R, G560R, F582R, M583R, S640R, E644G, K710R, and / or T747R relative to SEQ ID NO: 13.
[0287] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of A45R, G46R, E49R, and T747R relative to SEQ ID NO: 13.
[0288] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, M583R, and T747R relative to SEQ ID NO: 13.
[0289] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, F432R, S640R, and T747R relative to SEQ ID NO: 13.
[0290] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, S549R, M583R, S640R, and T747R relative to SEQ ID NO: 13.
[0291] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, F432R, M583R, and T747R relative to SEQ ID NO: 13.
[0292] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, Q138R, M583R, S640R, and T747R relative to SEQ ID NO: 13.
[0293] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, Q138R, F432R, and T747R relative to SEQ ID NO: 13.
[0294] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, S549R, M583R, and T747R relative to SEQ ID NO: 13.
[0295] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, F432R, M583R, and T747R relative to SEQ ID NO: 13.
[0296] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, I200R, F432R, M583R, and T747R relative to SEQ ID NO: 13.
[0297] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of G46R, I200R, F432R, F582R, and T747R relative to SEQ ID NO: 13.
[0298] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of G46R, I200R, F582R, and T747R relative to SEQ ID NO: 13.
[0299] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of G46R, D120R, I200R, F432R, F582R, and T747R relative to SEQ ID NO: 13.
[0300] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of G46R, D120R, I200R, F432R, G560R, F582R, M583R, and T747R relative to SEQ ID NO: 13.
[0301] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of G46R, D120R, I200R, F432R, G560R, F582R, and T747R relative to SEQ ID NO: 13.
[0302] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, D120R, F432R, M583R, E644R, and T747R relative to SEQ ID NO: 13.
[0303] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, F432R, G560R, M583R, and T747R relative to SEQ ID NO: 13.
[0304] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, F432R, M583R, E644G, and T747R relative to SEQ ID NO: 13.
[0305] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, F432R, M583R, K710R, and T747R relative to SEQ ID NO: 13.
[0306] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of E49R, Q138R, F432R, M583R, and T747R relative to SEQ ID NO: 13.
[0307] In some embodiments, the amino acid mutation leads to an increased endonuclease activity, or the polypeptide comprising the amino acid mutation has an increased endonuclease activity, compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more. In some embodiments, the amino acid mutation leads to a decreased endonuclease activity, or the polypeptide comprising the amino acid mutation has a decreased endonuclease activity, compared to the reference polypeptide, e.g., a decrease by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%.
[0308] In some embodiments, the amino acid mutation comprises (a) an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of D10, D386, H387, N410, H509, and / or D512 of SEQ ID NO: 13, and optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ; and (b) an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of 41, 42, 45, 46, 49, 50, 53, 56, 57, 58, 64, 66, 69, 72, 73, 80, 81, 82, 83, 84, 86, 87, 92, 95, 96, 102, 105, 106, 107, 108, 110, 112, 113, 114, 115, 116, 117, 118, 120, 121, 128, 135, 136, 138, 140, 141, 145, 153, 154, 157, 162, 163, 185, 187, 198, 200, 246, 250, 255, 268, 344, 349, 355, 368, 408, 412, 431, 432, 434, 443, 446, 449, 450, 479, 500, 523, 532, 535, 537, 549, 550, 553, 557, 560, 564, 575, 576, 577, 582, 583, 588, 594, 614, 618, 622, 626, 636, 640, 644, 645, 683, 684, 697, 698, 699, 710, 716, 718, 731, 732, 739, 740, 742, 743, 745, 747, 749, 751, 754, 758, and / or 759 of SEQ ID NO: 13, and optionally, the amino acid substitution is an amino acid substitution with a positively charged amino acid residue, such as, Arginine (R) or Gly (G) .
[0309] In some embodiments, the amino acid mutation comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of D10A, D386A, H387A, N410A, H509A, and / or D512A relative to SEQ ID NO: 13; and (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of A45R, G46R, E49R, D120R, Q138R, I200R, F432R, S549R, G560R, F582R, M583R, S640R, E644G, K710R, and / or T747R relative to SEQ ID NO: 13.
[0310] In some embodiments, the amino acid mutation comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of D10A or D386A relative to SEQ ID NO: 13; and (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of:
[0311] (1) A45R, G46R, E49R, and T747R relative to SEQ ID NO: 13;
[0312] (2) E49R, M583R, and T747R relative to SEQ ID NO: 13;
[0313] (3) E49R, F432R, S640R, and T747R relative to SEQ ID NO: 13;
[0314] (4) E49R, S549R, M583R, S640R, and T747R relative to SEQ ID NO: 13;
[0315] (5) E49R, F432R, M583R, and T747R relative to SEQ ID NO: 13;
[0316] (6) E49R, Q138R, M583R, S640R, and T747R relative to SEQ ID NO: 13;
[0317] (7) E49R, Q138R, F432R, and T747R relative to SEQ ID NO: 13;
[0318] (8) E49R, S549R, M583R, and T747R relative to SEQ ID NO: 13;
[0319] (9) E49R, F432R, M583R, and T747R relative to SEQ ID NO: 13;
[0320] (10) E49R, I200R, F432R, M583R, and T747R relative to SEQ ID NO: 13;
[0321] (11) G46R, I200R, F432R, F582R, and T747R relative to SEQ ID NO: 13;
[0322] (12) G46R, I200R, F582R, and T747R relative to SEQ ID NO: 13;
[0323] (13) G46R, D120R, I200R, F432R, F582R, and T747R relative to SEQ ID NO: 13;
[0324] (14) G46R, D120R, I200R, F432R, G560R, F582R, M583R, and T747R relative to SEQ ID NO: 13;
[0325] (15) G46R, D120R, I200R, F432R, G560R, F582R, and T747R relative to SEQ ID NO: 13;
[0326] (16) E49R, D120R, F432R, M583R, E644R, and T747R relative to SEQ ID NO: 13;
[0327] (17) E49R, F432R, G560R, M583R, and T747R relative to SEQ ID NO: 13;
[0328] (18) E49R, F432R, M583R, E644G, and T747R relative to SEQ ID NO: 13;
[0329] (19) E49R, F432R, M583R, K710R, and T747R relative to SEQ ID NO: 13; or
[0330] (20) E49R, Q138R, F432R, M583R, and T747R relative to SEQ ID NO: 13.
[0331] In some embodiments, the polypeptide comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of D10A or D386A relative to SEQ ID NO: 13; and (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of (i) E49R, F432R, M583R, and T747R or (ii) E49R, I200R, F432R, M583R, and T747R relative to SEQ ID NO: 13. In some embodiments, the polypeptide comprises an amino acid sequence of SEQ ID NO: 21 (SN007-D386A, E49R, F432R, M583R, T747R) , SEQ ID NO: 22 (SN007-D10A, E49R, F432R, M583R, T747R) , or SEQ ID NO: 23 (eMM762v2 (D10A+E49R+I200R+F432R+M583R+T747R) ) .
[0332] In some embodiments, the polypeptide comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13. In some embodiments, the amino acid substitution is an amino acid substitution with P, T, L, A, or D. In some embodiments, the polypeptide comprises an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of (i) R391P, A393T, I396L, and G400A (hpHNHv22) or (ii) R391D, N397D, and G400A (hpHNHv8) relative to SEQ ID NO: 13.
[0333] In some embodiments, the polypeptide comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of D10A or D386A relative to SEQ ID NO: 13; (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of (i) E49R, F432R, M583R, and T747R or (ii) E49R, I200R, F432R, M583R, and T747R relative to SEQ ID NO: 13; and (c) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of (i) R391P, A393T, I396L, and G400A (hpHNHv22) or (ii) R391D, N397D, and G400A (hpHNHv8) relative to SEQ ID NO: 13.
[0334] In some embodiments, the most N-terminal Met amino acid residue (at position 1) of the polypeptide is deleted. In some embodiments, the polypeptide comprises an amino acid sequence of SEQ ID NO: 24 (eMM762v2-hpHNHv22 (R391P, A393T, I396L, and G400A) ) .
[0335] Mutation of SN005 (MM745)
[0336] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of D10, D373, H374, N397, H496, and / or D499 of SEQ ID NO: 9.
[0337] In some embodiments, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) .
[0338] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of D10A, D373A, H374A, N397A, H496A, and / or D499A relative to SEQ ID NO: 9.
[0339] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of D10A or D373A relative to SEQ ID NO: 9.
[0340] In some embodiments, the polypeptide comprising the amino acid mutation is a nickase or has nickase activity. In some embodiments, the amino acid mutation leads to increased nickase activity, or the polypeptide comprising the amino acid mutation has increased nickase activity, compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0341] In some embodiments, the polypeptide comprises an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of E44, D128, S546, Y568, M569, S623, and / or E667 of SEQ ID NO: 9.
[0342] In some embodiments, the amino acid substitution is an amino acid substitution with a positively charged amino acid residue, such as, Arginine (R) .
[0343] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of E44R, D128R, S546R, Y568R, M569R, S623R, and / or E667R relative to SEQ ID NO: 9.
[0344] In some embodiments, the amino acid substitution is corresponding to an amino acid substitution or is an amino acid substitution of
[0345] (1) E44R+ Y568R E667R relative to SEQ ID NO: 9;
[0346] (2) E44R+S546R+Y568R+E667R relative to SEQ ID NO: 9;
[0347] (3) E44R+S546R+M569R+E667R relative to SEQ ID NO: 9;
[0348] (4) E44R+D128R+S546R+Y568R+E667R relative to SEQ ID NO: 9; or
[0349] (5) E44R+D128R+S546R+Y568R+S623R+E667R relative to SEQ ID NO: 9.
[0350] In some embodiments, the amino acid mutation leads to an increased endonuclease activity, or the polypeptide comprising the amino acid mutation has an increased endonuclease activity, compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more. In some embodiments, the amino acid mutation leads to a decreased endonuclease activity, or the polypeptide comprising the amino acid mutation has a decreased endonuclease activity, compared to the reference polypeptide, e.g., a decrease by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%.
[0351] In some embodiments, the amino acid mutation comprises (a) an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of D10, D373, H374, N397, H496, and / or D499 of SEQ ID NO: 9, and optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ; and (b) an amino acid mutation (e.g., substitution) at a position that is corresponding to a position or that is a position selected from the group consisting of E44, D128, S546, Y568, M569, S623, and / or E667 of SEQ ID NO: 9, and optionally, the amino acid substitution is an amino acid substitution with a positively charged amino acid residue, such as, Arginine (R) . In some embodiments, the amino acid mutation comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of D10A, D373A, H374A, N397A, H496A, and / or D499A relative to SEQ ID NO: 9; and (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution selected from the group consisting of E44R, D128R, S546R, Y568R, M569R, S623R, and / or E667R relative to SEQ ID NO: 9.
[0352] In some embodiments, the amino acid mutation comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of D10A or D373A relative to SEQ ID NO: 9; and (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of:
[0353] (1) E44R+ Y568R E667R relative to SEQ ID NO: 9;
[0354] (2) E44R+S546R+Y568R+E667R relative to SEQ ID NO: 9;
[0355] (3) E44R+S546R+M569R+E667R relative to SEQ ID NO: 9;
[0356] (4) E44R+D128R+S546R+Y568R+E667R relative to SEQ ID NO: 9; or
[0357] (5) E44R+D128R+S546R+Y568R+S623R+E667R relative to SEQ ID NO: 9.
[0358] In some embodiments, the polypeptide comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of D10A relative to SEQ ID NO: 9; and (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of E44R+D128R+S546R+Y568R+S623R+E667R relative to SEQ ID NO: 9.
[0359] In some embodiments, the polypeptide comprises an amino acid substitution at a position of SEQ ID NO: 9 that is corresponding to a position or that is a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13.
[0360] In some embodiments, the polypeptide comprises an amino acid substitution relative to SEQ ID NO: 9 that is corresponding to an amino acid substitution of (i) R391P, A393T, I396L, and G400A (hpHNHv22) or (ii) R391D, N397D, and G400A (hpHNHv8) relative to SEQ ID NO: 13.
[0361] In some embodiments, the polypeptide comprises (a) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of D10A relative to SEQ ID NO: 9; (b) an amino acid substitution that is corresponding to an amino acid substitution or is an amino acid substitution of E44R+D128R+S546R+Y568R+S623R+E667R relative to SEQ ID NO: 9; and (c) an amino acid substitution relative to SEQ ID NO: 9 that is corresponding to an amino acid substitution of (i) R391P, A393T, I396L, and G400A (hpHNHv22) or (ii) R391D, N397D, and G400A (hpHNHv8) relative to SEQ ID NO: 13.
[0362] In some embodiments, the most N-terminal Met amino acid residue (at position 1) of the polypeptide is deleted.
[0363] In some embodiments, the polypeptide comprises an amino acid sequence of SEQ ID NO: 71 (SN005-D10A+E44R+D128R+S546R+Y568R+S623R+E667R) (enSN005) .
[0364] Effect of mutation
[0365] In some embodiments, the amino acid mutation leads to increased nickase activity, or the polypeptide comprising the amino acid mutation has increased nickase activity, compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0366] In some embodiments, the amino acid mutation leads to increased editing efficiency (e.g., base editing efficiency, prime editing efficiency, epigenomic editing efficiency) , or the polypeptide comprising the amino acid mutation or the fusion protein comprising the polypeptide has increased editing efficiency, compared to the reference polypeptide or an otherwise identical fusion protein comprising the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0367] In some embodiments, the amino acid mutation leads to a decreased guide sequence-independent (off-target) endonuclease activity, or the polypeptide comprising the amino acid mutation has a decreased guide sequence-independent (off-target) endonuclease activity, compared to the reference polypeptide, e.g., a decrease by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%.
[0368] In some embodiments, the polypeptide is endonuclease deficient. In some embodiments, the polypeptide is catalytically inactive.
[0369] In some embodiments, the polypeptide is capable of recognizing a protospacer adjacent motif (PAM) comprising, consisting essentially of, or consisting of 5’ -NNN-3’ immediately 3’ to a protospacer sequence of a target DNA, wherein N is A, T, G, or C. In some embodiments, the polypeptide is capable of recognizing a PAM comprising, consisting essentially of, or consisting of 5’ -NGG-3’ immediately 3’ to a protospacer sequence of a target DNA, wherein N is A, T, G, or C.
[0370] Fusion proteins
[0371] The polypeptide can be fused to a functional domain to form a fusion protein. The functional domain is usually heterologous to the polypeptide. Therefore, in an aspect, the disclosure provides a fusion protein comprising the polypeptide and a functional domain. In some embodiments, the polypeptide is fused to a functional domain to form a fusion protein. The fusion protein can be considered as a mutant of the polypeptide with an amino acid addition at the N-terminal and / or C-terminal of the polypeptide or with an amino acid insertion within the polypeptide. The polypeptide in the disclosure may also refer to the fusion of the disclosure depending on the context.
[0372] In some embodiments, the functional domain is fused to the N-terminal of (N-terminally fused to) or the C-terminal of (C-terminally fused to) the polypeptide or inserted into (fused internally to) the polypeptide. In some embodiments, the functional domain is fused to the polypeptide via a linker, e.g., a XTEN linker, a GS linker. As used herein, the term “GS linker” refers to a linker consisting of one or more Gly (G; glycine) and one or more Ser (S; serine) in any sequence. In some embodiments, a GS linker may contain an additional sequence, e.g., a XTEN linker, a NLS, within the linker, such as, SEQ ID NOs: 29 and 39. In some embodiments, the linker is selected from GSG, SGGS, SEQ ID NOs: 29, 35, 39, and 56-63, and optionally the linker of SEQ ID NO: 63. In some embodiments, the fusion protein contains the polypeptide and more than one (e.g., 2, 3, 4, 6, or more) functional domain. In some embodiments, two functional domains are fused together via a linker in the disclosure. In some embodiments, the functional domain has transposase activity, methylase activity, demethylase activity, translation activation activity, translation repression activity, transcription activation activity, transcription repression activity, transcription release factor activity, chromatin modifying or remodeling activity, histone modification activity, nuclease activity, single-strand RNA cleavage activity, double-strand RNA cleavage activity, single-strand DNA cleavage activity, double-strand DNA cleavage activity, nucleic acid binding activity, detectable activity, or any combination thereof.
[0373] In some embodiments, the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a deaminase or a catalytic domain thereof, an uracil glycosylase inhibitor (UGI) , an uracil glycosylase (UNG) , a methylpurine glycosylase (MPG) , a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease or a catalytic domain thereof (e.g., T5 exonuclease) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having RNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment thereof, and any combination thereof.
[0374] In some embodiments, the fusion protein comprises a NLS at the N-terminal and / or the C-terminal of the polypeptide. In some embodiments, the fusion protein comprises one or two NLS at the N-terminal and / or the C-terminal of the polypeptide.
[0375] In some embodiments, the fusion protein comprises a NLS at the N-terminal and / or the C-terminal of the functional domain. In some embodiments, the fusion protein comprises one or two NLS at the N-terminal and / or the C-terminal of the functional domain.
[0376] In some embodiments, the NLS comprises or is SV40 NLS, bpSV40 NLS (BP NLS, bpNLS) (SEQ ID NO: 27 or 34) , NP NLS (Xenopus laevis Nucleoplasmin NLS, nucleoplasmin NLS) (SEQ ID NO: 28) , or c-myc NLS (SEQ ID NO: 38) . In some embodiments, the NLS is selected from the group consisting of SEQ ID NO: 27, 28, 34, and 38.
[0377] Base editor
[0378] In some embodiments, the functional domain comprises a deaminase or a catalytic domain thereof.
[0379] In some embodiments, the functional domain does not comprise a deaminase or a catalytic domain thereof.
[0380] In some embodiments, the functional domain comprises an uracil glycosylase (UNG) .
[0381] In some embodiments, the functional domain comprises a methylpurine glycosylase (MPG) .
[0382] The polypeptide of the disclosure can be used to replace the napDNAbp / napDNAbd in PCT / CN2023 / 094023 (AYBE) and PCT / CN2024 / 089874 (gBE) to constitute AYBE base editor and gBE base editor, respectively, which two PCT applications are incorporated herein by reference in their entities.
[0383] Adenine Base editor
[0384] In some embodiments, the deaminase or catalytic domain thereof is an adenine deaminase or a catalytic domain thereof (e.g., tRNA adenosine deaminase (TadA) , such as, TadA8e, TadA8.17, TadA8.20, TadA9, TadA8e-V106W, TadA8EV106W+D108Q TadA-CDa, TadA-CDb, TadA-CDc, TadA-CDd, TadA-CDe, TadA-dual, TADAC-1.2, TADAC-1.14, TADAC-1.17, TADAC-1.19, TADAC-2.5, TADAC-2.6, TADAC-2.9, TADAC-2.19, TADAC-2.23, TadA8e-N46L, TadA8e-N46P, TadA* (8.17m) , TadA (8.8m) ) .
[0385] In some embodiments, the adenine deaminase or a catalytic domain thereof is TadA8EV106W.
[0386] In some embodiments, the adenine deaminase or a catalytic domain thereof comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 26.
[0387] In some embodiments, the adenine deaminase or a catalytic domain thereof is set forth in SEQ ID NO: 26 (TadA8EV106W) .
[0388] In some embodiments, the fusion protein comprises the polypeptide and the adenine deaminase or a catalytic domain thereof.
[0389] In some embodiments, the fusion protein comprises, from N-to C-terminus, an optional NLS, the adenine deaminase or a catalytic domain thereof, an optional linker, the polypeptide, and an optional NLS.
[0390] In some embodiments, the fusion protein comprises a DNA binding domain. In some embodiments, the DNA binding domain comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 72 or 54 (Sto7d*) .
[0391] In some embodiments, the fusion protein comprises the DNA binding domain at the N-terminal or C-terminal of the adenine deaminase or a catalytic domain thereof, or at the N-terminal or C-terminal of the polypeptide.
[0392] In some embodiments, the fusion protein comprises, from N-to C-terminus, an optional NLS, the DNA binding domain, an optional linker, the adenine deaminase or a catalytic domain thereof, an optional linker, the polypeptide, an optional linker, and an optional NLS.
[0393] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 25, 49-53, 70, and 73.
[0394] Cytosine Base editor
[0395] In some embodiments, the deaminase or catalytic domain thereof is a cytosine deaminase or a catalytic domain thereof (e.g., an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase, an activation induced deaminase (AID) , a cytidine deaminase 1 from Petromyzon marinus (pmCDA1) , DddA, or a functional variant thereof, e.g., APOBEC1 (rAPOBEC1) , APOBEC2, APOBEC3, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, hAPOBEC3-W104A, CBE6c-V106W) .
[0396] In some embodiments, the cytidine deaminase or a catalytic domain thereof comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 32 (AjTadA cytosine deaminase) or 68 (CBE6c-V106W) .
[0397] In some embodiments, the cytidine deaminase or a catalytic domain thereof is set forth in SEQ ID NO: 32 or 68.
[0398] In some embodiments, the cytidine deaminase or a catalytic domain thereof is any deaminase mentioned in PCT / CN2024 / 078613 or any PCT application claims the priority of PCT / CN2024 / 078613.
[0399] In some embodiments, the functional domain comprises an uracil glycosylase inhibitor (UGI) domain.
[0400] In some embodiments, the UGI domain comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 33.
[0401] In some embodiments, the fusion protein comprises one, two, or three UGI domains.
[0402] In some embodiments, the fusion protein comprises the polypeptide, the cytidine deaminase or a catalytic domain thereof, and the UGI domain.
[0403] In some embodiments, the fusion protein comprises, from N-to C-terminus, an optional NLS, the cytidine deaminase or a catalytic domain thereof, an optional linker, the polypeptide, an optional linker, the UGI domain, an optional linker, optionally the UGI domain, an optional linker, and an optional NLS.
[0404] In some embodiments, the fusion protein comprises a DNA binding domain. In some embodiments, the DNA binding domain comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 72 or 54 (Sto7d*) .
[0405] In some embodiments, the fusion protein comprises the DNA binding domain at the N-terminal or C-terminal of the cytosine deaminase or a catalytic domain thereof, or at the N-terminal or C-terminal of the polypeptide.
[0406] In some embodiments, the fusion protein comprises, from N-to C-terminus, an optional NLS, the DNA binding domain, an optional linker, the cytidine deaminase or a catalytic domain thereof, an optional linker, the polypeptide, an optional linker, the UGI domain, an optional linker, optionally the UGI domain, an optional linker, and an optional NLS.
[0407] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 31 and 64-67.
[0408] Prime editor
[0409] In some embodiments, the functional domain comprises a reverse transcriptase or a catalytic domain thereof.
[0410] In some embodiments, the fusion protein comprises the polypeptide and the reverse transcriptase or a catalytic domain thereof.
[0411] In some embodiments, the reverse transcriptase or a catalytic domain thereof comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 37.
[0412] In some embodiments, the fusion protein comprises, from N-to C-terminus, an optional NLS, the polypeptide, an optional linker, the reverse transcriptase or a catalytic domain thereof, an optional linker, an optional NLS, an optional linker, and an optional NLS.
[0413] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 36.
[0414] Epigenomic editor
[0415] In some aspects, the disclosure provides a way to epigenomic modification of a target gene, e.g., methylation, to regulate the gene. The epigenomic modification, in some embodiments, silences the expression of the gene, leading to reduced level of a corresponding mRNA and / or reduced level of a corresponding protein.
[0416] In some embodiments, the fusion protein comprises a transcription inhibiting domain (e.g., KRAB domain or SID domain) .
[0417] In some embodiments, the fusion protein comprises a KRAB domain.
[0418] In some embodiments, the KRAB domain comprises, consists essentially of, or consists of a sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to the sequence of SEQ ID NO: 41.
[0419] In some embodiments, the fusion protein comprises a DNA methyltransferase, such as, DNMT3l, DNMT3a.
[0420] In some embodiments, the fusion protein comprises a DNMT3l domain and a DNMT3a domain.
[0421] In some embodiments, the DNMT3l domain comprises, consists essentially of, or consists of a sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to the sequence of SEQ ID NO: 42.
[0422] In some embodiments, the DNMT3a domain comprises, consists essentially of, or consists of a sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to the sequence of SEQ ID NO: 43.
[0423] In some embodiments, the fusion protein comprises the polypeptide, a DNMT3l domain, a DNMT3a domain, and a KRAB domain.
[0424] In some embodiments, the fusion protein comprises, from N-terminal to C-terminal, the polypeptide, the KRAB domain, the DNMT3l domain, and the DNMT3a domain.
[0425] In some embodiments, the fusion protein comprises, consists essentially of, or consists of a sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to the sequence of SEQ ID NO: 40.
[0426] Representative systems
[0427] The polypeptide (including the fusion) of the disclosure can be used in combination with a guide nucleic acid as described herein to constitute a system comprising the polypeptide (including the fusion) and the guide nucleic acid. The system in the disclosure is non-naturally occurring as it requires a guide sequence targeting to a target DNA heterologous to the scaffold sequence.
[0428] In an aspect, the disclosure provides a system comprising:
[0429] (1) the polypeptide of the disclosure or the fusion protein of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the polypeptide or the fusion protein, and
[0430] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0431] (i) a scaffold sequence capable of forming a complex with the polypeptide or the fusion protein; and
[0432] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.
[0433] In some embodiments, the target DNA is a dsDNA.
[0434] In some embodiments, the system is a complex comprising the polypeptide complexed with the guide nucleic acid.
[0435] In some embodiments, the complex further comprises the target DNA hybridized with the guide sequence.
[0436] In some embodiments, the system is a composition comprising (1) the polypeptide and (2) the guide nucleic acid.
[0437] In some embodiments, the scaffold sequence is 3’ to the guide sequence. In some embodiments, the guide sequence is 5’ to the scaffold sequence.
[0438] In some embodiments, the guide nucleic acid is a guide RNA (gRNA) .
[0439] In some embodiments, the system further comprises a donor polynucleotide for integration or insertion into the target DNA.
[0440] In the disclosure, for any polypeptide based on SN001 (SEQ ID NO: 1) (e.g., SN001 per se or a mutant (e.g., nickase, fusion protein) thereof) , the corresponding scaffold sequence can be any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, or 14, or a mutant thereof (e.g., any one of SEQ ID NOs: 44-48 and 10) .
[0441] In the disclosure, for any polypeptide based on SN002 (SEQ ID NO: 3) (e.g., SN002 per se or a mutant (e.g., nickase, fusion protein) thereof) , the corresponding scaffold sequence can be any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, or 14, or a mutant thereof (e.g., any one of SEQ ID NOs: 44-48 and 10) .
[0442] In the disclosure, for any polypeptide based on SN003 (SEQ ID NO: 5) (e.g., SN003 per se or a mutant (e.g., nickase, fusion protein) thereof) , the corresponding scaffold sequence can be any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, or 14, or a mutant thereof (e.g., any one of SEQ ID NOs: 44-48 and 10) .
[0443] In the disclosure, for any polypeptide based on SN004 (SEQ ID NO: 7) (e.g., SN004 per se or a mutant (e.g., nickase, fusion protein) thereof) , the corresponding scaffold sequence can be any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, or 14, or a mutant thereof (e.g., any one of SEQ ID NOs: 44-48 and 10) .
[0444] In the disclosure, for any polypeptide based on SN005 (SEQ ID NO: 9) (e.g., SN005 per se or a mutant (e.g., nickase, fusion protein) ) thereof, e.g., enSN005 (SEQ ID NO: 71) , the corresponding scaffold sequence can be any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, or 14, or a mutant thereof (e.g., any one of SEQ ID NOs: 44-48 and 10) . In the disclosure, for any polypeptide based on SN006 (SEQ ID NO: 11) (e.g., SN006 per se or a mutant (e.g., nickase, fusion protein) thereof) , the corresponding scaffold sequence can be any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, or 14, or a mutant thereof (e.g., any one of SEQ ID NOs: 44-48 and 10) .
[0445] In the disclosure, for any polypeptide based on SN007 (SEQ ID NO: 13) (e.g., SN007 per se or a mutant (e.g., nickase, fusion protein) thereof, e.g., any one of SEQ ID NOs: 21-24) , the corresponding scaffold sequence can be any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, or 14, or a mutant thereof (e.g., any one of SEQ ID NOs: 44-48 and 10) . For example, for enSN005 (SEQ ID NO: 71) , both the scaffold sequence of SEQ ID NO: 10 and the scaffold sequence of SEQ ID NO: 44 can be used to construct the gRNA for use in combination with enSN005.
[0446] Scaffold sequence
[0447] For the purpose of the disclosure, the scaffold sequence is compatible with the polypeptide of the disclosure and is capable of complexing with the polypeptide. The scaffold sequence may be a naturally occurring scaffold sequence identified along with the polypeptide, or a variant thereof maintaining the ability to complex with the polypeptide. Generally, the ability to complex with the polypeptide is maintained as long as the secondary structure of the variant is substantially identical to the secondary structure of the naturally occurring scaffold sequence. A nucleotide deletion, insertion, or substitution in the primary sequence of the scaffold sequence may not necessarily change the secondary structure of the scaffold sequence (e.g., the relative locations and / or sizes of the stems, bulges, and loops of the scaffold sequence do not significantly deviate from that of the original stems, bulges, and loops) . For example, the nucleotide deletion, insertion, or substitution may be in a bulge or loop region of the scaffold sequence so that the overall symmetry of the bulge and hence the secondary structure remains largely the same. The nucleotide deletion, insertion, or substitution may also be in the stems of the scaffold sequence so that the lengths of the stems do not significantly deviate from that of the original stems (e.g., adding or deleting one base pair in each of two stems correspond to 4 total base changes) . On the other hand, engineering of the scaffold sequence may be applied to improve the activity of system. The disclosure provides engineered scaffold sequence leading to improved endonuclease activity of the polypeptide when used together. In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0448] In some embodiments, the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0449] In some embodiments, the scaffold sequence leads to an increased guide sequence-specific (on-target) endonuclease activity compared to that led by SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48 when both are used in otherwise identical guide nucleic acid in combination with a same polypeptide (e.g., the polypeptide of the disclosure) , e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, or more.
[0450] In some embodiments, the scaffold sequence comprises a deletion in a stem-loop region of the scaffold sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0451] In some embodiments, the scaffold sequence comprises a deletion of about or at least about or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 nucleotides, or in a numerical range between any two of the preceding values, e.g., from about 5 to about 18 nucleotides. In some embodiments, the scaffold sequence comprises a deletion of about 15 nucleotides.
[0452] In some embodiments, the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48, or a polynucleotide sequence having a sequence identity of at least about 80%(e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the polynucleotide sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0453] In some embodiments, the scaffold sequence comprises a base pair substitution of a thermodynamically unstable base pair (e.g., a A-T base pair or a mismatched base pair (e.g., a A-G base pair, a T-G base pair) ) with a G-C base pair relative to the scaffold sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0454] In some embodiments, the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48, or a polynucleotide sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the polynucleotide sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0455] In some embodiments, the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0456] In yet another aspect, the disclosure provides a guide nucleic acid comprising the scaffold sequence of the disclosure.
[0457] In an aspect, the disclosure provides a reference scaffold sequence as set forth in SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, and 44-48.
[0458] Protospacer / target sequence
[0459] In some embodiments, the protospacer sequence comprises about or at least about 14 contiguous nucleotides of the target DNA, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more contiguous nucleotides of the target DNA, or in a numerical range between any two of the preceding values, e.g., from about 14 to about 50, or from about 19 to about 24 contiguous nucleotides of the target DNA. In some embodiments, the protospacer sequence comprises about 20 contiguous nucleotides of the target DNA. As used herein, in the context of a target dsDNA, the protospacer sequence is on the nontarget strand of the target dsDNA.
[0460] In some embodiments, the protospacer sequence is immediately 5’ to a protospacer adjacent motif (PAM) . In some embodiments, the PAM is 5’ -NGG-3’ , wherein N is A, T, G, or C.
[0461] In some embodiments, the target sequence comprises about or at least about 14 contiguous nucleotides of the target DNA, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more contiguous nucleotides of the target DNA, or in a numerical range between any two of the preceding values, e.g., from about 14 to about 50, or from about 19 to about 24 contiguous nucleotides of the target DNA. In some embodiments, the target sequence comprises about 20 contiguous nucleotides of the target DNA. As used herein, in the context of a target dsDNA, the target sequence is on the target strand of the target dsDNA.
[0462] In some embodiments, the target sequence is immediately 3’ to a PAM, or the reversely complementary sequence of the target sequence (i.e., the protospacer sequence) is immediately 5’ to a PAM. In some embodiments, the PAM is 5’ -NGG-3’ , wherein N is A, T, G, or C.
[0463] Guide sequence
[0464] In some embodiments, the guide sequence is about or at least about 14 nucleotides in length, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more nucleotides in length, or in a length of a numerical range between any two of the preceding values, e.g., in a length of from about 14 to about 50 nucleotides, or from about 19 to about 24 nucleotides. In some embodiments, the guide sequence is about 20 nucleotides in length.
[0465] In some embodiments, (1) the guide sequence is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% (fully) reverse complementary to the target sequence; (2) the guide sequence contains no more than 5, 4, 3, 2, or 1 mismatch or contains no mismatch with the target sequence; or (3) the guide sequence comprises no mismatch with the target sequence in the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nucleotides at the 3’ end of the guide sequence. In some embodiments, (1) the guide sequence is about 100% (fully) reverse complementary to the target sequence.
[0466] In some embodiments, the system comprises two or more guide nuclei acids comprising two or more guide sequences capable of hybridizing to two or more target sequences of the same target DNA or different target DNAs, wherein the two or more guide sequences are the same or different, and wherein the two or more target sequences are the same or different.
[0467] In some embodiments, the guide sequence is any one of SEQ ID NOs: 80-260.
[0468] In yet another aspect, the disclosure provides a guide nucleic acid comprising the guide sequence of the disclosure.
[0469] In yet another aspect, the disclosure provides a guide RNA comprising a guide sequence of any one of SEQ ID NOs: 80-260 5’ to a scaffold sequence of SpCas9.
[0470] In yet another aspect, the disclosure provides a guide RNA comprising a guide sequence of any one of SEQ ID NOs: 80-260 5’ to a scaffold sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48. For example, in some embodiments, the guide RNA comprises a guide sequence of any one of SEQ ID NOs: 80-260 5’ to a scaffold sequence of SEQ ID NOs: 44, e.g., the gRNA of SEQ ID NO: 261 or 262.
[0471] In yet another aspect, the disclosure provides a composition comprising two, three, four, or more of the guide RNAs of the disclosure. In some embodiments, the composition further comprises a Cas9 (e.g., SpCas9) or a mutant thereof, or the polypeptide or the fusion protein of the disclosure.
[0472] Target DNA
[0473] In some embodiments, the target DNA is a target dsDNA, such as, a eukaryotic dsDNA, e.g., a gene in a eukaryotic cell.
[0474] In some embodiments, the target dsDNA comprises a protospacer sequence on a nontarget strand of the target dsDNA, wherein the dsDNA comprises a target deoxyribonucleotide (e.g., dA, dT, dC, dG) at a position of the protospacer sequence selected from the group consisting of position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, position 11, position 12, position 13, position 14, and a combination thereof; or wherein the target deoxyribonucleotide is at a position of the protospacer sequence between position 2 and position 12 or between position 3 and position 14, both inclusive.
[0475] In some embodiments, the target deoxyribonucleotide is the N2 nucleotide in a motif of N1N2N3, wherein N1, N2, or N3 is A, T, G, or C.
[0476] Polynucleotides
[0477] In yet another aspect, the disclosure provides a polynucleotide comprising a sequence encoding the polypeptide or the fusion protein of the disclosure. In some embodiments, the polynucleotide encodes a guide nucleic acid as described herein.
[0478] In yet another aspect, the disclosure provides a polynucleotide comprising or encoding the guide nucleic acid of the disclosure.
[0479] Regulation of guide nucleic acid
[0480] In some embodiments, the polynucleotide encoding the guide nucleic acid is a DNA, a RNA, or a DNA / RNA mixture. By “DNA / RNA mixture” it refers to a nucleic acid comprising both one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, whether consecutive or not. However, by “DNA” or “RNA” it may also refer to a DNA containing one or more modified or unmodified ribonucleotides, whether consecutive or not, or an RNA containing one or more modified or unmodified deoxyribonucleotides, whether consecutive or not.
[0481] In some embodiments, the guide nucleic acid is operably linked to or under the regulation of a promoter.
[0482] In some embodiments, the promoter is a ubiquitous, tissue-specific, cell-type specific, constitutive, or inducible promoter.
[0483] Suitable promoters are known in the art and include, for example, a Cbh promoter, a Cba promoter, a pol I promoter, a pol II promoter, a pol III promoter, a T7 promoter, a U6 promoter, a H1 promoter, a retroviral Rous sarcoma virus LTR promoter, a cytomegalovirus (CMV) promoter, a SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, an elongation factor 1α short (EFS) promoter, a β glucuronidase (GUSB) promoter, a cytomegalovirus (CMV) immediate-early (Ie) enhancer and / or promoter, a chicken β-actin (CBA) promoter or derivative thereof such as a CAG promoter, CB promoter, a (human) elongation factor 1α-subunit (EF1α) promoter, a ubiquitin C (UBC) promoter, a prion promoter, a neuron-specific enolase (NSE) , a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a platelet-derived growth factor (PDGF) promoter, a platelet-derived growth factor B-chain (PDGF-β) promoter, a synapsin (Syn) promoter, a synapsin 1 (Syn1) promoter, a methyl-CpG binding protein 2 (MeCP2) promoter, a Ca2+ / calmodulin-dependent protein kinase II (CaMKII) promoter, a metabotropic glutamate receptor 2 (mGluR2) promoter, a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a β-globin minigene nβ2 promoter, a preproenkephalin (PPE) promoter, an enkephalin (Enk) promoter, an excitatory amino acid transporter 2 (EAAT2) promoter, a glial fibrillary acidic protein (GFAP) promoter, and a myelin basic protein (MBP) promoter.
[0484] Regulation of polypeptide
[0485] In some embodiments, the polynucleotide encoding the polypeptide is a DNA, a RNA, or a DNA / RNA mixture.
[0486] In some embodiments, the polynucleotide encoding the polypeptide is operably linked to or under the regulation of a promoter.
[0487] In some embodiments, the promoter is a ubiquitous, tissue-specific, cell-type specific, constitutive, or inducible promoter.
[0488] Suitable promoters are known in the art and include, for example, a Cbh promoter, a Cba promoter, a pol I promoter, a pol II promoter, a pol III promoter, a T7 promoter, a U6 promoter, a H1 promoter, a retroviral Rous sarcoma virus LTR promoter, a cytomegalovirus (CMV) promoter, a SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, an elongation factor 1α short (EFS) promoter, a β glucuronidase (GUSB) promoter, a cytomegalovirus (CMV) immediate-early (Ie) enhancer and / or promoter, a chicken β-actin (CBA) promoter or derivative thereof such as a CAG promoter, CB promoter, a (human) elongation factor 1α-subunit (EF1α) promoter, a ubiquitin C (UBC) promoter, a prion promoter, a neuron-specific enolase (NSE) , a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a platelet-derived growth factor (PDGF) promoter, a platelet-derived growth factor B-chain (PDGF-β) promoter, a synapsin (Syn) promoter, a human synapsin (hSyn) promoter, a synapsin 1 (Syn1) promoter, a methyl-CpG binding protein 2 (MeCP2) promoter, a Ca2+ / calmodulin-dependent protein kinase II (CaMKII) promoter, a metabotropic glutamate receptor 2 (mGluR2) promoter, a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a β-globin minigene nβ2 promoter, a preproenkephalin (PPE) promoter, an enkephalin (Enk) promoter, an excitatory amino acid transporter 2 (EAAT2) promoter, a glial fibrillary acidic protein (GFAP) promoter, a myelin basic protein (MBP) promoter, a OTOF promoter, a GRK1 promoter, a CRX promoter, a NRL promoter, a MECP2 promoter, a mMECP2 promoter, a hMECP2 promoter, an APP promoter, and a RCVRN promoter.
[0489] Delivery
[0490] Various ways of delivery can be applied to the polypeptide of the disclosure or the system of the disclosure as needed in practices.
[0491] In yet another aspect, the disclosure provides a delivery system comprising (1) the polypeptide of the disclosure, the fusion protein of the disclosure, the system of the disclosure, or the polynucleotide of the disclosure; and (2) a delivery vehicle.
[0492] In yet another aspect, the disclosure provides a vector comprising the polynucleotide of the disclosure. In some embodiments, the vector encodes a guide nucleic acid as described herein. In some embodiments, the vector is a plasmid vector, a viral vector (e.g., a recombinant AAV (rAAV) vector, a recombinant lentivirus vector) , a ribonucleoprotein (RNP) , or a lipid nanoparticle (LNP) .
[0493] In yet another aspect, the disclosure provides a recombinant AAV (rAAV) particle comprising the rAAV vector of the disclosure. In some embodiments, the rAAV vector is an RNA. A simple introduction of AAV for delivery may refer to “Adeno-associated Virus (AAV) Guide” (addgene. org / guides / aav / ) .
[0494] Adeno-associated virus (AAV) , when engineered to delivery, e.g., a protein-encoding sequence of interest, may be termed as a (r) AAV vector, a (r) AAV vector particle, or a (r) AAV particle, where “r” stands for “recombinant” . And the nucleic acid packaged in AAV vectors for delivery may be termed as a (r) AAV vector genome, vector genome, or vg for short, while viral genome may refer to the original viral genome of natural AAVs.
[0495] The serotypes of the capsids of rAAV particles can be matched to the types of target cells. For example, Table 2 of WO2018002719A1 lists exemplary cell types that can be transduced by the indicated AAV serotypes (incorporated herein by reference) .
[0496] In some embodiments, the rAAV particle comprising a capsid with a serotype suitable for delivery into ear cells (e.g., inner hair cells) . In some embodiments, the rAAV particle comprising a capsid with a serotype of AAV1, AAV2, AAV3A, AAV3B, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV-DJ, or AAV. PHP. eB, a member of the Clade to which any of the AAV1-AAV13 belong, or a functional variant (e.g., a functional truncation) thereof, encapsidating the rAAV vector genome. In some embodiments, the serotype of the capsid is AAV9 or a functional variant thereof.
[0497] General principles of rAAV particle production are known in the art. In some embodiments, rAAV particles may be produced using the triple transfection method (described in detail in U.S. Pat. No. 6,001,650) .
[0498] The vector titers are usually expressed as vector genomes per ml (vg / ml) . In some embodiments, the vector titer is above 1×109, above 5×1010, above 1×1011, above 5×1011, above 1×1012, above 5×1012, or above 1×1013 vg / ml. Instead of packaging a single strand (ss) DNA as a vector genome of a rAAV particle, systems and methods of packaging an RNA as a vector genome into a rAAV particle is recently developed and applicable herein. See PCT / CN2022 / 075366, which is incorporated herein by reference in its entirety.
[0499] When the vector genome is RNA as in, for example, PCT / CN2022 / 075366, for simplicity of description and claiming, sequence elements described herein for DNA vector genomes, when present in RNA vector genomes, should generally be considered to be applicable for the RNA vector genomes except that the deoxyribonucleotides in the DNA sequence are the corresponding ribonucleotides in the RNA sequence (e.g., dT is equivalent to U, and dA is equivalent to A) and / or the element in the DNA sequence is replaced with the corresponding element with a corresponding function in the RNA sequence or omitted because its function is unnecessary in the RNA sequence and / or an additional element necessary for the RNA vector genome is introduced.
[0500] As used herein, a coding sequence, e.g., as a sequence element of rAAV vector genomes herein, is construed, understood, and considered as covering and covers both a DNA coding sequence and an RNA coding sequence. When it is a DNA coding sequence, an RNA sequence can be transcribed from the DNA coding sequence, and optionally further a protein can be translated from the transcribed RNA sequence as necessary. When it is an RNA coding sequence, the RNA coding sequence per se can be a functional RNA sequence for use, or an RNA sequence can be produced from the RNA coding sequence, e.g., by RNA processing, or a protein can be translated from the RNA coding sequence.
[0501] For example, a polypeptide coding sequence encoding a polypeptide covers either a polypeptide DNA coding sequence from which a polypeptide is expressed (indirectly via transcription and translation) or a polypeptide RNA coding sequence from which a polypeptide is translated (directly) .
[0502] For example, a gRNA coding sequence encoding a gRNA covers either a gRNA DNA coding sequence from which a gRNA is transcribed or a gRNA RNA coding sequence (1) which per se is the functional gRNA for use, or (2) from which a gRNA is produced, e.g., by RNA processing.
[0503] In some embodiments for rAAV RNA vector genomes, 5’ -ITR and / or 3’ -ITR as DNA packaging signals may be unnecessary and can be omitted at least partly, while RNA packaging signals can be introduced. In some embodiments for rAAV RNA vector genomes, a promoter to drive transcription of DNA sequences may be unnecessary and can be omitted at least partly. In some embodiments for rAAV RNA vector genomes, a sequence encoding a polyA signal may be unnecessary and can be omitted at least partly, while a polyA tail can be introduced. Similarly, other DNA elements of rAAV DNA vector genomes can be either omitted or replaced with corresponding RNA elements and / or additional RNA elements can be introduced, in order to adapt to the strategy of delivering an RNA vector genome by rAAV particles.
[0504] In yet another aspect, the disclosure provides a ribonucleoprotein (RNP) comprising the polypeptide of the disclosure or the fusion protein of the disclosure and a guide nucleic acid. In some embodiments, the guide nucleic acid is as described herein.
[0505] In yet another aspect, the disclosure provides a lipid nanoparticle (LNP) comprising an RNA (e.g., mRNA) encoding the polypeptide of the disclosure or the fusion protein of the disclosure and a guide nucleic acid. In some embodiments, the guide nucleic acid is as described herein.
[0506] In yet another aspect, the disclosure provides a cell comprising the polypeptide of the disclosure, the fusion protein of the disclosure, , the system of the disclosure, the polynucleotide of the disclosure, or the vector of the disclosure.
[0507] Method of modifying
[0508] The system of the disclosure comprising the polypeptide of the disclosure has a wide variety of utilities, including modifying (e.g., cleaving, deleting, inserting, base editing, translocating, inactivating, or activating) a target DNA in a multiplicity of cell types. The system has a broad spectrum of applications requiring high activity / efficiency and small sizes, e.g., drug screening, disease diagnosis and prognosis, and treating various genetic disorders.
[0509] The method and / or the system of the disclosure can be used to modify a target DNA, for example, to modify the translation and / or transcription of one or more genes of the cells. For example, the modification may lead to increased transcription / translation / expression of a gene. In other embodiments, the modification may lead to decreased transcription / translation / expression of a gene.
[0510] In yet another aspect, the disclosure provides a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, or the lipid nanoparticle of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.
[0511] In some embodiments, the modification includes indel event, double-strand cleavage (double stranded break, DSB) , single-strand cleavage (nick) (e.g., on either target strand or nontarget stand of a dsDNA) , base editing (e.g., single base editing) , prime editing, and integration or insertion of exogenous donor (e.g., by homologous recombination) .
[0512] In some embodiments, the method is in vitro, in vivo, or ex vivo.
[0513] In some embodiments, the target DNA is in a cell.
[0514] In yet another aspect, the disclosure provides a cell comprising the system of the disclosure.
[0515] In yet another aspect, the disclosure provides a cell modified by the system of the disclosure or the method of the disclosure. In some embodiments, the cell is modified in vitro, in vivo, or ex vivo.
[0516] In some embodiments, the cell is a eukaryotic cell (e.g., an animal cell, a vertebrate cell, a mammalian cell, a non-human mammalian cell, a non-human primate cell, a rodent (e.g., mouse or rat) cell, a human cell, a plant cell, or a yeast cell) or a prokaryotic cell (e.g., a bacteria cell) .
[0517] In some embodiments, the cell is from a plant or an animal. In some embodiments, the cell is not from a plant.
[0518] In some embodiments, the cell is a non-human mammalian cell, such as a cell from a non-human primate (e.g., monkey) , an ox / cow / bull / cattle, sheep, goat, pig, horse, dog, cat, rodent (such as rabbit, mouse, rat, hamster, etc. ) , alpaca. In some embodiments, the cell is from fish (such as salmon, zebra fish) , bird (such as poultry bird, including chick, duck, goose) , reptile, shellfish (e.g., oyster, clam, lobster, shrimp) , insect, worm, yeast, etc. In some embodiments, the plant is a dicotyledon. In some embodiments, the dicotyledon is selected from the group consisting of soybean, cabbage (e.g., Chinese cabbage) , rapeseed, brassica, watermelon, melon, potato, tomato, tobacco, eggplant, pepper, cucumber, cotton, alfalfa, eggplant, grape. In some embodiments, the plant is a monocotyledon. In some embodiments, the monocotyledon is selected from the group consisting of rice, corn, wheat, barley, oat, sorghum, millet, grasses, Poaceae, Zizania, Avena, Coix, Hordeum, Oryza, Panicum (e.g., Panicum miliaceum) , Secale, Setaria (e.g., Setaria italica) , Sorghum, Triticum, Zea, Cymbopogon, Saccharum (e.g., Saccharum officinarum) , Phyllostachys, Dendrocalamus, Bambusa, Yushania.
[0519] In some embodiments, the cell is from a plant, such as monocot or dicot. In certain embodiment, the plant is a food crop such as barley, cassava, cotton, groundnuts or peanuts, maize, millet, oil palm fruit, potatoes, pulses, rapeseed or canola, rice, rye, sorghum, soybeans, sugar cane, sugar beets, sunflower, and wheat. In certain embodiment, the plant is a cereal (barley, maize, millet, rice, rye, sorghum, and wheat) . In certain embodiment, the plant is a tuber (cassava and potatoes) . In certain embodiment, the plant is a sugar crop (sugar beets and sugar cane) . In certain embodiment, the plant is an oil-bearing crop (soybeans, groundnuts or peanuts, rapeseed or canola, sunflower, and oil palm fruit) . In certain embodiment, the plant is a fiber crop (cotton) . In certain embodiment, the plant is a tree (such as a peach or a nectarine tree, an apple or pear tree, a nut tree such as almond or walnut or pistachio tree, or a citrus tree, e.g., orange, grapefruit or lemon tree) , a grass, a vegetable, a fruit, or an algae. In certain embodiment, the plant is a nightshade plant; a plant of the genus Brassica; a plant of the genus Lactuca; a plant of the genus Spinacia; a plant of the genus Capsicum; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.
[0520] In some embodiments, the cell is a stem cell. In some embodiments, the cell is an embryonic stem cell. In some embodiments, the cell is a primary human cell or an established human cell line.
[0521] In some embodiments, the cell is not a human or animal embryonic stem cell. In some embodiments, the cell is not a human or animal germ cell. In some embodiments, the cell is not a plant cell.
[0522] Pharmaceutical composition
[0523] In yet another aspect, the disclosure provides a pharmaceutical composition comprising (1) the system of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, or the cell of the disclosure; and (2) a pharmaceutically acceptable excipient.
[0524] In some embodiments, the pharmaceutical composition comprises the rAAV particle in a concentration selected from the group consisting of about 1×1010 vg / mL, 2×1010 vg / mL, 3×1010 vg / mL, 4×1010 vg / mL, 5×1010 vg / mL, 6×1010 vg / mL, 7×1010 vg / mL, 8×1010 vg / mL, 9×1010 vg / mL, 1×1011 vg / mL, 2×1011 vg / mL, 3×1011 vg / mL, 4×1011 vg / mL, 5×1011 vg / mL, 6×1011 vg / mL, 7×1011 vg / mL, 8×1011 vg / mL, 9×1011 vg / mL, 1×1012 vg / mL, 2×1012 vg / mL, 3×1012 vg / mL, 4×1012 vg / mL, 5×1012 vg / mL, 6×1012 vg / mL, 7×1012 vg / mL, 8×1012 vg / mL, 9×1012 vg / mL, 1×1013 vg / mL, or in a concentration of a numerical range between any of two preceding values, e.g., in a concentration of from about 9×1010 vg / mL to about 8×1011 vg / mL.
[0525] In some embodiments, the pharmaceutical composition is an injection.
[0526] In some embodiments, the volume of the injection is selected from the group consisting of about 1 microliter, 10 microliters, 50 microliters, 100 microliters, 150 microliters, 200 microliters, 250 microliters, 300 microliters, 350 microliters, 400 microliters, 450 microliters, 500 microliters, 550 microliters, 600 microliters, 650 microliters, 700 microliters, 750 microliters, 800 microliters, 850 microliters, 900 microliters, 950 microliters, 1000 microliters, and a volume of a numerical range between any of two preceding values, e.g., in a concentration of from about 10 microliters to about 750 microliters.
[0527] Method of diagnosing, preventing, or treating
[0528] In yet another aspect, the disclosure provides a method for diagnosing, preventing, or treating a disease in a subject in need thereof, comprising administering to the subject the system of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, the cell of the disclosure, or the pharmaceutical composition of the disclosure, wherein the disease is associated with a target DNA, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnose, prevents, or treats the disease.
[0529] In some embodiments, the disease is selected from the group consisting of Angelman syndrome (AS) , Alzheimer's disease (AD) , transthyretin amyloidosis (ATTR) , transthyretin amyloid cardiomyopathy (ATTR-CM) , cystic fibrosis (CF) , hereditary angioedema, diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD) , Becker muscular dystrophy (BMD) , spinal muscular atrophy (SMA) , alpha-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington’s disease (HTT) , fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis (ALS) , frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA) , sickle cell disease, thalassemia (e.g., β-thalassemia) , Parkinson's disease (PD) , myelodysplastic syndrome (MDS) , retinitis pigmentosa (RP) , age-related macular degeneration (AMD) , Hepatitis B, nonalcoholic fatty liver disease (NAFLD) , Acquired Immune Deficiency Syndrome, corneal dystrophy (CD) , hypercholesterolemia, familial hypercholesterolemia (FH) , heart disease (e.g., hypertrophic cardiomyopathy (HCM) ) , and cancer.
[0530] In some embodiments, the target DNA encodes a mRNA, a tRNA, a ribosomal RNA (rRNA) , a microRNA (miRNA) , a non-coding RNA, a long non-coding (lnc) RNA, a nuclear RNA, an interfering RNA (iRNA) , a small interfering RNA (siRNA) , a ribozyme, a riboswitch, a satellite RNA, a microswitch, a microzyme, or a viral RNA.
[0531] In some embodiments, the target DNA is a eukaryotic DNA.
[0532] In some embodiments, the eukaryotic DNA is a mammal DNA, such as a non-human mammalian DNA, a non-human primate DNA, a human DNA, a plant DNA, an insect DNA, a bird DNA, a reptile DNA, a rodent (e.g., mouse, rat) DNA, a fish DNA, a nematode DNA, or a yeast DNA.
[0533] In some embodiments, the target DNA is in a eukaryotic cell, for example, a human cell, a non-human primate cell, or a mouse cell.
[0534] In some embodiments, the administrating comprises local administration or systemic administration.
[0535] In some embodiments, the administrating comprises intrathecal administration, intramuscular administration, intravenous administration, transdermal administration, intranasal administration, oral administration, mucosal administration, intraperitoneal administration, intracranial administration, intracerebroventricular administration, or stereotaxic administration.
[0536] In some embodiments, the administration is injection or infusion.
[0537] In some embodiments, the subject is a human, a non-human primate, or a mouse.
[0538] In some embodiments, the level of the transcript (e.g., mRNA) of the target DNA is decreased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the transcript (e.g., mRNA) of the target DNA in the subject prior to the administration.
[0539] In some embodiments, the level of the transcript (e.g., mRNA) of the target DNA is increased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the transcript (e.g., mRNA) of the target DNA in the subject prior to the administration.
[0540] In some embodiments, the level of the expression product (e.g., protein) of the target DNA is decreased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the expression product (e.g., protein) of the target DNA in the subject prior to the administration.
[0541] In some embodiments, the level of the expression product (e.g., protein) of the target DNA is increased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the expression product (e.g., protein) of the target DNA in the subject prior to the administration. In some embodiments, the expression product is a functional mutant of the expression product of the target DNA.
[0542] In some embodiments, the median survival of the subject suffering from the disease but receiving the administration is 5 days, 10 days, 20 days, 30 days, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 1.5 year, 2 years, 2.5 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years or more longer than that of a subject or a population of subjects suffering from the disease and not receiving the administration.
[0543] The therapeutically effective dose may be either via a single dose, or multiple doses. One skilled in the art understands that the actual dose may vary greatly depending upon a variety of factors, such as the vector choices, the target cells, organisms, tissues, the general conditions of the subject to be treated, the degrees of transformation / modification sought, the administration routes, the administration modes, the types of transformation / modification sought, etc.
[0544] For example, the therapeutically effective dose of the rAAV particle may be about 1.0E+8, 2.0E+8, 3.0E+8, 4.0E+8, 6.0E+8, 8.0E+8, 1.0E+9, 2.0E+9, 3.0E+9, 4.0E+9, 6.0E+9, 8.0E+9, 1.0E+10, 2.0E+10, 3.0E+10, 4.0E+10, 6.0E+10, 8.0E+10, 1.0E+11, 2.0E+11, 3.0E+11, 4.0E+11, 6.0E+11, 8.0E+11, 1.0E+12, 2.0E+12, 3.0E+12, 4.0E+12, 6.0E+12, 8.0E+12, 1.0E+13, 2.0E+13, 3.0E+13, 4.0E+13, 6.0E+13, 8.0E+13, 1.0E+14, 2.0E+14, 3.0E+14, 4.0E+14, 6.0E+14, 8.0E+14, 1.0E+15, 2.0E+15, 3.0E+15, 4.0E+15, 6.0E+15, 8.0E+15, 1.0E+16, 2.0E+16, 3.0E+16, 4.0E+16, 6.0E+16, 8.0E+16, or 1.0E+17 vg, or within a range of any two of the those point values. vg stands for vector genomes of rAAV particles for administration.
[0545] Method of detecting
[0546] In yet another aspect, the disclosure provides a method of detecting a target DNA, comprising contacting the target DNA with the system of the disclosure, wherein the target DNA is modified by the complex, and wherein the modification detects the target DNA. In some embodiments, the modification generates a detectable signal, e.g., a fluorescent signal.
[0547] Kits
[0548] In yet another aspect, the disclosure provides a kit comprising the polypeptide of the disclosure, the system of the disclosure, the polynucleotide of the disclosure, the vector of the disclosure, the RNP of the disclosure, the LNP of the disclosure, the delivery system of the disclosure, the cell of the disclosure, or the pharmaceutical composition of the disclosure, or any one, two, or all components of the same.
[0549] In some embodiments, the kit further comprises an instruction to use the component (s) contained therein, and / or instructions for combining with additional component (s) that may be available or necessary elsewhere.
[0550] In some embodiments, the kit further comprises one or more buffers that may be used to dissolve any of the component (s) contained therein, and / or to provide suitable reaction conditions for one or more of the component (s) . Such buffers may include one or more of PBS, HEPES, Tris, MOPS, Na2CO3, NaHCO3, NaB, or combinations thereof. In some embodiments, the reaction condition includes a proper pH, such as a basic pH. In some embodiments, the pH is between 7-10.
[0551] In some embodiments, any one or more of the kit components may be stored in a suitable container or at a suitable temperature, e.g., 4 Celsius degree.
[0552] Further embodiments are illustrated in the following Examples which are given for illustrative purposes only and are not intended to limit the scope of the disclosure.
[0553] EXAMPLES
[0554] The following examples are provided to further illustrate some embodiments of the disclosure but are not intended to limit the scope of the invention. The following methods were used in the examples unless otherwise indicated. It will be understood by the exemplary nature of the examples that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.
[0555] Methods
[0556] Plasmid construction
[0557] Human codon-optimized sequences of programmable RNA-guided DNA endonucleases and their derived base editors were synthesized and cloned using pEASY-Basic Seamless Cloning and Assembly Kit (TransGen Biotech) under the control of a CBh promoter. For sgRNA expression, a phU6-BpiI-sgRNA scaffold vector was generated, in which the human U6 promoter drives expression of sgRNAs or ωRNAs. Spacer sequences were introduced by annealing complementary oligonucleotides and ligating them into BpiI-digested backbones. The MM762 protein mutant library was constructed as described previously.
[0558] Cloning of PAM library and E. coli PAM screen
[0559] A randomized PAM library containing a constant protospacer sequence followed by five randomized nucleotides was synthesized as single-stranded DNA (ssDNA) (HuaGene) . Double-stranded DNA (dsDNA) was generated by annealing with a short complementary oligonucleotide followed by second-strand synthesis using the Klenow fragment (New England Biolabs) . The resulting dsDNA fragments were assembled into pACYC184 vectors by pEASY-Basic Seamless Cloning and Assembly Kit (TransGen Biotech) and purified by isopropanol precipitation. Constructs were electroporated into DH5α Electroporation-Competent Cell (Weidibio) and plated on chloramphenicol-containing agar. Colonies were scraped after 15 h of growth at 37 ℃, and plasmid DNA was extracted using a NucleoBond Xtra Midi EF (Genetech) .
[0560] For bacterial PAM depletion assays, 200 ng of PAM library plasmids and 300 ng of plasmids encoding programmable RNA-guided DNA endonucleases with sgRNA expression cassettes were co-electroporated into DH5α Electroporation-Competent Cell (Weidibio) . Following 1 h recovery at 37 ℃ in antibiotic-free medium, cells were plated on 250 mm × 250 mm agar plates supplemented with carbenicillin and chloramphenicol. After 15 h of incubation at 37 ℃, colonies were collected and plasmid DNA was isolated (NucleoBond Xtra Midi EF, Genetech) . The PAM-containing regions were amplified using Phanta Max Super-Fidelity DNA polymerase (Vazyme) for 12 cycles, and Illumina sequencing adapters with sample-specific barcodes were added by a second PCR (18 cycles) . Amplicons were purified by gel extraction (Omega BioTek) and subjected to PE150 sequencing on an Illumina NovaSeq 6000 platform (AZENTA) .
[0561] Sequencing reads were processed to extract and enumerate PAM sequences, which were normalized to the total PAM counts per sample. Unique PAMs occurring only once were excluded. For each PAM, log fold change (logFC) values were calculated relative to non-targeting controls. PAMs with logFC < –3σ (s.d. ) were considered significantly depleted. A position weight matrix (PWM) was constructed from all significantly depleted PAMs using -logFC values as weights, and sequence logos were generated with WebLogo394.
[0562] Calculation of efficiency and specificity scores
[0563] For Fig. 6g, specificity scores were defined as 100 minus the mean normalized off-target editing of 20 mismatched sgRNAs (sm1–sm20) , normalized to wild-type SpCas9 sgRNA efficiency. Efficiency scores were calculated from GFP activation assays, normalized to wild-type SpCas9. For Fig. 9e, specificity scores were defined as 100 minus the mean indel frequency at off-target sites; sites with indel frequencies >1%were designated as bona fide off-targets. Efficiency scores were calculated as the mean editing efficiency across seven genomic loci. For Fig. 9h, specificity scores were defined as 100 minus the mean A·T-to-G·C conversion rate at 21 off-target sites, and efficiency scores were calculated as the mean A·T-to-G·C editing efficiency at the HEK293T site 4 locus.
[0564] Cell culture, transfection, and flow cytometry
[0565] HEK293T cells (Stem Cell Bank, Chinese Academy of Sciences) were maintained in Dulbecco’s modified Eagle’s medium (DMEM) supplemented with 10%fetal bovine serum (FBS) and penicillin–streptomycin at 37 ℃ in 5%CO2. Cells were seeded onto poly-D-lysine–coated 24-well plates (Corning) prior to transfection.
[0566] For GFP activation assays, cells were transfected using polyethylenimine (PEI, Polysciences) according to the manufacturer’s instructions with a total of 1.6 μg plasmids (0.8 μg reporter and 0.8 μg Cas-expressing constructs) and 3.2 μl PEI. Forty-eight hours post-transfection, GFP activation was quantified using a CytoFLEX flow cytometer (Beckman Coulter) , and data were analyzed with FlowJo v10.9.0.
[0567] For endogenous editing assays, HEK293T cells were transfected with 800 ng of Cas-expressing plasmids, 400 ng of sgRNA plasmids, and 2.4 μl PEI. At 48 h post-transfection, ~12,000 mCherry+GFP+ double-positive cells were isolated by fluorescence-activated cell sorting (FACS) for indel quantification, and cells collected at 72 h were used for base editing efficiency analysis.
[0568] Targeted deep sequencing and analysis
[0569] Sorted cells were lysed in 20 μl of lysis buffer containing proteinase K (Vazyme) according to the manufacturer’s instructions to extract genomic DNA. For EditR analysis, the genomic region surrounding each target site was amplified by nested PCR using Phanta Max Super-Fidelity DNA polymerase (Vazyme) . Purified amplicons were subjected to Sanger sequencing, and editing outcomes were quantified with EditR95. For deep sequencing, target loci were amplified by nested PCR, with barcoded primers introduced in the second round. Amplicons were purified using a gel extraction kit (Vazyme) and sequenced on an Illumina NovaSeq 6000 platform (AZENTA) with 150-bp paired-end reads. Raw reads were demultiplexed using Cutadapt, and editing outcomes (indels and base substitutions) were quantified with CRISPResso296. Target sites and primer sequences are listed in Figs. 31-36.
[0570] Purification of endonuclease-gRNA-target DNA R-loop complex
[0571] The wild type and engineered programmable RNA-guided DNA endonucleases were co-expressed with sgRNA from a single pRSFDuet-1 plasmid in E. coli BL21 (De3) cells. The resulting pellet from large scale expression was resuspended in 100 ml of cold lysis buffer (25 mM Hepes-NaOH, pH 7.5, 500 mM NaCl, 10%glycerol, and 1 mM 2-mercaptoethanol) supplemented with two tablets of cOmplete EDTA-free Protease Inhibitor Cocktail (Roche) before lysis by several passes through a microfluidizer (Microfluidics) at 18,000 psi. The lysate was clarified by centrifugation for 30 min at 70, 560g and 4 ℃ in an Optima XPN Ultracentrifuge (Beckman Coulter) using a Ti-45 rotor. The supernatant, which contained soluble Flag-tagged endonucleases in complex with sgRNA was incubated with 1 ml of M2 affinity gel (Millipore) pre-equilibrated with wash buffer (25 mM Hepes-NaOH, pH 7.5, 500 mM NaCl, 10%glycerol) in a gravity flow column. After incubation, the column was washed with 50 ml of wash buffer and further washed with buffer containing 25 mM Hepes-NaOH, pH 7.5, 200 mM NaCl, 10%glycerol. Following washing, the beads were eluted with 4 ml of elution buffer (25 mM Hepes-NaOH, pH 7.5, 200 mM NaCl, 10%glycerol, 120 μg ml-1 of 3× FLAG peptide and 1 mM 2-mercaptoethanol) .
[0572] Pooled elution fractions were concentrated to ~500 ul in 100K Amicon Ultra-15 concentrators (Millipore) and further purified by gel filtration chromatography on a 10 / 300 GL Superose 6 gel filtration column (Cytiva Life Sciences) in gel filtration buffer (25 mM Hepes-NaOH, pH 7.5, 150 mM NaCl and 1 mM dithiothreitol (DTT) ) . Peak fractions (as determined by the chromatograms with ultraviolet light of 280 nm) generated from the Unicorn software (v. 7.1) containing complete endonucleases and sgRNA complexes (as determined by sodium dodecylsulfate (SDS) –polyacrylamide gel electrophoresis (PAGE) analysis) , were pooled and concentrated to an absorbance at 280 nm of 3.0 to prepare cryo-EM grids. endonuclease-sgRNA RNP complex for use in biochemical experiments was purified in gel filtration buffer supplemented with 10%glycerol with peak fractions pooled, concentrated, flash-frozen and stored at -80 ℃.
[0573] For assembly of the ternary complex 200 μl of endonuclease-sgRNA RNP at 1.5 μM was incubated with 100 μl of double stranded target DNA at 5 μM. The sample was incubated at 37℃ for 30 mins before loading onto 10 / 300 GL Superose 6 gel filtration column. The peak fraction was collected and concentrated to A280 of 3.0 before use in cryoEM experiments.
[0574] Sample preparation of endonuclease-gRNA-target DNA R-loop complex
[0575] Cryo-EM grids were prepared by applying 3 μl of concentrated endonuclease-sgRNA-target DNA sample on to 400-mesh R1.2 / 1.3 UltrAuFoil grids (Quantifoil Micro Tools GmbH) , which had been rendered hydrophilic by glow discharging at 15 mA for 60 s with a PELCO easiGlow device (Ted Pella, Inc. ) . The sample was adsorbed for 30 s on the grids, followed by blotting and plunge freezing into liquid ethane using a Vitrobot Mark IV plunge freezer (Thermo Fisher Scientific) . Cryo-EM data were collected using the automated data acquisition software EPU (Thermo Fisher Scientific) on a Titan Krios G4 transmission electron microscope (Thermo Fisher Scientific) , operating at 300 kV and equipped with a cold field emission gun electron source and a Falcon4 direct detection camera. Images were recorded in counting mode at a nominal magnification at physical pixel size of 0.72 and Datasets were collected at a defocus range of 0.8–2.5 μm with a total electron dose of Image data were saved as electron event recordings.
[0576] In vitro target DNA cleavage assay
[0577] To investigate the target DNA cleavage activity of programmable RNA-guided DNA endonucleases, in vitro cleavage assays were performed using 708 bp of dsDNA containing the 20bp target site as substrate. Cleavage reactions were performed with 0.5 μM of target DNA and different concentration of endonuclease-sgRNA RNP in reaction buffer (25 mM Hepes-NaOH pH 7.5, 50 mM NaCl, 50 mM KCl, 1.5 mM MgCl2, 20%glycerol, and 1 mM DTT) . Reactions were incubated at 37 ℃ for 2 h and quenched by adding urea and proteinase K (Thermo Fisher Scientific) at final concentrations of 1 M and 1 μg μl-1, respectively, and incubated at 60 ℃ for 3 h. The sample was heated at 95 ℃ for 5 min before loading on a 6%Novex Native-PAGE (Thermo Fisher Scientific) . The gel was run at 200 V for 30 min followed by staining in GelRed Nucleic acid stain and imaged on an iBright FL1500 Imaging System. Gel images were processed and prepared on ImageJ (v. 1.53k) .
[0578] Cryo-EM image processing, model building, and refinement
[0579] The cryo-EM image processing was performed using cryoSPARC v4.7.1. The EM movie stacks were aligned and dose-weighted using patch-based motion correction (cryoSPARC implementation) . Contrast transfer function (CTF) estimation was also performed using the patch-based option. For the data of the MM762 (D10A) -sgRNA-target DNA ternary complex, a blob picker was used for initial particle selection. These particles were used for 2D classification to generate templates for template-based particles picking on the full dataset, resulting in 9,923,052 particles. Multiple rounds of 2D classifications, ab initio reconstruction, yielded multiple 3D classes from 2,550,392 particles. The best 3D class comprising 635,307 particles was used for 3D classification. The best 3D classes showing distinct densities for all regions of the complex, resulting in 154,740 particles, were selected (Fig. 17) . Further processing of these particles by non-uniform refinement resulted in a cryo-EM map at in C1 symmetry (Fig. 17) .
[0580] The 2D classes generated from the MM762 (D10A) -sgRNA-target DNA complex were used for template picking for the eMM762v2 (D10A) -sgRNA-target DNA complex followed by several rounds of 2D classification, resulting in 3,378,950 particles, which were subsequently used for ab-initio reconstruction. One 3D class containing 1,369,344 particles were used for non-uniform refinement, resulting in a final map of an overall resolution of in C1 symmetry (Fig. 17) .
[0581] Atomic models for both structures were built de novo into the cryo-EM map, in Coot 0.9.4. Real-space refinement for all built models was performed using Phenix, version 1.19.2-4158, using a general restraints setup.
[0582] Cryo-EM data collection, refinement and validation statistics for MM762-gRNA-target DNA R loop complexes.
[0583] GUIDE-seq
[0584] GUIDE-seq was performed with eMM762v2 and wild-type SpCas9 using five distinct sgRNAs, essentially as described previously74. Briefly, GUIDE-seq libraries were prepared following established protocols and sequenced on an Illumina NovaSeq 6000 platform (AZENTA) with 150-bp paired-end reads. Sequencing data were analyzed with the open-source GUIDE-seq software package (https: / / github. com / tsailabSJ / guideseq) , with genome-wide specificity profiling performed using an NGG PAM and a maximum of six mismatches. GUIDE-seq datasets have been deposited in the NCBI Sequence Read Archive.
[0585] Cas-dependent off-target analysis by targeted deep sequencing
[0586] Cas-dependent off-target activities of eMM762v2, SpCas9, HiFi Cas9, and eSpCas9 (1.1) were evaluated at the VEGFA, FANCF, and OFF-S1–OFF-S5 sites using previously reported off-target loci23, 25, 26, 33, 75. For analysis of eMM762ABEv3, SpCas9-ABE8e, and HiFi-ABE8e at the DMD exon 50 locus, potential off-target sites were predicted with Cas-OFFinder84 (http: / / www. rgenome. net / cas-offinder / ) using “NGG” PAM, up to three mismatches, and DNA / RNA bulge sizes of 1, with all other parameters set to default. Candidate sites were manually selected in order of increasing mismatch number. Off-target sites for the HEK293 site 4 and VEGFA loci were identified by Tracking-seq75 as described previously. All analyzed sites and primer sequences are provided in Figs. 31-36.
[0587] Orthogonal R-loop assay
[0588] An orthogonal R-loop assay was conducted to evaluate Cas-independent off-target editing as described previously. Briefly, HEK293T cells were co-transfected with 600 ng of ABE-expressing plasmids, 300 ng of corresponding sgRNA plasmids, and 600 ng of dSaCas9 plasmids with sgRNAs targeting five reported R-loop sites, using PEI. After 72 h, transfected cells were sorted by FACS, and genomic DNA was extracted using 20 μl of freshly prepared lysis buffer containing proteinase K (Vazyme) . Target loci and dSaCas9 R-loop sites were amplified and subjected to targeted deep sequencing. All target sequences and primers are listed in Figs. 31-36.
[0589] Genome editing in primary human T cells
[0590] Cryopreserved peripheral blood mononuclear cells (PBMCs) were thawed on day 0, and T cells were isolated using the EasySep Human T Cell Isolation Kit (STEMCELL Technologies) . On day 1, cells were activated with Dynabeads Human T-Activator CD3 / CD28 (Thermo Fisher) in CTS OpTmizer T medium (Gibco) supplemented with 10%human AB serum (Access Biologicals) , 100 IU ml-1 IL-2, and 1%GlutaMAX (Gibco) . On day 3, T cells were counted, washed with CTS OpTmizer T medium, and resuspended in P3 Primary Cell 4D-Nucleofector buffer (Lonza) at a density of 1 × 106 cells per 20 μl reaction. Cells (1 × 106) were electroporated with 1 μg of each sgRNA (GenScript, China) and 4 μg of mRNA encoding eMM762ABE or eMM762CBE using the Lonza 4D-Nucleofector system. mRNAs were synthesized with the EasyCap T7 Co-transcription Kit with CAG Trimer (Vazyme) incorporating N1-Me-pseudouridine modifications. On day 6, editing efficiencies were quantified by targeted deep sequencing. On day 8, edited cells were stained with antibodies against PD-1 (BioLegend, 379207) , CD52 (BioLegend, 316005) , B2M (BioLegend, 395727) , and CD3 (BioLegend, 300312) . For PD-1 detection, T cells were stimulated overnight with PMA (Sigma) and ionomycin (TopScience) prior to staining.
[0591] Study approval
[0592] All mice were maintained in a barrier facility under a 12-h light / dark cycle in accordance with the Guidelines for the Care and Use of Laboratory Animals of the Ministry of Science and Technology of China. Only male mice were used in the following examples. The number of independent biological replicates is provided in the figure legends. All procedures were approved by the Animal Care and Use Committee of the Zhongshan Institute for Drug Discovery, Shanghai Institute of Materia Medica, Chinese Academy of Sciences (Zhongshan, China) .
[0593] AAV9 production and delivery in DMDΔmE5051, KIhE50 / Y mice
[0594] Base editor constructs were packaged into AAV9 by PackGene Biotech (Guangzhou, China) . Briefly, HEK293 cells at 70–90%confluency were transfected with 20 μg pHelper, 10 μg pRepCap, and 10 μg gene-of-interest plasmid per 15-cm dish. Culture medium was replaced with pre-warmed growth medium immediately prior to transfection. Viral particles were harvested three days later and purified by iodixanol density gradient centrifugation. For intramuscular delivery, 3-week-old DMDΔmE5051, KIhE50 / Y mice were anesthetized and injected with 40 μl of AAV9 (1 × 1011 vg) or saline into the tibialis anterior (TA) muscle. Tissues were collected six weeks post-injection for genomic DNA extraction, RNA analysis, immunoblotting, and immunofluorescence.
[0595] RT-PCR
[0596] Total RNA was extracted from muscle tissues, and cDNA was synthesized using the HiScript II One Step RT–PCR Kit (Vazyme) according to the manufacturer’s instructions. PCR reactions (20 μl) contained ~2 μl of cDNA, 0.25 μM each primer, and 10 μl of Ex Taq (Takara) . Amplification was carried out on a C1000 Touch Thermal Cycler (Bio-Rad) with the following program: 95 ℃ for 5 min; 35 cycles of 95 ℃ for 30 s, 60 ℃ for 30 s, and 72 ℃ for 30 s. Products were analyzed by agarose gel electrophoresis.
[0597] Western blot analysis
[0598] Tissue samples were homogenized in RIPA buffer supplemented with protease inhibitors, and protein concentrations were determined with the Pierce BCA Protein Assay Kit (Thermo Fisher Scientific) . Equalized lysates were mixed with NuPAGE LDS sample buffer (Invitrogen) and 10%β-mercaptoethanol, denatured at 70 ℃ for 10 min, and resolved on 3–8%Tris-acetate gels (Invitrogen) at 200 V for 1 h. Proteins (10 μg per lane) were transferred to PVDF membranes by wet transfer (350 mA, 3.5 h) , blocked with 5%non-fat milk in TBST, and probed with primary antibodies against dystrophin (Sigma, D8168) or vinculin (Cell Signaling Technology, 13901S) . After incubation with HRP-conjugated secondary antibodies, signals were visualized using chemiluminescent substrates (Invitrogen) .
[0599] Immunofluorescence
[0600] Tissues were embedded in OCT compound, flash-frozen in liquid nitrogen, and cryosectioned at 10 μm. Sections were fixed at 37 ℃ for 2 h, permeabilized with PBS containing 0.4%Triton X-100 for 30 min and blocked with 10%goat serum for 1 h at room temperature. Samples were incubated overnight at 4 ℃ with primary antibodies against dystrophin (Abcam, ab15277) and spectrin (Millipore, MAB1622) , followed by incubation for 3 h at room temperature with species-specific secondary antibodies (Alexa Fluor 488 donkey anti-rabbit IgG, Jackson; Alexa Fluor 647 donkey anti-mouse IgG, Jackson) and DAPI. Slides were mounted with Fluoromount-G after final PBS washes. Images were acquired using Olympus FV3000 or Nikon C2 confocal microscopes. The proportion of dystrophin-positive fibers was quantified as a percentage of spectrin-positive fibers.
[0601] Statistical analysis
[0602] Data were presented as mean ± s.d., except for editing measurements with n = 2, which were shown as the mean. Statistical methods for individual experiments were indicated in the corresponding figure legends. For multiple comparisons, one-way ANOVA followed by Tukey’s post hoc test was applied. A P value < 0.05 was considered statistically significant. Experiments were not randomized, and investigators were not blinded to allocation or outcome assessment. Analyses were performed using GraphPad Prism v10.1.1.
[0603] Data availability
[0604] The cryo-EM structure of the MM762–sgRNA–DNA and eMM762v2–sgRNA–DNA complex was deposited in the Protein Data Bank (PDB; https: / / www. rcsb. org) . The human reference genome GRCh38 is available from the UCSC Genome Browser (https: / / genome. ucsc. edu / cgi-bin / hgGateway? db=hg38) . Plasmids are available on Addgene. Deep sequencing data generated in the following examples were deposited in the NCBI Sequence Read Archive under accession number.
[0605] Code availability
[0606] Sequencing analyses were performed using Cutadapt, CRISPResso2, and GUIDE-seq, while potential off-target sites were predicted in silico with Cas-OFFinder.
[0607] EXAMPLE 1. Identification of the programmable RNA-guided DNA endonucleases of the disclosure and characterization of their endonuclease activities
[0608] To identify more efficient Cas suitable for base editor development, the applicant developed a computational pipeline that surveyed over 2,000,000 bacterial genomes and 30,000 metagenomic assemblies. This search yielded a panel of Cas, including SN001-SN007 (SEQ ID NOs: 1, 3, 5, 7, 9, 11, and 13, respectively) , averaging ~750 amino acids in length, shared recognition of a 3′-NGG PAM (Fig. 15) , rendering them broadly compatible with canonical SpCas9 target sites while remaining amenable to single-AAV packaging.
[0609] Design and Construction:
[0610] Seven (7) Cas systems (SN001-SN007 systems) were developed by the applicant, each composed of one of SN001-SN007 (SEQ ID NOs: 1, 3, 5, 7, 9, 11, and 13, respectively) in Table 1 and a corresponding gRNA consisting of the corresponding scaffold sequence (SEQ ID NOs: 2, 4, 6, 8, 10, 12, and 14, respectively) in the same row of Table 1 and a guide sequence (spacer sequence) 5’ to scaffold sequence without a linker in-between (i.e., 5’ -guide sequence-scaffold sequence-3’ ) . The guide sequence was designed to target the single-target GFP activation reporter (Fig. 13a) for testing endonuclease activity or designed for E. coli PAM screen assay as in the applicant’s previous published work.
[0611] As an example, the SN007 (MM762) system for testing endonuclease activity was composed of (1) SN007 (SEQ ID NO: 13) encoded by a codon-optimized nucleic acid (SEQ ID NO: 20) , and (2) a guide RNA (gRNA) . The gRNA was composed of (1) the guide sequence as designed, and (2) the scaffold sequence of SEQ ID NO: 14, wherein the guide sequence is 5’ to the scaffold sequence without a linker between the guide sequence and the scaffold sequence.
[0612] As another example, the SN005 (MM745) system for testing endonuclease activity was composed of (1) SN005 (SEQ ID NO: 9) encoded by a codon-optimized nucleic acid (SEQ ID NO: 18) , and (2) a guide RNA (gRNA) . The gRNA was composed of (1) the guide sequence as designed, and (2) the scaffold sequence of SEQ ID NO: 10, wherein the guide sequence is 5’ to the scaffold sequence without a linker between the guide sequence and the scaffold sequence.
[0613] Results:
[0614] The experimental results as shown in Fig. 6h demonstrated the endonuclease activity of all the tested systems in single-target GFP activation reporter assay.
[0615] The experimental results as shown in Fig. 15 demonstrated the endonuclease activity of all the tested systems and their PAM recognition property in E. coli PAM screen assay. All the tested systems were demonstrated to have the same recognition of a 3′-NGG PAM as SpCas9 (N = A, T , C, or G) .
[0616] Among all the systems, SN007 (MM762) exhibited the highest endonuclease activity in single-target GFP activation reporter assay and was prioritized for further development (Fig. 6h) .
[0617] EXAMPLE 2. Optimal length of guide sequence for SN007 (MM762)
[0618] Following the evaluation procedure in Example 1, SN007 system was evaluated with various lengths of the guide sequence of the gRNA for optimal length of guide sequence.
[0619] Results:
[0620] In both reporter-based and endogenous site assays in HEK293T cells, the experimental results of Fig. 2 and Fig. 3 (and also Fig. 16d and Fig. 16e) show that SN007 system exhibited obvious endonuclease activity with a guide sequence in a length of 19-24 nt, and most suitable in a length of 20 nt.
[0621] EXAMPLE 3. Engineering of scaffold sequence of SN007 (MM762)
[0622] The sgRNA scaffold of MM762 spans 134 nucleotides-substantially longer than those of conventional Cas systems. To streamline its architecture and potentially improve performance, the applicant applied a truncation strategy targeting the crRNA-tracrRNA duplex. Specifically, the applicant generated a variant bearing a 14-nucleotide deletion in the duplex region, termed SL1-Δ14 (SEQ ID NO: 44) (Fig. 16a) . Unexpectedly, SL1-Δ14 (SEQ ID NO: 44) yielded higher editing activity than the full-length scaffold (SEQ ID NO: 14) in the presence of the same MM762 in a GFP activation assay (Fig. 16b) and was therefore adopted for subsequent MM762-based constructs unless otherwise indicated.
[0623] Sequence of original scaffold sequence of MM762 (134 nt; SEQ ID NO: 14) , in which the underlined nucleotides were deleted, resulting in SL1-Δ14 (SEQ ID NO: 44) :
[0624] Sequence of SL1-Δ14 scaffold of MM762 (120nt; SEQ ID NO: 44) :
[0625] EXAMPLE 4. SN007 (MM762) as intrinsically high‐fidelity Cas scaffolds for base editors
[0626] Cas nickase activity is essential for efficient base editing, as it enables cleavage of the non-edited strand and guides the cellular DNA repair machinery to install the desired base conversion1, 2. Type II CRISPR–Cas orthologues, as well as the Cas9 ancestor IscB, contain both HNH and RuvC-like nuclease domains42, which cleave the target strand (TS, complementary to the guide RNA) and the non-target strand (NTS) , respectively. These dual nuclease domain–containing systems can be converted into nickases by inactivating one of the nuclease domains, providing suitable scaffolds for the development of base editors.
[0627] To identify high-fidelity programmable RNA-guided DNA endonucleases suitable for compact base editors, the applicant selected the IscB orthologue OgeuIscB38, 42 (496 amino acids (aa) ) and representative compact Cas. The selected candidates included SlugCas945 (type II-A, 1, 052 aa) , Nme2Cas946 (type II-C, 1, 081 aa) , MG34-1644 (type II-D, 748 aa) , and the widely used SpCas947-49 as a benchmark (type II-A, 1, 367 aa) (Fig. 6a) . Editing specificity was assessed using a single-target GFP activation reporter-based mismatch tolerance assay (Fig. 13a, b) . OgeuIscB (Fig. 6b) and SpCas9 (Fig. 6c) exhibited high tolerance to single-nucleotide mismatches across the spacer region, indicating relatively low specificity. In contrast, SlugCas9 displayed minimal tolerance, with detectable activity only at the first position of the spacer (Fig. 6d) , while eNme2-C. NR demonstrated intermediate mismatch tolerance that was substantially lower than that of SpCas9 and OgeuIscB (Fig. 6e) . MG34-16 showed markedly reduced activity in the presence of any mismatch within the spacer (Fig. 6f) , indicative of high specificity. As expected, HiFi Cas925, a high-fidelity SpCas9 variant, exhibited greatly reduced mismatch tolerance compared to wild-type SpCas9 (Fig. 13c) . To enable quantitative comparison of specificity, efficiency, and molecular size across candidate nucleases, the applicant calculated an efficiency score and a specificity score for each, plotted alongside their respective protein lengths (Fig. 6g) . SlugCas9 and MG34-16 emerged as high-fidelity candidates, exhibiting greater specificity than HiFi Cas9, albeit with relatively low on-target editing efficiency. Notably, the smaller size of MM762 (762 amino acids) renders it more attractive for therapeutic applications than SlugCas9 (1,052 amino acids) , particularly for delivery-constrained contexts such as single-AAV packaging.
[0628] In the single-target GFP activation reporter-based mismatch tolerance assays, MM762 exhibited high specificity (Fig. 6i) . Unlike SpCas9, which tolerated certain double-nucleotide mismatches across the spacer, MM762 showed no detectable activity at doubly mismatched sites (Fig. 16c) , reinforcing its designation as a high-fidelity effector. Taken together, these findings nominate MM762 as a compact, high-fidelity Cas with high endonuclease activity, making it an ideal candidate for further development into base editors (Fig. 6g) .
[0629] EXAMPLE 5. Design of nickase based on the programmable RNA-guided DNA endonucleases of the disclosure and characterization of nickase activities
[0630] Designs:
[0631] By aligning the amino acid sequences of SN001-SN007 with CLUSTAL O (1.2.4) multiple sequence alignment (EMBL-EBI) , conservative motif (at positions corresponding to positions 6-18 of SN007 of SEQ ID NO: 13) was identified in the RuvC I domain (at positions corresponding to positions 1-38 of SN007 of SEQ ID NO: 13) of all SN001-SN007, conservative motif (at positions corresponding to positions 508-518 of SN007 of SEQ ID NO: 13) was identified in the RuvC III domain (at positions corresponding to positions 461-549 of SN007 of SEQ ID NO: 13) of all SN001-SN007, and conservative motifs DHI (at positions corresponding to positions 386-388 of SN007 of SEQ ID NO: 13) and (at positions corresponding to positions 409-410 of SN007 of SEQ ID NO: 13) were identified in the HNH domain (at positions corresponding to positions 353-448 of SN007 of SEQ ID NO: 13) of all SN001-SN007. By “conserved amino acid residue” it means that the amino acid residue is the same (constant) (not changed) at each indicated and corresponding position across all SN001-SN007. By “conservative motif” it means that the motif (consisting of multiple amino acid residues) is the same (constant) (not changed) across all SN001-SN007.
[0632] Mutation was then made at one or more of the conserved amino acid residues in one or more of the conservative motifs of any one of SN001-SN007 to generate engineered polypeptides (mutants) and evaluate the possible outcome.
[0633] To convert SN007 (MM762) into a nickase-compatible scaffold for base1, 2 or prime editing53, the applicant engineered mutants of SN007, in which key catalytic residues in the RuvC (e.g., D10, N323, H509, D512) or HNH (e.g., D386, H387, N410) domains were mutated to alanine for inactivation of the single-domain RuvC or HNH.
[0634] Results:
[0635] Using a dual-target GFP activation assay for testing nickase activity (Fig. 16f) , the applicant confirmed that the single-domain inactivation yielded multiple SN007 nickases suitable for DNA strand-specific editing applications (Fig. 16g) .
[0636] The flow cytometry results of SN007 mutants compared with the negative control ( “NT” ) and wild type SN007 positive control ( “WT” ) are shown in FIG. 4 and Fig. 16g. The mutants include SN007-D10A, SN007-N323A, SN007-D386A, SN007-H387A, SN007-N410A, SN007-H509A, and SN007-D512A.
[0637] The results demonstrate that by introducing a substitution (e.g., Alanine (A) ) at a conserved amino acid residue (e.g., D10) into the identified conservative motif (e.g., in the RuvC I nuclease domain) of SN007, the resulting mutant of SN007 (e.g., SN007-D10A) with such a substitution showed significantly reduced (almost eliminated) endonuclease activity (e.g., 1.5%) as compared with SN007 ( “WT” ) (43.9%) and showed significant nickase activity (e.g., 73.9%) , suggesting that such a substitution inactivating the RuvC nuclease domain and generating a nickase that is believed to nick the target strand and suitable for base editing.
[0638] The results demonstrate that by introducing a substitution (e.g., Alanine (A) ) at a conserved amino acid residue (e.g., D386) into the identified conservative motif (e.g., in the HNH nuclease domain) of SN007, the resulting mutant of SN007 (e.g., SN007-D386A) with such a substitution showed significantly reduced (almost eliminated) endonuclease activity (e.g., 1.2%) as compared with SN007 ( “WT” ) (43.9%) and showed significant nickase activity (e.g., 23.7%) , suggesting that such a substitution inactivating the HNH nuclease domain and generating a nickase that is believed to nick the nontarget strand and suitable for prime editing.
[0639] In contrast, the substitution N323A at N323 that is not conserved across SN001-SN007 failed to eliminate the endonuclease activity of SN007.
[0640] EXAMPLE 6. Cryo-EM structure of the MM762-sgRNA-target DNA ternary complex
[0641] To investigate the structural basis of target DNA recognition by MM762, the applicant purified MM762 and the catalytically inactivated MM762-D10A mutant, which abolishes NTS cleavage and favors R-loop stabilization, in complex with cognate full-length sgRNA (154 nt; consisting of 20 nt guide sequence and 134 nt scaffold sequence of SEQ ID NO: 14) (Fig. 17a) . In vitro cleavage assays confirmed that the MM762–sgRNA ribonucleoprotein (RNP) efficiently cleaved a 708 bp linear DNA substrate containing the target site (Fig. 17b) . For structural analysis, the MM762-D10A–sgRNA RNP was assembled with target DNA harboring a 20 bp protospacer and a 3′-TGG PAM, and its structure was resolved by cryo-electron microscopy (cryo-EM) at resolution (Fig. 7 and Fig. 17c) .
[0642] The cryo-EM structure revealed a bilobed architecture characteristic of the MM762-sgRNA system, comprising distinct recognition (REC) and nuclease (NUC) lobes (Fig. 7a and Fig. 7b) . The R-loop formed by the MM762-sgRNA RNP and target DNA was clearly resolved, as were most Cas domains, except for portions of the HNH and REC domains, which appeared disordered. The sgRNA scaffold organized the Cas domains through flexible linkers, supporting the domain mobility required for target binding and catalysis (Fig. 7b) .
[0643] MM762 engages both sgRNA and target DNA through extensive networks of specific and non-specific interactions. The bridge helix (BH) , enriched in basic residues, inserts between stem 1 of the sgRNA and the sgRNA–DNA duplex, forming electrostatic contacts with the RNA phosphate backbone and the target DNA (Fig. 7b) . The WED and PI domains clamp the 3′PAM region of the target DNA, underscoring their role in PAM recognition and target engagement (Fig. 7c) . Multiple contacts are also made with the deoxyribose–phosphate backbones of both TS and NTS. In the NTS, guanine bases at positions dG–2 and dG–3 are recognized in the major groove through interactions with Lys728, while Lys626 engages dG–2 and Asn664 interacts with dC–2 in the minor groove, providing base-specific recognition (Fig. 7c) .
[0644] The REC domain engages the sgRNA stem loops, stabilizing the RNA architecture and likely contributing to stringent mismatch discrimination (Fig. 7d) . Notably, an extensive interaction network between MM762 and the sgRNA–DNA heteroduplex formed a continuous interface along the R-loop (Fig. 7e) , exceeding the contact footprint observed in SpCas924, 55. This structural feature may underlie the stringent base-pairing requirement and low mismatch tolerance of MM762.
[0645] The MM762 sgRNA tested in this Example consists of a 20-nt guide sequence and a 134-nt scaffold (SEQ ID NO: 14) . The scaffold folds into a tightly organized tertiary architecture comprising two stems, two stem loops, and a pseudoknot (Fig. 7f and Fig. 7g) . Immediately downstream of the guide, the crRNA–tracrRNA duplex forms stem loop 1 (SL1) , which functions as a structural adaptor. This is followed by a central stem 1 (63A: 145A–80G: 130C) that extends into a 14-bp stem 2. Stem 2 diverges into stem loop 2 and a pseudoknot. Stem loop 2 forms a short 4-bp duplex (117G: 128C–120A: 125U) that extends outward from the scaffold core, whereas the pseudoknot stacks back-to-back against stems 1 and 2, reinforcing scaffold stability. The loop-proximal regions of stem loops 1 and 2 form only partial contacts with the protein, with much of these elements remaining solvent-exposed, suggesting they are not strictly required for MM762 activity-consistent with the observation in Example 3 above that a 14-nt truncation within SL1 enhanced editing efficiency (Fig. 16a and Fig. 16b) .
[0646] Together, these findings establish the structural basis for target recognition and intrinsic high fidelity of MM762-a property likely conserved across SN001-SN007-and indicate that its editing activity can be enhanced by engineering without compromising specificity.
[0647] EXAMPLE 7. Directed evolution of MM762 to enhance DNA editing activity
[0648] The applicant next constructed prototype base editors by fusing the MM762-D10A nickase to either TadA8e-V106W19 or CBE6c-V106W54 to generate MM762ABE and MM762CBE, respectively, which showed modest base editing efficiency.
[0649] To enhance the DNA editing activity of MM762, the applicant hypothesized that introducing positively charged residues such as lysine or arginine might increase its affinity for nucleic acids. To explore this, the applicant performed an unbiased arginine scanning mutagenesis across MM762, excluding naturally positively charged residues (lysine, arginine, and histidine) , using a previously described mutagenesis method56. This approach was favored over structure-guided design, as structural data may not fully capture the conformational plasticity required for activity enhancement.
[0650] The mutagenesis library was constructed on the MM762-D386A mutant, because both the wild-type MM762 and the MM762-D10A mutant exhibited high GFP activation efficiencies (>40%) , which could obscure the identification of mutants with enhanced editing activity. In contrast, MM762-D386A mutant displayed lower baseline editing activity (Fig. 4 and Fig. 16g) , allowing for improved resolution in detecting beneficial mutations. Each mutant was co-transfected with a GFP activation reporter into HEK293T cells, and editing activity was quantified by flow cytometry (Fig. 8a) .
[0651] Results:
[0652] Based on MM762-D386A in Example 3, further mutants were designed by introducing into MM762-D386A one additional substitution with R substantially across the full length of MM762-D386A. The fold change of the editing activity of the resulting SN007 mutants relative to MM762-D386A (set as 1 fold) are listed in the table below and shown in Fig. 5 and Fig. 8b.
[0653] In total, the applicant assessed 589 single amino acid substitutions. Among these, 120 mutants exhibited higher editing activity (> 1-fold) relative to SN007-D386A, and 83 mutants exhibited even higher editing activity (>1.2-fold) (Fig. 8b) , indicating that the single substitutions of those mutants have potential for further explore. The beneficial mutations were distributed across multiple domains, with notable enrichment in the BH, REC, WED, and PI regions, domains known to mediate critical contacts with both the sgRNA and the DNA duplex. Activity-enhancing variants in these regions exhibited 1.8 to 3.2-fold increases relative to MM762-D386A [i.e., wtMM762 (D386A) in Fig. 8b] .
[0654] EXAMPLE 8. Combinational mutations based on MM762-D386A to further enhance DNA editing activity
[0655] To further enhance editing activity, the applicant evaluated the combinatorial effects of top-performing single mutations. Representative substitutions were selected from six distinct functional domains, including BH, REC, HNH, RuvCIII, WED, and PI regions, with the aim of maximizing synergistic effects.
[0656] The single substitution with R at position 45, 46, 49, 138, 432, 549, 583, 640, or 747 of MM762-D386A was combined for further evaluation, including MM762 mutants v1-v8 in the table below. Editing activity was assessed using a dual-target GFP activation reporter (n=3) . The results in the table below show that v5 exhibited the highest editing activity compared to the other MM762 mutants and SpCas9-H840A. These mutants are suitable for primer editing.
[0657] The v5 mutant MM762-D386A, E49R, F432R, M583R, T747R (SEQ ID NO: 21) can be used to establish prime editors.
[0658] Based on the MM762 nickase (D386A, E49R, F432R, M583R, T747R) (SEQ ID NO: 21) , a prime editor (PE) (SEQ ID NO: 36) was designed and tested for base editing efficiency.
[0659] As shown, the PE contains a reverse transcriptase (SEQ ID NO: 37)
[0660] As shown, the PE contains bpSV40 NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 27) at its N-terminal and a GS linker containing bpSV40 NLS (SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS, SEQ ID NO: 39) between the SN007 nickase and the reverse transcriptase. The PE also contains a GS linker (SGGS) between the reverse transcriptase and the C-terminal bpSV40 NLS and a GS linker (GSG) between the C-terminal bpSV40 NLS and the C-terminal c-myc NLS (PAAKRVKLD, SEQ ID NO: 38) .
[0661] Based on the MM762 nickase (D386A, E49R, F432R, M583R, T747R) (SEQ ID NO: 21) , an epigenomic editor (SEQ ID NO: 40) was designed and tested for epigenomic editing efficiency.
[0662] As shown, the epigenomic editor contains KRAB domain (SEQ ID NO: 41) , Mus DNMT3l domain (SEQ ID NO: 42) , and Mus DNMT3a domain (SEQ ID NO: 43) .
[0663] SEQ ID NO: 41, KRAB, 62 aa
[0664] SEQ ID NO: 42, Mus DNMT3l, 214 aa
[0665] SEQ ID NO: 43, Mus DNMT3a, 302 aa
[0666] A corresponding mutant based on MM762-D10A instead of MM762-D386A was constructed as MM762-D10A, E49R, F432R, M583R, T747R (SEQ ID NO: 22) that can be used to establish base editors.
[0667] Based on the MM762 nickase (D10A, E49R, F432R, M583R, T747R) (SEQ ID NO: 22) , an adenine base editor (ABE) (SEQ ID NO: 25) was designed and tested for base editing efficiency.
[0668] As shown, the ABE contains TadA8e-V106W deaminase (SEQ ID NO: 26) .
[0669] As shown, the ABE contains a N-terminal bpSV40 NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 27) and a C-terminal NP NLS (KRPAATKKAGQAKKKK; SEQ ID NO: 28) , and a XTEN-GS linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGS; SEQ ID NO: 29) between the deaminase and the nickase. It is noted that the N-terminal NLS, the deaminase, and the XTEN-GS linker were inserted between the N-terminal starting Met (M) and the remaining sequence of the SN007 nickase.
[0670] A SpCas9-based adenine base editor (SEQ ID NO: 30) was used as control.
[0671] Based on the MM762 nickase (D10A, E49R, F432R, M583R, T747R) (SEQ ID NO: 22) , a cytosine base editor (CBE) (SEQ ID NO: 31) was designed and tested for base editing efficiency.
[0672] As shown, the CBE contains a cytosine deaminase (SEQ ID NO: 32) .
[0673] As shown, the CBE contains two (2) UGI domains (SEQ ID NO: 33) .
[0674] As shown, the CBE contains a N-terminal bpSV40 NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 27) and a C-terminal bpNLS (KRTADGSEFEPKKKRKV; SEQ ID NO: 34) , and a XTEN-GS linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGS; SEQ ID NO: 29) between the deaminase and the nickase. It is noted that the N-terminal NLS, the deaminase, and the XTEN-GS linker were inserted between the most N-terminal Met (M) and the remaining sequence of the MM762 nickase. The CBE also contains a GS linker (SGGSGGSGGS, SEQ ID NO: 35) between the nickase and the first UGI domain and between the two UGI domains, and a GS linker (SGGS) between the second UGI domain and the C-terminal bpNLS.
[0675] EXAMPLE 9. Combinational mutations based on MM762-D10A to further enhance DNA editing activity
[0676] On the other hand, the single substitutions with R at position G46, E49, D120, Q138, I200, F432, G560, F582, M583, E644, K710, or T747 were combined on the basis of MM762-D10A for further evaluation, resulting in twelve combinatorial variants eMM762v1–v12 (Fig. 8c) , which is compatible with base editor development. Editing activity was assessed using a dual-target GFP activation reporter (n=3) , revealing that several mutants exhibited markedly enhanced performance, achieving up to a 31.5-fold increase in GFP activation relative to the parental MM762-D10A (normalized to 1) (Fig. 8c) .
[0677] eMM762v2 (D10A+E49R+I200R+F432R+M583R+T747R) (SEQ ID NO: 23) :
[0678] EXAMPLE 10. Development of compact base editor eMM762ABEv1 based on eMM762v1
[0679] To generate a compact adenine base editor, the applicant fused TadA8e-V106W to eMM762v1, yielding eMM762ABEv1 (2.9 kb) (SEQ ID NO: 49) (Fig. 8d) .
[0680] For benchmarking, ABE8e variants were assembled from other compact Cas44-46 and IscB-derived nickases38 using the same design strategy, substituting eMM762v1 with each Cas or IscB scaffold (Fig. 18a) . A panel of TAM / PAM-matched target sites, including both previously reported and disease-relevant loci38, 44-46, 60, were selected for comparative analysis (Fig. 8e and Figs. 31-36) . In HEK293T cells, eMM762ABEv1 achieved a mean A·T-to-G·C editing efficiency of 48.2±16.3%, outperforming IminiIscB-ABE8e (16.4±17.8%) , MG34-16-ABE8e (8.5±9.2%) and SlugCas9-ABE8e (26.0±16.1%) and performing better than eNme-C-ABE8e (42.1±23.4%) (Fig. 8f) .
[0681] Consistent with previous observations36-38, 40, IscB-derived base editors exhibited a broad editing window spanning protospacer positions 1–13 (TAM at positions 17–22) (Fig. 18b) . For eMM762ABEv1, this window extended to positions 1–15 (PAM at positions 21–23) , with optimal activity between positions 2 and 15, and included PAM-proximal nucleotides that were refractory to editing by SpCas9ABE8e (Fig. 18c) . MG34-16-ABE8e, despite low overall activity, displayed a similar profile (Fig. 18d) . By contrast, SpCas9-ABE8e edited positions 1–13 (PAM at positions 21–23) with maximal activity between positions 2 and 12. The SlugCas9-NNG-ABE8e window typically spanned positions 3–15 (PAM at positions 22–24) (Fig. 18e) , whereas eNme2-C-ABE8e exhibited the broadest range, covering positions 1–17 (PAM at positions 24–29) and peaking between positions 4 and 15 (Fig. 18f) . Together, these results demonstrate that eMM762ABEv1 expands PAM-proximal editing beyond the reach of SpCas9ABE8e, thereby increasing the accessible sequence space for efficient adenine base editing.
[0682] EXAMPLE 11. Optimized engineering strategies to maximize eMM762ABE efficiency
[0683] To further enhance the activity of eMM762ABEv1 (SEQ ID NO: 49) , the applicant implemented a multi-tiered engineering strategy involving five key modifications.
[0684] Strategy i: Fusion of a DNA-binding domain.
[0685] The applicant fused the archaeal 7 kDa DNA-binding protein Sto7d from Sulfolobus tokodaii, incorporating a K12L mutation (Sto7d*) (SEQ ID NO: 54) to abolish ribonuclease activity65, 66, to base editor to improve base editing efficiency. Sto7d*was appended to the N terminus (nSto7d*) , C terminus (cSto7d*) , or positioned internally between TadA8e and eMM762v1 (inSto7d*) of eMM762ABEv1 (Fig. 9a) . All three configurations enhanced A·T-to-G·C conversion at endogenous sites in HEK293T cells, with nSto7d*conferring the highest base editing efficiency, and this variant was designated eMM762ABEv2 (SEQ ID NO: 50) (Fig. 9b and Fig. 19) .
[0686] Strategy ii: Computational design of a high-performance HNH domain (hpHNH) .
[0687] The applicant applied FuncLib68 to redesign the MM762 HNH domain (residues 385–400) . Among 12, 870 in silico-generated variants, the 29 lowest-energy designs were synthesized and evaluated using a base-editing GFP reporter (Fig. 9c, Fig. 20a, and Fig. 20b) . The high performance HNH domain variant hpHNHv22 demonstrated the highest activity (Fig. 9d and Fig. 20c) and was selected for further engineering. The mutations of the tested high performance HNH domain variants are shown in Fig. 20a. hpHNHv22 contains four substitutions R391P, A393T, I396L, and G400A.
[0688] Strategy iii: Selection of optimized eMM762 variants to enhance base editing efficiency.
[0689] Given that eMM762ABEv1 was based on eMM762v1, the applicant evaluated additional eMM762 variants (eMM762v2~eMM762v12) with higher nickase activity (Fig. 8c) . the applicant constructed corresponding ABE8e variants by fusing Sto7d*to their N-termini (Fig. 21a) and measured base editing using the BE–GFP activation reporter. Among the eMM762v1–v12 variants, eMM762v2-based ABE8e exhibited the highest base editing efficiency (Fig. 21b) and was designated eMM762ABEv3 (SEQ ID NO: 51) . Introduction of hpHNHv22 into the eMM762ABEv3 scaffold yielded a new variant, eMM762ABEv4 (SEQ ID NO: 52) (Fig. 9e) .
[0690] eMM762v2 (hpHNHv22) (SEQ ID NO: 24) :
[0691] To investigate the structural basis underlying the enhanced activity of eMM762v2 relative to wild-type MM762, the applicant determined the cryo-EM structure of the eMM762v2 variant carrying five substitutions (E49R, I200R, F432R, M583R, and T747R) in complex with sgRNA and target DNA. The structure was resolved at resolution (Fig. 17d) . Globally, eMM762v2 adopted a conformation highly similar to the WT complex, indicating that these substitutions do not introduce major structural rearrangements (Fig. 21c) .
[0692] Local differences, however, revealed clear functional implications. Substitutions introducing positively charged residues (E49R in the BH domain, I200R in the REC domain, and M583R in the WED domain) established new electrostatic interactions with the sgRNA, likely enhancing RNA–protein binding and complex stability. The F432R substitution, located within the HNH domain, may alter conformational dynamics of this domain and thereby influence nuclease activity. Finally, the T747R substitution in the PI domain formed a new contact with the target DNA, potentially stabilizing DNA recognition (Fig. 21c) .
[0693] Together, these structural insights suggest that the engineered mutations in eMM762v2 enhance MM762 activity primarily by reinforcing interactions with sgRNA and target DNA, thereby promoting efficient target engagement and cleavage.
[0694] Strategy iv: Linker optimization.
[0695] To improve the spatial configuration between deaminase and eMM762, the applicant tested linkers of varying lengths (5–19 amino acids) and rigidity (Fig. 21d and Fig. 21e) . The sequences of the linkers are shown in Fig. 21d. Compared with the commonly used XTEN flexible linker (32 aa) , linker1, 5, and 6 reduced efficiency, linker2, 3, and 4 were comparable, and linker7 and 8 conferred improved activity (Fig. 21f) . Replacing XTEN in eMM762ABEv4 with linker8 resulted in eMM762ABEv5 (SEQ ID NO: 53) (Fig. 9e) .
[0696] Overall comparison
[0697] To evaluate the impact of iterative engineering strategies, the applicant compared the base editing efficiencies of eMM762ABEv1 through eMM762ABEv5 at three endogenous loci (ATXN2, B2M and TRAC) . Base editing efficiencies increased progressively from eMM762ABEv1 (69.7±10.9%) to eMM762ABEv2 (71.1±12.4%) , eMM762ABEv3 (82.3±8.3%) , eMM762ABEv4 (84.5±7.6%) and eMM762ABEv5 (83.5±6.5%) (Fig. 9f and Fig. 22a) . eMM762ABEv3, v4, and v5 outperformed SpCas9-ABE8e (71.4±19.7%) and substantially exceeded the activity of HiFiABE8e (37.4±36.1%) , an ABE8e variant constructed on the high-fidelity SpCas9 scaffold (HiFi Cas925) . To further benchmark performance, the applicant extended the analysis to seven additional genomic loci (Fig. 22b) . eMM762ABEv3 and eMM762ABEv5 achieved mean editing rates of 80.4±8.6%and 75.1±15.4%, respectively, higher than SpCas9ABE8e (74.4±13.1%) , and markedly greater than eMM762ABEv1 (57.0±16.8%) and HiFiABE8e (44.5±26.4%) (Fig. 9g) . In addition, all eMM762ABE variants induced lower levels of unintended indels relative to SpCas9ABE8e and HiFiABE8e (Fig. 22c) . Collectively, these results demonstrate that eMM762ABEv3, eMM762ABEv4, and eMM762ABEv5 are highly efficient, compact adenine base editors with activities surpassing those of SpCas9-based ABE8e. Importantly, incorporation of high-fidelity SpCas9 mutations into the SpCas9ABE8e scaffold markedly reduced activity, consistent with previous observations that enhanced specificity often compromises on-target efficiency33.
[0698] eMM762ABEv1 = bpNLS-TadA8e-V106W-XTEN-eMM762v1-nlpNLS (SEQ ID NO: 49)
[0699] eMM762ABEv2 = v1 + nSto7d* = bpNLS-Sto7d*-TadA8e-V106W-XTEN-eMM762v1-nlpNLS (SEQ ID NO: 50)
[0700] eMM762ABEv3 = v2 + eMM762v2 = bpNLS-Sto7d*-TadA8e-V106W-XTEN-eMM762v2-nlpNLS (SEQ ID NO: 51)
[0701] eMM762ABEv4 = v3 + hpHNHv22 = bpNLS-Sto7d*-TadA8e-V106W-XTEN-eMM762v2 (hpHNHv22) -nlpNLS (SEQ ID NO: 52)
[0702] eMM762ABEv5 = v4 + linker8 = bpNLS-Sto7d*-TadA8e-V106W-linker8-eMM762v2 (hpHNHv22) -nlpNLS (SEQ ID NO: 53)
[0703] Sto7d* (SEQ ID NO: 54)
[0704] nlpNLS (SEQ ID NO: 55)
[0705] Strategy v: Structure-guided sgRNA scaffold optimization.
[0706] Cryo-EM structural analysis (Fig. 7) , together with preliminary optimization showing that the truncated MM762 scaffold SL1-Δ14 (SEQ ID NO: 44) improves MM762 editing efficiency (Fig. 16a and Fig. 16b) , established SL1-Δ14 as a streamlined yet functional scaffold for further engineering. The sequences of the engineered scaffold sequences are listed in Fig. 23a, which were tested in combination with eMM762ABEv3.
[0707] Progressive truncation of SL1 and SL2 revealed that most modifications reduced MM762 activity, with the exception of a 4 nucleotides deletion in SL2 (SL1-Δ14+SL2-Δ4) , which retained activity comparable to SL1-Δ14 alone (Fig. 23a, Fig. 23b, and Fig. 23c) . To further enhance sgRNA stability, the applicant introduced G–C base pairs in place of A–U or mismatched nucleotides-astrategy previously shown to improve Cas nuclease activity56, 73. Among these variants, v1 (U98C) , v2 (A43C+U99G) , and v3 (U9C+A20G+U98C) exhibited enhanced editing (Fig. 23a and Fig. 23d) . When co-transfected with eMM762ABEv3, variants v2, v3, and a combinatorial design v4 (A43C+U98C+U99G) achieved editing efficiencies of 81.0±11.6%, 81.2±10.9%, and 82.3±9.1%, respectively-exceeding the 77.6±10.7%observed with the sgRNA SL1-Δ14 (Fig. 9h, Fig. 9i, and Fig. 23e) . Together, these results demonstrate that structure-guided sgRNA engineering-via scaffold truncation and G–C base pair substitution-can reduce sgRNA length by over 13%while further enhancing the activity of eMM762ABE variants.
[0708] Scaffold variant v1 (U98C) (SEQ ID NO: 45)
[0709] Scaffold variant v2 (A43C+U99G) (SEQ ID NO: 46)
[0710] Scaffold variant v3 (U9C+A20G+U98C) (SEQ ID NO: 47)
[0711] Scaffold variant v4 (A43C+U98C+U99G) (SEQ ID NO: 48)
[0712] EXAMPLE 12. Development of eMM762CBE
[0713] Building on the optimization strategies established for eMM762ABE, the applicant generated compact cytosine base editors (eMM762CBEs) by replacing TadA8e-V106W with CBE6c-V106W54 (SEQ ID NO: 68) and tethering two uracil glycosylase inhibitors (2× UGI) . This yielded four eMM762CBE variants (v1–v4) (SEQ ID NOs: 64-67) , each with a compact coding size of =<3.8 kb (Fig. 9j) . All variants exhibited robust cytosine editing, with efficiencies comparable to or exceeding those of SpCas9CBE6c at the TRAC and CIITA loci in human cells (Fig. 9k) . Across an expanded panel of eight additional genomic loci (Fig. 24) , eMM762CBEv1 and eMM762CBEv4 achieved average C·G-to-T·Aediting efficiencies of 70.5±7.1%and 68.8±9.2%, respectively, surpassing the 63.5 ± 16.9%observed for SpCas9CBE6c (Fig. 9l) . Collectively, these results establish eMM762CBEs as highly efficient, single-AAV–compatible cytosine base editors. The scaffold sequence SL1-Δ14 (SEQ ID NO: 44) was used in combination with eMM762CBEv1-v4.
[0714] eMM762CBEv1 = bpNLS-Sto7d*-CBE6c-V106W-XTEN-eMM762v2-UGI-UGI-bpNLS (SEQ ID NO: 64)
[0715] eMM762CBEv2 = bpNLS-CBE6c-V106W-XTEN-eMM762v2 (hpHNHv22) -UGI-UGI-bpNLS (SEQ ID NO: 65)
[0716] eMM762CBEv3 = bpNLS-Sto7d*-CBE6c-V106W-XTEN-eMM762v2 (hpHNHv22) -UGI-UGI-bpNLS (SEQ ID NO: 66)
[0717] eMM762CBEv4 = bpNLS-Sto7d*-CBE6c-V106W-linker8-eMM762v2 (hpHNHv22) -UGI-UGI-bpNLS (SEQ ID NO: 67)
[0718] CBE6c-V106W (SEQ ID NO: 68)
[0719] EXAMPLE 13. eMM762 and its derived base editors exhibit high fidelity
[0720] The applicant systematically evaluated the genome-wide specificity of eMM762v2 using GUIDE-seq74 across five endogenous sites and observed markedly fewer off-target sites compared to SpCas9 (Fig. 10a and Fig. 25a) . For example, at the FANCF-1, FANCF-2, and VEGFA loci, eMM762v2 resulted in much lower off target activity than SpCas9, as reflected by both the number of detected off-target sites as well as the total modification frequency at each detected off-target site (Fig. 10b) . The editing specificity, quantified by the log2 ratio of (on-target + 1) to (off-target + 1) reads, was consistently higher for eMM762v2 than for SpCas9 (mean log2 ratio: 8.9 versus 5.4) , further underscoring its superior fidelity (Fig. 10c) .
[0721] Targeted sequencing at FANCF and VEGFA loci further confirmed that eMM762v2 maintained robust on-target activity (53.2±1.9%–60.9±3.0%) , comparable to or higher than the high-fidelity variants HiFi Cas925 (53.5±3.0%–62.9±5.0%) and eSpCas9 (1.1) 23 (3.0±1.8%–57.1±5.0%) . Importantly, eMM762v2 exhibited the highest specificity, with up to 29.5-fold lower off-target editing than wild-type SpCas9 (Fig. 10d) . Across five additional genomic loci, eMM762v2 maintained high on-target activity close to wild-type SpCas9 and outperformed HiFi Cas9 and eSpCas9 (1.1) in specificity, illustrating that eMM762v2 achieves an improved balance between efficiency and fidelity compared to existing Cas9 variants (Fig. 10e and Fig. 25b) .
[0722] To evaluate base editors derived from eMM762v2, the applicant compared eMM762ABEv3 with SpCas9ABE8e and HiFiABE8e at two representative genomic loci. At the HEK293 site4 locus, eMM762ABEv3 achieved 87.0±2.9%A·T-to-G·C editing-comparable to SpCas9ABE8e (88.3±1.3%) and exceeding HiFiABE8e (78.5±5.5%) . Across 21 previously reported off-target sites75, eMM762ABEv3 showed reduced editing activity at 19 sites relative to SpCas9ABE8e and at 17 sites compared to HiFiABE8e (Fig. 10f) . On average, eMM762ABEv3 reduced off-target editing by 27.6-fold compared to SpCas9ABE8e, with reductions of up to 117.4-fold. When benchmarked against HiFiABE8e, eMM762ABEv3 exhibited a 3.7-fold average decrease in off-target activity, reaching up to 19.0-fold at individual sites (Fig. 10g) . Similar trends were observed at the VEGFA locus, where eMM762ABEv3 showed 4.2-fold and 2.9-fold reductions in off-target editing compared to SpCas9ABE8e and HiFiABE8e, respectively (Fig. 25c) . Taken together, these results demonstrate that eMM762ABEv3 delivers SpCas9-like on-target performance while achieving specificity equal to or greater than HiFiABE8e (Fig. 10h) .
[0723] Finally, to determine whether Sto7d*fusion impacts Cas-independent off-target editing, the applicant performed an orthogonal R-loop assay16 using catalytically inactive dSaCas9 to generate non-cognate ssDNA regions (Fig. 25d) . Although eMM762ABEv2 (Sto7d*fusion) showed an increase in on-target efficiency (77.7±0.7%) relative to eMM762ABEv1 (72.5±1.3%) and SpCas9ABE8e (75.7±1.6%) , it did not exhibit obviously increased off-target editing across five R-loop sites (0.5±0.2%–4.2±0.2%) , comparable to eMM762ABEv1 (0%–2.3±0.1%) and lower than SpCas9ABE8e (1.3±0.3%–6.3±0.1%) (Fig. 25e) .
[0724] Together, these results demonstrate that activity-enhanced eMM762 variants preserve the intrinsic high fidelity of wild-type MM762, leading to base editors with robust on-target activity and dramatically reduced Cas-dependent off-target editing. This establishes eMM762ABE as a compact and highly efficient base editor with superior specificity.
[0725] EXAMPLE 14. eMM762ABE and eMM762CBE enable efficient multiplex editing in T cells with minimal off-target effects
[0726] Base editors can introduce precise point mutations to disrupt gene function by generating premature stop codons or abrogating splice sites, thereby inducing exon skipping or nonsense-mediated decay76-78. Unlike nucleases, base editors operate without creating DNA double-strand breaks, making them particularly well suited for multiplex genome engineering in allogeneic CAR-T cells, where safety and predictability are paramount. SpCas9-based ABEs and CBEs have already progressed from laboratory development to clinical translation in CAR-T cell therapies79-81.
[0727] To evaluate the multiplex editing capabilities of eMM762ABE and eMM762CBE, the applicant co-delivered mRNA encoding each editor alongside three synthetic sgRNAs into primary human T cells via a single electroporation (Fig. 11a) . eMM762ABEv3 efficiently induced A·T-to-G·C editing at B2M (88.3%) , CD52 (97.3%) , and PD1 (95.8%) , comparable to SpCas9ABE8e (96.2%, 98.4%, and 96.9%, respectively) (Fig. 11b) . Notably, eMM762ABEv3 generated unintended indels at only 0.4–0.6%, near background levels (0.2–0.4%) , whereas SpCas9ABE8e induced higher indel frequencies of 1.2–2.3% (Fig. 11b) , indicating enhanced safety of eMM762ABEv3 for allogeneic CAR-T cell engineering. Similarly, eMM762CBEv1 achieved robust C·G-to-T·Aediting at B2M (88.1%) , TRAC (93.9%) , and PD1 (86.7%) , comparable to SpCas9CBE6c (93.9%, 86.7%, and 89.8%, respectively) , with indel rates of 0.5–1.4%for eMM762CBEv1 versus 0.2–1.9%for SpCas9CBE6c (Fig. 11c) . Flow cytometry confirmed efficient protein knockout, with eMM762ABEv3 and eMM762CBEv1 achieving 91.6–97.0%and 90.5–96.9%knockout efficiencies, respectively-comparable to SpCas9-based editors (95.7–99.0%for SpCas9ABE8e and 90.2–94.4%for SpCas9CBE6c) (Fig. 11d and Fig. 11e) .
[0728] Because off-target activity is a critical safety concern for clinical translation6, the applicant next examined off-target editing at sites previously reported for B2M, TRAC and CD5282. Using targeted amplicon sequencing, eMM762ABEv3 achieved high on-target efficiency at B2M while showing no detectable activity at four tested off-targets, in contrast to SpCas9ABE8e, which edited B2M-OT1 (10.0%) and OT2 (40.4%) (Fig. 11f) . At the CD52 locus, eMM762ABEv3 also demonstrated higher fidelity than SpCas9ABE8e (Fig. 26a) . Likewise, eMM762CBEv1 exhibited no detectable off-target activity at B2M or TRAC, whereas SpCas9CBE6c edited B2M-OT2 at 1% (Fig. 11g and Fig. 26b) .
[0729] Together, these findings demonstrate that eMM762ABE and eMM762CBE enable efficient multiplex genome editing in primary human T cells, achieving on-target efficiencies comparable to SpCas9-derived base editors while exhibiting substantially reduced off-target activity and unintended indels, underscoring their potential for safer allogeneic CAR-T engineering.
[0730] EXAMPLE 15. Single-AAV delivery of eMM762ABE efficiently restores dystrophin expression in vivo in humanized Duchenne muscular dystrophy mouse model
[0731] The applicant next evaluated eMM762ABE for therapeutic base editing in vivo using the applicant’s previously established humanized DMDΔmE5051, KIhE50 / Y mouse model83, in which exons 50 and 51 of the mouse Dmd gene are replaced with human DMD exon 50 (Fig. 27b) . In the applicant’s prior work, SpCas9ABE8e delivered via dual AAV vectors efficiently disrupted the exon 50 splice donor site to restore dystrophin expression. Given that MM762 shares the same 3′-NGG PAM preference as SpCas9, the applicant used the identical DMD E50 target sequence to compare in vivo editing by eMM762ABE.
[0732] In HEK293T cells, eMM762ABEv3 and SpCas9ABE8e displayed comparable on-target editing efficiencies and both outperformed HiFiABE8e (Fig. 12a) . At 19 Cas-OFFinder84 predicted off-target sites, SpCas9ABE8e and HiFiABE8e induced detectable editing at Off-10 (8.6±1.5%and 0.9±0.4%, respectively) , whereas eMM762ABEv3 showed no measurable off-target activity (0.3±0.2%) (Fig. 12a) . Consistently, eMM762ABEv3 demonstrated superior editing fidelity in N2A cells compared with SpCas9ABE8e and HiFiABE8e (Fig. 27a) .
[0733] The compact size of eMM762ABEs enabled packaging of both the editor and sgRNA within a single AAV vector (4.6 kb) , whereas SpCas9ABE8e required dual-AAV vector delivery (Fig. 12b) . Following intramuscular injection of tibialis anterior (TA) muscle and six weeks of expression (Fig. 12c) , eMM762ABEv3 and SpCas9ABE8e induced similar A·T-to-G·C conversion at the DMD E50 site (8.8±1.9%and 9.3±2.7%, respectively) (Fig. 12d) . RT–PCR revealed efficient exon 50 skipping, with eMM762ABEv3 achieving 61.2±3.8%compared with 54.1±7.5%for SpCas9ABE8e (Fig. 12e and Fig. 27c) . Histological staining showed dystrophin restoration in 91.2±2.3%of fibers with eMM762ABEv3 and 82.8±5.1%with SpCas9ABE8e (Fig. 12f and Fig. 12g) , and western blotting confirmed robust dystrophin re-expression (18.9±6.5%and 19.8±3.9%of wild-type levels, respectively) (Fig. 12h and Fig. 12i) . In vivo evaluation of eMM762ABEv1-v3 also demonstrated stepwise improvements in efficiency (Fig. 27d and Fig. 27h) .
[0734] Taken together, these results demonstrate that eMM762ABEv3 enables efficient, high-fidelity, and therapeutically relevant base editing in vivo through single-AAV delivery.
[0735] EXAMPLE 16. Design and screen of high-efficiency ABE based on SN005 (MM745)
[0736] SN005-D10A nickase (SEQ ID NO: 69; 744 aa) was designed based on SN005 (SEQ ID NO: 9; 745 aa) without N-terminal Met.
[0737] A prototype SN005-based ABE was designed (SEQ ID NO: 70) , consisting of, from N-terminal to C-terminal, Met, N-terminal bpSV40 NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 27) , TadA8e-V106W deaminase (SEQ ID NO: 26) , XTEN-GS linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGS; SEQ ID NO: 29) , Ser, SN005-D10A nickase (SEQ ID NO: 69) , and C-terminal NP NLS (KRPAATKKAGQAKKKK; SEQ ID NO: 28) .
[0738] Further ABEs were designed by introducing one or more substitutions with R (e.g., E44R, D128R, S546R, Y568R, M569R, S623R, and E667R) into SN005-D10A nickase (SEQ ID NO: 69) contained in the prototype SN005-based ABE (SEQ ID NO: 70) . All the ABEs were tested for base editing efficiency, with the efficiency of the prototype SN005-based ABE (SEQ ID NO: 70) setting as 1-fold for comparison.
[0739] The results of Fig. 28 and Fig. 29 show that multiple single and combinational substitutions with Arg (R) resulted in improved base editing efficiency than the prototype SN005-based ABE (SEQ ID NO: 70; marked as “WT745” in Fig. 28 and Fig. 29) . Since the other elements are the same, it is believed that the improved base editing efficiency is due to the improved nickase activity of those SN005 mutants contained in the ABEs. Among those mutants, the mutant SN005-D10A+E44R+D128R+S546R+Y568R+S623R+E667R (SEQ ID NO: 71) is termed as enSN005 nickase for further explore.
[0740] To further improve the base editing efficiency, a DNA binding domain (SEQ ID NO: 72) was introduced to generate additional ABEs containing the DNA binding domain (DBD) at its N-terminal.
[0741] A representative DBD-containing ABE is set forth in SEQ ID NO: 73, consisting of, from N-terminal to C-terminal, Met, N-terminal bpSV40 NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 27) , DNA binding domain (SEQ ID NO: 72) , short linker (SGGS) , TadA8e-V106W deaminase (SEQ ID NO: 26) , XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGS; SEQ ID NO: 29) , enSN005 nickase (SEQ ID NO: 71) , short linker (SGGS) , and C-terminal NP NLS (KRTADGSEFEPKKKRKV; SEQ ID NO: 34) .
[0742] The results of Fig. 29 show that the ABE (SEQ ID NO: 73) containing both DBD (SEQ ID NO: 72) and enSN005 nickase (SEQ ID NO: 71) achieved the highest base editing efficiency, which was termed “enABE” . PDCD1, B2M, and CD52 were selected to test the base editing efficiency of enABE at endogenous sites of cellular genome. The guide sequences for PDCD1, B2M, and CD52 are set forth in SEQ ID NOs: 74-76 respectively. The gRNAs for PDCD1, B2M, and CD52 are set forth in SEQ ID NOs: 77-79, respectively, with the scaffold sequence (SEQ ID NO: 10) of SN005.
[0743] The results of Fig. 30 show that enABE achieved high base editing efficiency for all the three endogenous sites of cellular genome.
[0744] EXAMPLE 17. Design and identification of guide sequences for SN001-SN007
[0745] The guide sequences (SEQ ID NOs: 80-260) in the table below were designed on each indicated gene for use in combination with the scaffold sequences for each of SN001-SN007. The guide sequences can be combined with the scaffold sequence of any one of SN001-SN007. For example, the TRAC-1 guide sequence TTCGTATCTGTAAAACCAAG (SEQ ID NO: 212) can be combined with the scaffold sequence (SEQ ID NO: 10) of SN005 (MM745) to form a full length gRNA (SEQ ID NO: 261) .
[0746] Alternatively, the TRAC-1 guide sequence TTCGTATCTGTAAAACCAAG (SEQ ID NO: 212) can be combined with the scaffold sequence (SEQ ID NO: 44) of SN007 (MM762) to form a full length gRNA (SEQ ID NO: 262) .
[0747] References
[0748] 1. Komor, A.C., Kim, Y.B., Packer, M.S., Zuris, J.A. &Liu, D.R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016) .
[0749] 2. Gaudelli, N.M. et al. Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551, 464-471 (2017) .
[0750] 3. Song, Y. et al. Large-Fragment Deletions Induced by Cas9 Cleavage while Not in the BEs System. Mol Ther Nucleic Acids 21, 523-526 (2020) .
[0751] 4. Leibowitz, M.L. et al. Chromothripsis as an on-target consequence of CRISPR-Cas9 genome editing. Nat Genet 53, 895-905 (2021) .
[0752] 5. Kosicki, M., Tomberg, K. &Bradley, A. Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. Nat Biotechnol 36, 765-771 (2018) .
[0753] 6. Tsuchida, C.A. et al. Mitigation of chromosome loss in clinical CRISPR-Cas9-engineered T cells. Cell 186, 4567-4582 e4520 (2023) .
[0754] 7. Liu, M. et al. Global detection of DNA repair outcomes induced by CRISPR-Cas9. Nucleic Acids Res 49, 8732-8742 (2021) .
[0755] 8. Anzalone, A.V., Koblan, L.W. &Liu, D. R. Genome editing with CRISPR-Cas nucleases, base editors, transposases and prime editors. Nat Biotechnol 38, 824-844 (2020) .
[0756] 9. Porto, E.M., Komor, A.C., Slaymaker, I.M. &Yeo, G.W. Base editing: advances and therapeutic opportunities. Nat Rev Drug Discov 19, 839-859 (2020) .
[0757] 10. Musunuru, K. et al. Patient-Specific In Vivo Gene Editing to Treat a Rare Genetic Disease. N Engl J Med 392, 2235-2243 (2025) .
[0758] 11. Zhang, H., Li, T., Sun, Y. &Yang, H. Perfecting Targeting in CRISPR. Annu Rev Genet 55, 453-477 (2021) .
[0759] 12. Zuo, E. et al. Cytosine base editor generates substantial off-target single-nucleotide variants in mouse embryos. Science 364, 289-292 (2019) .
[0760] 13. Grunewald, J. et al. Transcriptome-wide off-target RNA editing induced by CRISPR-guided DNA base editors. Nature 569, 433-437 (2019) .
[0761] 14. Zhou, C. et al. Off-target RNA mutation induced by DNA base editing and its elimination by mutagenesis. Nature 571, 275-278 (2019) .
[0762] 15. Jin, S. et al. Cytosine, but not adenine, base editors induce genome-wide off-target mutations in rice. Science 364, 292-295 (2019) .
[0763] 16. Doman, J.L., Raguram, A., Newby, G.A. &Liu, D.R. Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nat Biotechnol 38, 620-628 (2020) .
[0764] 17. Liu, Y. et al. Elimination of Cas9-dependent off-targeting of adenine base editor by using TALE to separately guide deaminase to target sites. Cell Discov 8, 28 (2022) .
[0765] 18. Rees, H.A., Wilson, C., Doman, J.L. &Liu, D.R. Analysis and minimization of cellular RNA editing by DNA adenine base editors. Sci Adv 5, eaax5717 (2019) .
[0766] 19. Richter, M.F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat Biotechnol 38, 883-891 (2020) .
[0767] 20. Li, J. et al. Structure-guided engineering of adenine base editor with minimized RNA off-targeting activity. Nat Commun 12, 2287 (2021) .
[0768] 21. Grunewald, J. et al. CRISPR DNA base editors with reduced RNA off-target and self-editing activities. Nat Biotechnol 37, 1041-1048 (2019) .
[0769] 22. Tsai, S.Q. &Joung, J.K. Defining and improving the genome-wide specificities of CRISPR-Cas9 nucleases. Nat Rev Genet 17, 300-312 (2016) .
[0770] 23. Slaymaker, I.M. et al. Rationally engineered Cas9 nucleases with improved specificity. Science 351, 84-88 (2016) .
[0771] 24. Chen, J.S. et al. Enhanced proofreading governs CRISPR-Cas9 targeting accuracy. Nature 550, 407-410 (2017) .
[0772] 25. Vakulskas, C.A. et al. A high-fidelity Cas9 mutant delivered as a ribonucleoprotein complex enables efficient gene editing in human hematopoietic stem and progenitor cells. Nat Med 24, 1216-1224 (2018) .
[0773] 26. Kleinstiver, B.P. et al. High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects. Nature 529, 490-495 (2016) .
[0774] 27. Casini, A. et al. A highly specific SpCas9 variant is identified by in vivo screening in yeast. Nat Biotechnol 36, 265-271 (2018) .
[0775] 28. Hu, J.H. et al. Evolved Cas9 variants with broad PAM compatibility and high DNA specificity. Nature 556, 57-63 (2018) .
[0776] 29. Lee, J.K. et al. Directed evolution of CRISPR-Cas9 to increase its specificity. Nat Commun 9, 3048 (2018) .
[0777] 30. Fu, Y., Sander, J.D., Reyon, D., Cascio, V.M. &Joung, J.K. Improving CRISPR-Cas nuclease specificity using truncated guide RNAs. Nat Biotechnol 32, 279-284 (2014) .
[0778] 31. Kocak, D.D. et al. Increasing the specificity of CRISPR systems with engineered RNA secondary structures. Nat Biotechnol 37, 657-666 (2019) .
[0779] 32. Bravo, J.P.K. et al. Structural basis for mismatch surveillance by CRISPR-Cas9. Nature 603, 343-347 (2022) .
[0780] 33. Schmid-Burgk, J.L. et al. Highly Parallel Profiling of Cas9 Variant Specificity. Mol Cell 78, 794-800 e798 (2020) .
[0781] 34. Kim, D., Luk, K., Wolfe, S.A. &Kim, J.S. Evaluating and Enhancing Target Specificity of Gene-Editing Nucleases and Deaminases. Annu Rev Biochem 88, 191-220 (2019) .
[0782] 35. Kim, N. et al. Prediction of the sequence-specific cleavage activity of Cas9 variants. Nat Biotechnol 38, 1328-1336 (2020) .
[0783] 36. Han, D. et al. Development of miniature base editors using engineered IscB nickase. Nat Methods 20, 1029-1036 (2023) .
[0784] 37. Yan, H. et al. Assessing and engineering the IscB-omegaRNA system for programmed genome editing. Nat Chem Biol 20, 1617-1628 (2024) .
[0785] 38. Han, L. et al. Engineering miniature IscB nickase for robust base editing with broad targeting range. Nat Chem Biol 20, 1629-1639 (2024) .
[0786] 39. Xiao, Q. et al. Engineered IscB-omegaRNA system with expanded target range for base editing. Nat Chem Biol 21, 100-108 (2025) .
[0787] 40. Xue, N. et al. Engineering IscB to develop highly efficient miniature editing tools in mammalian cells and embryos. Mol Cell 84, 3128-3140 e3124 (2024) .
[0788] 41. Guo, R. et al. Engineered IscB-omegaRNA system with improved base editing efficiency for disease correction via single AAV delivery in mice. Cell Rep 43, 114973 (2024) .
[0789] 42. Altae-Tran, H. et al. The widespread IS200 / IS605 transposon family encodes diverse programmable RNA-guided endonucleases. Science 374, 57-65 (2021) .
[0790] 43. Zhao, C. et al. Evaluation of the effects of sequence length and microsatellite instability on single-guide RNA activity and specificity. Int J Biol Sci 15, 2641-2653 (2019) .
[0791] 44. Aliaga Goltsman, D.S. et al. Compact Cas9d and HEARO enzymes for genome editing discovered from uncultivated microbes. Nat Commun 13, 7602 (2022) .
[0792] 45. Qi, T. et al. Phage-assisted evolution of compact Cas9 variants targeting a simple NNG PAM. Nat Chem Biol 20, 344-352 (2024) .
[0793] 46. Huang, T.P. et al. High-throughput continuous evolution of compact Cas9 variants targeting single-nucleotide-pyrimidine PAMs. Nat Biotechnol 41, 96-107 (2023) .
[0794] 47. Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013) .
[0795] 48. Jinek, M. et al. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816-821 (2012) .
[0796] 49. Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013) .
[0797] 50. Yang, J. et al. Insights into the compact CRISPR-Cas9d system. Nat Commun 16, 2462 (2025) .
[0798] 51. Ocampo, R.F. et al. DNA targeting by compact Cas9d and its resurrected ancestor. Nat Commun 16, 457 (2025) .
[0799] 52. Kim, D.Y. et al. Efficient CRISPR editing with a hypercompact Cas12f1 and engineered guide RNAs delivered by adeno-associated virus. Nat Biotechnol 40, 94-102 (2022) .
[0800] 53. Anzalone, A.V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019) .
[0801] 54. Zhang, E., Neugebauer, M.E., Krasnow, N.A. &Liu, D.R. Phage-assisted evolution of highly active cytosine base editors with enhanced selectivity and minimal sequence context preference. Nat Commun 15, 1697 (2024) .
[0802] 55. Nishimasu, H. et al. Crystal structure of Cas9 in complex with guide RNA and target DNA. Cell 156, 935-949 (2014) .
[0803] 56. Kong, X. et al. Engineered CRISPR-OsCas12f1 and RhCas12f1 with robust activities and expanded target range for genome editing. Nat Commun 14, 2046 (2023) .
[0804] 57. Xu, X. et al. Engineered miniature CRISPR-Cas system for mammalian genome regulation and editing. Mol Cell 81, 4333-4345 e4334 (2021) .
[0805] 58. Kleinstiver, B.P. et al. Engineered CRISPR-Cas12a variants with increased activities and improved targeting ranges for gene, epigenetic and base editing. Nat Biotechnol 37, 276-282 (2019) .
[0806] 59. Strecker, J. et al. Engineering of CRISPR-Cas12b for human genome editing. Nat Commun 10, 212 (2019) .
[0807] 60....
Claims
1.A polypeptide comprising an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11.2.The polypeptide of claim 1, wherein the polypeptide has at least one of endonuclease activity, nickase activity, and DNA binding property.3.The polypeptide of claim 1 or 2, wherein the polypeptide comprises:(a) an amino acid substitution at a position of any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 that is corresponding to a position selected from the group consisting of D10, D386, H387, N410, H509, and D512 of SEQ ID NO: 13, for example, the positions of SEQ ID NO: 9 corresponding to the positions of D10, D386, H387, N410, H509, and D512 of SEQ ID NO: 13 are positions 10, D373, H374, N397, H496, and D499 of SEQ ID NO: 9, respectively, and optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ; or(b) an amino acid substitution relative to any one of SEQ ID NOs: 9, 13, 1, 3, 5, 7, and 11 that is corresponding to an amino acid substitution selected from the group consisting of D10A, D386A, H387A, N410A, H509A, and D512A relative to SEQ ID NO: 13, for example, the amino acid substitutions relative to SEQ ID NO: 9 corresponding to the amino acid substitutions of D10A, D386A, H387A, N410A, H509A, and D512A relative to SEQ ID NO: 13 are D10A, D373A, H374A, N397A, H496A, and D499A relative to SEQ ID NO: 9, respectively.4.The polypeptide of any one of claims 1-3, wherein the polypeptide comprises:(a) an amino acid substitution at a position selected from the group consisting of D10, D386, H387, N410, H509, and D512 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ;(b) an amino acid substitution at a position selected from the group consisting of 41, 42, 45, 46, 49, 50, 53, 56, 57, 58, 64, 66, 69, 72, 73, 80, 81, 82, 83, 84, 86, 87, 92, 95, 96, 102, 105, 106, 107, 108, 110, 112, 113, 114, 115, 116, 117, 118, 120, 121, 128, 135, 136, 138, 140, 141, 145, 153, 154, 157, 162, 163, 185, 187, 198, 200, 246, 250, 255, 268, 344, 349, 355, 368, 408, 412, 431, 432, 434, 443, 446, 449, 450, 479, 500, 523, 532, 535, 537, 549, 550, 553, 557, 560, 564, 575, 576, 577, 582, 583, 588, 594, 614, 618, 622, 626, 636, 640, 644, 645, 683, 684, 697, 698, 699, 710, 716, 718, 731, 732, 739, 740, 742, 743, 745, 747, 749, 751, 754, 758, and 759 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with Arginine (R) or Gly (G) ;(c) an amino acid substitution at a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with P, T, L, A, or D; or(d) a combination of any two or three of (a) , (b) , and (c) .5.The polypeptide of any one of claims 1-4, wherein the polypeptide comprises:(a) an amino acid substitution at a position of D10 and / or D386 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ;(b) an amino acid substitution at a position selected from the group consisting of 45, 46, 49, 120, 138, 200, 432, 549, 560, 582, 583, 640, 644, 710, and 747 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with Arginine (R) or Gly (G) ;(c) an amino acid substitution at a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with P, T, L, or A; or(d) a combination of any two or three of (a) , (b) , and (c) .6.The polypeptide of any one of claims 1-5, wherein the polypeptide comprises:(a) an amino acid substitution of D10A and / or D386A relative to SEQ ID NO: 13;(b) an amino acid substitution selected from the group consisting of A45R, G46R, E49R, D120R, Q138R, I200R, F432R, S549R, G560R, F582R, M583R, S640R, E644G, K710R, and T747R relative to SEQ ID NO: 13;(c) a combinational substitution of (i) R391P, A393T, I396L, and G400A or (ii) R391D, N397D, and G400A, relative to SEQ ID NO: 13; or(d) a combination of any two or three of (a) , (b) , and (c) .7.The polypeptide of any one of claims 1-6, wherein the polypeptide comprises:(a) D10A or D386A;(b) a combinational substitution selected from the group consisting of:(1) A45R, G46R, E49R, and T747R;(2) E49R, M583R, and T747R;(3) E49R, F432R, S640R, and T747R;(4) E49R, S549R, M583R, S640R, and T747R;(5) E49R, F432R, M583R, and T747R;(6) E49R, Q138R, M583R, S640R, and T747R;(7) E49R, Q138R, F432R, and T747R;(8) E49R, S549R, M583R, and T747R;(9) E49R, F432R, M583R, and T747R;(10) E49R, I200R, F432R, M583R, and T747R;(11) G46R, I200R, F432R, F582R, and T747R;(12) G46R, I200R, F582R, and T747R;(13) G46R, D120R, I200R, F432R, F582R, and T747R;(14) G46R, D120R, I200R, F432R, G560R, F582R, M583R, and T747R;(15) G46R, D120R, I200R, F432R, G560R, F582R, and T747R;(16) E49R, D120R, F432R, M583R, E644R, and T747R;(17) E49R, F432R, G560R, M583R, and T747R;(18) E49R, F432R, M583R, E644G, and T747R;(19) E49R, F432R, M583R, K710R, and T747R; or(20) E49R, Q138R, F432R, M583R, and T747R;(c) R391P, A393T, I396L, and G400A; or(d) a combination of any two or three of (a) , (b) , and (c) ,relative to SEQ ID NO: 13.8.The polypeptide of any one of claims 1-7, wherein the polypeptide comprises (a) D10A; (b) E49R, I200R, F432R, M583R, and T747R; and (c) R391P, A393T, I396L, and G400A, relative to SEQ ID NO: 13.9.The polypeptide of any one of claims 1-8, wherein the polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 21-24.10.The polypeptide of claim 1 or 2, wherein the polypeptide comprises:(a) an amino acid substitution at a position selected from the group consisting of D10, D373, H374, N397, H496, and D499 of SEQ ID NO: 9; optionally, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) ;(b) an amino acid substitution at a position selected from the group consisting of E44, D128, S546, Y568, M569, S623, and E667 of SEQ ID NO: 9; optionally, the amino acid substitution is an amino acid substitution with Arginine (R) ;(c) an amino acid substitution at a position of SEQ ID NO: 9 that is corresponding to a position selected from the group consisting of R391, A393, I396, N397, and G400 of SEQ ID NO: 13; optionally, the amino acid substitution is an amino acid substitution with P, T, L, A, or D; or(d) a combination of any two or three of (a) , (b) , and (c) .11.The polypeptide of any one of claims 1-2 and 10, wherein the polypeptide comprises:(a) an amino acid substitution of D10A and / or D373A relative to SEQ ID NO: 9;(b) an amino acid substitution selected from the group consisting of E44R, D128R, S546R, Y568R, M569R, S623R, and E667R relative to SEQ ID NO: 9;(c) an amino acid substitution relative to SEQ ID NO: 9 that is corresponding to a combinational substitution of (i) R391P, A393T, I396L, and G400A or (ii) R391D, N397D, and G400A relative to SEQ ID NO: 13; or(d) a combination of any two or three of (a) , (b) , and (c) .12.The polypeptide of any one of claims 1-2 and 10-11, wherein the polypeptide comprises:(a) D10A or D373A relative to SEQ ID NO: 9;(b) a combinational substitution selected from the group consisting of:(1) E44R+ Y568R E667R relative to SEQ ID NO: 9;(2) E44R+S546R+Y568R+E667R relative to SEQ ID NO: 9;(3) E44R+S546R+M569R+E667R relative to SEQ ID NO: 9;(4) E44R+D128R+S546R+Y568R+E667R relative to SEQ ID NO: 9; or(5) E44R+D128R+S546R+Y568R+S623R+E667R relative to SEQ ID NO: 9;(c) an amino acid substitution relative to SEQ ID NO: 9 that is corresponding to a combinational substitution of R391P, A393T, I396L, and G400A relative to SEQ ID NO: 13; or(d) a combination of any two or three of (a) , (b) , and (c) .13.The polypeptide of any one of claims 1-2 and 10-12, wherein the polypeptide comprises (a) D10A and (b) E44R+D128R+S546R+Y568R+S623R+E667R, relative to SEQ ID NO: 9.14.The polypeptide of any one of claims 1-2 and 10-13, wherein the polypeptide comprises an amino acid sequence of SEQ ID NO: 71.15.A fusion protein comprising the polypeptide of any preceding claim and a functional domain.16.The fusion protein of claim 15, wherein the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a deaminase or a catalytic domain thereof, an uracil glycosylase inhibitor (UGI) , an uracil glycosylase (UNG) , a methylpurine glycosylase (MPG) , a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease or a catalytic domain thereof (e.g., T5 exonuclease) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having RNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment thereof, and any combination thereof.17.The fusion protein of any preceding claim, wherein the deaminase or catalytic domain thereof is an adenine deaminase or a catalytic domain thereof (e.g., tRNA adenosine deaminase (TadA) , such as, TadA8e, TadA8.17, TadA8.20, TadA9, TadA8e-V106W, TadA8EV106W+D108Q TadA-CDa, TadA-CDb, TadA-CDc, TadA-CDd, TadA-CDe, TadA-dual, TADAC-1.2, TADAC-1.14, TADAC-1.17, TADAC-1.19, TADAC-2.5, TADAC-2.6, TADAC-2.9, TADAC-2.19, TADAC-2.23, TadA8e-N46L, TadA8e-N46P, TadA* (8.17m) , TadA (8.8m) ) ; or the deaminase or catalytic domain thereof is a cytosine deaminase or a catalytic domain thereof (e.g., an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase, an activation induced deaminase (AID) , a cytidine deaminase 1 from Petromyzon marinus (pmCDA1) , DddA, or a functional variant thereof, e.g., APOBEC1 (rAPOBEC1) , APOBEC2, APOBEC3, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, hAPOBEC3-W104A, CBE6c-V106W) .18.The fusion protein of any preceding claim, wherein the fusion protein further comprises a DNA binding domain; optionally, the DNA binding domain comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 72 or 54.19.The fusion protein of any preceding claim, wherein the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 25, 49-53, 70, and 73; or the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 31 and 64-67.20.The fusion protein of any preceding claim, wherein the fusion protein comprises the polypeptide and a reverse transcriptase or a catalytic domain thereof; optionally, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 36.21.The fusion protein of any preceding claim, wherein the fusion protein comprises the polypeptide, a DNMT3l domain, a DNMT3a domain, and a KRAB domain; optionally, the fusion protein comprises, consists essentially of, or consists of a sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to the sequence of SEQ ID NO: 40.22.A system comprising:(1) the polypeptide of any preceding claim or the fusion protein of any preceding claim, or a polynucleotide (e.g., a DNA, an RNA) encoding the polypeptide or the fusion protein, and(2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:(i) a scaffold sequence capable of forming a complex with the polypeptide or the fusion protein; and(ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.23.The system of any preceding claim, wherein the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48; or the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 80% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48; or the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, and 44-48.24.The system of any preceding claim, wherein the guide sequence is in a length of from about 19 to about 24 nucleotides; optionally, the guide sequence is about 20 nucleotides in length.25.A polynucleotide comprising a sequence encoding the polypeptide of any preceding claim or the fusion protein of any preceding claim.26.A vector comprising the polynucleotide of any preceding claim; optionally, the vector is a plasmid vector, a viral vector (e.g., a recombinant AAV (rAAV) vector, a recombinant lentivirus vector) , a ribonucleoprotein (RNP) , or a lipid nanoparticle (LNP) .27.A cell comprising the polypeptide of any preceding claim, the fusion protein of any preceding claim, the system of any preceding claim, the polynucleotide of any preceding claim, or the vector of any preceding claim.28.A method for modifying a target DNA, comprising contacting the target DNA with the system of any preceding claim, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.
Citation Information
Patent Citations
Endonuclease system
CN118265783A