Deaminases and uses thereof

Optimized TadA variants with amino acid substitutions enhance cytosine base editing efficiency, addressing off-target issues and achieving effective exon skipping and dystrophin restoration in a humanized mouse model.

WO2025180412A1PCT designated stage Publication Date: 2025-09-04HUIGENE THERAPEUTICS CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/079327
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-26
Filing Date
2025-02-26
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing cytidine deaminases used in DNA base editors exhibit off-target effects, motif bias, and insufficient cytosine editing activity, limiting their effectiveness in gene editing applications.

Method used

Development of TadA variants, particularly from Acinetobacter junii, with optimized amino acid substitutions to enhance cytosine base editing efficiency and reduce adenosine activity, combined with aTdCBE delivery using AAV for dystrophin restoration in a genetically humanized mouse model.

Benefits of technology

The optimized TadA variants demonstrate improved specificity and editing efficiency, enabling robust exon skipping and dystrophin restoration, providing a valuable tool for gene editing therapy and research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025079327-FTAPPB-I100001
    Figure PCTCN2025079327-FTAPPB-I100001
  • Figure PCTCN2025079327-FTAPPB-I100002
    Figure PCTCN2025079327-FTAPPB-I100002
  • Figure PCTCN2025079327-FTAPPB-I100003
    Figure PCTCN2025079327-FTAPPB-I100003
Patent Text Reader

Abstract

Provided herein are deaminases and fusion proteins or systems comprising the same, and uses thereof.
Need to check novelty before this filing date? Find Prior Art

Description

DEAMINASES AND USES THEREOF

[0001] REFERENCE TO RELATED APPLICATIONS

[0002] The instant application claims the priority to and the benefit of the filing date of PCT / CN2024 / 078613, filed on February 26, 2024, the entire contents of which, including any drawings and sequence listing, are incorporated herein by reference.

[0003] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0004] The disclosure contains a Sequence Listing XML file which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on February 25, 2025, by software “WIPO Sequence” according to WIPO Standard ST. 26, is named HGP041PCT. xml, and is 180, 090 bytes in size.

[0005] According to WIPO Standard ST. 26, symbol “t” is used to denote both T in DNA and U in RNA. Thus, in the instant sequence listing prepared according to ST. 26, wherever a sequence is an RNA, the T in the sequence shall be deemed as U.BACKGROUND

[0006] DNA base editors composed of Cas9 nickase and deaminases can perform specific single base conversion on target sequences without requiring DNA double-stranded breaks. The main effector domains of the cytosine base editor (CBE) and adenine base editor (ABE) are cytidine and adenine deaminase, catalyzing the base conversion from C·G to T·Aand A·T to G·C, respectively. Cytidine deaminases derived from proteins such as APOBEC, AID, and CDA naturally possess cytosine deamination activity. The adenosine deaminases TadA of ABEs were engineered from tRNA-specific adenosine deaminase for adenosine base editing of DNA. After removing uracil glycosylase inhibitor (UGI) from CBEs, C-to-G base editors (CGBEs) were developed to catalyze the conversion of bases from C·G to T·A.

[0007] By introducing stop codons (CAA / CAG / CGA to TAA / TAG / TGA) , CBEs have shown great potential in the treatment of common diseases such as T-cell acute lymphoblastic leukemia, Hepatitis B, and acquired immunodeficiency syndrome. Although high on-target DNA editing efficiency can be achieved with traditional CBEs, they also exhibit different types of DNA and RNA off-target effects. Compared to natural cytidine deaminases with high intrinsic single-stranded DNA (ssDNA) affinity, TadA of ABE exhibited undetectable guide-independent off-target effects at DNA and RNA levels. Previous studies reported that the TadA8e variant derived from E. coli TadA exhibited partial cytidine deaminase activity for TC sequences within the editing window. Recently, various CBEs, including Td-CBEmax and TadCBEd, have been developed by engineering TadA8e into a cytosine deaminase, demonstrating the engineering plasticity of TadA for versatile base editor development. However, these TadA-based CBEs still have defects such as motif bias, unwanted adenosine deaminase activity, or insufficient cytosine editing activity. Moreover, unwanted A-to-G base editing can cause the termination codon to become TGG, limiting the disruption effect of the target gene. Therefore, it would be desired to further develop suitable deaminases for base editing.

[0008] Citation or identification of any document in the disclosure is not an admission that such a document is available as prior art to the disclosure. Each of the references mentioned or cited in the disclosure is incorporated by reference in its entirety.SUMMARY

[0009] In the disclosure, the inventors utilized both TadA orthologs screening and protein mutagenesis strategies to obtain TadA variants, in which the TadA from Acinetobacter junii exhibited the most potent activity for cytosine base editing. Furthermore, by optimizing its activity through protein amino acid substitution, aTdCBE with better performance was obtained, including reduced motif preference and reduced adenosine activity, The inventors then established a genetically humanized DMD mouse model and demonstrated the potential of exon skipping and dystrophin restoration using AAV delivery of the base editor, which may inform future pre-clinical research. This study provides a practical strategy for the development of base editors with distinctive features, which is valuable for expanding the functional diversity of gene editing tools and their potential applications for gene editing therapy and research use.

[0010] In an aspect, the disclosure provides a cytidine deaminase, wherein the cytidine deaminase is a mutant of a reference polypeptide of SEQ ID NO: 19 and comprises an amino acid mutation (e.g., substitution) relative to (compared to) the reference polypeptide at a position selected from the group consisting of 8, 9, 19, 25, 41, 43, 45, 46, 47, 105, 107, 145, 146, and / or 154 of the reference polypeptide.

[0011] In another aspect, the disclosure provides a fusion protein comprising the cytidine deaminase of the disclosure fused to a heterogeneous functional domain.

[0012] In yet another aspect, the disclosure provides a system comprising:

[0013] (1) the fusion protein of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the fusion protein, and

[0014] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:

[0015] (i) a scaffold sequence capable of forming a complex with the fusion protein; and

[0016] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.

[0017] In yet another aspect, the disclosure provides a polynucleotide encoding the cytidine deaminase or the fusion protein of the disclosure.

[0018] In yet another aspect, the disclosure provides a delivery system comprising (1) the cytidine deaminase of the disclosure, the fusion of the disclosure, the polynucleotide of the disclosure, or the system of the disclosure; and (2) a delivery vehicle.

[0019] In yet another aspect, the disclosure provides a vector comprising the polynucleotide of the disclosure.

[0020] In yet another aspect, the disclosure provides a cell comprising the cytidine deaminase of the disclosure, the fusion of any preceding claim, the system of the disclosure, the polynucleotide of the disclosure, or the vector of the disclosure.

[0021] In yet another aspect, the disclosure provides a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, or the lipid nanoparticle of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.

[0022] The details of one or more embodiments of the disclosure are set forth in the description below. Other features or advantages of the disclosure will be apparent from the following drawings and detailed description of several embodiments, and also from the appended claims. It is understood that any aspect or embodiment of the disclosure can be combined with any other one or more aspects or embodiments of the disclosure, including aspects or embodiments only described in one sub-section, only in the examples, or only in the claims, to constitute another embodiment explicitly or implicitly disclosed herein unless otherwise indicated.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] An understanding of the features and advantages of the disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure may be utilized, and the accompanying drawings of which:

[0024] FIG. 1: Screening for TadA orthologs with cytosine deaminase activity in mammalian cells. FIG. 1a. The experimental workflow for detecting base editing activity with fluorescence reporter system in HEK293T cells. FIG. 1b. The Phylogenetic Tree of TadA orthologs. FIG. 1c, Screening for TadA orthologs using fluorescence reporter system. Data are presented as means ± s. d. Values and error bars represent mean and s. d., n=3 independent biological replicates. All of the above base editors do not contain UGI. PAM stands for protospacer adjacent motif. NC represents negative control.

[0025] FIG. 2: Engineering of AjTadA to improve base editing efficiency in mammalian cells. FIG. 2a. The experimental workflow for engineering of AjTadA. FIG. 2b. Substitutions of non-positively charged amino acids of AjTadA at the AHC sites. Each dot represents activity for a single variant. FIG. 2c. Substitutions of 25E and 46P of AjTadA. v1 with seven amino acid residues. Values and error bars represent mean and s. d., n=3 independent biological replicates. FIG. 2d. Combination of mutations at E25 and P46 of AjTadA. v1. Values and error bars represent mean and s. d., n=3. FIG. 2e. Combination of single amino acid mutation with AjTadA. v2 at three endogenous loci. Values present mean, n=2 independent biological replicates. All of the above base editors do not contain UGI.

[0026] FIG. 3: aTdCBE enables robust genomic base editing in mammalian cells. FIG. 3a. Comparison of C-to-T conversion efficiency for endogenous loci by base editors derived from aTdCBE, Td-CBEmax, TadCBEd, B3PCY2-CBE, hA3A*-CBE and YE1-BE4max in HEK293T cells. FIG. 3b. Base editing activity window plots showing mean C-to-T editing at all tested target positions. FIG. 3c. Base editing activity window plots showing mean A-to-G editing at all tested target positions. FIG. 3d. Comparison of motif preference for 25 endogenous loci by base editors derived from aTdCBE, Td-CBEmax, TadCBEd, B3PCY2-CBE, hA3A*-CBE and YE1-BE4max in HEK293T cells. The editing window of CBEs was primarily concentrated between bases 4-7. Consequently, the analysis focused on preference motifs exhibiting higher editing efficiency within this range. FIG. 3e. High-throughput library experiments to evaluate the motif preference of aTdCBE, Td-CBEmax, TadCBEd and B3PCY2-CBE. FIG. 3f. Introducing premature termination codons into the PCSK9 coding region using CBEs. The numbers in the grid represent the average editing efficiency of cytosines with corresponding motifs at positions 4-7 for each sgRNA. Data are presented as means ± s.d. Values represent n=3 independent biological replicates. P-values determined by one-sided Mann-Whitney U-test and adjusted by Benjamini-Hochberg procedure. All of the above base editors contain UGI. NC represents negative control.

[0027] FIG. 4: aTdCBE treatment robustly rescued dystrophin expression in TA six weeks after AAV injection. FIG. 4a.Overview for the in vivo intramuscular (IM) injection of the AAV9-aTdCBE construct into the tibialis anterior (TA) muscle of the right leg of 3-week-old DMDΔ54 mdx mice. The DMDΔE54 mdx mice were derived from the mating of the humanized DMDΔE54 mice with mdx mice. Left leg was injected with saline as a control. Black arrows indicate time points for tissue collection after injection. AjTadA. v2-Y146R, UGIs, Cas9n split by Rma intein and gRNA were packaged into 2 separated AAV particles. A muscle-specific promoter Spc5.12 was used to drive AjTadA-Cas9n-N or Cas9n-C. FIG. 4b. Schematic illustrating exon skipping strategies to restore the correct open reading frame (ORF) of the DMD transcript. The shape and color of the boxes representing DMD exons indicate the reading frame. Specifically, the deletion of exon 54 in the DMD gene results in a premature stop codon in exon 55 (depicted in red) . The in-frame ORF can be achieved by editing exon 55 splice acceptor sites (SAS) (ag base pair) . FIG. 4c. The DNA editing efficiency was analyzed after 6-week treatment. FIG. 4d. RT-PCR products from muscle of DMDΔE54 mdx mice were analyzed by gel electrophoresis. FIG. 4e. Dystrophin (Abcam, ab15277) is shown in green. Scale bar, 100 μm. FIG. 4f. Quantification of Dys+ fibers in cross sections of TA muscles. FIG. 4g. Western blot analysis of dystrophin (Sigma, D8168) and vinculin (CST, 13901S) expression in TA muscles 6 weeks after injection with AAV-aTdCBE or saline. For comparative analyses, wild-type (WT) mice were derived from crosses between STOCK Tg (DMD) 72Thoen / J mice (#018900) and mdx mice, which carry a c. 2977C>T, p. Gln993*mutation in exon 23 on Chr. X. Dilutions of protein extract from WT mice were used to standardize dystrophin expression (25%, 50%and 75%) . Each band (#1-4) represents an individual mouse sample. Vinculin was used as the loading control. Data are represented as mean ± s.e.m. (n=4 independent biological replicates) . Each dot represents an individual mouse.

[0028] FIG. 5: Sanger sequencing of AjTadA. v1 at PIK3CA loci.

[0029] FIG. 6: Conservative site analysis of TadA protein sequence based on multiple sequence alignment. The red triangle indicates highly conserved amino acids. The yellow, green, purple and blue squares represent adjacent amino acids with different distances from highly conserved sites. The amino acid frequency visualized by webogo (https:  / / weblogo. berkeley. edu / logo. cgi) .

[0030] FIG. 7: aTdCBE enables robust C-to-T genomic editing in mammalian cells. Data are presented as means ± s.d. Values represent n = 3 independent biological replicates. All of the above base editors contain UGI.

[0031] FIG. 8: A-to-G genomic editing of base editors in mammalian cells. Data are presented as means ± s.d. Values represent n = 3 independent biological replicates. All of the above base editors contain UGI.

[0032] FIG. 9: AjTadA-derived base editors with high specificity in mammalian cells. FIG. 9a. The gRNA-dependent off-target levels at the potential off-target sites. FIG. 9b. The gRNA-independent off-target activity at five R-loops formed by dSaCas9. FIG. 9c. Transcriptome-wide off-target analysis of aTdCBE and TadCBEd. Data are presented as means ± s. d. Values represent n = 3 independent biological replicates. All of the above base editors contain UGI.

[0033] FIG. 10: Comparison of the editing products of the base editors without UGI at three endogenous loci. Data are presented as means ± s. d. Values represent n = 3 independent biological replicates. All of the above base editors do not contain UGI.

[0034] FIG. 11: Comparing the editing efficiency of CBEs at the exon 55 SAS of DMD gene in HEK293T cells. SpG Cas9 were used to targeting the SAS-containing sequence with 3’-TGT PAM. Data are presented as means ± s. d. Values represent n = 3 independent biological replicates. All of the above base editors contain UGI.

[0035] FIG. 12: Establishment and characterization of a humanized DMD mouse model. FIG. 12a. Strategy for generating a humanized DMD mouse model. CRISPR-Cas9 editing was employed to delete human DMD exon54. FIG. 12b. RT-PCR analysis of TA muscles to validate deletion of human exon 54. FIG. 12c. Dystrophin immunohistochemistry from indicated muscles of WT and DMDΔE54 mdx mice. WT mice were derived from crosses between STOCK Tg (DMD) 72Thoen / J mice (#018900) and mdx mice, which carry a c. 2977C>T, p. Gln993*mutation in exon 23 on Chr. X. Dystrophin (Abcam, ab15277) and spectrin (Millipore, MAB1622) are shown in green and mangenta, respectively. FIG. 12d. Western blot confirming the absence of dystrophin (Sigma, D8168) in indicated muscle tissues. FIG. 12e. Sirius red staining and HE staining of TA, DI, and heart muscle of WT and DMDΔE54 mdx mice. FIG. 12f Serum CK, a marker of muscle damage and membrane leakage, was measured in WT and DMDΔE54 mdx mice. FIG. 12g. WT and DMDΔE54 mdx mice were subjected to forelimb grip strength testing to measure muscle performance. All mice were 4 weeks old at the time of the experiment. Data are represented as mean ± s.e.m. (n=6 independent biological replicates) . Each dot represents an individual mouse. **P < 0.01 using unpaired two-tailed Student’s t test. Scale bar, 100 μm.

[0036] FIG. 13: Rescue of dystrophin expression following intramuscular injection of aTdCBE after 6 weeks. Dystrophin immunohistochemistry of TA muscle. Control mice were injected with saline. Images shown in both FIG. 4e and FIG. 12 were obtained from the same tissue at 20x magnification. FIG. 4e showed the local region staining image rather than the reconstituted whole-tissue scanning image in FIG. 12, and highlighted with white boxes. Dystrophin is shown in green. Scale bar, 500 μm.

[0037] FIG. 14: Uncropped images. The red rectangles indicate the cropping location.

[0038] FIG. 15: Flow cytometry gating strategy. Cell singletons were first gated out via forward scatter (FSC) and side scatter (SSC) parameters. Fluorescent cells were then gated for gene editing analysis.

[0039] FIG. 16 illustrates an exemplary target DNA, and an exemplary base editing system comprising (1) an exemplary guide nucleic acid comprising a guide sequence and a scaffold sequence and (2) an exemplary base editor comprising IscB nickase fused to a deaminase.

[0040] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.DETAILED DESCRIPTION

[0041] The disclosure will be described with respect to particular embodiments, but the disclosure is not limited thereto in any respect. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which this disclosure belongs. Terms as set forth hereinafter are generally to be understood in their plain and ordinary meaning or common sense unless indicated otherwise. The definitions and explanations of terms in WO2022087494A1 are incorporated herein by reference in their entireties except for the extent that a different definition or explanation of a term is specifically provided herein.

[0042] Definition

[0043] Nucleic acid programmable DNA binding protein (napDNAbp) , for example, IscB, Cas9, Cas12, is capable of binding to a DNA (e.g., a dsDNA) as guided by a guide nucleic acid (e.g., a guide RNA) comprising a guide sequence targeting the DNA. napDNAbp may be associated with the guide nucleic acid (e.g., a guide RNA) , which localizes  / targets the napDNAbp to a DNA that comprises a DNA strand (i.e., a target strand) that is reversely complementary to the guide nucleic acid, or a portion thereof (e.g., the guide sequence of a guide RNA) . In other words, the guide nucleic acid “programs” the napDNAbp to localize and bind to the DNA. Binding of the napDNAbp to the DNA enables the napDNAbp or a construct comprising the napDNAbp to access to and function on the DNA.

[0044] Without wishing to be bound by theory, in some embodiments, the guide nucleic acid comprises a scaffold sequence responsible for forming a complex with the napDNAbp, and a guide sequence that is intentionally designed to be responsible for hybridizing to a target sequence of the DNA, thereby guiding the complex comprising the napDNAbp and the guide nucleic acid to the DNA such that the napDNAbp is indirectly bound to the DNA.

[0045] Referring to FIG. 16, an exemplary dsDNA is depicted to comprise a 5’ to 3’ single DNA strand and a 3’ to 5’ single DNA strand, the 5’ to 3’ single DNA strand comprises an exemplary first deoxyribonucleotide dA, and the 3’ to 5’ single DNA strand comprises an exemplary second deoxyribonucleotide dC that base pairs with the dT.

[0046] An exemplary guide nucleic acid is depicted to comprise a guide sequence and a scaffold sequence. The guide sequence is designed to hybridize to a part of the 3’ to 5’ single DNA strand, and so the guide sequence “targets” that part. And thus, the 3’ to 5’ single DNA strand is referred to as a “target strand (TS) ” of the dsDNA, while the opposite 5’ to 3’ single DNA strand is referred to as a “nontarget strand (NTS) ” of the dsDNA. That part of the target strand based on which the guide sequence is designed and to which the guide sequence may hybridize is referred to as a “target sequence” , while the opposite part on the nontarget strand corresponding to that part is referred to as the “protospacer sequence” , which is typically 100%(fully) reversely complementary to the target sequence, if there is no intentional or unintentional mismatch.

[0047] Generally, as is conventional in the art, a nucleic acid sequence (e.g., a DNA sequence) is written in 5’ to 3’ direction  / orientation unless explicitly indicated otherwise.

[0048] For example, for a DNA sequence of ATGC, it is usually understood as 5’-ATGC-3’ unless otherwise indicated. Its reverse sequence is 5’-CGTA-3’ . Its fully complementary sequence is 5’-TACG-3’ . Its fully reverse complementary sequence is 5’-GCAT-3’ . Note that the fully complementary sequence usually does not have the ability to base-pair  / hybridize with the original sequence.

[0049] Generally, the double-strand sequence of a dsDNA may be represented with the sequence of its 5’ to 3’ single DNA strand conventionally written in 5’ to 3’ direction  / orientation unless otherwise indicated.

[0050] For example, for a dsDNA having a 5’ to 3’ single DNA strand of 5’-ATGC-3’ and a 3’ to 5’ single DNA strand of 3’-TACG-5’ as shown below, the dsDNA may be simply represented as 5’-ATGC-3’.

[0051] 5’-----ATGC -----3’

[0052] 3’-----TACG -----5’

[0053] It should be noted that either the 5’ to 3’ single DNA strand or the 3’ to 5’ single DNA strand of a dsDNA can be a nontarget strand from which a protospacer sequence is selected.

[0054] In the sense of base editing, the strand on which the target nucleotide (e.g., deoxyribonucleotide dA) to be edited is located is termed as an edited strand, and the opposite strand is termed as a non-edited strand. As used herein, the nontarget strand is the edited strand, and the target strand is the non-edited strand.

[0055] Typically for a gene, the 5’ to 3’ single DNA strand of the gene is the sense strand, and the 3’ to 5’ single DNA strand of the gene is the antisense strand. Either the sense strand or the antisense strand can be a nontarget strand from which a protospacer sequence is selected.

[0056] To hybridize to a dsDNA, such as, a dsDNA 5’-ATGC-3’, the guide sequence of a guide nucleic acid, in one embodiment, is designed to have a sequence of 5’-AUGC-3’ that is fully reversely complementary to the 3’ to 5’ strand of the dsRNA (3’-TACG-5’ ) , which would be set forth in ATGC in the electric sequence listing and marked as an RNA sequence according to WIPO standard ST. 26; and in another embodiment, the guide sequence of a guide nucleic acid is designed to have a sequence of 5’-GCAU-3’ that is fully reversely complementary to the 5’ to 3’ strand of the dsDNA (5’-ATGC-3’ ) , which would be set forth in GCAT in the electric sequence listing and marked as an RNA sequence according to WIPO standard ST. 26.

[0057] In the case that the guide sequence of a guide nucleic acid is fully reversely complementary to the target sequence and the target sequence is fully reversely complementary to the protospacer sequence, the guide sequence is identical to the protospacer sequence except for the difference between the U in the guide sequence due to its RNA nature and the corresponding T in the protospacer sequence due to its DNA nature. According to WIPO standard ST. 26, symbol “t” is used to denote both T in DNA and U in RNA (See “Table 1: List of nucleotides symbols” , the definition of symbol “t” is “thymine in DNA / uracil in RNA (t / u) ” ) . Thus, in the electronic sequence listing of the disclosure prepared according to WIPO standard ST. 26, such a guide sequence could be set forth in the same sequence as a corresponding protospacer sequence. For convenience, a single SEQ ID NO in the electronic sequence listing can be used to denote both such guide sequence and protospacer sequence, although such a single SEQ ID NO may be marked as either DNA or RNA in the electronic sequence listing. When a reference is made to such a SEQ ID NO that sets forth a protospacer  / guide sequence, it refers to either a protospacer sequence that is a DNA sequence or a guide sequence that is an RNA sequence depending on the context, no matter whether it is marked as a DNA or an RNA in the electronic sequence listing.

[0058] As used herein, the terms “nucleic acid” , “nucleic acid molecule” , or “polynucleotide” are used interchangeably. They refer to a polymer of deoxyribonucleotides or ribonucleotides or their mixtures of any length in either single-or double-stranded form, and, unless otherwise stated, encompass known analogs of natural nucleotides that can function in a similar manner as naturally occurring nucleotides. The terms encompass nucleic acid-like structures with synthetic backbones, as well as amplification products. DNAs and RNAs are both polynucleotides. The polymer may include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine) , nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O (6) -methylguanine, and 2-thiocytidine) , chemically modified bases, biologically modified bases (e.g., methylated bases) , intercalated bases, modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose) , or modified phosphate groups (e.g., phosphorothioates and 5′-N-phosphoramidite linkages) .

[0059] As used herein, the phrase “polynucleotide encodes / encoding polypeptide X” or a similar phrase refers to a polynucleotide that is translated to express polypeptide Y comprising polypeptide X, meaning that polypeptide X is all or part of polypeptide Y. For example, polynucleotide encoding a Cas protein can refer to (i) a polynucleotide that is translated to express a fusion protein comprising the Cas protein and one or more additional amino acids, wherein the Cas protein is part of the fusion protein; or alternatively, (ii) a polynucleotide that is translated to express the Cas protein per se without any additional amino acid.

[0060] As used herein, the phrase “polynucleotide encodes / encoding RNA X” or a similar phrase refers to a polynucleotide that is transcribed to RNA Y comprising RNA X, meaning that RNA X is all or part of RNA Y. For example, polynucleotide encoding a gRNA can refer to (i) a polynucleotide that is transcribed to an RNA comprising the gRNA and one or more additional nucleotides, wherein the gRNA is part of the transcribed RNA; or alternatively, (ii) a polynucleotide that is transcribed to the gRNA per se without any additional nucleotide.

[0061] As used herein, the term “polypeptide” and “protein” are used interchangeably to refer to a polymer of amino acids of any length. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified; for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component.

[0062] As used herein, a “fusion protein” refers to a protein created through the joining of two or more originally separate proteins, or portions thereof. In some embodiments, a linker may be present between each protein.

[0063] As used herein, the term “heterologous, ” in reference to polypeptide domains, refers to the fact that the polypeptide domains do not naturally occur together (e.g., in the same polypeptide) . For example, in fusion proteins generated by the hand of man, a polypeptide domain from one polypeptide may be fused to a polypeptide domain from a different polypeptide. The two polypeptide domains would be considered “heterologous” with respect to each other, as they do not naturally occur together.

[0064] As used herein, the term “heterologous, ” in reference to nucleotide sequences, refers to the fact that the nucleotide sequences do not naturally occur together (e.g., in the same polynucleotide) . For example, in a guide nucleic acid generated by the hand of man, a guide sequence intentionally designed to target a human gene locus may be fused to a scaffold sequence from a microorganism. The two nucleotide sequences would be considered “heterologous” with respect to each other, as they do not naturally occur together.

[0065] As used herein, the term “guide nucleic acid” refers to a nucleic acid-based molecule capable of forming a complex with a nucleic acid programmable protein, for example, an IscB polypeptide (e.g., via a scaffold sequence of the guide nucleic acid) , and comprises a sequence (e.g., a guide sequences) that is sufficient to hybridize to a target nucleic acid and guides the complex to the target nucleic acid, which includes but is not limited to RNA-based molecules, e.g., a guide RNA. As used herein, the terms “guide RNA (gRNA) ” , “omega RNA” , “ωRNA” , and “RNA guide” are used interchangeably. As used in the disclosure, the term “guide sequence” is used interchangeably with the term “spacer sequence” .

[0066] As used herein, the term “complex” refers to a grouping of two or more molecules. In some embodiments, the complex comprises a nucleic acid and a polypeptide interacting with (e.g., binding to, coming into contact with, adhering to) one another. As used herein, the term “complex” can refer to a grouping of a guide nucleic acid and a polypeptide (e.g., an IscB polypeptide) . As used herein, the term “complex” can refer to a grouping of a guide nucleic acid, a polypeptide (e.g., an IscB polypeptide) , and a target nucleic acid (e.g., a target DNA) .

[0067] As used herein, if a DNA sequence, for example, 5’-ATGC-3’ is transcribed to an RNA sequence, with each dT (deoxythymidine, or “T” for short) in the primary sequence of the DNA sequence replaced with a U (uridine) and each dA (deoxyadenosine, or “A” for short) , dG (deoxyguanosine, or “G” for short) , and dC (deoxycytidine, or “C” for short) replaced with A (adenosine) , G (guanosine) , and C (cytidine) , respectively, for example, 5’-AUGC-3’, it is said in the disclosure that the DNA sequence “encodes” the RNA sequence.

[0068] As used herein, the terms “protospacer adjacent motif (PAM) ” and “target adjacent motif (TAM) ” are used interchangeably and refer to a short sequence (or a motif) adjacent to a protospacer sequence on the nontarget strand of a dsDNA recognizable by an napDNAbp or a complex comprising an napDNAbp and a guide nucleic acid. In some embodiments, the PAM or TAM is immediately 3’ to a protospacer sequence.

[0069] As used herein, the term “adjacent” includes instances wherein there is no nucleotide between the protospacer sequence and the PAM and also instances wherein there are a small number (e.g., 1, 2, 3, 4, or 5) of nucleotides between the protospacer sequence and the PAM. As used herein, A “immediately adjacent (to) ” B, A “immediately 5’ to” B, and A “immediately 3’ to” B mean that there is no nucleotide between A and B.

[0070] As described herein, the guide sequence is so designed to be substantially capable of hybridizing to a target sequence. As used herein, the term “hybridize” , “hybridizing” , or “hybridization” refers to a reaction in which one or more polynucleotide sequences react to form a complex that is stabilized via hydrogen bonding between the bases of the one or more polynucleotide sequences. The hydrogen bonding may occur by Watson Crick base pairing, Hoogstein binding, or in any other sequence specific manner. A polynucleotide sequence capable of hybridizing to a given polynucleotide sequence is referred to as the “complement” of the given polynucleotide sequence. As used herein, the hybridization of a guide sequence and a target sequence is so stabilized to permit an napDNAbp that is complexed with a guide nucleic acid comprising the guide sequence or a function domain (e.g., a deaminase domain) associated (e.g., fused) with the napDNAbp to act (e.g., cleave, deaminize) at or near the target sequence or its complement (e.g., a sequence of a target DNA or its complement) .

[0071] For the purpose of hybridization, in some embodiments, the guide sequence is reversely complementary to a target sequence. As used herein, the term “reverse complementary” refers to the ability of nucleobases of a first polynucleotide sequence, such as a guide sequence, to base pair with nucleobases of a second polynucleotide sequence, such as a target sequence, by traditional Watson-Crick base-pairing. Two reverse complementary polynucleotide sequences are able to non-covalently bind under appropriate temperature and solution ionic strength conditions. In some embodiments, a first polynucleotide sequence (e.g., a guide sequence) comprises 100% (fully) reverse complementarity to a second nucleic acid (e.g., a target sequence) . In some embodiments, a first polynucleotide sequence (e.g., a guide sequence) is reverse complementary to a second polynucleotide sequence (e.g., a target sequence) if the first polynucleotide sequence comprises at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%complementarity to the second nucleic acid (i.e., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%of the nucleotides of the first polynucleotide sequence can base-pair with the nucleotides of the second polynucleotide sequence) . As used herein, the term “substantially complementary” refers to a polynucleotide sequence (e.g., a guide sequence) that has a certain level of complementarity to a second polynucleotide sequence (e.g., a target sequence) (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%of the guide sequence can base-pair with the polynucleotide sequence of the target sequence, or at most 1, 2, 3, 4, or 5 contiguous or non-contiguous nucleotides of the guide sequence mismatch the nucleotides of the target sequence) . In some embodiments, the level of complementarity is such that the first polynucleotide sequence (e.g., a guide sequence) can hybridize to the second polynucleotide sequence (e.g., a target sequence) with sufficient affinity to permit an napDNAbp that is complexed with the first polynucleotide sequence or a nucleic acid comprising the first polynucleotide sequence or a function domain (e.g., a deaminase domain) associated (e.g., fused) with the napDNAbp to act (e.g., cleave, deaminize) on the target sequence or its complement (e.g., a sequence of a target DNA or its complement) . In some embodiments, a guide sequence that is substantially complementary to a target sequence has 100%or less than 100%complementarity to the target sequence. In some embodiments, a guide sequence that is substantially complementary to a target sequence has at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%complementarity to the target sequence, and / or has at most 1, 2, 3, 4, or 5 contiguous or non-contiguous nucleotide mismatches from the target sequence.

[0072] As used herein, the term “identity” refers to the overall relatedness between polymeric molecules, e.g., between nucleic acids (e.g., DNA and / or RNA) and / or between polypeptides. In some embodiments, polymeric molecules are considered to be “substantially identical” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%identical. Calculation of the percent identity of two nucleic acids or polypeptides, for example, can be performed by aligning the two sequences for optimal comparison purpose (e.g., gaps can be introduced in one or both of a first and a second sequences for optimal alignment and non-identical sequences can be disregarded for comparison purposes) . In certain embodiments, the length of a sequence aligned for comparison purpose is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or substantially 100%of the length of a reference sequence. The nucleotides at corresponding positions are then compared. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. As is well known in the art, nucleic acids or polypeptides may be compared using any of a variety of algorithms, including those available in commercial computer programs such as BLASTN for nucleotide sequences and BLASTP, gapped BLAST, and PSI-BLAST for amino acid sequences. In some embodiments, the sequence identity is calculated by global alignment, for example, using the Needleman-Wunsch algorithm and an online tool at ebi. ac. uk / Tools / psa / emboss_needle / . In some embodiments, the sequence identity is calculated by local alignment, for example, using the Smith-Waterman algorithm and an online tool at ebi. ac. uk / Tools / psa / emboss_water / .

[0073] As used herein, the term “variant” refers to an entity that shows significant structural identity with a reference entity (e.g., a wild-type sequence) but differs structurally from the reference entity in the presence or level of one or more chemical moieties as compared with the reference entity. In many embodiments, a variant also differs functionally from its reference entity. In general, whether a particular entity is properly considered to be a “variant” of a reference entity is based on its degree of structural identity with the reference entity. As will be appreciated by those skilled in the art, any biological or chemical reference entity has certain characteristic structural elements. A variant, by definition, is a distinct chemical entity that shares one or more such characteristic structural elements. To give but a few examples, a polypeptide may have a characteristic sequence element comprising a plurality of amino acids having designated positions relative to one another in linear or three-dimensional space and / or contributing to a particular biological function; a nucleic acid may have a characteristic sequence element comprising a plurality of nucleotide residues having designated positions relative to one another in linear or three-dimensional space. For example, a variant polypeptide may differ from a reference polypeptide as a result of one or more differences in amino acid sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, etc. ) covalently attached to the polypeptide backbone. In some embodiments, a variant polypeptide shows an overall sequence identity with a reference polypeptide (e.g., a nuclease described herein) that is at least about 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%or 99%. Alternatively or additionally, in some embodiments, a variant polypeptide does not share at least one characteristic sequence element with a reference polypeptide, for example, an IscB nickase as a variant of a reference IscB polypeptide does not share an active RuvC domain or an active HNH domain with the reference IscB polypeptide. In some embodiments, the reference polypeptide has one or more biological activities. In some embodiments, a variant polypeptide shares one or more of the biological activities of the reference polypeptide, e.g., nuclease activity. In some embodiments, a variant polypeptide lacks one or more of the biological activities of the reference polypeptide, for example, an IscB nickase as a variant of a reference IscB polypeptide does not share the endonuclease activity of the reference IscB polypeptide. In some embodiments, a variant polypeptide shows a reduced level of one or more biological activities (e.g., nuclease activity, e.g., off-target nuclease activity) as compared with the reference polypeptide. In some embodiments, a polypeptide of interest is considered to be a “variant” of a reference polypeptide if the polypeptide of interest has an amino acid sequence that is identical to that of the reference polypeptide but for a small number of sequence alterations at particular positions. Typically, fewer than about 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%of the residues in the variant are substituted as compared with the reference polypeptide. In some embodiments, a variant has about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residue as compared with a reference polypeptide. Often, a variant has a very small number (e.g., fewer than about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1) of substituted functional residues (i.e., residues that participate in a particular biological activity) . In some embodiments, a variant has not more than about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 additions or deletions, and often has no additions or deletions, as compared with the reference polypeptide. Moreover, any additions or deletions are typically fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 12, about 11, about 10, about 9, about 8, about 7, about 6, and commonly are fewer than about 5, about 4, about 3, or about 2 residues. In some embodiments, the reference polypeptide is a wild type polypeptide. A variant of a polynucleotide may be naturally occurring such as an allelic variant, or it may be a variant that is not known to occur naturally. Non-naturally occurring variants of a polynucleotide may be made by mutagenesis techniques, by direct synthesis, and by other recombinant methods known to skilled artisans.

[0074] As used herein, the terms “non-naturally occurring” and “engineered” are used interchangeably and refer to artificial participation. When these terms are used to describe a nucleic acid or a polypeptide, it is meant that the nucleic acid or polypeptide is at least substantially freed from at least one other component of its association in nature or as found in nature.

[0075] In some embodiments, a “conservative substitution” refers to a substitution of an amino acid made among amino acids within one of the following four groups:

[0076] (1) non-polar amino acids, including Glycine (Gly / G) , Alanine (Ala / A) , Valine (Val / V) , Cysteine (Cys / C) , Proline (Pro / P) , Leucine (Leu / L) , Isoleucine (Ile / I) , Methionine (Met / M) , Tryptophan (Trp / W) , and Phenylalanine (Phe / F) ;

[0077] (2) negatively charged amino acids, including Aspartic Acid (Asp / D) and Glutamic Acid (Glu / E) ;

[0078] (3) polar amino acids, including Serine (Ser / S) , Threonine (Thr / T) , Tyrosine (Tyr / Y) , Asparagine (Asn / N) , and Glutamine (Gln / Q) ; and

[0079] (4) positively charged amino acids, including Lysine (Lys / K) , Arginine (Arg / R) , and Histidine (His / H) .

[0080] As used herein, the term “wild type” has the meaning commonly understood by those skilled in the art to mean a typical form of an organism, a strain, a gene, or a feature that distinguishes it from a mutant or variant when it exists in nature. It can be isolated from sources in nature and not intentionally modified.

[0081] As used herein, the description of a mutant  / engineered polypeptide (e.g., of WT OgeuIscB) “comprising an amino acid mutation (e.g., substitution) at a position corresponding to a given position (e.g., D61) of a given polypeptide (e.g., WT OgeuIscB” or similar description means that the given polypeptide serves as a parent or reference polypeptide that does not comprises an amino acid mutation at the given position, and the mutant is a mutant of the parent or reference polypeptide and comprises an amino acid mutation at a position of the amino acid sequence of the mutant corresponding to the given position of the amino acid sequence of the given polypeptide. The position of the amino acid mutation in the amino acid sequence of the mutant may be the same as the given position of the given polypeptide, for example, when the mutant has exactly the same length as the given polypeptide. The position of the amino acid mutation in the amino acid sequence of the mutant may be different from the given position of the given polypeptide, for example, when the mutant does not have exactly the same length as the given polypeptide, for example, when the mutant comprises a N-terminal truncation as compared with the given polypeptide and thus the first N-terminal amino acid of the mutant is not corresponding to the first N-terminal amino acid of the given polypeptide but to an internal amino acid within the given polypeptide, but the position of the amino acid mutation in the mutant can be determined by alignment of the mutant and the given polypeptide to identify the corresponding amino acids in the two sequences as understood by a skilled in the art. For example, if the mutant has a N-terminal truncation of 20 amino acids as compared with the given polypeptide, then the mutant comprising an amino acid mutation at a position corresponding to D61 of a given polypeptide means that the mutant comprises an amino acid mutation at position 41 of the mutant since position 41 in the mutant is corresponding to D61 in the given polypeptide as determined by alignment of the mutant and the given polypeptide.

[0082] As used herein, the description of a mutant  / engineered polypeptide (e.g., of WT OgeuIscB) “comprising an amino acid substitution corresponding a given amino acid substitution (e.g., D61A) relative to a given polypeptide (e.g., WT OgeuIscB) ” means that the given polypeptide serves as a parent or reference polypeptide that does not comprise the given amino acid substitution, and the mutant is a mutant of the parent or reference polypeptide and comprises the same type of amino acid substitution (e.g., D-to-Asubstitution) as the given amino acid substitution at a position in the mutant corresponding to the position (e.g., D61) of the given amino acid substitution (e.g., D61A) numbered according to the given polypeptide. For example, an engineered IscB polypeptide comprising an amino acid substitution corresponding to D61A relative to wild type OgeuIscB refers to the fact that the wild type OgeuIscB comprises amino acid D (Asp) at position 61, and the engineered IscB polypeptide comprises amino acid A (Ala) at a position corresponding to D61 of wild type OgeuIscB. The corresponding relationship of positions in the two amino acid sequences as determined by alignment is explained in the previous paragraph.

[0083] As used herein, the terms “upstream” and “downstream” refer to relative positions within a single nucleic acid (e.g., DNA) sequence in a nucleic acid. “Upstream” and “downstream” relate to the 5’ to 3’ direction, respectively, in which transcription occurs. For a first sequence and a second sequence present on the same strand of a single nucleic acid written in 5’ to 3’ direction, the first sequence is upstream of the second sequence when the 3’ end of the first sequence is on the left side of the 5’ end of the second sequence, and the first sequence is downstream of the second sequence when the 5’ end of the first sequence is on the right side of the 3’ end of the second sequence. For example, a promoter is usually at the upstream of a sequence under the regulation of the promoter; and on the other hand, a sequence under the regulation of a promoter is usually at the downstream of the promoter.

[0084] As used herein, the term “regulatory element” refers to a DNA sequence that controls or impacts one or more aspects of transcription and / or expression and is intended to include promoters, enhancers, silencers, termination signals, internal ribosome entry sites (IRES) , and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences) . Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences) . Regulatory elements may also direct expression in a time-dependent manner, e.g., in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type specific.

[0085] As used herein, the term “operably linked” refers to a juxtaposition wherein the components described are in a relationship permitting them to function in their intended manner. A regulatory element “operably linked” to a functional element is associated in such a way that transcription, expression, and / or activity of the functional element is achieved under conditions compatible with the regulatory element. In some embodiments, “operably linked” regulatory elements are contiguous (e.g., covalently linked) with the functional elements of interest; in some embodiments, regulatory elements act in trans to or otherwise at a distance from the functional elements of interest.

[0086] As used herein, the term “cell” is understood to refer not only to a particular individual cell, but to the progeny or potential progeny of the cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term.

[0087] As used herein, the term “in vivo” means inside the body of an organism, and the terms “ex vivo” or “in vitro” means outside the body of an organism.

[0088] As used herein, the term “treat” , “treatment” , or “treating” is an approach for obtaining beneficial or desired results including clinical results. For purposes of the disclosure, the beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms resulting from a disease, diminishing the extent of a disease, stabilizing a disease (e.g., delaying the worsening of a disease) , delaying the spread (e.g., metastasis) of a disease, delaying the recurrence of a disease, reducing recurrence rate of a disease, delay or slowing the progression of a disease, ameliorating a disease state, providing a remission (partial or total) of a disease, decreasing the dose of one or more other medications required to treat a disease, delaying the progression of a disease, increasing the quality of life, and prolonging survival. Also encompassed by the term is a reduction of pathological consequence of a disease (such as cancer) . The methods of the disclosure contemplate any one or more of these aspects of treatment.

[0089] As used herein, the term “disease” includes the terms “disorder” and “condition” and is not limited to those specific diseases that have been medically or clinically defined.

[0090] As used herein, reference to “not” a value or parameter generally means and describes “other than” a value or parameter. For example, the method is not used to treat cancer of type X means the method may be used to treat cancer of types other than X.

[0091] As used herein, the singular forms “a” , “an” , and “the” include plural referents unless the context clearly dictates otherwise. That is, articles “a / an” and “the” are used herein to refer to one or more than one (i.e., at least one) grammatical object of the article. For example, “an element” means one element or more than one element, e.g., two elements.

[0092] As used herein, the term “and / or” in a phrase such as “Aand / or B” is intended to mean either or both of the alternatives, including both A and B, A or B, A (alone) , and B (alone) . Likewise, the term “and / or” in a phrase such as “A, B, and / or C” is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone) ; B (alone) ; and C (alone) .

[0093] As used herein, when the term “about” is ahead of a serious of numbers (for example, about 1, 2, 3) , it is understood that each of the serious of numbers is modified by the term “about” (that is, about 1, about 2, about 3) . The term “about X-Y” used herein has the same meaning as “about X to about Y. ”

[0094] As used herein, a numerical range includes the end values of the range, and each specific value within the range, for example, “16 to 100 nucleotides” includes 16 nucleotides and 100 nucleotides, and each specific value between 16 and 100, e.g., 17, 23, 34, 52, 78.

[0095] It is understood that embodiments of the disclosure described herein include “consisting” and / or “consisting essentially of” embodiments.

[0096] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely” , “only” , and the like in connection with the recitation of claim elements, or use of a “negative” limitation.

[0097] As used herein, the term “reference polypeptide” is used in the context of designing and developing a new polypeptide based on an original polypeptide (e.g., a wild-type polypeptide) , for example, the original polypeptide is mutated to generate the new polypeptide. In that case, the original polypeptide is a reference of the new polypeptide and termed as a reference polypeptide. The properties of a new polypeptide can be evaluated with a reference polypeptide as a reference from which the new polypeptide is derived. For example, one or more of the properties (e.g., endonuclease activity, nickase activity) of the new polypeptide can be compared with the reference polypeptide from which the new polypeptide is derived. As used herein, the term “engineered polypeptide” refers to a polypeptide artificially designed and developed based on a reference polypeptide (e.g., a wild-type polypeptide) , for example, by introducing an amino acid mutation.

[0098] As used herein, the term “endonuclease activity” is used interchangeably with “dsDNA cleavage activity” herein, and the term “nickase activity” is used interchangeably with “ssDNA cleavage activity” herein. As used herein, the term “nick” is used interchangeably with “ssDNA cleavage” . Unless otherwise indicated, the term “endonuclease activity” refers to guide sequence specific (on-target) endonuclease activity. Unless otherwise indicated, the term “nickase activity” refers to guide sequence specific (on-target) nickase activity.

[0099] As used herein, the phrase “polynucleotide encodes / encoding polypeptide X” or a similar phrase refers to a polynucleotide that is translated to express polypeptide Y comprising polypeptide X, meaning that polypeptide X is all or part of polypeptide Y. For example, polynucleotide encoding a Cas protein can refer to (i) a polynucleotide that is translated to express a fusion protein comprising the Cas protein and one or more additional amino acids, wherein the Cas protein is part of the fusion protein; or alternatively, (ii) a polynucleotide that is translated to express the Cas protein per se without any additional amino acid.

[0100] As used herein, the phrase “polynucleotide encodes / encoding RNA X” or a similar phrase refers to a polynucleotide that is transcribed to RNA Y comprising RNA X, meaning that RNA X is all or part of RNA Y. For example, polynucleotide encoding a gRNA can refer to (i) a polynucleotide that is transcribed to an RNA comprising the gRNA and one or more additional nucleotides, wherein the gRNA is part of the transcribed RNA; or alternatively, (ii) a polynucleotide that is transcribed to the gRNA per se without any additional nucleotide.

[0101] Overview

[0102] The engineered TadA variants used in the current cytosine base editors (CBEs) present distinctive advantages, including smaller sizes and fewer off-target effects compared to cytosine base editors that rely on natural deaminases. However, those engineered TadA variants demonstrate a preference for base editing in DNA with specific motif sequences and possess dual deaminase activity, acting on both cytosine and adenosine in adjacent positions, limiting their application scope. To address these issues, the inventors employed TadA orthologs screening and multi sequence alignment (MSA) -guided protein engineering techniques to create highly effective cytosine base editors (aTdCBE) without motif and adenosine deaminase activity limitations. Notably, as demonstrated, the delivery of aTdCBE to a humanized mouse model of Duchenne muscular dystrophy (DMD) mice achieved robust exon 55 skipping and restoration of dystrophin expression. The advancement of the disclosure in engineering TadA ortholog for cytosine editing enrich the base editing toolkits for gene-editing therapy and other potential applications.

[0103] In this study, through TadA orthologs screening and MSA-based protein engineering, the inventors obtained the cytosine base editor aTdCBE with robust cytosine editing activity in mammalian cells. The inventors found that by replacing adjacent amino acids of conserved residues with arginine, over 18%variants achieved enhanced activity with more than 1.5-fold. Compared to other strategies, protein engineering methods used in this study are much more convenient and have the potential to be applied to the engineering of other proteins. Compared with the previously reported cytosine base editors, aTdCBE has significant advantages in editing target sequences with AC motif. In addition, aTdCBE carrying no significant adenosine deaminase activity avoids generating unwanted editing products in applications. Considering its high efficiency with DNA base editing and high specificity at DNA level, aTdCBE may have more advantages in application compared to other cytosine base editors. aTdCBE has been employed in in vivo gene editing therapy studies for DMD in mice, and has demonstrated that aTdCBE is an efficient strategy to modulate exon skipping and restores dystrophin expression. Overall, aTdCBE offers a platform with highly efficient DNA base editing in mammalian cells, broadening applications in fundamental research, and has the potential to be applied in the field of gene editing therapy.

[0104] Representative Cytidine Deaminases

[0105] In an aspect, the disclosure provides a cytidine deaminase, wherein the cytidine deaminase is a mutant of a reference polypeptide of SEQ ID NO: 19, 1-9, and 73 and comprises an amino acid mutation (e.g., substitution) relative to (compared to) the reference polypeptide at a position selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, and / or 167 of the reference polypeptide.

[0106] In an aspect, the disclosure provides a cytidine deaminase, wherein the cytidine deaminase is a mutant of a reference polypeptide of SEQ ID NO: 19 and comprises an amino acid mutation (e.g., substitution) relative to (compared to) the reference polypeptide at a position selected from the group consisting of 8, 9, 19, 25, 41, 43, 45, 46, 47, 105, 107, 145, 146, and / or 154 of the reference polypeptide.

[0107] In some embodiments, the amino acid mutation leads to an increased deaminase activity, or wherein the cytidine deaminase has an increased deaminase activity compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.

[0108] In some embodiments, the amino acid mutation leads to an increased cytidine deaminase activity, or wherein the cytidine deaminase has an increased cytidine deaminase activity compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.

[0109] In some embodiments, the cytidine deaminase leads to increased cytosine base editing efficiency compared to the reference polypeptide, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.

[0110] In some embodiments, the cytidine deaminase comprises an amino acid mutation (e.g., substitution) at a position selected from the group consisting of 25, 46, 105, 107, 146, and / or 154 of the reference polypeptide.

[0111] In some embodiments, the amino acid substitution is a conservative amino acid substitution or a non-conservative amino acid substitution.

[0112] In some embodiments, the amino acid substitution is an amino acid substitution with an amino acid residue that is not the amino acid residue at the position of the reference polypeptide.

[0113] In some embodiments, the amino acid substitution is an amino acid substitution with

[0114] a) a non-polar amino acid residue (such as, Glycine (Gly / G) , Alanine (Ala / A) , Valine (Val / V) , Cysteine (Cys / C) , Proline (Pro / P) , Leucine (Leu / L) , Isoleucine (Ile / I) , Methionine (Met / M) , Tryptophan (Trp / W) , Phenylalanine (Phe / F) ,

[0115] b) a polar amino acid residue (such as, Serine (Ser / S) , Threonine (Thr / T) , Tyrosine (Tyr / Y) , Asparagine (Asn / N) , Glutamine (Gln / Q) ) ,

[0116] c) a positively charged amino acid residue (such as, Lysine (Lys / K) , Arginine (Arg / R) , Histidine (His / H) ) , or

[0117] d) a negatively charged amino acid residue (such as, Aspartic Acid (Asp / D) , Glutamic Acid (Glue / E) ) . In some embodiments, the cytidine deaminase comprises an amino acid substitution selected from the group consisting of E25A, P46G, T105V, E107N, D146R, E154V, and a combination thereof, wherein the position is numbered according to SEQ ID NO: 19.

[0118] In some embodiments, the cytidine deaminase comprises a combination substitution of E25A, P46G, T105V, E107N, D146R, and E154V, wherein the position is numbered according to SEQ ID NO: 19.

[0119] In some embodiments, the cytidine deaminase comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the reference polypeptide.

[0120] In some embodiments, the cytidine deaminase comprises, consists essentially of, or consists of the amino acid sequence of SEQ ID NO: 20 or 21, or a N-terminal truncation thereof without the N-terminal Met (e.g., SEQ ID NO: 51) .

[0121] Fusion protein

[0122] In another aspect, the disclosure provides a fusion protein comprising the cytidine deaminase of the disclosure fused to a heterogeneous functional domain.

[0123] In some embodiments, the heterogeneous functional domain comprises a polypeptide selected from the group consisting of a DNA binding domain, an RNA binding domain, an UGI domain, and a nuclear localization signal (NLS) .

[0124] In some embodiments, the DNA binding domain is a nucleic acid programmable DNA binding domain (napDNAbd) . In some embodiments, the RNA binding domain is a nucleic acid programmable RNA binding domain (napRNAbd) . In some embodiments, the RNA binding domain is a MS2 coat protein (MCP) .

[0125] In some embodiments, the fusion protein comprises a UGI domain.

[0126] In some embodiments, the fusion protein comprises a nucleic acid programmable DNA binding domain (napDNAbd) and two UGI domains.

[0127] In some embodiments, the fusion protein comprises, from N-to C-terminus, the cytidine deaminase, an optional linker, the napDNAbd, an optional linker, a UGI domains, an optional linker, and a UGI domain.

[0128] In some embodiments, the napDNAbd substantially lacks dsDNA cleavage activity.

[0129] In some embodiments, the napDNAbd substantially lacks dsDNA cleavage activity and nickase activity.

[0130] In some embodiments, the napDNAbd has nickase activity.

[0131] In some embodiments, the napDNAbd has nickase activity to nick the target strand.

[0132] In some embodiments, the napDNAbd comprises a Cas nickase or a dead Cas of a Cas protein.

[0133] In some embodiments, the Cas protein is selected from a group consisting of a Cas9 protein (such as, SpCas9, SaCas9, GeoCas9, CjCas9, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago) , SmacCas9, Spy-macCas9, xCas9, SpCas9-NG, ) ; a Cas12 protein (such as, Cas12a, AsCas12a, LbCas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f (Cas14) , Cas12g, Cas12h, Cas12i, xCas12i, Cas12Max, hfCas12Max, Cas12j, Cas12k, Cas12l, Cas12m, Cas12n, Cas12o, Cas12p, Cas12q, Cas12r, Cas12s, Cas12t, Cas12u, Cas12v, Cas12w, Cas12x, Cas12y, Cas12z) ; a Cas13 protein (such as, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, Cas13x, Cas13y) ; Csn2; and a mutant thereof.

[0134] In some embodiments, wherein the Cas nickase is a Cas9 nickase (nCas9) , such as SpCas9 nickase (SpCas9-D10A) .

[0135] In some embodiments, wherein the dead Cas is a dead Cas9 (dCas9) , such as dead SpCas9 (SpCas9-D10A+H840A) .

[0136] In some embodiments, wherein the Cas nickase is a Cas12i nickase (nCas12i) or dead Cas12i (dCas12i) , such as a deadCas12i of xCas12i polypeptide.

[0137] In some embodiments, the napDNAbd comprises a IscB nickase (nIscB) or a dead IscB (dIscB) of a IscB protein (e.g., OgeuIscB) .

[0138] In some embodiments, the napDNAbd comprise an amino acid sequence having a sequence identity of at least about 60%(e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 28 or 52.

[0139] In some embodiments, the UGI domain comprise an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 54.

[0140] In some embodiments, the napDNAbd comprises a TnpB nickase or a dead TnpB of a TnpB protein.

[0141] Base editor

[0142] In some embodiments, the fusion protein comprises an NLS at the N-terminal and / or C-terminal of the napDNAbp. In some embodiments, the fusion protein comprises an NLS at the N-terminal and / or C-terminal of the cytidine deaminase.

[0143] In some embodiments, the NLS is a SV40 NLS, a bpSV40 NLS (e.g., SEQ ID NO: 24 or 49) , or a NP NLS (Xenopus laevis Nucleoplasmin NLS, nucleoplasmin NLS) .

[0144] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 45 or 50.

[0145] System

[0146] In yet another aspect, the disclosure provides a system comprising:

[0147] (1) the fusion protein of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the fusion protein, and

[0148] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:

[0149] (i) a scaffold sequence capable of forming a complex with the fusion protein; and

[0150] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.

[0151] In some embodiments, the scaffold sequence is 5’ or 3’ to the guide sequence.

[0152] In some embodiments, the guide nucleic acid is a guide RNA (gRNA) .

[0153] In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of SEQ ID NO: 43.

[0154] In some embodiments, the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 43.

[0155] In some embodiments, the target sequence comprises about or at least about 14 contiguous nucleotides of the target DNA, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more contiguous nucleotides of the target DNA, or in a numerical range between any two of the preceding values, e.g., from about 14 to about 50, or from about 17 to about 22 contiguous nucleotides of the target DNA. In some embodiments, the target sequence comprises about 16 contiguous nucleotides of the target DNA.

[0156] In some embodiments, the target sequence is immediately 5’ or 3’ to a target adjacent motif (TAM) , or wherein the reversely complementary sequence of the target sequence (i.e., the protospacer sequence) is immediately 5’ or 3’ to a target adjacent motif (TAM) .

[0157] In some embodiments, the TAM is 5’-NGG’-3’ (e.g., for SpCas9) or 5’-AGA-3’ (e.g., for SpG Cas9) , wherein N is A, T, G, or C. In some embodiments, the TAM is 5’-NT’-3’ (e.g., for Cas12) , wherein N is A, T, G, or C. In some embodiments, the TAM is 5’-NNNNNN-3’, wherein N is A, T, G, or C. In some embodiments, the TAM is 5’-NNNGAN-3’, wherein N is A, T, G, or C.

[0158] In some embodiments, the guide sequence is about or at least about 14 nucleotides in length, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more nucleotides in length, or in a length of a numerical range between any two of the preceding values, e.g., in a length of from about 14 to about 50 nucleotides, or from about 17 to about 22 nucleotides. In some embodiments, the guide sequence is about 16 nucleotides in length.

[0159] In some embodiments, (1) the guide sequence is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% (fully) , optionally about 100% (fully) , reversely complementary to the target sequence; (2) the guide sequence contains no more than 5, 4, 3, 2, or 1 mismatch or contains no mismatch with the target sequence; or (3) the guide sequence comprises no mismatch with the target sequence in the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nucleotides at the 3’ end of the guide sequence.

[0160] In some embodiments, the system comprises two or more guide nuclei acids comprising two or more guide sequences capable of hybridizing to two or more target sequences of the same target DNA or different target DNAs, wherein the two or more guide sequences are the same or different, and wherein the two or more target sequences are the same or different.

[0161] In some embodiments, the target DNA is a target dsDNA, such as, a eukaryotic dsDNA, e.g., a gene in a eukaryotic cell.

[0162] In some embodiments, the target DNA is a target dsDNA, and wherein the target dsDNA comprises a protospacer sequence on a nontarget strand of the target dsDNA, wherein the dsDNA comprises a target deoxyribonucleotide (e.g., dA, dT, dC, dG) at a position of the protospacer sequence selected from the group consisting of position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, position 11, position 12, position 13, position 14, and a combination thereof; or wherein the target deoxyribonucleotide is at a position of the protospacer sequence between position 2 and position 12 or between position 3 and position 14, both inclusive.

[0163] In yet another aspect, the disclosure provides a polynucleotide encoding the cytidine deaminase or the fusion protein of the disclosure. In some embodiments, the polynucleotide encodes a guide nucleic acid of the disclosure.

[0164] In yet another aspect, the disclosure provides a delivery system comprising (1) the cytidine deaminase of the disclosure, the polynucleotide of the disclosure, or the system of the disclosure; and (2) a delivery vehicle.

[0165] In yet another aspect, the disclosure provides a vector comprising the polynucleotide of the disclosure. In some embodiments, the vector encodes a guide nucleic acid of the disclosure. In some embodiments, the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentivirus vector.

[0166] In yet another aspect, the disclosure provides a recombinant AAV (rAAV) particle comprising the rAAV vector of the disclosure. In some embodiments, the rAAV vector is an RNA.

[0167] In yet another aspect, the disclosure provides a ribonucleoprotein (RNP) comprising the cytidine deaminase of the disclosure and a guide nucleic acid.

[0168] In yet another aspect, the disclosure provides a lipid nanoparticle (LNP) comprising an RNA (e.g., mRNA) encoding the cytidine deaminase of the disclosure and a guide nucleic acid.

[0169] In yet another aspect, the disclosure provides a cell comprising the cytidine deaminase of the disclosure, the system of the disclosure, the polynucleotide of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the RNP of the disclosure, or the LNP of the disclosure.

[0170] In yet another aspect, the disclosure provides a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, or the lipid nanoparticle of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.

[0171] In some embodiments, the target DNA is in a cell.

[0172] In some embodiments, the cell is a eukaryotic cell (e.g., an animal cell, a vertebrate cell, a mammalian cell, a non-human mammalian cell, a non-human primate cell, a rodent (e.g., mouse or rat) cell, a human cell, a plant cell, or a yeast cell) or a prokaryotic cell (e.g., a bacteria cell) .

[0173] In some embodiments, the cell is from a plant or an animal.

[0174] In some embodiments, the plant is a dicotyledon. In some embodiments, the dicotyledon is selected from the group consisting of soybean, cabbage (e.g., Chinese cabbage) , rapeseed, brassica, watermelon, melon, potato, tomato, tobacco, eggplant, pepper, cucumber, cotton, alfalfa, eggplant, grape.

[0175] In some embodiments, the plant is a monocotyledon. In some embodiments, the monocotyledon is selected from the group consisting of rice, corn, wheat, barley, oat, sorghum, millet, grasses, Poaceae, Zizania, Avena, Coix, Hordeum, Oryza, Panicum (e.g., Panicum miliaceum) , Secale, Setaria (e.g., Setaria italica) , Sorghum, Triticum, Zea, Cymbopogon, Saccharum (e.g., Saccharum officinarum) , Phyllostachys, Dendrocalamus, Bambusa, Yushania.

[0176] In some embodiments, the animal is selected from the group consisting of pig, ox, sheep, goat, mouse, rat, alpaca, monkey, rabbit, chicken, duck, goose, fish (e.g., zebra fish) .

[0177] In yet another aspect, the disclosure provides a cell modified by the method of the disclosure.

[0178] In yet another aspect, the disclosure provides a pharmaceutical composition comprising (1) the system of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, or the cell of the disclosure; and (2) a pharmaceutically acceptable excipient.

[0179] In yet another aspect, the disclosure provides a method for diagnosing, preventing, or treating a disease in a subject in need thereof, comprising administering to the subject the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, the cell of the disclosure, or the pharmaceutical composition of the disclosure, wherein the disease is associated with a target DNA, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex, and wherein the modification of the target DNA diagnose, prevents, or treats the disease.

[0180] In some embodiments, the disease is selected from the group consisting of Angelman syndrome (AS) , Alzheimer's disease (AD) , transthyretin amyloidosis (ATTR) , transthyretin amyloid cardiomyopathy (ATTR-CM) , cystic fibrosis (CF) , hereditary angioedema, diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD) , Becker muscular dystrophy (BMD) , spinal muscular atrophy (SMA) , alpha-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington’s disease (HTT) , fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis (ALS) , frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA) , sickle cell disease, thalassemia (e.g., β-thalassemia) , Parkinson's disease (PD) , myelodysplastic syndrome (MDS) , retinitis pigmentosa (RP) , age-related macular degeneration (AMD) , Hepatitis B, nonalcoholic fatty liver disease (NAFLD) , Acquired Immune Deficiency Syndrome, corneal dystrophy (CD) , hypercholesterolemia, familial hypercholesterolemia (FH) , heart disease (e.g., hypertrophic cardiomyopathy (HCM) ) , and cancer.

[0181] In yet another aspect, the disclosure provides a guide nucleic acid comprising (i) a scaffold sequence capable of forming a complex with a nucleic acid programmable DNA binding protein; and (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.

[0182] In some embodiments, the guide sequence comprises or encodes the polynucleotide sequence of any one of SEQ ID NOs: 69-71 or a polynucleotide sequence difference from the polynucleotide sequence of any one of SEQ ID NOs: 69-71 by no more than 1, 2, 3, 4, or 5 nucleotides.

[0183] In some embodiments, the nucleic acid programmable DNA binding domain comprises a Cas9 protein.

[0184] In some embodiments, the scaffold sequence is 3’ to the guide sequence.

[0185] In some embodiments, the scaffold sequence has substantially the same secondary structure as that of the polynucleotide sequence of SEQ ID NO: 43 or comprises or encodes the polynucleotide sequence of SEQ ID NO: 43 or a polynucleotide sequence difference from the polynucleotide sequence of SEQ ID NO: 43 by no more than 1, 2, 3, 4, or 5 nucleotides.

[0186] In yet another aspect, the disclosure provides a polynucleotide encoding the guide nucleic acid of the disclosure. In yet another aspect, the disclosure provides a system comprising:

[0187] (1) a nucleic acid programmable DNA binding protein (napDNAbp) or a fusion protein comprising the napDNAbp, or a polynucleotide (e.g., a DNA, an RNA) encoding the napDNAbp or the fusion protein, and

[0188] (2) the guide nucleic acid of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid.

[0189] In yet another aspect, the disclosure provides a method of modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, whereby the target DNA is modified.

[0190] In some embodiments, the target DNA is DMD gene.

[0191] In yet another aspect, the disclosure provides a cell comprising a target DNA modified by the method of the disclosure.

[0192] In yet another aspect, the disclosure provides an animal model comprising the cell of the disclosure.

[0193] In some embodiments, the animal model is a mouse model or a non-human primate model.

[0194] Further embodiments are illustrated in the following Examples which are given for illustrative purposes only and are not intended to limit the scope of the disclosure.

[0195] EXAMPLES

[0196] The following examples are provided to further illustrate some embodiments of the disclosure but are not intended to limit the scope of the invention; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.

[0197] Methods and Materials

[0198] Study approval

[0199] Exclusively male mice were utilized for all experiments, including grip strength tests, creatine kinase (CK) analysis, and AAV injections. All animal experiments were performed and approved by the Institutional Animal Care and Use Committee (IACUC) of HuidaGene Therapeutics Co., Ltd., Shanghai, China and Lingang Laboratory, Shanghai, China.

[0200] Computational analysis of TadA orthologs

[0201] Firstly, the inventors downloaded 15, 167 TadA protein sequences from NCBI database. The inventors further used BLASTP (v2.2.21) to remove redundant proteins with identity over than 90% (Altschul, S. F., et al., Basic local alignment search tool. J. Mol. Biol. 215, 403-410, doi: 10.1016 / S0022-2836 (05) 80360-2 (1990) . ) . Then, the inventors performed multiple sequence alignment using MAFFT (v7.429) (Nakamura, T., et al., Parallelization of MAFFT for large-scale multiple sequence alignments. Bioinformatics. 34, 2490-2492, doi: 10.1093 / bioinformatics / bty121 (2018) . ) . MEGA11 were used to construct phylogenetic tree (Tamura, K., et al., MEGA11: molecular evolutionary genetics analysis version 11. Mol. Biol. Evol. 38, 3022-3027, doi: 10.1093 / molbev / msab120 (2021) . ) . "Calculate_AHC. pl" were used to identify highly conservative residues and AHC residues. MCL (v14-137) (Enright, A. J., et al., An efficient algorithm for large-scale detection of protein families. Nucleic Acids Res. 30, 1575-1584, doi: 10.1093 / nar / 30.7.1575 (2002) . ) was used to cluster the redundant proteins with identity over than 70%, and nine TadA orthologs were randomly selected from different clusters for experimental screening.

[0202] Plasmid constructions

[0203] Human Codon-optimized orthologous TadAs were synthesized commercially (GenScript Co., Ltd) and cloned to generate pT7_NLS-TadA-Cas9-NLS_pA_pCBH_mCherry_pA plasmid by NEBuilder (New England Biolabs) . The dual-AAV delivery system was designed to express two separate fragments of a base editor, which are subsequently spliced into the full-length protein via the Rhodothermus marinus (Rma) intein, as reported in a previous study (Levy, J. M., et al., Cytosine and adenine base editing of the brain, liver, retina, heart and skeletal muscle of mice via adeno-associated viruses. Nat Biomed Eng. 4, 97-110, doi: 10.1038 / s41551-019-0501-5 (2020) ; Jin, M., et al., Correction of human nonsense mutation via adenine base editing for Duchenne muscular dystrophy treatment in mouse. Mol Ther Nucleic Acids. 35, 102165, doi: 10.1016 / j. omtn. 2024.102165 (2024) ) . The intein sequences, specifically the N-and C-terminal segments, were synthesized by Genewiz (Suzhou, China) . These segments were then integrated into the 573 and 574 amino acid residues of the aTdCBE backbones using Gibson cloning of PCR-amplified inserts.

[0204] Mammalian cell culture, transfection, and flow cytometry analysis

[0205] HEK293T cells were cultured in Dulbecco’s Modified Eagle’s Medium (Gibco, 11965-092) supplemented with 10%fetal bovine serum (Gibco, 10099-141C) , and 1%Pen-Strep-Glutamine (100×) (Gibco, 10378-016) at 37℃ with 5%CO2 in a cell incubator. For TadA variants screening, HEK293T cells cultured in 24-well plates were co-transfected with 1.0 μg of tagBFP-*EGFP reporter plasmid and TadA-mCherry plasmid in a molar ratio of 1: 1 with Polyetherimide (PEI) . After 48 hours, mCherry, BFP and EGFP fluorescence were analyzed by Beckman CytoFlex flow-cytometer. To evaluate genome editing in endogenous sites, cells were harvested at 48 hours after transfection and sorted by BD FACS Aria III flow cytometer. FACS data were analyzed with FlowJo X (v10.0.7) .

[0206] Detection of gene editing frequency.

[0207] 20 μL of lysis buffer with proteinase K (Vazyme Biotech) were used to lysis about ten thousand sorted cells following the manufacturer’s manual. Targeted amplifications were produced by Phanta Max Super-Fidelity DNA Polymerase (Vazyme Biotech) . For targeted amplicon sequencing, PCR reactions were performed using primers with different barcodes. The DNA products were purified with Gel extraction kit (Omega) and analyzed by 150-bp paired-end reads Illumina NovaSeq 6000 platform (Genewiz Co. Ltd. ) . The deep sequencing data were first de-multiplexed by Cutadapt (v. 2.8) based on sample barcodes. The de-multiplexed reads were then processed by CRISPResso2 for the quantification of editing efficiency, including indels, A-to-G or C-to-T conversions at each target site (Clement, K., et al., CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 224-226, doi: 10.1038 / s41587-019-0032-3 (2019) . ) .

[0208] High-throughput library experiments

[0209] The cell line previously constructed using 11, 868 pairs of sgRNA lentivirus plasmid library was used to detect motif preference in the base editor (Yuan, T., et al., Deep learning models incorporating endogenous factors beyond DNA sequences improve the prediction accuracy of base editing outcomes. Cell Discov. 10, 20, doi: 10.1038 / s41421-023-00624-1 (2024) ) . For each 10-cm dish, 35 μg plasmids that encode CBEs and mCherry were transfected using PEI. After 48 hours, transfected cells were harvested using FACS followed by genomic DNA extraction. The PCR products were sequenced using a 150-bp paired-end Illumina NovaSeq 6000 platform (Genewiz Co. Ltd. ) . High-throughput sequencing datasets were processed using CRISPResso2 to calculate editing efficiency of each target. The target sites were excluded with a coverage depth of less than 100 in each sample. Cytosines in positions 4-7 of the target sequences were used to statistically analyze motif preferences.

[0210] Off-target analysis with in-silico prediction

[0211] To evaluate the specificity of TadA base editors, the Cas-OFFinder was employed to predict the potential off-target sites (Bae, S., et al., Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics. 30, 1473-1475, doi: 10.1093 / bioinformatics / btu048 (2014) . ) . Search queries covered both Cas9 spacer sequence and PAM of the on-target site. The PAM of research was set to “NGG” and the mismatches were set to less than 5. All other parameters were left as default. The potential off-target sites were amplified and deep sequenced for analysis.

[0212] Orthogonal R-loop assay.

[0213] Orthogonal R-loop assay was performed to detect the nuclease-independent off-target editing as described previously (Doman, J. L., et al., Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nat. Biotechnol. 38, 620-628, doi: 10.1038 / s41587-020-0414-6 (2020) . ) . 1.5 μg plasmids that encode aTdCBE / TadCBEd and an on-target sgRNA for aTdCBE / TadCBEd, along with plasmids expressing dSaCas9 and a SaCas9 sgRNA that targets the genome locus previously reported were co-transfected using PEI. After 48 hours, transfected cells were harvested using FACS followed by genomic DNA extraction with 20 μL of freshly prepared lysis buffer (Vazyme) with proteinase K added. The targeted loci by dSaCas9 were amplified and deep sequenced.

[0214] Generation of humanized DMDΔE54 mdxmice

[0215] Mice were housed in a barrier facility with a 12-hour light / dark cycle and maintained in compliance with the guidelines outlined in the Instructive Notions with Respect to Caring for Laboratory Animals issued by the Ministry of Science and Technology of China. To generate the humanized DMDΔE54 mice, the inventors employed the CRISPR-Cas9 system on the embryos obtained from mating STOCK Tg (DMD) 72Thoen / J male and female mice (#018900) . Specifically, two sgRNAs targeting the flanking introns of human DMD exon54 on Chr. 5 were designed. The guide sequences of these sgRNAs are g1: TTTCTGCAAGTGCAGAGAGG (SEQ ID NO: 69) and g2: GGTGTGTGGAGTGAGATACT (SEQ ID NO: 70) . Each sgRNA template was appended with the T7 promoter sequence (TAATACGACTCACTATAG (SEQ ID NO: 72) ) for efficient transcription. The PCR product was then purified directly using the Omega gel extraction kit (Omega, D2500-02) , and the templates were used for in vitro transcription with the MEGAshortscript T7 Kit (Invitrogen, AM1354) . The sgRNAs were purified using a MEGAclear Kit (Invitrogen, AM1908) and eluted with nuclease-free water. The concentration of target sgRNA was measured using a NanoDrop instrument. For cytoplasmic injection, spCas9 mRNA (100 ng / μl) , sgRNA-L (50 ng / μl) and sgRNA-R (50 ng / μl) were mixed and then injected into fertilized eggs using a FemtoJet microinjector (Eppendorf) with constant flow settings. The injected zygotes were cultured in KSOM medium for 12 hours and surgically transferred to the oviduct of recipient mice 24 hours after estrus was observed. Genomic DNA from the tail tissue of founder (F0) mice was isolated according to manufacturer’s instructions for the OMEGA Kit (Omega, D3396-02) for PCR, followed by gel electrophoresis.

[0216] AAV9 production and delivery to DMDΔE54 mdxmice

[0217] AAVs used in this study were produced by HuidaGene Therapeutics Co., Ltd. The transfection process involved achieving a confluency of 70–90%, after which the media was replaced with fresh pre-warmed growth media prior to transfection. For each 15-cm dish, a mixture of 20 μg of pHelper, 10 μg of pRepCap, and 10 μg of GOI plasmid was transferred dropwise to the cell media. Following a three-day incubation, the AAVs were purified using iodixanol density gradient centrifugation. The DMDΔE54 mdx mice were derived by mating the humanized DMDΔE54 mice with mdx mice carrying stop mutation in mouse exon 23 on Chr. X. In the case of intramuscular injection, 3-week-old DMDΔE54 mdx mice were anesthetized, and their tibialis anterior (TA) muscle was injected with 50 μL of AAV9 (2.5 × 1011 vg per virus) preparations or with an equivalent volume of saline solution. Tissues were divided into distinct segments for targeted assessment. Specifically, the distal region was allocated for evaluating DNA editing and exon skipping efficiency, the middle portion was dedicated to Western blot analysis of dystrophin expression, and the proximal segment was reserved for immunofluorescent analysis of dystrophin levels at six weeks after treatment.

[0218] Western blot analysis

[0219] The samples were homogenized using RIPA buffer supplemented with protease inhibitor cocktail. The lysate supernatants were quantified using a Pierce BCA protein assay kit (Thermo Fisher Scientific, 23225) and adjusted to an identical concentration using H2O. Equal amounts of the sample were mixed with NuPAGE LDS sample buffer (Invitrogen, NP0007) and 10%β-mercaptoethanol, then boiled at 70 ℃ for 10 min. Ten μg of total protein per lane was loaded into 3%to 8%tris-acetate gels (Invitrogen, EA03752BOX) and electrophoresed for 1 hour at 200 V. Protein was transferred onto a PVDF membrane under wet conditions at 350 mA for 3.5 hours. Subsequently, the membrane was blocked in 5%non-fat milk in TBST buffer and then incubated with primary antibody to label the specific protein. After washing three times with TBST, the membrane was incubated with an HRP-conjugated secondary antibody specific to the IgG of the species of primary antibody against dystrophin (Sigma, D8168) or vinculin (CST, 13901S) . Finally, the target proteins were visualized using Chemiluminescent substrates (Invitrogen, WP20005) .

[0220] Histology and Immunofluorescence

[0221] Tissue samples were collected and immersed into preconditioned 4%paraformaldehyde. The fixed tissues underwent dehydration through a series of alcohol concentrations, followed by treatment with xylene and embedding in melted paraffin wax. Subsequently, the paraffin-embedded tissues were deparaffinized using xylene, followed by a series of alcohol washes ranging from high to low concentrations, and finally placed in distilled water. For hematoxylin and eosin (H&E) staining, the slides were stained with hematoxylin for 3-8 minutes, followed by color separation using acid water and ammonia water. After dehydration using 70%and 90%alcohol for 10 minutes each, the tissues were stained in eosin staining solution for 1-3 minutes, and dehydrated in ascending alcohol solutions (50%, 70%, 80%, 95%, 100%) . Coverslips were then mounted onto the glass slides with neutral resin.

[0222] For Sirius red staining, the slides were stained with picrosirius red for one hour, washed in two changes of acidified water. Physical removal of most of the water from the slides was accomplished by vigorous shaking. Then, slides were dehydrated in three changes of 100%ethanol, cleared in xylene, and finally mounted in neutral resin.

[0223] For immunofluorescence, the tissues were embedded in optimal cutting temperature (OCT) compound and snap-frozen in liquid nitrogen. Serial frozen cryosections (10 μm) were fixed for two hours at 37 ℃ followed by permeabilization with PBS + 0.4%Triton-X for 30 min. After washing with PBS, the samples were blocked with 10%goat serum for 1 hour at room temperature. Next, the slides were incubated overnight at 4 ℃ with primary antibodies against dystrophin (Abcam, ab15277) and spectrin (Millipore, MAB1622) . The next day, samples were extensively washed with PBS and incubated with compatible secondary antibodies (Alexa 488 AffiniPure donkey anti-rabbit IgG (Jackson ImmunoResearch labs, 711-545-152) or Alexa Fluor 647 AffiniPure donkey anti-mouse IgG (Jackson ImmunoResearch labs, 715-605-151) ) and DAPI for 3 hours at room temperature. Following a 15-minute wash with PBS, the slides were sealed with fluoromount-G mounting medium. All images were visualized using Nikon C2. The number of Dys+ muscle fibers is represented as a percentage of total spectrin-positive muscle fibers.

[0224] RNA-seq for off-target analysis

[0225] To quantify the transcriptome deaminases off-target edits, HEK293T cells were cultured in 10-cm dishes with 80%confluence and transfected with 35 μg plasmids containing base editors and gRNA. After 48 hours, about 600, 000 transfected cells were sorted by FACS, and RNA was extracted using Trizol (Ambion) for RNA-seq library preparation. An RNA-seq library was generated with a TruSeq Stranded Total RNA library preparation kit according to the standard protocol. The transcriptome libraries were sequenced using a 150-bp paired-end Illumina NovaSeq 6000 platform (Genewiz Co. Ltd. ) .

[0226] The calculation analysis referred to previously published methods (Richter, M. F., et al., Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat Biotechnol. 38, 883-891, doi: 10.1038 / s41587-020-0453-z (2020) ) . Trimmomatic (v. 0.39-2) were using to filter the RNAseq raw data (Bolger, A.M., et al., Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 30, 2114-2120, doi: 10.1093 / bioinformatics / btu170 (2014) ) . The clean reads were aligned to the hg38 reference genome with Hisat2 (v.2.2.1) (Kim, D., et al., Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat. Biotechnol. 37, 907-915, doi: 10.1038 / s41587-019-0201-4 (2019) ) . RNA editing sites were calculated using REDItools (v1.2.1) with “-e -d -p -u -m 60 -T 5-5 -W -n 0.0” parameters (Flati, T., et al., HPC-REDItools: a novel HPC-aware tool for improved large scale RNA-editing analysis. BMC. Bioinformatics. 21, 353, doi: 10.1186 / s12859-020-03562-x (2020) ) . The edited adenosines divided by total adenosines and the edited cytosines divided by total cytosines were calculated separately.

[0227] Statistics &Reproducibility

[0228] All cell experimental results are presented as mean ± s. d, while all animal experimental results are presented as mean ± s. e. m. The one-sided Mann-Whitney U test was utilized for comparisons. The number of independent biological replicates are shown in figure legends. No data were excluded from the analyses. We randomly selected cells for test group and control group. DMD mice used for gene editing therapy were allocated to control or AAV9 treated group randomly. Differences in means were considered statistically significant when they reached P < 0.05. Significance levels are: *P < 0.05. **P < 0.01.

[0229] Example 1. Screening of TadA orthologs using fluorescent reporter system.

[0230] To sensitively detect DNA base editing activity in mammalian cells, the inventors designed a fluorescent reporter system, termed as BFP-*EGFP, which contains BFP and inactive EGFP carrying stop codon TAG that can be corrected through C-to-G editing (FIG. 1a) . Due to the weakest base editing preference for ACT motif by cytosine base editor based on the ecTadA variant (Neugebauer, M. E., et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity. Nat. Biotechnol. 41, 673-685, doi: 10.1038 / s41587-022-01533-6 (2023) . ) , the inventors introduced ACT motif in the target sequence to screen for TadA orthologs and variants thereof that can efficiently edit AC motif sequences. When the target sequence is edited by cytosine deaminase of CGBE, the stop codon TAG within the mutant *EGFP coding sequence will be corrected into TAC (codon for Tryosine  / Try  / Y) to restore the translation of EGFP protein.

[0231] The BFP-*EGFP fluorescent reporter system was composed of an expression plasmid and a reporter plasmid.

[0232] The expression plasmid (see the upper construct in FIG. 1) comprised, from 5’ to 3’, (1) a CBE coding polynucleotide comprising start codon ATG, a polynucleotide (SEQ ID NO: 23) encoding a N-terminal bpNLS (SEQ ID NO: 24) , a polynucleotide (e.g., one of SEQ ID NOs: 11-18) encoding a representative cytidine deaminase (e.g., one of SEQ ID NOs: 2-9 and 73) , a polynucleotide (SEQ ID NO: 25) encoding a XTEN&GS linker (SEQ ID NO: 26) , a polynucleotide (SEQ ID NO: 27) encoding SpG Cas9-D10A nickase (SEQ ID NO: 28) , a GS linker (SEQ ID NO: 46; encoding SEQ ID NO: 47) , and a polynucleotide (SEQ ID NO: 48) encoding a C-terminal bpNLS (SEQ ID NO: 49) , operably linked to CAG promoter and followed by SV40 polyA signal coding sequence (SEQ ID NO: 29) ; and (2) a polynucleotide (SEQ ID NO: 31) encoding mCherry operably linked to CMV promoter (SEQ ID NO: 33) and followed by bGH polyA signal coding sequence (SEQ ID NO: 32) . Red fluorescent signals generated by the expression of the mCherry indicated successful transfection and expression of the expression plasmid in host cells.

[0233] The reporter plasmid (see the lower construct in FIG. 1) comprised, from 5’ to 3’, (1) a polynucleotide encoding sgRNA (SEQ ID NO: 44) composed of insertion sequence-targeting guide sequence (SEQ ID NO: 42) 5’ to scaffold sequence (SEQ ID NO: 43) operably linked to U6 promoter (SEQ ID NO: 41) , and (2) a polynucleotide encoding deactivated BFP-P2A-EGFP expression cassette operably linked to CMV promoter (SEQ ID NO: 33) and followed by bGH polyA signal coding sequence (SEQ ID NO: 40) . Blue fluorescent signals generated by the expression of the BFP indicated successful transfection and expression of the reporter plasmid in host cells.

[0234] The polynucleotide encoding deactivated BFP-P2A-EGFP expression cassette comprised, from 5’ to 3’, TagBFP coding sequence (SEQ ID NO: 34) , P2A coding sequence (SEQ ID NO: 35) , insertion sequence (SEQ ID NO: 36) containing premature stop codon TAG, and EGFP coding sequence (SEQ ID NO: 39) . The premature stop codon prevented the translation of the EGFP coding sequence.

[0235] The insertion sequence was on the sense strand of the reporter plasmid, and the reverse complement (SEQ ID NO: 37) of the insertion sequence on the antisense strand of the reporter plasmid was intended to be as an edited strand (nontarget strand) containing target deoxynucleotide dC-bearing protospacer sequence (SEQ ID NO: 38) (corresponding to the guide sequence (SEQ ID NO: 42) ) 5’ to PAM  (suitable PAM for SpG Cas9) . With successful C-to-G base editing on the antisense strand, the stop codon TAG on the sense strand would be converted to non-stop codon TAC as illustrated in FIG. 1, eliminating the stop codon-induced prevention of EGFP translation and hence initiating emission of green fluorescent signals. SEQ ID NO: 36, insertion sequence,  is the premature stop codon.

[0236] Complement of insertion sequence (SEQ ID NO: 36)

[0237] SEQ ID NO: 37, reverse complement of insertion sequence (SEQ ID NO: 36) ,  is PAM,  is the reverse complement of the premature stop codon, and the therein is the target deoxynucleotide dC for conversion to G by C-to-G base editing, thereby converting the premature stop codon on the sense strand to non-stop codon TAC (Tyr, Y) .

[0238] SEQ ID NO: 38, protospacer sequence, 20 nt

[0239] SEQ ID NO: 42, guide sequence, 20 nt

[0240] SEQ ID NO: 43, scaffold sequence, 76 nt

[0241] SEQ ID NO: 44, sgRNA, 96 nt

[0242] Based on different sequence homology, the inventors synthesized nine representative TadA orthologs (SEQ ID NOs: 2-9 and 73, respectively) (FIG. 1b) , and constructed base editors (SEQ ID NOs: 56-64, respectively) together with SpG Cas9-D10A nickase (SEQ ID NO: 28) and NLSs (SEQ ID NOs: 24 and 49) . As an example, AjTadA. v1 (SEQ ID NO: 3) contains four substitutions T105V, E107N, D146R, and E154V relative to reference (wild type; wt) AjTadA (SEQ ID NO: 19) . A negative control ( “NC” ) was prepared by replacing the reporter plasmid with an otherwise identical reporter plasmid without the insertion sequence-targeting guide sequence (SEQ ID NO: 42) . Td-CGBE (SEQ ID NO: 55) was used as a positive control to also edit the BFP-*EGFP reporter by co-transfection of Td-CGBE-mCherry and reporter plasmids targeted by the *EGFP-specific sgRNA in HEK293T cells. The cytosine base editing efficiency of the tested CBE was calculated as the percentage of EGFP positive cells ( “EGFP+” ) in BFP &mCherry dual-positive cells ( “mCherry+ BFP+” ) . The higher the %EGFP+  / mCherry+ BFP+ is, the higher the cytosine base editing efficiency would be.

[0243] The results are plotted in FIG. 1c. The TadA names in the x axis refer to the corresponding CBEs. Two days after transfection, the inventors observed that about 5%of successfully transfected cells had EGFP activated by Td-CGBE. Notably, it was found that the CBE (SEQ ID NO; 57) with AjTadA. v1 (SEQ ID NO: 3) engineered from wt AjTadA (SEQ ID NO: 19) of Acinetobacter junii showed the highest editing activity among the CBEs and especially higher than Td-CGBE (FIG. 1c) . In another independent test, the CBE with AjTadA. v1 (SEQ ID NO: 3) was demonstrated to exhibit about 2-fold cytosine base editing efficiency (3.005%) compared to the CBE with the reference AjTadA (SEQ ID NO: 19) (1.525%) . The inventors further used AjTadA. v1 to target PIK3CA loci and found that AjTadA. v1 showed cytosine deamination activity (FIG. 5) .

[0244] Example 2. Engineering of AjTadA based on multiple sequence alignment.

[0245] Multiple sequence alignment (MSAs) is commonly used for phylogenetic inference, protein structure prediction, and functional effects of mutations prediction. Through multiple sequence alignment, the inventors identified 20 highly conserved amino acid sites with over 99%conservation in the TadA family (FIG. 6) . Highly conserved residues are under greater selection pressure throughout the evolution process to be indispensable for normal deaminase activity of TadA and survival advantage of the organism. The inventors assume that adjacent amino acids of highly conserved sites (AHC) can assist highly conserved amino acids in performing their functions, and replacing these AHC amino acids may potentially enhance or decrease their activity. The inventors performed arginine scanning mutagenesis and individually substituted all non-positively charged AHC residues with arginine (FIG. 2a, FIG. 6) . The inventors divided AHC amino acids with distances of 1, 2, 3, and 4 from highly conserved amino acids into four groups: D1, D2, D3, and D4. In total, 65 arginine single substitutions of AjTadA. v1 (SEQ ID NO: 3) were screened using BFP-*EGFP reporter assay. Based on the threshold of over 1.5-fold for the EGFP signal increased by AjTadA variants relative to the AjTadA. v1 control, the inventors identified 12 variants (W9R, Y43R, and A45R in D1 group; Y8R, A19R, N25R, E46R, D146R (relative to SEQ ID NO: 19; or Y146R relative to SEQ ID NO: 3) , and E154R in D2 group; S41R, I47R, and S145R in D3 group) with enhanced base editing activity (FIG. 2b) . Remarkably, 40% (6 / 15) of D2 variants showed enhanced activity. In particular, E25R and P46R variants in D2 variants showed over 4-fold improvement relative to AjTadA. v1 (FIG. 2b) .

[0246] To test whether replacing the two sites E25 and P46 with other amino acids can further enhance activity, the inventors selected six different amino acids and found that the P46G variant performed best, producing over 60%of EGFP positive cells (FIG. 2c) . The inventors then combined P46G with E25R, E25W, E25F, E25Y, E25G, E25A, or E25D in AjTadA. v1 for reporter assay and found that the variant carrying P46G+E25A exhibited the highest editing activity, named AjTadA. v2 (SEQ ID NO: 20) (FIG. 2d) . Considering the upper detection limit of the fluorescent reporter system, the inventors next selected three different endogenous gene loci to further evaluate the editing activity of the variants. On the basis of AjTadA. v2, the inventors further introduced amino acid substitutions in different regions (Y8R, Y43R, Y146R) that performed well in the first round of screening. The inventors found that the CBE (SEQ ID NO: 45) with the AjTadA variant carrying P46G+E25A+Y146R (SEQ ID NO: 21; also known as “AjTadA. v3” ) showed relatively high editing activity at all three endogenous sites (FIG. 2e) . Apart from the above tested CBEs containing SpG Cas9-D10A, the inventors further fused AjTadA. v3 (SEQ ID NO: 51; without N-terminal Met compared to SEQ ID NO: 21) with SpCas9-D10A (SEQ ID NO: 52) and two UGI domains (SEQ ID NO: 54) and NLS, and obtained a cytosine editor named as “aTdCBE” (SEQ ID NO: 50) .

[0247] Example 3. aTdCBE enables robust genomic editing in mammalian cells with broad target scope and high specificity

[0248] In order to compare the editing efficiency of aTdCBE according to the disclosure with other CBEs, e.g., Td-CBEmax (SEQ ID NO: 65) , TadCBEd (SEQ ID NO: 66) , B3PCY2-CBE (SEQ ID NO: 67) , hA3AW104A (hA3A*) -CBE, and YE1-BE4max (SEQ ID NO: 68) were selected due to their diverse properties, including differences in editing efficiency, off-target effects, substrate preferences, and applicability across various cell types and model organisms. By including these specific CBEs in this study, the inventors aimed to provide a comprehensive analysis of the current state-of-the-art cytidine base editing tools and facilitate a comparative assessment of their performance under standardized experimental conditions. The inventors selected 25 different genomic loci for activity assessment and found that aTdCBE and TadCBEd showed comparable cytosine editing efficiency (FIG. 3a, FIG. 7) . It is noteworthy that aTdCBE exhibited no significant adenosine deaminase activity. In contrast, TadCBEd had detectable adenosine deaminase activity at multiple sites (FIG. 8) . At almost all loci, aTdCBE and TadCBEd performed better than Td-CBEmax, B3PCY2-CBE, hA3A*-CBE, and YE1-BE4max.

[0249] Moreover, the non-evolved TadA deaminase B3PCY2 only produces effective editing for very few targets. Except for hA3A*-CBE, the aTdCBE and TadCBEd exhibit a broad cytosine editing window than other cytosine base editors, approximately between target position 4 and 7 (FIG. 3b) . In addition, TadCBEd exhibits weak editing activity towards adenosine in the fifth to seventh positions of the target sequence (FIG. 3c) .

[0250] Previous studies have shown that CBE editors YE1-BE4max and TadCBEd based on Apopec and ecTadA have weaker editing abilities for GC and AC motif in library experiments, respectively. Consistent with previous reports, the average editing efficiency of Apobec and ecTadA in GC and AC context of endogenous loci is indeed weaker than in other motifs, respectively (FIG. 3d) . Another ecTadA-derived cytosine base editor Td-CBEmax exhibits weak editing efficiency in both AC and GC. To verify the motif preference of the TadA deaminases, the inventors performed high-throughput library experiments using a library containing 11, 868 paired sgRNA. Consistent with previous results, the aTdCBE has better editing efficiency for AC sites (FIG. 3e) .

[0251] The inventors then tested the effects of aTdCBE, TadCBEd, and recently reported CBE6b on introducing the PCSK9 termination codon, using four sites containing the CAG codon (FIG. 3f) . The inventors found that TadCBEd produced A-to-G editing at all four sites, while aTdCBE and CBE6b hardly caused A-to-G editing. aTdCBE performs better at the AC motif site than aTdCBE and CBE6b. aTdCBE, TadCBEd, and CBE6b can generate C-to-T editing efficiencies of up to approximately 72%, 57%, and 60%for these CAG codon sites, respectively.

[0252] The inventors further evaluated the specificity of aTdCBE, Td-CBEmax, TadCBEd, B3PCY2-CBE, and YE1-BE4max in HEK293T cells at two target sites. After deep-sequencing analysis, it was found that all the base editors displayed low level of off-target effects at the predicted off-target sites by Cas-OFFinder (FIG. 9a) . To detect nuclease-independent off-target effect of aTdCBE and TadCBEd, the inventors performed orthogonal R-loop assay. Using five previously reported SaCas9 target sites, it was found that aTdCBE exhibited comparably low nuclease-independent off-target events with TadCBEd (FIG. 9b) .

[0253] TadA8e was previously reported to have significant RNA off-target editing, and the introduction of V106W mutations resulted in significant reduction of RNA editing with slightly decreased on-target editing activity. To investigate whether TadA variants with cytosine deaminase activity can cause transcriptome-wide off-target effects, the inventors conducted RNA-seq to evaluate RNA off-target effect for aTdCBE and TadCBEd. Compared to the HEK293T control without the base editor transfection, both aTdCBE and TadCBEd showed very low off-target effects on cytosine in RNA (FIG. 9c) .

[0254] Since the engineered AjTadA was obtained through C-to-G fluorescence reporter screening, to investigate whether the edited product of engineered AjTadA exhibits C-to-G preference compared to other deaminase enzymes, the inventors removed the UGIs of aTdCBE, TadCBEd, Td-CBEmax, B3PCY2-CBE, hA3A*-CBE, and CBE6, and generated 6 CGBEs. Compared to the effect of sites and sequences on editing product preference, the effect of deaminase variants on C-to-G preference is relatively low (FIG. 10) .

[0255] Target sequences for target loci and off-targets above are set forth in SEQ ID NOs: 82-127.

[0256] Example 4. Local muscle administration of cytosine base editor aTdCBE restores dystrophin expression

[0257] DMD is a fatal X-linked recessive disease affecting 1 out of 3500-5000 newborn males resulting from thousands of pathogenic mutations in the human X chromosome-linked DMD gene. While there are thousands of documented clinical mutations, most DMD-causing mutations occur in a “hotspot” region encompassing exons 45 to 55 of the DMD gene, and skipping of exon 55 can provide therapeutic benefits to approximately 2%of DMD patients. Engineering TadA with broad targeting scope and stringent cytosine deaminase activity may be a potential strategy for gene editing correction of DMD.

[0258] To evaluate the in vivo activity of cytosine base editors, the inventors firstly compared the cytosine base editors in HEK293T cells and found that only aTdCBE and TadCBEd exhibited high activity at exon 55 splice acceptor site (FIG. 11) . The inventors further generate a genetically humanized DMD mouse model (DMDΔE54 mdx mice) (FIG. 4a, FIG. 12a) . Sequencing of the RT-PCR product confirmed proper splicing between human DMD exon53 and exon55 in the DMDΔE54 mdx mice (FIG. 12b) . Immunostaining and western blot revealed a complete loss of dystrophin expression in DMDΔE54 mdx mice (FIG. 12c and 12d) . Additionally, muscular histology, creatine kinase (CK) activity, and motor function also suggested that DMDΔE54 mdx mice presented severe DMD symptoms (FIG. 12e, 12f and 12g) .

[0259] Given the vector size beyond the genome packaging capacity of AAV, the inventors have developed a strategy where each fragment of the base editor is expressed individually by two separate AAV vectors. The AAV coding sequence for AjTadA-Cas9n-N is set forth in SEQ ID NO: 128, and the AAV coding sequence for Cas9n-C is set forth in SEQ ID NO: 129. The Rhodothermus marinus (Rma) intein facilitated the precise autocatalytic splicing of the two fragments, thereby reconstructing the full-length, active form of aTdCBE within the target cells. Then, the inventors performed intramuscular (IM) injection of dual-AAV9 particles in tibialis anterior (TA) muscle of 3-week-old DMDΔE54 mdx mice. The DMD-targeting guide sequence of the sgRNA used with aTdCBE is TCACCCTGCAAAGGACCAAA (SEQ ID NO: 71) . Six weeks after IM injection, TA muscle samples were collected for subsequent analysis (FIG. 4a) . Deletion of exon 54 in the DMD gene resulted in the introduction of a downstream premature stop codon in exon 55, resulting in the production of a nonfunctional truncated dystrophin protein. The DMD open reading frame (ORF) can be restored by editing exon 55 SAS, allowing for the splicing of exons 53 to 56 in the case of exon 55 skipping (FIG. 4b) . The findings herein revealed that aTdCBE exhibited over 40%DNA base editing efficiency (FIG. 4c) . The RT-PCR results indicated successful splicing alteration to skip human DMD exon 55 following aTdCBE-induced base conversion, as confirmed by gel analysis (FIG. 4d) . Furthermore, immunostaining results demonstrated a remarkable rescue of dystrophin expression following local injection of aTdCBE (FIG. 4e, FIG. 13) . The percentage of dystrophin-positive fibers after aTdCBE treatment reached up to 99%of the wildtype level (FIG. 4f) . To quantify the level of dystrophin restoration, the inventors conducted western blot analysis, revealing that aTdCBE system restored 60%of dystrophin expression (FIG. 4g) . Together, these results indicate that the cytosine base editor aTdCBE, as a highly effective base editing tool with broad targeting scope, provides a promising approach for basic research and therapeutic applications.

[0260] ***

[0261] Various modifications and variations of the described products, methods, and uses of the disclosure will be apparent to those skilled in the art without departing from the scope and spirit of the disclosure. Although the disclosure has been described in connection with specific embodiments, it will be understood that it is capable of further modifications and that the disclosure as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the disclosure that are obvious to those skilled in the art are intended to be within the scope of the disclosure. This application is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles of the disclosure and including such departures from the present disclosure come within known customary practice within the art to which the disclosure pertains and may be applied to the essential features herein before set forth.

Claims

1.A cytidine deaminase, wherein the cytidine deaminase is a mutant of a reference polypeptide of SEQ ID NO: 19 and comprises an amino acid mutation (e.g., substitution) relative to (compared to) the reference polypeptide at a position selected from the group consisting of 8, 9, 19, 25, 41, 43, 45, 46, 47, 105, 107, 145, 146, and / or 154 of the reference polypeptide.2.The cytidine deaminase of any preceding claim, wherein the cytidine deaminase comprises an amino acid mutation (e.g., substitution) at a position selected from the group consisting of 25, 46, 105, 107, 146, and / or 154 of the reference polypeptide.3.The cytidine deaminase of any preceding claim, wherein the cytidine deaminase comprises an amino acid substitution selected from the group consisting of E25A, P46G, T105V, E107N, D146R, E154V, and a combination thereof.4.The cytidine deaminase of any preceding claim, wherein the cytidine deaminase comprises a combination substitution of E25A, P46G, T105V, E107N, D146R, and E154V.5.The cytidine deaminase of any preceding claim, wherein the cytidine deaminase comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the reference polypeptide.6.The cytidine deaminase of any preceding claim, wherein the cytidine deaminase comprises, consists essentially of, or consists of the amino acid sequence of SEQ ID NO: 20 or 21, or a N-terminal truncation thereof without the N-terminal Met (e.g., SEQ ID NO: 51) .7.A fusion protein comprising the cytidine deaminase of any preceding claim fused to a heterogeneous functional domain.8.The fusion protein of any preceding claim, wherein the heterogeneous functional domain comprises a polypeptide selected from the group consisting of a DNA binding domain (e.g., a nucleic acid programmable DNA binding domain (napDNAbd) , an RNA binding domain (e.g., a nucleic acid programmable RNA binding domain (napRNAbd) , an MCP) , a UGI domain, and a nuclear localization signal (NLS) .9.The fusion protein of any preceding claim, wherein the fusion protein comprises a nucleic acid programmable DNA binding domain (napDNAbd) and two UGI domains.10.The fusion protein of any preceding claim, comprising, from N-to C-terminus, the cytidine deaminase, an optional linker, the napDNAbd, an optional linker, a UGI domains, an optional linker, a UGI domain, and an NLS.11.The fusion protein of any preceding claim, wherein the napDNAbd comprises a Cas9 nickase or an IscB nickase.12.The fusion protein of any preceding claim, wherein the napDNAbd comprise an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 28 or 52.13.The fusion protein of any preceding claim, wherein the UGI domain comprise an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 54.14.The fusion protein of any preceding claim, wherein the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 45 or 50.15.A system comprising:(1) the fusion protein of any preceding claim, or a polynucleotide (e.g., a DNA, an RNA) encoding the fusion protein, and(2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:(i) a scaffold sequence capable of forming a complex with the fusion protein; and(ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.16.A polynucleotide encoding the cytidine deaminase or the fusion protein of any preceding claim.17.A delivery system comprising (1) the cytidine deaminase of any preceding claim, the polynucleotide of any preceding claim, or the system of any preceding claim; and (2) a delivery vehicle.18.A vector comprising the polynucleotide of any preceding claim.19.A cell comprising the cytidine deaminase of any preceding claim, the system of any preceding claim, the polynucleotide of any preceding claim, or the vector of any preceding claim.20.A method for modifying a target DNA, comprising contacting the target DNA with the system of any preceding claim, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the system.

Citation Information

Patent Citations

  • Fusion protein and base editing tool and method and application thereof

    CN110467679A

  • Nucleobase editors comprising nucleic acid programmable DNA binding proteins

    US20180312828A1

  • Improved cytosine base editing system

    WO2021175288A1