Cytosine deaminases and their use in base editing

JP2025510586A5Pending Publication Date: 2026-06-02INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
Filing Date
2023-03-07
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Current base editing systems lack diversity in deaminases, limiting their ability to accurately manipulate target DNA sequences.

Method used

Identification and utilization of novel cytosine deaminases from different clades, such as the APOBEC/AID clade and the SCP1.201 clade, to expand the capabilities of base editing systems.

Benefits of technology

Enhances the efficiency and accuracy of base editing by providing a broader range of deaminases that can target specific DNA sequences, improving the manipulation of genomic DNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2023169410000001
    Figure 2023169410000001
  • Figure 2023169410000002
    Figure 2023169410000002
Patent Text Reader

Abstract

The present invention relates to the field of genetic engineering.Specifically, the present invention relates to cytosine deaminase and its use in base editing.More specifically, the present invention relates to a method for screening and identifying deaminase, a base editing system based on the newly identified cytosine deaminase, a method for using the base editing system to base edit target sequences in the genome of organisms (e.g., plants), and the genetically modified organisms (e.g., plants) and their progeny produced by the method.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the field of genetic engineering.Specifically, the present invention relates to cytosine deaminase and its use in base editing.More specifically, the present invention relates to a base editing system based on a newly identified cytosine deaminase, a method for base editing a target sequence in the genome of an organism (e.g., a plant) using the base editing system, and the genetically modified organism (e.g., a plant) and its progeny produced by the method.

[0002] 2. Background of the Invention Modification of a specific sequence in the genome of an organism can endow the organism with new stable genetic traits. Here, a single nucleotide mutation at a specific site can cause changes in the amino acid sequence or premature termination of a gene, or cause changes in regulatory sequences, resulting in superior traits. Genome editing technologies such as the CRISPR / Cas9 system can realize the function of targeting a target sequence in the genome. Taking advantage of the property that the genome editing system is bound to the target sequence, the base editing system developed by combining this genome editing system with a deaminase can precisely deaminate the target nucleotide on the genome. The cytosine base editing system can realize the conversion of cytosine (C) to uracil (U) at the target site by fusing the APOBEC / AID family and APOBEC / AID family-like deaminase, and then realize the conversion of cytosine to thymine (T) with the help of the associated repair pathway in the cell. In addition, the efficiency of base editing can be greatly improved by introducing a nick into the non-deaminated single strand on the opposite side and cutting it.

[0003] Based on the structural comparison of deaminases, Iyer et al. searched for proteins with potential deamination functions and classified the above proteins into at least 20 clades (Iyer, LM, Zhang, D., Rogozin, IB, & Aravind, L. (2011). Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems. Nucleic acids research, 39 (22), 9473-9497.). They found that deaminases from different clades differ greatly in structure and sequence. The functions of several clades have been analyzed, including the "dCMP deaminase and ComE" clade, which can convert dCMP to dUMP, the "guanine deaminase" clade, which can convert guanine (G) to xanthine (I), the "RibD-like" clade, which has diaminohydroxyphosphoribosylaminopyrimidine deaminase function, the "Tad1 / ADAR" clade, which has an RNA editing enzyme function that converts RNA adenine (A) to xanthine (I), and the "PurH / AICAR transformylase" clade, which has formyltransferase activity. The functions of some clades, such as the bacterial SCP1.201 clade, XOO2897 clade, MafB19 clade, and Pput_2613 clade, such as the presence or absence of deamination activity and what substrates they can deaminate, have not been analyzed or demonstrated. Currently, only a few deaminases from the APOBEC / AID clade, including APOBEC1, APOBEC3, CDA, AID, CDA1L1 and CDA1L2, have been demonstrated to act on single-stranded DNA and are therefore applicable to cytosine base editing systems.

[0004] There is a need in the art for more deaminases that can be used in base editing systems to expand the systems and improve their ability to precisely manipulate target DNA sequences. [Brief description of the drawings]

[0005] [Figure 1] The cryptic deaminase No. 182 (SEQ ID NO: 1) of the APOBEC / AID clade achieves cytosine base editing in a reporter system. [Diagram 2] Cryptic deaminase No. 182 (SEQ ID NO: 1) of the APOBEC / AID clade mediates cytosine base editing at endogenous sites. [Diagram 3] Cryptic deaminase No. 69 (SEQ ID NO: 2) of the SCP1.201 clade achieves cytosine base editing at endogenous sites. [Figure 4] Cytosine base editing efficiency of eight deaminases that show high editing efficiency at endogenous sites in rice OsACC-T1. [Diagram 5] Cytosine base editing efficiency of eight deaminases that show high editing efficiency at endogenous sites in rice CDC48-T2. [Figure 6] Cytosine base editing efficiency of eight deaminases showing intermediate editing efficiency at endogenous sites in rice OsACC-T1. [Figure 7] Cytosine base editing efficiency of eight deaminases that show intermediate editing efficiency at endogenous sites in rice CDC48-T2. [Figure 8] This is a protein clustering process based on AlphaFold2 predicted structures, where the structures of candidate sequences were predicted using AlphaFold2 and clustered based on structural similarity, and subsequently the cytidine deamination activity of proteins from each structural clade against ssDNA and dsDNA was experimentally tested in plant and human cells. [Figure 9]The process of re-annotation and synthesis of candidate deaminases involves obtaining the full length of the gene encoding the deaminase using Protein BLAST (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi) of the NCBI database, then re-annotating the deaminase domain sequence using hmmscan (https: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan), and the obtained domain sequence was used for structural classification. To confirm that they have deaminase activity, we synthesized several candidate deaminases with extended N- and C-terminal sequences and evaluated their cytidine deaminase activity using a reporter system or at endogenous sites. [Figure 10] Structural similarity matrix reflecting the similarity between 242 predicted protein structures from 16 deaminase families (238) and one outgroup, JAB (4), where proteins from different families are distinguished by different numbers and the color intensity of the heatmap indicates the degree of similarity. [Figure 11A] Based on the protein structure, proteins are classified into different deaminase families, and different families are distinguished by different numbers. [Figure 11B] Representative predicted structures for each of the 16 deaminase clades. [Figure 12] Representative structural alignments of two clades of LmjF365940, APOBEC, dCMP, and MafB19 families corresponding to Fig. 11. Although the two clade systems of each of these four families have partially similar structures, the two clade systems are classified as different clade systems because there are relatively large differences in their overall structures. [Figure 13A]Classification of SCP1.201 deaminases based on protein structure, where the JAB family is considered an outgroup; tested deaminases are shown as single-strand editing (ssDNA), double-strand editing (dsDNA), or non-double-strand / single-strand editing (non_ds / ss) according to their function; undefined deaminases in light grey require further functional analysis; single-strand edited deaminase domains in the figure are SCP356, SCP020, SCP051, SCP170, S The deaminase domains that have undergone double-strand editing are SCP271, SCP103, SCP009, SCP006, SCP004, SCP234, and SCP177, while the remaining deaminases that have been annotated are not edited. [Figure 13B] Predict the core structure of DddA using AlphaFold2. [Figure 13C] This is a typical structural feature of Ddd proteins (proteins with double-stranded deaminase activity). [Figure 13D] Predict the core structure of Sdd7 using AlphaFold2. [Figure 13E] This is a typical structural feature of Sdd proteins (proteins with single-strand deaminase activity). [Figure 14] Identification of ssDNA and dsDNA cytosine deamination activity at endogenous sites in animal cells. (A) Schematic diagram of ssDNA base editing vector for endogenous site editing, (B) Schematic diagram of DdCBE vector and its dyad, (C) DdCBE editing activity on dsDNA and CBE editing activity on ssDNA are detected in HEK293T cells, respectively, and high-throughput sequencing is performed. [Figure 15]Experimental evaluation of the dsDNA deamination activity of Ddd at two endogenous sites in HEK293T cells. For base editing sites used in the calculations, the color intensity represents the level of editing efficiency. [Figure 16A] Experimental evaluation of the ssDNA deamination activity of Sdd at two endogenous sites in HEK293T cells. For base editing sites used in the calculations, color intensity indicates the level of editing efficiency. [Figure 16B] Experimental evaluation of the ssDNA deaminase activity of Sdd at HsJAK2 and HsSIRT6 sites. Data are the average of three replicates and independent experiments. [Figure 17] Evaluation of the editing properties of newly discovered Ddd proteins for use as base editors. (A) Editing efficiency and editing window of dsDNA deaminases Ddd1, Ddd7, Ddd8, Ddd9 and DddA of SCP1.201 at two genomic targets in HEK293T cells. (B) Plasmid library analysis to analyze the context preference of each Ddd protein in mammalian cells, where the candidate proteins targeted and edited the "NC10N" motif. (C) Motif logo plot summarizing the context preference of Ddd1, Ddd7, Ddd8, Ddd9 and DddA from the plasmid library analysis. In the plots, dots indicate single biological replicates, column heights indicate the mean value of editing efficiency, and error bars indicate the standard deviation of three independent biological experiments. [Figure 18] Heatmap of the editing efficiency and editing window of SCP1.201 dsDNA deaminase on two targets in HEK293T cells. [Figure 19] Percentage of editing efficiency of various Ddd deaminases' context preferences in 16 plasmid libraries. Data are represented by the average of three independent experiments. [Figure 20]For the newly discovered Sdd proteins for use as base editors in plants, we assessed the overall editing efficiency of 10 Sdd proteins and rAPOBEC1 at six endogenous targets in rice protoplasts, where the average editing frequency of APOBEC1 at each target was set to 1 and the editing efficiency observed for each Sdd was normalized accordingly. [Figure 21] Editing behavior of Sdd deaminases and APOBEC1 at six endogenous targets in rice protoplasts. Heat maps (A-F) show the editing efficiency and editing window of ten Sdd deaminases and APOBEC1 at OsAAT (A), OsACC1 (B), OsCDC48-T1 (C), OsCDC48-T2 (D), OsDEP1 (E) and OsODEV (F) sites in rice protoplasts. Values ​​designated in cells of heat maps indicate C-to-T editing efficiency, color intensity represents the level of editing efficiency, target sequences are shown at the top of heat maps, dark boxes indicate the position of C-to-T edits, and the last three light color fonts indicate PAMs. Data are represented by the average of three independent experiments. [Figure 22] Editing behavior of SCP1.201 ssDNA deaminases and APOBEC deaminases at three endogenous targets in HEK293T cells. Heat maps (A-C) show the editing efficiency and editing window of four Sdd deaminases and APOBEC1, APOBEC3A, APOBEC1-YE1 and APOBEC1-YEE at HsEMX1 (A), HsHEK2 (B) and HsWFS1 (C) sites in HEK293T cells. Values ​​specified in cells of heat maps indicate C-to-T editing efficiency. Color intensity represents the level of editing efficiency. Target sequences are shown at the top of the heat maps. Dark boxes indicate the position of C-to-T edits. The last three light fonts indicate PAMs. Data are represented by the average of three independent experiments. [Diagram 23]Comparison of editing efficiencies of Sdd7, APOBEC1 and APOBEC3A at five sites in rice protoplasts. (A-E) Comparison of the efficiencies of Sdd7, APOBEC1 and APOBEC3A base editors at five endogenous targets: (A) OsACTG, (B) OsALS-T1, (C) OsALS-T2, (D) OsCDC48-T3 and (E) OsMPK16. Data are representative of three independent experiments. Column heights indicate the average editing efficiency and error bars indicate the standard deviation of three independent biological experiments. [Figure 24] Sequence preferences of Sdd deaminases and APOBEC1 at five endogenous targets in rice protoplasts. Stacked graphs are the context preferences of 10 Sdd deaminases and APOBEC1 at five endogenous targets OsAAT, OsACC1, OsCDC48-T1, OsCDC48-T2 and OsDEP1. Bar graphs represent C-to-T editing preferences of TC, AC, GC and CC from bottom to top, respectively. Data are the results of three independent experiments. [Figure 25A] Overview of high-throughput quantification of Sdd and rAPOBEC1 activity and properties in HEK293T cells using 12K-TRAPseq libraries. [Figure 25B] Evaluation of Sdd and rAPOBEC1 editing preferences and patterns with the 12K-TRAP library. The left panel shows the editing efficiency and editing window of the deaminases, and the right panel shows the logo map of sequence motifs reflecting the context preference of the deaminases. [Figure 26](A) Evaluation of off-target effects using orthogonal R-loop analysis in rice protoplasts. Dots indicate the average frequency of on-target C-to-T conversions for each base editor in six targets in rice (Figure 20) and sgRNA-independent off-target C-to-T conversions in two ssDNAs (OsDEP1-SaT1 and OsDEP1-SaT2). (B) On-target:off-target editing ratios for each base editor in Figure 26A. (C) On-target:off-target editing ratios for Sdd6, rAPOBEC1-YE1, rAPOBEC1-YEE, rAPOBEC1 and hAPOBEC3A tested at two on-target and three off-target sites in HEK293T cells. Dots in the figure indicate single biological replicates, column height indicates average value, and error bars indicate standard deviation of three independent biological replicates. [Figure 27] Specific off-target frequencies of Sdd deaminase and APOBEC1 at two endogenous targets in rice protoplasts (Figures 26A and 26B). Off-targets were assessed using the orthogonal R-loop method. (A,B) are the off-target frequencies of Sdd deaminase and APOBEC1 at the OsDEP1-SaT1 (A) and OsDEP1-SaT2 (B) sites in rice protoplasts. Data are the results of three independent experiments. [Figure 28]Specific on-target and off-target editing efficiencies of Sdd6 and APOBEC base editors tested at two on-target and four off-target sites in HEK293T cells (Figure 26C), including on-target and off-target efficiencies of Sdd6, APOBEC1-YE1, APOBEC1-YEE, APOBEC1, and APOBEC3A at off-target sites HsJAK2-Sa and HsSIRT6-Sa corresponding to on-target sites in HsHEK2, and on-target and off-target editing efficiencies at off-target sites HsRNF2-Sa and HsFANCF-SaT1 corresponding to on-target sites in HsHEK3. Data are the results of three independent experiments. [Figure 29] Conserved protein structures of highly active Sdd deaminases predicted by AlphaFold2, showing the core structure of Sdd deaminases with high deaminating activity, and that α4 is not an essential structure in some active deaminases. [Figure 30A] Engineered truncated Sdd proteins for use in animals and plants.Engineered truncated Sdd proteins.The top figures are the structures of Sdd6, Sdd7, Sdd3 and Sdd9 predicted by AlphaFold2, with conserved regions in darker shades and truncated regions in lighter shades.The bottom figures are the editing efficiencies of Sdd and its minimized versions relative to the original length protein at two endogenous sites in two endogenous rice protoplasts and HEK293T cells, respectively.Dots indicate single biological replicates, column heights and line points indicate the mean, and error bars indicate the standard deviation of three independent biological experiments. [Figure 30B]The Sdd protein has been engineered to be truncated for use in animals and plants, and the SaCas9-based CBE vector could theoretically be packaged into a single AAV. The top figure shows a schematic of the APOBEC / AID-like deaminase, the minimized version of Sdd, and its AAV vector, among which APOBEC3G, hAPOBEC3B, rAPOBEC1, PmCDA1, APOBEC3A and hAID deaminase are too large to be packaged in a single AAV. The bottom figure shows a schematic of the AAV vector based on the minimized mini version of Sdd. [Figure 30C] Figure 1. Editing efficiency of mini-Sdd6, an engineered truncated Sdd protein for use in animals and plants, at two endogenous target sites in the MmHPD gene in mouse N2a cells. Dots indicate single biological replicates, column heights and line points indicate average values, and error bars indicate standard deviations of three independent biological experiments. [Figure 30D] Engineered truncated Sdd proteins for use in animals and plants. Editing efficiencies of mini-Sdd7, rAPOBEC1, hAPOBECA and human AID base editors at five endogenous targets in soybean hairy roots. Dots indicate single biological replicates, column heights and line points indicate average values, and error bars indicate standard deviations of three independent biological experiments. [Figure 30E] A truncated Sdd protein engineered for use in animals and plants, and the frequency of mini-Sdd7-induced mutations in T0 generation soybean plant lines. [Figure 30F] Engineered truncated Sdd proteins for use in animals and plants, and base-edited genotypes of soybean plant lines. [Figure 30G] A truncated Sdd protein engineered for use in animals and plants. Phenotypes of soybean plant lines treated with carfentrazone-ethyl for 10 days. On the left is a wild-type soybean plant line (R98) and on the right is a base-edited soybean plant line (C98). [Diagram 31]Base editing efficiency in regenerated rice. (A) Schematic diagram of Agrobacterium-mediated transformation of rice base editing binary vectors. (B) Efficiency of mini-Sdd7 and hAPOBEC3A base editors in inducing mutations in T0 rice plant lines. [Diagram 32] FIG. 1 is a schematic diagram of a base-edited binary vector for Agrobacterium-mediated transformation in soybean.

[0006] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENT 1.Definition In the present invention, unless otherwise specified, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and experimental operation steps related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are terms and common steps widely used in the corresponding fields. In addition, the definitions and explanations of related terms are provided below to facilitate a better understanding of the present invention.

[0007] As used herein, the term "and / or" includes all combinations of the items connected by this term, each combination being considered as if listed individually herein. For example, "A and / or B" includes "A," "A and B," and "B." For example, "A, B and / or C" includes "A," "B," "C," "A and B," "A and C," "B and C," and "A, B and C."

[0008] "Cytosine deaminase" refers to a deaminase that can accept a nucleic acid, such as single-stranded DNA, as a substrate and can catalyze the deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively.

[0009] As used herein, "genome" encompasses not only the chromosomal DNA present in the cell nucleus, but also organelle DNA present in the subcellular components of the cell (eg, mitochondria, plastids).

[0010] As used herein, "organism" includes any organism suitable for genome editing, preferably eukaryotic organism.Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, and cats; poultry such as chickens, ducks, and geese; plants including monocotyledonous and dicotyledonous plants, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis, and the like.

[0011] "Genetically modified organism" or "genetically modified cell" refers to an organism or cell that contains an exogenous polynucleotide, modified gene or expression control sequence in its genome. For example, an exogenous polynucleotide can be stably integrated into the genome of an organism or cell and passed on to successive generations. An exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant DNA construct. A modified gene or expression control sequence is one whose sequence contains single or multiple deoxynucleotide substitutions, deletions and additions in the genome of an organism or cell.

[0012] "Exogenous" with respect to a sequence means a sequence that originates from a foreign species or, if from the same species, has been significantly altered in composition and / or genetic locus from its original form by deliberate human intervention.

[0013] "Polynucleotide", "nucleic acid sequence", "nucleotide sequence" or "nucleic acid fragment", used interchangeably, is a single- or double-stranded RNA or DNA polymer, optionally containing synthetic, non-natural or modified nucleotide bases. Nucleotides are referred to by their single alphabetical letter names as follows: "A" for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), "C" for cytidine or deoxycytidine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purine (A or G), "Y" for pyrimidine (C or T), "K" for G or T, "H" for A, C or T, "I" for inosine, and "N" for any nucleotide.

[0014] "Polypeptide", "peptide" and "protein" are used interchangeably in the present invention to refer to a polymer of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide", "peptide", "amino acid sequence" and "protein" also include modified forms, including, but not limited to, glycosylation, lipid attachment, sulfation, gamma carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.

[0015] Sequence "identity" has an art-recognized meaning, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using published techniques. Sequence identity can be measured along the entire length of a polynucleotide or polypeptide, or along a region of the molecule (see, e.g., Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). Although there are many methods for measuring identity between two polynucleotides or polypeptides, the term "identity" is familiar to those skilled in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48:1073 (1988)).

[0016] When the term "comprises" is used herein to describe a protein or nucleic acid sequence, the protein or nucleic acid may consist of the sequence or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in the present invention. Furthermore, it is clear to those skilled in the art that the methionine encoded by the start codon at the N-terminus of a polypeptide may be retained under certain practical circumstances (e.g., when expressed in a certain expression system), but without substantially affecting the function of the polypeptide. Thus, when the amino acid sequence of a particular polypeptide is described in the specification and claims of this application, it may not contain the N-terminal methionine encoded by the start codon, but even in that case, the sequence including the methionine is included. Accordingly, the coding nucleotide sequence may include the start codon, and vice versa.

[0017] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art and can generally be made without altering the biological activity of the resulting molecule.Those skilled in the art generally recognize that single amino acid substitutions in non-essential regions of a polypeptide do not substantially alter biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub.co., p.224).

[0018] As used herein, "expression construct" refers to a vector, such as a recombinant vector, suitable for expressing a target nucleotide sequence in vivo. "Expression" refers to producing a functional product. For example, expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to produce mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein.

[0019] An "expression construct" of the present invention may be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (eg, mRNA).

[0020] An "expression construct" of the present invention may contain regulatory sequences and target nucleotide sequences from different sources, or regulatory sequences and target nucleotide sequences from the same source but arranged in a manner different from that normally found in nature.

[0021] "Regulatory sequence" and "regulatory element" are used interchangeably and refer to nucleotide sequences located upstream (5' non-coding sequences), in the middle or downstream (3' non-coding sequences) of a coding sequence that influence the transcription, RNA processing, stability, translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences.

[0022] "Promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the present invention, the promoter is a promoter capable of controlling the transcription of a gene in a cell, regardless of whether it originates from said cell. The promoter may be a constitutive promoter, a tissue-specific promoter, a developmentally regulated promoter, or an inducible promoter.

[0023] A "constitutive promoter" generally refers to a promoter that causes expression of a gene in most cell types under most circumstances. A "tissue-specific promoter" and a "tissue-preferential promoter" are used interchangeably and refer to a promoter that can be expressed primarily, but not necessarily exclusively, in one tissue or organ, as well as in one particular cell or cell type. A "developmentally regulated promoter" refers to a promoter whose activity is determined by developmental events. An "inducible promoter" selectively expresses an operably linked DNA sequence in response to endogenous or exogenous stimuli (environment, hormones, chemical signals, etc.).

[0024] Examples of promoters include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. Examples of pol I promoters include chicken RNA pol I promoter. Examples of pol II promoters include, but are not limited to, cytomegalovirus immediate early (CMV) promoter, Rous sarcoma virus long terminal repeat (RSV-LTR) promoter, and simian virus 40 (SV40) immediate early promoter. Examples of pol III promoters include U6 and H1 promoters. Inducible promoters such as metallothionein promoters can be used. Other examples of promoters include T7 phage promoter, T3 phage promoter, β-galactosidase promoter, and Sp6 phage promoter. When used in plants, the promoter can be cauliflower mosaic virus 35S promoter, maize Ubi-1 promoter, wheat U6 promoter, rice U3 promoter, maize U3 promoter, rice actin promoter.

[0025] As used herein, the term "operably linked" means that regulatory elements (e.g., but not limited to, promoter sequences, transcription termination sequences, etc.) are linked to a nucleic acid sequence (e.g., a coding sequence, an open reading frame, etc.) such that transcription of the nucleotide sequence is controlled and regulated by the transcriptional regulatory elements. Techniques for operably linking regulatory element regions to nucleic acid molecules are known in the art.

[0026] "Introducing" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism refers to transforming a cell of the organism with the nucleic acid or protein so that the nucleic acid or protein can function within the cell. "Transformation" as used herein includes stable transformation and transient transformation.

[0027] "Stable transformation" refers to the introduction of an exogenous nucleotide sequence into a genome, resulting in the stable inheritance of the foreign gene. When stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and subsequent successive generations.

[0028] "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell to carry out a function without stable inheritance of the foreign gene. In transient transformation, the exogenous nucleic acid sequence is not integrated into the genome.

[0029] 2. Protein clustering and function prediction methods based on three-dimensional structures In one aspect, the present invention provides a method for protein clustering, comprising: (1) obtaining sequences of a plurality of candidate proteins from a database; (2) predicting the three-dimensional structure of each of the plurality of candidate proteins using a protein prediction program; (3) performing a multiple structural alignment of the three-dimensional structures of the plurality of candidate proteins using a scoring function to obtain a structural similarity matrix; and (4) clustering the plurality of candidate proteins based on the structural similarity matrix by a phylogenetic tree construction method.

[0030] In some embodiments, in step (1), the sequences of the multiple candidate proteins are obtained according to annotation information in a database. For example, when clustering deaminases, the sequences of multiple candidate proteins annotated as "deaminase" may be selected from a database.

[0031] In some embodiments, in step (1), the sequences of the plurality of candidate proteins are obtained by searching a database based on sequence identity / similarity using the sequence of a reference protein. For example, the sequences of the plurality of candidate proteins can be obtained by searching a database using a BLAST program based on the sequence of a reference protein with a known function. In some embodiments, the plurality of candidate proteins have at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with the reference protein.

[0032] In some embodiments, the candidate protein is a deaminase, hi some preferred embodiments, the candidate protein is a cytosine deaminase.

[0033] In some embodiments, the database is an InterPro database.

[0034] In some embodiments, the protein structure prediction program in step (2) is selected from AlphaFold2, RoseTT, or other programs capable of predicting protein structures (John Jumper and others, 'Highly Accurate Protein Structure Prediction with AlphaFold', Nature, 596.7873 (2021), 583-89).

[0035] In some embodiments, the scoring function used in step (3) includes TM-score, RMSD, LDDT, GDT score, QSC, FAPE, or other scoring functions capable of scoring protein structure similarity (John Jumper and others, 'Highly Accurate Protein Structure Prediction with AlphaFold', Nature, 596.7873 (2021), 583-89).

[0036] In some embodiments, when the scoring function is a TM-score, the TM-score is at least 0.6, at least 0.7, at least 0.75, at least 0.8, at least 0.85, or more. For example, the calculation of the TM-score can refer to the formula and method described in the "Materials and Methods" section of the Examples of this application.

[0037] In some embodiments, the phylogenetic tree construction method in step (4) is the Unweighted Pair Group Method with Arithmetic Mean (UPGMA) (CP Kurtzman, Jack W. Fell, and T. Boekhout, The Yeasts: A Taxonomic Study, 5th ed (Amsterdam: Elsevier, 2011); 'A Statistical Method for Evaluating Systematic Relationships-Robert Reuven Sokal, Charles Duncan Michener-Google Books').

[0038] In some embodiments, step (4) comprises obtaining a clustering dendrogram of the plurality of candidate proteins.

[0039] In one aspect, the present invention provides a method for predicting a function of a protein based on a three-dimensional structure, comprising: clustering a plurality of candidate proteins based on the protein clustering method of the present invention; and then predicting a function of the candidate proteins based on the clustering results.

[0040] In some embodiments, the plurality of candidate proteins includes at least one reference protein of known function.

[0041] In some embodiments, the function of other candidate proteins of the same clade or subclade is predicted by their position in the cluster (dendrogram) of reference proteins with known functions. In some embodiments, other candidate proteins located in the same clade or subclade as a reference protein are predicted to have the same or similar function as the reference protein. In some embodiments, the TM-score between different candidate proteins in the same clade or subclade is at least 0.6, at least 0.7, at least 0.75, at least 0.8, at least 0.85, or more. In some embodiments, the TM-score between candidate proteins of different clades or subclades is less than 0.85, less than 0.8, less than 0.75, less than 0.7, less than 0.6, or less.

[0042] In some embodiments, the reference protein is a deaminase. In some preferred embodiments, the reference protein is a cytosine deaminase. In some embodiments, the reference protein is a reference cytosine deaminase, and the reference cytosine deaminase is rAPOBEC1, the sequence of which is shown in SEQ ID NO: 64, or DddA, the sequence of which is shown in SEQ ID NO: 65. In some embodiments, the TM-score between different candidate proteins in the same clade or subclade as the reference protein, or the TM-score with the reference protein is at least 0.7. In some embodiments, the candidate proteins in a different clade or subclade from the reference protein have a TM-score with the reference protein of less than 0.7.

[0043] In another aspect, the present invention provides a method for producing a composition comprising: a) aligning the structures of a plurality of candidate proteins clustered by the method of the present invention, e.g. clustered into the same clade or subclade, to determine a conserved core structure; b) identifying said conserved core structure as a minimal functional domain of a protein based on its three-dimensional structure.

[0044] As used herein, a "minimal functional domain" refers to the smallest portion of a protein that is capable of substantially retaining a function of the full-length protein.

[0045] In some embodiments, the plurality of candidate proteins includes at least one reference protein of known function.

[0046] In some embodiments, the reference protein is a deaminase. In some preferred embodiments, the reference protein is a cytosine deaminase. In some embodiments, the reference protein is a reference cytosine deaminase, and the reference cytosine deaminase is rAPOBEC1, whose sequence is shown in SEQ ID NO: 64, or DddA, whose sequence is shown in SEQ ID NO: 65.

[0047] In another aspect, the present invention provides a cytosine deaminase identified by the protein function prediction method of the present invention.

[0048] In another aspect, the present invention provides a truncated cytosine deaminase comprising or consisting of a minimal functional domain of a cytosine deaminase identified by the methods of the present invention.

[0049] In one aspect, the present invention further provides for the use of a cytosine deaminase or a truncated cytosine deaminase in gene editing, such as base editing, in an organism or an organism cell.

[0050] 3. Cytosine deaminase and base-editing fusion protein containing it In one aspect, the present invention provides a cytosine deaminase capable of deaminating cytosine bases of deoxycytidine in DNA, hi some embodiments, the cytosine deaminase is derived from a bacterium.

[0051] In some embodiments, the cytosine deaminase has a TM score of 0.6 or more, 0.7 or more, 0.75 or more, 0.8 or more, or 0.85 or more with AlphaFold2 of the three-dimensional structure of a reference cytosine deaminase, and has 20 to 70%, 20 to 60%, 20 to 50%, 20 to 45%, 20 to 40%, or 20 to 35% sequence identity, or at least 20%, at least 30%, at least 40%, or at least The cytosine deaminase has an action of deaminating a cytosine base of deoxycytidine in DNA, and the cytosine deaminase has an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the cytosine deaminase.

[0052] In some embodiments, the canonical cytosine deaminase is (a) rAPOBEC1, the sequence of which is set forth in SEQ ID NO: 64, or (b) DddA, the sequence of which is set forth in SEQ ID NO: 65; or (c) Sdd7, the sequence of which is shown in SEQ ID NO:4.

[0053] In some embodiments, the TM-score with the three-dimensional structure of AlphaFold2 of rAPOBEC1, whose sequence is set forth in SEQ ID NO:64, is 0.6 or more, 0.7 or more, 0.75 or more, 0.8 or more, or 0.85 or more, and the cytosine deaminase comprises an amino acid sequence having 20 to 70%, 20 to 60%, 20 to 50%, 20 to 45%, 20 to 40%, or 20 to 35% sequence identity, or at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO:64, and the cytosine deaminase has the effect of deaminating cytosine bases in deoxycytidine of DNA.

[0054] In some embodiments, the sequence has a TM score of 0.6 or more, 0.7 or more, 0.75 or more, 0.8 or more, or 0.85 or more with respect to the three-dimensional structure of AlphaFold2 of DddA shown in SEQ ID NO: 65, and includes an amino acid sequence having 20 to 70%, 20 to 60%, 20 to 50%, 20 to 45%, 20 to 40%, or 20 to 35% sequence identity with SEQ ID NO: 65, or at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, and the cytosine deaminase has the effect of deaminating the cytosine base of deoxycytidine in DNA.

[0055] In some embodiments, the TM-score with the three-dimensional structure of AlphaFold2 of Sdd7, whose sequence is set forth in SEQ ID NO: 4, is 0.6 or more, 0.7 or more, 0.75 or more, 0.8 or more, or 0.85 or more, and the cytosine deaminase comprises an amino acid sequence having 20 to 70%, 20 to 60%, 20 to 50%, 20 to 45%, 20 to 40%, or 20 to 35% sequence identity, or at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 4, and the cytosine deaminase has the effect of deaminating cytosine bases in deoxycytidine of DNA.

[0056] In some embodiments, the cytosine deaminase is from the AID / APOBEC clade, the SCP1.201 clade, the MafB19 clade, a novel AID / APOBEC-like clade, the TM1506 clade, or the XOO2897 clade.

[0057] In the present specification, the cytosine deaminase clade is determined by the contents described in Iyer, LM, Zhang, D., Rogozin, IB, & Aravind, L. (2011). Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems. Nucleic acids research, 39(22), 9473-9497.

[0058] In some embodiments, the cytosine deaminase is from the AID / APOBEC clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to SEQ ID NO:1.

[0059] In some embodiments, the cytosine deaminase is from the SCP1.201 clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to any of SEQ ID NOs: 28-40. In some embodiments, the cytosine deaminase is capable of deaminating cytosine bases in double-stranded DNA. In some embodiments, the amino acid sequence of the cytosine deaminase consists of any of the amino acid sequences of SEQ ID NOs: 28, 33, 34, and 35.

[0060] In some embodiments, the cytosine deaminase is from the SCP1.201 clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to any of SEQ ID NOs: 2-18, 41-49. In some embodiments, the cytosine deaminase is capable of deaminating cytosine bases in single-stranded DNA. In some embodiments, the amino acid sequence of the cytosine deaminase consists of any of the amino acid sequences of SEQ ID NOs: 2-18, 41-49. In some embodiments, the amino acid sequence of the cytosine deaminase consists of any of the amino acid sequences of SEQ ID NOs: 2-7, 12, and 17.

[0061] In some embodiments, the cytosine deaminase is a truncated cytosine deaminase capable of deaminating cytosine bases of deoxycytidines of DNA. In some embodiments, the truncated cytosine deaminase ranges in length from 130 to 160 amino acids. In some embodiments, the truncated cytosine deaminase may be packaged individually within an AAV particle.

[0062] In some embodiments, the truncated cytosine deaminase comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with any of SEQ ID NOs: 50-55. In some embodiments, the truncated cytosine deaminase can deaminate cytosine bases in single-stranded DNA. In some embodiments, the truncated cytosine deaminase consists of any of the amino acid sequences of SEQ ID NOs: 50-55.

[0063] In some embodiments, the cytosine deaminase is from the MafB19 clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to any of SEQ ID NOs: 19, 56, 57, 58.

[0064] In some embodiments, the cytosine deaminase is from a novel AID / APOBEC-like clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to any of SEQ ID NOs:20, 21.

[0065] In some embodiments, the cytosine deaminase is from the TM1506 clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to SEQ ID NO:22.

[0066] In some embodiments, the cytosine deaminase is from the XOO2897 clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to SEQ ID NOs:23, 24, 59-62.

[0067] In some embodiments, the cytosine deaminase is from the toxin deaminase clade and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to SEQ ID NO: 74 or 75.

[0068] In one aspect, the present application relates to the use of the cytosine deaminase of the present invention in gene editing, such as base editing, in an organism or a cell of an organism.

[0069] In some embodiments, the cytosine deaminase is used to produce a base editing fusion protein or a base editing system for performing base editing in an organism or a cell of an organism.

[0070] In another aspect, the invention provides a base editing fusion protein comprising a nucleic acid targeting domain and a cytosine deamination domain, wherein the cytosine deamination domain comprises at least one (e.g., one or two) cytosine deaminase polypeptides of the invention.

[0071] In embodiments herein, "fusion protein," "base editing fusion protein," and "base editor" are used interchangeably to refer to a protein having one or more nucleotide substitutions that can mediate targeting of a sequence in a genome in a sequence-specific manner, such as a C to T substitution.

[0072] As used herein, a "nucleic acid targeting domain" refers to a domain that can mediate the attachment of a base editing fusion protein to a specific target sequence in a genome in a sequence-specific manner (e.g., by a guide RNA). In some embodiments, the nucleic acid targeting domain may include one or more zinc finger protein domains (ZFPs) or transcription factor effector domains (TALEs) that are directed to a specific target sequence. In some embodiments, the nucleic acid targeting domain includes at least one (e.g., one) CRISPR effector protein (CRISPR effector) polypeptide.

[0073] A "zinc finger protein domain (ZFP)" typically contains 3-6 individual zinc finger repeats, each of which can recognize a unique sequence of, for example, 3 bp. Different zinc finger repeats can be combined to target different genomic sequences.

[0074] A "transcription activator-like effector domain" is the DNA-binding domain of a transcription activator-like effector (TALE). TALEs can be engineered to bind to almost any DNA sequence.

[0075] As used herein, the term "CRISPR effector protein" generally refers to the nuclease (CRISPR nuclease) present in naturally occurring CRISPR system or its functional variant. This term encompasses any effector protein based on the CRISPR system that enables sequence-specific targeting in cells.

[0076] As used herein, "functional variant" in relation to CRISPR nuclease means that it at least retains the sequence-specific targeting ability mediated by guide RNA. Preferably, said functional variant is a nuclease inactive variant, i.e., lacks double-stranded nucleic acid cleavage activity. However, CRISPR nucleases lacking double-stranded nucleic acid cleavage activity also include nickases that form nicks in double-stranded nucleic acid molecules but do not completely cleave double-stranded nucleic acid. In some preferred embodiments of the present invention, the CRISPR effector protein according to the present invention has nickase activity. In some embodiments, said functional variant recognizes a different PAM (protospacer adjacent motif) sequence than wild-type nuclease.

[0077] The "CRISPR effector protein" may be derived from Cas9 nuclease, including Cas9 nuclease or functional variants thereof. The Cas9 nuclease may be derived from different species, such as spCas9 from S.pyogenes or SaCas9 from S.aureus. "Cas9 nuclease" and "Cas9" are used interchangeably herein and refer to an RNA-guided nuclease that includes Cas9 protein or a fragment thereof (e.g., a protein that includes the active DNA cleavage domain of Cas9 and / or the gRNA binding domain of Cas9). Cas9 is a component of the CRISPR / Cas (clustered regularly interspaced short palindromic repeats and associated systems) genome editing system, which can target and cleave DNA target sequences under the guidance of guide RNA to form DNA double-strand breaks (DSBs). An exemplary amino acid sequence of wild-type SpCas9 is shown in SEQ ID NO:25.

[0078] The "CRISPR effector protein" may be derived from Cpf1 nuclease, including Cpf1 nuclease or a functional variant thereof. The Cpf1 nuclease may be derived from different species, such as Cpf1 nuclease from Francisella novicida U112, Acidaminococcus sp. BV3L6 and Lachnospiraceae bacterium ND2006.

[0079] Available "CRISPR effector proteins" include Cas3, Cas8a, Cas5, Cas8b, Cas8c, Cas10d, Cse1, Cse2, Csy1, Csy2, Csy3, GSU0054, Cas10, Csm2, Cmr5, Cas10, Csx11, Csx10, Csf1, Csn2, Cas4, C2c1 (Cas12b), C2c3, C2c2, Cas The nuclease may be derived from nucleases such as Cas12c, Cas12d (i.e., CasY), Cas12e (i.e., CasX), Cas12f (i.e., Cas14), Cas12g, Cas12h, Cas12i, Cas12j (i.e., CasΦ), Cas12k, Cas12l, and Cas12m, for example, including these nucleases or functional variants thereof.

[0080] In some embodiments, the CRISPR effector protein is a nuclease-inactive Cas9. It is known that the DNA cleavage domain of Cas9 nuclease contains two subdomains: HNH nuclease subdomain and RuvC subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC subdomain cleaves the non-complementary strand. Mutation of these subdomains can inactivate the nuclease activity of Cas9 to form a "nuclease-inactive Cas9". The nuclease-inactive Cas9 retains gRNA-guided DNA binding ability.

[0081] The nuclease-inactive Cas9 according to the present invention may be derived from different species of Cas9, such as S. pyogenes Cas9 (SpCas9) or S. aureus Cas9 (SaCas9). When the HNH nuclease subdomain and the RuvC subdomain of Cas9 are mutated simultaneously (e.g., containing the mutations D10A and H840A), the Cas9 nuclease becomes inactive, resulting in nuclease-dead Cas9 (dCas9). When one of the subdomains is mutationally inactivated, Cas9 has nickase activity, i.e., Cas9 nickase (nCas9), e.g., nCas9 with only the mutation D10A is obtained.

[0082] Thus, in some embodiments of each aspect of the invention, the nuclease-inactive Cas9 variants described herein comprise the amino acid substitutions D10A and / or H840A compared to wild-type Cas9, the amino acid numbering of which is set forth in SEQ ID NO: 25. In some preferred embodiments, the nuclease-inactive Cas9 comprises the amino acid substitution D10A compared to wild-type Cas9, the amino acid numbering of which is set forth in SEQ ID NO: 25. In some embodiments, the nuclease-inactive Cas9 comprises the amino acid sequence set forth in SEQ ID NO: 26 (nCas9(D10A)).

[0083] When Cas9 nuclease is used for gene editing, it usually requires a PAM (protospacer adjacent motif) sequence of 5'-NGG-3' at the 3' end of the target sequence.However, the present inventors have surprisingly found that the frequency of this PAM sequence in some species, such as rice, is very low, which significantly limits gene editing in these species, such as rice.Therefore, the present invention preferably uses CRISPR effector proteins that recognize different PAM sequences, such as functional variants of Cas9 nuclease with different PAM sequences.

[0084] In some embodiments of the present invention, the cytidine deamination domain of the fusion protein converts cytidine deamination in single-stranded DNA generated upon formation of the fusion protein-guide RNA-DNA complex to U, and further achieves C to T base substitution by base mismatch repair.

[0085] In some embodiments of the invention, the nucleic acid targeting domain and the cytosine deamination domain are fused by a linker.

[0086] As used herein, a "linker" may be a non-functional amino acid sequence having a length of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids, with no secondary or higher structure. For example, the linker may be a flexible linker.

[0087] In some embodiments, the base editing fusion protein comprises, from N-terminus to C-terminus, a cytosine deamination domain, and a nucleic acid targeting domain, in that order.

[0088] Furthermore, in cells, uracil DNA glycosylase catalyzes the removal of U from DNA, initiating base excision repair (BER) and causing the repair of U:G to C:G. Thus, without being bound by any theory, the base editing fusion proteins of the invention can be combined with a uracil DNA glycosylase inhibitor (UGI) to increase the efficiency of C to T base editing.

[0089] In some embodiments, the base editing fusion protein is co-expressed with a uracil DNA glycosylase inhibitor (UGI).

[0090] In some embodiments, the base editing fusion protein further comprises a uracil DNA glycosylase inhibitor (UGI).

[0091] In some embodiments, the UGI is linked to other portions of the base editing fusion protein via a linker.

[0092] In some embodiments, the UGI is linked to other portions of the base editing fusion protein via a "self-cleaving peptide."

[0093] As used herein, "self-cleaving peptide" refers to a peptide that can achieve self-cleavage within a cell. For example, the self-cleaving peptide can include a protease recognition site, which allows it to be recognized and specifically cleaved by a protease within the cell. Alternatively, the self-cleaving peptide can be a 2A polypeptide. The 2A polypeptide is a short peptide derived from a virus that undergoes self-cleavage during translation. When two different target polypeptides are linked using a 2A polypeptide and expressed in the same reading frame, the two target polypeptides are generated in a ratio of approximately 1:1. Commonly used 2A polypeptides are P2A from porcine techovirus-1, T2A from Thosea asigna virus, E2A from equine rhinitis A virus, and F2A from foot-and-mouth disease virus. Various functional variants of these 2A polypeptides are also known in the art and can be used in the present invention.

[0094] Preferably, the self-cleaving peptide is not present between or within the nucleic acid targeting domain, the cytosine deamination domain, In some embodiments, UGI is located at the N-terminus or C-terminus, preferably the C-terminus, of the base editing fusion protein.

[0095] In some specific embodiments, the uracil DNA glycosylase inhibitor (UGI) comprises the amino acid sequence set forth in SEQ ID NO:27.

[0096] In some embodiments of the present invention, the fusion protein of the present invention may further comprise a nuclear localization sequence (NLS). In general, one or more NLSs in the fusion protein should be strong enough to promote the accumulation of the fusion protein in the cell nucleus in an amount that allows its base editing function to be realized. In general, the strength of the nuclear localization activity depends on the number, position, and specific NLS or NLSs used in the fusion protein, or a combination of these factors.

[0097] In some embodiments of the invention, the NLS of the fusion protein of the invention may be located at the N-terminus and / or the C-terminus. In some embodiments of the invention, the NLS of the fusion protein of the invention may be located between the adenine deamination domain, the cytosine deamination domain, the nucleic acid targeting domain and / or the UGI. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs located at or near the N-terminus. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs located at or near the C-terminus. In some embodiments, the polypeptide comprises a combination thereof, for example, one or more NLSs at the N-terminus and one or more NLSs at the C-terminus. When multiple NLSs are present, each can be selected independently from the other NLSs.

[0098] Generally, NLSs consist of one or more short sequences of positively charged lysines or arginines exposed on the protein surface, although other types of NLSs are known. Non-limiting examples of NLSs include KKRKV, PKKKRKV, or KRPAATKKAGQAKKKK.

[0099] Additionally, depending on the location of the DNA requiring editing, the fusion proteins of the present invention may contain other localization sequences, such as a cytoplasmic localization sequence, a chloroplast localization sequence, a mitochondrial localization sequence, etc.

[0100] 4. Base editing system In another aspect, the invention provides a base editing system comprising i) a cytosine deaminase or base editing fusion protein of the invention, and / or an expression construct comprising a nucleotide sequence encoding said cytosine deaminase or base editing fusion protein.

[0101] In some embodiments, the base editing system is used to modify a nucleic acid target region.

[0102] In some embodiments, the base editing system further comprises ii) at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA. However, a person skilled in the art will appreciate that if the base editing fusion protein is not based on a CRISPR effector protein, the system may not require a guide RNA or an expression construct encoding same.

[0103] In some embodiments, the at least one guide RNA is capable of binding to a nucleic acid targeting domain of the fusion protein. In some embodiments, the guide RNA is directed to at least one target sequence within the nucleic acid target region.

[0104] As used herein, a "base editing system" refers to a combination of components required for base editing of a nucleic acid sequence, such as a genomic sequence, in a cell or organism. Each component of the system, such as a cytosine deaminase, a base editing fusion protein, and one or more guide RNAs, can exist independently or in any combination as a composition.

[0105] In some embodiments, it comprises a cytosine deaminase of the invention or a fusion protein of the invention and a guide RNA capable of binding to a nucleic acid targeting binding protein.

[0106] As used herein, "guide RNA" and "gRNA" are used interchangeably and refer to an RNA molecule that can form a complex with a CRISPR effector protein and has a certain degree of identity with the target sequence, thereby targeting the complex to the target sequence. The guide RNA targets the target sequence by base pairing with the complementary strand of the target sequence. For example, the gRNA used for Cas9 nuclease or its functional variants is usually composed of crRNA and tracrRNA molecules that are partially complementary to form a complex, where the crRNA contains a guide sequence (also called a seed sequence) that has sufficient identity with the target sequence to hybridize to the complementary strand of the target sequence and guide the CRISPR complex (Cas9+crRNA+tracrRNA) to specifically bind to the target sequence. However, it is known in the art that single guide RNAs (sgRNAs) can be designed that contain features of both crRNA and tracrRNA. The gRNA used for Cpf1 nuclease or its functional variants is usually composed of only a mature crRNA molecule, also called sgRNA. It is within the capabilities of one of ordinary skill in the art to design an appropriate gRNA based on the CRISPR nuclease being used and the target sequence to be edited.

[0107] In some embodiments, the guide RNA is 15-100 nucleotides in length and comprises at least 10, at least 15, or at least 20 contiguous nucleotide sequences complementary to a target sequence.

[0108] In some embodiments, the guide RNA comprises a sequence of 15 to 40 contiguous nucleotides complementary to a target sequence.

[0109] In some embodiments, the guide RNA is 15-50 nucleotides in length.

[0110] In some embodiments, the target sequence is a DNA sequence.

[0111] In some embodiments, the target sequence is in the genome of an organism. In some embodiments, the organism is a prokaryote. In some embodiments, the prokaryote is a bacterium. In some embodiments, the organism is a eukaryote. In some embodiments, the organism is a plant or a fungus. In some embodiments, the organism is a vertebrate. In some embodiments, the vertebrate is a mammal. In some embodiments, the mammal is a mouse, a rat, or a human. In some embodiments, the organism is a cell. In some embodiments, the cell is a mouse cell, a rat cell, or a human cell. In some embodiments, the cell is a HEK-293 cell.

[0112] In some embodiments, after the base editing system of the present invention is introduced into a cell, the base editing fusion protein and the guide RNA can form a complex, and this complex specifically targets the target sequence under the mediation of the guide RNA, and one or more Cs are replaced with Ts and / or one or more As are replaced with Gs within the target sequence.

[0113] In some embodiments, the at least one guide RNA may target a target sequence located on the sense strand (e.g., protein-coding strand) and / or antisense strand within a target nucleic acid region in a genome. When a guide RNA targets a sense strand (e.g., protein-coding strand), the base editing composition of the present invention can replace one or more Cs with Ts and / or replace one or more As with Gs in a target sequence on the sense strand (e.g., protein-coding strand). When a guide RNA targets an antisense strand, the base editing composition of the present invention can replace one or more Gs with As and / or replace one or more Ts with Cs in a target sequence on the sense strand (e.g., protein-coding strand).

[0114] To obtain efficient expression within a cell, in some embodiments of the invention, the nucleotide sequence encoding the cytosine deaminase or base editing fusion protein is codon-optimized for the organism whose genome is to be modified.

[0115] Codon optimization refers to modifying a nucleic acid sequence to replace at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of a native sequence with a codon that is more frequently or most frequently used in the host cell's genes while maintaining this native amino acid sequence and enhancing expression in the target host cell. Different species show specific preferences for specific codons for specific amino acids. Codon preferences (differences in codon usage between organisms) are often related to the efficiency of messenger RNA (mRNA) translation, which is believed to depend on the nature of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell usually reflects the codons most frequently used for peptide synthesis. Thus, genes can be tailored to obtain optimal gene expression in a particular organism based on codon optimization. Codon usage tables can be found, for example, in www.kazusa.orjp / codon / Codon usage tables are readily available, for example from the Codon Usage Database available at http: / / www.ncbi.nlm.nih.gov / pubmed / 1024741, and these tables may be adjusted in various ways (see Nakamura Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000)).

[0116] The organisms whose genomes can be modified by the base editing system of the present invention include any organisms suitable for base editing, preferably eukaryotic organisms.Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, and cats; poultry such as chickens, ducks, and geese; and plants including monocotyledonous and dicotyledonous plants, such as crop plants including, but not limited to, wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0117] 5. Base editing methods In another aspect, the present invention provides a base editing method comprising contacting a base editing system of the present invention with a target sequence of a nucleic acid molecule.

[0118] In some embodiments, the nucleic acid molecule is a DNA molecule. In some preferred embodiments, the nucleic acid molecule is a double-stranded DNA molecule or a single-stranded DNA molecule.

[0119] In some embodiments, the target sequence of the nucleic acid molecule comprises a sequence associated with a trait or expression in a plant.

[0120] In some embodiments, the target sequence of the nucleic acid molecule comprises a sequence or point mutation associated with a disease or disorder.

[0121] In some embodiments, the base editing system contacts a target sequence of a nucleic acid molecule and exerts a deamination effect, which results in the substitution of one or more nucleotides in the target sequence.

[0122] In some embodiments, the target sequence comprises the DNA sequence 5'-MCN-3', where M is A, T, C or G, and N is A, T, C or G, and the central C of the 5'-MCN-3' sequence is deaminated.

[0123] In some embodiments, the deamination action introduces or removes a splice site.

[0124] In some embodiments, the deamination effect introduces a mutation into a gene promoter, which increases or decreases transcription of a gene operably linked to the gene promoter.

[0125] In some embodiments, the deamination effect introduces a mutation into a gene repressor, the mutation increasing or decreasing transcription of a gene operably linked to the gene repressor.

[0126] In some embodiments, the contacting occurs in vivo.

[0127] In some embodiments, the contacting occurs in vitro.

[0128] 6. Methods for producing genetically modified cells In another aspect, the present invention also provides a method of producing at least one genetically modified cell, comprising introducing a base editing system of the present invention into at least one said cell, thereby generating one or more nucleotide substitutions in a target nucleic acid region of said at least one cell. In some embodiments, the one or more nucleotide substitutions are C to T substitutions.

[0129] In some embodiments, the method further comprises screening for cells from said at least one cell that have one or more desired nucleotide substitutions.

[0130] In some embodiments, the methods of the invention are performed in vitro, e.g., the cell is an isolated cell, or a cell within an isolated tissue or organ.

[0131] In another aspect, the present invention also provides a genetically modified organism comprising a genetically modified cell produced by the method of the present invention or a progeny thereof, preferably said genetically modified cell or a progeny thereof having one or more desired nucleotide substitutions.

[0132] In the present invention, the target nucleic acid region to be modified may be located anywhere in the genome, such as in a functional gene, such as a gene encoding a protein, or may be located in a gene expression regulatory region, such as a promoter region or enhancer region. This achieves the modification of the gene function or the modification of the gene expression. In some embodiments, the desired nucleotide substitution results in the desired modification of the gene function or the gene expression.

[0133] In some embodiments, the target nucleic acid region is associated with a trait of the cell or organism. In some embodiments, a mutation in the target nucleic acid region results in a change in the trait of the cell or organism. In some embodiments, the target nucleic acid region is located in a coding region of a protein. In some embodiments, the target nucleic acid region encodes a motif or domain associated with a function of a protein. In some preferred embodiments, one or more nucleotide substitutions in the target nucleic acid region result in an amino acid substitution in the amino acid sequence of the protein. In some embodiments, the one or more nucleotide substitutions result in a change in the function of the protein.

[0134] In the methods of the present invention, the base editing system can be introduced into a cell by various methods known to those skilled in the art.

[0135] Methods that can be used to introduce the base editing system of the present invention into cells include, but are not limited to, calcium phosphate transfection, protoplast fusion, electroporation, lipofectamine transfection, microinjection, viral infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun techniques, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation, and the like.

[0136] Cells that can be base edited by the methods of the present invention may be derived from mammals, such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, and cats; poultry, such as chickens, ducks, and geese; and plants, including monocotyledonous and dicotyledonous plants, preferably crop plants, including, but not limited to, wheat, rice, corn, soybeans, sunflowers, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0137] 7. Use in plants The base editing fusion protein, base editing system, and method for producing genetically modified cells of the present invention are particularly suitable for genetic modification of plants.Preferably, the plant is a crop plant, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.More preferably, the plant is rice.

[0138] In another aspect, the present invention provides a method for producing a genetically modified plant, comprising introducing a base editing system of the present invention into at least one said plant, thereby causing one or more nucleotide substitutions in a target nucleic acid region in the genome of said at least one plant.

[0139] In some embodiments, the method further comprises screening for plants having one or more desired nucleotide substitutions from the at least one plant.

[0140] In the method of the present invention, the base editing composition can be introduced into a plant by various methods well known to those skilled in the art. Methods that can be used to introduce the base editing system of the present invention into a plant include, but are not limited to, gene gun method, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation, plant virus-mediated transformation, pollen tube channel method, and ovary injection method. Preferably, the base editing composition is introduced into a plant by transient transformation.

[0141] In the method of the present invention, modification of a target sequence can be achieved simply by introducing or producing the base editing fusion protein and guide RNA into a plant cell, and the modification can be stably inherited without stably transforming the plant with an exogenous polynucleotide encoding a component of the base editing system, thereby avoiding potential off-target effects of a stably present (continuously produced) base editing composition and also avoiding the integration of exogenous nucleotide sequences into the plant genome, thereby providing greater biological safety.

[0142] In some preferred embodiments, the introduction is performed in the absence of selective pressure, thereby avoiding integration of the exogenous nucleotide sequence into the plant genome.

[0143] In some embodiments, the introduction comprises transforming the base editing system of the present invention into an isolated plant cell or tissue, and then regenerating the transformed plant cell or tissue into an intact plant. Preferably, the regeneration is performed in the absence of selection pressure, i.e., without the use of any selection agent against the selection gene carried on the expression vector during tissue culture. The absence of a selection agent increases the efficiency of plant regeneration and allows for the production of modified plants that do not contain exogenous nucleotide sequences.

[0144] In some other embodiments, the base editing system of the present invention can be transformed into specific parts of an intact plant, such as leaves, shoot tips, pollen tubes, young spikes, or hypocotyls, which is particularly suitable for transforming plants that are difficult to regenerate by tissue culture.

[0145] In some embodiments of the invention, the in vitro expressed protein and / or the in vitro transcribed RNA molecule (e.g., the expression construct is an in vitro transcribed RNA molecule) is directly transformed into the plant, where the protein and / or RNA molecule can undergo base editing within the plant cell and is subsequently degraded by the cell, avoiding integration of the exogenous nucleotide sequence into the plant genome.

[0146] Thus, in some embodiments, the genetic modification and breeding of plants using the methods of the present invention can result in plants that are free of integration of exogenous polynucleotides into their genome, i.e., transgene-free modified plants.

[0147] In some embodiments of the invention, the modified target nucleic acid region is associated with a plant trait, such as an agronomic trait, such that the one or more nucleotide substitutions cause the plant to have an altered (preferably improved) trait, e.g., an agronomic trait, compared to a wild-type plant.

[0148] In some embodiments, the method further comprises screening plants for one or more desired nucleotide substitutions and / or desired traits, such as agronomic traits.

[0149] In some embodiments of the invention, the method further comprises obtaining a progeny of the genetically modified plant. Preferably, the genetically modified plant or its progeny has one or more desired nucleotide substitutions and / or desired traits, such as agronomic traits.

[0150] In another aspect, the present invention further provides a genetically modified plant or its progeny or part thereof obtained by the above method of the present invention. In some embodiments, the genetically modified plant or its progeny or part thereof is transgene-free. Preferably, the genetically modified plant or its progeny has a desired genetic modification and / or a desired trait, such as an agronomic trait.

[0151] In another aspect, the present invention further provides a plant breeding method comprising hybridizing a genetically modified first plant comprising one or more nucleotide substitutions in a target nucleic acid region obtained by the method of the present invention described above with a second plant not comprising said one or more nucleotide substitutions, thereby introducing said one or more nucleotide substitutions into the second plant. Preferably, said genetically modified first plant has a desired trait, such as an agronomic trait.

[0152] 8. Therapeutic Use The present invention further encompasses the use of the base editing system of the present invention in the treatment of disease.

[0153] By modifying a disease-associated gene using the base editing system of the present invention, it is possible to achieve upregulation, downregulation, inactivation, activation, correction of mutations, etc. of the disease-associated gene, thereby preventing and / or treating the disease. For example, the target nucleic acid region described in the present invention can be located in the protein coding region of the disease-associated gene or in a gene expression regulatory region such as a promoter region or enhancer region, thereby modifying the function or expression of the disease-associated gene. Thus, the modification of a disease-associated gene described herein includes not only the modification of the disease-associated gene itself (e.g., protein coding region), but also the modification of its expression regulatory region (e.g., promoter, enhancer, intron, etc.).

[0154] A "disease-associated" gene refers to any gene that produces a transcription or translation product at an abnormal level or in an abnormal form in cells derived from disease-affected tissues compared to non-disease control tissues or cells. If the change in expression is associated with the appearance and / or progression of a disease, it may be a gene that is expressed at an abnormally high level, or it may be a gene that is expressed at an abnormally low level. A disease-associated gene also refers to a gene that has one or more mutations, or that has a genetic variation that is directly involved in the pathogenesis of a disease, or that is in linkage disequilibrium with one or more genes that are involved in the pathogenesis of a disease. The mutation or genetic variation is, for example, a single nucleotide variation (SNV). The transcription or translation product may be known or unknown, and may be normal or abnormal in level.

[0155] Thus, the present invention further provides a method of treating a disease in a subject in need thereof, comprising delivering an effective amount of a base editing system of the present invention to the subject to modify a gene associated with the disease (e.g., deaminating mitochondrial DNA with a fusion protein or multiple fusion proteins). The present invention further provides a use of a base editing system in the manufacture of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the base editing system is used to modify a gene associated with the disease. The present invention further provides a pharmaceutical composition for treating a disease in a subject in need thereof, comprising the base editing system of the present invention and, optionally, a pharma- ceutically acceptable carrier, wherein the base editing system is used to modify a gene associated with the disease.

[0156] In some embodiments, the fusion protein or base editing system described herein is used to introduce a point mutation into a nucleic acid by deaminating a target nucleobase (e.g., a C residue). In some embodiments, deamination of the target nucleobase results in the correction of a genetic defect, such as the correction of a point mutation that results in loss of function of a gene product. In some embodiments, the genetic defect is associated with a disease or disorder (e.g., a lysosomal storage disease or a metabolic disease, e.g., type I diabetes). In some embodiments, the methods provided herein can be used to introduce an inactivating point mutation into a gene or allele that encodes a gene product associated with a disease or disorder.

[0157] In some embodiments, the purpose of the embodiments described herein is to restore the function of a dysfunctional gene by genome editing. The nucleobase editing proteins provided herein can be used for ex vivo gene editing of human cells, such as correcting disease-related mutations in human cell cultures. The nucleobase editing proteins provided herein, such as nucleic acid editing DNA proteins (e.g., CRISPR effector protein Cas9) and fusion proteins containing a cytosine deaminase domain, can be used to correct any T to C or A to G single point mutation. In the former case, the mutant C corrects the mutation by deamination, while in the latter case, the C paired with the mutant A corrects the mutation by deamination and subsequent replication.

[0158] In some embodiments, the objective of the present invention is to treat a patient having a disease associated with or caused by a point mutation that can be corrected by the DNA base editing fusion protein provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neoplastic disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.

[0159] In some embodiments, the purpose of the embodiments described herein is to treat mitochondrial diseases or disorders. As used herein, "mitochondrial disease" refers to diseases caused by abnormal mitochondria, such as mitochondrial gene mutations, enzyme pathways, etc. Examples of diseases include, but are not limited to, neurological disorders, loss of motor control, muscle weakness and pain, gastrointestinal disorders and swallowing difficulties, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, vision / hearing problems, lactic acidosis, developmental delays, and high susceptibility to infections.

[0160] Examples of diseases described in the present invention include, but are not limited to, genetic diseases, cardiovascular diseases, muscle diseases, brain, central nervous system diseases, immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism spectrum disorders, trinucleotide repeat expansion disorders, hearing disorders, gene targeted therapy of non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, beta thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, and schizophrenia. Other diseases that are treated by correcting point mutations or introducing inactivating mutations into disease-associated genes are known to those skilled in the art, and the present disclosure is not limited in this respect. In addition to the diseases exemplarily described in the present invention, the strategies and fusion proteins provided by the present invention can also be used to treat other related diseases, which uses will be apparent to those skilled in the art. For diseases or targets applicable to the present invention, refer to the relevant diseases to which the base editing systems described in International Publication No. 2015089465 (PCT / US2014 / 070135), International Publication No. 2016205711 (PCT / US2016 / 038181), International Publication No. 2018141835 (PCT / EP2018 / 052491), International Publication No. 2020191234 (PCT / US2020 / 023713), International Publication No. 2020191233 (PCT / US2020 / 023712), International Publication No. 2019079347 (PCT / US2018 / 056146), and International Publication No. 2021155065 (PCT / US2021 / 015580) are applicable.

[0161] Administration of the base editing system or pharmaceutical composition of the present invention can be adjusted to the weight and species of the patient or subject. The frequency of administration is within the range permitted by medical or veterinary medicine. It depends on general factors including the age, sex, general health, other conditions, and the specific medical condition or symptom being addressed of the patient or subject.

[0162] 9. Adeno-associated virus (AAV) The base editing fusion protein according to the present invention and / or an expression construct comprising a nucleotide sequence encoding said base editing fusion protein, or one or more gRNAs comprising the base editing system of the present invention can be delivered using adeno-associated virus (AAV), lentivirus, adenovirus, or other plasmid or viral vector types. As AAV has a packaging limit of 4.5-4.75 Kb. This indicates that both a promoter and a transcription terminator must be incorporated into the same viral vector. Constructs exceeding 4.5-4.75 Kb result in a significant decrease in viral delivery efficiency. Cytosine deaminase is difficult to package into AAV due to its large size. Thus, an embodiment of the present invention provides the use of a truncated cytosine deaminase for packaging into AAV to achieve base editing.

[0163] 10. Nucleic Acids, Cells and Compositions In another aspect, the present invention provides a nucleic acid molecule encoding a cytosine deaminase of the present invention or a fusion protein of the present invention.

[0164] In another aspect, the present invention provides a cell comprising a cytosine deaminase of the present invention, or a fusion protein of the present invention, or a base editing system of the present invention, or a nucleic acid molecule of the present invention.

[0165] In another aspect, the present invention provides a composition comprising a cytosine deaminase of the present invention, or a fusion protein of the present invention, or a base editing system of the present invention, or a nucleic acid molecule of the present invention.

[0166] In some embodiments, the cytosine deaminase, fusion protein, base editing system or nucleic acid molecule is packaged into a virus, virus-like particle, virion, liposome, vesicle, exosome, or liposomal nanoparticle (LNP).

[0167] In some embodiments, the virus is an adeno-associated virus (AAV) or a recombinant adeno-associated virus (rAAV).

[0168] 11. Kit The present invention further includes a kit for use in the method of the present invention, the kit comprising a base-editing fusion protein of the present invention and / or an expression construct comprising a nucleotide sequence encoding said base-editing fusion protein, or a base editing system of the present invention. The kit typically includes a label indicating the desired purpose and / or method of use of the contents of the kit. The term label includes any written or recorded material that is included on or provided with the kit or is otherwise provided with the kit. The kit of the present invention may further include suitable materials for constructing an expression vector in the base editing system of the present invention. The kit of the present invention may further include reagents suitable for transforming the base-editing fusion protein or base editing composition of the present invention into a cell.

[0169] In one aspect, the present invention provides a method for producing a method for treating a cancer cell comprising: (a) a nucleic acid sequence encoding the cytosine deaminase of the present invention; (b) a heterologous promoter driving expression of the sequence of (a).

[0170] In one aspect, the present invention provides a method for producing a method for treating a cancer cell comprising: (a) a nucleic acid sequence encoding a fusion protein of the invention; (b) a heterologous promoter driving expression of the sequence of (a).

[0171] In some embodiments, the expression construct further comprises an expression construct encoding a guide RNA backbone, said construct comprising a cloning site that allows for cloning of a nucleic acid sequence identical or complementary to a target sequence into said guide RNA backbone.

[0172] Working Example In order to facilitate understanding of the present invention, the present invention will be described in more detail below with reference to the relevant specific embodiments and the accompanying drawings. Preferred embodiments of the present invention are shown in the accompanying drawings. However, the present invention may be embodied in many different forms and is not limited to the embodiments set forth herein. Rather, the purpose of providing these embodiments is to provide a complete understanding of the disclosed subject matter of the present invention.

[0173] material and method 1. Vector Construction The identified novel deaminase sequence was constructed into pJIT63-nCas9-PBE backbone (Addgene No.=#98164) after double codon optimization for rice and wheat by Nanjing Jinsirui Co., Ltd. The plasmid of the reporter system used in the examples was constructed in our laboratory.

[0174] For sgRNA, the pOsU3 vector (Addgene No.=#170132) was used for expression.

[0175] 2. Protoplast Isolation and Transformation The protoplasts used in the present invention are derived from 11 medium-flowering rice varieties.

[0176] 2.1 Rice seedling cultivation First, rice seeds were washed with 75% ethanol for 1 minute, then treated with 4% sodium hypochlorite for 30 minutes, and washed with sterile water at least five times. They were placed on M6 medium and cultured at 26°C in the dark for 3 to 4 weeks.

[0177] 2.2 Isolation of protoplasts (1) Cut the rice stalk, cut the central part into 0.5 to 1 mm strips with a blade, place in a 0.6 M mannitol solution, and treat for 10 minutes in the dark. Then, filter the solution and place it in 50 mL of enzymatic hydrolysis solution (filtered through a 0.45 μm membrane), vacuum suction for 30 minutes (pressure approximately 15 Kpa), remove it, place it on a shaker (10 rpm), and perform enzymatic hydrolysis at room temperature for 5 hours.

[0178] (2) 30 to 50 mL of W5 was added to dilute the enzymatic hydrolysate, and the enzymatic hydrolysate was filtered into a round-bottom centrifuge tube (50 mL) using a 75 μm nylon filtration membrane.

[0179] (3) The mixture was centrifuged at 23°C and 250g (rcf) for 3 minutes at three speed levels, with three slower speed levels, and the supernatant was discarded.

[0180] (4) The cells were gently suspended in 20 mL of W5, and step (3) was repeated.

[0181] (5) An appropriate amount of MMG was added to the suspension and transformation was performed.

[0182] 2.3 Rice protoplast transformation (1) 10 μg of each of the required transformation vectors was added to a 2 mL centrifuge tube and mixed uniformly. Then, 200 μL of protoplasts were aspirated using a pipette tip with the sharp tip removed, and mixed uniformly by pipetting. 220 μL of PEG4000 solution was added, and mixed uniformly by pipetting. Transformation was induced for 20 to 30 minutes at room temperature in the dark.

[0183] (2) 880 μL of W5 was added, gently inverted to mix uniformly, and centrifuged at 250 g (rcf) for 3 minutes with 3 levels of increasing speed and 3 levels of decreasing speed, and the supernatant was discarded.

[0184] (3) 1 mL of WI solution was added, gently inverted to mix evenly, gently transferred to a flow tube, and incubated in the dark at room temperature for 48 hours.

[0185] 3. Observation of cell fluorescence using a flow cytometer Protoplast GFP-negative and positive populations were analyzed by flow cytometry on a FACSAria III (BD Biosciences) instrument.

[0186] 4. Protoplast and Plant DNA Extraction and Amplicon Sequencing Analysis The protoplasts were collected in a 2 mL centrifuge tube, protoplast DNA (approximately 30 μL) was extracted using the CTAB method, its concentration (30–60 ng / μL) was measured using a NanoDrop ultramicrospectrophotometer, and the DNA was stored at -20 °C.

[0187] PCR amplification of protoplast DNA templates was performed using genomic primers specific for the target site. The 20 μL amplification system contained 4 μL 5× Fastpfu buffer, 1.6 μL dNTPs (2.5 mM), 0.4 μL Forward primer (10 μM), 0.4 μL Reverse primer (10 μM), 0.4 μL FastPfu polymerase (2.5 U / μL), and 2 μL DNA template (approximately 60 ng). Amplification conditions: pre-denaturation at 95 °C for 5 min, denaturation at 95 °C for 30 s, annealing at 50–64 °C for 30 s, extension at 72 °C for 30 s, 35 cycles; fully extended at 72 °C for 5 min and stored at 12 °C.

[0188] The above amplification product was diluted 10-fold, and 1 μL was used as a template for the second PCR amplification. The amplification primer was a sequencing primer containing Barcode. The 50 μL amplification system contained 10 μL 5×Fastpfu buffer, 4 μL dNTPs (2.5 mM), 1 μL forward primer (10 μM), 1 μL reverse primer (10 μM), 1 μL FastPfu polymerase (2.5 U / μL) and 1 μL DNA template. The amplification conditions were as described above, and the number of amplification cycles was 35 cycles.

[0189] The PCR products were separated by 2% agarose gel electrophoresis, the gel was extracted from the target fragments using the AxyPrep DNA Gel Extraction kit, and the extracted products were quantitatively analyzed by NanoDrop ultra-microspectrophotometer. 100ng of the extracted products were mixed and sent to Nucleobase for amplicon sequencing library construction and amplicon sequencing analysis.

[0190] 5. Transfection of Human and Animal Cells Human HEK293T cells (ATCC, CRL-3216) and mouse N2a cells (ATCC, CCL131) were cultured in a humidified incubator at 37 °C with 5% CO2 in Dulbecco's modified Eagle's medium (DMEM, Gibco) supplemented with 10% (vol / vol) fetal bovine serum (FBS, Gibco) and 1% (vol / vol) penicillin-streptomyces (Gibco). All cells were routinely detected for mycoplasma contamination using a mycoplasma detection kit (Transgen Biotech). Cells were seeded on poly-D-lysine-coated 48-well plates (Corning) without antibiotics. After 16–24 h, cells were incubated with 1 uL of Lipofectamine 2000 (ThermoFisher Scientific), 300 ng of deaminase vector and 100 ng of sgRNA expression vector. For transfection with the cytosine base editing system, cells were incubated with 1uL Lipofectamine2000, 300ng TALE-L and 300ng TALE-R. After 72 hours, cells were washed with PBS and DNA was extracted. To detect off-target effects using the R-loop method, cells were co-transfected with the BE4max vector, SaCas9BE4max vector and the corresponding sgRNA vector (Koblan, LW, Doman, JL, Wilson, C., Levy, JM, Tay, T., Newby, GA, Maianti, JP, Raguram, A., & Liu, DR (2018). Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat. Biotechnol., 36, 843-846.).

[0191] 6.TRAPseq Library The performance of the deaminase base editing system was evaluated using the sgRNA 12K-TRAPseq library. 20 hours prior to viral transduction, 2 × 106 cells were seeded in 100 mm culture dishes. 500 μL of sgRNA lentivirus was transduced. Stably integrated cells were screened using 1 μg / mL puromycin (Gibco). For each base editor, 2 × 106 cells were seeded in six culture dishes 24 hours prior to transfection. Each CBE member was transfected with 15 μg of plasmid DNA and 15 μg of Tol2 DNA using 60 μL of Lipofectamine 2000. 24 hours after transfection, the medium was replaced with fresh medium containing 10 μg / mL blasticidin (Gibco). After 3 days, the cells were washed, resuspended, and seeded in blasticidin medium at a concentration of 10 μg / mL. After 6 days, all cells were collected by washing with PBS, centrifuged, and DNA was extracted using the Cell / Tissue DNA Isolation Mini Kit (Vazyme). Each deaminase base editor sample was sequenced by next-generation sequencing.

[0192] 7. DNA Extraction Genomic DNA from HEK293T cells and N2a cells was extracted by treatment with Lysis Buffer and Proteinase K and then with the Triumfi Mouse Tissue Direct Amplification Kit (Beijing Jinsha Biology).

[0193] Plant protoplasts were cultured for 72 h, and then genomic DNA was extracted using a Plant Genomic DNA Kit (Tiangen Biochemical Technology). All DNA samples were quantified using a NanoDrop 2000 spectrophotometer (Thermo Fisher).

[0194] 8. Protein structure analysis and clustering Protein structure analysis was performed using AlphaFold v2.2.0 (John Jumper and others, 'Highly Accurate Protein Structure Prediction with AlphaFold', Nature, 596.7873 (2021), 583-89).

[0195] The TM-score was calculated from the analysis results using TM-align software. The specific calculation formula for the TM-score is as follows (Reference: Zhang, Yang, and Jeffrey Skolnick. (2004). Scoring function for automated assessment of protein structure template quality. Proteins 57(4), 702-710.).

[0196]

number

[0197] After converting the TM-scores, we used the APE and phangorn packages in the R language to perform clustering calculations using the UPGMA method (CP Kurtzman, Jack W. Fell, and T. Boekhout, The Yeasts: A Taxonomic Study, 5th ed (Amsterdam: Elsevier, 2011).; 'A Statistical Method for Evaluating Systematic Relationships-Robert Reuven Sokal, Charles Duncan Michener-Google Books'). First, we obtained the distance between any two pairs using the following formula:

[0198]

number

[0199] The formula for calculating the average distance used in the clustering process is as follows: If C1, C2 are terminal classification clusters containing sets n1 and n2, respectively, that are to be merged into a new set C, then the average distance to any other cluster D is calculated by the following formula:

[0200]

number

[0201] Example 1. Identification of novel deaminases in the APOBEC / AID clade that can be used for base editing To find a novel deaminase that is different from the deaminases used in existing base editing systems, we first tested deaminases of the APOBEC / AID clade that have low sequence similarity to existing deaminases with the list of representative deaminases listed in the study by Iyer et al. (Iyer, LM, Zhang, D., Rogozin, IB, & Aravind, L. (2011). Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems. Nucleic acids research, 39 (22), 9473-9497.). Among them, deaminase No. 182 (SEQ ID NO: 1) has very low similarity to existing deaminases, and its amino acid sequence has only 34% sequence identity with the most similar mouse rAPOBEC1. Deaminase No. 182 was constructed in the pJIT163-nCas9-PBE backbone. In other words, deaminase No. 182 was used instead of rAPOBEC1 and fused to nCas9. Evaluation of the reporter system revealed that 182-PBE could undergo base editing in cells (Figures 1, 6, and 7).

[0202] To further confirm its editing ability, the 182-PBE construct and the targeting endogenous sgRNA construct were co-transformed into rice protoplasts. Analysis of the editing results of six endogenous sites showed that 182-PBE could effectively realize base editing, and its editing window was significantly larger than that of the commonly used rAPOBEC1-based cytosine base editing system (Figure 2, Figure 6, and Figure 7). Therefore, protein No. 182 has the function of deaminating cytosine on single-stranded DNA, and a novel cytosine base editing system can be established based on this protein.

[0203] Example 2. Detection of cytosine deamination activity of deaminases in different clades Iyer et al. searched the database for proteins with folding modes similar to known deaminases and divided the above proteins into at least 21 clades based on their domains (Table 1). The cytosine deaminases APOBEC1, APOBEC3, AID and CDA1, which are currently widely used for base editing, all belong to the APOBEC / AID-like clade. In addition to the above clades, there are also clades with proven functions, such as the "dCMP deaminase and ComE" clade that can convert dCMP to dUMP, the "guanine deaminase" clade that can convert guanine (G) to xanthine (I), the "RibD-like" clade with diaminohydroxyphosphoribosylaminopyrimidine deaminase function, the "Tad1 / ADAR" clade with RNA editing enzyme function that converts RNA adenine (A) to xanthine (I), and the "PurH / AICAR transformylase" clade with formyltransferase activity. Furthermore, the deamination functions of some bacterial clades named based on protein domains, such as the SCP1.201, XOO2897, MafB19, and Pput_2613 clades, remain to be elucidated.

[0204] Table 1. Deaminase classification families (Iyer et al., 2011) [Table 1]

[0205] To detect whether the above clades have cytosine deaminase activity, a total of 48 deaminase proteins distributed in 14 clades, including Bd3614, CDD / CDA-like, DYW-like, FdhD, MafB19, novel AID / APOBEC-like, OTT1508, PurH / AICAR transformylase, RibD-like, TM1506, SCP1.201, Imm1 immunity protein related to SCP1.201 deaminase, YwqJ and XOO2897, except for APOBEC / AID clade, were selected from the representative deaminase list listed in Iyer et al. All proteins were constructed into pJIT163-nCas9-PBE backbone and their activity of binding and deaminating single-stranded DNA was evaluated by BFP-to-GFP reporter system (Zong, Y. et al. Nat. Biotechnol. 35, 438-440 (2017)). As a result, a total of 23 proteins from five clades were found to have cytosine deaminase activity, which were derived from a novel AID / APOBEC-like clade (No. 2-1479, and No. 2-1478), as well as the bacterial SCP1.201 clade (No. 69, No. 55, No. 57, No. 64, No. 76, No. 2-1146, No. 2-1160, No. 54, No. 56, No. 59, No. 60, No. 61, No. 72, No. 74, No. 75, No. 63, No. 2-1158), XOO2897 clade (No. 2-1429, No. 2-1442), TM1506 clade (No. 2-39), and MafB19 clade (No. 101m). In particular, in the SCP1.201 clade, cytosine deaminase activity was detected in 18 of the 19 proteins tested. By testing the two endogenous sites in rice, the total of 23 proteins with cytosine deaminase activity listed above can be classified into eight high-editing efficiency deaminases (Figures 4 and 5), eight medium-editing efficiency deaminases (Figures 6 and 7), and seven low-editing efficiency deaminases.

[0206] To further confirm the editing ability of the newly discovered deaminase, we selected and tested candidate deaminase No. 69 from the group that could illuminate the reporter system. This protein belongs to the SCP1.201 clade. To further confirm its editing ability, we co-transformed the 69-PBE construct and the targeting endogenous sgRNA construct into rice protoplasts. Analysis of the editing results of six endogenous sites revealed that 69-PBE could effectively achieve base editing, and its editing efficiency was significantly higher than that of the commonly used rAPOBEC1-based cytosine base editing system (Figure 3). Therefore, the newly identified proteins can have the function of deaminating cytosine on single-stranded DNA, and a new cytosine base editing system can be established based on these proteins.

[0207] Example 3. Protein structural analysis, clustering, and discovery of a novel cytosine deaminase Based on the above examples, it is necessary to propose an effective protease function identification and screening method to efficiently analyze protein functions. Because the three-dimensional structure of a protein has a crucial effect on its function, comparative analysis and classification clustering of known or predicted protein structures may be an effective method to classify deaminases into functional clades. Therefore, we combined AI-assisted protein structure prediction, structure calibration, and clustering to generate a new protein classification relationship among deaminases (Figure 8).

[0208] From the InterPro database, we selected 238 protein sequences annotated as containing deaminase domains and four outgroup candidate protein sequences from the JAB domain family (Figure 9). Specifically, we selected 15 candidate genes with a length of at least 100 amino acids from each of the 16 deaminase families, and predicted their protein structures using AlphaFold2. We performed multiple structural alignments (MSA) for all candidate proteins using the standardized scoring model TM-score. The specific calculation formula for TM-score is as follows (Reference: Zhang, Yang, and Jeffrey Skolnick. (2004). Scoring function for automated assessment of protein structure template quality. Proteins 57(4), 702-710.):

[0209]

number

[0210] Based on the MSA results, we generated structural similarity matrices that reflect the overall structural relatedness among proteins. We then condensed these similarity matrices into structure-based dendrograms (Figure 10) using the unweighted pair group method with arithmetic average (UPGMA). In the dendrogram, 238 proteins were clustered into 20 unique structural clades, and deaminases in each clade have distinct conserved protein domains (Figure 11A and Figure 11B). We found that accurate protein cluster classifications can be generated based on protein structure without using context information such as conserved gene neighborhoods or domain architectures. When using structure-based hierarchical clustering, different clades reflected unique structures, implying that they have distinct catalytic functions and properties (Figure 11A and Figure 11B). Interestingly, we also found that this structure-based clustering method was more effective in ranking functional similarities than traditional one-dimensional amino acid sequence clustering methods. For example, adenine deaminase (A_deamin, PFO2137 in the InterPro database), which is involved in purine metabolism, is classified into different clades by clustering methods based on amino acid sequence, whereas it is classified into the same deaminase clade by clustering methods based on structure.

[0211] Furthermore, using the structure-based clustering method, four deaminase families (dCMP, MafB19, LmjF365940 and APOBEC (annotated by InterPro)) were classified into two independent clades (Figure 11A and Figure 11B). As a result of comparing the protein structures, the two clades of these four deaminase families have completely different structures, which may be opposite to their InterPro naming and sequence-based classification (Figure 11B and Figure 12). In other words, artificial intelligence-assisted protein clustering classification based on protein three-dimensional structure provides reliable clustering results and is a more convenient and effective protein relationship generation strategy than other methods because only one amino acid sequence is required and no other genomic inference is required.

[0212] Example 4: Identification of the functions of deaminases of the SCP1.201 clade that can be used for base editing using 3D structural trees As a result of evaluating the function of deaminases in each clade in Example 2, we surprisingly found that some deaminases in the SCP1.201 clade have the ability to catalyze the deamination of single-stranded DNA substrates. Previously, these deaminases were annotated as double-stranded DNA deaminase toxin A-like (DddA-like) deaminases in the InterPro database (PF14428). Of these, the DddA enzyme is a deaminase that has recently been used in non-CRISPR double-stranded DNA cytosine base editors (DdCBEs) and can be used to deaminate double-stranded DNA cytosine bases (NCBI Reference Sequence: WP_006498588.1) (BYMok, MHde Moraes, J.Zeng, DEBosch, AVKotrys, A.Raguram, F.Hsu, MCRadey, SBPeterson, VKMootha, JDMougous, DRLiu, A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature 583, 631-637 (2020).). Indeed, the presence of DddA led to all proteins in the SCP1.201 clade to which it belongs being annotated as double-stranded DNA deaminases (Ddd).

[0213] In light of this issue, we used the function prediction based on the three-dimensional structure of the protein in Example 3 to carry out the following work. To reanalyze this SCP1.201 clade, we selected all 489 SCP1.201 deaminases from the InterPro database. We also included seven other proteins that were found to have 35%-50% similarity to DddA by BLAST alignment but were listed separately in InterPro. After recognition and coverage screening, we performed a new artificial intelligence-assisted protein structural classification for the 332 SCP1.201 deaminases. As a result of the structural cluster analysis, the SCP1.201 deaminases are clustered into different subclades, each with its own core domain motif (Figures 13A-E).

[0214] Importantly, DddA and 10 other proteins are clustered in the same subclade of SCP1.201. By analyzing the 3D predicted structures of all 11 proteins in this subclade, we found that they have a core structure similar to DddA. Given the structural similarity to DddA, we predicted that the other proteins in this subclade also have the function of double-stranded DNA cytosine deamination.

[0215] Example 5. Verification of deaminases including the DddA subclade in base editing in animal cells To evaluate whether the SCP1.201 candidate proteins of the subclade in which DddA exists, obtained using the prediction method of the present invention in Example 4, have functional similarity to DddA, i.e., whether they have a deaminating effect on dsDNA, we designed DdCBEs formed by each deaminase of this subclade individually or by splitting the deaminase-like DddA structure into two parts at the allelic sites of two residues and linking them with the dual TALE system. (For the method, see BY Mok, MH de Moraes, J. Zeng, DE Bosch, AV Kotrys, A. Raguram, F. Hsu, MC Radey, SB Peterson, VK Mootha, JD Mougous, DR Liu, A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature 583, 631-637 (2020)) (Figure 14, Table 2). We assessed proteins from this Ddd subclade at JAK2 and SIRT6 sites in HEK293T cells and observed 13 proteins capable of dsDNA base editing (Table 2). Hereafter, we name these deaminases double-stranded DNA deaminases (Ddd) and classify them into this newly discovered Ddd subclade.

[0216] Table 2. Catalytic activities of proteins from the subclade in which DddA resides [Table 2] The symbol ++ denotes high catalytic activity, + denotes low catalytic activity, and - denotes no catalytic activity.

[0217] Example 6. Validation of deaminases that do not contain the DddA subclade in base editing in plant and animal cells For comparison, in this experiment, we further evaluated the deamination activity of other SCP1.201 candidate proteins that do not contain the DddA subclade. From there, 24 proteins were randomly selected and placed in the CBE fluorescent reporter system. As a result, 22 of these proteins showed detectable fluorescence. Of these, 13 proteins were selected to evaluate base editing at endogenous sites under CBE conditions in mammalian cells (Figure 16A, Table 3). Although these proteins were previously annotated as DddA-like proteins, our experimental results showed that these proteins showed cytosine base editing activity only on ssDNA (Figure 13A, Figure 16A and Table 3), but not on dsDNA (Figure 16B). Based on these functions and actions, we named these proteins from the SCP1.201 clade with ssDNA-targeting action as single-stranded DNA deaminases (Sdd) in future work.

[0218] According to the above experimental results, we surprisingly found that most of the protein members from the SCP1.201 clade are Sdd proteins, not DddA-like proteins annotated in the InterPro database (PF14428). We also found that these Sdd proteins are similar to each other and clearly distinguished from the structure of Ddd proteins, as shown for example in the structure of Sdd7 (Figure 13D, Figure 13E). Sdd7 is one of the cytosine base editors with the highest ssDNA editing efficiency. Therefore, the method of the present invention indicates that the DddA-like deaminases annotated in the InterPro database (PF14428) should be further subdivided and appropriately re-annotated.

[0219] As a control, we further clustered proteins from the SCP1.201 clade based on their one-dimensional amino acid sequences and validated the structural tree with the JAB outgroup, finding that members of the JAB outgroup were distributed throughout the tree. These results demonstrate the validity and importance of using structure-based classification of proteins to compare and evaluate protein relationships.

[0220] Table 3. Catalytic activities of proteins from the DddA-free subclade [Table 3] The symbol ++ denotes high editing preference, + denotes low editing preference, and - denotes no editing preference.

[0221] In addition, the verification results of the protein functions in Examples 5 and 6 were comprehensively analyzed. The results show that artificial intelligence-assisted protein clustering classification based on the three-dimensional structure of the protein provides reliable clustering results, and the three-dimensional structural tree constructed using the method of the present invention can accurately identify and predict the detailed functions of proteins. In addition, since only one amino acid sequence is required and no other genome inference is required, it is a more convenient and effective protein relationship generation strategy than other methods. In the three-dimensional structural tree, when the clustering condition is a TM-score of 0.7 or more, the prediction result of the catalytic function of the protein by the clustering result is consistent with the conclusion of the experimental verification (Table 4). That is, clustering is performed with the annotation that the TM-score by TM scoring with the reference protein is 0.7 or more, and the obtained subclade has the same or similar catalytic function as the reference protein, and the method of the present invention has significantly improved the efficiency of identification and prediction.

[0222] Table 3. Catalytic activities of proteins from the DddA-free subclade [Table 4]

[0223] Example 7. Different editing preferences of new Ddd proteins and DddA DddA has a strict preference for 5'-TC motifs, so the use of DddA-based dsDNA base editors is primarily limited to TC targets (BY Mok, MH de Moraes, J. Zeng, DE Bosch, AV Kotrys, A. Raguram, F. Hsu, MC Radey, SB Peterson, VK Mootha, JD Mougous, DR Liu, A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature 583, 631-637 (2020).). The recently evolved DddA11 shows more general applicability and can be used to deaminate 5'-HC (H = A, C or T) motifs to achieve cytosine base editing, but the editing efficiency for AC, CC and GC targets still needs to be improved (BYMok, AVKotrys, A.Raguram, TP Huang, VKMootha, DRLiu, CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat.Biotechnol.40,1378-1387). The newly discovered Ddd proteins of the present invention were evaluated to determine whether they could enhance the efficiency and expand the targeting scope of DdCBE. Thirteen deaminases belonging to the Ddd subclade were constructed into DdCBE and the status of dsDNA base editing at endogenous JAK2 and SIRT6 sites in HEK293T cells was evaluated (Figure 15, Figure 18 and Table 2). Interestingly, Ddd1, Ddd7, Ddd8 and Ddd9 were found to have similar or higher editing efficiency compared to DddA (Figure 17A and Figure 18). Importantly, Ddd1 and Ddd9 have much higher editing activity against 5'-GC motifs than DddA (Figure 17A and Figure 18).It is noteworthy here that at the C10 (5'-GC) residue of JAK2 and the C11 (5'-GC) residue of SIRT6, the editing ratios of DddA were only 21.1% and 0.6%, respectively, whereas the editing ratios of Ddd9 were 65.7% and 45.7%, respectively (Figure S17A).

[0224] Some Ddd proteins appear to show different editing patterns compared to DddA. Therefore, we decided to evaluate the motif preferences of all of these Ddd proteins. First, we constructed several plasmids (BYMok, AVKotrys, A.Raguram, TP Huang, VKMootha, DRLiu, CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol. 40, 1378-1387) that encode JAK2 target sequences. We also constructed plasmids (BYMok, AVKotrys, A.Raguram, TP Huang, VKMootha, DRLiu, CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol. 40, 1378-1387) that encode JAK2 target sequences. We also constructed plasmids (BYMok, AVKotrys, A.Raguram, TP Huang, VKMootha, DRLiu, CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol. 40, 1378-1387) that encode JAK2 target sequences. We also constructed plasmids (BYMok, AVKotrys, A.Raguram, TP Huang, VKMootha, DR Li ... C C to M C The 9th to 11th positions of N (M / N = A, T, C, and G) were changed to obtain 16 different plasmids, and each plasmid was co-transfected with its respective DdCBE variant (Figure 17B). C After comparative analysis of the C·G-to-T·A transversion frequencies of N, we generated corresponding motif logo plots reflecting the sequence context preferences of each dsDNA deaminase (Figure S17C). As previously described, we found that DddA and its structural homolog Ddd7 strongly prefer the 5'-TC motif (Figure S17C, Figure S19). On the other hand, Ddd1 and Ddd9 prefer the 5'-G motif (Figure S17C, Figure S19). C The motif of Ddd8 is a 5'-W C We found that Ddd proteins tend to edit substrates with (W=A or T) motifs. Thus, through deeper analysis of the Ddd subclade, we discovered a series of new Ddd proteins that can be used to edit different motifs. These proteins significantly expand the targeting scope and utility of DdCBEs, showing great application potential (Figure 17C, Figure 19).

[0225] Example 8. Base editing of Sdd deaminase in human cells and plants Next, we decided to check whether the newly discovered Sdd proteins could also be used for more precise or efficient base editing. To this end, we selected and evaluated the six most active Sdds and four weaker Sdds, and compared their activities using a fluorescent reporter system (Table 3). We designed plant CBEs for each of the ten Sdds and evaluated their endogenous base editing at six sites in rice protoplasts (Figures 20 and 21). We found that seven of the deaminases (Sdd7, Sdd9, Sdd5, Sdd6, Sdd4, Sdd76, and Sdd10) had higher activity than rat APOBEC1 (rAPOBEC1)-based CBEs. The cytosine base editing ratio of the most active Sdd7 base editor was as high as 55.6%, which is more than 3.5-fold higher than that of rAPOBEC1.

[0226] To validate the versatility of these deaminases, we also constructed corresponding human cell-targeting BE4max vectors (LW Koblan, JL Doman, C. Wilson, JM Levy, T. Tay, GA Newby, JP Maianti, A. Raguram, DR Liu, Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat. Biotechnol. 36, 843-846 (2018).) and evaluated their editing efficiency in three endogenous targets in HEK293T cells. The results in HEK293T cells were consistent with those in rice, with Sdd7 found to have the highest editing activity (Figure 22).

[0227] Previously, we found that human APOBEC3A (A3A) has a large editing window in plants and exhibits high editing activity (Y. Zong, Q. Song, C. Li, S. Jin, D. Zhang, Y. Wang, J.-L. Qiu, C. Gao, Efficient C-to-T base editing in plants using a fusion of nCas9and human APOBEC3A. Nat. Biotechnol. 36, 950-953 (2018)., Q. Lin, Z. Zhu, G. Liu, C. Sun, D. Lin, C. Xue, S. Li, D. Zhang, C. Gao, Y. Wang, J.-L. Qiu, Genome editing in plants with MAD7 nuclease. J. Genet. Genomics 48, 444-451 (2021).). Therefore, we compared the editing activities of A3A and Sdd7 in human cells (Figure 22) and plants (Figure 23). Interestingly, Sdd7 had comparable editing activity to A3A at all three target sites in HEK293T cells (Figure 22) and at five endogenous sites in rice protoplasts (Figure 23). These results confirm that Sdd7 is a powerful cytosine base editor that can be universally applied in plant and human cells.

[0228] Example 9. Unique base editing properties of Sdd proteins When assessing endogenous base editing, distinct editing patterns of different Sdd-CBEs were observed at genomic target sites tested in human and rice cells. For example, Sdd7, Sdd9, and Sdd6 did not show any specific motif editing preference, whereas Sdd3 seemed to prefer editing 5'-GC and 5'-AC motifs and significantly dislike editing 5'-TC and 5'-CC motifs (Figure 24). To better analyze the editing patterns of each deaminase, we used Targeted Reporter Anchored Positional Sequencing (TRAP-seq), a high-throughput method for parallel quantification of base editing outcomes. (Xi Xiang, Kunli Qu, Xue Liang, Xiaoguang Pan, Jun Wang, Peng Han, Zhanying Dong, Lijun Liu, Jiayan Zhong, Tao Ma, Yiqing Wang, Jiaying Yu, Xiaoying Zhao, Siyuan Li, Zhe Xu, Jinbao Wang, Xiuqing Zhang, Hui Jiang, Fengping Xu, Lijin Zou, Huajing Teng, Xin Liu, Xun Xu, Jian Wang, Huanming Yang, Lars Bolund, George M. Church, Lin Lin, Yonglun Luo. (2020). Massively parallel quantification of CRISPR. Editing in cells by TRAP-seq enables better design of Cas9, ABE, CBE gRNAs of high efficiency and accuracy. bioRxiv 2020.05.20.103614). A 12K TRAP-seq library containing 12,000 TRAP constructs, each containing a unique gRNA expression cassette and corresponding alternative target sites, was stably integrated into HEK293T cells by lentiviral transduction.Following cell culture and antibody selection, base editors were transiently transfected into this 12K-TRAP cell line, followed by puromycin and blasticidin selection for 10 days (Figure 25A). At day 11 post-transfection, genomic DNA was extracted and deep amplicon sequencing was performed to evaluate the editing products of each deaminase (Figure 25A). We found that Sdd7 and Sdd6 did not show high sequence context preference, while rAPOBEC1 has a high preference for 5'-TC and 5'-CC bases, while showing no interest in 5'-GC and 5'-AC bases (Figure 25B). On the other hand, Sdd3 showed a perfectly complementary pattern, preferring the editing of 5'-GC and 5'-AC bases, but showing little activity against 5'-TC and 5'-CC bases (Figure 25B). Interestingly, we found that Sdd6 and Sdd3 have distinct editing windows compared to rAPOBEC1 and Sdd7, with a stronger preference for editing the +1 to +3 positions on the distal PAM side (Figure 25B). In summary, the newly recognized Sdd base editors exhibit unique base editing properties compared to conventional cytosine base editors, including enhanced editing efficiency, distinct deamination preferences, and altered editing windows.

[0229] Example 10. High-fidelity editing properties of Sdd proteins Previous reports that CBE may result in genome-wide Cas9-based off-target editing results have raised concerns about the safety of these high-precision genome editing techniques in clinical applications. We believe that these off-target mutations may be the result of overexpression of cytidine deaminases. We decided to confirm whether the newly discovered Sdd proteins could provide a more favorable balance between off-target and on-target editing. Therefore, we evaluated the Cas9-independent off-target effects of the 10 Sdds using an orthogonal R-loop assay in rice protoplasts. We found that six of the ten deaminases (Sdd2, Sdd3, Sdd4, Sdd6, Sdd10 and Sdd59) had lower off-target activity than rAPOBEC1. Interestingly, although Sdd6 showed little off-target editing activity, it still had strong on-target base editing ability when tested at six endogenous sites in rice and human cells (Figure 26A and Figure 27). The on-target:off-target ratios of these 10 deaminases were analyzed, and Sdd6 showed the highest on-target:off-target editing ratio, which was 37.6-fold higher than that of APOBEC1 (Figure 26B). Furthermore, we compared the on-target and off-target editing of Sdd6 and rAPOBEC1 and its two high-fidelity deaminase variants, YE1 and YEE, in HEK293T cells. The important thing here is that Sdd6 has the highest on-target:off-target editing ratio. The calculation results showed that it was 2.8-fold, 2.1-fold, and 2.5-fold higher than rAPOBEC1, YE1, and YEE, respectively (Figure 26B, Figure 26C, and Figure 28), and 10.4-fold higher than hA3A (Figure 26C and Figure 28). Note that the on-target activity of Sdd6 was comparable to that of rAPOBEC1 and much higher than that of YE1 and YEE (Figure 28). We therefore determine that the SCP1.201 clade contains unique and more accurate Sdd proteins that can be used as high-fidelity base editors.

[0230] Example 11. Rational design of Sdd proteins with the aid of AlphaFold2 structure prediction The use of viruses for CBE delivery has enormous potential in disease treatment, but the large size of APOBEC / AID-like deaminases limits their ability to be packaged into a single adeno-associated virus (AAV) particle for in vivo editing applications (31). Other researchers have developed dual AAV strategy delivery methods, splitting CBEs into amino- and carboxyl-terminal fragments and packaging them into single AAV particles. However, such avenues of delivery have challenges in mass production capacity and higher viral loads, and there are potential safety concerns for their use. Recently, single AAV packaged CBEs were developed using truncated lamprey CBEs based on CDA-1, but these vectors showed little editing activity in HEK293T cells. The standard compactness and conservation of SCP1.201 deaminases suggest that they may be ideal proteins for the development of single AAV packaged CBEs. The present invention attempted to further engineer and shorten the size of the newly discovered Sdd proteins by artificial intelligence-assisted three-dimensional structural protein modeling.

[0231] First, we compared the AlphaFold2 predicted structures for all active Sdd deaminases and found that they had a conserved core structure (Figure S13D, S13E, and S13F). Next, we generated multiple truncated variants of Sdd7, Sdd6, Sdd3, Sdd9, Sdd10, and Sdd4, respectively, and tested the endogenous base editing of these variants at two sites in rice protoplasts. As a result, we found that mini-Sdd7, mini-Sdd6, mini-Sdd3, mini-Sdd9, mini-Sdd10, and mini-Sdd4 are new minimized deaminases. All of them are very small (approximately 130-160 aa) and have editing efficiencies equal to or higher than those of the full-length protein in rice protoplasts and human cells (Figure 30A). It is noteworthy here that all mini-deaminases enabled the construction of single AAV-packaged SaCas9-based CBEs (<4.7 kb) (Figure 30B). We constructed a SaCas9 vector packaged in a single AAV using mini-Sdd6, and found that its editing efficiency was about 60% at two sites in the HPD gene (4-hydroxyphenylpyruvate dioxygenase in mouse neuroblastoma N2a cells by transient transfection (Figure 30C). These results indicate that Sdd protein has a great advantage over APOBEC / AID deaminase in AAV-based CRISPR base editing delivery. The success of further shortening Sdd protein for AAV packaging indicates that it is highly advantageous for the three-dimensional structure-based protein function prediction method of the present invention.

[0232] Example 12. Base editing capabilities of the new Sdd-based CBE Next, we explored the use of the new Sdd engineered proteins in base editing in plants. First, we evaluated the feasibility of using mini-Sdd7 in Agrobacterium-mediated genome editing in rice. More rice-positive plant lines were observed with mini-Sdd7-based CBEs compared to CBEs based on human A3A (hA3A), which is the most commonly used in agricultural applications. Furthermore, a higher number of edited plants and higher editing efficiency were observed, reflecting higher CBE efficiency and lower toxicity than hA3A (Figure 31).

[0233] Soybean is one of the most important crops cultivated worldwide and is an important source of vegetable oil and protein. Although base editing has been demonstrated in soybean, most test sites in soybean crops are still difficult to edit and have low editing efficiency. To understand whether the newly developed Sdd-based CBE produces better cytosine base editing effects in soybean, we constructed a vector with the AtU6 promoter driving sgRNA expression and the CaMV 2x35S promoter driving CBE expression, and evaluated transgenic soybean hairy roots after Agrobacterium-mediated transformation (Figure 32). We found that APOBEC / AID deaminase had low editing activity at all five sites evaluated, including two sites, GmALS1-T2 and GmPPO2, which are particularly difficult to edit by other CBEs in soybean (Figure 30D). In addition, the cytosine base editing levels of mini-Sdd7 at five sites were 26.3-fold, 28.2-fold, and 10.8-fold higher than those of other deaminases rAPOBEC1, hA3A, and hAID, respectively, with a high editing efficiency of 67.4% (Figure 30D). Therefore, by focusing on these newly discovered Sdd proteins, we can break through the limitations of efficient cytosine base editing in soybean crops.

[0234] Next, we attempted to use mini-Sdd7 for base editing to obtain Agrobacterium-mediated transgenic soybean plant lines. The endogenous GmPPO2 gene was edited to generate the R98C mutation, which resulted in soybean plant lines with resistance to carfentrazone-ethyl. We obtained 77 transgenic soybean seedlings from three independent transformation experiments, of which 21 were heterozygous for base editing (Figure 30E, Figure 30F). It could be notably observed that after 10 days of treatment with carfentrazone-ethyl, the wild-type plant lines were sensitive to wilting and could not take root, while the mutant plant lines grew well and normally (Figure 30G). The development of an efficient cytosine base editor for use in soybean plants will enable various applications in the future.

[0235] Example 13. Base editing properties of proteins from other families In addition to the detailed classification and validation of the proteins of the SCP1.201 family, the deaminase functions and preferences of other families in the Iyer deaminase classification family (Table 1) were also verified. According to the analysis and validation method, a series of deaminases with similar Sdd activity (specific deaminases are shown in Table 5) were found in other families such as MafB19, AID / APOBEC, novel AID / APOBEC-like, TM1506, toxin deaminase, and XOO2897. Taking the MafB19 family as an example, in Example 2, it was found that some of the proteins in the MafB19 clade (No. 101m) have the function of single-stranded deaminase. In addition, in Example 3, based on the artificial intelligence-assisted clustering classification of the three-dimensional structure of proteins, it was found that there are two clades in the MafB19 deaminase family with completely different structures (Figure 11B and Figure 12). By applying the deaminase screening and identification method of the present invention, we found that three proteins of the MafB19 family (No. 2-1241, No. 2-1231 and No. 99) also have Sdd catalytic activity, and obtained their base editing sequence motif preferences (Table 5). The various novel cytosine base deaminases with different editing properties selected by the method of the present invention have enriched the base editing means, expanded the base editing system, and enhanced the ability to precisely manipulate target DNA sequences.

[0236] Table 5. Catalytic and editing properties of proteins from different families [Table 5] The symbol ++ denotes high editing preference, + denotes low editing preference, and - denotes no editing preference.

[0237] Conclusion of the experiment Conventional deaminase-based CBEs have the disadvantages of low editing efficiency, small editing window, significant preference, etc. Using the three-dimensional structure-based protein function prediction method of the present invention, a series of novel cytosine deaminases were obtained.

[0238] These cytosine deaminases have been proven to have great application potential and uses. For example, in Agrobacterium-mediated transformed transgenic soybean hairy roots, APOBEC / AID deaminases were found to have low editing activity at all five sites evaluated, including two sites, GmALS1-T2 and GmPPO2, which are particularly difficult to edit by other CBEs in soybean. Compared with rAPOBEC1, hA3A and hAID, mini-Sdd7 showed 26.3-, 28.2- and 10.8-fold higher cytosine base editing levels at the five sites, respectively, with a high editing efficiency of 67.4%. Therefore, we focused on utilizing these newly discovered Sdd proteins to break through the limitations of efficient cytosine base editing in soybean crops. We then attempted to use mini-Sdd7 for base editing to obtain Agrobacterium-mediated transgenic soybean plant lines. The endogenous GmPPO2 gene was edited to generate the R98C mutation, which resulted in a soybean plant line with resistance to carfentrazone-ethyl. Two heterozygotes for base editing were obtained from 30 transgenic soybean seedlings. It was notable that after 10 days of treatment with carfentrazone-ethyl, the wild-type plant line was sensitive to wilting and failed to root, while the mutant plant line grew well and normally. The development of efficient cytosine base editors for use in soybean plants will enable a variety of applications in the future. We believe that in the future, sequencing efforts in parallel with structure prediction will lead to great advances in discovering, tracking, classifying, and designing functional proteins. Currently, only a few cytosine deaminases are used as cytosine base editors. Standard efforts based solely on protein engineering and directed evolution can help diversify editing properties, but these efforts are often difficult to establish. A clustering prediction method based on three-dimensional structures was used to discover and analyze a series of deaminases with distinct properties. For example, of the newly discovered deaminases, both Sdd7 and Sdd6 were found to have great potential for therapeutic and agricultural applications.Sdd7 has strong base editing capabilities in all species tested, with higher editing activity than the most commonly used APOBEC / AID-like deaminases. Surprisingly, we found that Sdd7 is capable of efficient editing in soybean plant lines where cytosine base editing has been previously difficult (plant genes usually have a high GC sequence content). Compared to mammalian APOBEC / AID deaminases, we speculate that Sdd7 from the bacterium Actinosynnema mirum may have higher activity under temperatures suitable for soybean growth. Analysis of Sdd6 revealed that this deaminase is by default more specific than other deaminases while maintaining high target editing activity. Interestingly, we found that AlphaFold2-based modeling further enables protein engineering efforts to minimize protein size, which is crucial for using these editing techniques for viral delivery in in vivo therapeutic applications.

[0239] The above is only a preferred embodiment of the present invention, but those skilled in the art can make some improvements and supplements without departing from the method of the present invention, and these improvements and supplements are naturally considered to be within the protection scope of the present invention.

[0240] Related sequences and brief descriptions

[0241] [ka]

[0242] [ka]

[0243] [ka]

[0244] [ka]

[0245]

change

[0246]

change

[0247]

change

[0248]

change

[0249]

change

[0250]

change

Claims

1. A method for clustering proteins based on a three-dimensional structure, (1) A step of obtaining sequences of multiple candidate proteins from a database, (2) A step of predicting the three-dimensional structure of each of the multiple candidate proteins using a protein prediction program, (3) A step of performing multi-structure alignment on the three-dimensional structures of the multiple candidate proteins using a scoring function and obtaining a structural similarity matrix, (4) The step of clustering the plurality of candidate proteins based on the structural similarity matrix by a phylogenetic tree construction method, The aforementioned candidate protein is cytosine deaminase, The scoring function used in step (3) is the TM-score, and the calculation formula is as follows: [Math 1] Here, LN is the length of the amino acid sequence of the target protein, LT is the length of the amino acid sequence appearing in both the template and target structures, di is the distance between the i-th residue pair of the template and target structures, d0 is the scale of the normalized matching difference, and Max represents the maximum value after optimal spatial superposition.

2. The method according to claim 1, wherein in step (1), the sequences of the plurality of candidate proteins are obtained through annotation information in a database, or in step (1), the sequences of the plurality of candidate proteins are obtained by searching the database based on sequence identity / similarity using the sequence of a reference protein.

3. The method according to claim 1, wherein the protein prediction program in step (2) is selected from AlphaFold2, RoseTT, or other programs capable of predicting protein structures.

4. The method according to claim 1, wherein the scoring function used in step (3) includes a TM score, RMSD, LDDT, GDT score, QSC, FAPE, or other scoring function capable of scoring protein structure similarity.

5. The method for constructing the phylogenetic tree in step (4) is the unweighted pair group method (UPGMA) using the arithmetic mean, The aforementioned UPGMA includes obtaining the distance between any two proteins using the following formula: [Math 2] The method according to claim 1, wherein d(ABX) is the distance between two points.

6. The method according to claim 1, wherein in step (4), a clustering dendrogram of the plurality of candidate proteins is obtained.

7. A method for predicting the function of a protein based on a three-dimensional structure, comprising clustering a plurality of candidate proteins based on the method of any one of claims 1 to 6, and then predicting the function of the candidate proteins based on the clustering results.

8. The method according to claim 7, wherein the plurality of candidate proteins include at least one reference protein with a known function.

9. The method according to claim 8, wherein the function of other candidate proteins in the same clade or subclade is predicted by the position of the functional reference protein in the clustering dendrogram.

10. The method according to claim 8, wherein the reference protein is cytosine deaminase.

11. The method according to claim 10, wherein the reference protein is a reference cytosine deaminase, and the reference cytosine deaminase is rAPOBEC1 having the sequence shown in SEQ ID NO: 64, or DddA having the sequence shown in SEQ ID NO:

65.

12. A method for identifying the minimal functional domain of a protein based on its three-dimensional structure, a) a step of determining a conserved core structure by aligning the structures of multiple candidate proteins clustered into the same clade or subclade by the method of any one of claims 1 to 6, b) A method comprising the step of identifying the preserved core structure as a minimal functional domain.

13. The method according to claim 12, wherein the plurality of candidate proteins include at least one reference protein with a known function.

14. The method according to claim 13, wherein the reference protein is cytosine deaminase.

15. The method according to claim 14, wherein the reference protein is a reference cytosine deaminase, and the reference cytosine deaminase is rAPOBEC1 having the sequence shown in SEQ ID NO: 64, or DddA having the sequence shown in SEQ ID NO: 65.