Cytosine deaminases and their use in base editing
Novel cytosine deaminases were identified using three-dimensional structure prediction and phylogenetic tree construction methods. An efficient cytosine base editing system was developed, which solved the problem of insufficient deaminases in existing systems and enabled precise editing of the genomes of various organisms.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
- Filing Date
- 2023-03-07
- Publication Date
- 2026-05-26
AI Technical Summary
Existing base editing systems lack a variety of deaminases, limiting the ability to precisely manipulate target DNA sequences.
By using AlphaFold2-based three-dimensional structure prediction and phylogenetic tree construction methods, novel cytosine deaminases were identified and clustered, and a highly efficient cytosine base editing system was developed. Combined with APOBEC/AID family deaminases, precise editing of target nucleotides was achieved.
It improves the efficiency and accuracy of base editing, expands the application range of base editing systems, and is suitable for targeted editing of genomes in a variety of organisms.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on March 7, 2023, with application number 202310220057.1 and invention title "Cytosine deaminase and its use in base editing". Technical Field
[0002] This invention relates to the field of genetic engineering. Specifically, it relates to cytosine deaminases and their use in base editing. More specifically, it relates to a base editing system based on a newly identified cytosine deaminase, a method for base editing a target sequence in the genome of an organism (e.g., a plant) using this system, and genetically modified organisms (e.g., plants) and their offspring produced by said method. Background of the Invention
[0003] Modifying specific sequences in an organism's genome can endow it with new, stably heritable traits. Variations in single nucleotides at specific sites can alter the amino acid sequence of a gene, cause premature termination, or alter regulatory sequences, leading to the development of desirable traits. Genome editing technologies, such as the CRISPR / Cas9 system, can target specific sequences in the genome. Base editing systems developed by combining genome editing systems with deaminases, leveraging the binding properties of these systems to target sequences, can precisely deaminate target nucleotides in the genome. For example, cytosine base editing systems, by fusing APOBEC / AID family and APOBEC / AID-like deaminases, can achieve the conversion of cytosine (C) to uracil (U) at the target site, followed by the conversion to thymine (T) with the aid of relevant cellular repair pathways. Furthermore, introducing nicks into single strands that have not undergone deamination on the contralateral side can significantly improve the efficiency of base editing.
[0004] Based on structural comparisons of deaminases, Iyer et al. identified proteins with potential deamination functions and categorized these proteins into at least 20 branches (Iyer, LM, Zhang, D., Rogozin, IB, & Aravind, L. (2011). Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems. Nucleic acids research, 39(22), 9473-9497.). They found significant differences in structure and sequence among the deaminases in different branches. The functions of some branches have been elucidated, including the "dCMP deaminase and ComE" branch, which converts dCMP to dUMP; the "Guanine deaminase" branch, which converts guanine (G) to xanthine (I); the "RibD-like" branch, which functions as a diaminohydroxyphosphonoribonucleoylpyrimidine deaminase; the "Tad1 / ADAR" branch, which functions as an RNA editing enzyme, converting RNA adenine (A) to xanthine (I); and the "PurH / AICAR transformylase" branch, which exhibits formyltransferase activity. However, the functions of some branches, such as whether they specifically possess deamination activity and what substrates they can deaminate, have not yet been elucidated or confirmed. Examples include the bacterial branches SCP1.201, XOO2897, MafB19, and Pput_2613. Currently, only a few types of deaminases derived from the APOBEC / AID branch, namely APOBEC1, APOBEC3, CDA, AID, CDA1L1, and CDA1L2, have been shown to act on single-stranded DNA, thus making them applicable to cytosine base editing systems.
[0005] There is still a need in the field for more deaminases that can be used in base editing systems to expand base editing systems and improve the ability to precisely manipulate target DNA sequences. Brief description of the attached diagram
[0006] Figure 1 The potential deaminase No. 182 (SEQ ID NO: 1) in the APOBEC / AID branch enables cytosine base editing in the reporter system.
[0007] Figure 2The potential deaminase No. 182 (SEQ ID NO: 1) in the APOBEC / AID branch enables cytosine base editing at an endogenous site.
[0008] Figure 3 The potential deaminase No. 69 (SEQ ID NO:2) in the SCP1.201 branch enables cytosine base editing at endogenous sites.
[0009] Figure 4 The editing efficiency of cytosine bases at the endogenous site of OsACC-T1 in rice by eight deaminases with high editing efficiency.
[0010] Figure 5 The editing efficiency of cytosine bases at the endogenous site of CDC48-T2 in rice by eight deaminases with high editing efficiency.
[0011] Figure 6 Editing efficiency of cytosine bases at the endogenous site of OsACC-T1 in rice by eight deaminases with moderate editing efficiency.
[0012] Figure 7 Editing efficiency of cytosine bases at the endogenous site of CDC48-T2 in rice by eight deaminases with moderate editing efficiency.
[0013] Figure 8 A protein clustering workflow based on AlphaFold2 predicted structure was implemented. The structure of candidate sequences was predicted using AlphaFold2, and clustering was performed based on structural similarity. Subsequently, the cytidine deamination activities of proteins from each structural branch on ssDNA and dsDNA were experimentally tested in plant and human cells.
[0014] Figure 9 The relabeling and synthesis process for candidate deaminases. We used ProteinBLAST from the NCBI database. https: / / blast.ncbi.nlm.nih.gov / Blast.cgi The full-length gene encoding deaminase was obtained, and then the deaminase domain sequence was re-annotated using hmmscan. https: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan The resulting domain sequences were used for structural classification. To confirm their deaminase activity, we synthesized several candidate deaminases with extended N-terminal and C-terminal sequences, and then assessed their cytidine deaminase activity using a reporter system or at endogenous sites.
[0015] Figure 10 The structural similarity matrix reflects the similarity between 242 predicted protein structures from 16 deaminase families (238 proteins) and one outgroup JAB (4 proteins). Proteins from different families are distinguished by different numbers; the intensity of the heatmap color indicates the degree of similarity.
[0016] Figure 11: (A) Proteins are classified into different deaminase families based on their structure, with different families distinguished by different numbers; (B) Representative predicted structures of each of the 16 deaminase branches.
[0017] Figure 12 Figure 11 shows a comparison of representative structures of two branches of the LmjF365940(A), APOBEC(B), dCMP(C), and MafB19(D) families. Although the two branches of these four families have partially similar structures, the overall structures of the two branches show relatively large differences, leading them to be classified as different branches.
[0018] Figure 13: (A) Classification of SCP1.201 deaminases based on protein structure. The JAB family is considered an outgroup, and the tested deaminases are shown as single-stranded editing (ssDNA), double-stranded editing (dsDNA), or non-double-stranded / single-stranded editing (non-ds / ss) based on function. Light gray undefined deaminases await further functional analysis. The single-stranded edited deaminase domains in the diagram are: SCP356, SCP020, SCP051, SCP170, SCP014, SCP273, SCP158, SCP013, SCP008, SCP157, SCP315, SCP183, SCP044, SCP012, SCP011, SCP018, SCP038, SCP016, SCP017; the double-stranded edited deaminase domains are: SCP271, SCP103, SCP009, SCP006, SCP004, SCP234, SCP177; the remaining labeled deaminases are unedited deaminases. (B) Prediction of the core structure of DddA using AlphaFold2. (C) Typical structural features of Ddd protein (a protein with double-stranded deaminase activity). (D) Prediction of the core structure of Sdd7 using AlphaFold2. (E) Typical structural features of Sdd protein (a protein with single-stranded deaminase activity).
[0019] Figure 14 Identification of cytosine deamination activity of ssDNA and dsDNA at endogenous sites in animal cells. (A) Schematic diagram of ssDNA base editing vector used for endogenous site editing. (B) Schematic diagram of DdCBE vector and its binary form. (C) Detection of DdCBE editing activity on dsDNA and CBE editing activity on ssDNA in HEK293T cells, followed by high-throughput sequencing.
[0020] Figure 15Experimental evaluation of dsDNA deamination activity at two endogenous sites in HEK293T cells using Ddd. Base editing sites were used for calculation; color intensity represents editing efficiency.
[0021] Figure 16: (A) Experimental assessment of ssDNA deamination activity of Sdd at two endogenous sites in HEK293T cells. Base editing sites used for calculation; color intensity represents editing efficiency. (B) Experimental assessment of ssDNA deamination activity of Sdd at HsJAK2 and HsSIRT6 sites. Data are the average of three replicate independent experiments.
[0022] Figure 17 Evaluation of the editing properties of newly discovered Ddd proteins as base editors. (A) Editing efficiency and editing window of SCP1.201 dsDNA deaminases Ddd1, Ddd7, Ddd8, Ddd9, and DddA at two genomic targets in HEK293T cells. (B) Plasmid library analysis to analyze the contextual preference of each Ddd protein in mammalian cells. Candidate proteins target and edit “NC 10 N” motif. (C) Through plasmid library analysis, the context-preference motif logo diagrams of Ddd1, Ddd7, Ddd8, Ddd9, and DddA were summarized. In the figure, dots represent single biological replicates, bar heights represent the average editing efficiency, and error bars represent the standard deviations of three independent biological experiments.
[0023] Figure 18 SCP1.201 dsDNA deaminase targets two sites in HEK293T cells ( Figure 18 Editing efficiency and heatmap of editing window for A and B).
[0024] Figure 19 : Percentage of context-biased editing efficiency of different Ddd deaminases in 16 plasmid libraries. Data are expressed as the average of three independent experiments.
[0025] Figure 20 To evaluate the role of newly discovered Sdd proteins as base editors in plants. The overall editing efficiency of ten Sdd proteins and rAPOBEC1 at six endogenous target sites in rice protoplasts was investigated. The average editing frequency of APOBEC1 at each target site was set to 1, and the observed editing efficiency of each Sdd protein was normalized accordingly.
[0026] Figure 21Editing behavior of Sdd deaminases and APOBEC1 at six endogenous target sites in rice protoplasts. The (AF) heatmap shows the editing efficiency and editing window of 10 Sdd deaminases and APOBEC1 at OsAAT (A), OsACC1 (B), OsCDC48-T1 (C), OsCDC48-T2 (D), OsDEP1 (E), and OsODEV (F) sites in rice protoplasts. Values given in the heatmap cells represent C-to-T editing efficiency, with color intensity indicating higher efficiency. Target sequences are listed above the heatmap, with dark boxes marking C-to-T editing locations, and the last three light-colored text indicating PAMs. Data are expressed as the average of three independent experiments.
[0027] Figure 22 Editing behavior of SCP1.201's ssDNA deaminase and APOBEC deaminase at three endogenous targets in HEK293T cells. The (AC) heatmap illustrates the editing efficiency and editing window of four Sdd deaminases, as well as APOBEC1, APOBEC3A, APOBEC1-YE1, and APOBEC1-YEE, at HsEMX1(A), HsHEK2(B), and HsWFS1(C) sites in HEK293T cells. Values given in the heatmap cells represent C-to-T editing efficiency, with color intensity indicating higher efficiency. Target sequences are listed above the heatmap, with dark boxes marking C-to-T editing locations and the last three light-colored text indicating PAM. Data are expressed as the average of three independent experiments.
[0028] Figure 23 This study compared the editing efficiencies of Sdd7, APOBEC1, and APOBEC3A at five sites in rice protoplasts. (AE) The efficiency of the Sdd7, APOBEC1, and APOBEC3A base editors at five endogenous targets was compared: (A) OsACTG, (B) OsALS-T1, (C) OsALS-T2, (D) OsCDC48-T3, and (E) OsMPK16. Data are representative from three independent experiments. Column height represents the mean editing efficiency, and error bars represent the standard deviation of the three independent biological experiments.
[0029] Figure 24 Sequence preferences of Sdd deaminases and APOBEC1 at five endogenous targets in rice protoplasts. The stacked plot shows the contextual preferences of 10 Sdd deaminases and APOBEC1 at five endogenous targets: OsAAT, OsACC1, OsCDC48-T1, OsCDC48-T2, and OsDEP1. The bars, from bottom to top, represent the C-to-T editing preferences of TC, AC, GC, and CC. Data are from three independent experiments.
[0030] Figure 25: (A) Overview of high-throughput quantification of the activity and properties of Sdd and rAPOBEC1 in HEK293T cells using a 12K-TRAPseq library. (B) Evaluation of Sdd and rAPOBEC1 editing preferences and patterns using a 12K-TRAP library. The left panel shows the editing efficiency and editing window of the deaminases. The right panel shows the sequence motif logo diagram reflecting the contextual preferences of the deaminases.
[0031] Figure 26 (A) Off-target effects were assessed in rice protoplasts using orthogonal R-loop analysis. The points represent the average frequency of target C-to-T transitions at the six target sites in rice for each base editor. Figure 20 (B) and the off-target C-to-T switching frequencies in two ssDNAs (OsDEP1-SaT1 and OsDEP1-SaT2) that are not dependent on sgRNA. Figure 26 A. On-target:off-target editing ratios for each base editor. (C) On-target:off-target editing ratios of Sdd6, rAPOBEC1-YE1, rAPOBEC1-YEE, rAPOBEC1, and hAPOBEC3A tested at two on-target and three off-target sites in HEK293T cells. Points in the figure represent individual biological replicates, bar heights represent the mean, and error bars represent the standard deviations of three independent biological replicates.
[0032] Figure 27 The specific off-target frequencies of Sdd deaminase and APOBEC1 at two endogenous targets in rice protoplasts ( Figure 26 Off-target effects were assessed using an orthogonal R-loop method. (A, B) Off-target frequencies of Sdd deaminase and APOBEC1 at the OsDEP1-SaT1 (A) and OsDEP1-SaT2 (B) sites in rice protoplasts. Data are from three independent experiments.
[0033] Figure 28 The specific editing efficiencies of Sdd6 and APOBEC base editors at two target sites and four off-target sites were tested in HEK293T cells for both on-target and off-target effects. Figure 26 C). The targeting and off-target efficiencies of Sdd6, APOBEC1-YE1, APOBEC1-YEE, APOBEC1, APOBEC1, and APOBEC3A at the HsJAK2 target site and the corresponding HsJAK2-Sa and HsSIRT6-Sa off-target sites, respectively, and the targeting and off-target editing efficiencies at the HsRNF2-Sa and HsFANCF-SaT1 off-target sites, respectively, at the HsHEK3 target site. The data are from three independent experiments.
[0034] Figure 29 Conserved protein structures of highly active Sdd deaminases predicted by AlphaFold2 are presented. The core structure of Sdd deaminases with high deaminase activity is given. For some active deaminases, α4 is not an essential structure.
[0035] Figure 30: Engineered truncated Sdd proteins for use in animals and plants. (A) Engineered truncated Sdd proteins. The top figure shows the structures of Sdd6, Sdd7, Sdd3, and Sdd9 predicted by AlphaFold2. Conserved regions are shown in dark colors, and truncated regions in light colors. The bottom figure shows the editing efficiency of Sdd and its minimized version at two endogenous sites in two endophytic rice protoplasts and HEK293T cells for the relative original length of Sdd proteins. (B) SaCas9-based CBE vectors are theoretically packaged in a single AAV. The top figure shows a schematic diagram of APOBEC / AID-like deaminases, the minimized version of Sdd, and their AAV vectors. Among them, APOBEC3G, hAPOBEC3B, rAPOBEC1, PmCDA1, APOBEC3A, and hAID deaminases are too large for packaging in a single AAV. The bottom figure shows a schematic diagram of an AAV vector based on the minimized mini version of Sdd. (C) Editing efficiency of mini-Sdd6 at two endogenous targets of the MmHPD gene in mouse N2a cells. (D) Editing efficiency of mini-Sdd7, rAPOBEC1, hAPOBECA, and humanAID base editors at five endogenous targets in soybean hairy roots. (E) Frequency of mutations induced by mini-Sdd7 in T0 generation soybean plants. (F) Genotypes of base-edited soybean plants. (G) Phenotypes of soybean plants treated with carfentraone ethyl for 10 days. The left figure shows wild-type soybean plants (R98). The right figure shows base-edited soybean plants (C98). For figures A, C, and D, dots represent single biological replicates, column heights and broken line points represent means, and error bars represent the standard deviations of three independent biological experiments.
[0036] Figure 31 : Base editing efficiency in regenerated rice. (A) Schematic diagram of rice base editing binary vectors transformed by Agrobacterium. (B) Efficiency of mini-Sdd7 and hAPOBEC3A base editors in inducing mutations in T0 rice plants.
[0037] Figure 32 Schematic diagram of a base-editing binary vector for Agrobacterium-mediated transformation in soybeans. Invention Details
[0038] I. Definition
[0039] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are all widely used terms and routine procedures in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.
[0040] As used herein, the term “and / or” covers all combinations of items connected by the term and should be regarded as if each combination had been listed separately herein. For example, “A and / or B” covers “A,” “A and B,” and “B.” For example, “A, B, and / or C” covers “A,” “B,” “C,” “A and B,” “A and C,” “B and C,” and “A and B and C.”
[0041] "Cytosine deaminase" refers to a deaminase that can accept nucleic acids, such as single-stranded DNA, as substrates and catalyze the deamination of cytidine or deoxycytidine into uracil or deoxyuracil, respectively.
[0042] The term "genome," as used in this article, encompasses not only chromosomal DNA located in the cell nucleus but also organelle DNA located in subcellular components of the cell, such as mitochondria and plastids.
[0043] As used herein, “organism” includes any organism suitable for genome editing, preferably eukaryotes. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants including monocots and dicots such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, and Arabidopsis thaliana.
[0044] "Genetically modified organism" or "genetically modified cell" refers to an organism or cell whose genome contains exogenous polynucleotides or modified genes or expression regulatory sequences. For example, exogenous polynucleotides can be stably integrated into the genome of an organism or cell and inherited across generations. Exogenous polynucleotides can be integrated into the genome alone or as part of a recombinant DNA construct. Modified genes or expression regulatory sequences are sequences in the genome of an organism or cell that contain single or multiple deoxynucleotide substitutions, deletions, and additions.
[0045] In relation to a sequence, “exogenous” means a sequence that originates from a foreign species, or, if from the same species, a sequence whose composition and / or loci have been significantly altered from its natural form through deliberate human intervention.
[0046] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.
[0047] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably in this invention to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.
[0048] Sequence “identity” has a generally accepted meaning in the art, and the percentage of sequence similarity between two nucleic acid or polypeptide molecules or regions can be calculated using publicly available techniques. Sequence similarity can be measured along the full length of a polynucleotide or polypeptide or along a region of that molecule. (See, for example: Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., Stockton Press, New York, 1991). Although there are many methods for measuring the similarity between two polynucleotides or polypeptides, the term "similarity" is well known to those skilled in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48:1073 (1988)).
[0049] When the term "comprising" is used herein to describe a protein or nucleic acid sequence, the protein or nucleic acid may consist of the stated sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, while still possessing the activities described in this invention. Furthermore, those skilled in the art will understand that the methionine encoded by the start codon at the N-terminus of a polypeptide may be retained in certain practical situations (e.g., when expressed in a specific expression system) without substantially affecting the polypeptide's function. Therefore, when describing a specific polypeptide amino acid sequence in this specification and claims, although it may not contain the methionine encoded by the start codon at the N-terminus, the sequence containing that methionine is still included, and correspondingly, its encoding nucleotide sequence may also contain the start codon; and vice versa.
[0050] In peptides or proteins, suitable conserved amino acid substitutions are known to those skilled in the art and can generally be performed without altering the biological activity of the resulting molecule. Typically, those skilled in the art recognize that single amino acid substitutions in non-essential regions of a polypeptide do not substantially alter its biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub.co., p. 224).
[0051] As used in this invention, "expression construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to generate mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein.
[0052] The "expression construct" of the present invention may be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (such as mRNA).
[0053] The "expression construct" of the present invention may contain regulatory sequences and nucleotide sequences of interest from different sources, or regulatory sequences and nucleotide sequences of interest from the same source but arranged in a manner different from those normally found in nature.
[0054] "Regulatory sequence" and "regulatory element" are used interchangeably, referring to nucleotide sequences located upstream (5' non-coding sequence), in the middle, or downstream (3' non-coding sequence) of a coding sequence that affect the transcription, RNA processing, or stability or translation of the relevant coding sequence. Regulatory sequences may include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences.
[0055] A "promoter" refers to a nucleic acid fragment that controls the transcription of another nucleic acid fragment. In some embodiments of the present invention, a promoter is a promoter capable of controlling gene transcription in a cell, regardless of whether it originates from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmental regulatory promoter, or an inducible promoter.
[0056] "Constraint promoters" refer to promoters that generally cause gene expression in most cell types and under most conditions. "Tissue-specific promoters" and "tissue-preferred promoters" are used interchangeably and refer to promoters that are primarily, but not necessarily, expressed specifically in one tissue or organ, and may also be expressed in a specific cell type. "Developmental regulatory promoters" are promoters whose activity is determined by developmental events. "Inducible promoters" selectively express manipulated DNA sequences in response to endogenous or exogenous stimuli (environment, hormones, chemical signals, etc.).
[0057] Examples of promoters include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. Examples of pol I promoters include the chicken RNA pol I promoter. Examples of pol II promoters include, but are not limited to, the cytomegalovirus Immediate Early (CMV) promoter, the Rous sarcoma virus long terminal repeat (RSV-LTR) promoter, and the simian virus 40 (SV40) Immediate Early promoter. Examples of pol III promoters include the U6 and H1 promoters. Inducible promoters such as metallothionein promoters can be used. Other examples of promoters include the T7 phage promoter, the T3 phage promoter, the β-galactosidase promoter, and the Sp6 phage promoter. When used for plants, promoters can be the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, and the rice actin promoter.
[0058] As used herein, the term "operably linked" refers to the linking of a regulatory element (e.g., but not limited to, promoter sequences, transcription termination sequences, etc.) to a nucleic acid sequence (e.g., coding sequences or open reading frames) such that transcription of the nucleotide sequence is controlled and regulated by the transcriptional regulatory element. Techniques for operably linking regulatory element regions to nucleic acid molecules are known in the art.
[0059] "Introducing" nucleic acid molecules (such as plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the cells of an organism with the nucleic acid or protein, enabling the nucleic acid or protein to function within the cell. The term "transformation" as used in this invention includes both stable transformation and transient transformation.
[0060] "Stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome, resulting in the stable inheritance of the foreign gene. Once stable transformation occurs, the foreign nucleic acid sequence is stably integrated into the genome of the organism and its subsequent generations.
[0061] "Transient conversion" refers to the introduction of nucleic acid molecules or proteins into cells to perform their functions without the foreign gene being stably inherited. In transient conversion, the foreign nucleic acid sequence does not integrate into the genome.
[0062] II. Protein Clustering and Function Prediction Methods Based on Three-Dimensional Structures
[0063] In one aspect, the present invention provides a protein clustering method, comprising:
[0064] (1) Obtain the sequences of multiple candidate proteins from the database;
[0065] (2) Predict the three-dimensional structure of each of the multiple candidate proteins using a protein prediction program;
[0066] (3) The three-dimensional structures of the multiple candidate proteins are compared using a scoring function to obtain a structural similarity matrix.
[0067] (4) Cluster the multiple candidate proteins based on the structural similarity matrix using the phylogenetic tree construction method.
[0068] In some implementations, the sequences of the plurality of candidate proteins are obtained in step (1) using annotation information in a database. For example, if deaminases are clustered, the sequences of a plurality of candidate proteins annotated as “deaminase” can be selected from the database.
[0069] In some embodiments, step (1) involves obtaining the sequences of the plurality of candidate proteins by searching a database based on sequence identity / similarity using the sequence of a reference protein. For example, the sequences of the plurality of candidate proteins can be obtained by searching a database using a BLAST procedure based on the sequence of a reference protein with known functions. In some embodiments, the plurality of candidate proteins have at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity with the sequence of the reference protein.
[0070] In some embodiments, the candidate protein is a deaminase. In some preferred embodiments, the candidate protein is a cytosine deaminase.
[0071] In some implementations, the database is an InterPro database.
[0072] In some implementations, the protein structure prediction program in step (2) is selected from AlphaFold2, RoseTT, or other programs capable of predicting protein structures. (John Jumper and others, 'Highly Accurate Protein Structure Prediction with AlphaFold', Nature, 596.7873(2021), 583–89)
[0073] In some implementations, the scoring function used in step (3) includes TM-score, RMSD, LDDT, GDT score, QSC, FAPE, or other scoring functions that can score protein structure similarity. (John Jumper and others, 'Highly Accurate Protein Structure Prediction with AlphaFold', Nature, 596, 7873 (2021), 583–89)
[0074] In some implementations, when the scoring function is TM-score, the TM-score value is at least 0.6, at least 0.7, at least 0.75, at least 0.8, at least 0.85, or higher. The calculation of the TM-score can be referenced, for example, to the formulas and methods described in the "Materials and Methods" section of the embodiments of this application.
[0075] In some implementations, the phylogenetic tree construction method in step (4) is the Unweighted Pair Group Method with Arithmetic Mean (UPGMA) (CPKurtzman, Jack W. Fell, and T. Boekhout, The Yeasts: A Taxonomic Study, 5th ed (Amsterdam: Elsevier, 2011). 'A Statistical Method for Evaluating Systematic Relationships-Robert Reuven Sokal, Charles Duncan Michener-Google Books').
[0076] In some implementations, step (4) involves obtaining a clustering dendrogram of the plurality of candidate proteins.
[0077] In one aspect, the present invention provides a protein function prediction method based on three-dimensional structure, the method comprising clustering multiple candidate proteins using the protein clustering method according to the present invention, and then predicting the function of the candidate proteins based on the clustering results.
[0078] In some implementations, the plurality of candidate proteins includes at least one reference protein with a known function.
[0079] In some embodiments, the function of other candidate proteins in the same branch or subbranch is predicted by the position of a reference protein with known function in a cluster (dendritic diagram). In some embodiments, other candidate proteins located in the same branch or subbranch as the reference protein are predicted to have the same or similar function as the reference protein. In some embodiments, the TM-score between different candidate proteins within the same branch or subbranch is at least 0.6, at least 0.7, at least 0.75, at least 0.8, at least 0.85, or higher. In some embodiments, the TM-score between candidate proteins in different branches or subbranchs is less than 0.85, less than 0.8, less than 0.75, less than 0.7, less than 0.6, or lower.
[0080] In some embodiments, the reference protein is a deaminase. In some preferred embodiments, the reference protein is a cytosine deaminase. In some embodiments, the reference protein is a reference cytosine deaminase, wherein the reference cytosine deaminase is rAPOBEC1 as shown in SEQ ID No: 64 or DddA as shown in SEQ No: 65. In some embodiments, the TM-score between different candidate proteins within the same branch or subbranch as the reference protein, or the TM-score with the reference protein, is at least 0.7. In some embodiments, the TM-score between the candidate protein and the reference protein in different branches or subbranchs is less than 0.7.
[0081] In another aspect, the present invention provides a method for identifying the minimum functional domain of a protein based on three-dimensional structure, comprising:
[0082] a) Compare the structures of multiple candidate proteins clustered together by the method of the present invention, for example, clustered in the same branch or subbranch, to determine the conserved core structure.
[0083] b) Identify the conservative core structure as a minimal functional domain.
[0084] As used in this article, "minimum functional domain" refers to the smallest part of a protein that can essentially maintain the function of the entire protein.
[0085] In some implementations, the plurality of candidate proteins includes at least one reference protein with a known function.
[0086] In some embodiments, the reference protein is a deaminase. In some preferred embodiments, the reference protein is a cytosine deaminase. In some embodiments, the reference protein is a reference cytosine deaminase, wherein the reference cytosine deaminase is rAPOBEC1 as shown in SEQ ID No: 64 or DddA as shown in SEQ No: 65.
[0087] In another aspect, the present invention provides a cytosine deaminase identified by the protein function prediction method of the present invention.
[0088] In another aspect, the present invention provides a truncated cytosine deaminase comprising or consisting of a minimal functional domain of cytosine deaminase identified by the method of the present invention.
[0089] In one aspect, the present invention also provides the use of the cytosine deaminase or truncated cytosine deaminase in gene editing, such as base editing, in organisms or somatic cells.
[0090] III. Cytosine deaminase and base-editing fusion proteins containing it
[0091] In one aspect, the present invention provides a cytosine deaminase, wherein the cytosine deaminase is capable of deaminating the cytosine base of deoxycytidine in DNA. In some embodiments, the cytosine deaminase is derived from bacteria.
[0092] In some embodiments, the cytosine deaminase has an AlphaFold2 TM-score of at least 0.6, 0.7, 0.75, 0.8, and 0.85 for its three-dimensional structure, and contains an amino acid sequence with 20-70%, 20-60%, 20-50%, 20-45%, 20-40%, and 20-35% sequence identity with the reference cytosine deaminase, respectively. Or an amino acid sequence having at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity; the cytosine deaminase has the function of deamination of the cytosine bases of deoxycytidine in DNA.
[0093] In some embodiments, the reference cytosine deaminase is:
[0094] (a) The sequence shown in SEQ ID No: 64 is rAPOBEC1; or
[0095] (b) The sequence shown in SEQ ID No: 65, DddA; or
[0096] (c) The sequence is shown as Sdd7 of SEQ ID No: 4.
[0097] In some embodiments, the TM-score of the AlphaFold2 three-dimensional structure of rAPOBEC1 shown in SEQ ID No: 64 is not less than 0.6, 0.7, 0.75, 0.8, or 0.85, and it contains an amino acid sequence having 20-70%, 20-60%, 20-50%, 20-45%, 20-40%, or 20-35% sequence identity with SEQ ID No: 64, or having at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity; the cytosine deaminase has the function of deamination of the cytosine bases of deoxycytidine in DNA.
[0098] In some embodiments, the TM-score of the AlphaFold2 three-dimensional structure of DddA shown in SEQ ID No: 65 is not less than 0.6, 0.7, 0.75, 0.8, or 0.85, and it contains an amino acid sequence having 20-70%, 20-60%, 20-50%, 20-45%, 20-40%, or 20-35% sequence identity with SEQ ID No: 65, or having at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity; the cytosine deaminase has the function of deamination of the cytosine bases of deoxycytidine in DNA.
[0099] In some embodiments, the TM-score of the AlphaFold2 three-dimensional structure of Sdd7 shown in SEQ ID No: 4 is not less than 0.6, 0.7, 0.75, 0.8, or 0.85, and it contains an amino acid sequence having 20-70%, 20-60%, 20-50%, 20-45%, 20-40%, or 20-35% sequence identity with SEQ ID No: 4, or having at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity; the cytosine deaminase has the function of deamination of the cytosine bases of deoxycytidine in DNA.
[0100] In some embodiments, the cytosine deaminase is derived from the AID / APOBEC branch, the SCP1.201 branch, the MafB19 branch, the Novel AID / APOBEC-like branch, the TM1506 branch, or the XOO2897 branch.
[0101] In this paper, the cytosine deaminase branch is defined according to the description in Iyer, LM, Zhang, D., Rogozin, IB, & Aravind, L. (2011). Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxinsystems. Nucleic acids research, 39(22), 9473-9497.
[0102] In some embodiments, the cytosine deaminase is derived from the AID / APOBEC branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with SEQ ID NO:1.
[0103] In some embodiments, the cytosine deaminase is derived from the SCP1.201 branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with any of SEQ ID Nos. In some embodiments, the cytosine deaminase is capable of deaminating cytosine bases in double-stranded DNA. In some embodiments, the amino acid sequence of the cytosine deaminase consists of the amino acid sequences of any of SEQ ID Nos. 28-40. In some embodiments, the amino acid sequence of the cytosine deaminase consists of the amino acid sequences of any of SEQ ID Nos. 28, 33, 34, and 35.
[0104] In some embodiments, the cytosine deaminase is derived from the SCP1.201 branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with any of SEQ ID Nos: 2-18, 41-49. In some embodiments, the cytosine deaminase is capable of deaminating cytosine bases in single-stranded DNA. In some embodiments, the amino acid sequence of the cytosine deaminase consists of the amino acid sequence of any of SEQ ID Nos: 2-18, 41-49. In some embodiments, the amino acid sequence of the cytosine deaminase consists of the amino acid sequence of any of SEQ ID Nos: 2-7, 12, 17.
[0105] In some embodiments, the cytosine deaminase is a truncated cytosine deaminase capable of deamination of the cytosine base of deoxycytidine in DNA. In some embodiments, the truncated cytosine deaminase is 130-160 amino acids in length. In some embodiments, the truncated cytosine deaminase may be individually packaged in AAV particles.
[0106] In some embodiments, the truncated cytosine deaminase comprises an amino acid sequence having at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with any one of SEQ ID Nos: 50-55. In some embodiments, the truncated cytosine deaminase is capable of deaminating cytosine bases in single-stranded DNA. In some embodiments, the truncated cytosine deaminase consists of an amino acid sequence of any one of SEQ ID Nos: 50-55.
[0107] In some embodiments, the cytosine deaminase is derived from the MafB19 branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with any of SEQ ID No: 19, 56, 57, 58.
[0108] In some embodiments, the cytosine deaminase is derived from the NovelAID / APOBEC-like branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with any of SEQ ID No: 20, 21.
[0109] In some embodiments, the cytosine deaminase is derived from the TM1506 branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with SEQ ID No: 22.
[0110] In some embodiments, the cytosine deaminase is derived from the XOO2897 branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with SEQ ID No: 23, 24, 59-62.
[0111] In some embodiments, the cytosine deaminase is derived from the Toxin deam branch and comprises an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with SEQ ID No: 74 or 75.
[0112] In one aspect, this application relates to the use of the cytosine deaminase of the present invention in gene editing, such as base editing, in organisms or somatic cells.
[0113] In some embodiments, the cytosine deaminase is used to prepare a base-editing fusion protein or base-editing system for base editing in an organism or somatic cells.
[0114] In another aspect, the present invention provides a base editing fusion protein comprising a nucleic acid targeting domain and a cytosine deamination domain, wherein the cytosine deamination domain comprises at least one (e.g., one or two) of the cytosine deamination peptide of the present invention.
[0115] In the embodiments described herein, the terms "fusion protein," "base-editing fusion protein," and "base editor" are used interchangeably and refer to proteins that can mediate the substitution of one or more nucleotides in a target sequence in the genome in a sequence-specific manner. The one or more nucleotide substitutions are, for example, C-to-T substitutions.
[0116] As used herein, a “nucleic acid targeting domain” refers to a domain capable of mediating the attachment of the base-editing fusion protein to a specific target sequence in the genome in a sequence-specific manner (e.g., via guide RNA). In some embodiments, the nucleic acid targeting domain may include one or more zinc finger protein domains (ZFP) or transcription factor effector domains (TALE) targeting a specific target sequence. In some embodiments, the nucleic acid targeting domain comprises at least one (e.g., one) CRISPR effector polypeptide.
[0117] The zinc finger desmin domain (ZFP) typically contains 3-6 individual zinc finger repeat sequences, each of which can identify a unique sequence of, for example, 3 bp. By combining different zinc finger repeat sequences, different genomic sequences can be targeted.
[0118] The "transcription activator-like effector domain" is the DNA-binding domain of a transcription activator-like effector (TALE). TALEs can be engineered to bind to almost any desired DNA sequence.
[0119] As used herein, the term "CRISPR effector protein" generally refers to a nuclease (CRISPR nuclease) or a functional variant thereof that is present in the naturally occurring CRISPR system. The term encompasses any CRISPR-based effector protein capable of sequence-specific targeting within cells.
[0120] As used herein, a “functional variant” of a CRISPR nuclease means one that retains at least the guide RNA-mediated sequence-specific targeting ability. Preferably, the functional variant is a nuclease-inactivating variant, i.e., lacking double-stranded nucleic acid cleavage activity. However, CRISPR nucleases lacking double-stranded nucleic acid cleavage activity also encompass nickases, which form a nick in a double-stranded nucleic acid molecule but do not completely cleave the double-stranded nucleic acid. In some preferred embodiments of the invention, the CRISPR effector protein of the invention possesses nickase activity. In some embodiments, the functional variant recognizes a different PAM (pre-intermediate sequence adjacent motif) sequence relative to the wild-type nuclease.
[0121] “CRISPR effector proteins” can be derived from Cas9 nucleases, including Cas9 nucleases or functional variants thereof. The Cas9 nuclease can be a Cas9 nuclease from a different species, such as spCas9 from *S. pyogenes* or SaCas9 derived from *S. aureus*. “Cas9 nuclease” and “Cas9” are used interchangeably herein, referring to RNA-guided nucleases comprising the Cas9 protein or fragments thereof (e.g., proteins containing the active DNA-cutting domain of Cas9 and / or the gRNA-binding domain of Cas9). Cas9 is a component of the CRISPR / Cas (clustered, regularly spaced short palindromic repeats and related systems) genome editing system, capable of targeting and cleaving DNA target sequences to form DNA double-strand breaks (DSBs) under the guidance of guide RNA. An exemplary amino acid sequence of wild-type spCas9 is shown in SEQ ID NO:25.
[0122] “CRISPR effector proteins” can also be derived from Cpf1 nucleases, including Cpf1 nucleases or their functional variants. The Cpf1 nucleases can be Cpf1 nucleases from different species, such as those from Francisella novicida U112, Acidaminococcus sp. BV3L6, and Lachnospiraceae bacterium ND2006.
[0123] Available “CRISPR effector proteins” can also be derived from nucleases such as Cas3, Cas8a, Cas5, Cas8b, Cas8c, Cas10d, Cse1, Cse2, Csy1, Csy2, Csy3, GSU0054, Cas10, Csm2, Cmr5, Cas10, Csx11, Csx10, Csf1, Csn2, Cas4, C2c1 (Cas12b), C2c3, C2c2, Cas12c, Cas12d (i.e., CasY), Cas12e (i.e., CasX), Cas12f (i.e., Cas14), Cas12g, Cas12h, Cas12i, Cas12j (i.e., CasΦ), Cas12k, Cas12l, and Cas12m, including, for example, these nucleases or their functional variants.
[0124] In some embodiments, the CRISPR effector protein is a nuclease-inactivated Cas9. The DNA-cutting domain of the Cas9 nuclease is known to contain two subdomains: the HNH nuclease subdomain and the RuvC subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC subdomain cleaves the non-complementary strand. Mutations in these subdomains can inactivate the nuclease activity of Cas9, forming a "nuclease-inactivated Cas9." The nuclease-inactivated Cas9 still retains its gRNA-directed DNA binding ability.
[0125] The nuclease-inactivated Cas9 described in this invention can be derived from Cas9 of different species, such as Cas9 derived from *Streptococcus pyogenes* (SpCas9) or Cas9 derived from *Staphylococcus aureus* (SaCas9). Simultaneously, mutations in the HNH nuclease subdomain and RuvC subdomain of Cas9 (e.g., containing mutants D10A and H840A) render the Cas9 nuclease inactive, resulting in a nuclease-dead Cas9 (dCas9). Mutating and inactivating one of the subdomains can impart nickase activity to Cas9, thus obtaining Cas9 nickase (nCas9), for example, nCas9 containing only the mutant D10A.
[0126] Therefore, in some embodiments of various aspects of the present invention, the nuclease-inactivated Cas9 variant of the present invention comprises, relative to wild-type Cas9, the amino acid substitutions D10A and / or H840A, wherein the amino acid numbers refer to SEQ ID NO:25. In some preferred embodiments, the nuclease-inactivated Cas9 comprises, relative to wild-type Cas9, the amino acid substitution D10A, wherein the amino acid numbers refer to SEQ ID NO:25. In some embodiments, the nuclease-inactivated Cas9 comprises the amino acid sequence (nCas9(D10A)) shown in SEQ ID NO:26.
[0127] When used for gene editing, Cas9 nucleases typically require a 5'-NGG-3' PAM (pre-intercalation sequence adjacent motif) at the 3' end of the target sequence. However, the inventors have surprisingly discovered that this PAM sequence occurs very infrequently in certain species, such as rice, significantly limiting gene editing in these species. Therefore, this invention preferably utilizes CRISPR effector proteins that recognize different PAM sequences, such as functional variants of Cas9 nucleases with different PAM sequences.
[0128] In some embodiments of the present invention, the cytidine deamination domain in the fusion protein can deamination the cytidine in the single-stranded DNA generated during the formation of the fusion protein-guide RNA-DNA complex and convert it to U, and then achieve C to T base substitution through base mismatch repair.
[0129] In some embodiments of the present invention, the nucleic acid targeting domain and the cytosine deamination domain are fused via a linker.
[0130] As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids without secondary or higher structures. For example, the linker can be a flexible linker.
[0131] In some embodiments, the base-editing fusion protein comprises, from the N-terminus to the C-terminus, a cytosine deamination domain and a nucleic acid targeting domain in the following order.
[0132] Furthermore, in cells, uracil DNA glycosyltransferase catalyzes the removal of U from DNA and initiates base excision repair (BER), resulting in the repair of U:G to C:G. Therefore, without any theoretical limitations, the combination of the base editing fusion protein of this invention with a uracil DNA glycosyltransferase inhibitor (UGI) will be able to increase the efficiency of C to T base editing.
[0133] In some embodiments, the base-editing fusion protein is co-expressed with a uracil DNA glycosylation inhibitor (UGI).
[0134] In some embodiments, the base-editing fusion protein further comprises a uracil DNA glycosylation inhibitor (UGI).
[0135] In some implementations, the UGI is connected to other parts of the base-editing fusion protein via a connector.
[0136] In some implementations, UGI is linked to other parts of the base-editing fusion protein via a "self-cleaving peptide".
[0137] As used herein, "self-cleaving peptide" refers to a peptide that can self-cleave within a cell. For example, the self-cleaving peptide may contain a protease recognition site, thereby being recognized and specifically cleaved by intracellular proteases. Alternatively, the self-cleaving peptide may be a 2A peptide. 2A peptides are a class of short peptides derived from viruses whose self-cleavage occurs during translation. When two different target peptides are expressed in the same reading frame using a 2A peptide, the two target peptides are generated in an almost 1:1 ratio. Commonly used 2A peptides include P2A from porcine techovirus-1, T2A from the β-tetrasomatic moth virus (Thosea asigna virus), E2A from equine rhinitis A virus, and F2A from foot-and-mouth disease virus. Various functional variants of these 2A peptides are also known in the art and can be used in this invention.
[0138] Preferably, the self-cleaving peptide is not present between or within the nucleic acid targeting domain and the cytosine deamination domain. In some embodiments, the UGI is located at the N-terminus or C-terminus of the base editing fusion protein, preferably the C-terminus.
[0139] In some specific embodiments, the uracil DNA glycosylation inhibitor (UGI) comprises the amino acid sequence shown in SEQ ID NO:27.
[0140] In some embodiments of the invention, the fusion protein may further comprise a nuclear localization sequence (NLS). Generally, one or more NLSes in the fusion protein should have sufficient strength to drive the fusion protein to accumulate in the nucleus of the cell in an amount sufficient to enable its base-editing function. Generally, the strength of nuclear localization activity is determined by the number and location of the NLSes in the fusion protein, the use of one or more specific NLSes, or a combination of these factors.
[0141] In some embodiments of the invention, the NLS of the fusion protein of the invention may be located at the N-terminus and / or the C-terminus. In some embodiments of the invention, the NLS of the fusion protein of the invention may be located between the adenine deamination domain, the cytosine deamination domain, the nucleic acid targeting domain, and / or the UGI. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS at or near the N-terminus. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS at or near the C-terminus. In some embodiments, the polypeptide comprises a combination of these, such as one or more NLS at the N-terminus and one or more NLS at the C-terminus. When more than one NLS is present, each may be selected independently of the other NLS.
[0142] Generally, NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the surface of a protein, but other types of NLS are also known. Non-limiting examples of NLS include: KKRKV, PKKKRKV, or KRPAATKKAGQAKKKK.
[0143] Furthermore, depending on the location of the DNA to be edited, the fusion protein of the present invention may also include other localization sequences, such as cytoplasmic localization sequences, chloroplast localization sequences, mitochondrial localization sequences, etc.
[0144] IV. Base Editing System
[0145] In another aspect, the present invention provides a base editing system comprising: i) the cytosine deaminase or base editing fusion protein of the present invention, and / or an expression construct containing a nucleotide sequence encoding the cytosine deaminase or base editing fusion protein.
[0146] In some implementations, the base editing system is used to modify nucleic acid target regions.
[0147] In some embodiments, the base editing system further comprises ii) at least one guide RNA and / or at least one expression construct containing a nucleotide sequence encoding said at least one guide RNA. However, those skilled in the art will appreciate that if the base editing fusion protein is not based on a CRISPR effector protein, the system may not require a guide RNA or an expression construct encoding it.
[0148] In some embodiments, the at least one guide RNA can bind to the nucleic acid targeting domain of the fusion protein. In some embodiments, the guide RNA targets at least one target sequence within the nucleic acid target region.
[0149] As used herein, a "base editing system" refers to a combination of components required for base editing of nucleic acid sequences, such as genomic sequences in cells or organisms. The individual components of such a system, such as cytosine deaminase, base editing fusion proteins, and one or more guide RNAs, may exist independently or in any combination as a composition.
[0150] In some embodiments, it comprises the cytosine deaminase of the present invention or the fusion protein of the present invention and a guide RNA that can bind to a nucleic acid targeting protein.
[0151] As used herein, "guide RNA" and "gRNA" are used interchangeably and refer to RNA molecules capable of forming a complex with a CRISPR effector protein and targeting the target sequence by means of a certain degree of identity with the target sequence. Guide RNA targets the target sequence through base pairing with the complementary strand of the target sequence. For example, the gRNA used by Cas9 nuclease or its functional variants typically consists of partially complementary crRNA and tracrRNA molecules forming a complex, wherein the crRNA contains a guide sequence (also called a seed sequence) that is sufficiently identical with the target sequence to hybridize with the complementary strand of the target sequence and guide the CRISPR complex (Cas9 + crRNA + tracrRNA) to specifically bind to the target sequence. However, it is known in the art that single guide RNAs (sgRNAs) can be designed that simultaneously contain the characteristics of both crRNA and tracrRNA. The gRNA used by Cpf1 nuclease or its functional variants typically consists only of mature crRNA molecules, which can also be referred to as sgRNA. Designing suitable gRNAs based on the CRISPR nuclease used and the target sequence to be edited is within the capabilities of those skilled in the art.
[0152] In some embodiments, the guide RNA is 15-100 nucleotides in length and contains a sequence of at least 10, at least 15, or at least 20 consecutive nucleotides complementary to the target sequence.
[0153] In some embodiments, the guide RNA comprises 15 to 40 consecutive nucleotide sequences complementary to the target sequence.
[0154] In some implementations, the guide RNA is 15-50 nucleotides in length.
[0155] In some implementations, the target sequence is a DNA sequence.
[0156] In some embodiments, the target sequence is located in the genome of an organism. In some embodiments, the organism is a prokaryote. In some embodiments, the prokaryote is a bacterium. In some embodiments, the organism is a eukaryote. In some embodiments, the organism is a plant or fungus. In some embodiments, the organism is a vertebrate. In some embodiments, the vertebrate is a mammal. In some embodiments, the mammal is a mouse, rat, or human. In some embodiments, the organism is a cell. In some embodiments, the cell is a mouse cell, rat cell, or human cell. In some embodiments, the cell is a HEK-293 cell.
[0157] In some embodiments, after the base editing system of the present invention is introduced into the cell, the base editing fusion protein and the guide RNA are able to form a complex, and the complex specifically targets the target sequence under the guidance of the guide RNA, resulting in one or more Cs being replaced by Ts and / or one or more Aes being replaced by Gs in the target sequence.
[0158] In some embodiments, the at least one guide RNA can target a target sequence located on a sense strand (e.g., a protein-coding strand) and / or an antisense strand within a target nucleic acid region of the genome. When the guide RNA targets the sense strand (e.g., a protein-coding strand), the base editing composition of the present invention can cause one or more Cs within the target sequence on the sense strand (e.g., a protein-coding strand) to be replaced by Ts and / or one or more As to be replaced by Gs. When the guide RNA targets the antisense strand, the base editing composition of the present invention can cause one or more Gs within the target sequence on the sense strand (e.g., a protein-coding strand) to be replaced by As and / or one or more Ts to be replaced by Cs.
[0159] In order to achieve effective expression in cells, in some embodiments of the present invention, the nucleotide sequence encoding the cytosine deaminase or base editing fusion protein is codon-optimized for the organism whose genome is to be modified.
[0160] Codon optimization refers to the modification of nucleic acid sequences to enhance expression in host cells of interest by replacing at least one codon of the natural sequence with codons that are used more frequently or most frequently in the gene in the host cell (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) while maintaining the natural amino acid sequence. Different species exhibit specific preferences for certain codons of specific amino acids. Codon preference (differences in codon use between organisms) is often associated with the translation efficiency of messenger RNA (mRNA), which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be customized to achieve optimal gene expression in a given organism based on codon optimization. Codon utilization tables are readily available, for example in... www.kazusa.orjp / codon / The codons used are available in the Codon Usage Database, and these tables can be adapted in different ways. See Nakamura Y. et al., “Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).
[0161] Organisms whose genomes can be modified using the base editing system of the present invention include any organism suitable for base editing, preferably eukaryotes. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, for example, crop plants, including but not limited to wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, tobacco, cassava, and potatoes.
[0162] V. Base Editing Methods
[0163] In another aspect, the present invention provides a base editing method, which includes contacting the base editing system of the present invention with a nucleic acid molecular target sequence.
[0164] In some embodiments, the nucleic acid molecule is a DNA molecule. In some preferred embodiments, the nucleic acid molecule is a double-stranded DNA molecule or a single-stranded DNA molecule.
[0165] In some embodiments, the nucleic acid molecular target sequence includes a sequence associated with plant traits or expression.
[0166] In some implementations, the nucleic acid molecular target sequence includes sequences or point mutations associated with a disease or condition.
[0167] In some embodiments, the base editing system contacts the target sequence of a nucleic acid molecule to perform deamination, which results in the substitution of one or more nucleotides in the target sequence.
[0168] In some embodiments, the target sequence comprises the DNA sequence 5'-MCN-3', where M is A, T, C, or G; N is A, T, C, or G; and the C in the middle of the 5'-MCN-3' sequence is deamination.
[0169] In some embodiments, the deamination process results in the introduction or removal of splice sites.
[0170] In some embodiments, the deamination results in the introduction of a mutation in the gene promoter, which leads to an increase or decrease in transcription of a gene operatively linked to the gene promoter.
[0171] In some embodiments, the deamination results in the introduction of a mutation in the gene repressor, the mutation leading to an increase or decrease in transcription of a gene operatively linked to the gene repressor.
[0172] In some implementations, the contact takes place within a living organism.
[0173] In some implementations, the contact is performed outside the biological body.
[0174] VI. Methods for producing genetically modified cells
[0175] In another aspect, the present invention also provides a method for generating at least one genetically modified cell, comprising introducing the base editing system of the present invention into at least one said cell, thereby causing substitution of one or more nucleotides in a target nucleic acid region of said at least one cell. In some embodiments, said one or more nucleotide substitutions are C to T substitutions.
[0176] In some embodiments, the method further includes the step of screening cells from the at least one cell for cells having one or more desired nucleotide substitutions.
[0177] In some embodiments, the method of the present invention is performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.
[0178] In another aspect, the present invention also provides genetically modified organisms comprising genetically modified cells or their progeny cells produced by the methods of the present invention. Preferably, the genetically modified cells or their progeny cells have one or more desired nucleotide substitutions.
[0179] In this invention, the target nucleic acid region to be modified can be located anywhere in the genome, such as within a functional gene like a protein-coding gene, or in a gene expression regulatory region such as a promoter or enhancer region, thereby achieving modification of the gene function or gene expression. In some embodiments, the desired nucleotide substitution leads to the desired modification of gene function or gene expression.
[0180] In some embodiments, the target nucleic acid region is associated with a trait of the cell or organism. In some embodiments, mutations in the target nucleic acid region lead to an alteration in the trait of the cell or organism. In some embodiments, the target nucleic acid region is located in a coding region of a protein. In some embodiments, the target nucleic acid region encodes a function-related motif or domain of the protein. In some preferred embodiments, substitution of one or more nucleotides in the target nucleic acid region results in amino acid substitution in the amino acid sequence of the protein. In some embodiments, the substitution of one or more nucleotides leads to an alteration in the function of the protein.
[0181] In the method of the present invention, the base editing system can be introduced into cells using various methods well known to those skilled in the art.
[0182] Methods for introducing the base editing system of the present invention into cells include, but are not limited to: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun method, PEG-mediated protoplast transformation, and Agrobacterium tumefaciens-mediated transformation.
[0183] Cells that can be base-edited by the method of the present invention can be derived from, for example, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, preferably crop plants, including but not limited to wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, tobacco, cassava, and potatoes.
[0184] VII. Applications in Plants
[0185] The base-editing fusion protein, base-editing system, and method for generating genetically modified cells of the present invention are particularly suitable for genetic modification of plants. Preferably, the plant is a crop plant, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato. More preferably, the plant is rice.
[0186] In another aspect, the present invention provides a method for producing genetically modified plants, comprising introducing the base editing system of the present invention into at least one of the plants, thereby causing substitution of one or more nucleotides in a target nucleic acid region in the genome of the at least one plant.
[0187] In some embodiments, the method further includes screening from the at least one plant for plants having one or more desired nucleotide substitutions.
[0188] In the method of this invention, the base editing composition can be introduced into plants using various methods well known to those skilled in the art. Methods that can be used to introduce the base editing system of this invention into plants include, but are not limited to: gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the base editing composition is introduced into plants via transient transformation.
[0189] In the method of this invention, modification of the target sequence can be achieved simply by introducing or generating the base-editing fusion protein and guide RNA in plant cells, and the modification can be stably inherited without the need for stable transformation of the plant with exogenous polynucleotides encoding the components of the base-editing system. This avoids the potential off-target effects of a stably existing (continuously generated) base-editing composition and also avoids the integration of exogenous nucleotide sequences into the plant genome, thus providing higher biosafety.
[0190] In some preferred embodiments, the introduction is performed without selection pressure, thereby avoiding the integration of exogenous nucleotide sequences into the plant genome.
[0191] In some embodiments, the introduction includes converting the base editing system of the present invention into isolated plant cells or tissues, and then regenerating the converted plant cells or tissues into complete plants. Preferably, the regeneration is performed without selection pressure, that is, without using any selection agents targeting the selection genes carried on the expression vector during tissue culture. Not using selection agents can improve the regeneration efficiency of the plants, resulting in modified plants free of exogenous nucleotide sequences.
[0192] In other embodiments, the base editing system of the present invention can be transferred to specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young spikelets, or hypocotyls. This is particularly suitable for the transfer of plants that are difficult to regenerate through tissue culture.
[0193] In some embodiments of the present invention, in vitro expressed proteins and / or in vitro transcribed RNA molecules (e.g., the expression construct is an in vitro transcribed RNA molecule) are directly transformed into the plant. The proteins and / or RNA molecules enable base editing in plant cells and are subsequently degraded by the cells, avoiding the integration of exogenous nucleotide sequences into the plant genome.
[0194] Therefore, in some embodiments, using the methods of the present invention to genetically modify and breed plants can yield plants whose genomes are free of foreign polynucleotide integration, i.e., non-transgene-free modified plants.
[0195] In some embodiments of the invention, the modified target nucleic acid region is associated with plant traits such as agronomic traits, whereby the substitution of one or more nucleotides results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.
[0196] In some embodiments, the method further includes the step of screening plants having one or more desired nucleotide substitutions and / or desired traits such as agronomic traits.
[0197] In some embodiments of the invention, the method further includes obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring has one or more desired nucleotide substitutions and / or desired traits such as agronomic traits.
[0198] In another aspect, the present invention also provides genetically modified plants or their offspring or portions thereof, wherein said plants are obtained by the methods described above. In some embodiments, the genetically modified plants or their offspring or portions thereof are non-GMO. Preferably, the genetically modified plants or their offspring have the desired genetic modification and / or desired traits such as agronomic traits.
[0199] In another aspect, the present invention also provides a plant breeding method, comprising crossing a genetically modified first plant, obtained by the method described above, containing one or more nucleotide substitutions in a target nucleic acid region, with a second plant not containing the one or more nucleotide substitutions, thereby introducing the one or more nucleotide substitutions into the second plant. Preferably, the genetically modified first plant has desired traits such as agronomic traits.
[0200] VIII. Therapeutic Applications
[0201] This invention also covers the application of the base editing system of this invention in disease treatment.
[0202] By modifying disease-related genes using the base editing system of this invention, it is possible to achieve upregulation, downregulation, inactivation, activation, or mutation correction of disease-related genes, thereby achieving disease prevention and / or treatment. For example, the target nucleic acid region described in this invention can be located within the protein-coding region of the disease-related gene, or, for example, within gene expression regulatory regions such as promoter regions or enhancer regions, thereby enabling modification of the function or expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes modification of the disease-related gene itself (e.g., protein-coding region), as well as modification of its expression regulatory regions (e.g., promoters, enhancers, introns, etc.).
[0203] "Disease-associated" genes are any genes that produce transcriptional or translational products at abnormal levels or in abnormal forms in cells derived from tissues affected by a disease, compared to tissues or cells from non-disease control groups. In cases where altered expression is associated with the onset and / or progression of the disease, it can be a gene expressed at abnormally high levels; it can also be a gene expressed at abnormally low levels. Disease-associated genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or linked to one or more genes responsible for the etiology of the disease in disequilibrium. Such mutations or genetic variations are, for example, single nucleotide variants (SNVs). The transcribed or translated products can be known or unknown and can be at normal or abnormal levels.
[0204] Therefore, the present invention also provides a method for treating a disease in a subject of need, comprising delivering an effective amount of the base editing system of the present invention to the subject to modify a gene associated with the disease (e.g., deamination of mitochondrial DNA via a fusion protein or multiple fusion proteins). The present invention also provides the use of the base editing system in the preparation of a pharmaceutical composition for treating a disease in a subject of need, wherein the base editing system is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject of need, comprising the base editing system of the present invention, and optionally a pharmaceutically acceptable vector, wherein the base editing system is used to modify a gene associated with the disease.
[0205] In some embodiments, the fusion protein or base editing system described in this invention is used to introduce point mutations into nucleic acids by deaminating a target nucleobase (e.g., a C residue). In some embodiments, the deamination of the target nucleobase results in the correction of a genetic defect, such as in the correction of a point mutation that results in loss of function in a gene product. In some embodiments, the genetic defect is associated with a disease or condition (e.g., lysosomal storage disease or metabolic disease, such as, for example, type 1 diabetes). In some embodiments, the methods provided herein can be used to introduce inactive point mutations into a gene or allele encoding a gene product associated with a disease or condition.
[0206] In some embodiments, the purpose of the schemes described in this invention is to restore the function of dysfunctional genes via genome editing. The nucleobase editing proteins provided herein are intended for use in vitro gene editing in human cells, such as correcting disease-related mutations in human cell cultures. The nucleobase editing proteins provided herein, such as fusion proteins containing nucleic acid-editable DNA proteins (e.g., CRISPR effector protein Cas9) and cytosine deaminase domains, can be used to correct any single-point T-to-C or A-to-G mutation. In the first case, the mutant C corrects the mutation via deamination, while in the latter case, the C paired with mutant A corrects the mutation via deamination and a subsequent round of replication.
[0207] In some embodiments, the purpose of the schemes described in this invention is to treat diseases associated with or caused by point mutations, which can be corrected by the DNA base-editing fusion proteins provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neonatal disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.
[0208] In some embodiments, the purposes of the solutions described in this invention are for the treatment of mitochondrial diseases or disorders. As used herein, "mitochondrial disease" refers to diseases caused by abnormal mitochondria, such as mitochondrial gene mutations, enzyme pathways, etc. Examples of diseases include, but are not limited to: neurological disorders, loss of motor control, muscle weakness and pain, gastrointestinal disorders and dysphagia, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, visual / hearing problems, lactic acidosis, developmental delay, and susceptibility to infection.
[0209] Examples of diseases described in this invention include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous system and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat amplification disorders, hearing disorders, gene-targeted therapy for non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, β-thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, and schizophrenia. Other diseases that can be treated by correcting point mutations or introducing inactive mutations into disease-related genes are known to those skilled in the art, and therefore this disclosure is not limited in this respect. In addition to the diseases exemplarily described in this invention, other related diseases can also be treated with the strategies and fusion proteins provided in this invention, and this application will be apparent to those skilled in the art. The diseases or targets to which this invention can be applied refer to the base editing systems listed in WO2015089465A1 (PCT / US2014 / 070135), WO2016205711A1 (PCT / US2016 / 038181), WO2018141835A1 (PCT / EP2018 / 052491), WO2020191234A1 (PCT / US2020 / 023713), WO2020191233A1 (PCT / US2020 / 023712), WO2019079347A1 (PCT / US2018 / 056146), and WO2021155065A1 (PCT / US2021 / 015580) for applicable diseases.
[0210] The administration of the base editing system or pharmaceutical composition of the present invention can be tailored to the patient's or subject's weight and species. The frequency of administration is within medically or veterinary limits. It depends on conventional factors including the patient's or subject's age, sex, general health condition, other conditions, and the specific symptom or condition being addressed.
[0211] 9. Adeno-associated virus (AAV)
[0212] The base-editing fusion protein and / or expression constructs containing nucleotide sequences encoding the base-editing fusion protein provided by this invention, or one or more gRNAs containing the base-editing system of this invention, can be delivered using adeno-associated virus (AAV), lentivirus, adenovirus, or other plasmid or viral vector types. Due to the 4.5-4.75 kb packaging limitation of AAV, both the promoter and transcription terminator must be contained within the same viral vector. Constructs larger than 4.5-4.75 kb will result in a significant reduction in viral delivery efficiency. Cytosine deaminases are relatively large, making them difficult to package into AAV. Therefore, embodiments of this invention provide the use of truncated cytosine deaminases packaged into AAV to achieve base editing.
[0213] 10. Nucleic Acids, Cells, and Compositions
[0214] In another aspect, the present invention provides a nucleic acid molecule that encodes the cytosine deaminase of the present invention, or the fusion protein of the present invention.
[0215] In another aspect, the present invention provides a cell comprising the cytosine deaminase of the present invention, or the fusion protein of the present invention, or the base editing system of the present invention, or the nucleic acid molecule of the present invention.
[0216] In another aspect, the present invention provides a composition comprising the cytosine deaminase of the present invention, or the fusion protein of the present invention, or the base editing system of the present invention, or the nucleic acid molecule of the present invention.
[0217] In some embodiments, the cytosine deaminase, fusion protein, base editing system, or nucleic acid molecule is packaged into a virus, virus-like particle, virion, liposome, vesicle, exosome, or liposome nanoparticle (LNP).
[0218] In some implementations, the virus is described as adeno-associated virus (AAV) or recombinant adeno-associated virus (rAAV).
[0219] XI. Reagent Kit
[0220] This invention also includes kits for use with the methods of the invention, the kits comprising the base-editing fusion protein of the invention and / or expression constructs containing a nucleotide sequence encoding said base-editing fusion protein, or comprising the base-editing system of the invention. The kits generally include labels indicating the intended use and / or method of use of the kit contents. Terminology labels include any written or documented material provided on or with the kit or otherwise with the kit. The kits of the invention may also contain suitable materials for constructing expression vectors in the base-editing systems of the invention. The kits of the invention may also contain reagents suitable for converting the base-editing fusion protein or base-editing composition of the invention into cells.
[0221] In one aspect, the present invention provides a kit containing a nucleic acid construct, wherein the nucleic acid construct comprises:
[0222] (a) The nucleic acid sequence encoding the cytosine deaminase of the present invention; and
[0223] (b) Heterogeneous promoters that drive the expression of the sequence in (a).
[0224] In one aspect, the present invention provides a kit containing a nucleic acid construct, wherein the nucleic acid construct comprises:
[0225] (a) The nucleic acid sequence encoding the fusion protein of the present invention; and
[0226] (b) Heterogeneous promoters that drive the expression of the sequence in (a).
[0227] In some embodiments, the embodiment further includes an expression construct encoding a guide RNA backbone, wherein the construct contains a cloning site that allows the cloning of a nucleic acid sequence that is identical to or complementary to the target sequence into the guide RNA backbone. Example
[0228] To facilitate understanding of the present invention, a more complete description will be given below with reference to specific embodiments and accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.
[0229] Materials and Methods
[0230] 1. Carrier Construction
[0231] The sequence of the novel deaminase identified was optimized using rice and wheat double codons by Nanjing GenScript Biotech Co., Ltd., and constructed into the pJIT63-nCas9-PBE backbone (Addgene #98164). The plasmid of the reporter system used in the examples was previously constructed in our laboratory.
[0232] For sgRNA, expression was performed using the pOsU3 vector (Addgene #170132).
[0233] 2. Protoplast isolation and transformation
[0234] The protoplasts used in this invention are derived from the rice variety Zhonghua 11.
[0235] 2.1 Rice seedling cultivation
[0236] Rice seeds were first rinsed with 75% ethanol for 1 minute, then treated with 4% sodium hypochlorite for 30 minutes, and washed with sterile water at least 5 times. They were then cultured on M6 medium for 3-4 weeks at 26°C in the dark.
[0237] 2.2 Protoplast Isolation
[0238] (1) Cut off the rice stalks, cut the middle part into 0.5-1mm shreds with a blade, put them into 0.6M Mannitol solution and treat them in the dark for 10min, then filter them with a filter screen, put them into 50mL of enzymatic hydrolysis solution (0.45μm filter membrane), vacuum (pressure about 15Kpa) for 30min, take them out and place them on a shaker (10rpm) at room temperature for 5h of enzymatic hydrolysis.
[0239] (2) Add 30-50 mL of W5 to dilute the enzyme digestion product, and filter the enzyme digestion solution into a round-bottom centrifuge tube (50 mL) using a 75 μm nylon filter membrane;
[0240] (3) Centrifuge at 23℃, 250g (rcf), increase by 3°C and decrease by 3°C for 3 min, and discard the supernatant;
[0241] (4) Gently suspend the cells in 20 mL of W5 solution and repeat step (3).
[0242] (5) Add an appropriate amount of MMG to suspend the mixture and wait for conversion.
[0243] 2.3 Rice protoplast transformation
[0244] (1) Add 10 μg of each of the required transformation vectors to a 2 mL centrifuge tube, mix well, then use a de-pointed pipette tip to take 200 μL of protoplasts, gently tap to mix, add 220 μL of PEG4000 solution, gently tap to mix, and induce transformation at room temperature in the dark for 20-30 min.
[0245] (2) Add 880 μL W5 and gently invert to mix. Centrifuge at 250 g (rcf) for 3 min, then centrifuge at 3°C and 3°C for 3 min. Discard the supernatant.
[0246] (3) Add 1 mL of WI solution, gently invert to mix, gently transfer to a flow cytometer, and incubate at room temperature in the dark for 48 hours.
[0247] 3. Observe cell fluorescence using flow cytometry.
[0248] Protoplast GFP-negative and GFP-positive populations were analyzed using a FACSAria III (BD Biosciences) instrument.
[0249] 4. Protoplast and plant DNA extraction and amplicon sequencing analysis
[0250] Protoplasts were collected in 2 mL centrifuge tubes, and protoplast DNA (~30 μL) was extracted using the CTAB method. The concentration of DNA was determined using a NanoDrop micro-spectrophotometer (30-60 ng / μL), and the samples were stored at -20 °C.
[0251] PCR amplification of protoplast DNA templates was performed using genomic primers specifically targeting the target sites. A 20 μL amplification system contained 4 μL 5×FastPfu buffer, 1.6 μL dNTPs (2.5 mM), 0.4 μL forward primer (10 μM), 0.4 μL reverse primer (10 μM), 0.4 μL FastPfu polymerase (2.5 U / μL), and 2 μL DNA template (~60 ng). Amplification conditions: 95℃ pre-denaturation for 5 min; 95℃ denaturation for 30 s, 50-64℃ annealing for 30 s, 72℃ extension for 30 s, 35 cycles; 72℃ final extension for 5 min; storage at 12℃.
[0252] The amplified product was diluted 10-fold, and 1 μL was used as the template for the second round of PCR amplification. The amplification primers were sequencing primers containing barcodes. The 50 μL amplification system contained 10 μL 5×Fastpfu buffer, 4 μL dNTPs (2.5 mM), 1 μL forward primer (10 μM), 1 μL reverse primer (10 μM), 1 μL FastPfu polymerase (2.5 U / μL), and 1 μL DNA template. The amplification conditions were as described above, and the number of amplification cycles was 35.
[0253] PCR products were separated by 2% agarose gel electrophoresis, and the target fragments were recovered using the AxyPrep DNA Gel Extraction kit. The recovered products were quantitatively analyzed using a NanoDrop micro-volume spectrophotometer. 100 ng of the recovered products were mixed and sent to Novogene for amplicon sequencing library construction and amplicon sequencing analysis.
[0254] 5. Human and animal cell transfection
[0255] Human HEK293T cells (ATCC, CRL-3216) and mouse N2a cells (ATCC, CCL131) were cultured in Dulbecco's Modified Eagle's medium (DMEM, Gibco) supplemented with 10% (vol / vol) fetal bovine serum (FBS, Gibco) and 1% (vol / vol) penicillin-streptomycin (Gibco) at 37°C in a humidified incubator with 5% CO2. All cells were routinely tested for mycoplasma contamination using a mycoplasma detection kit (Transgen Biotech). In the absence of antibiotics, cells were seeded into 48-well Poly-D-Lysinecoated plates (Corning). After 16–24 hours, cells were incubated with 1 μL Lipofectamine 2000 (Thermo Fisher Scientific), 300 ng deaminase vector, and 100 ng sgRNA expression vector. For transfection with the cytosine base editing system, cells were incubated with 1 μL Lipofectamine 2000, 300 ng TALE-L, and 300 ng TALE-R. After 72 hours, cells were washed with PBS, and DNA was extracted. To detect off-target effects using the R-loop method, cells were co-transfected with the vectors BE4max, SaCas9BE4max, and their corresponding sgRNA vectors (Koblan, LW, Doman, JL, Wilson, C., Levy, JM, Tay, T., Newby, GA, Maianti, JP, Raguram, A., & Liu, DR).
[0256] (2018). Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat. Biotechnol., 36, 843–846.).
[0257] 6. TRAPseq Library
[0258] We evaluated the performance of the deaminase base editing system using an sgRNA 12K-TRAPseq library. Twenty hours before viral transduction, we seeded 2 × 10⁶ cells into 100 mm culture dishes. We transduced 500 μL of sgRNA lentivirus. For stably integrated cells, we selected them using 1 μg / mL puromycin (Gibco). For each base editor, we seeded 2 × 10⁶ cells into 6 culture dishes 24 hours before transfection. We transfected each CBE member plasmid DNA (15 μg) and Tol₂ DNA with 60 μL Lipofectamine 2000. Twenty-four hours after transfection, we replaced the medium with fresh medium containing 10 μg / mL blasticidin (Gibco). Three days later, we washed, resuspended, and seeded the cells into medium containing 10 μg / mL blasticidin. Six days later, we washed all cells with PBS, then centrifuged and extracted DNA using the Cell / Tissue DNA Isolation Mini Kit (Vazyme). For each deaminase base editor sample, we identified the sequences using next-generation sequencing.
[0259] 7. DNA extraction
[0260] For HEK293T and N2a cells, genomic DNA was extracted using Lysis Buffer and Proteinase K treatment and Triumfi mouse tissue direct amplification kit (Beijing Jinsha Biotechnology).
[0261] For plant protoplasts, genomic DNA was extracted using a plant genomic DNA kit (Tiangen Biotech) after 72 hours of culture. All DNA samples were quantified using a NanoDrop 2000 spectrophotometer (Thermo Fisher Scientific).
[0262] 8. Protein structure analysis and clustering
[0263] Protein structure analysis was performed using AlphaFold v2.2.0 (John Jumper and others, 'Highly Accurate Protein Structure Prediction with AlphaFold', Nature, 596.7873 (2021), 583–89).
[0264] The TM-align software was used to calculate the TM-score based on the analysis results. The specific formula for calculating the TM-score is as follows (Reference: Zhang, Yang, and Jeffrey Skolnick. (2004). Scoring function for automated assessment of protein structure template quality. Proteins 57(4), 702-710.)
[0265] Where L N It is the length of the target protein's amino acid sequence; LT is the amino acid that appears simultaneously in both the template and the target structure.
[0266]
[0267] The sequence length, di is the distance between the i-th pair of residues in the template and the target structure, and d0 is the scale of the normalized matching difference. "Max" represents the maximum value after optimal spatial superposition.
[0268] After converting the TM-Score, the APE and phangorn packages in R were used to perform clustering calculations using the UPGMA method (CPKurtzman, Jack W. Fell, and T. Boekhout, The Yeasts: A Taxonomic Study, 5th ed (Amsterdam: Elsevier, 2011); 'A Statistical Method for Evaluating Systematic Relationships-Robert Reuven Sokal, Charles Duncan Michener-Google Books'). First, the distance between any two pairs of data was obtained using the following formula.
[0269]
[0270] Where d (ABX) The distance between two of these points is shown.
[0271] Then, during the clustering process, the formula for calculating the average distance is as follows. If C1 and C2 are terminal taxa containing sets n1 and n2 respectively, which will be merged into the new set C, then the average distance to any other cluster D is calculated by the following formula:
[0272]
[0273] Example 1: Identification of novel deaminases in the APOBEC / AID branch that can be used for base editing
[0274] To find novel deaminases different from those used in existing base editing systems, we first tested deaminases from the APOBEC / AID branch that showed low sequence similarity to existing deaminases, based on a list of representative deaminases provided by Iyer et al. (Iyer, LM, Zhang, D., Rogozin, IB, & Aravind, L. (2011). Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems. Nucleic acids research, 39(22), 9473-9497.). Among them, deaminase No. 182 (SEQ ID NO: 1) showed very low similarity to existing deaminases, with its amino acid sequence sharing only 34% with the most similar mouse rAPOBEC1. Deaminase No. 182 was constructed onto the pJIT163-nCas9-PBE backbone, i.e., deaminase No. 182 was used to replace rAPOBEC1 in the fusion with nCas9. Evaluation using a report system revealed that 182-PBE can undergo base editing in cells. Figure 1 , Figure 6 and Figure 7 ).
[0275] To further confirm its editing capabilities, the 182-PBE construct and a sgRNA-targeting construct were co-transformed into rice protoplasts. Analysis of the editing results at six endogenous sites revealed that 182-PBE can effectively achieve base editing, and its editing window is significantly larger than that of commonly used rAPOBEC1-based cytosine base editing systems. Figure 2 , Figure 6 and Figure 7 Therefore, protein 182 has the function of deamination of cytosine in single-stranded DNA, and a novel cytosine base editing system can be established based on this protein.
[0276] Example 2: Detection of cytosine deamination activity of deaminases in different branches
[0277] Iyer et al. searched for proteins in databases with folding patterns similar to known deaminases and divided these proteins into at least 21 clades based on their domains (Table 1). Cytosine deaminases APOBEC1, APOBEC3, AID, and CDA1, which are currently widely used for base editing, were all classified into the APOBEC / AID-like clade. In addition to these clades, there are also clades with proven functions, such as "dCMP," which can convert dCMP to dUMP.
[0278] The "deaminase and ComE" branch, which converts guanine (G) to xanthine (I), is a type of guanine.
[0279] The deaminase branch includes the "RibD-like" branch, which functions as a diaminohydroxyphosphonoribonucleoylpyrimidine deaminase; the "Tad1 / ADAR" branch, which functions as an RNA editing enzyme that converts RNA adenine (A) to xanthine (I); and the "PurH / AICAR transformylase" branch, which has formyltransferase activity. In addition, the deaminase functions of some branches are still unclear, such as the bacterial branches named based on protein domains: SCP1.201, XOO2897, MafB19, and Pput_2613.
[0280] Table 1. Deaminase families (Iyer et al., 2011)
[0281]
[0282]
[0283] To detect whether the above branches possess cytosine deaminase activity, 48 representative deaminase proteins, excluding the APOBEC / AID branch, were selected from the list of representative deaminases compiled by Iyer et al., distributed across 14 branches: Bd3614, CDD / CDA-like, DYW-like, FdhD, MafB19, Novel AID / APOBEC-like, OTT1508, PurH / AICARtransformylase, RibD-like, TM1506, SCP1.201, Imm1immunity protein associated with SCP1.201 deaminases, YwqJ, and XOO2897. All proteins were constructed onto the pJIT163-nCas9-PBE backbone, and their deaminase activity by binding to single-stranded DNA was evaluated using the BFP-to-GFP reporter system (Zong, Y. et al. Nat. Biotechnol. 35, 438-440 (2017)). Five branches containing a total of 23 proteins were found to possess cytosine deaminase activity. These branches originated from the Novel AID / APOBEC-like branch (No.2-1479 and No.2-1478), and the bacterial branches SCP1.201 (No.69, No.55, No.57, No.64, No.76, No.2-1146, No.2-1160, No.54, No.56, No.59, No.60, No.61, No.72, No.74, No.75, No.63, No.2-1158), XOO2897 (No.2-1429, No.2-1442), TM1506 (No.2-39), and MafB19 (No.101m). Specifically, for the SCP1.201 branch, 18 out of 19 proteins tested showed cytosine deaminase activity. Through testing at two endogenous sites in rice, the aforementioned 23 proteins with cytosine deaminase activity were classified into 8 highly efficient (…) Figure 4 and Figure 5 ), 8 medium ( Figure 6 and Figure 7 ) and 7 deaminases with low editing efficiency.
[0284] To further confirm the editing capabilities of the newly discovered deaminase, candidate deaminase No. 69, belonging to the SCP1.201 branch, was selected from the group that enabled reporter luminescence. To further confirm its editing ability, the 69-PBE construct and an endogenous sgRNA construct were co-transformed into rice protoplasts. Analysis of the editing results at six endogenous sites revealed that 69-PBE can effectively achieve base editing, and its editing efficiency is significantly higher than that of commonly used rAPOBEC1-based cytosine base editing systems. Figure 3 Therefore, the newly identified proteins may have the function of deamination of cytosine from single-stranded DNA, and novel cytosine base editing systems can be established based on these proteins.
[0285] Example 3: Discovery of novel cytosine deaminases through protein structure analysis, clustering, and differentiation.
[0286] Based on the above embodiments, effective methods for protease function identification and screening are needed to efficiently analyze protein functions. Given the determining role of protein three-dimensional structure in its function, comparative analysis and clustering of known or predicted protein structures may be an effective method for classifying deaminases into functional branches. Therefore, we combine AI-assisted protein structure prediction, structure calibration, and clustering to generate new protein classification relationships among deaminases. Figure 8 ).
[0287] We selected 238 protein sequences annotated with deaminase domains and 4 candidate protein sequences from distant outgroups of the JAB-domain family from the InterPro database. Figure 9 Specifically, we selected 15 candidate genes with a length of at least 100 amino acids from each of the 16 deaminase families and used AlphaFold2 to predict their protein structures. We used the standardized scoring model TM-score to perform multiple structural alignments (MSA) on all candidate proteins. The specific formula for calculating TM-score is as follows (Reference: Zhang, Yang, and Jeffrey Skolnick. (2004). Scoring function for automated assessment of protein structure template quality. Proteins 57(4), 702-710.)
[0288]
[0289] Where L NIt is the length of the target protein's amino acid sequence, L T It is the length of the amino acid sequence that appears simultaneously in both the template and the target structure, d i d is the distance between the i-th pair of residues in the template and the target structure, and d0 is the scale of the normalized matching difference. “Max” represents the maximum value after optimal spatial superposition.
[0290] Based on the MSA results, we generated structural similarity matrices reflecting the overall structural correlations between proteins. Then, we used the weighted pairwise pairing algorithm (UPGMA) to organize these similarity matrices into a structure-based dendrogram. Figure 10 The dendrogram groups 238 proteins into 20 unique structural branches, each branch containing deaminases with different conserved protein domains. Figure 11A and 11B We found that even without using contextual information, such as conserved gene neighborhood and domain architectures, accurate protein clustering classifications can be generated based on protein structure. When using structure-based hierarchical clustering, different branches reflect unique structures, implying different catalytic functions and properties. Figure 11A and 11B Interestingly, we also found that this structure-based clustering method is more effective than traditional one-dimensional amino acid sequence clustering methods in ranking functional similarity. For example, in the amino acid sequence-based clustering method, adenine deaminases involved in purine metabolism (A_deamin, PFO2137 in the InterPro database) were divided into different branches, while in the structure-based clustering method, they were grouped into a single deaminase branch.
[0291] Furthermore, we used a structure-based clustering method to divide the four deaminase families (dCMP, MafB19, LmjF365940, and APOBEC (labeled by InterPro)) into two independent branches. Figure 11A and 11B Comparison of protein structures reveals that the two branches of these four deaminase families have distinctly different structures, which may contradict their InterPro nomenclature and sequence-based classification. Figure 11B and Figure 12 In summary, protein clustering and classification based on AI-assisted protein three-dimensional structure provides reliable clustering results and requires only one amino acid sequence without any other genomic reasoning, making it a more convenient and effective strategy for generating protein relationships than other methods.
[0292] Example 4: Identifying the function of deaminases available for base editing in the SCP1.201 branch using a three-dimensional structure tree.
[0293] In Example 2, by evaluating the function of deaminases in each branch, we were surprised to find that some deaminases in the SCP1.201 branch possess the ability to catalyze the deamination of single-stranded DNA substrates. Previously, these deaminases were annotated in the InterPro database (PF14428) as double-stranded DNA deaminase toxin A-like (DddA-like) deaminases. Among them, DddA enzyme is a deaminase recently applied to a non-CRISPR double-stranded DNA cytosine base editor (CRISPR-free double-stranded DNA cytosine base editor, DdCBE), which can be used to deaminate cytosine bases in double-stranded DNA (NCBI Reference Sequence: WP_006498588.1)(BYMok,MHde Moraes,J.Zeng,DEBosch,AVKotrys,A.Raguram,F.Hsu,MCRadey,SBPeterson,VKMootha,JDMougous,DRLiu,Abacterial cytidine deaminase toxin enables CRISPR-free mitochondrial baseediting.Nature 583,631-637(2020).). It is precisely because of the existence of DddA that all proteins in its SCP1.201 branch are annotated as double-stranded DNA deaminases (Ddd).
[0294] Based on this problem, we performed the following work using protein function prediction based on the three-dimensional structures in Example 3. To reanalyze this SCP1.201 branch, we selected all 489 SCP1.201 deaminases from the InterPro database. We also included seven other proteins that showed 35% to 50% similarity to DddA in BLAST alignment, but were described separately in InterPro. After identification and coverage screening, we performed a novel AI-assisted protein structure classification on 332 SCP1.201 deaminases. The structural clustering analysis showed that SCP1.201 deaminases clustered into different sub-branches, each with its own unique core structural domain motif. Figure 13A -E).
[0295] Importantly, we found that DddA and 10 other proteins clustered in a subbranch of the same SCP1.201. Analysis of the 3D predicted structures of all 11 proteins in this subbranch revealed that they share a similar core structure with DddA. Given this structural similarity to DddA, we predict that other proteins in this subbranch also possess double-stranded DNA cytosine deamination capabilities.
[0296] Example 5: Validation of DddA-containing deaminases in animal cell base editing.
[0297] To evaluate whether the SCP1.201 candidate protein of the subbranch containing DddA obtained using the prediction method of this invention in Example 4 has functional similarity to DddA, i.e., it has deamination activity on dsDNA, we designed a DdCBE formed by each deaminase in this subbranch individually, or by separating the deaminase into two parts by dividing it into two allele sites similar to the DddA structure and linking them together with a double TALE system to form a DdCBE (see: BYMok, MHde Moraes, J. Zeng, DEBosch, AVKotrys, A. Raguram, F. Hsu, MCRadey, SBPeterson, VKMootha, JDMougous, DRLiu, A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing. Nature 583, 631-637 (2020). Figure 14 (Table 2). We evaluated proteins from this Ddd subbranch at the JAK2 and SIRT6 sites in HEK293T cells and observed 13 proteins capable of dsDNA base editing (Table 2). Hereinafter, we will name these deaminases double-stranded DNA deaminases (Ddd) and classify them into this newly discovered Ddd subbranch.
[0298] Table 2. Catalytic activity of proteins in the subbranch containing DddA
[0299] name GeneID SEQ ID No. dsDNA catalytic activity Ddd1 SCP-177 28 ++ Ddd2 / 29 + Ddd3 SCP001 30 ++ Ddd4 / 31 + Ddd6 SCP009 32 + Ddd7 SCP103 33 ++ Ddd8 SCP234 34 ++ Ddd9 SCP003 35 ++ Ddd10 SCP004 36 + Ddd11 SCP271 37 + Ddd12 SCP005 38 + Ddd13 SCP006 39 ++ Ddd14 SCP007 40 +
[0300] The symbol ++ indicates strong catalytic activity, + indicates weak catalytic activity, and - indicates no catalytic activity.
[0301] Example 6: Validation of DddA-free deaminases in base editing of plant and animal cells.
[0302] In comparison, this experiment further evaluated the deamination of other SCP1.201 candidate proteins that do not contain the DddA subbranch. We randomly selected 24 proteins and placed them into our CBE fluorescent reporter system. We found that 22 of these proteins showed detectable fluorescence, and selected 13 of them to evaluate base editing at endogenous sites under CBE conditions in mammalian cells. Figure 16A (Table 3). Although these proteins were previously annotated as DddA-like proteins, experimental results show that they only exhibit cytosine base editing activity on ssDNA. Figure 13A , Figure 16A (and Table 3), but no activity was shown against dsDNA (and Table 3). Figure 16B Based on their function and role, in subsequent work we will name these proteins from the SCP1.201 branch that have ssDNA targeting activity as single-stranded DNA deaminases (Sdd).
[0303] Based on the above experimental results, we were surprised to find that most of the protein members from the SCP1.201 branch were Sdd proteins, rather than the DddA-like proteins annotated in the InterPro database (PF14428). We also observed that these Sdd proteins have structures similar to each other but clearly distinct from Ddd proteins, such as the structure of Sdd7 (Figures 13D, 13E). Sdd7 is one of the most efficient cytosine base editors for ssDNA editing. Therefore, the method of this invention demonstrates that the DddA-like deaminases already annotated in the InterPro database (PF14428) should be further subdivided and re-annotated accordingly.
[0304] As a control, we also clustered proteins from the SCP1.201 branch based on their one-dimensional amino acid sequences and validated the tree structure with JAB outgroups, finding that JAB outgroup members were scattered throughout the tree. These results demonstrate the effectiveness and importance of using protein structure-based classification to compare and assess protein relationships.
[0305]
[0306] Furthermore, a comprehensive analysis of the protein function verification results in Examples 5 and 6 shows that protein clustering classification based on artificial intelligence-assisted three-dimensional protein structures provides reliable clustering results. The three-dimensional structure tree constructed using the method of this invention can accurately identify and predict the detailed functions of proteins. Moreover, it requires only one amino acid sequence and no other genomic reasoning, making it a more convenient and effective protein relationship generation strategy than other methods. In the three-dimensional structure tree, when a TM-score of not less than 0.7 is used as the clustering condition, the clustering results predicting protein catalytic functions are consistent with the experimental verification conclusions (Table 4). That is, clustering with a TM-score of not less than 0.7 compared to the reference protein as a criterion results in subbranchings with the same or similar catalytic functions as the reference protein. The method of this invention significantly improves the identification and prediction efficiency.
[0307] Table 4. Comparison of sequence and structural similarity of some deaminases in the SCP1.201 family.
[0308]
[0309] Example 7: The novel Ddd protein has different editing preferences than DddA.
[0310] Due to DddA's strict preference for the 5'-TC motif, the application of DddA-based dsDNA base editors is mainly limited to TC targets (BYMok, MHde Moraes, J. Zeng, DEBosch, AVKotrys, A. Raguram, F. Hsu, MCRadey, SBPeterson, VKMootha, JDMougous, DRLiu, Abacterial cytidinedeaminase toxin enables CRISPR-free mitochondrial base editing. Nature 583, 631-637 (2020).). Although the recently evolved DddA11 has shown greater general applicability, enabling deamination of 5'-HC (H = A, C, or T) motifs for cytosine base editing, its editing efficiency for AC, CC, and GC targets still needs improvement (BYMok, AVKotrys, A. Raguram, T.P. Huang, V.K. Mootha, DR. Liu, CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol. 40, 1378-1387). We evaluated the newly discovered Ddd proteins of this invention to determine whether they could expand the efficiency and targeting scope of DdCBE. We constructed 13 deaminases belonging to the Ddd subbranch into DdCBE and evaluated dsDNA base editing at endogenous JAK2 and SIRT6 sites in HEK293T cells. Figure 15 , Figure 18 (and Table 2). Interestingly, we found that Ddd1, Ddd7, Ddd8, and Ddd9 had similar or higher editing efficiency compared to DddA (and Table 2). Figure 17 A and Figure 18 Importantly, we found that Ddd1 and Ddd9 exhibit significantly higher editing activity on the 5'-GC motif than DddA. Figure 17 A and Figure 18 Notably, in the C10 (5'-GC) residues of JAK2 and the C11 (5'-GC) residues of SIRT6, we found that the editing rates of DddA were only 21.1% and 0.6%, respectively, while the editing rates of Ddd9 were 65.7% and 45.7%, respectively. Figure 17 A).
[0311] Since some Ddd proteins appear to exhibit different editing modes compared to DddA, we sought to evaluate all motif preferences of these Ddd proteins. We first constructed several plasmids encoding the JAK2 target sequence (BYMok, AVKotrys, A. Raguram, TPHuang, VKMootha, DRLiu, CRISPR-free base editors with enhanced activity and expanded targeting scope in mitochondrial and nuclear DNA. Nat. Biotechnol. 40, 1378-1387), and inserted the G at positions 9-11. C C changed to M C Positions 9-11 of N (M / N = A, T, C, and G) were used to obtain 16 different plasmids, and each plasmid was co-transfected with the DdCBE variant. Figure 17 B). In the comparative analysis of each M C After determining the C·G-to-T·A base transition frequency of N, we generated corresponding motif logo diagrams to reflect the sequence context preference of each dsDNA deaminase. Figure 17 C). As mentioned earlier, we found that DddA and its structural homolog Ddd7 strongly favor the 5'-TC motif ( Figure 17 C Figure 19 Conversely, we found that Ddd1 and Ddd9 tend to edit 5'-G. C The substrate of the motif, while Ddd8 tends to edit 5'-W. C (W = A or T) motif substrates. Therefore, through further analysis of the Ddd sub-branch, we discovered a whole suite of novel Ddd proteins that can be used for editing different motifs. These proteins greatly expand the targeting range and practicality of DdCBE and show great application potential. Figure 17 C, Figure 19 ).
[0312] Example 8: Base Editing of Sdd Deaminase in Human Cells and Plants
[0313] Next, we wanted to know if the newly discovered Sdd proteins could also be used for more precise or efficient base editing. To this end, we selected six of the most active Sdd and four of the weaker ones for evaluation and compared their activities using a fluorescent reporter system (Table 3). We designed plant CBEs for each of the 10 Sdd and evaluated their endogenous base editing at six sites on rice protoplasts. Figure 20 and Figure 21We found that seven of these deaminases (Sdd7, Sdd9, Sdd5, Sdd6, Sdd4, Sdd76, and Sdd10) exhibited higher activity than rat APOBEC1 (rAPOBEC1)-based CBEs. The most active Sdd7 base editor showed a cytosine base editing rate of up to 55.6%, which is more than 3.5 times that of rAPOBEC1.
[0314] To test the universality of these deaminases, we also constructed corresponding human cell-targeting BE4max vectors (LWKoblan, JLDoman, C. Wilson, JMLevy, T. Tay, GANewby, JPMaianti, A. Raguram, DRLiu, Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat. Biotechnol. 36, 843-846 (2018).) and evaluated their editing efficiency at three endogenous targets in HEK293T cells. The results in HEK293T cells were consistent with those in rice, and we found that Sdd7 exhibited the highest editing activity (…). Figure 22 ).
[0315] We previously discovered that human APOBEC3A (A3A) has a large editing window in plants and exerts strong editing activity (Y. Zong, Q. Song, C. Li, S. Jin, D. Zhang, Y. Wang, J.-L. Qiu, C. Gao, Efficient C-to-T baseediting in plants using a fusion of nCas9 and human APOBEC3A. Nat. Biotechnol. 36, 950-953 (2018)., Q. Lin, Z. Zhu, G. Liu, C. Sun, D. Lin, C. Xue, S. Li, D. Zhang, C. Gao, Y. Wang, J.-L. Qiu, Genome editing in plants with MAD7 nuclease. J. Genet. Genomics 48, 444-451 (2021).). Therefore, we compared A3A and Sdd7 in human cells ( Figure 22 ) and plants ( Figure 23 Editing activity of Sdd7 in HEK293T cells. Interestingly, Sdd7 targets all three sites ( Figure 22) and five endogenous sites of rice protoplasts ( Figure 23 It exhibits editing activity comparable to A3A. These results confirm that Sdd7 is a powerful cytosine base editor, widely applicable to both plant and human cells.
[0316] Example 9: Sdd protein has unique base editing properties.
[0317] In evaluating endogenous base editing, we observed different editing patterns at genomic target sites tested in human and rice cells for different Sdd-CBEs. For example, while Sdd7, Sdd9, and Sdd6 did not show a specific motif editing preference, Sdd3 appeared to prefer editing 5'-GC and 5'-AC motifs, but strongly disliked editing 5'-TC and 5'-CCmotifs. Figure 24To better analyze the editing patterns of each deaminase, we used Targeted Reporter Anchored Positional Sequencing (TRAP-seq), a high-throughput method for parallel quantification of base editing results (Xi Xiang, Kunli Qu, Xue Liang, Xiaoguang Pan, Jun Wang, Peng Han, Zhanying Dong, Lijun Liu, Jiayan Zhong, Tao Ma, Yiqing Wang, Jiaying Yu, Xiaoying Zhao, Siyuan Li, Zhe Xu, Jinbao Wang, Xiuqing Zhang, Hui Jiang, Fengping Xu, Lijin Zou, Huaijing Teng, Xin Liu, Xun Xu, Jian Wang, Huanming Yang, Lars Bolund, George M. Church, Lin Lin, Yonglun Luo. (2020). Massively parallel quantification of CRISPR editing in cells by TRAP-seq enables better design of Cas9, ABE, CBE gRNAs of high efficiency and accuracy.bioRxiv 2020.05.20.103614. A 12K TRAP-seq library containing 12,000 TRAP structures was stably integrated into HEK293T cells via lentiviral transduction. Each TRAP structure contained a unique gRNA expression cassette and a corresponding alternative target site. After cell culture and antibody selection, a base editor was transiently transfected into this 12K-TRAP cell line, followed by 10 days of selection with puromycin and blastomycin (…). Figure 25A On day 11 post-transfection, we extracted genomic DNA and performed deep amplicon sequencing to evaluate the editing products of each deaminase. Figure 25A We found that while Sdd7 and Sdd6 did not exhibit strong sequence context preference, rAPOBEC1 showed a strong preference for 5'-TC and 5'-CC bases, but not for 5'-GC and 5'-AC bases. Figure 25B Conversely, Sdd3 exhibits a fully complementary pattern, favoring editing 5'-GC and 5'-AC bases, while showing little activity towards 5'-TC and 5'-CC bases. Figure 25BInterestingly, we found that Sdd6 and Sdd3 have different editing windows compared to rAPOBEC1 and Sdd7, and we preferred editing positions +1 to +3 on the far PAM end. Figure 25B In summary, the newly identified Sdd base editor exhibits unique base editing properties compared to traditional cytosine base editors, such as increased editing efficiency, different deamination preferences, and a modified editing window.
[0318] Example 10: High-fidelity editing properties of Sdd protein
[0319] Previous reports have suggested that CBE may lead to genome-wide Cas9-based off-target editing, raising concerns about the safety of these precise genome editing technologies for clinical applications. We believe these off-target mutations may result from the overexpression of cytidine deaminases. We wanted to know if the newly discovered Sdd proteins could provide a more favorable balance between off-target and targeted editing. Therefore, we evaluated the Cas9-independent off-target effects of 10 Sdd proteins using orthogonal R-loop assays in rice protoplasts. We found that 6 of the 10 deaminases (Sdd2, Sdd3, Sdd4, Sdd6, Sdd10, and Sdd59) had lower off-target activity than rAPOBEC1. Interestingly, while Sdd6 showed almost no off-target editing activity, it still exhibited strong targeted base editing capabilities when tested at six endogenous sites in rice and human cells. Figure 26 A and Figure 27 When we analyzed the on-target:off-target ratio of these 10 deaminases, Sdd6 showed the highest on-target:off-target editing ratio, which was 37.6 times that of APOBEC1. Figure 26 B). We further compared the on-target and off-target editing of Sdd6 with rAPOBEC1 and its two high-fidelity deaminase variants YE1 and YEE in HEK293T cells. Importantly, we found that Sdd6 had the highest on-target:off-target editing ratio, calculated to be 2.8-fold, 2.1-fold, and 2.5-fold higher than rAPOBEC1, YE1, and YEE, respectively. Figure 26 B, C and Figure 28 It is 10.4 times higher than hA3A. Figure 26 C and Figure 28 It is worth noting that the on-target activity of Sdd6 is comparable to that of rAPOBEC1, and significantly higher than that of YE1 and YEE. Figure 28 Therefore, we determined that the SCP1.201 branch contains a unique and more precise Sdd protein that can be used as a high-fidelity base editor.
[0320] Example 11: Rational Design of Sdd Protein Assisted by AlphaFold2 Structure Prediction
[0321] While viral delivery of CBEs holds great potential for disease treatment, the large size of APOBEC / AID-like deaminases limits their ability to be packaged into single adeno-associated virus (AAV) particles for in vivo editing applications (31). Dual AAV delivery strategies have been developed in other works, splitting CBEs into N-terminal and C-terminal fragments and packaging them into individual AAV particles. However, these delivery efforts will face challenges in terms of large-scale production capacity, higher viral doses, and potential safety for human use. Recently, a CDA-1-based truncated lamprey CBE has been used to develop single-AAV-encapsulated CBEs, but these vectors exhibited little HEK293T cell editing activity. Given the canonical density and conservation of SCP1.201 deaminases, we believe they may be ideal proteins for developing single-AAV-encapsulated CBEs. This invention attempts to further design and shorten the size of the newly discovered Sdd protein using AI-assisted three-dimensional structural protein modeling.
[0322] We first compared the AlphaFold2 predicted structures of all active Sdd deaminases and found that they possess a conserved core structure (Figures 13D, 13E, and 13D). Figure 29 We then generated several truncated variants of Sdd7, Sdd6, Sdd3, Sdd9, Sdd10, and Sdd4, and tested the endogenous base editing of these variants in rice protoplasts at two sites. We found that mini-Sdd7, mini-Sdd6, mini-Sdd3, mini-Sdd9, mini-Sdd10, and mini-Sdd4 are novel, minimized deaminases, all very small (~130–160 aa), and exhibit comparable or higher editing efficiency compared to full-length proteins in rice protoplasts and human cells. Figure 30A Notably, all mini deaminases support the construction of a single AAV-encapsulated SaCas9-based CBE (<4.7kb). Figure 30B We constructed a single-AAV SaCas9 vector using mini-Sdd6 and, through transient transfection, found that it achieved approximately 60% editing efficiency at two sites of the HPD gene (mouse 4-hydroxyphenylpyruvate dioxygenase) in mouse neuroblastoma N2a cells. Figure 30C These results demonstrate that the Sdd protein has a greater advantage than APOBEC / AID deaminase in AAV-based CRISPR base editing delivery. The further successful shortening of Sdd protein packaging for AAV further highlights the significant potential of this invention's method for predicting protein function based on three-dimensional structure.
[0323] Example 12: Base editing capability of CBE based on novel Sdd
[0324] Next, we explored the application of novel Sdd engineered proteins in plant base editing. We first evaluated the ability to use mini-Sdd7 in Agrobacterium-mediated rice genome editing, comparing it to the most commonly used human A3A-based (hA3A) CBE in agricultural applications. Mini-Sdd7-based CBE showed more positive rice plants, a larger number of edited plants, and higher editing efficiency, reflecting its higher efficiency and lower toxicity compared to hA3A-based CBE. Figure 31 ).
[0325] Soybean is one of the most important staple crops grown worldwide and is a fundamental source of vegetable oil and protein. Although base editing has been demonstrated in soybean, it remains difficult and inefficient at most test sites in the soybean crop. To understand whether our newly developed Sdd-based CBE would produce better cytosine base editing in soybean, we constructed vectors using the AtU6 promoter driving sgRNA expression and the CaMV2×35S promoter driving CBE expression, and evaluated them using Agrobacterium-mediated transformation of transgenic soybean hairy roots. Figure 32 We found that APOBEC / AID deaminases exhibited low editing activity at all five assessed sites, including the GmALS1-T2 and GmPPO2 sites, which are particularly difficult to edit by other CBEs in soybean. Figure 30D Notably, mini-Sdd7 exhibited cytosine base editing levels at five sites that were 26.3, 28.2, and 10.8 times higher than those of other deaminases rAPOBEC1, hA3A, and hAID, respectively, with an editing efficiency as high as 67.4%. Figure 30D Therefore, we can leverage these newly discovered Sdd proteins to overcome the limitations of soybean crops in efficient cytosine base editing.
[0326] Next, we attempted to use mini-Sdd7 for base editing to obtain Agrobacterium-mediated transgenic soybean plants. We chose to edit the endogenous GmPPO2 gene to produce the R98C mutation, which would result in soybean plants resistant to carfentraone ethyl. From three independent transformation experiments, we obtained 77 transgenic soybean seedlings, of which 21 were heterozygous for base editing (Fig. 30E, F). After 10 days of treatment with cytosine, we could clearly observe that wild-type plants were susceptible to wilting and unable to root, while mutant plants grew well and normally (Fig. 30G). Developing a highly efficient cytosine base editor for soybean plants could enable a variety of applications in the future.
[0327] Example 13: Base editing properties of proteins from other families
[0328] In addition to a detailed classification and validation of proteins in the SCP1.201 family, we also validated the deaminase functions and preferences of other families in the Iyer deaminase classification family (Table 1). Based on our analytical and validation methods, we also discovered a series of deaminases with similar Sdd activities in other families such as MafB19, AID / APOBEC, Novel AID / APOBEC-like, TM1506, Toxin deam, and XOO2897 (see Table 5 for specific deaminases). Taking the MafB19 family as an example, in Example 2, we found that some proteins in the MafB19 branch (No. 101m) have single-chain deaminase functions. Furthermore, in Example 3, based on AI-assisted clustering of protein three-dimensional structures, we discovered two branches with distinctly different structures within the MafB19 deaminase family (…). Figure 11B and Figure 12 Using the deaminase screening and identification method of this invention, we discovered that three proteins (No. 2-1241, No. 2-1231, and No. 99) in the MafB19 family also possess Sdd catalytic activity, and obtained their motif preferences for base editing sequences (Table 5). The various novel cytosine base deaminases with different editing properties screened by the method of this invention enrich base editing tools, expand base editing systems, and enhance the ability to precisely manipulate target DNA sequences.
[0329]
[0330] Experimental conclusions: Traditional CBE based on deaminases suffers from drawbacks such as low editing efficiency, small editing window, and significant bias. A series of novel cytosine deaminases were obtained using the protein function prediction method based on three-dimensional structure of this invention.
[0331] These cytosine deaminases have demonstrated promising application potential and a wide range of applications. For example, in the hairy roots of transgenic soybeans transformed using Agrobacterium-mediated transformation. We found that the APOBEC / AID deaminase exhibited low editing activity at all five evaluation sites, including the GmALS1-T2 and GmPPO2 sites, which are particularly difficult to edit by other CBEs in soybean. Notably, compared to rAPOBEC1, hA3A, and hAID, mini-Sdd7 showed 26.3-fold, 28.2-fold, and 10.8-fold cytosine base editing levels at the five sites, respectively, with an editing efficiency as high as 67.4%. Therefore, we emphasize the use of these newly discovered Sdd proteins to overcome the limitations of efficient cytosine base editing in soybean crops. Next, we attempted to use mini-Sdd7 for base editing to obtain Agrobacterium-mediated transgenic soybean plants. We chose to edit the endogenous GmPPO2 gene to generate the R98C mutation, which will produce soybean plants resistant to cyclohexane. We obtained two heterozygous base-editing plants from 30 transgenic soybean seedlings. After treatment with cytosine for 10 days, we observed that wild-type plants were significantly susceptible to wilting and unable to root, while mutant plants grew well and normally. Developing efficient cytosine base editors for soybean plants could enable a variety of applications in the future. We believe that future sequencing efforts in parallel with structure prediction will greatly advance the discovery, tracking, classification, and design of functional proteins. Currently, only a few cytosine deaminases are used as cytosine base editors. Canonical efforts based solely on protein engineering and directed evolution can help diversify editing properties; however, these efforts are often difficult to establish. Using our three-dimensional structure-based clustering prediction method, we discovered and analyzed a set of deaminases with different properties. For example, among the newly discovered deaminases, we found that Sdd7 and Sdd6 show great promise for therapeutic and agricultural applications. Sdd7 exhibited strong base-editing capabilities in all tested species and showed higher editing activity than the most commonly used APOBEC / AID-like deaminases. Surprisingly, we found that Sdd7 can be efficiently edited in soybean plants, where editing cytosine bases has previously been challenging (plant genes typically have high GC sequence content). We hypothesize that Sdd7, derived from the bacterium *Actinosynnema mirum*, may exhibit higher activity at temperatures suitable for soybean growth compared to mammalian APOBEC / AID deaminases. In analyzing Sdd6, we found that this deaminase, while maintaining high targeted editing activity, is more specific by default than other deaminases. Interestingly, we found that AlphaFold2-based modeling further enables our protein engineering work to minimize protein size, which is crucial for using these editing techniques for viral delivery in in vivo therapeutic applications.
[0332] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and additions without departing from the method of the present invention, and these improvements and additions should also be considered within the scope of protection of the present invention.
[0333] Related sequence and brief description > SEQ ID NO:1 No.182 (AID / APOBEC clade)
[0334] MICKLDSVLMTQKKFIFHYKNVRWARGRHETYLCFVVKRRIGPDSLSFDFGHLRNRSGCHVELLF
[0335] LRHLGALCPGLSASSVDGARLCYSVTWFCSWSPCSKCAQQLAHFLSQTPNLRLRIFVSRLYFCDEE
[0336] DSVEREGLRHLKRAGVQISVMTYKDFFYCWQTFVARRERSFKAWDGLHENSVRLVRKLNRILQP
[0337] CETEDLRDVFALLGL
[0338] >SEQ ID NO:2 Sdd9 / SCP044 / No.69(SCP1.201 clade)
[0339] MLDVDDPHTFYVLAGKTPVLVHNSECPIWVKNALQELVGRKETSGKVFDVDGKPIGPDIQSGYKD
[0340] RELRRGVYETLRKSPFFQKHFPSTATWYVSLHVEAQYAVWMLRNRIKHATVVINNTYVCSDMNR
[0341] LHDNCMTAVPHILPEGYTMTVWKADRTEVTLRGKAPKE
[0342] >SEQ ID NO:3 Sdd5 / SCP017 / No.55(SCP1.201 clade)
[0343] MGPLLDGIAARLEAVRAALLGEGGGAGDDEPPAVPWDRVERLRRELPPPVVPNTGQKTHGRWIG
[0344] PDGQARPIVSGRDDKSVLVNPLLRGKGAPGPTRRDSDVEMKLAAHMAARGIRHATVVINNTPCR
[0345] GPLRCDTLVPILLPEGSTLTVHGINENGTRTRIRYTGGARPWWS
[0346] >SEQ ID NO:4 Sdd7 / SCP016 / No.57(SCP1.201 clade)
[0347] MLEAVRARLIGEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVR
[0348] AIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCD
[0349] TLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWS
[0350] >SEQ ID NO:5 Sdd4 / SCP014 / No.64(SCP1.201 clade)
[0351] MLDAMDAYLSEIAGGNAPARAGPKAPEPKQPGGSSSPRARDGRIDFRALLERLKAQGVVGLEGRS
[0352] DDPIPDFDPKKQNPACYQGLAPRQKGKPVRGNLFFPDGRRWNDVALESSRGEPAFDLNIIKPEYRS
[0353] LSPARGHLEGNVAAWMRSTFHQEMVLYINESPCRKHGKGCLYTLEHFLPRGYVLHVWSRNDRGE
[0354] WRGNTFRGSGEAFTEGA
[0355] >SEQ ID NO:6 Sdd76 / SCP012 / No.76(SCP1.201 clade)
[0356] MYVLAGNTPVLVHNTGPGCGEPGFVSDAANSLSGRRITTGQIFDASGNPIGPEITSGGGSLADRAQ
[0357] SYLADSPNIRNLPAKARYASADHVEAQYAVWMRENGVTDASVVINQNYVCGLPLGCQAAVPAIL
[0358] PRGSTMTVWYPGSGSPIVLRGVG
[0359] >SEQ ID NO:7 Sdd6 / SCP273 / No.2-1146(SCP1.201 clade)
[0360] MVETRDKIIAAKSRSDAGLLAFQQATNGSIDSRPAEAIANLQRAKTHLDEAQRLVANSDAAVDNY
[0361] INAILGGASAATAQPSAVIPASKPSRFKPMRTDPAKADEIRPHVGKDRAVATLWDADGNRVLGLH
[0362] SADDDGPAATAAWKPPWRDYVRLRRHVEAHAAARMHQDGHKTMVMYINLPPCKYFDGCKLNL
[0363] EDILPKGSTLWMHRVFQNGGTKIYQFNGTGRAYV
[0364] >SEQ ID NO:8 SCP021 / No.2-1160(SCP1.201 clade)
[0365] MPIRLSGGLNLYQYAPETNNWIDPLGCSGHRRRHEKMPPEGARLTTGNLFRHAQDEGISPILSSKN
[0366] DEFYYRLIKIYAGTGMLNANIRGIASHVEPKAGLILNEDGSGWKIGSLYINYPNGPCLDCRTLMPFI
[0367] LNDGSILYVTFPTLGLDGYSYGHFHGREPGFFREGTPCNLPHPE
[0368] >SEQ ID NO:9 SCP038 / No.54(SCP1.201 clade)
[0369] MADTPQPDSNPLPARDGQRLDREQAEAIRAQLPPTVKPGTGQKTHGCWVDEQGQPQSVTSGQDN
[0370] SAAAVWARLQALGIPLSGPPTATADVEQKVAIQMIQQGRQHVDVVINNEPCRGRFSCDTLVPIILP
[0371] EGSSLTVHGTNGFRKTYTGGAKPPWSR
[0372] >SEQ ID NO:10 SCP051 / No.56(SCP1.201 clade)
[0373] MRAGCGVPDGRTRSAISGKDEGFALALRTIRELGMTRGFPLRAADVEMKVASEMRANGITSATLV
[0374] INHVPCDDGMFSCDRMVPVLLPAGSTLTVFGAGGFRMTYHGGEQLPCPTP
[0375] >SEQ ID NO:11 Sdd59 / SCP183 / No.59(SCP1.201 clade)
[0376] MLLTPPPRPAAPPTTRPKPLVARTGDAYPPGTEWALPLIVQPHPPVGGTVPVEGHVRALRPESQISH
[0377] VFHPGGGHWTEQARARLRVLPGFGWAVNLGHHVELQIAAWMTACGIHHAELVLNRPPCGERYG
[0378] LGCHQALPVLLPRGYRLTVSSTRGGPQPYQHHYEGKA
[0379] >SEQ ID NO:12 Sdd10 / SCP018 / No.60(SCP1.201 clade)
[0380] MLDAALGAVRRIIAALGTSGAERASPGANGSERVDELAERLPPTVVPNTSAKTHGWWFTGQGAA
[0381] QELISGEGPDARAAYEALREEGYPRPGMPFVAMHVEIKLAAHMRRNDIEHATVVINNIPCPLVWG
[0382] CENLIGVVLPEGSSLTVHGSNGYERTFTGGRKPPWPR
[0383] >SEQ ID NO:13 SCP011 / No.61(SCP1.201 clade)
[0384] MLGGVLPARSVMFPGHVEPDAHFGPNPERHHPALVEVPIVWAGRQEDRTSTWARRVQRGFPRYT
[0385] VGAKTAGMFYNAGSQSWELLSGVDHRGGLTRKASQHISRMLSSGFFDGKPLDTKSDHLRMLNYT
[0386] STHVETKAAIWARDSDQETIDVVTNRNYVCGESYDPDDVDEPPGCYQAVESVLREGQTMRVWTT
[0387] DPENRVITIHGKGM
[0388] >SEQ ID NO:14 SCP157 / No.72(SCP1.201 clade)
[0389] MAAGGGVSRPPATRDGAANPTRVPNPPEWLPGWLTEAARDLPRRQAKDPTSGVALINGERIPMRS
[0390] GRDPAAAADLKAAYKLIATTTDHLEAKLAARMRRDQVMHAEVLTNNPPCDYEPYGCEKILSRLL
[0391] PAGAQLSVYVRDDDDQVRLWRTYIGNGKAIA
[0392] >SEQ ID NO:15 SCP008 / No.74(SCP1.201 clade)
[0393] MTTGGGSDISRPPATRDTSATATEAPAQPELVPEGLTDAARDLPRRQAKDPTSGVALIGGERIPMR
[0394] SGRDPDAAADLKPAYKLIATTTDHLEAKLAARMRRDHITQAAVVTNNPPCDYTPYGCEKILSRLL
[0395] PAGARLAVYVRDDDGQVRHWRTYTGNGKAIA
[0396] >SEQ ID NO:16 SCP013 / No.75(SCP1.201 clade)
[0397] MIRARDRLTAVTASSRHPLVDQALQHVTAAIERLQVADRDAALAASALVAYGRTLGISLPVPPPVS
[0398] APTRGAAPVPSWIRQTGQDLPTRPDDHGPTHGQAFDSTGRPLSAEPWRSGRNIASTSDLRPIPGLK
[0399] GFPWTLTDHVESRAAQQMRRPGAPREVSLVVNKEPCTDDPYGCDRILRHIIPAGSRLTIYVRDPDA
[0400] PAGVRTVGQYEG
[0401] >SEQ ID NO:17 Sdd3 / SCP170 / No.63(SCP1.201 clade)
[0402] MSASAQLNTYLAAIGNSTTTVEAQPEAAPPPAAAESLDSTPRLPDGGIDFHALAKRLGLLEARPTE
[0403] QPPFDPRRFNPACWQGLKPYDQAGTAEGNLFIAPGKRWNTRPMQASKLEVGPQSDLHPQWRSRK
[0404] APWHIEGKIAAYMRQKGFTDGCVYLNARPCSGPDGCARNLPDLLPVGSTLHVHARYIDRTGETRF
[0405] YYREYRGTGKALT
[0406] >SEQ ID NO:18 No.2-1158(SCP1.201 clade)
[0407] MPIRLSGGLNLYQYAPETNNWIDPLGCSGYRRRHEKMPPEGARLTTGNLFRHAQDEGIPPIFSSEN
[0408] DEFYHRLIEIYAGTGVLNAYIRGIASHVEPKAGLILNEDGSGWKIGSLYINYPDGPCLGCRTLMPFIL
[0409] NDGSILYVTFPTLGLDGYSYGHFHGGVSGFFREGTPCNLRHPE
[0410] >SEQ ID NO:19 No.101m(MafB19 clade)
[0411] MCGWSELASYRAREGMPARGSADDTFTAARLQIDGQVFFGRNAHGRPVDIRVNAQTKTHAEAD
[0412] VFQQAKDAGATGTRAVLHVVRDFCRSCGATGGVGSLMRGLGVEELLVHSPSGIFTINAVRRPSTP
[0413] RPLG
[0414] >SEQ ID NO:20 No.2-1479(Novel AID / APOBEC-like clade)
[0415] MRTRPAFGRCDCAGSDRGWEVAGGYTSEASHVRRSPTPDGNSLLGGVAQLCHAFFHCSPTPELSS
[0416] HPDELCLRVACDPAGRKCETKGVVVVAALRDRAGDLRFLSRYSNCPLSSHAEEYVVRDEELVRA
[0417] VMEMAPEDDARSSTKTPGSAGTLTLYQRLQPCHGSSDNRGPLWSCSDALVAGLHRELLGPRGVSL
[0418] RVAVSYTYRAHWDVRGFESERERRWWGPKVEAAREGIRVFAAAAKDGVTLEALNAEDWAFLVS
[0419] LCDEDVARDYAAAFGEGGEGFGSKGCGVFTAPAVAHRRAMDEFVAEQIRRHSASGPTNGRATNA
[0420] AG
[0421] >SEQ ID NO:21 No.2-1478(Novel AID / APOBEC-like clade)
[0422] MADLDDVQDPLLDTALDSTKDEADDSVLSEIAVNDTSVDDGVEDPDHEHKKIAKAGDKVLGNKK
[0423] EFCGAFYHVPRSKSGCLDKQSCAIAKRGHDATPLTAVALVKYEQQESSEWAIKSVRRYTNCSDKM
[0424] KHAEEFFLMDIDCQLEARHKGEEGFLDFWNKKKWQITMYLTMQPCHLSTDTGGTKEDQSCCEV
[0425] MIKAKEKLGDNVEIVIKPTHLCQVGWYKGKPREKPKNAEKGVRKLFKTTGIELECMKEGDWKYL
[0426] LQYAQPEVENKLPDYDTSRRKTEDEKIGEELHNQQLEQLAPELAQLSVNEKRRK
[0427] >SEQ ID NO:22 No.2-39(TM1506 clade)
[0428] MVYSMDQKNKVTEKLKEGGYSFVLYKDGEWSTSEKRGIAPIMELLKENKELLRGAYVADKVIGK
[0429] AAALLLIEGGISYLHAEIISEHAIEVLQNSNIEYEYQELVPYIVNRSGDGMCPMEETVLDVTDTKIAF
[0430] ELLQEKIKKMQAAMQAQNMK
[0431] >SEQ ID NO:23 No.2-1429(XOO2897-like clade)
[0432] MISDAAVAGIASKMAEKYYSACKKLSRSIPISTLGVIGKPVPEYSCDGIVPYNSTDLGRMAYKARV
[0433] EAGFGIFGGRNVAVARVPGWDDPKTGDLVVGFSQGNGFHAEDHVLEQLTKKDISPKKITELYSER
[0434] QPCAACGPNLENHLSPGTEITWSVQWGSDLEMNSAFTELLGKLIQQQ
[0435] >SEQ ID NO:24 No.2-1442(XOO2897-like clade)
[0436] MAPDSLVWFDPLGLIVLQQVPYNDHPLFGAVSEFIQGKSRSDLRGRNVAAVLLDDGTVIVRASEG
[0437] GGNHAERVLMGLSEVDPAKVAVYTERSPCTGRINCHDLLDSSLGADVPVYYTHEMIRGQEGKT
[0438] AQQIEADRNQFCRGG
[0439] >SEQ ID NO:25 SpCas9
[0440] MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKR
[0441] TARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKY
[0442] PTIYHLRKKLVDSTDKADLRIYLALAMHIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEE
[0443] NPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQ
[0444] LSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQD
[0445] LTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRED
[0446] LLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAW
[0447] MTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKY
[0448] VTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYH
[0449] DLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAHLFDDKVMKQLKRRRYTGWGR
[0450] LSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLA
[0451] GSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQ
[0452] ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTR
[0453] SDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVE
[0454] TRQITKHVAQILDSRMNTKYDENDKLIVEKVITLSKLVSDFRKDFQFYKVREINNYHHAHDAY
[0455] LNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLAN
[0456] GEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA
[0457] RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKG
[0458] YKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPED
[0459] NEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLG
[0460] APAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD
[0461] >SEQ ID NO:26 nCas9(D10A)
[0462] MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKR
[0463] TARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKY
[0464] PTIYHLRKKLVDSTDKADLRIYLALAMHIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEE
[0465] NPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQ
[0466] LSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQD
[0467] LTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRED
[0468] LLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAW
[0469] MTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKY
[0470] VTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYH
[0471] DLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAHLFDDKVMKQLKRRRYTGWGR
[0472] LSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLA
[0473] GSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQ
[0474] ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTR
[0475] SDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVE
[0476] TRQITKHVAQILDSRMNTKYDENDKLIVEKVITLSKLVSDFRKDFQFYKVREINNYHHAHDAY
[0477] LNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLAN
[0478] GEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA
[0479] RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKG
[0480] YKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPED
[0481] NEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLG
[0482] APAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD
[0483] >SEQ ID NO:27 UGI
[0484] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKP
[0485] WALVIQDSNGENKIKML
[0486] >SEQ ID NO:28 Ddd1 / SCP177
[0487] MSLPEYDGTTTHGVLVLDDGTTQIGFTSGNGDPRYTNYRNNGHVEQKSALYMRENNISNATVYHN
[0488] NTNGTCGYCNTMTATFLPEGATLTVVPPENAVANNSRAIDYVKTYTGTSNDPKISPRYKGN
[0489] >SEQ ID No:29 Ddd2
[0490] MEDFHTYHVGKCRLLVHNANCNQEKPVLPKYDGKTTEGVMVTPDGKQISFKSGNSSTPSYPQYK
[0491] AQSASHVEGKAALYMRENGINEATVFHNNPNGTCGFCDRQVPALLPKGAKLTVVPPSNSVANNV
[0492] RAIPVPKTYIGNSTVPKIK
[0493] >SEQ ID No:30 Ddd3 / SCP001
[0494] MLSSSYNAFALTYGVLILDDGKQYSFNSGKPDPIYRNYIPASHVEGKAAIYMRENKIQSGTVYHNN
[0495] TDGTCPYCDKMLPTLLEKDSTLKVVPPQNATSSKKGWITNEKIYIGNDKIPKTAR
[0496] >SEQ ID No:31 Ddd4
[0497] MGGDEEEENLTSNNEKKKNANKQKIELPPYDGKTTYGVLILDDGKQYSFNSGKPAPIYRNYIPASH
[0498] VEGKAAIYMRENKIQSGTVYHNNTDGTCPYCDKMLPTLLEKDSTLKVVPPQNATSSKKGWITNEK
[0499] IYIGNDKIPKTAR
[0500] >SEQ ID No:32 Ddd6 / SCP009
[0501] MALLREAYPSMEGATLPPFDGKTTIGLMFYTDASGQYQVKKLFSGEKVLSNYDATGHVEGKAALI
[0502] MRNEKITEAVVMHNHPSGTCNYCDKQVETLLPKNATLRVIPPENAKAPTSYWNDQPTTYRGDGK
[0503] DPKAPSKK
[0504] >SEQ ID No:33 Ddd7 / SCP103
[0505] MIGLMGGLNLYQYAPNSIAWTDWWGLAGSYTLGSYQISAPQLPAYNGQTVGTFYYVNGAGGLE
[0506] SRTFSSGGPTPYPNYANAGHVEGQSALFMRDNGISDGLVFHNNPEGTCGFCVNMTETLLPENSKL
[0507] TVVPPEGAIPVKRGATGETRTFTGNSKSPKSPVKGEC
[0508] >SEQ ID No:34 Ddd8 / SCP234
[0509] MQDNTNIIDNRPKLPDYDGKTTHGILVTPNSEHIPFSSGNPNPNYKNYIPASHVEGKSAIYMRENGI
[0510] TSGTIYYNNTDGTCPYCDKMLSTLLEEGSVLEVIPPINAKAPKPSWVDKPKTYIGNNKVPKPNK>SEQID No:35 Ddd9 / SCP003
[0511] MGKSLSESQATLSVAQRLLATIGEEGKTAGVLELDGELIPLVSGKSSLPNYAASGHVEGQAALIMR
[0512] DRGATSGRLLIDNPSGICGYCKSQVATLLPENATLQVGTPLGTVTPSSRWSASRTFTGNDRDPKPW
[0513] PR
[0514] >SEQ ID No:36 Ddd10 / SCP004
[0515] MASPAVGTNAAGSSGKNVRMPRDYASELPEYDGKTTHGVLVTNEGKVIQLRSGGKEEPYTGYKA
[0516] VSASHVEGKAAIWIRENGSSGGTVYHNNTTGTCGYCNSQVKALLPEGVELKIVPPTNAVAKNAQA
[0517] RAVPTINVGNGTQPGRKQK
[0518] >SEQ ID No:37 Ddd11 / SCP271
[0519] MQGTSSDTIAEMLNSASQPGRTAGVLDIDGELTPLTSGRPSLPNYIASGHVEGQAAMIMRQQQVQS
[0520] ATVYHDNPNGTCGYCYSQLPTLLPEGAALDVVPPAGTVPPSNRWHNGGPSFIGNSSEPKPWPR>SEQID No:38 Ddd12 / SCP005
[0521] MGVAGGAATNADAQALLGSIRQAGKTAGVLNIDGDLMPLVSRKSSLPNYAASGHVEGQAALIMR
[0522] ERGVSSAELLIDNPNGICSYCTSQVPTLLPEGAQLMVRPPLGTVPTQWWFNGRTFLGNAANPKPSP
[0523] W
[0524] >SEQ ID No:39 Ddd13 / SCP006
[0525] MHNINGCGPSAVQQLLSANGEPGKTAGVLDLNGELTSLVSGKGELPNYAASGHVEGQAAMMMR
[0526] AEKATSATLYIDNPNGICGYCRSQIATLLPEGATLEVVTPLGTVEPTARWSSSKVFTGNERYPKGW
[0527] VE
[0528] >SEQ ID No:40 Ddd14 / SCP007
[0529] MGSAVVGGGIAATGAKALTTGKKLTESPGTLNAAQRLLASIGEEGKTAGVLEVDGALFPLVSGKS
[0530] VLPNYAASGHVEGQAALLMQGMGATNGRLLIDNPNGICGYCTSQVPTLLPENAVLEVGTPLGTVT
[0531] PSARWSASKPFIGNDREPKPWPR
[0532] >SEQ ID No:41 No.2-1157(SCP1.201 clade)
[0533] MDPIRLSGGLNLYQYAPETNNWIDPLGCSGHRRRHEKMPPEGAPLTTGNLFRHAQDEGIPPIFSRK
[0534] DDEFYHRLIEIYAGTGVLNAYIRGIASHVEPKAGLILNEDGNGWKIGSLYINYPDGPCPGCRRLMPF
[0535] ILNDGSILYVTFPTLGLDGYSYGHFHGGVSGFFREGTPCNLRHPE
[0536] >SEQ ID No:42 No.73 / SCP158(SCP1.201 clade)
[0537] MPGGGEINRPPTTQADDVPQPVGSEWERTEPDALPGTVRAAVERLQPRPAGSTRPTLGVFNGEEIT
[0538] SGGGDRSLAADLDHDPLRGPPVTFYDHVESKAAARMRRTGSTESDLAIDNTVCGTNDRDQSYPW
[0539] TCDKILPAILPNGSRLRVWVTRDGGVTWWHRVYIGTGERITK
[0540] >SEQ ID No:43 No.2-1145 / SCP315(SCP1.201 clade)
[0541] MSPKKPTASSDLKAIGERLGLKPCEGLLGTLPAMKPNSGQRTRGRWHKHPDRELTSGANDRDWE
[0542] HVKDFWHHNIWSGTAEDTEPRWLAHLELKFAMTMRRTRTKSEPVDQVHEEITINHPDGPCPQCQL
[0543] LLPYFLEEGSSLTIHWPAGSATYIGRPYFDRPLRDVKPYINEEQQ
[0544] >SEQ ID No:44 No.2-1157 / SCP020(SCP1.201 clade)
[0545] MDPIRLSGGLNLYQYAPETNNWIDPLGCSGHRRRHEKMPPEGAPLTTGNLFRHAQDEGIPPIFSRK
[0546] DDEFYHRLIEIYAGTGVLNAYIRGIASHVEPKAGLILNEDGNGWKIGSLYINYPDGPCPGCRRLMPF
[0547] ILNDGSILYVTFPTLGLDGYSYGHFHGGVSGFFREGTPCNLRHPE
[0548] >SEQ ID No:45 No.2-1156(SCP1.201 clade)
[0549] MDPIRLSGGLNLYQYAPETNNWIDPLGCSGYRRRHEKMPPEGARLTTGNLFRHAQDEGIPPIFSSE
[0550] NDEFYYRLIKIYAGTGILNANIRGIASHVEPKAGLILNEDGSGWKIGSLYINYPNGPCLDCRRLMPFI
[0551] LNEGSILYVTFPTLGLDGYSYGHFHGREPGFFREGTPCNLRHPE
[0552] >SEQ ID No:46 Sdd2(SCP1.201 clade)
[0553] MAPDSLVWFDPLGLIVLQQVPYNDHPLFGAVSEFIQGKSRSDLRGRNVAAVLLDDGTVIVRASEG
[0554] GGNHAERVLMGLSEVDPAKVVAVYTERSPCTGRINCHDLLDSSLGADVPVYYTHEMIRGQEGKT
[0555] AQQIEADRNQFCRGG
[0556] >SEQ ID No:47 SCP008(SCP1.201 clade)
[0557] MTTGGGSDISRPPATRDTSATATEAPAQPELVPEGLTDAARDLPRRQAKDPTSGVALIGGERIPMR
[0558] SGRDPDAAADLKPAYKLIATTTDHLEAKLAARMRRDHITQAAVVTNNPPCDYTPYGCEKILSRLL
[0559] PAGARLAVYVRDDDGQVRHWRTYTGNGKAIA
[0560] >SEQ ID No:48 SCP011(SCP1.201 clade)
[0561] MLGGVLPARSVMFPGHVEPDAHFGPNPERHHPALVEVPIVWAGRQEDRTSTWARRVQRGFPRYT
[0562] VGAKTAGMFYNAGSQSWELLSGVDHRGGLTRKASQHISRMLSSGFFDGKPLDTKSDHLRMLNYT
[0563] STHVETKAAIWARDSDQETIDVVTNRNYVCGESYDPDDVDEPPGCYQAVESVLREGQTMRVWTT
[0564] DPENRVITIHGKGM
[0565] >SEQ ID No:49 SCP013(SCP1.201 clade)
[0566] MIRARDRLTAVTASSRHPLVDQALQHVTAAIERLQVADRDAALAASALVAYGRTLGISLPVPPPVS
[0567] APTRGAAPVPSWIRQTGQDLPTRPDDHGPTHGQAFDSTGRPLSAEPWRSGRNIASTSDLRPIPGLK
[0568] GFPWTLTDHVESRAAQQMRRPGAPREVSLVVNKEPCTDDPYGCDRILRHIIPAGSRLTIYVRDPDA
[0569] PAGVRTVGQYEG
[0570] >SEQ ID No:50 miniSdd7
[0571] MEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDA
[0572] ALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGC
[0573] TLTVHGQTDKGMRVRVRYTGGARPWWS
[0574] >SEQ ID No:51 miniSdd4
[0575] MDPKKQNPACYQGLAPRQKGKPVRGNLFFPDGRRWNDVALESSRGEPAFDLNIIKPEYRSLSPAR
[0576] GHLEGNVAAWMRSTFHQEMVLYINESPCRKHGKGCLYTLEHFLPRGYVLHVWSRNDRGEWRGN
[0577] TFRGSGEAFTEGA
[0578] >SEQ ID No:52 miniSdd9
[0579] MCPIWVKNALQELVGRKETSGKVFDVDGKPIGPDIQSGYKDRELRRGVYETLRKSPFFQKHFPSTA
[0580] TWYVSLHVEAQYAVWMLRNRIKHATVVINNTYVCSDMNRLHDNCMTAVPHILPEGYTMTVWK
[0581] ADRTEVTLRGKAPKE
[0582] >SEQ ID No:53 miniSdd6
[0583] MPASKPSRFKPMRTDPAKADEIRPHVGKDRAVATLWDADGNRVLGLHSADDDGPAATAAWKPP
[0584] WRDYVRLRRHVEAHAAARMHQDGHKTMVMYINLPPCKYFDGCKLNLEDILPKGSTLWMHRVF
[0585] QNGGTKIYQFNGTGRAYV
[0586] >SEQ ID No:54 miniSdd10
[0587] MPGANGSERVDELAERLPPTVVPNTSAKTHGWWFTGQGAAQELISGEGPDARAAYEALREEGYP
[0588] RPGMPFVAMHVEIKLAAHMRRNDIEHATVVINNIPCPLVWGCENLIGVVLPEGSSLTVHGSNGYE
[0589] RTFTGGRKPPWPR
[0590] >SEQ ID No:55 miniSdd3
[0591] MPRRFNPACWQGLKPYDQAGTAEGNLFIAPGKRWNTRPMQASKLEVGPQSDLHPQWRSRKAPW
[0592] HIEGKIAAYMRQKGFTDGCVYLNARPCSGPDGCARNLPDLLPVGSTLHVHARYIDRTGETRFYYR
[0593] EYRGTGKALT
[0594] >SEQ ID NO:56 No.2-1241(MafB19 clade)
[0595] MGLEGTPCDGFGALAARRKSLGLPAAGSEGDTSTLSLLRINGQSFEGINSSDQNPKTPITLDRVNAQ
[0596] TKTHAEAEAVQKAVNAGMAGKASHAEMWVDRDPCRACGIPGAGGLRSLARNLGCPITVHSPSGT
[0597] QVYTPTK
[0598] >SEQ ID NO:57 No.2-1231(MafB19 clade)
[0599] MLGPPLDLNPANRAPEFGRCDGTSWIDSYRTINNATDLFGRPVWPNHRGTVAVARIDGDIYFGVN
[0600] SKAPGYSDADWNLAAGLRDQMALEHPELIRGESRGSRPLDAVFHAEANLLIRASRYVGSLVKRSI
[0601] EVQVDRPVCWSCEQALPKVGLELGDPYVTIREVRSGRASVMWQGEWLVWRKK
[0602] >SEQ ID NO:58 No.99(MafB19 clade)
[0603] MGWVDPLGLVSGGAWDAISFFRDQNSLLSVVDEDLLAASGAKNAQNTVALLRVGDREFIGVNSR
[0604] IQNPKNPFTAGPINNITKFHAEGNAAQQAIDAGMVGKHRIAEMWVDRDLCHACGPSNGVGSLTRA
[0605] LGLDAIIVHTPAGTRKFNAPCAG
[0606] >SEQ ID NO:59 No.2-1430(XOO2897-like clade)
[0607] MVKGLDGFVKTCTRTRGKLTARASTAVGGCPVGLVAYNSEEMSHWAYRYRTESEYFEGDHNVA
[0608] VAKVPGWNDPRTGDFIIANSKFSGHSETEILGKLEAKGFTPGQITALYTERQPCPACASVLTGSLKE
[0609] GTPVTWSVPYHPDYAKESRSLLDSYVRQANGQQRARPTTTQRLTEGNEAHD
[0610] >SEQ ID NO:60 No.2-1440(XOO2897-like clade)
[0611] MRLGQSVDPRLLEMAKEARVTQAGISREAFASYNVATARVRVGTEIRYLDAGNSPGRLMHSEDW
[0612] LITQVEELRRVHGRESVALEQLFSERIPCGECLPKLERLFNAEVFYAVAKRGTRATDLMKAYGLR>SEQID NO:61 No.2-1432(XOO2897-like clade)
[0613] MPDPLGLAPAANDRAYVPNPLTWADPYGLACTGTTEPGSTDLSQAVIQERLRLGKKGNNFAAAR
[0614] YIDDNGVEQIAVAASSKGQFMHAERKLVRQYGDKITEVYSEFEPCIGTNQCRKTLGDMGIKYTYS
[0615] WAWTLSKDGVAANAARKAYVDQIFDDAEAGNWAAPWAD
[0616] >SEQ ID NO:62 No.2-1437(XOO2897-like clade)
[0617] MNIDGDSYLVPGALAAMFLHKPGRGGKGKVGYGTTDLGQSVRLQRLIDKNRGMTNYAAARLDD
[0618] GDVIVGKSKKHVHAEEHLFQQAGKRKIVELYSEREPCSNKCEDLVKDIPFVSWSFKWNHPDRIKQ
[0619] DAIRDKANADLKDAVRSLFNSP
[0620] >SEQ ID NO:63 No.181(AID / APOBEC clade)
[0621] MQGYIVDESGRVLDANGLPIASLPPADDLSKWANYTVESGLHDDLAALENRLDLLYRQQFGLPM
[0622] APPAWHLETQLAYRVATREVALRDSTLRLVMNNPGGVCDAVPLTGTGPDRQRQAVAGCIQAVK
[0623] MLLPAGTTMIIYYPDPDDPAELLEITVRGVGRWLD
[0624] >SEQ ID NO:64 rAPOBEC
[0625] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFI
[0626] EKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLI
[0627] SSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQP
[0628] QLTFFTIALQSCHYQRLPPHILWATGLK
[0629] >SEQ ID NO:65 DddA
[0630] GSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALF
[0631] MRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNS
[0632] PKSPTKGGC
[0633] >SEQ ID NO:66 SCP090(SCP1.201 clade)
[0634] MDGPHGTPVLDRIAKLREELPPPAVPGKGQKTDGRWFDGNGAVRDSVSGKDVDSEEAWRLLRES
[0635] GIPLPRPPVVAHAEMKVAAAMRRLNVRHAVLVITNVPCDERWSCENLLPAVLPVGCSLSVHGPG
[0636] YQRTFHGRTPKW
[0637] >SEQ ID NO:67 SCP015(SCP1.201 clade)
[0638] MRPTTPPGPHARWRPDPPSAPHVAAIRRVGWPKKPQSDDDVRARGQLYHRDGTPWNASMLIASR
[0639] RGPASQRTDLKEPWASDPGYTTGWHIEGNTAALMVKHQQRDAVLYINQAVCGAEEPQDPKRCHS
[0640] NIVAMLSVGYALYVHSVQESGWLRRRVYKGTGEAIR
[0641] >SEQ ID NO:68 SCP278(SCP1.201 clade)
[0642] MQRHGGRVRLLSGENDDPHSWQQQAARFLRETFPDKGPGLAVLSRHVEIQLAVRLRHRPTNEVV
[0643] HEVLVIDRVVCGRDPRTQGREYTCDTVLPFVLDEGATLTVVEHDGARVTYRGRGRR
[0644] >SEQ ID NO:69 SCP287(SCP1.201 clade)
[0645] MKILYIKSAEGYPSLMLKNNPRIPTNARSYTHVEGRAASIMRQSGIKSAKLTINNTNGVCDPCRGN
[0646] MEKSLLPDGGKLNVRYPDGQGGYTQRILI
[0647] >SEQ ID NO:70 SCP341(SCP1.201 clade)
[0648] MREFGLPIEPPAWHLETQLAYRVSKREVALRENTLRLVMNNPGGVCDAVPAKDRGPDGQRQVVA
[0649] GCVQAIQMLLPAGTTMIIYYPDPANPAKLLQVTVRGVGRWLD
[0650] >SEQ ID NO:71 SCP353(SCP1.201 clade)
[0651] MLQGWRLAVGDGSSDRELASGTKLASGQTDPSYTAAVQRARELGLARGGFVPDIARHIEIKEAST
[0652] MTAGETRTIVIGKDPCGIDPVTNVSCHPFLRYFLPPGATLIVYGPRGEPYRYEGKRTS
[0653] >SEQ ID NO:72 SCP357(SCP1.201 clade)
[0654] MVGATTDSVPTQPGGRARSAPPRPPRPGKTHGRWCDSDGNAVVLESGKGGEYYEATRARGVAL
[0655] GLAKGIPNAEPSIARHVETQFVSRMIDQGIEYAEIEINRPVCGTTPKDQQ
[0656] >SEQ ID NO:73 SCP372(SCP1.201 clade)
[0657] MQGYIVDESGRVLDANGLPIASLPPADDLSKWANYTVESGLHDDLAALENRLDLLYRQQFGLPM
[0658] APPAWHLETQLAYRVATREVALRDSTLRLVMNNPGGVCDAVPLTGTGPDRQRQAVAGCIQAVK
[0659] MLLPAGTTMIIYYPDPDDPAELLEITVRGVGRWLD
[0660] >SEQ ID NO:74 No.2-1128(Toxin deam clade)
[0661] MQGYIVDESGRVLDANGLPIASLPPADDLSKWANYTVESGLHDDLAALENRLDLLYRQQFGLPM
[0662] APPAWHLETQLAYRVATREVALRDSTLRLVMNNPGGVCDAVPLTGTGPDRQRQAVAGCIQAVK
[0663] MLLPAGTTMIIYYPDPDDDPAELLEITVRGVGRWLD
[0664] >SEQ ID NO:75 No.2-1114(Toxin deam clade)
[0665] MGAPR™GNMGVAQISIPGVQSKMAASSQIPDPTAAQRALGFVGEVNETFPSASSVWTGGDTPYLL
[0666] NRKVDSEAKILNNIAAQLGDNTSASGTINLFTERPPCESCSNTIIKFQEKYPNIKINVMDSNGVIRPS
[0667] KR
[0668] >SEQ ID NO:76 No.2-1223(MafB19 clade)
[0669] MGQGFDGVQPANEFIKAWGEAMVAEAAGLGIVAGLGRFGLWGAKGATPVVTAEGRIGNSVFTD
[0670] VNQTARPAAQANPNQPTLIADRVDAKIAAKGTPHPNGNMADAHAEIGVIQKAFNEGKTVGSDMT
[0671] MNVVGKDVCGYCRGDIAAAASKSGLKSLTIQAKDDITGLPKTYWEVGMKSIREKKI
[0672] >SEQ ID NO:77 No.2-1224(MafB19 clade)
[0673] MYLPRGTSVTSKETVAKDPVSLAGQADNEAGILVDRNVIVGGTKGVSPVVTAEGKIGGKTFTDFN
[0674] QTARPASEANASQPTLISDRVTAKADASGKVLPNGNMADAHAEIGVIQQAYTAGKTMGASMELT
[0675] VSGKAVCGYCRGDIAMAEXGLTSLEVKEVATGKTLYWQPGMRALRERN
[0676] >SEQ ID NO:78 No.2-1228(MafB19 clade)
[0677] MKKPTGSIVSPETTIVQESSKILDKKTHTSIPKVEAELIDKETGKIFKDTNQGNRPDYFLGDKSRPTLI
[0678] NDRIEAKVEKNPSKYLPNGNMASAHAEVGTIQQAFEDGITVGRDMNMKVTKEAVCGYCRGDIAA
[0679] MADKAGLKSLTVYEESTGKTLYWNPGMKSLKEKK
[0680] >SEQ ID NO:79 No.2-1229(MafB19 clade)
[0681] MAGVLAPEVYLAKTPRVGDNSARGVGDGRSTTPKVTAEAEVDGVKFNDTNQNARPSEAANPNIP
[0682] TLISDDIQVKIDKNPDKPFPNGNMATAHAEVGAIQQAYDAGKTQGKNMTMRVTGEDVCDYCRSD
[0683] LRKAADKSGLNSLSVYEETTGRTLTWTRREDGTIGKVKIIEPEG
[0684] >SEQ ID NO:80 No.2-1230(MafB19 clade)
[0685] MGVDRKTAQGYAETKQGMDTIVASVTPILGAAAAKQLSKVVDANIKVVAEGNVNGAKFSDTNQ
[0686] GARPSNLADVNKPTLIDGRIQAKIDKQNKPLPNGNMATAHAEVGVIQQAFEKGMSQGREMTMSV
[0687] SKEPVCGYCRSDIAAMADKAGLKSLTIYEETTGSVLYWQPGMKSLKIRD
[0688] >SEQ ID NO:81 No.2-1234(MafB19 clade)
[0689] MDRQTAESYTETKQGLEIIAASVTPILGSVAAKQLSKIVDANLKVVARGNVDGARFSDTNQGVRP
[0690] SQLADFNKPTLINDVVQAKIDKRPDKNYPNGNMGTAHAEVGVIQQAFDKGMTQGREMAMSVGG
[0691] KEVCNYCLSDVRIMAEKAGLKSLTIYEEATGNVLFWQQGMKKIENRGPAK
[0692] >SEQ ID NO:82 No.2-1235(MafB19 clade)
[0693] MTKSALGEIVIVVSDLVIPTNYVEILPVGKLSKVAKILKIGEDGTKSAGRLAEELAELQKVDIKFGK
[0694] TLPGAKAPITVTAESNIGGKHMFDTNQTARPEVNRTNTPTLAAGNAKIDPSNPNLTMKNAHAEIAL
[0695] IQRAYDAGLTKGETMQVLVRGKEVCDHCGQVMKTMYERSGLSKLIIHDTTSGTTTTYYKVIDAK
[0696] TKIATTKIEV
[0697] >SEQ ID NO:83 No.3-2107_cl21x07(MafB19 clade)
[0698] MDDSYYMKQALLEAQKAGERGEVPVGAVVVCKDRIIARAHNLTETLTDVTAHAEMQAITAAAST
[0699] LGGKYLNECALYVTVEPCVMCAGAIAWAQTGKLVFGAEDEKRGYQRYAPQALHPKTMVVKGV
[0700] LADECAALMKNFFAAKRK
[0701] >SEQ ID NO:84 No.4-2130A_cl21x30(MafB19 clade)
[0702] MTKSALGEIVIVVSDLVIPTNYVEILPVGKLSKVAKILKIGEDGTKSAGRLAEELAELQKVDIKFGK
[0703] TLPGAKAPITVTAESNIGGKHMFDTNQTARPEVNRTNTPTLAAGNAKIDPSNPNLTMKNAHAEIAL
[0704] IQRAYDAGLTKGETMQVLVRGKEVCDHCGQVMKTMYERSGLSKLIIHDTTSGTTTTYYKVIDAK
[0705] TKIATTKIEV
[0706] >SEQ ID NO:85HsJAK2(5'-3') Figure 15 ,17,18AGGCAGGCCATTCCCA
[0707] >SEQ ID NO:86HsJAK2(3'-5') Figure 15 ,17,18TCCGTCCGGTAAGGGT
[0708] >SEQ ID NO:87HsSIRT6(5'-3') Figure 15 ,17,18TCGCCGTACGCGGACAAGGG
[0709] >SEQ ID NO:88HsSIRT6(3'-5') Figure 15 ,17,18AGCGGCATGCGCCTGTTCCC
[0710] >SEQ ID NO:89HsHEK2(5'-3') Figure 16A GAACACAAAGCATAGACTGCGGG
[0711] >SEQ ID NO:90HsHEK2(3'-5') Figure 16A CTTGTGTTTCGTATCTGACGCCC
[0712] >SEQ ID NO:91HsWFS1(5'-3') Figure 16A CAGCAGTATGGTGCGCTGTGCGG
[0713] >SEQ ID NO:92HsWFS1(3'-5') Figure 16A GTCGTCATACCACGCGACACGCC
[0714] >SEQ ID NO:93OsAAT Figure 21
[0715] ACAAGGATCCCAGCCCCGTGAAGG
[0716] >SEQ ID NO:94OsACC1 or OsACC-T1 Figure 21 TCTCAGCATAGCACTCAATGCGGTCTGGG>SEQID NO:95OsCDC48-T1 Figure 21 TAGCACCCATGACAATGACATGG
[0717] >SEQ ID NO:96OsCDC48-T2 Figure 21 GACCAGCCAGCGTCTGGCGCCGG
[0718] >SEQ ID NO:97OsDEP1 Figure 21
[0719] CTAGCACATGAGAGAACAATATTGGG
[0720] >SEQ ID NO:98OsODEV Figure 21
[0721] GCACACACACACTAGTACCTCTGG
[0722] >SEQ ID NO:99HsEMX1 Figure 22
[0723] ACAAAGTACAAACGGCAGAAGCTGGAGG>SEQ ID NO:100HsHEK2 Figure 22
[0724] ACTGGAACACAAAGCATAGACTGCGGG
[0725] >SEQ ID NO:101HsWFS1 Figure 22 CCTGGCAGCAGTATGGTGCGCTGTGCGG
[0726] >SEQ ID NO:102MmHPD-T1 Figure 30C GCAACCAACCCGACCAAGAAATGCAGT
[0727] >SEQ ID NO:103MmHPD-T2 Figure 30CAGTCATTCAACGTCACAACCACCAGGT
[0728] >SEQ ID NO:104 GmALS1-T1 Figure 30D TCTCCATCGACGCACCGCCGGGG
[0729] >SEQ ID NO:105 GmALS1-T2 Figure 30D CAGGTCCCCCGCCGGATGATCGG
[0730] >SEQ ID NO:106 GmALS1-T3 Figure 30D GATCCATTACTGGGAATCATCGG
[0731] >SEQ ID NO:107 GmPPO2 Figure 30D AAGCGCTATATTGTGAAAAATGG
[0732] >SEQ ID NO:108 GmEPSPS Figure 30D CAATGCGTCCTTTGACAGCAGCTGTGG
[0733] >SEQ ID NO:109 GmPPO2 wild-type Figure 30 FCATAAGCGCTATATTGTGAAAAATGGGGCA > SEQID NO:110 GmPPO2 edited Figure 30 FCATAAGTGCTATATTGTGAAAAATGGGGCA
Claims
1. A protein clustering method based on three-dimensional structure, comprising: (1) Obtain the sequences of multiple candidate proteins from the database; (2) Predict the three-dimensional structure of each of the candidate proteins using a protein prediction program; (3) Use a scoring function to perform multiple structure comparisons on the three-dimensional structures of the multiple candidate proteins to obtain a structure similarity matrix; (4) Cluster the multiple candidate proteins based on the structural similarity matrix using a phylogenetic tree construction method. The candidate protein is cytosine deaminase. The scoring function used in step (3) is TM-score, and its calculation formula is as follows: Where L N It is the length of the target protein's amino acid sequence, L T It is the length of the amino acid sequence that appears simultaneously in both the template and the target structure, d i d0 is the distance between the i-th pair of residues in the template and the target structure, d0 is the scale of the normalized matching difference, and "Max" represents the maximum value after optimal spatial superposition.
2. The protein clustering method of claim 1, wherein in step (1) the sequences of the plurality of candidate proteins are obtained by means of annotation information in a database; or wherein in step (1) the sequences of the plurality of candidate proteins are obtained by means of ...
3. The protein clustering method of claim 1 or 2, wherein the protein structure prediction program in step (2) is selected from AlphaFold2, RoseTT or other programs capable of predicting protein structure.
4. The protein clustering method of claim 1 or 2, wherein the scoring function used in step (3) includes TM-score, RMSD, LDDT, GDT score, QSC, FAPE or other scoring functions that can score protein structural similarity.
5. The protein clustering method of claim 1 or 2, wherein the phylogenetic tree construction method in step (4) is the weighted pairwise pairwise UPGMA method. The UPGMA mentioned above includes obtaining the distance between any two proteins using the following formula: in d (ABX) The distance between two of these points is shown.
6. The protein clustering method of claim 1 or 2, wherein step (4) obtains a clustering dendrogram of the plurality of candidate proteins.
7. A method for predicting protein function based on three-dimensional structure, the method comprising clustering multiple candidate proteins using the protein clustering method according to any one of claims 1-6, and then predicting the function of the candidate proteins based on the clustering results.
8. The protein function prediction method of claim 7, wherein the plurality of candidate proteins includes at least one reference protein with a known function.
9. The protein function prediction method of claim 8, wherein the function of other candidate proteins in the branch or subbranch to which the reference protein is located is predicted by the position of the reference protein with known function in the clustering dendrogram.
10. The protein function prediction method of claim 8 or 9, wherein the reference protein is cytosine deaminase.
11. The protein function prediction method of claim 10, wherein the reference protein is a reference cytosine deaminase, and the reference cytosine deaminase is rAPOBEC1 with the sequence shown in SEQ ID No: 64 or DddA with the sequence shown in SEQ No:
65.
12. A method for identifying the minimum functional domain of a protein based on three-dimensional structure, comprising: a) By comparing the structures of multiple candidate proteins clustered together by the protein clustering method of any one of claims 1-6 and clustered in the same branch or subbranch, a conserved core structure is determined. b) Identify the conservative core structure as a minimal functional domain.
13. The method of claim 12, wherein the plurality of candidate proteins includes at least one reference protein with a known function.
14. The method of claim 13, wherein the reference protein is cytosine deaminase.
15. The method of claim 14, wherein the reference protein is a reference cytosine deaminase, and the reference cytosine deaminase is rAPOBEC1 with the sequence shown in SEQ ID No: 64 or DddA with the sequence shown in SEQ No: 65.