Cytidine deaminase derived from metagenomic mining and biological materials and applications thereof
By discovering and validating a novel cytidine deaminase L70, the problem of low efficiency in cytosine base editing in existing technologies has been solved, enabling efficient gene editing in maize organelles and nuclei.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA AGRI UNIV
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-24
AI Technical Summary
Existing cytosine base editing technologies based on DddA deaminase have low editing efficiency in maize organelles or nuclei, making it difficult to achieve widespread gene editing.
A novel cytidine deaminase, L70, was discovered and validated. It has a structural similarity to the known DddA protein but a low sequence similarity. Through metagenomic big data bioinformatics mining, a fusion protein containing a His tag was constructed, purified, and expressed for application in cytosine base editing.
It improves the efficiency and scope of cytosine base editing, enabling extensive gene editing both in vitro and in cells, and has good application potential.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This application belongs to the field of protein technology, specifically relating to cytosine deaminases derived from metagenomic mining and their biomaterials and applications. Background Technology
[0002] Genome editing is a genetic engineering technique that uses sequence-specific nucleases to modify specific locations in an organism's genome. The process involves creating double-strand breaks (DSBs) in the genome, which then triggers the cell's endogenous repair mechanisms, such as non-homologous end joining repair or homologous recombination repair, to perform targeted artificial modifications such as knockout, replacement, or insertion at target sites in the genome. Currently, commonly used genome editing technologies include zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and clustered regularly spaced short palindromic repeats (CRISPR / Cas). Compared to the first two gene editing technologies, CRISPR / Cas technology is more precise, convenient, and efficient.
[0003] Single-nucleotide polymorphisms (SNPs) are the genetic basis of agronomic traits in crops, and changes in many important crop traits are often caused by variations in a single base. Currently, novel base editing systems based on CRISPR, such as cytosine and adenine base editing systems, have been developed and widely applied in various organisms, including plants and animals. Cytidine deaminases successfully applied in plant base editing systems include APOBEC1, APOBEC3, and DddA, while adenosine deaminases include TadA7.10 and TadA8e.
[0004] Because DddA cytidine deaminase is toxic, cytosine base editing systems based on DddA deaminase primarily utilize two different Cas9 cells or two TALENs to bind to the split DddA halves, forming a fusion protein. Guided by sgRNA, the two halves only regain catalytic activity upon binding to the target gene region, deaminating the C base within the editing window to dU. Through DNA replication and repair, base editing from CG to TA in the target gene within organelles and the cell nucleus is ultimately achieved. However, currently, cytosine base editing technology based on DddA deaminase in maize organelles or the cell nucleus still suffers from low editing efficiency. Summary of the Invention
[0005] The purpose of this application is to discover proteins with natural sequence-free characteristics and improve the efficiency and scope of base editing. Specifically, the technical problem to be solved by this application is how to improve the editing efficiency and scope of cytosine base editors.
[0006] To address the aforementioned technical problems, this application provides a protein, which may be at least one of the following:
[0007] A1) The amino acid sequence of the protein is SEQ ID NO.2;
[0008] A2) The amino acid sequence is the protein consisting of positions 15-147 of SEQ ID NO.2;
[0009] A3) Proteins whose amino acid sequences shown in A1) or A2) have undergone substitution and / or deletion and / or addition of amino acid residues, have at least 70% similarity to proteins shown in A1) or A2), and are associated with deaminases.
[0010] A4) is a fusion protein obtained by attaching a tag to the N-terminus and / or C-terminus of A1), A2), or A3).
[0011] Furthermore, the connection described in A4) can be linked via peptide bonds.
[0012] Furthermore, the tags mentioned in this article include, but are not limited to: GST (glutathione thiotransferase) tag protein, His tag protein (His-tag), MBP (maltose-binding protein) tag protein, Flag tag protein, SUMO tag protein, HA tag protein, Myc tag protein, eGFP (enhanced green fluorescent protein), eCFP (enhanced cyan fluorescent protein), eYFP (enhanced yellow-green fluorescent protein), mCherry (monomer red fluorescent protein), or AviTag tag protein.
[0013] In some embodiments of this application, the label is a His label.
[0014] This application also provides biomaterials related to the above-mentioned proteins, said biomaterials may be at least one of the following:
[0015] B1) Nucleic acid molecules that encode the above proteins;
[0016] B2) An expression cassette containing the nucleic acid molecule described in B1);
[0017] B3) A recombinant vector containing the nucleic acid molecule described in B1) or a recombinant vector containing the expression cassette described in B2);
[0018] B4) Recombinant microorganisms containing the nucleic acid molecules described in B1), recombinant microorganisms containing the expression cassette described in B2), or recombinant microorganisms containing the recombinant vector described in B3);
[0019] B5) A transgenic plant cell line containing the nucleic acid molecule described in B1), a transgenic plant cell line containing the expression cassette described in B2), or a transgenic plant cell line containing the recombinant vector described in B3);
[0020] B6) Transgenic plant tissue containing the nucleic acid molecules described in B1), transgenic plant tissue containing the expression cassette described in B2), or transgenic plant tissue containing the recombinant vector described in B3);
[0021] B7) A transgenic plant organ containing the nucleic acid molecule described in B1), a transgenic plant organ containing the expression cassette described in B2), or a transgenic plant organ containing the recombinant vector described in B3).
[0022] Furthermore, in the biological material described, the nucleic acid molecule in B1) may be at least one of the following:
[0023] g1) The coding sequence of its coding strand is the cDNA molecule or DNA molecule shown in positions 1-444 of SEQ ID NO.1;
[0024] g2) Its coding strand has a nucleotide sequence that is the cDNA molecule or DNA molecule shown at positions 43-444 of SEQ ID NO.1;
[0025] g3) has at least 70% identity with the nucleotide sequence defined by g1) or g2) and is a cDNA molecule or DNA molecule encoding the protein.
[0026] In some embodiments of this application, the expression cassette utilizes the T7 promoter to initiate the transcriptional translation of a portion of the start codon of the vector backbone with deaminase L70.
[0027] In the above text, the recombinant microorganisms may specifically be bacteria, yeast, algae, and fungi.
[0028] This application also provides the use of the above-mentioned protein as a cytidine deaminase.
[0029] This application also provides the use of the above-mentioned protein in the preparation of cytidine deaminase.
[0030] This application also provides the application of the above-mentioned biomaterials in the preparation of cytidine deaminase.
[0031] This application also provides the application of the aforementioned proteins in single-base editing.
[0032] This application also provides the use of the above-mentioned proteins in the preparation of single-base edited products.
[0033] This application also provides the application of the aforementioned biological materials in single-base editing.
[0034] This application also provides the application of the above-mentioned biomaterials in the preparation of single-base edited products.
[0035] In this application, the single-base editing may be cytosine base editing in cytosine deoxynucleoside (or cytosine nucleoside).
[0036] Furthermore, the cytosine base editing can be catalyzed by deamination of cytosine bases to form uracil bases.
[0037] In this application, the product of single-base editing can be a base editor. More specifically, the product of single-base editing can be a cytosine base editor.
[0038] In this application, the base editor may be a fusion protein containing cytidine deaminase, Cas protein and uracil glycosylation inhibitor, wherein the cytidine deaminase may be the aforementioned protein.
[0039] Furthermore, the Cas protein may be nCas9.
[0040] Furthermore, the fusion protein is a protein composed of the cytidine deaminase, the Cas protein, the uracil glycosylation inhibitor, and a nuclear localization signal.
[0041] In this application, the term cytidine deaminase is equivalent to cytosine deaminase.
[0042] This application also provides a method for deamination of cytosine, the method comprising using the above-described protein (as a cytosine deaminase) to catalyze the deamination reaction of cytosine.
[0043] Furthermore, the cytosine in the method can be cytosine in nucleic acids (nucleotides).
[0044] In some embodiments of this application, the method includes the following steps: placing the above-mentioned protein (as cytosine deaminase) and substrate DNA in the same reaction system (deamination reaction system).
[0045] In some embodiments of this application, the deamination reaction system is as follows: 50 μL of the above-mentioned protein (i.e., L70 protein (20 nM)) is mixed with 10 μL of a deamination buffer solution containing DNA substrate. The deamination buffer solution containing DNA substrate has the following composition: 20 mM MES, 200 mM NaCl, 1 mM DTT, 8% Ficoll 70, and 1 μM DNA substrate.
[0046] The deamination reaction conditions were as follows: the above mixture was deaminated at 37°C for 1 hour; then 3 μL of UDG and 7 μL of 10×UDG buffer were added to the system, and the enzyme digestion reaction was carried out at 37°C for 30 minutes; 7 μL of 1M NaOH was added to the system, and the reaction was terminated by placing it at 95°C for 3 minutes. The deamination reaction sample obtained was used for subsequent electrophoresis verification.
[0047] In this application, the single-base editing occurs either extracellularly or intracellularly.
[0048] In this application, the cell may be a plant cell.
[0049] The technical effects achieved by this application are as follows:
[0050] This application utilizes metagenomic big data to conduct a comprehensive and systematic bioinformatics mining analysis of novel cytidine deaminases, and selects protein L70, which has a high structural similarity to the known DddA protein but low sequence similarity (identity < 62.92%), for functional verification. As a cytidine deaminase, protein L70 exhibits high catalytic activity and has no target sequence bias, making it widely applicable for gene editing in vitro and in cells / tissues, demonstrating significant application potential and development value. Attached Figure Description
[0051] Figure 1 Phylogenetic analysis of deaminase families.
[0052] Figure 2 This section compares the structures of novel deaminases with known DddA proteins. In the protein structure alignment results, yellow represents novel deaminases, and pink represents known DddA proteins. The RMSD value represents the root mean square deviation between the ligand structure and the reference structure; it is an indicator used to measure the structural similarity between two structures, with a smaller value indicating greater structural similarity. The TM-Score is a template modeling score; a higher score indicates more similar protein structures.
[0053] Figure 3 This is a schematic diagram of the deaminase purification system.
[0054] Figure 4 This is a system for analyzing the deamination activity of deaminases.
[0055] Figure 5 These are the results of two in vitro verification tests of deamination activity.
[0056] Figure 6 To validate DNA substrates in a target sequence preference system in vitro.
[0057] Figure 7 The results of in vitro validation of target sequence preference for candidate proteins. Detailed Implementation
[0058] I. Terminology in this application:
[0059] Examples of resources describing many of the molecular biology-related terms used in this article can be found in the following literature: Alberts et al., Molecular Biology of The Cell, 5th ed., Garland Science Publishing, Inc.: New York, 2007; Rieger et al., Glossary of Genetics: Classical and Molecular, 5th ed., Springer-Verlag: New York, 1991; King et al., A Dictionary of Genetics, 6th ed., Oxford University Press: New York, 2002; and Lewin, GenesIX, Oxford University Press: New York, 2007.
[0060] Any references cited in this article, including, for example, all patents, published patent applications and non-patent publications, are incorporated in their entirety by reference.
[0061] For ease of understanding this application, several terms and abbreviations used herein are defined as follows:
[0062] In this application, "identity" refers to the similarity of amino acid or nucleotide sequences. The similarity of amino acid sequences (or nucleotide sequences) can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, by using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, and setting the Gap existence cost, Perresidue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, and performing a search for the similarity of a pair of amino acid sequences, the similarity value (%) can be obtained.
[0063] Specifically, the consistency of 70% or more can be 75% or more. Specifically, the consistency of 75% or more can be 80% or more. Specifically, the consistency of 80% or more can be 85% or more. Specifically, the consistency of 85% or more can be 90% or more. Specifically, the consistency of 90% or more can be 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more. More specifically, the consistency of 70% or more can be at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% consistency.
[0064] When used in a list of two or more items, the term "and / or" means that any of the listed items can be used alone or in combination with any one or more of the listed items. For example, the expression "A and / or B" is intended to mean at least one or both of A and B, i.e., A alone, B alone, or a combination of A and B. The expression "A, B and / or C" means A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B and C.
[0065] The term "comprising" is not intended to be restrictive, but rather inclusive and implies the presence of other elements besides those listed, and can be interpreted as "including but not limited to". The term "comprising" also encompasses the terms "consisting of" and "substantially consisting of". In this document, the terms "including" and "comprise" are used interchangeably.
[0066] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, proteins, peptides, or polypeptides are at least 3 amino acids in length. Proteins, peptides, or polypeptides can refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified, for example, by adding chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. Proteins, peptides, or polypeptides can also be single molecules or can be multi-molecular complexes. Proteins, peptides, or polypeptides can simply be fragments of naturally occurring proteins or peptides. Proteins, peptides, or polypeptides can be naturally occurring, recombinant, or synthetic, or any combination thereof. Any protein provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers.
[0067] As used in this article, the term "fusion protein" refers to a hybrid polypeptide containing protein domains from at least two different proteins. One protein may be located at the N-terminal (N-terminal) portion or the C-terminal (C-terminal) portion of the fusion protein, thus forming an "N-terminal fusion protein" or a "C-terminal fusion protein," respectively.
[0068] The term "biomaterial" refers to any material that carries genetic information and is capable of self-replication or replication within a biological system, such as genes, plasmids, microorganisms, animals, and plants.
[0069] As used in this article, "plant" includes explants, plant parts, seedlings, plantlets, or whole plants at any stage of regeneration or development.
[0070] As used in this article, “cereals” refers to monocotyledonous crops of the Poaceae or Gramineae family, and is typically harvested for their seeds, including, for example, corn, wheat, rice, millet, barley, sorghum, oats, and rye.
[0071] As used herein, "plant part" can refer to any organ or intact tissue of a plant, such as meristematic tissue, bud organs / structures (e.g., leaves, stems, or nodes), roots, flowers or floral organs / structures (e.g., flowers, bracts, sepals, petals, stamens, carpels, anthers, and ovules), seeds (e.g., embryo, endosperm, and seed coat), fruits (e.g., mature ovaries), propagules, or other plant tissues (e.g., vascular tissue, dermal tissue, ground tissue, etc.) or any part thereof. The plant part in this application can be viable, non-viable, renewable, and / or non-renewable. "Propagule" can include any plant part that can grow into a whole plant.
[0072] Plant cells are biological cells of plants, derived from plants or derived from cultures obtained by culturing cells taken from plants. As used herein, “transgenic plant cell” means any plant cell transformed with a stably integrated recombinant DNA molecule, construct, expression cassette, or sequence. Transgenic plant cells can include original transformed plant cells, transgenic plant cells regenerated or developed from R0 generation transgenic plant cells, transgenic plant cells cultured from another transgenic plant cell, or transgenic plant cells from any progeny or offspring of a transformed R0 generation plant, including cells of plant seeds or embryos, or cultured plant cells, callus cells, etc.
[0073] As is commonly understood in the art, the term "promoter" generally refers to a DNA containing an RNA polymerase binding site, a transcription start site, and / or a TATA box that assists or promotes the transcription of transcribed DNA. Promoters can be artificially synthesized, modified, or derived from known or naturally occurring promoters. Promoters can also include chimeric promoters comprising combinations of two or more heterologous sequences. Therefore, the promoters of this application may include variants of promoter sequences that are compositionally similar but not identical to other promoter sequences provided herein.
[0074] Promoters can be classified according to various criteria related to the expression patterns of the associated coding or transcribed sequences or genes (including transgenes) operably linked to them, such as constitutive, developmental, tissue-specific, and inducible promoters. A promoter that drives expression in all or most tissues of a plant is called a "constitutive" promoter. A promoter that drives expression at certain times or stages of development is called a "developmental" promoter. A promoter that drives enhanced expression in certain tissues of a plant relative to other tissues is called a "tissue-enhancing" or "tissue-preferred" promoter. Therefore, a "tissue-preferred" promoter elicits relatively high or preferential expression in a specific tissue of the plant, but lower expression levels in other tissues. A promoter that is expressed in a specific tissue of the plant but rarely or not expressed in other tissues is called a "tissue-specific" promoter. An "inducible" promoter is a promoter that initiates transcription in response to environmental stimuli (e.g., cold, drought, or light) or other stimuli (e.g., injury or chemical application). Promoters can also be classified according to their origin, such as heterologous, homologous, chimeric, synthetic, etc.
[0075] The term "transcribed DNA" refers to DNA that can be transcribed into RNA molecules.
[0076] The term "operationally ligated" can refer to a functional connection between a promoter and transcribed DNA, enabling the promoter to function and initiate transcription of the transcribed DNA. The term "operationally ligated" can also refer to a functional connection between other regulatory elements and a target gene to regulate the transcription and / or expression of the target gene.
[0077] As used herein, an "expression cassette" refers to a cassette containing at least transcribed DNA operatively linked to one or more regulatory elements, typically at least a promoter and a 3' UTR (such as a terminator).
[0078] In this application, the expression cassette refers to DNA capable of expressing the protein in a host cell (such as a microbial cell). This DNA may include not only a promoter to initiate transcription of the protein gene, but also a terminator to terminate transcription. Furthermore, the expression cassette may also include an enhancer sequence. Promoters that can be used in this application include, but are not limited to: constitutive promoters, tissue-, organ-, and development-specific promoters, and inducible promoters. Examples of promoters include, but are not limited to: the Ubiquitin promoter from maize; the constitutive promoter 35S from cauliflower mosaic virus; the wound-inducible promoter from tomato, leucine aminopeptidase ("LAP", Chao et al. (1999) Plant Physiology 120:979-992); the chemically induced promoter from tobacco, pathogenesis-associated protein 1 (PR1, induced by salicylic acid and BTH (benzothiadiazole-7-thiohydroxy acid S-methyl ester)); the tomato protease inhibitor II promoter (PIN2) or the LAP promoter (both induced by methyl jasmonic acid); the heat shock promoter (US Patent 5,187,267); the tetracycline-inducible promoter (US Patent 5,057,422); and seed-specific promoters, such as the millet seed-specific promoter pF128 (CN101063139B (Chinese Patent 2007 1)). 0099169.7), seed storage protein-specific promoters (e.g., promoters of beta-conglycin, napin, oleosin, and soybean beta-conglycin (Beachy et al. (1985) EMBO J.4:3047-3053)). They can be used alone or in combination with other plant promoters. All references cited herein are cited in full.Suitable transcription terminators include, but are not limited to: Agrobacterium carmine synthase terminator (NOS terminator), cauliflower mosaic virus CaMV 35S terminator, tml terminator, pea rbcS E9 terminator, and carmine and octopine synthase terminators (see, for example: Odell et al. (1985) Nature 313:810; Rosenberg et al. (1987) Gene, 56:125; Guerineau et al. (1991) Mol. Gen. Genet, 262:141; Proudfoot (1991) Cell, 64:671; Sanfacon et al. Genes Dev., 5:141; Mogen et al. (1990) Plant Cell, 2:1261; Munroe et al. (1990) Gene, 91:151; Ballad et al. (1989) Nucleic Acids Res. 17:7891; Joshi et al. (1987) Nucleic Acid Res., 15:9627.
[0079] As used herein, the term "vector" refers to any construct that can be used for transformation purposes, i.e., to introduce heterologous DNA into a host cell. Examples include plasmids, granules, viruses, bacteriophages, or linear or circular DNA.
[0080] As used herein, the term "bacterial solution" has the same meaning as "culture," which refers to a liquid or solid product (all substances within the culture container) that has grown a microbial community after artificial inoculation and cultivation. That is, it is a product obtained by growing and / or amplifying microorganisms; it can be a biologically pure culture of microorganisms, or it can contain a certain amount of culture medium, metabolites, or other components produced during the cultivation process. It can also be a mixture containing a certain amount of culture medium, microbial cell metabolites, and with the microbial cells removed.
[0081] II. Implementation Examples
[0082] This application utilizes metagenomic big data to conduct a comprehensive and systematic bioinformatics mining analysis of novel cytosine deaminases, and selects protein L70, which has a high structural similarity to the known DddA protein but low sequence similarity (identity < 62.92%), for subsequent experimental verification. In vitro experiments show that L70 protein has high activity and no sequence bias, and can deaminate C from AC / TC / GC / CC.
[0083] The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.
[0084] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0085] Unless otherwise specified, the quantitative experiments in the following examples were performed in triplicate, and the results were averaged.
[0086] Example 1: Novel Cytidine Deaminase Discovery Process
[0087] 1.1. All sequenced metagenomic assembly nucleic acid sequence information were retrieved and downloaded from the JGI and NCBI biological databases; all proteins were obtained by annotating protein sequences using Prodigal software; and domain annotations were performed on the predicted coding genes using the hidden Markov model in the Pfam database.
[0088] 1.2 Because DddA belongs to the SCP1.201 family, its full length is >1400 aa, and the part that actually performs deamination is the latter half of the protein—DddAtox, whose annotation entry is SCP1.201-deam. Therefore, this entry was used as a keyword for screening. All the selected candidate proteins were aligned using Maft software. Based on prior knowledge, the active site for DddAtox is glutamate E at position 1347. Therefore, proteins with the same site were screened. Since DddA participates in the type VI secretion system (T6SS) of the intercellular protein delivery system, its downstream is the immune protein DddIA. The two interact to inhibit the toxicity of DddA protein and prevent cell inactivation. To facilitate the purification of DddA protein in subsequent experiments, we searched for the presence of a corresponding immune protein, DddIA, within a 10kb range upstream and downstream of the screened DddA homologous proteins. The lengths of the screened proteins varied, but only the DddAtox portion was required. Therefore, the sequences were aligned and truncated again to obtain 128 truncated sequences that enable the DddA homologous proteins to function.
[0089] 1.3. Using IQ-TREE software, a phylogenetic tree was constructed by combining all the mined proteins with other deaminase family protein sequences reported in the literature (Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems), and visualized using ITOL software. The figure shows that the 128 novel deaminases mined through the established bioinformatics workflow were classified into the SCP1.201 clade; therefore, these proteins do indeed belong to the SCP1.201 family. Figure 1 ).
[0090] 1.4 To further screen for novel active deaminases, AlphaFold 2 software was used to predict the structure of all discovered proteins, and TM-align structures were compared pairwise with known DddA proteins. Finally, 27 proteins with structures similar to known proteins were selected as novel deaminases for subsequent experimental verification and analysis. These 27 proteins were numbered as deaminases L8, L64, L65, L67, L70, L71, L72, L74, L78, L79, L80, L81, L83, L84, L85, L86, L88, L89, L90, L91, L115, L135, L136, L138, S45, S57, and S76. Figure 2 ).
[0091] 1.5. The mined deaminase sequence was optimized using maize codons, and a company was commissioned to synthesize the encoding genes for the deaminase and downstream immunoproteins. Taking deaminase L70 as an example, deaminase L70 and the immunoprotein encoding gene were ligated into the pET-duet-1 vector to construct a co-expression vector for the deaminase and downstream immunoprotein. This vector was then transformed into *E. coli* TSR2566 prokaryotic cells for induced expression and protein purification. The nucleotide sequences of deaminase L70 and the immunoprotein encoding gene are SEQ ID NO.1.
[0092] The co-expression vector for the deaminase and downstream immunoprotein was named pETduet1-L70. The structure of pETduet1-L70 is as follows: the fragment between the BamHI and XhoI sites in the pET-duet-1 vector was replaced with a double-stranded DNA molecule containing the nucleotide sequence from positions 42 to 1024 of SEQ ID NO. 1, while keeping the other nucleotide sequences of the pET-duet-1 vector unchanged. Specifically, positions 43-444 of SEQ ID NO. 1 contain the coding gene for the deaminase L70, positions 507-525 contain the T7 promoter, and positions 593-982 contain the coding gene for the immunoprotein.
[0093] The pETduet1-L70 vector can express an L70 fusion protein and an L70 immunoprotein carrying the vector sequence and a His tag. The amino acid sequence of the L70 fusion protein is SEQ ID NO.2, where positions 1-14 of SEQ ID NO.2 are the amino acid sequence containing the His tag encoded by the start codon on the pET-duet-1 vector backbone, and positions 15-147 are the amino acid sequence of the L70 deaminase. The amino acid sequence of the L70 immunoprotein is SEQ ID NO.3.
[0094] A schematic diagram of the deaminase purification system is shown below. Figure 3 As shown. The steps for inducing expression and protein purification are as follows: The constructed pETduet1-L70 vector was transformed into E. coli competent cells TSR2566 to obtain recombinant bacteria TSR2566 / pETduet1-L70. The bacteria were shaken in LB liquid medium until OD600 = 0.6-0.7, and IPTG was added to induce protein expression. The bacterial culture was collected. The deaminase L70 fusion protein (amino acid sequence is SEQ ID NO.2) obtained by induction expression in the bacterial culture was combined with an immunoprotein (amino acid sequence is SEQ ID NO.3) to form a protein complex. The bacterial culture was then purified using a nickel column to obtain the protein complex of deaminase L70 and the immunoprotein. The protein complex was then denatured and renatured using urea solutions of different concentrations (8M, 6M, 4M, 2M, and 1M). The eluent containing only deaminase L70 was then purified by ultrafiltration to obtain deaminase L70 (also known as L70 protein).
[0095] The preparation of the other 26 deaminases is the same as that of deaminase L70, the only difference being the coding sequence of the deaminase.
[0096] Using deaminase DddA as a control, the preparation method of deaminase DddA was the same as that of deaminase L70, except that the DNA molecule in SEQ ID NO.1 of the pETduet1-L70 vector was replaced with the DNA molecule shown in SEQ ID NO.4. In SEQ ID NO.4, positions 2-418 are the coding gene for deaminase DddA, positions 484-502 are the T7 promoter, and positions 570-1001 are the coding gene for the immunoprotein of deaminase DddA. The deaminase DddA obtained by induced expression (amino acid sequence SEQ ID NO.5) and the immunoprotein (amino acid sequence SEQ ID NO.6) formed a protein complex.
[0097] Example 2: In vitro activity verification
[0098] 2.1 In vitro deamination activity
[0099] DNA substrates labeled with 5'-FAM fluorophores were used to prepare a deamination reaction system with 20 nM deaminase. Deamination activity was analyzed. If the deaminase showed deamination activity, the DNA substrate would cleave from the cytosine residue, and the reaction product of the deamination reaction system would show a cleavage band around 18 nt. Figure 4 ).
[0100] The deamination reaction system was prepared by mixing 50 μL of the purified L70 protein (20 nM) obtained in Example 1 with 10 μL of a deamination buffer solution containing DNA substrate. The deamination buffer solution containing DNA substrate consisted of: 20 mM MES, 200 mM NaCl, 1 mM DTT, 80 g / L Ficoll 70, and 1 μM DNA substrate.
[0101] The deamination reaction conditions were as follows: the above reaction mixture was deaminated at 37°C for 1 h; then 3 μL of UDG (Uracil-DNA Glycosylase) and 7 μL of 10×UDG buffer were added to the system, and the enzyme digestion reaction was carried out at 37°C for 30 min; 7 μL of 1M NaOH was added to the system, and the reaction was terminated by placing it at 95°C for 3 min. The deamination reaction sample obtained was used for subsequent electrophoresis verification.
[0102] Under the action of deaminase Figure 4 In the double-stranded DNA molecule shown in Figure A, the cytosine (C) at position 18 of the upper strand undergoes a deamination reaction, becoming... Figure 4As shown in B, uracil (U) is further added to the reaction system with uracil-DNA glycosylase (UDG), causing the N-glycosidic bond between the uracil base and the sugar phosphate backbone to break, resulting in the uracil being removed from the nucleotide chain and forming a base-free site. Figure 4 (C); With the addition of NaOH, under high temperature, the double-stranded DNA molecule unwinds into single strands. The upper strand undergoes phosphodiester bond breakage at base-free sites, resulting in single-stranded DNA breaks, producing one 17nt single-stranded DNA and one 18nt single-stranded DNA. Figure 4 (D).
[0103] Figure 4 In the two nucleotide sequences A to D, the first A at the 5' end of the upper chain is modified by 5'FAM (5-Carboxyfluorescein); A(14) represents 14 A's and T(14) represents 14 T's.
[0104] Figure 4 The nucleotide sequence of the upper chain of A is as follows:
[0105] 5'-AAAAAAAAAAAAAAAACTCGCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 7).
[0106] Figure 4 The cytosine at position 18 of SEQ ID NO. 8 undergoes deamination to form uracil, and its nucleotide sequence in the sequence listing is as follows. In SEQ ID NO. 8, the T at position 18 represents uracil:
[0107] 5'-AAAAAAAAAAAAAAAACTTGCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 8).
[0108] Figure 4 Uracil is lost from the upper chain of the C-chain at the procytosine position, resulting in a base-free site. Figure 4 The nucleotide sequence of the chain above the C-chain is as follows, where N represents a base-free site:
[0109] 5'-AAAAAAAAAAAAAAAACTNGCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 9).
[0110] Figure 4 The upper chain of D breaks into two strands, with the nucleotide sequences as follows:
[0111] 5'FAM-AAAAAAAAAAAAAAAACT-3' (SEQ ID NO.10)
[0112] 5'-GCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 11);
[0113] Figure 4 The nucleotide sequences of the lower chains from A to D are as follows:
[0114] 5'-TTTTTTTTTTTTTGGCGAGTTTTTTTTTTTTTTT-3' (SEQ ID NO. 12).
[0115] Electrophoresis validation: Before electrophoresis, the marker system consisted of 1 μL of 36 nt substrate, 1 μL of 18 nt substrate, 5 μL of 2×RNA loading buffer, and 13 μL of ddH2O. 77 μL of the deamination reaction sample was mixed with 10 μL of 2×RNA loading buffer and incubated at 100°C for 5 min to denature. During gel electrophoresis, the voltage was initially set to 100V for 30 min, then increased to 150V and run for 44 min.
[0116] Results Analysis: DNA fragments labeled with 5'-FAM fluorophores showed bands when scanned using a gel scanning system (iBright FL1500 imaging system). DNA fragments without 5'-FAM fluorophores did not show bands when scanned using a gel scanning system. That is, if the deaminase has in vitro deamination activity, the 18nt fluorescently labeled DNA single strands in the deamination and denaturation products of the fluorescently labeled DNA substrate will show a band near the 18nt nucleotide level when scanned using a gel scanning system; if the deaminase has no in vitro deamination activity (or very weak activity), the reaction products of the deamination reaction system will not show a band near the 18nt nucleotide level when scanned using a gel scanning system.
[0117] The deamination activity of the 28 candidate proteins was verified, and 18 deaminases with deamination activity were successfully verified (L8, L70, L71, L72, L78, L79, L80, L91, L115, L138, S57, L83, L84, L85, L86, L89, L90, and S45). The verification was repeated, and based on the combined results of the two verifications, the six deaminases with the highest deamination activity (L8, L70, L85, L91, L90, and S45) were finally selected for further verification. Figure 5 ).
[0118] 2.2 Validation of target sequence preference
[0119] The deamination reaction was carried out according to the preparation of the deamination reaction system in 2.1. The only difference from 2.2 is that the DNA substrate is different. In the DNA substrate, N represents the four different bases A, T, G, C. Figure 6This was used to determine the in vitro preference of candidate deaminases for target sequence C. The results showed that deaminase L70 can efficiently deaminate cytosine (C) in vitro in an NC (N represents the four different bases ATGC) background. Figure 7 ).
[0120] Figure 6 The first A at the 5' end of the two nucleotide sequences is modified by 5'FAM (5-Carboxyfluorescein); A(15) represents 15 A's, T(15) represents 15 T's, and N is a, c, g, or t.
[0121] Figure 6 The nucleotide sequence of the upper chain in the middle chain is as follows:
[0122] 5'-AAAAAAAAAAAAAAAGGNCGGAAAAAAAAAAAAAAA-3' (SEQ ID NO. 13).
[0123] Figure 6 The nucleotide sequences of the middle and lower chains are as follows:
[0124] 5'-TTTTTTTTTTTTTCCGNCCTTTTTTTTTTTTTTT-3' (SEQ ID NO. 14).
[0125] The sequence in this application is as follows:
[0126] SEQ ID NO.1, positions 1-41 of SEQ ID NO.1 are the pETduet1-L70 vector backbone sequence, positions 43-444 of SEQ ID NO.1 are the coding gene for deaminase L70, positions 1-444 of SEQ ID NO.1 are the coding sequence for the L70 fusion protein, positions 507-525 are the T7 promoter, and positions 593-982 are the coding gene for an immune protein.
[0127]
[0128] In SEQ ID NO.2, positions 1-10 represent the His-tagged amino acid sequence encoded by the start codon to the sequence on the vector; positions 11-14 are the linker; and positions 15-147 are the amino acid sequence of deaminase L70.
[0129] MGSSHHHHHHSQDPYQPANIPSSEVVLPEFDGKTTYGELRTPDGKSIPLQSGDPDPQYSNYVSSSHVEGKAAQYMRENGIEQATVYHNNANGTCGYCDKMLPTLLPDGSELTVIPPASAVPNNPQAVAAPKTYTGNSAVPKTNPRFK.
[0130] The amino acid sequence of the immune protein of SEQ ID NO.3, L70:
[0131] MLAVSHFGGTEECSSADDLKALLGMRFENDSNEFWLNHENASYPCMSMMVVGGYACLHYFPDDTSTGFVSSGIENGLDPDGITVFYTNTSSEEIEVFNDLVVDAQAGVEALIGFFDTPTMPDSIEWLEL.
[0132]
[0133] SEQ ID NO.5, amino acid sequence of the deaminase DddA fusion protein:
[0134] MGSSHHHHHHSQDPGSYALGPYQISAPQLPAYNGQTVGTFYYVNDAGGLESKVFSSGGPTPYPNYANAGHVEGQSALFMRDNGISEGLVFHNNPEGTCGFCVNMTETLLPENAKMTVVPPEGAIPVKRGATGETKVFTGNSNSPKSPTKGGC.
[0135] SEQ ID NO.6, DddA immune protein amino acid sequence:
[0136] MYADDFDGEIEIDEVDSLVEFLSRRPAFDANNFVLTFEESGFPQLNIFAKNDIAVVYYMDIGENFVSKGNSASGGTEKFYENKLGGEVDLSKDCVVSKEQMIEAAKQFFATKQRPEQLTWSELGSGSKRPAATKKAGQAKKKK.
[0137] The present application has been described in detail above. Those skilled in the art will recognize that the present application can be implemented in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments are given in this application, it should be understood that further modifications can be made to the present application. In summary, in accordance with the principles of this application, this application is intended to include any changes, uses, or improvements to the present application, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.
Claims
1. A protein, characterized by: The protein is at least one of the following: A1) The amino acid sequence of the protein is SEQ ID NO.2; A2) The amino acid sequence is the protein consisting of positions 15-147 of SEQ ID NO.2; A3) is a fusion protein obtained by attaching a tag to the N-terminus and / or C-terminus of A1) or A2).
2. A biomaterial relating to the protein of claim 1, wherein the biomaterial is at least one of the following: B1) A nucleic acid molecule encoding the protein of claim 1; B2) An expression cassette containing the nucleic acid molecule described in B1); B3) A recombinant vector containing the nucleic acid molecule described in B1) or a recombinant vector containing the expression cassette described in B2); B4) Recombinant microorganisms containing the nucleic acid molecules described in B1), recombinant microorganisms containing the expression cassette described in B2), or recombinant microorganisms containing the recombinant vector described in B3); B5) A transgenic plant cell line containing the nucleic acid molecule described in B1), a transgenic plant cell line containing the expression cassette described in B2), or a transgenic plant cell line containing the recombinant vector described in B3); B6) Transgenic plant tissue containing the nucleic acid molecules described in B1), transgenic plant tissue containing the expression cassette described in B2), or transgenic plant tissue containing the recombinant vector described in B3); B7) A transgenic plant organ containing the nucleic acid molecule described in B1), a transgenic plant organ containing the expression cassette described in B2), or a transgenic plant organ containing the recombinant vector described in B3).
3. The biomaterial according to claim 2, characterized in that: B1) The nucleic acid molecule is at least one of the following: g1) The coding sequence of its coding strand is the cDNA molecule or DNA molecule shown in positions 1-444 of SEQ ID NO.1; g2) Its coding strand has a nucleotide sequence that is the cDNA molecule or DNA molecule shown at positions 43-444 of SEQ ID NO.1; The nucleotide sequence defined by g3) has at least 70% identity with g1) or g2) and encodes a cDNA molecule or DNA molecule of the protein of claim 1.
4. The use of the protein according to claim 1 as a cytidine deaminase.
5. The use of the protein according to claim 1 in the preparation of cytidine deaminase.
6. The use of the biomaterial described in claim 2 or 3 in the preparation of cytidine deaminase.
7. The application of the protein according to claim 1 in single-base editing.
8. The use of the protein of claim 1 in the preparation of products with single-base editing.
9. The application of the biomaterial described in claim 2 or 3 in single-base editing.
10. The use of the biomaterial of claim 2 or 3 in the preparation of products with single-base editing.
Citation Information
Patent Citations
Seed specificity highly effective promoter and its application
CN101063139A
Seed specific highly effective promoter and its application
CN101063139B
Recombinant DNA: transformed microorganisms, plant cells and plants: a process for introducing an inducible property in plants, and a process for producing a polypeptide or protein by means of plants or plant cells
US5057422A
Plant proteins, promoters, coding sequences and use
US5187267A
Cytosine deaminase and related biological material and application thereof
CN119913132A