Single-base editor, deaminase used therein, and use thereof

By constructing a fusion protein containing cytidine deaminase L8, Cas protein, and uracil glycosylation inhibitor, the problem of low basic base editing efficiency of DddA deaminase was solved, and efficient base editing in maize organelles and nuclei was achieved.

WO2025260911A1PCT designated stage Publication Date: 2025-12-26CHINA AGRI UNIV

Patent Information

Application Number
PCT/CN2025/087668
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2025-05-07
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing DddA deaminase-based cytosine base editing technologies have low editing efficiency in maize organelles or nuclei, making it difficult to effectively improve the efficiency and scope of base editing.

Method used

A fusion protein was constructed, comprising cytidine deaminase L8, Cas protein, and uracil glycosylation inhibitor, linked by peptide bonds to improve the efficiency and scope of base editing.

Benefits of technology

This improved the efficiency and scope of cytosine base editing, enabling highly efficient base editing in maize organelles and the nucleus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025087668-FTAPPB-I100001
    Figure PCTCN2025087668-FTAPPB-I100001
  • Figure PCTCN2025087668-FTAPPB-I100002
    Figure PCTCN2025087668-FTAPPB-I100002
  • Figure PCTCN2025087668-FTAPPB-I100003
    Figure PCTCN2025087668-FTAPPB-I100003
Patent Text Reader

Abstract

The present invention belongs to the technical field of genetic engineering. Provided are a single-base editor, a deaminase used therein, and the use thereof. The technical problems to be solved are to identify a naturally occurring cytosine deaminase without sequence preference, construct a base editor, and improve the efficiency and scope of base editing. In order to solve the technical problems above, a cytosine base editor is provided. The cytosine base editor is a fusion protein, wherein the fusion protein is a protein containing a cytidine deaminase, a Cas protein and a uracil-DNA glycosylase inhibitor, and the cytidine deaminase is a protein having an amino acid sequence of positions 28-164 of SEQ ID NO. 2. Further provided is the use of the fusion protein above and a biomaterial related thereto in plant single-base editing. The single-base editor can improve the efficiency of cytosine base editing and accurately mediate the base mutation of a target, and is widely applicable in the cells of maize and even other plants.
Need to check novelty before this filing date? Find Prior Art

Description

Single-base editors and their deaminases and applications Technical Field

[0001] This application belongs to the field of genetic engineering technology, specifically relating to single-base editors and the deaminases used therein, and their applications. This application claims priority to Chinese patent application No. 202410776475.3, filed on June 17, 2024, the entire contents of which are incorporated herein by reference. Background Technology

[0002] Genome editing is a genetic engineering technique that uses sequence-specific nucleases to modify specific locations in an organism's genome. The process involves creating double-strand breaks (DSBs) in the genome, which in turn triggers the cell's endogenous repair mechanisms, such as non-homologous end joining repair or homologous recombination repair, resulting in sequence insertion, deletion, and substitution. Currently, commonly used genome editing technology platforms include ZFN, TALEN, and CRISPR / Cas systems. Among these, the CRISPR / Cas genome editing platform is the most efficient.

[0003] Single nucleotide polymorphisms (SNPs) are the genetic basis of agronomic traits in crops, and changes in many important crop traits are often caused by variations in a single base. Currently, novel base editing systems based on CRISPR, such as cytosine and adenine base editing systems, have been developed and widely applied in various organisms, including plants and animals. Cytidine deaminases successfully applied in plant base editing systems include APOBEC1, APOBEC3, and DddA, while adenosine deaminases include TadA7.10 and TadA8e.

[0004] Because DddA cytidine deaminase is toxic, cytosine base editing systems based on DddA deaminase primarily utilize two different Cas9 cells or two TALENs to bind to the split DddA halves, forming a fusion protein. Guided by sgRNA, the two halves only regain catalytic activity upon binding to the target gene region, deaminating the C base within the editing window to dU. Through DNA replication and repair, base editing from CG to TA in the target gene within organelles and the cell nucleus is ultimately achieved. However, currently, cytosine base editing technology based on DddA deaminase in maize organelles or the cell nucleus still suffers from low editing efficiency.

[0005] Invention Overview

[0006] The technical problem to be solved by this application is: to discover naturally occurring cytosine deaminases without sequence bias, to construct a base editor, and to improve the efficiency and scope of base editing.

[0007] To address the aforementioned technical problems, this application provides a cytosine base editor, wherein the cytosine base editor is a fusion protein (construct).

[0008] The fusion protein (construct) is a protein containing cytidine deaminase, Cas protein, and a uracil glycosylation inhibitor, wherein the cytidine deaminase is protein L8, and protein L8 is any of the following:

[0009] A1) The amino acid sequence is the protein consisting of positions 28-164 of SEQ ID NO.2;

[0010] A2) A protein that is more than 70% identical to the protein shown in A1) and is related to deaminases, obtained by substituting, deleting and / or adding amino acid residues of the amino acid sequence shown in A1).

[0011] A3) is a fusion protein obtained by attaching a tag to the N-terminus and / or C-terminus of A1) or A2).

[0012] Furthermore, the connection described in A3) can be linked via peptide bonds.

[0013] Furthermore, the fusion protein is a protein composed of the cytidine deaminase, the Cas protein, the uracil glycosylation inhibitor, and a nuclear localization signal.

[0014] In this application, the Cas protein refers to an RNA (sgRNA)-guided DNA endonuclease associated with the CRISPR system. Non-limiting examples of the Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, their homologs or modified forms thereof.

[0015] Furthermore, the Cas protein may be the Cas9 protein.

[0016] Furthermore, the Cas9 protein is not limited to a specific protein, as long as it can be used in conjunction with the sgRNA of this application. Furthermore, the Cas9 protein described herein is selected from Streptococcus pyogenes Cas9 (spCas9, subtype II-A), spCas9HF (high fidelity), nickase Cas9 (nCas9), Staphylococcus aureus Cas9 (saCas9, subtype II-A), Neisseria meningitidis Cas9 (NmCas9, subtype II-C), Francisella novicida Cas9 (FnCas9, subtype II-B), Streptococcus thermophilus Cas9 (St1Cas9, St3Cas9), Campylobacter jejuni Cas9 (CjCas9), and Treponema sp. Cas9, as well as orthologs of Cas9 from other organisms, but not limited to these. The Cas9 protein may also include high-fidelity Cas9 mutants (such as SpCas9-HF1, eSpCas9-1.1, and TrueCut). TM HiFiCas9 protein, etc.

[0017] Furthermore, in the fusion protein, the Cas protein may be nCas9.

[0018] Furthermore, in the fusion protein, nCas9 may be a protein whose amino acid sequence is SEQ ID NO.2, positions 182-1548.

[0019] Furthermore, in the fusion protein, the uracil glycosylase inhibitor may be a protein whose amino acid sequence is SEQ ID NO. 2, positions 1559-1641 and / or SEQ ID NO. 2, positions 1652-1734.

[0020] Furthermore, in the fusion protein, the amino acid sequence of the nuclear localization signal may be positions 3-9 of SEQ ID NO.2 and / or positions 1766-1781 of SEQ ID NO.2.

[0021] Furthermore, the fusion protein may be a protein with the amino acid sequence of SEQ ID NO.2.

[0022] This application also provides biomaterials related to the above-described fusion proteins, said biomaterials being at least one of the following D1)-D7):

[0023] D1) The nucleic acid molecule encoding the fusion protein;

[0024] D2) An expression cassette containing the nucleic acid molecules described in D1);

[0025] D3) A recombinant vector containing the nucleic acid molecule described in D1) or a recombinant vector containing the expression cassette described in D2);

[0026] D4) Recombinant microorganisms containing the nucleic acid molecules described in D1), recombinant microorganisms containing the expression cassette described in D2), or recombinant microorganisms containing the recombinant vector described in D3);

[0027] D5) A transgenic plant cell line containing the nucleic acid molecule described in D1), a transgenic plant cell line containing the expression cassette described in D2), or a transgenic plant cell line containing the recombinant vector described in D3);

[0028] D6) Transgenic plant tissue containing the nucleic acid molecules described in D1), transgenic plant tissue containing the expression cassette described in D2), or transgenic plant tissue containing the recombinant vector described in D3);

[0029] D7) Transgenic plant organs containing the nucleic acid molecules described in D1), transgenic plant organs containing the expression cassette described in D2), or transgenic plant organs containing the recombinant vector described in D3).

[0030] Furthermore, in the biological material, the DNA molecule in D1) may be at least one of the following:

[0031] D11) Its coding strand has a nucleotide sequence that is the cDNA molecule or DNA molecule shown in positions 3589-8899 of SEQ ID NO.1;

[0032] D12) Its coding strand has a nucleotide sequence that is the cDNA molecule or DNA molecule shown in positions 3514-9040 of SEQ ID NO.1;

[0033] The nucleotide sequence defined by D13) has 70% or more identity with D11) and / or D12) and is a cDNA molecule or DNA molecule encoding the fusion protein.

[0034] To address the aforementioned technical problems, this application also provides the aforementioned protein L8 and related biological materials.

[0035] Protein L8 is described in any of the following:

[0036] A1) The amino acid sequence is the protein consisting of positions 28-164 of SEQ ID NO.2;

[0037] The proteins obtained by substituting and / or deleting and / or adding amino acid residues of the amino acid sequences shown in A2) and A1) have more than 70% identity with the protein shown in A1) and are related to deaminases.

[0038] A3) is a fusion protein obtained by attaching a tag to the N-terminus and / or C-terminus of A1) or A2).

[0039] Furthermore, the connection described in A3) can be linked via peptide bonds.

[0040] Furthermore, the tags mentioned in this article include, but are not limited to: GST (glutathione thiotransferase) tag protein, His tag protein (His-tag), MBP (maltose-binding protein) tag protein, Flag tag protein, SUMO tag protein, HA tag protein, Myc tag protein, eGFP (enhanced green fluorescent protein), eCFP (enhanced cyan fluorescent protein), eYFP (enhanced yellow-green fluorescent protein), mCherry (monomer red fluorescent protein), or AviTag tag protein.

[0041] In some embodiments of this application, the label is a His label.

[0042] This application also provides biomaterials related to the aforementioned protein L8, said biomaterials may be at least one of the following:

[0043] B1) The nucleic acid molecule that encodes the protein;

[0044] B2) An expression cassette containing the nucleic acid molecule described in B1);

[0045] B3) A recombinant vector containing the nucleic acid molecule described in B1) or a recombinant vector containing the expression cassette described in B2);

[0046] B4) Recombinant microorganisms containing the nucleic acid molecules described in B1), recombinant microorganisms containing the expression cassette described in B2), or recombinant microorganisms containing the recombinant vector described in B3);

[0047] B5) A transgenic plant cell line containing the nucleic acid molecule described in B1), a transgenic plant cell line containing the expression cassette described in B2), or a transgenic plant cell line containing the recombinant vector described in B3);

[0048] B6) Transgenic plant tissue containing the nucleic acid molecules described in B1), transgenic plant tissue containing the expression cassette described in B2), or transgenic plant tissue containing the recombinant vector described in B3);

[0049] B7) Transgenic plant organs containing the nucleic acid molecules described in B1), transgenic plant organs containing the expression cassette described in B2), or transgenic plant organs containing the recombinant vector described in B3).

[0050] Furthermore, in the aforementioned biological material, the nucleic acid molecule in B1) may be at least one of the following:

[0051] B11) Its coding sequence is the cDNA molecule or DNA molecule shown in positions 2-412 of SEQ ID NO.4;

[0052] B12) Its coding strand has a nucleotide sequence that is the cDNA molecule or DNA molecule shown at positions 3589-4189 of SEQ ID NO.1;

[0053] The nucleotide sequence defined by B13) has 70% or more identity with B11) and / or B11) and is a cDNA molecule or DNA molecule encoding the protein L8.

[0054] In this application, the expression cassette refers to DNA capable of expressing the protein in a host cell (such as a plant cell). This DNA may include not only a promoter to initiate transcription of the protein gene, but also a terminator to terminate transcription. Furthermore, the expression cassette may also include an enhancer sequence. Promoters that can be used in this application include, but are not limited to: constitutive promoters, tissue-, organ-, and development-specific promoters, and inducible promoters. Examples of promoters include, but are not limited to: the Ubiquitin promoter from maize; the constitutive promoter 35S from cauliflower mosaic virus; the wound-inducible promoter from tomato, leucine aminopeptidase ("LAP", Chao et al. (1999) Plant Physiology 120:979-992); the chemically inducible promoter from tobacco, pathogenesis-associated protein 1 (PR1) (induced by salicylic acid and BTH (benzothiadiazole-7-thiohydroxy acid S-methyl ester)); the tomato protease inhibitor II promoter (PIN2) or the LAP promoter (both induced by methyl jasmonic acid); the heat shock promoter (US Patent 5,187,267); the tetracycline-inducible promoter (US Patent 5,057,422); and seed-specific promoters, such as the millet seed-specific promoter pF128 (CN101063139B (Chinese Patent 2007 1)). 0099169.7)), seed-specific promoters for storage proteins (e.g., promoters of beta-conglycin, napin, oleosin, and soybean beta-conglycin (Beachy et al. (1985) EMBO J.4:3047-3053)). They can be used alone or in combination with other plant promoters. All references cited herein are cited in full. Suitable transcription terminators include, but are not limited to: Agrobacterium carmine synthase terminator (NOS terminator), cauliflower mosaic virus CaMV 35S terminator, tml terminator, pea rbcS E9 terminator, and carmine and octopine synthase terminators (see, for example: Odell et al. (1985) Nature 313:810; Rosenberg et al. (1987) Gene, 56:125; Guerineau et al. (1991) Mol. Gen. Genet, 262:141; Proudfoot (1991) Cell, 64:671; Sanfacon et al. Genes Dev., 5:141; Mogen et al. (1990) Plant Cell, 2:1261; Munroe et al. (1990) Gene, 91:151; Ballad et al. (1989) Nucleic Acids Res. 17:7891; Joshi et al. (1987) Nucleic Acid Res., 15:9627.

[0055] In some embodiments of this application, the expression cassette utilizes the Zmubi promoter (nucleotide sequence SEQ ID NO. 1, positions 1506-3498) to initiate the expression of the fusion protein of deaminase L8, DddA, linker, nCas9 (D10A), and glycosylation inhibitor protein (UGI), and terminates the expression with the RBCS E9T terminator (nucleotide sequence SEQ ID NO. 1, positions 9044-9678). The expression cassette is composed of the Zmubi promoter, a nucleic acid molecule encoding the cytosine base editor, and the RBCS E9T terminator.

[0056] In the above text, the recombinant microorganisms may specifically be bacteria, yeast, algae, and fungi.

[0057] In some embodiments of this application, the term "bacterial solution" is synonymous with "culture." The term "culture" refers to a liquid or solid product (all substances within the culture container) that has grown a microbial community after artificial inoculation and cultivation. That is, it is a product obtained by growing and / or amplifying microorganisms; it can be a biologically pure culture of microorganisms, or it can contain a certain amount of culture medium, metabolites, or other components produced during the cultivation process. It can also be a mixture containing a certain amount of culture medium, microbial cell metabolites, and with the microbial cells removed.

[0058] This application also provides the use of the aforementioned protein L8, biomaterials associated with protein L8, the fusion protein, and / or biomaterials associated with the fusion protein in single-base editing and / or the preparation of single-base edited products.

[0059] Furthermore, the product of the single-base editing can be a single-base editor.

[0060] Furthermore, the single-base editing product may be a reagent or kit containing the single-base editor.

[0061] Furthermore, in this application, the single-base editing can occur extracellularly or intracellularly.

[0062] Furthermore, in this application, the single-base editing can occur within plant cells.

[0063] Furthermore, in this application, the single-base editing can occur within organelles of plant cells.

[0064] This application also provides a method for mutating C:G to T:A in a plant genome.

[0065] The method provided in this application for mutating the C:G base pair in a plant genome to T:A includes the following steps: introducing a DNA molecule expressing the cytosine base editor and sgRNA into a recipient plant to obtain a target plant with the C:G mutated to T:A; the target sequence of the sgRNA is 5′-N19-20PAM-3′, where N19-20 consists of 19-20 Ns.

[0066] In the above method, the Cas protein in the cytosine base editor can be nCas9, the PAM (protospacer adjacent motif) is NGG, and N is A, G, C, or T.

[0067] In the above text, the plant can be a dicotyledonous plant or a monocotyledonous plant. The monocotyledonous plant can be maize. The single base editing can be replacing cytosine C with uracil.

[0068] When introducing DNA molecules expressing the cytosine base editor and sgRNA into recipient plants, PEG-mediated transformation, gene gun method, or Agrobacterium infection method can be used to introduce the gene editing toolkit into maize protoplasts or callus tissue, as is readily understood by those skilled in the art. It is well known to those skilled in the art that maize genomic DNA consists of two strands; therefore, the target nucleotide sequence can be on either of the complementary strands. For example, when the target nucleotide sequence is located in the sense strand of a functional gene, if a site-directed mutation of C to T and / or G to A occurs at a specific site in the functional gene, and if one of these mutations yields the desired amino acid in its corresponding functional protein, this system can also be used. That is, a direct base substitution on the sense strand can replace C in the triplet codon with T and / or G with A, thus obtaining a maize gene function "correction" mutant. Alternatively, when the target nucleotide sequence is located in the antisense strand of a functional gene, if a site-directed mutation of C to T occurs at a specific site in the functional gene, and if one of these mutations yields the desired amino acid in its corresponding functional protein, this system can also be used. That is, a site-directed mutation of G in the antisense strand to A occurs, thereby replacing the corresponding complementary C in the sense strand with T, thus altering the amino acid encoded by the triplet codon in the sense strand, resulting in a maize gene function "correction" mutant. Attached Figure Description

[0069] Figure 1 shows the in vitro deamination activity analysis system of deaminase.

[0070] Figure 2 shows the in vitro deamination activity of deaminase L8.

[0071] Figure 3 shows the DNA substrate structure in the in vitro validation system for target sequence preference.

[0072] Figure 4 shows the results of target sequence preference verification for in vitro deamination of deaminase L8.

[0073] Figure 5 shows the location and structure of the target fragment in the L8 vector.

[0074] Figure 6 shows the fluorescence detection of maize protoplasts L8 and DddA.

[0075] Figure 7 shows the single-base editing vectors of the L8 deaminase and DddA system.

[0076] Figure 8 shows the editing window and editing efficiency of deaminases L8 and DddA on the Bx9 gene.

[0077] Figure 9 shows the sequence preference of deaminases L8 and DddA in the Bx9 gene. Embodiments of the present invention

[0078] I. Terms used in this application:

[0079] Examples of resources describing many of the molecular biology-related terms used in this article can be found in the following literature: Alberts et al., Molecular Biology of The Cell, 5th ed., Garland Science Publishing, Inc.: New York, 2007; Rieger et al., Glossary of Genetics: Classical and Molecular, 5th ed., Springer-Verlag: New York, 1991; King et al., A Dictionary of Genetics, 6th ed., Oxford University Press: New York, 2002; and Lewin, GenesIX, Oxford University Press: New York, 2007.

[0080] Any references cited in this article, including, for example, all patents, published patent applications and non-patent publications, are incorporated in their entirety by reference.

[0081] For ease of understanding this application, several terms and abbreviations used herein are defined as follows:

[0082] In this application, "identity" refers to the identity of an amino acid sequence or nucleotide sequence. The identity of an amino acid sequence (or nucleotide sequence) can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, by using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, and setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values), respectively, and performing an identity search on a pair of amino acid sequences, the identity value (%) can be obtained.

[0083] Specifically, the 70% or more similarity can be 75% or more similarity. Specifically, the 75% or more similarity can be 80% or more similarity. Specifically, the 80% or more similarity can be 85% or more similarity. Specifically, the 85% or more similarity can be 90% or more similarity. Specifically, the 90% or more similarity can be 91% or more similarity, 92% or more similarity, 93% or more similarity, 94% or more similarity, 95% or more similarity, 96% or more similarity, 97% or more similarity, 98% or more similarity, or 99% or more similarity. More specifically, the 70% or more identity can be at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity.

[0084] When used in a list of two or more items, the term "and / or" means that any of the listed items can be used alone or in combination with any one or more of the listed items. For example, the expression "A and / or B" is intended to mean either or both of A and B, i.e., A alone, B alone, or a combination of A and B. The expression "A, B and / or C" means A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B and C.

[0085] As used in this article, "plant" includes explants, plant parts, seedlings, plantlets, or whole plants at any stage of regeneration or development.

[0086] As used herein, "plant part" can refer to any organ or intact tissue of a plant, such as meristem, bud organs / structures (e.g., leaves, stems, or nodes), roots, flowers or floral organs / structures (e.g., flowers, bracts, sepals, petals, stamens, carpels, anthers, and ovules), seeds (e.g., embryo, endosperm, and seed coat), fruits (e.g., mature ovaries), propagules, or other plant tissues (e.g., vascular tissue, dermal tissue, ground tissue, etc.) or any part thereof. The plant part in this application can be viable, non-viable, renewable, and / or non-renewable. "Propagule" can include any plant part that can grow into a whole plant.

[0087] Plant cells are biological cells of plants, derived from plants or derived from cultures obtained by culturing cells taken from plants. As used herein, “transgenic plant cell” means any plant cell transformed with a stably integrated recombinant DNA molecule, construct, expression cassette, or sequence. Transgenic plant cells can include original transformed plant cells, transgenic plant cells regenerated or developed from R0 generation transgenic plant cells, transgenic plant cells cultured from another transgenic plant cell, or transgenic plant cells from any progeny or offspring of a transformed R0 generation plant, including cells of plant seeds or embryos, or cultured plant cells, callus cells, etc.

[0088] As is commonly understood in the art, the term "promoter" generally refers to a DNA containing an RNA polymerase binding site, a transcription start site, and / or a TATA box that assists or promotes the transcription of transcribed DNA. Promoters can be artificially synthesized, modified, or derived from known or naturally occurring promoters. Promoters can also include chimeric promoters comprising combinations of two or more heterologous sequences. Therefore, the promoters of this application may include variants of promoter sequences that are compositionally similar but not identical to other promoter sequences provided herein.

[0089] Promoters can be classified according to various criteria related to the expression patterns of the associated coding or transcribed sequences or genes (including transgenes) operably linked to them, such as constitutive, developmental, tissue-specific, and inducible promoters. A promoter that drives expression in all or most tissues of a plant is called a "constitutive" promoter. A promoter that drives expression at certain times or stages of development is called a "developmental" promoter. A promoter that drives enhanced expression in certain tissues of a plant relative to other tissues is called a "tissue-enhancing" or "tissue-preferred" promoter. Therefore, a "tissue-preferred" promoter elicits relatively high or preferential expression in a specific tissue of the plant, but lower expression levels in other tissues. A promoter that is expressed in a specific tissue of the plant but rarely or not expressed in other tissues is called a "tissue-specific" promoter. An "inducible" promoter is a promoter that initiates transcription in response to environmental stimuli (e.g., cold, drought, or light) or other stimuli (e.g., injury or chemical application). Promoters can also be classified according to their origin, such as heterologous, homologous, chimeric, synthetic, etc.

[0090] The term "transcribed DNA" refers to DNA that can be transcribed into RNA molecules.

[0091] The term "operationally ligated" can refer to a functional connection between a promoter and transcribed DNA, enabling the promoter to function and initiate transcription of the transcribed DNA. The term "operationally ligated" can also refer to a functional connection between other regulatory elements and a target gene to regulate the transcription and / or expression of the target gene.

[0092] As used herein, an "expression cassette" refers to a cassette containing at least transcribed DNA operatively linked to one or more regulatory elements, typically at least a promoter and a 3' UTR (such as a terminator).

[0093] As used herein, the term "vector" refers to any construct that can be used for transformation purposes, i.e., to introduce heterologous DNA into a host cell. Examples include plasmids, granules, viruses, bacteriophages, or linear or circular DNA.

[0094] In this application, "editing" or "genome editing" means using targeted genome editing technology to produce a targeted mutation, deletion, inversion, or substitution of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1000, at least 2500, at least 5000, or at least 10,000 nucleotides of endogenous plant genome nucleic acid sequence.

[0095] In this application, "editing" or "genome editing" may also cover the targeted insertion or site-specific integration of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides into the endogenous genome of a plant using targeted genome editing technology.

[0096] In this application, a “target site” for genome editing refers to a location within a plant genome of a polynucleotide sequence that is targeted and cleaved by a site-specific nuclease, thereby introducing a double-strand break (or single-strand nick) into the nucleic acid backbone and / or its complementary DNA strand. The site-specific nuclease may bind to the target site, for example, via a non-coding guide RNA (e.g., but not limited to CRISPR RNA (crRNA) or single-strand guide RNA (sgRNA)). The non-coding guide RNA provided herein may be complementary to the target site (e.g., complementary to the strand of a double-stranded nucleic acid molecule or the chromosome of the target site). A “target site” also refers to a location within the plant genome of a polynucleotide sequence that is bound and cleaved by another site-specific nuclease, which may not be guided by a non-coding RNA molecule, such as a broad-spectrum nuclease, zinc finger nuclease (ZFN), or transcription activator-like effector nuclease (TALEN), to introduce a double-strand break (or single-strand nick) into the polynucleotide sequence and / or its complementary DNA strand.

[0097] In this application, the terms “guide RNA,” “gRNA,” or “sgRNA” are short RNA sequences comprising (1) a structural or scaffold RNA sequence required to bind to or interact with RNA-guided nucleases and / or other RNA molecules (e.g., tracrRNA), and (2) an RNA sequence that is identical to or complementary to a target sequence or site (referred to herein as a “guide sequence”). A “single-stranded guide RNA” (or “sgRNA”) is an RNA molecule comprising tracrRNA and crRNA covalently linked by a linker sequence, which may be expressed as a single RNA transcript or molecule. Guide RNA contains a guide or target sequence (“guide sequence”) that is identical to or complementary to a target site within the plant genome, for example at or near the GA oxidase gene. An interstitial sequence adjacent motif (PAM) may be present immediately adjacent to the 5' end of a genomic target sequence complementary to the target sequence of the guide RNA and upstream of it in the genome, i.e., downstream (3') of the sense (+) strand immediately adjacent to the genomic target (relative to the target sequence of the guide RNA), as is known in the art. The genomic PAM sequence (relative to the target sequence of the guide RNA) on the sense (+) strand adjacent to the target site may contain 5'-NGG-3'. However, the corresponding sequence of the guide RNA (i.e., immediately downstream (3') of the target sequence of the guide RNA) is typically not complementary to the genomic PAM sequence. The guide RNA can typically be a non-coding RNA molecule that does not encode a protein.

[0098] II. Technical Solution Provided in this Application

[0099] The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.

[0100] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0101] Unless otherwise specified, the quantitative experiments in the following examples were performed in triplicate, and the results were averaged.

[0102] The following examples use GraphPad Prism statistical software to process the data. The experimental results are expressed as mean ± standard deviation. A two-way ANOVA method was used. * represents P < 0.05, ** represents P < 0.01, *** represents P < 0.001, and **** represents P < 0.0001.

[0103] Example 1: Novel Cytidine Deaminase Discovery Process

[0104] 1.1. All sequenced metagenomic assembly nucleic acid sequence information were retrieved and downloaded from the JGI and NCBI biological databases; all proteins were obtained by annotating protein sequences using Prodigal software; and domain annotations were performed on the predicted coding genes using the hidden Markov model in the Pfam database.

[0105] 1.2 Because DddA belongs to the SCP1.201 family, its full length is >1400aa, and the part that actually performs deamination is the latter half of the protein—DddAtox, whose annotation entry is SCP1.201-deam. Therefore, this entry was used as a keyword for screening. All the selected candidate proteins were aligned using Maft software. Based on prior knowledge, the active site for DddAtox is glutamate E at position 1347. Therefore, proteins with the same site were screened. Since DddA participates in the type VI secretion system (T6SS) of the intercellular protein delivery system, its downstream is the immune protein DddIA. The two interact to inhibit the toxicity of DddA protein and prevent cell inactivation. To facilitate the purification of DddA protein in subsequent experiments, we searched for the presence of a corresponding immune protein, DddIA, within a 10kb range upstream and downstream of the screened DddA homologous proteins. The lengths of the screened proteins varied, but only the DddAtox portion was required. Therefore, the sequences were aligned and truncated again to obtain 128 truncated sequences that enable the DddA homologous proteins to function.

[0106] 1.3. Using IQ-TREE software, a phylogenetic tree was constructed by combining all the mined proteins with other deaminase family protein sequences reported in the literature (Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems), and visualized using ITOL software. The conclusion is that the 128 novel deaminases mined through the established bioinformatics pipeline were classified into the SCP1.201 clade; therefore, these proteins do indeed belong to the SCP1.201 family.

[0107] 1.4 To further screen for novel active deaminases, AlphaFold 2 software was used to predict the structure of all discovered proteins, and their structures were compared pairwise with known DddA proteins using TM-align. Finally, 27 proteins were selected as novel deaminases with structures similar to known proteins for subsequent experimental verification and analysis. These 27 proteins were numbered as deaminases L8, L64, L65, L67, L70, L71, L72, L74, L78, L79, L80, L81, L83, L84, L85, L86, L88, L89, L90, L91, L115, L135, L136, L138, S45, S57, and S76.

[0108] 1.5. The mined deaminase sequence was optimized using maize codons, and a company was commissioned to synthesize the encoding genes for the deaminase and downstream immunoproteins. Taking deaminase L8 as an example, the encoding genes for deaminase L8 and the immunoprotein were ligated into the pET-duet-1 vector to construct a co-expression vector for the deaminase and downstream immunoprotein. This vector was then transformed into *E. coli* TSR2566 prokaryotic cells for induced expression, and the protein was purified. The encoding genes for deaminase L8 and the immunoprotein are DNA molecules with the nucleotide sequence SEQ ID NO. 4.

[0109] The co-expression vector for the deaminase and downstream immunoprotein was named pETduet1-L8. The structure of pETduet1-L8 is as follows: the segment between the BamHI and XhoI sites in the pET-duet-1 vector was replaced with a DNA molecule whose nucleotide sequence is SEQ ID NO.4, while keeping the other nucleotide sequences of the pET-duet-1 vector unchanged. Specifically, positions 2-412 of SEQ ID NO.4 contain the coding gene for deaminase L8, positions 478-496 contain the T7 promoter, and positions 564-944 contain the coding gene for the immunoprotein.

[0110] The pETduet1-L8 vector can express an L8 fusion protein (amino acid sequences shown in SEQ ID NO.8) with a vector sequence and a His tag, as well as an L8 immunoprotein (amino acid sequence shown in SEQ ID NO.9). The amino acid sequence of the L8 fusion protein consists of the start codon on the vector to the His-tagged amino acid sequence encoded by the sequence, the linker on positions 11-14, and the amino acid sequence of the L8 deaminase on positions 15-151.

[0111] The steps for inducing expression and purifying the protein were as follows: The constructed pETduet1-L8 vector was transformed into E. coli competent cells TSR2566 to obtain recombinant bacteria TSR2566 / pETduet1-L8. The bacteria were shaken in LB broth until OD600 = 0.6-0.7, and IPTG was added to induce protein expression. The bacterial culture was then collected. The bacterial culture was then purified using a nickel column to obtain a protein complex consisting of the L8 fusion protein and the L8 immunoglobulin. The protein complex was then denatured and renatured using urea solutions of different concentrations (8M, 6M, 4M, 2M, and 1M). The final eluent containing only the L8 deaminase was purified by molecular sieve to obtain the L8 fusion protein.

[0112] The preparation of the other 26 deaminases is the same as that of deaminase L8, the only difference being the coding sequence of the deaminase.

[0113] Using deaminase DddA as a control, the preparation method of deaminase DddA was the same as that of deaminase L8, except that the DNA molecule in SEQ ID NO.4 of the pETduet1-L8 vector was replaced with the DNA molecule shown in SEQ ID NO.5. Positions 2-415 of SEQ ID NO.5 contain the coding gene for deaminase DddA, positions 484-502 contain the T7 promoter, and positions 570-1001 contain the coding gene for the immunoprotein of deaminase DddA. The deaminase DddA fusion protein obtained by induced expression (amino acid sequence as shown in SEQ ID NO.10) was combined with the immunoprotein of DddA (amino acid sequence as shown in SEQ ID NO.11) to form a protein complex.

[0114] Example 2: In vitro activity verification

[0115] 2.1 In vitro deamination activity

[0116] DNA substrates labeled with 5'-FAM fluorophores were used to prepare a deamination reaction system with 20 nM concentration of deaminase. Deamination activity was analyzed. If the deaminase had deamination activity, the DNA substrate would be cleaved from the deamination cytosine. The reaction product of the deamination reaction system showed a cleavage band around 18 nt (Figure 1).

[0117] The deamination reaction system was prepared by mixing 50 μL of L8 protein (20 nM) with 10 μL of deamination buffer solution containing DNA substrate. The deamination buffer solution containing DNA substrate consisted of: 20 mM MES, 200 mM NaCl, 1 mM DTT, 80 g / L Ficoll 70, and 1 μM DNA substrate.

[0118] The deamination reaction conditions were as follows: the above deamination reaction mixture was deaminated at 37°C for 1 h; then 3 μL of UDG (Uracil-DNA Glycosylase) and 7 μL of 10×UDG buffer were added to the system, and the enzyme digestion reaction was carried out at 37°C for 30 min; 7 μL of 1M NaOH was added to the system, and the reaction was terminated by placing it at 95°C for 3 min. The deamination reaction sample obtained was used for subsequent electrophoresis verification.

[0119] Under the action of deaminase, cytosine (C) at position 18 of the cytosine deoxyribonucleotide in the double-stranded DNA molecule shown in Figure 1A undergoes a deamination reaction to become uracil (U) as shown in Figure 1B. Further addition of uracil-DNA glycosylase (UDG) to the reaction system causes the N-glycosidic bond between the uracil base and the sugar phosphate backbone to break, resulting in the uracil being removed from the nucleotide chain and forming a base-free site (C in Figure 1). Further addition of NaOH, under high temperature, causes the double-stranded DNA molecule to unwind into single strands. Phosphodiester bonds break at the base-free site in the upper strand, resulting in single-stranded DNA breaks, producing one 17nt single-stranded DNA and one 18nt single-stranded DNA.

[0120] In Figure 1, the first A at the 5' end of the two nucleotide sequences A to D is modified with 5'FAM (5-Carboxyfluorescein); A(14) represents 14 A's and T(14) represents 14 T's.

[0121] The nucleotide sequence of the upper chain of A in Figure 1 is as follows:

[0122] 5'-AAAAAAAAAAAAAAAACTCGCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 16).

[0123] In Figure 1, the cytosine at position 18 above B undergoes deamination to form uracil. The nucleotide sequence in the sequence listing is as follows, where T at position 18 of SEQ ID NO. 17 represents uracil:

[0124] 5'-AAAAAAAAAAAAAAAACTTGCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 17).

[0125] In Figure 1, the upper chain of C undergoes uracil shedding at the protocytosine position, becoming a baseless site. The nucleotide sequence of the upper chain of C in Figure 1 is as follows, where N represents a baseless site:

[0126] 5'-AAAAAAAAAAAAAAAACTNGCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 18).

[0127] In Figure 1, the upper chain of D breaks into two strands, and the nucleotide sequences are as follows:

[0128] 5'FAM-AAAAAAAAAAAAAAAACT-3'(SEQ ID NO.19)

[0129] 5'-GCCAAAAAAAAAAAAAAA-3' (SEQ ID NO. 20);

[0130] The nucleotide sequences of the lower chains A through D in Figure 1 are as follows:

[0131] 5'-TTTTTTTTTTTTTGGCGAGTTTTTTTTTTTTTTT-3' (SEQ ID NO. 21).

[0132] Electrophoresis validation: Before electrophoresis, the marker system consisted of 1 μL of 36 nt substrate, 1 μL of 18 nt substrate, 5 μL of 2×RNA loading buffer, and 13 μL of ddH2O. 77 μL of the deamination reaction sample was mixed with 10 μL of 2×RNA loading buffer and incubated at 100°C for 5 min for denaturation. During gel electrophoresis, the voltage was initially set to 100V for 30 min, then increased to 150V and run for 44 min.

[0133] Results Analysis: DNA fragments labeled with a 5'-FAM fluorophore showed bands when scanned using a gel scanning system (iBright FL1500 imaging system). DNA fragments without a 5'-FAM fluorophore label did not show bands when scanned using a gel scanning system. That is, if the deaminase has in vitro deamination activity, the 17nt fluorescently labeled DNA single strands in the deamination and denaturation product of the fluorescently labeled DNA substrate will show a band near 18nt when scanned using a gel scanning system; if the deaminase has no in vitro deamination activity (or very weak activity), the reaction product of the deamination reaction system will not show a band near 18nt when scanned using a gel scanning system.

[0134] The deamination activity of the 28 candidate proteins was verified, and 18 deaminases with deamination activity were successfully verified (L8, L70, L71, L72, L78, L79, L80, L91, L115, L138, S57, L83, L84, L85, L86, L89, L90, and S45). The verification was repeated, and based on the combined results of the two verifications, the five deaminases with the highest deamination activity—L8, L70, L85, L91, L90, and S45—were selected for further verification. The deamination activity verification results for deaminase L8 are shown in Figure 2.

[0135] 2.2 Validation of target sequence preference

[0136] The deamination reaction system was prepared according to 2.1, with the only difference from 2.2: the DNA substrate was different. In the DNA substrate, N represented one of the four different bases ATGC (Figure 3), which was used to determine the preference of the candidate deaminase for the target sequence C in vitro. The results showed that deaminase L8 could effectively deaminate cytosine in the NC (N is any base) background (Figure 4). Deaminase L8 exhibited better activity. Therefore, deaminase L8 was used for further experimental verification.

[0137] In Figure 3, the first A at the 5' end of the two nucleotide sequences is modified with 5'FAM (5-Carboxyfluorescein); A(15) represents 15 A's, T(15) represents 15 T's, and N is a, c, g, or t.

[0138] The nucleotide sequence of the upper chain in Figure 3 is as follows:

[0139] 5'-AAAAAAAAAAAAAAAGGNCGGAAAAAAAAAAAAAAA-3' (SEQ ID NO. 22);

[0140] The nucleotide sequence of the lower chain in Figure 3 is as follows:

[0141] 5'-TTTTTTTTTTTTTCCGNCCTTTTTTTTTTTTTTT-3' (SEQ ID NO. 23).

[0142] Example 3: Functional verification of deaminase L8 in eukaryotes

[0143] 3.1 Verification of the expression of deaminase L8 in eukaryotic cells

[0144] Because the discovered deaminase L8 is a double-stranded cytidine deaminase, it does not require DNA strand breaks during deamination. However, this presents a significant challenge: L8 is cytotoxic, potentially leading to vector mutations or even cell death during vector construction. To address this issue, a cat1 intron sequence was inserted between the L8 and DddA genes. This ensures that neither the L8 nor DddA genes are expressed during vector construction, allowing for successful vector creation.

[0145] To verify that the intron sequences inserted in L8 and DddA could be correctly cleaved in maize LH244 protoplasts, a vector expressing L8 and eGFP proteins with introns was constructed using a 35S promoter (Figure 5), named the L8 vector. The L8 vector was used to transform maize protoplasts, and the resulting positive maize protoplasts were denoted as L8. The protoplast transformation steps are described in the reference "Targeted large fragment deletion in plants using paired crRNAs with type I CRISPR system".

[0146] A vector expressing DddA and eGFP proteins with an intron via a 35S promoter was constructed and named the DddA vector. Maize protoplasts were transformed using the pUC-eGFP and DddA vectors as controls. Positive maize protoplasts obtained from the transformation were represented by pUC-eGFP and DddA, respectively.

[0147] The L8 vector structure is as follows: the fragment between the BamHI and XhoI restriction enzyme recognition sites of pUC-eGFP is replaced with a DNA molecule with the nucleotide sequence of SEQ ID NO. 6, while keeping the other nucleotide sequences of the backbone vector unchanged. Positions 4-21 of SEQ ID NO. 6 are the nucleotide sequence of the 6×His tag, positions 22-33 are the nucleotide sequence of the linker, positions 34-334 are the N-terminal coding gene of the L8 deaminase, positions 335-524 are the nucleotide sequence of the cat1 intron, and positions 525-634 are the C-terminal coding gene of the L8 deaminase. The nucleotide sequence of the pUC-eGFP vector is shown in SEQ ID NO. 12, where positions 1137-1142 are the BamHI restriction enzyme recognition site, positions 1161-1166 are the XhoI restriction enzyme recognition site, positions 744-1089 are the 35S promoter, and positions 1182-1901 are eGFP.

[0148] The only structural difference between the DddA vector and the L8 vector is that the DNA molecule in the L8 vector whose nucleotide sequence is from position 34 to 634 of SEQ ID NO. 6 (encoding the N-terminal gene of deaminase L8, the cat1 intron, and the C-terminal gene of deaminase L8) is replaced with a DNA molecule whose nucleotide sequence is from SEQ ID NO. 3. Specifically, positions 1-217 of SEQ ID NO. 3 represent the nucleotide sequence of the N-terminal gene encoding deaminase DddA, positions 218-407 represent the nucleotide sequence of the cat1 intron, and positions 408-604 represent the nucleotide sequence of the C-terminal gene encoding deaminase DddA.

[0149] If the intron sequence can be correctly cleaved, the GFP protein can be expressed and emit light under a fluorescence microscope. Conversely, if the intron sequence is not properly cleaved, GFP will be frameshifted and will not show green fluorescence under a fluorescence microscope. The results show that green fluorescence can be observed in maize protoplasts L8, DddA, and pUC-eGFP under a microscope (Figure 6), indicating that the introns in the L8 and DddA genes can be cleaved in maize protoplasts.

[0150] 3.2 Construction of single-base editing vectors

[0151] Subsequently, single-base editing vectors for the L8 and DddA deaminase systems were designed, using DddA as a control to verify whether L8 deaminase has deaminase activity in plants. The expression of sgRNA was initiated using the Osu3 promoter; the expression of fusion proteins of L8 deaminase, DddA, linker, nCas9 (D10A), and glycosylation inhibitor (UGI) was initiated using the Zmubi promoter. These proteins were then integrated into the pUC19 vector (NEB product, catalog number N3041S) via homologous recombination, resulting in recombinant vectors named pUC19-L8 and pUC19-DddA, respectively.

[0152] The nucleotide sequence of the pUC19-L8 vector is SEQ ID NO.1. pUC19-L8 can express a fusion protein with the amino acid sequence SEQ ID NO.2.

[0153] SEQ ID NO. 1, positions 57-1493, represent the nucleotide sequence of the sgRNA gene expression cassette. Specifically: positions 57-436 of SEQ ID NO. 1 represent the nucleotide sequence of the OsU3 promoter; positions 438-444 and 1120-1127 of SEQ ID NO. 1 represent the BsaI restriction enzyme recognition site; positions 1127-1202 of SEQ ID NO. 1 represent the sgRNA scaffold; and positions 1203-1493 of SEQ ID NO. 1 represent the nucleotide sequence of the OsU3 terminator.

[0154] The nucleotide sequence of the expression cassette of the fusion protein of deaminase L8, nCas9 (D10A) and glycosylation inhibitor protein (UGI) is located at positions 3508-9043 of SEQ ID NO.1. Specifically: SEQ ID NO. 1, positions 1506-3498, is the nucleotide sequence of the Zmubi promoter; positions 3514-3534 of SEQ ID NO. 1 is the nucleotide sequence of SV40 NLS; positions 3535-3558 are the nucleotide sequence of linker1; positions 3559-3576 are the nucleotide sequence of the 6×His tag; positions 3577-3588 are the nucleotide sequence of linker2; positions 3589-3889 are the nucleotide sequence of the N-terminal encoding gene of deaminase L8; positions 3890-4079 are the nucleotide sequence of the cat1 intron; positions 4080-4189 are the nucleotide sequence of the C-terminal encoding gene of deaminase L8; positions 4190-4240 are the nucleotide sequence of the A(EAAAK)3A linker; positions 4241-8341 are the nucleotide sequence of the nCas9 encoding gene; and positions 8342-8371 are 10aa. The nucleotide sequence of the linker is as follows: positions 8372-8620 are the nucleotide sequence of the UGI-encoding gene; positions 8621-8650 are the nucleotide sequence of the 10aa linker; positions 8651-8899 are the nucleotide sequence of the UGI-encoding gene; positions 8900-8911 are the nucleotide sequence of the 4aa linker; positions 8912-8992 are the nucleotide sequence of the 3×HA(81) linker; positions 8993-9040 are the nucleotide sequence of the NLS linker; and positions 9044-9678 are the nucleotide sequence of the RBCS E9T terminator.

[0155] In SEQ ID NO.2, positions 3-9 are the amino acid sequence of SV40 NLS, positions 10-17 are the amino acid sequence of linker1, positions 18-23 are the amino acid sequence of the 6×His tag, positions 24-27 are the amino acid sequence of linker2, positions 28-164 are the amino acid sequence of deaminase L8, positions 165-181 are the amino acid sequence of A(EAAAK)3A linker, positions 182-1548 are the amino acid sequence of nCas9, positions 1549-1558 are the amino acid sequence of 10aa linker, positions 1559-1641 are the amino acid sequence of UGI, positions 1642-1651 are the amino acid sequence of 10aa linker, positions 1652-1734 are the amino acid sequence of UGI, and positions 1735-1738 are the amino acid sequence of 4aa linker. The amino acid sequence of the linker is as follows: 1739-1765 is the amino acid sequence of 3×HA(81), and 1766-1781 is the amino acid sequence of NLS.

[0156] The structure of the pUC19-DddA vector is the same as that of the pUC19-L8 vector. The only difference is that the DNA molecule with nucleotide sequence 3589-4189 of SEQ ID NO.1 in the pUC19-L8 vector is replaced with the DNA molecule with nucleotide sequence SEQ ID NO.3. The pUC19-DddA vector expresses a base editor containing the DddA protein (a fusion protein containing the DddA protein), and the amino acid sequence of the base editor containing the DddA protein is shown in SEQ ID NO.13.

[0157] 3.3 To detect the single-base editing activity of L8, an NGG target (SEQ ID NO.7: CTGGCCCGCACCGTCACAG) was designed on the maize Bx9 (Zm00001eb033030) gene. Annealing primers were designed to synthesize DNA fragments with sticky ends for later use.

[0158] The nucleotide sequences of the annealing primers are as follows:

[0159] Bx9T-F: ATATATGGTCTCTGGCGCTGGCCCGCACCGTCACAGGTT (SEQ ID NO. 14);

[0160] Bx9T-R: ATTATTGGTCTCTAAACCTGGTGACGGTGCGGGCCAGC (SEQ ID NO. 15).

[0161] The pUC19-L8 vector was digested with BsaI to obtain the digestion product. The obtained DNA fragment and the digestion product were ligated with T4 enzyme to obtain the recombinant vector pUC19-L8-Bx9 targeting the Bx9 gene. The structure of the pUC19-L8-Bx9 vector is as follows: the fragment between the BsaI restriction recognition sites of the pUC19-L8 vector (the small fragment between the BsaI restriction recognition sites) is replaced by a DNA molecule with the nucleotide sequence of SEQ ID NO.7, while keeping the other nucleotide sequences of the pUC19-L8 vector unchanged. The pUC19-L8 vector can express the sgRNA with the target site of SEQ ID NO.7 and the fusion protein with the amino acid sequence of SEQ ID NO.2 (i.e., the cytosine base editor). The sgRNA and the cytosine base editor form a complex to perform single-base editing of the genomic sequence near the target site. After transfection of the pUC19-L8-Bx9 vector into a plasmid, it was transformed into maize protoplasts. After culturing for 48 hours, DNA was extracted from the maize protoplasts. Target sites were designed upstream and downstream of the target site, and a fragment of about 250 bp was amplified. A DNA library was constructed for next-generation sequencing.

[0162] A cytosine base editor constructed using the deaminase DddA was used as a control. The preparation method of the pUC19-DddA-Bx9 vector was the same as that of the pUC19-L8-Bx9 vector, except that the pUC19-L8 vector was replaced with the pUC19-DddA vector.

[0163] After analyzing the next-generation sequencing results, the proportions of C to T and G to A within the target site range were statistically analyzed. N in PAM NGG was defined as 0, with upstream values ​​recorded as negative and downstream values ​​as positive. The results showed that on the Bx9 gene, the editing window range of deaminases L8 and DddA was approximately 140 bp. The editing efficiency of deaminase L8 was significantly higher than that of deaminase DddA at most sites, especially at G-74, where L8 exhibited the highest editing efficiency of 19.25%, with an average of 18.2% (Figure 8).

[0164] In summary, it can be demonstrated that the mined L8 is active and exhibits higher editing activity than DddA at multiple sites of the Bx9 target gene.

[0165] Table 1. Statistical results of editing efficiency between L8 and DddA

[0166] Note: CK refers to protoplasts that have not undergone any treatment.

[0167] 3.4 To detect the sequence bias of L8 and DddA editing on the Bx9 gene, 3 bp upstream and downstream of the effective editing site were taken, and the sequences were aligned and the base distribution was statistically analyzed using Weblogo software. The results showed that, as previously reported, DddA exhibited TC sequence bias on the Bx9 gene. L8, on the other hand, could deaminate HC (H being A, C, or T) (Figure 9).

[0168] Table 2. Some sequences in this application

[0169] The present application has been described in detail above. Those skilled in the art will recognize that the present application can be implemented in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments are given in this application, it should be understood that further modifications can be made to the present application. In summary, in accordance with the principles of this application, this application is intended to include any changes, uses, or improvements to the present application, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Industrial applicability

[0170] This application utilizes metagenomic big data to conduct a comprehensive and systematic bioinformatics mining analysis of novel cytidine deaminases. A protein, L8, with a structure similar to the known DddA protein but low sequence similarity (identity < 63.04%), was selected for functional verification, and a novel CRISPR / Cas gene editing toolkit was developed using it. The cytosine base editing technology mediated by the fusion of the deaminase L8 with nCas9 (D10A) in this application exhibits significantly higher editing efficiency than DddA-mediated cytosine base editing technology, and can extend its editing activity window, improving the efficiency of cytosine base editing technology and broadening its application scope. This provides researchers in plant research and crop genetic improvement with an important tool for gene function research and correction. The single-base editor in this application can improve the efficiency of cytosine base editing and accurately mediate target base mutations, and can be widely applied in maize and even other plant cells.

Claims

1. A fusion protein, characterized in that: The fusion protein is a protein containing a cytidine deaminase, a Cas protein, and a uracil glycosylase inhibitor, the cytidine deaminase is a protein L8, the protein L8 is any one of the following A1) to A3): A1) a protein having an amino acid sequence of SEQ ID NO. 2 from 28 to 164; A2) a protein having 80% or more identity to the protein of A1) and related to deaminase, obtained by substitution, deletion, and / or addition of amino acid residues to the amino acid sequence of A1); A3) a fusion protein obtained by linking a tag to the N terminus and / or C terminus of A1) or A2).

2. The fusion protein of claim 1, wherein: The Cas protein is nCas9.

3. The fusion protein of claim 2, wherein: The nCas9 is a protein having an amino acid sequence of SEQ ID NO. 2 from 182 to 1548.

4. The fusion protein of claim 1, wherein: The uracil glycosylase inhibitor is a protein having an amino acid sequence of SEQ ID NO. 2 from 1559 to 1641 and / or SEQ ID NO. 2 from 1652 to 1734.

5. The fusion protein of claim 1, wherein: The fusion protein is a protein obtained by linking the cytidine deaminase, the Cas protein, the uracil glycosylase inhibitor, and a nuclear localization signal.

6. The fusion protein of claim 5, wherein: The nuclear localization signal can have an amino acid sequence of SEQ ID NO. 2 from 3 to 9 and / or SEQ ID NO. 2 from 1766 to 1781.

7. The fusion protein according to any one of claims 1 to 6, characterized in that: The fusion protein is a protein having an amino acid sequence of SEQ ID NO.

2.

8. A biological material related to the fusion protein of any one of claims 1 to 7 is at least one of the following D1) to D7): D1) a nucleic acid molecule encoding the fusion protein; D2) an expression cassette containing the nucleic acid molecule of D1); D3) a recombinant vector containing the nucleic acid molecule of D1) or an expression cassette of D2); D4) a recombinant microorganism containing the nucleic acid molecule of D1), an expression cassette of D2), or a recombinant vector of D3); D5) a transgenic plant cell line containing the nucleic acid molecule of D1), an expression cassette of D2), or a recombinant vector of D3); D6) a transgenic plant tissue containing the nucleic acid molecule of D1), an expression cassette of D2), or a recombinant vector of D3); D7) a transgenic plant organ containing the nucleic acid molecule of D1), an expression cassette of D2), or a recombinant vector of D3).

9. The biomaterial of claim 8, wherein: The nucleic acid molecule of D1) is at least one of the following: D11) a cDNA molecule or a DNA molecule having a nucleotide sequence of SEQ ID NO. 1 from 3589 to 8899 on the coding strand; D12) a cDNA molecule or a DNA molecule having a nucleotide sequence of SEQ ID NO. 1 from 3514 to 9040 on the coding strand; D13) a cDNA molecule or a DNA molecule having 80% or more identity to the nucleotide sequence defined in D11) and / or D12), and encoding the fusion protein.

10. A protein characterized by: The protein is the protein L8 as defined in claim 1.

11. A biological material related to the protein of claim 10, the biological material being at least one of: B1) a nucleic acid molecule encoding the protein; B2) an expression cassette comprising the nucleic acid molecule of B1); B3) a recombinant vector comprising the nucleic acid molecule of B1) or the expression cassette of B2); B4) a recombinant microorganism comprising the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3); B5) a transgenic plant cell line comprising the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3); B6) a transgenic plant tissue comprising the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3); B7) a transgenic plant organ comprising the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3).

12. The biomaterial of claim 11, wherein: B1) the nucleic acid molecule is at least one of: B11) a cDNA molecule or a DNA molecule whose coding sequence of the coding strand is represented by positions 2-412 of SEQ ID NO. 4; B12) a cDNA molecule or a DNA molecule whose nucleotide sequence of the coding strand is represented by positions 3589-4189 of SEQ ID NO. 1; B13) a cDNA molecule or a DNA molecule having 80% or more identity to the nucleotide sequence defined in B11) and / or B11), and encoding the protein.

13. Use, characterized in that, The application is at least one of: C1) use of the fusion protein of any one of claims 1-7 in single base editing and / or in preparing a product of single base editing; C2) use of the biological material of claim 8 or 9 in single base editing and / or in preparing a product of single base editing; C3) use of the protein of claim 10 in single base editing and / or in preparing a product of single base editing; C4) use of the biological material of claim 11 or 12 in single base editing and / or in preparing a product of single base editing.

14. Use according to claim 13, characterized in that, The single base editing is replacing cytosine C with uracil.

15. A method of mutating a C:G to a T:A on a plant genome, characterized by: The method comprises the following steps: introducing a DNA molecule expressing a cytosine base editor and an sgRNA into a recipient plant to obtain a target plant in which C:G is mutated to T:A; the target sequence of the sgRNA is 5'-N19-20PAM-3', wherein N19-20 is 19-20 N, and N is A, G, C or T; and the cytosine base editor is the fusion protein of claim 1.

16. The method of claim 15, wherein: The Cas protein in the cytosine base editor is nCas9, and the PAM is NGG.

17. The method according to claim 15 or 16, characterized in that, The plant is a dicot or a monocot.

18. The method of claim 17, wherein: The monocot is corn.

Citation Information

Patent Citations

  • Bipartite base editor (BBE) architectures and type-ii-c-cas9 zinc finger editing

    CN110997728A

  • Mutation method of rat mitochondrial gene G14098A, and application of mutation method

    CN113699160A

  • Cytosine deaminases and their use in base editing

    CN116555237A

  • Base editing tool and application thereof

    CN117327679A

  • Plant chloroplast and mitochondrial monomer cytosine base editor and application method thereof

    CN117534772A

Cited By

  • Cytosine deaminase derived from metagenome mining as well as biological material and application of cytosine deaminase

    CN121555485A