Plant basic group editing system mediated by Cas12i fusion protein

By constructing fusion proteins A and B, and combining them with DNA-binding proteins and the transcriptional activation domain VP64, the Cas12i-mediated base editing system was optimized, solving the problem of low Cas12i editing efficiency and achieving efficient C-to-T and A-to-G editing, thus promoting gene editing and genetic improvement in crops.

CN121895461APending Publication Date: 2026-04-21INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
Filing Date
2025-12-10
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The low efficiency of Cas12i-mediated base editing in soybeans limits its widespread application in plants.

Method used

Fusion proteins A and B were constructed, containing cytosine deaminases APOBEC3A and dCas12i3-5M, and adenine deaminases TadA8e and dCas12i3-5M, respectively. They were then combined with DNA single-strand binding protein Rad51 or DNA double-strand binding protein HMG-D and transcription activation domain VP64 to optimize editing efficiency and expand the editing range.

Benefits of technology

It significantly improved the editing efficiency of C to T and A to G in rice, reaching up to 32.35% and 38.24% respectively, and created new herbicide-resistant germplasm, providing technical support for crop gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a Cas12i fusion protein mediated plant basic group editing system, and belongs to the technical field of biology. The technical problem to be solved by the invention is how to improve the base editing efficiency of Cas12i. The Cas12i fusion protein disclosed by the invention is a fusion protein A or a fusion protein B, the fusion protein A contains APOBEC3A as shown in the 9th site to the 207th site of SEQ ID No.2, dCas12i3-5M as shown in the 240th site to the 1287th site of SEQ ID No.2 and UGI as shown in the 1298th site to the 1380th site of SEQ ID No.2, and the fusion protein B contains TadA8e and dCas12i3-5M as shown in the 9th site to the 174th site of SEQ ID No.4. According to the plant basic group editing system, the editing efficiency of Cas12i is optimized and improved, and the editing range of Cas12i is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to a Cas12i fusion protein-mediated plant base editing system. Background Technology

[0002] Base editing is a novel target gene modification technology based on the CRISPR system. It precisely edits single bases at target sites without inducing DNA double-strand breaks using cytosine deaminase or artificially evolved adenine deaminase, thereby achieving single nucleotide alterations to produce gain-of-function mutations. This can accelerate de novo domestication or directed evolution of crops. The earliest work on base editing came from David Liu's laboratory, which used a fusion of an inactive CRISPR / Cas nuclease (lacking the ability to induce DNA double-strand breaks) with cytosine deaminase and uracil glycosylation inhibitor (UGI) to achieve C-to-T editing (Komor et al. 2016). The cytosine base editor (CBE), which fuses cytosine deaminase and UGI, can catalyze the conversion of C to T bases. The CBE system catalyzes the deamination of the base C to U, and U is recognized as T during DNA replication. After DNA repair and replication, the C-to-T conversion is achieved (Komor et al. 2016). CBE is usually fused with UGI, which prevents intracellular uracil glycosylation of the base U, thus preventing base excision repair and preventing UG from reversing to CG, thereby improving the efficiency of C-to-T conversion. The adenine base editor (ABE) fused with adenine deaminase can achieve A-to-G base conversion. The ABE system deaminizes A to I under the action of adenine deaminase, and after DNA repair and replication, the A-to-G conversion is achieved. In practical applications, two adenine deaminases (TadA) are often used in tandem as a dimer (Gaudelli et al. 2017). Building upon this, a novel base editing system, AKBE, capable of A-to-K base editing, was constructed by linking OsMPG or hMPG to the C-terminus of nCas9 (Li et al. 2023). To expand the editing range of single-base editing systems, researchers used uracil glycosyltransferase (UNG) instead of UGI. UNG can remove the U base produced by deaminase, forming a purine-free and pyrimidine-free site, initiating the DNA repair process, and achieving C-to-G transversion. This novel base editor was named the CGBE system (C to G base editor) (Zhao et al. 2021; Tian et al. 2022).

[0003] To improve the efficiency of single-base editing systems, David Liu's lab mutated eight amino acid sites in TadA (TadA8e), resulting in a deamination rate in mammalian cells that was nearly a thousand times faster than the previous wild-type deaminase TadA, significantly improving base editing efficiency (Richter et al. 2020). Subsequent application of Tad8e to rice showed a significant increase in editing efficiency, with some sites achieving up to 100% efficiency (Wei et al. 2021). Furthermore, to expand the application of base editors in plants, Qi Yiping's team successfully developed a highly efficient, broadly targeted plant cytosine single-base editor, hA3A-Y130F-dCas12a, mediated by dCas12a, demonstrating high editing efficiency at the TTTV-PAM target site in rice (Cheng et al. 2023). By utilizing the fusion of deaminase-nCas9 or deaminase-dCas12a, single-base editing systems have been successfully applied to a variety of important crops, such as rice, wheat, corn, soybeans, and cotton.

[0004] The transcriptional activation domain VP64 plays a crucial role in gene editing. Its unique molecular structure and functional properties make it a core tool for precisely regulating gene expression and optimizing editing efficiency. VP64 originates from the VP16 protein of herpes simplex virus (HSV), a viral transcriptional activator. Researchers extracted its core functional domain—the minimal transcriptional activation module (approximately 12 amino acids)—and tandemly replicated it in quadruple replication to form the artificial chimeric domain VP64 (4×VP16). This design significantly enhances transcriptional activation capacity, increasing activation efficiency by tens of times compared to a single VP16. VP64's low molecular weight facilitates fusion with other functional proteins (such as Cas9 and base editors) without affecting its targeting ability (Li et al., 2023). The single-stranded binding protein Rad51 is a core recombinase in DNA repair and a key factor in improving the efficiency of gene editing, especially precise editing relying on homologous recombination repair. Rad51 can indirectly improve the efficiency and fidelity of gene editing by stabilizing the transient single-stranded DNA structure generated during gene editing or protecting gRNA / pegRNA templates from degradation (Zheng et al., 2024). Furthermore, HMG-D (High Mobility Group-D) proteins are a class of high-mobility group (HMG) proteins known for their ability to bind to double-stranded DNA non-sequence-specifically and induce DNA bending (Xue et al. 2024). In the field of gene editing, DNA-binding proteins such as HMG-D and Rad51 can optimize editing effects by enhancing the binding efficiency of the editing complex to the target site or regulating DNA repair pathways.

[0005] Cas12i3 is an emerging type VI nuclease in the CRISPR-Cas system, characterized by its small size and recognition of PAMs different from Cas9 (Cas12i3: TTN; Cas9: NGG), greatly expanding the scope of gene editing. It has already attracted widespread attention in the fields of gene editing in plants and animals. Zhu Jiankang's team predicted the structure of Cas12i3 using AlphaFold2, screened five mutation sites with significantly enhanced activity, and combined them into a five-mutant variant (S7R / D233R / D267R / N369R / S433R), named Cas-SF01. This variant showed a four-fold increase in editing efficiency, exceeding 80% in mammalian cells, superior to SpCas9 and Cas12a (Duan et al. 2024). Xia Lanqin's team fused Cas12i3-5M (a Cas12i3 pentamutant) with four exonucleases (T5E, UL12, PapE, ME15) and combined it with an optimized crRNA expression strategy to construct a highly efficient Cas12i3-5M-mediated gene editing system in wheat and rice (Wang et al. 2025). Wheat gene editing was performed using two different crRNA expression strategies. When using the optimized expression strategy (Opt) with a composite promoter to initiate crRNA expression and selecting bar as the selectable marker gene, the fusion of T5E and Cas12i3-5M (Opt-T5E-Cas12i3-5M) significantly improved editing efficiency, achieving an average editing efficiency of up to 88.99% at four endogenous target sites in stable lines of three superior Chinese wheat varieties (Zhengmai 7698, Zhengmai 1860, and Zhengshi 9179). Furthermore, Opt-T5E-Cas12i3-5M not only enables highly efficient gene editing in wheat but also induces a higher proportion of larger deletions. The developed Opt-T5E-Cas12i3-5M system not only expands the scope of gene editing and enriches the wheat gene editing toolkit but will also promote its application in gene editing of wheat and other important polyploid crops in agriculture, as well as for biological research and crop genetic improvement. Further, using three superior varieties as recipient materials, three new high-resistant starch germplasm were obtained (Wang et al. 2025a). In addition, the study found that fusing Cas12i3-5M with four exonucleases (T5E, UL12, PapE, ME15) all improved gene editing efficiency, with the UL12-Cas12i3-5M fusion showing the highest efficiency, significantly higher than the Cas12i3-5M variant.Furthermore, the fusion of four different exonucleases with the Cas12i3-5M variant enhanced multi-gene editing efficiency when simultaneously targeting three, four, five, and six genes. The UL12:Cas12i3-5M variant exhibited the highest simultaneous multi-gene editing efficiency, reaching 82.76%, 61.36%, 52.94%, and 51.07%, respectively, and also demonstrated high cleavage activity. In addition, the fusions of the other three exonucleases with Cas12i3-5M induced a higher proportion of large fragment deletions (>100 bp). This study provides a novel and efficient tool for multi-gene editing in rice, offering important tools and technical support for rapidly combining multiple superior agronomic traits in rice using gene editing tools (Wang et al. 2025b). However, to date, the efficiency of Cas12i-mediated base editing remains low, especially in soybean where the base editing efficiency is only 2.16% (Niu et al., 2024), limiting the widespread application of Cas12i-mediated base editing systems. Summary of the Invention

[0006] The technical problem to be solved by this invention is how to improve the base editing efficiency of Cas12i.

[0007] To solve the above-mentioned technical problems, the present invention first provides a fusion protein, wherein the fusion protein is fusion protein A or fusion protein B, wherein fusion protein A contains cytosine deaminase APOBEC3A, dCas12i3-5M and uracil glycosylation inhibitor protein UGI, and fusion protein B contains adenine deaminase TadA8e and dCas12i3-5M.

[0008] Furthermore, the fusion protein A and the fusion protein B also contain the DNA single-strand binding protein Rad51 or the DNA double-strand binding protein HMG-D.

[0009] Furthermore, the fusion protein A and the fusion protein B also contain a transcriptional activation domain VP64.

[0010] Furthermore, the fusion protein A and the fusion protein B also contain nuclear localization signals and / or linker peptides.

[0011] Further, the fusion protein A is fusion protein A1, fusion protein A2, fusion protein A3, fusion protein A4, fusion protein A5, or fusion protein A6. Fusion protein A1 is obtained by sequentially linking a nuclear localization signal, the transcriptional activation domain VP64, the cytosine deaminase APOBEC3A, the DNA single-strand binding protein Rad51, dCas12i3-5M, and the uracil glycosylase inhibitor UGI. Fusion protein A2 is obtained by sequentially linking a nuclear localization signal, the transcriptional activation domain VP64, the cytosine deaminase APOBEC3A, the DNA double-strand binding protein HMG-D, dCas12i3-5M, and the uracil glycosylase inhibitor UGI. Fusion protein A3 is obtained by sequentially linking a nuclear localization signal, the transcriptional activation domain VP64, and the cytosine deaminase APOBEC3A, the DNA double-strand binding protein HMG-D, dCas12i3-5M, and the uracil glycosylase inhibitor UGI. The fusion protein A4 is obtained by sequentially linking the cytosine deaminase APOBEC3A, dCas12i3-5M, and the uracil glycosylase inhibitor UGI; the fusion protein A5 is obtained by sequentially linking the nuclear localization signal, the cytosine deaminase APOBEC3A, the DNA single-strand binding protein Rad51, dCas12i3-5M, and the uracil glycosylase inhibitor UGI; the fusion protein A6 is obtained by sequentially linking the nuclear localization signal, the cytosine deaminase APOBEC3A, dCas12i3-5M, and the uracil glycosylase inhibitor UGI. The fusion protein B is fusion protein B1, fusion protein B2, fusion protein B3, fusion protein B4, fusion protein B5, or fusion protein B6. Fusion protein B1 is obtained by sequentially linking a nuclear localization signal, the transcription activation domain VP64, the adenine deaminase TadA8e, the DNA single-strand binding protein Rad51, and dCas12i3-5M. Fusion protein B2 is obtained by sequentially linking a nuclear localization signal, the transcription activation domain VP64, the adenine deaminase TadA8e, the DNA double-strand binding protein HMG-D, and dCas12i3-5M. Fusion protein B3 is obtained by sequentially linking a nuclear localization signal, the transcription activation domain VP64, the adenine deaminase TadA8e, the DNA double-strand binding protein HMG-D, and dCas12i3-5M. The fusion protein B4 is obtained by sequentially linking the nuclear localization signal, the transcriptional activation domain VP64, the adenine deaminase TadA8e, and dCas12i3-5M; the fusion protein B5 is obtained by sequentially linking the nuclear localization signal, the adenine deaminase TadA8e, the DNA single-strand binding protein Rad51, and dCas12i3-5M; the fusion protein B6 is obtained by sequentially linking the nuclear localization signal, the adenine deaminase TadA8e, the DNA double-strand binding protein HMG-D, and dCas12i3-5M.

[0012] In the fusion protein A, there may be two uracil glycosylation inhibitor proteins (UGIs). The uracil glycosylation inhibitor proteins (UGIs) can be linked to each other, or to dCas12i3-5M, via a 10aa linker peptide or directly, as long as it does not affect the function of the fusion protein A.

[0013] The fragments in the fusion protein can be linked by linker peptides or directly, as long as it does not affect the function of the fusion protein.

[0014] The linker peptide is a 32aa linker peptide or a 10aa linker peptide. In one embodiment of the present invention, the 32aa linker peptide is positions 208-239 of SEQ ID No. 2. The 10aa linker peptide is positions 1288-1297 of SEQ ID No. 2.

[0015] The nuclear localization signal can be nuclear localization signal 1 and nuclear localization signal 2, which are located at the N-terminus and C-terminus of the fusion protein, respectively. Nuclear localization signal 1 is bits 2-8 of SEQ ID No. 2. Nuclear localization signal 2 is bits 1478-1495 of SEQ ID No. 2.

[0016] Furthermore, the cytosine deaminase APOBEC3A is as follows: A1) or A2): A1) The amino acid is the protein of positions 9-207 of SEQ ID No. 2; A2) The protein of positions 9-207 of SEQ ID No. 2, except for position 130 (position 130 of cytosine deaminase APOBEC3A, position 138 of SEQ ID No. 2), has one or more amino acid residues substituted and / or deleted and / or added, and has the same function; dCas12i3-5M is as follows (B1) or (B2): B1) A protein whose amino acid is from position 240 to 1287 of SEQ ID No. 2; B2) A protein that has the same function by substituting and / or deleting and / or adding one or more amino acid residues from position 240 to 1287 of SEQ ID No. 2. The uracil glycosylation inhibitor protein UGI is as follows (C1) or (C2): C1) The amino acid is the amino acid at positions 1298-1380 of SEQ ID No. 2; C2) The protein having the same function by substituting and / or deleting and / or adding one or more amino acid residues at positions 1298-1380 of SEQ ID No. 2. The DNA single-stranded binding protein Rad51 is either D1 or D2 as follows: D1) A protein whose amino acid is the first to last 114 of SEQ ID No. 8; D2) A protein that has the same function by substituting and / or deleting and / or adding one or more amino acid residues from the first to last 114 of SEQ ID No. 8. The DNA double-strand binding protein HMG-D is either E1 or E2). E1) A protein whose amino acid is at positions 1-112 of SEQ ID No. 10; E2) A protein that has the same function by substituting and / or deleting and / or adding one or more amino acid residues at positions 1-112 of SEQ ID No. 10. The transcriptional activation domain VP64 is as follows (F1) or (F2): F1) A protein whose amino acid is from position 1 to 50 of SEQ ID No. 6; F2) A protein that has the same function by substituting and / or deleting and / or adding one or more amino acid residues from position 1 to 50 of SEQ ID No. 6. The adenine deaminase TadA8e is either G1 or G2 as follows: G1) A protein whose amino acid is position 9-174 of SEQ ID No. 4; G2) A protein that has the same function by substituting and / or deleting and / or adding one or more amino acid residues at position 9-174 of SEQ ID No. 4.

[0017] In the above text, proteins obtained through "substitution and / or deletion and / or addition of one or more amino acid residues" are proteins with 75% or more identity with the corresponding protein and have the same function. Identity refers to the similarity of the amino acid sequence. The identity of amino acid sequences can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, and performing an identity search on a pair of amino acid sequences, the identity value (%) can then be obtained. The phrase "having 75% or more of the sameness" means having 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the sameness.

[0018] The proteins in A2), B2), C2), D2), E2), F2), and G2) mentioned above can be synthesized artificially, or their encoding genes can be synthesized first and then expressed biologically.

[0019] The present invention also provides biomaterials related to the fusion protein, said biomaterials being at least one of the following (H1)-H4): H1) A nucleic acid molecule encoding the fusion protein; H2) An expression cassette containing the nucleic acid molecule of H1); H3) A recombinant vector containing the nucleic acid molecule of H1 or a recombinant vector containing the expression cassette of H2); H4) A recombinant microorganism containing the nucleic acid molecule of H1, a recombinant microorganism containing the expression cassette of H2, or a recombinant microorganism containing the recombinant vector of H3).

[0020] The nucleic acid molecule can be DNA, such as cDNA, genomic DNA, or recombinant DNA; the nucleic acid molecule can also be RNA, such as mRNA or hnRNA.

[0021] Specifically, in H1), the nucleic acid molecule encoding the cytosine deaminase APOBEC3A is the DNA molecule shown in positions 2034-2630 of SEQ ID No. 1. The nucleic acid molecule encoding dCas12i3-5M is the DNA molecule shown in positions 2727-5870 of SEQ ID No. 1. The nucleic acid molecule encoding the uracil glycosylase inhibitor protein UGI is the DNA molecule shown in positions 5901-6149 of SEQ ID No. 1. The nucleic acid molecule encoding the DNA single-strand binding protein Rad51 is the DNA molecule shown in positions 1-342 of SEQ ID No. 7. The nucleic acid molecule encoding the DNA double-strand binding protein HMG-D is the DNA molecule shown in positions 1-336 of SEQ ID No. 9. The nucleic acid molecule encoding the transcription activation domain VP64 is the DNA molecule shown in positions 1-150 of SEQ ID No. 5. The nucleic acid molecule encoding the adenine deaminase TadA8e is the DNA molecule shown at positions 2034-2531 of SEQ ID No. 3.

[0022] In the aforementioned biological materials, the expression cassette refers to DNA capable of expressing the fusion protein in a host cell (such as a plant cell). This DNA may include not only a promoter to initiate transcription of the fusion protein-coding gene, but also a terminator to terminate transcription of the fusion protein-coding gene. Furthermore, the expression cassette may also include an enhancer sequence.

[0023] In one embodiment of the present invention, the promoter in the expression cassette described in H2) is a ubiquitous promoter.

[0024] In the aforementioned biological materials, the carrier may be a plasmid, a granule, a bacteriophage, or a viral vector.

[0025] In the aforementioned biological materials, the microorganisms may be yeast, bacteria, algae, or fungi. Among them, bacteria may be Agrobacterium.

[0026] The application of the fusion protein or the biomaterial in plant gene editing is also within the scope of protection of this invention.

[0027] The application of the fusion protein or the biomaterial in the preparation of plant gene editing products is also within the scope of protection of this invention.

[0028] The application of the fusion protein or the biomaterial in the preparation of herbicide-resistant plants is also within the scope of protection of this invention.

[0029] The plant can be a monocotyledonous or dicotyledonous plant. Further, the plant can be a grass (Poaceae). In one embodiment of the invention, the plant is rice.

[0030] In this invention, fusion protein A is used for editing from C to T. Fusion protein B is used for editing from A to G.

[0031] This invention constructs Cas12i-mediated CBE and ABE base editing systems by fusing a five-mutant of inactivated Cas12i3 (dCas12i3-5M) with either the cytosine deaminase human APOBEC3A (hA3A-Y130F) or the adenine deaminase TadA8e. Furthermore, the editing efficiency and scope are optimized and improved by fusing transcription activator VP64, DNA single-strand binding protein (Rad51), or DNA double-strand binding protein (HMG-D). This is applied to endogenous targets in rice. OsARF4, OsSBEIIb, OsALS, OsACC, OsDEP1 and OsBADH2 Test results showed that the editing efficiency of the novel base editing systems VP64-CBE-HMG-D-dCas12i3-5M and VP64-ABE-HMG-D-dCas12i3-5M can be improved by up to approximately 3 times. Specifically, the highest editing efficiency of C-to-T mediated by VP64-CBE-HMG-D-dCas12i3-5M reached 32.35%, and the highest efficiency of A-to-G mediated by VP64-ABE-HMG-D-dCas12i3-5M reached 38.24%, leading to the creation of new herbicide-resistant rice germplasm. The construction, optimization, and application of novel Cas12i fusion protein-mediated CBE and ABE base editors will provide important technical support for biological research and genetic improvement of important crop genes. Attached Figure Description

[0032] Figure 1 Schematic diagram of Cas12i-mediated base editing system. (A) Schematic diagram of Cas12i-mediated cytosine base editing system (CBE); (B) Schematic diagram of Cas12i-mediated adenine base editing system (ABE).

[0033] Figure 2 Editing efficiency and editing window of the Cas12i-mediated base editing system. (A) Editing efficiency and editing window of the Cas12i-mediated cytosine base editing system in rice protoplasts; (B) Editing efficiency and editing window of the Cas12i-mediated adenine base editing system in rice protoplasts.

[0034] Figure 3 Heatmaps showing the distribution of editing efficiency in the Cas12i-mediated base editing system. (A) Heatmap of the C to T editing efficiency distribution in the Cas12i-mediated cytosine base editing system; (B) Heatmap of the A to G editing efficiency distribution in the Cas12i-mediated adenine base editing system. The darker the color, the higher the editing efficiency.

[0035] Figure 4 Novel herbicide-resistant rice germplasm obtained using the Cas12i-mediated base editing system. (A) Resistance of the OsALS-R190H allele to bispyribac-sodium (BS) obtained using the Cas12i-mediated cytosine base editing system. The OsALS-R190H mutant survived for two weeks after field treatment with 1x and 2x BS concentrations, while wild-type (WT) seedlings died after field treatment with 1x BS concentration and after field treatment with 4x BS concentration. (B) Resistance of the OsACC-W2097R allele to haloxyfop obtained using the Cas12i-mediated adenine base editing system. The mutant survived for two weeks after field treatment with a 1x concentration of haloxyfop-R-methyl, while wild-type seedlings died after field treatment with a 1x concentration of haloxyfop-R-methyl, and the OsACC-W2097R mutant died after field treatment with 2x and 4x concentrations of haloxyfop-R-methyl. The PAM sequence and target bases in the editing window are highlighted in blue and red, respectively, and mutation sites are marked with red arrows. Scale bar = 7 cm. Detailed Implementation

[0036] In this document, unless otherwise defined herein, terms should be understood according to their common usage by those skilled in the art. Examples of resources describing many of the molecular biology-related terms used herein can be found in the following references: Alberts et al., Molecular Biology of The Cell, 5th ed., Garland Science Publishing, Inc.: New York, 2007; Rieger et al., Glossary of Genetics: Classical and Molecular, 5th ed., Springer-Verlag: New York, 1991; King et al., A Dictionary of Genetics, 6th ed., Oxford University Press: New York, 2002; and Lewin, GenesIX, Oxford University Press: New York, 2007.

[0037] Unless otherwise stated, the terms “nucleic acid,” “nucleotide,” and “polynucleotide” cover both DNA and RNA.

[0038] The term "vector" refers to a nucleic acid molecule that allows the insertion of foreign nucleic acids without disrupting its ability to replicate and / or integrate into the host cell. A vector may include a nucleic acid sequence that allows it to replicate in the host cell, such as an origin of replication. A vector may also include one or more optional marker genes and other genetic elements. An integration vector is capable of integrating itself into the host nucleic acid. An expression vector is a vector that contains the necessary regulatory sequences to allow the transcription and translation of the inserted gene. Vectors can be plasmid vectors or viral vectors. The term "plasmid" refers to a structure composed of genetic material designed to guide the transformation of target cells. A plasmid includes a plasmid backbone. The plasmid backbone contains multiple genetic elements that are localized and sequentially oriented with other essential genetic elements, enabling the nucleic acids in the nucleic acid cassette to be transcribed in the transfected cell and translated where necessary. The plasmid backbone may contain one or more unique nucleic acid restriction sites. A plasmid is capable of autonomous replication in a host or organism, thereby replicating a cloned sequence. A plasmid can confer certain phenotypes on the host organism that are selective or easily detectable. A plasmid or plasmid backbone may have a linear or circular conformation. The constituent elements of a plasmid may include, but are not limited to, the following DNA molecules: (1) DNA; (2) the plasmid backbone; (3) the sequence encoding the target gene; and (4) regulatory elements responsible for transcription, translation, RNA stability, and replication.

[0039] The term "nucleic acid restriction site" or "restriction site" refers to a deoxyribonucleic acid sequence at which a specific restriction endonuclease cleaves the molecule.

[0040] The term "transfection" refers to the process of introducing foreign DNA into cultured cells by exposing them to foreign DNA. Transfection methods include, but are not limited to, microinjection, electroporation, calcium phosphate precipitation, liposome fusion (e.g., liposome transfection), or gene gun. As an example, the transfection is performed using a Lip2000 transfection device. Transformation can occur through various mechanisms, such as transfection, electroporation, or particle bombardment.

[0041] The term "operably linked" refers to the operative connection of nucleic acid or amino acid sequences that are functionally related to each other. When used to refer to nucleic acids, it refers to the functional connection between a promoter or other regulatory element and a gene-associated transcribed DNA sequence or coding sequence (CDS), enabling the promoter or other regulatory element to function in initiating, assisting, and / or promoting the transcription and expression of the associated transcribed DNA sequence or coding sequence, at least in certain cells, tissues, developmental stages, and / or diseases. For example, operatively linked promoter, enhancer, coding sequence (CDS), open reading frame (ORF), 5' and 3' UTR, and terminator sequences result in the accurate production of nucleic acid molecules (e.g., RNA). In some instances, operatively linked nucleic acid elements lead to the transcription of open reading frames and ultimately the production of peptides (i.e., expression of open reading frames). In other instances, operatively linked peptides are peptides in which functional domains are placed at appropriate distances from each other to confer the intended function of each domain.

[0042] The term "transcribed DNA" refers to DNA that can be transcribed into RNA molecules.

[0043] The term "promoter" generally refers to a DNA molecule that contains an RNA polymerase binding site, a transcription start site, and / or a TATA box and assists or promotes transcription of transcribed DNA. Promoters can be synthesized artificially or derived from known or naturally occurring promoters. Promoters can also include chimeric promoters comprising combinations of two or more heterologous sequences. Promoters can be constitutively active promoters (i.e., promoters that are continuously active / "on"), inducible promoters (i.e., promoters whose state (active / "on" or inactive / "off") is controlled by external stimuli (e.g., the presence of a specific temperature, compound, or protein), spatially restricted promoters (e.g., tissue-specific promoters, cell-type-specific promoters, etc.), or time-restricted promoters (i.e., promoters that are "on" or "off" at a specific stage of embryonic development or a specific stage of a biological process). Examples of inducible promoters include, but are not limited to, the T7 RNA polymerase promoter, the T3 RNA polymerase promoter, isopropyl-β-D-thiogalactoside (IPTG)-regulated promoters, lactose-induced promoters, heat shock promoters, tetracycline-regulated promoters, steroid-regulated promoters, metal-regulated promoters, and estrogen receptor-regulated promoters. Inducers of inducible promoters include, but are not limited to, molecular regulators such as doxycycline, RNA polymerases (e.g., T7 RNA polymerase), estrogen receptors, and estrogen receptor fusion proteins.

[0044] The term "transcription terminator sequence" (also known as "transcription terminator element," "transcription terminator," "terminator," or "terminator sequence") refers to a nucleotide sequence that terminates transcription. In this invention, the sequence is located within the 3' adapter region of the polyclonal sequence, but it can also be located at other sites on the plasmid. In one embodiment, the terminator is derived from the *E. coli* rrnB operon. These sequences ensure that transcription of the target nucleic acid sequence does not read into other functional regions of the plasmid.

[0045] The term "transcription" or "transcriptionalization" refers to the process by which RNA molecules are formed on a DNA template through complementary base pairing. This process is mediated by RNA polymerase.

[0046] The term "import" refers to the operation of transferring the gene or a recombinant vector containing the gene into a cell so that the gene can be expressed in the cell.

[0047] The expression can be transient, continuous, or stable.

[0048] The term "transient expression" refers to the introduction of genetic material that does not integrate into the host cell genome or is not replicated, and therefore can be degraded or transferred to other compartments over a period of time.

[0049] The term "sustained expression" refers to the introduction of a target gene along with genetic elements into a cell that enable the genetic material to replicate and / or be maintained outside the cell (i.e., extrachromosomally). This can lead to a remarkably stable transformation of the cell without integrating the new genetic material into the host cell's chromosome.

[0050] The term "stable expression" refers to the introduction of genetic material into the chromosome of a target cell, where it integrates and becomes a permanent component of the cell's genetic material. Stable gene expression after introduction can permanently alter the characteristics of the cell and its offspring produced through replication, leading to stable transformation.

[0051] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0052] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials, reagents, instruments, etc., used in the following examples are commercially available.

[0053] In the following examples, unless otherwise specified, the first position of each nucleotide sequence in the sequence listing is the 5′ terminal nucleotide of the corresponding DNA / RNA, and the last position is the 3′ terminal nucleotide of the corresponding DNA / RNA.

[0054] Example 1 1. Materials and Methods 1.1 Experimental Materials The rice material used for rice conversion was Zhonghua 11, provided by the Institute of Crop Science, Chinese Academy of Agricultural Sciences.

[0055] The basic vector backbone pCXUN-Ubi-dCas12i3-5M was preserved by the Gene Editing New Technology and New Material Creation Laboratory of the Institute of Crop Science, Chinese Academy of Agricultural Sciences. dCas12i3-5M is a mutant protein obtained by mutating amino acid E at position 844 to A based on Cas12i3-5M. The basic cytosine base editing vector pCXUN-Ubi-mhA3A-nCas9-2×UGI was preserved by the Gene Editing New Technology and New Material Creation Laboratory of the Institute of Crop Science, Chinese Academy of Agricultural Sciences. The basic adenine base editing vector pCXUN-Ubi-Tad8e-nCas9 was preserved by the Gene Editing New Technology and New Material Creation Laboratory of the Institute of Crop Science, Chinese Academy of Agricultural Sciences. The VP64, RAD51, and HMG-D gene sequences were synthesized by Shenzhen BGI Genomics Co., Ltd.

[0056] 1.2 Experimental Methods 1.2.1 Construction of the base editing system The cytosine deaminase human APOBEC3A (hA3A-Y130F, hereinafter referred to as mhA3A) was fused to the N-terminus of dCas12i3-5M, and UGI was fused to the C-terminus of dCas12i3-5M in the form of a dimer to construct a dCas12i3-5M-mediated cytosine base editing system (CBE-dCas12i3-5M). Figure 1 ).

[0057] The basic vector pCXUN-Ubi-mhA3A-dCas12i3-5M-2×UGI for expressing CBE-dCas12i3-5M was prepared. This vector is a double-stranded DNA consisting of 10051 bp. The nucleotide sequence of the first 7351st position of one strand is SEQ ID No. 1, and the nucleotide sequence of the last 7352nd to the last 10051st position is SEQ ID No. 11. That is, the nucleotide sequence (5′→3′) of the pCXUN-Ubi-mhA3A-dCas12i3-5M-2×UGI vector is obtained by connecting the last position of SEQ ID No. 1 with the first position of SEQ ID No. 11.

[0058] In SEQ ID No. 1, positions 1-1991 represent the Ubi promoter, positions 2013-2033 represent the gene encoding the NLS nuclear localization signal, positions 2034-2630 represent the gene encoding mhA3A, positions 2631-2726 represent the gene encoding the 32aa linker peptide, positions 2727-5870 represent the gene encoding dCas12i3-5M, and positions 5871-5900 represent the 10a... The gene encoding the α-linker peptide is shown in positions 5901-6149, which is the gene encoding UGI. Positions 6150-6179 are the gene encoding the 10aa linker peptide. Positions 6180-6428 are the gene encoding UGI. Positions 6441-6494 are the gene encoding the NLS nuclear localization signal. Positions 6502-6716 are PloyA. Positions 6717-7351 are the NOS terminator.

[0059] In SEQ ID No. 11, positions 40-420 represent the OsU3 promoter, positions 630-1307 represent the 35S promoter, and positions 1352-2377 represent the HptII gene.

[0060] The sequence of CBE-dCas12i3-5M is shown in SEQ ID No. 2. In SEQ ID No. 2, positions 2-8 represent the NLS nuclear localization signal, positions 9-207 represent mhA3A, positions 208-239 represent the 32aa linker peptide, positions 240-1287 represent dCas12i3-5M, positions 1288-1297 represent the 10aa linker peptide, positions 1298-1380 represent UGI, positions 1381-1390 represent the 10aa linker peptide, positions 1391-1473 represent UGI, and positions 1478-1495 represent the NLS nuclear localization signal.

[0061] Based on the basic vector pCXUN-Ubi-mhA3A-dCas12i3-5M-2×UGI, a DNA single-strand binding protein (Rad51) or a DNA double-strand binding protein (HMG-D) was fused to the C-terminus of mhA3A and the N-terminus of dCas12i3-5M (Rad51 or HMG-D is separated from both mhA3A and dCas12i3-5M by a 32-amino acid linker peptide (32aa linker peptide)), constructing a Cas12i-mediated cytosine base editing system (CBE-Rad51-dCas12i3-5M, CBE-HMG-D-dCas12i3-5M). Figure 1 The resulting vectors are pCXUN-Ubi-mhA3A-Rad51-dCas12i3-5M-2×UGI and pCXUN-Ubi-mhA3A-HMG-D-dCas12i3-5M-2×UGI.

[0062] The Rad51 and HMG-D involved in linking the 32aa peptide are shown in SEQ ID No. 8 and SEQ ID No. 10, respectively, and their encoding gene sequences are shown in SEQ ID No. 7 and SEQ ID No. 9, respectively. Positions 1-114 of SEQ ID No. 8 represent RAD51, and positions 1-112 of SEQ ID No. 10 represent HMG-D.

[0063] Based on the above vectors, VP64 was fused between mhA3A and its upstream nuclear localization signal NLS in CBE-dCas12i3-5M, CBE-Rad51-dCas12i3-5M, and CBE-HMG-D-dCas12i3-5M, respectively. The N-terminus of VP64 was directly connected to the C-terminus of NLS, and the C-terminus was connected to mhA3A via a 32aa linker peptide, thus constructing Cas12i-mediated cytosine base editing systems (VP64-CBE-dCas12i3-5M, VP64-CBE-Rad51-dCas12i3-5M, VP64-CBE-HMG-D-dCas12i3-5M). Figure 1 The resulting vectors are pCXUN-Ubi-VP64-mhA3A-dCas12i3-5M-2×UGI, pCXUN-Ubi-VP64-mhA3A-Rad51-dCas12i3-5M-2×UGI, and pCXUN-Ubi-VP64-mhA3A-HMG-D-dCas12i3-5M-2×UGI.

[0064] The VP64 linker peptide involved is shown in SEQ ID No. 6, and its encoding gene sequence is shown in SEQ ID No. 5. Positions 1-50 of SEQ ID No. 6 represent VP64.

[0065] Adenine deaminase TadA8e was fused to the N-terminus of dCas12i3-5M to construct a Cas12i-mediated adenine base editing system (ABE-dCas12i3-5M). Figure 1 ).

[0066] The basic vector pCXUN-Ubi-TadA8e-dCas12i3-5M for expressing ABE-dCas12i3-5M was prepared. This vector is a double-stranded DNA consisting of 9394 bp. The nucleotide sequence of positions 1-6694 of one strand is SEQ ID No. 3, and the nucleotide sequence of positions 6695-9394 is SEQ ID No. 11. That is, the nucleotide sequence (5′→3′) of the pCXUN-Ubi-TadA8e-dCas12i3-5M vector is obtained by connecting the last position of SEQ ID No. 3 with the first position of SEQ ID No. 11.

[0067] In SEQ ID No. 3, positions 1-1991 represent the Ubi promoter, positions 2013-2033 represent the gene encoding the NLS nuclear localization signal, positions 2034-2531 represent the gene encoding TadA8e, positions 2532-2627 represent the gene encoding the 32aa linker peptide, positions 2628-5771 represent the gene encoding dCas12i3-5M, positions 5784-5837 represent the gene encoding the NLS nuclear localization signal, positions 5845-6059 represent PloyA, and positions 6060-6694 represent the NOS terminator.

[0068] The sequence of CBE-dCas12i3-5M is shown in SEQ ID No. 4. In SEQ ID No. 4, positions 2-8 represent the NLS nuclear localization signal, positions 9-174 represent TadA8e, positions 175-206 represent the 32aa linker peptide, positions 207-1254 represent dCas12i3-5M, and positions 1259-1276 represent the NLS nuclear localization signal.

[0069] Based on the basic vector pCXUN-Ubi-TadA8e-dCas12i3-5M, a DNA single-strand binding protein (Rad51) or a DNA double-strand binding protein (HMG-D) was fused to the C-terminus of TadA8e and the N-terminus of dCas12i3-5M (Rad51 or HMG-D is separated from both TadA8e and dCas12i3-5M by 32 amino acids (32aa linker peptide)), constructing a Cas12i-mediated adenine base editing system (ABE-Rad51-dCas12i3-5M, ABE-HMG-D-dCas12i3-5M). Figure 1 The resulting vectors were pCXUN-Ubi-TadA8e-Rad51-dCas12i3-5M and pCXUN-Ubi-TadA8e-HMG-D-dCas12i3-5M.

[0070] The Rad51 and HMG-D involved in linking the 32aa peptide are shown in SEQ ID No. 8 and SEQ ID No. 10, respectively, and their encoding gene sequences are shown in SEQ ID No. 7 and SEQ ID No. 9, respectively.

[0071] Based on the above vectors, VP64 was fused between TadA8e and its upstream nuclear localization signal NLS in ABE-dCas12i3-5M, ABE-Rad51-dCas12i3-5M, and ABE-HMG-D-dCas12i3-5M, respectively. The N-terminus of VP64 was directly connected to the C-terminus of NLS, and the C-terminus was connected to TadA8e via a 32aa linker peptide, thus constructing Cas12i-mediated adenine base editing systems (VP64-ABE-dCas12i3-5M, VP64-ABE-Rad51-dCas12i3-5M, VP64-ABE-HMG-D-dCas12i3-5M). Figure 1 The resulting vectors are pCXUN-Ubi-VP64-TadA8e-dCas12i3-5M, pCXUN-Ubi-VP64-TadA8e-Rad51-dCas12i3-5M, and pCXUN-Ubi-VP64-TadA8e-HMG-D-dCas12i3-5M.

[0072] The VP64 linker peptide involved is shown in SEQ ID No. 6, and its encoding gene sequence is shown in SEQ ID No. 5.

[0073] After constructing the base editing system vector, the editing window and editing efficiency of each editing system were evaluated using a protoplast transient transformation system and high-throughput sequencing. The base editing system with the highest editing efficiency was selected for further research.

[0074] 1.2.2 Detection of rice protoplast isolation, transformation and editing efficiency To examine the base efficiency and window of these Cas12i-mediated base editing systems fused with different elements in rice, 13 sgRNAs targeting different gene coding regions were designed, including... OsARF4, OsDEP1, OsSBEIIb OsBADH2, OsHRC, OsEPSPS, OsALS, OsACC, OsROC5, OsPDS, OsmiR528, OsGRF4 and OsSD1 (Table 2) and cloned them into the OsU3-sgRNA module of the base editing system, and then transiently expressed in rice protoplasts. Target fragments of the selected target genes / coding regions were amplified by PCR and then subjected to Hi-tom high-throughput sequencing.

[0075] The rice variety was Zhonghua 11. Seeds were first rinsed with 75% ethanol for 1 minute, then treated with 2.5% sodium hypochlorite for 20 minutes, and washed with sterile water at least 5 times. They were then cultured on 1 / 2 MS medium for approximately 2 weeks at 26℃ under 12-hour light (150 μmol / L) conditions. -2 s -1 Using large glass culture cups, 15 seeds can be placed in each cup; 40-60 seedlings can be used for one experiment, and the amount of protoplasts isolated can be transformed into about 6 plasmids.

[0076] Protoplasts were isolated from the stems and leaf sheaths of seedlings and cut into thin filaments approximately 0.5 mm wide using a sharp blade, 20-30 filaments at a time. The filaments were placed in 0.6 M Mannitol solution, wrapped in aluminum foil to protect from light for 10 minutes, then filtered through 75 μm nylon cloth. The filaments were then placed in 50 ml of enzymatic hydrolysate, protected from light, and vacuum-pumped for 30 minutes at approximately 50 kPa. After removal, the filaments were placed on a shaker at room temperature for 5-6 hours at 10-20 rpm. An equal volume of W5 was added to dilute the enzymatic hydrolysate. The hydrolysate was filtered through nylon cloth and collected in a 50 ml centrifuge tube. The tubes were centrifuged at 250 g at 23°C for 3 min (3°C + 3°F), and the supernatant was discarded. The tubes were gently resuspended in 10 ml of W5 solution and centrifuged at 250 g at 23°C for 3 min (3°C + 3°F), and the supernatant was discarded. The tubes were resuspended in an appropriate amount of MMG solution to achieve a protoplast concentration of 2*10⁻⁶. 6 / ml, blood cell counter count.

[0077] Add 20 μg of each of the 12 vector plasmids obtained in step one (the plasmid expressing GFP serves as a control) to a 2 ml centrifuge tube. Using a pipette tip with the tip removed, slowly add 200 μl of protoplasts along the tube wall to the centrifuge tube, gently aspirate and mix. Add 220 μl of PEG-4000, gently invert and mix, and induce transformation in the dark for 10–20 min. Add 800 μl of W5 (room temperature), invert and mix, centrifuge at 250 g horizontally at 23°C, centrifuge at 3°C ​​and 3°C for 3 min, and discard the supernatant. Add 1 ml of WI, invert and mix, and gently transfer to a 6-well plate (pre-added with 1 ml of WI). Incubate at 28°C or room temperature in the dark.

[0078] Twelve hours after transformation, the expression of GFP in the transformed GFP plasmid protoplasts was observed, and the transformation efficiency was calculated. Protoplasts were collected 36–48 hours after transformation. The protoplasts were suspended in 2 ml centrifuge tubes by pipetting, centrifuged at 12,000 rpm for 1 min, and the supernatant was discarded. The protoplast genome was extracted using the same method as for plant genome extraction. The target fragments of the selected target genes were amplified by PCR and subjected to Hi-tom high-throughput sequencing. The target sites of the selected genes were all reported to have been successfully edited in rice cells (Table 2).

[0079] 1.2.3 Stable genetic transformation in rice Based on six endogenous rice gene targets (Table 3), the inventors evaluated the editing efficiency of each gene-editing vector obtained through Agrobacterium-mediated rice genetic transformation experiments. An editing vector based on the most efficient vector obtained in step one was constructed for editing the six rice gene targets in Table 3 and transformed into Zhonghua 11. Resistant T0 generation plants were obtained through hygromycin (50 mg / L) screening. Target site mutation analysis was performed on the regenerated plants obtained from the screening. Genomic DNA was extracted from the T0 generation plants, and PCR amplification was performed using the T0 generation plant genomic DNA as a template to obtain the target gene fragment (Table 1).

[0080] The PCR amplification products were subjected to Sanger sequencing and compared with the wild-type gene sequence to identify the mutation type; the mutations of the amino acids encoded by the endogenous target genes in the regenerated plants were evaluated to identify the protein variation type. Table 1. List of primers used

[0081] Table 2. Selected target points for testing vector editing efficiency

[0082] 2. Results 2.1 Detection of editing activity of Cas12i-mediated base editing system Hi-tom high-throughput sequencing results showed that transcription activator VP64, single-stranded DNA binding protein (Rad51), and double-stranded DNA binding protein (HMG-D) all improved the editing efficiency of the Cas12i-mediated base editing system to varying degrees. Among them, the fusion protein composed of transcription activator VP64 and double-stranded DNA binding protein (HMG-D) with cytosine deaminase mhA3A or adenine deaminase TadA8e significantly improved the base editing efficiency from C to T and from A to G. Figure 2-3 The highest editing efficiency of C to T mediated by CBE-dCas12i3-5M was 4.59%, with a base editing window of C8-C15; the highest editing efficiency of A to G mediated by ABE-dCas12i3-5M was 4.62%, with a base editing window of A8-A13. Figure 2 Compared with the controls CBE-dCas12i3-5M and ABE-dCas12i3-5M, VP64-CBE-Rad51-dCas12i3-5M mediated a maximum C-to-T editing efficiency of 10.32%, with a base editing window of C7-C15, representing an efficiency improvement of approximately 2.3 times; VP64-ABE-Rad51-dCas12i3-5M mediated a maximum A-to-G editing efficiency of 10.92%, with a base editing window of A7-A15, representing an efficiency improvement of approximately 2.4 times. Figure 2 The highest editing efficiency of C to T mediated by VP64-CBE-HMG-D-dCas12i3-5M was 12.80%, with a base editing window of C7-C16. Compared with the control, the editing efficiency was improved by approximately 2.8 times, and compared with VP64-CBE-Rad51-dCas12i3-5M, the editing efficiency was improved by approximately 1.3 times. The highest editing efficiency of A to G mediated by VP64-ABE-HMG-D-dCas12i3-5M was 12.20%, with a base editing window of A7-A15. Compared with the control, the editing efficiency was improved by approximately 2.7 times, and compared with VP64-ABE-Rad51-dCas12i3-5M, the editing efficiency was improved by approximately 1.2 times. Figure 2 In summary, not only are the base editing efficiency mediated by VP64-CBE-HMG-D-dCas12i3-5M and VP64-ABE-HMG-D-dCas12i3-5M improved, but their editing window is also widened, enabling effective editing of bases distal to PAM. Figure 3 ).

[0083] 2.2 Stable transformation and genotype identification of rice Based on Hi-tom sequencing analysis of rice protoplasts, the most efficient base editing systems, VP64-CBE-HMG-D-dCas12i3-5M and VP64-ABE-HMG-D-dCas12i3-5M, were selected for stable transformation and compared with the control systems CBE-dCas12i3-5M and ABE-dCas12i3-5M to further verify the effect of this novel fusion protein on the base editing efficiency of endogenous target genes in rice. OsARF4, OsSBEIIb, OsALS, OsACC, OsDEP1 and OsBADH2 As an endogenous target gene (Table 3), the constructed targeting plasmid containing the endogenous target gene was transformed into Agrobacterium and then transformed into rice callus. After hygromycin selection and regeneration culture, regenerated plants were obtained. PCR amplification of the hygromycin resistance gene on the base editing vector was performed using forward primer Hyg-F and reverse primer Hyg-R to detect whether the T0 generation plants contained exogenous transgenic elements. Specific primers were designed to amplify the editing site of the endogenous target gene using PCR. Sanger sequencing was used to compare the edited plant sequences with the wild-type gene sequences to identify the mutation type and to calculate the base editing efficiency (Tables 4-5). OsARF4 The C-to-T base editing efficiency increased from 9.38% in the control to 32.35%. OsSBEIIb The C-to-T base editing efficiency increased from 11.76% in the control to 22.58%. OsALS The C-to-T base editing efficiency increased from 5.71% in the control to 27.27% (Table 4). OsACC The A-to-G base editing efficiency increased from 7.88% in the control to 20.51%. OsDEP1 The A-to-G base editing efficiency increased from 11.43% in the control to 38.24%. OsBADH2 The A-to-G base editing efficiency increased from 9.38% in the control to 30.00% (Table 5). Compared with the control, the editing efficiency of the novel base editing systems VP64-CBE-HMG-D-dCas12i3-5M and VP64-ABE-HMG-D-dCas12i3-5M can be improved by up to 3.4 times (Tables 4-5).

[0084] 2.3 Identification of herbicide resistance in rice After genetic segregation in the progeny, homozygous lines of OsALS-R190H and OsACC-W2097R without transgenes were obtained in the T1 generation. To further verify the herbicide resistance of these rice lines, OsALS-R190H and OsACC-W2097R seedlings at the three-leaf stage were sprayed with different field-approved concentrations of bispyribac-sodium (BS) and haloxyfop-R-methyl. OsALS-R190H mutant seedlings survived and grew normally two weeks after treatment with 1x and 2x concentrations of BS, while OsALS-R190H seedlings treated with 4x concentration of BS and wild-type seedlings treated with 1x concentration of BS all died (Figure 4). The results indicate that these T1 generation homozygous lines have moderate resistance to bispyribac-sodium. The OsACC-W2097R mutant survived treatment with a 1x concentration of flupyridine, while wild-type seedlings died from the same treatment. Furthermore, the mutants died under higher concentrations (2× and 4×) of the herbicide (Figure 4). These results indicate that W2097R is a weak herbicide resistance allele. In conclusion, Cas12i-mediated CBE and ABE can precisely generate herbicide tolerance alleles in rice. This newly developed Cas12i-mediated CBE and ABE system not only provides an effective tool for targeted mutagenesis of important agronomical traits in rice but also holds great potential for accelerating precision breeding across a wide range of crop varieties.

[0085] Table 3. Selected endogenous target genes and sites for stable transformation of rice

[0086] Table 4. Statistical analysis of the editing efficiency of the Cas12i-mediated CBE system on endogenous target genes in regenerated rice plants.

[0087] Table 5. Statistical analysis of the editing efficiency of the Cas12i-mediated ABE system on endogenous target genes in regenerated rice plants.

[0088] 3. Conclusion of the entire text This invention discloses a fusion protein and base editing system for improving the efficiency of Cas12i-mediated base editing, and their applications. The invention utilizes a novel five-mutant form of the nuclease Cas12i3, which is fused with transcription activator VP64, cytosine deaminase human APOBEC3A (hA3A-Y130F) or adenine deaminase TadA8e, DNA single-strand binding protein (Rad51), or DNA double-strand binding protein (HMG-D), respectively. The aim is to improve the base editing efficiency of endogenous target genes in rice, expand its editing range, and construct a novel Cas12i-mediated base editing system. Using protoplast transformation and high-throughput sequencing technologies, the editing efficiency of different editing systems on the coding sequences of endogenous genes in rice was tested. The highest editing efficiency of C to T mediated by the control system CBE-dCas12i3-5M was 4.59%, while the highest editing efficiency of C to T mediated by the optimized system VP64-CBE-HMG-D-dCas12i3-5M was 12.80%, an improvement of approximately 2.8 times. The highest editing efficiency of A to G mediated by the control system ABE-dCas12i3-5M was 4.62%, while the highest editing efficiency of A to G mediated by the optimized system VP64-ABE-HMG-D-dCas12i3-5M was 12.20%, an improvement of approximately 2.7 times. Figure 2 Furthermore, using rice Zhonghua 11 as the recipient material and rice endogenous genes as the research object, the highest editing efficiency of C to T mediated by the control CBE-dCas12i3-5M was 11.76%, while the highest editing efficiency of C to T mediated by the optimized VP64-CBE-HMG-D-dCas12i3-5M was 32.35%, an increase of approximately 2.8 times. Similarly, the highest editing efficiency of A to G mediated by the control ABE-dCas12i3-5M was 11.43%, while the highest editing efficiency of A to G mediated by the optimized VP64-ABE-HMG-D-dCas12i3-5M was 38.24%, an increase of approximately 3.4 times (Tables 4-5). These results demonstrate the establishment of a novel Cas12i fusion protein-mediated CBE and ABE base editing system, providing important technical support for biological research and genetic improvement of important genes in crops.

[0089] 4. References 1.Duan, Z., Liang, Y., Sun, J., Zheng, H., Lin, T., Luo, P., Wang,M., Liu, R., Chen, Y., Guo, S., Jia, N., Xie, H., Zhou, M., Xia, M., Zhao,K., Wang, S., Liu, N., Jia, Y., Si, W., Chen, Q., Hong, Y., Tian, R., andZhu, J.K. (2024). An engineered Cas12i nuclease that is an efficient genomeediting tool in animals and plants. Innovation (Camb) 5:100564. 2.Gaudelli NM, Komor AC, Rees HA, Packer MS, Badran AH, Bryson DI,Liu DR (2017) Programmable base editing of A•T to G•C in genomic DNA withoutDNA cleavage. Nature 551: 464-471 3.Komor AC, Kim YB, Packer MS, Zuris JA, Liu DR (2016) Programmableediting of a target base in genomic DNA without double-stranded DNA cleavage.Nature 533: 420-424 4.Li, Y., Li, S., Li, C., Zhang, C., Yan, L., Li, J., He, Y., Guo,Y., Lin, Y., Zhang, Y., and Xia, L. (2023). Engineering a plant A-to-K baseeditor with improved performance by fusion with a transactivation module.Plant Commun 4:100667. 5.Niu, Q., Xie, H., Cao, X., Song, M., Wang, X., Li, S., Pang, K.,Zhang, Y., Zhu, J.K., and Zhu, J. (2024). Engineering soybean with highlevels of herbicide resistance with a Cas12‐SF01‐based cytosine base editor.Plant Biotechnol J 22, 2435-2437. 6.Richter MF, Zhao KT, Eton E, Lapinaite A, Newby GA, Thuronyi BW,Wilson C, Koblan LW, Zeng J, Bauer DE, Doudna JA, Liu DR (2020) Phage-assisted evolution of an adenine base editor with improved Cas domaincompatibility and activity. Nature Biotechnol 38: 883-891 7.Tian Y, Shen R, Li Z, Yao Q, Zhang X, Zhong D, Tan X, Song M, HanH, Zhu JK, Lu Y (2022) Efficient C-to-G editing in rice using an optimizedbase editor. Plant Biotechnol J 20: 1238-1240 8.Wang, W., Yan, L., Li, J., Zhang, C., He, Y., Li, S., and Xia, L.(2025a). Engineering a robust Cas12i3 variant-mediated wheat genome editingsystem. Plant Biotechnol J 23:860-873. 9.Wang, W., Li, S., Yang, J., Li, J., Yan, L., Zhang, C., He, Y., andXia, L. (2025b). Exploiting the efficient Exo:Cas12i3-5M fusions for robustsingle and multiplex gene editing in rice. J Integr Plant Biol 67:1246-1253. 10.Wei C, Wang C, Jia M, Guo HX, Luo PY, Wang MG, Zhu JK, Zhang H(2021) Efficient generation of homozygous substitutions in rice in onegeneration utilizing an rABE8e base editor. J Integr Plant Biol 63: 1595-1599 11.Xue, N., Hong, D., Zhang, D., Wang, Q., Zhang, S., Yang, L., Chen,X., Li, Y., Han, H., Hu, C., Liu, M., Song, G., Guan, Y., Wang, L., Zhu, Y.,and Li, D. (2024). Engineering IscB to develop highly efficient miniatureediting tools in mammalian cells and embryos. Mol Cell 84:3128-3140 e3124. 12.Zhao D, Li J, Li S, Xin X, Hu M, Price MA, Rosser SJ, Bi C, ZhangX (2021) Glycosylase base editors enable C-to-A and C-to-G base changes. NatBiotechnol 39: 35-40 13. Zheng Z., Liu, T., Chai, N., Zeng, D., Zhang, R., Wu, Y., Hang, J., Liu, Y., Deng, Q., Tan, J., Liu, J., Xie, X., Liu, YG, and Zhu, Q. (2024). PhieDBEs: a DBD‐containing, PAM‐flexible, high‐efficiency dual baseeditor toolbox with wide targeting scope for use in plants. Plant BiotechnolJ 22, 3164-3174. The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.

Claims

1. A fusion protein, namely fusion protein A or fusion protein B, wherein fusion protein A contains cytosine deaminase APOBEC3A, dCas12i3-5M and uracil glycosylation inhibitor UGI, and wherein fusion protein B contains adenine deaminase TadA8e and dCas12i3-5M.

2. The protein according to claim 1, characterized in that: The fusion protein A and the fusion protein B also contain either the DNA single-strand binding protein Rad51 or the DNA double-strand binding protein HMG-D.

3. The protein according to claim 1 or 2, characterized in that: The fusion protein A and the fusion protein B also contain a transcriptional activation domain VP64.

4. The protein according to any one of claims 1-3, characterized in that: The fusion protein A and the fusion protein B also contain nuclear localization signals and / or linker peptides.

5. The protein according to any one of claims 1-4, characterized in that: The fusion protein A is fusion protein A1, fusion protein A2, fusion protein A3, fusion protein A4, fusion protein A5, or fusion protein A6. The fusion protein A1 is obtained by sequentially linking the nuclear localization signal, the transcription activation domain VP64, the cytosine deaminase APOBEC3A, the DNA single-strand binding protein Rad51, dCas12i3-5M, and the uracil glycosylation inhibitor protein UGI. The fusion protein A2 is obtained by sequentially linking the nuclear localization signal, the transcription activation domain VP64, the cytosine deaminase APOBEC3A, the DNA double-strand binding protein HMG-D, dCas12i3-5M, and the uracil glycosylation inhibitor protein UGI. The fusion protein A3 is obtained by sequentially linking the nuclear localization signal, the transcriptional activation domain VP64, the cytosine deaminase APOBEC3A, dCas12i3-5M and the uracil glycosylation inhibitor protein UGI. The fusion protein A4 is obtained by sequentially linking the nuclear localization signal, the cytosine deaminase APOBEC3A, the DNA single-strand binding protein Rad51, dCas12i3-5M and the uracil glycosylation inhibitor protein UGI; The fusion protein A5 is obtained by sequentially linking the nuclear localization signal, the cytosine deaminase APOBEC3A, the DNA double-strand binding protein HMG-D, dCas12i3-5M and the uracil glycosylation inhibitor protein UGI. The fusion protein A6 is obtained by sequentially linking the nuclear localization signal, the cytosine deaminase APOBEC3A, dCas12i3-5M and the uracil glycosylation inhibitor protein UGI; The fusion protein B is fusion protein B1, fusion protein B2, fusion protein B3, fusion protein B4, fusion protein B5, or fusion protein B6. The fusion protein B1 is obtained by sequentially linking the nuclear localization signal, the transcription activation domain VP64, the adenine deaminase TadA8e, the DNA single-strand binding protein Rad51, and dCas12i3-5M. The fusion protein B2 is obtained by sequentially linking the nuclear localization signal, the transcription activation domain VP64, the adenine deaminase TadA8e, the DNA double-strand binding protein HMG-D, and dCas12i3-5M. The fusion protein B3 is obtained by sequentially linking the nuclear localization signal, the transcriptional activation domain VP64, the adenine deaminase TadA8e, and dCas12i3-5M. The fusion protein B4 is obtained by sequentially linking the nuclear localization signal, the adenine deaminase TadA8e, the DNA single-strand binding protein Rad51, and dCas12i3-5M. The fusion protein B5 is obtained by sequentially linking the nuclear localization signal, the adenine deaminase TadA8e, the DNA double-strand binding protein HMG-D, and dCas12i3-5M. The fusion protein B6 is obtained by sequentially linking the nuclear localization signal, the adenine deaminase TadA8e, and dCas12i3-5M.

6. The protein according to any one of claims 1-5, characterized in that: The cytosine deaminase APOBEC3A is either A1) or A2) below. A1) The amino acid is the protein at positions 9-207 of SEQ ID No. 2; A2) A protein that has the same function by substitution and / or deletion and / or addition of one or more amino acid residues, except for position 130, in positions 9-207 of SEQ ID No. 2; dCas12i3-5M is as follows (B1) or (B2): B1) The amino acid is the protein located at positions 240-1287 of SEQ ID No. 2; B2) A protein having the same function by substituting and / or deleting and / or adding one or more amino acid residues from position 240 to 1287 of SEQ ID No. 2; The uracil glycosylation inhibitor protein UGI is as follows (C1) or (C2): C1) Amino acid is the protein at positions 1298-1380 of SEQ ID No. 2; C2) A protein having the same function by substituting and / or deleting and / or adding one or more amino acid residues at positions 1298-1380 of SEQ ID No. 2; The DNA single-stranded binding protein Rad51 is either D1 or D2 as follows: D1) The amino acid is the protein consisting of positions 1-114 of SEQ ID No. 8; D2) Proteins with the same function by substitution and / or deletion and / or addition of one or more amino acid residues at positions 1-114 of SEQ ID No. 8; The DNA double-strand binding protein HMG-D is either E1 or E2). E1) The amino acid is the protein at positions 1-112 of SEQ ID No. 10; E2) A protein that has the same function as SEQ ID No. 10 by substituting and / or deleting and / or adding one or more amino acid residues from position 1 to 112. The transcriptional activation domain VP64 is as follows (F1) or (F2): F1) The amino acid is the protein at positions 1-50 of SEQ ID No. 6; F2) Proteins that have the same function by substituting and / or deleting and / or adding one or more amino acid residues at positions 1-50 of SEQ ID No. 6; The adenine deaminase TadA8e is either G1 or G2 as follows: G1) Amino acid is the protein at positions 9-174 of SEQ ID No. 4; G2) A protein that has the same function as SEQ ID No. 4, with one or more amino acid residues substituted and / or deleted and / or added at positions 9-174.

7. The biomaterial associated with any of the fusion proteins described in claims 1-6 is at least one of the following (H1)-H4): H1) is a nucleic acid molecule encoding any of the fusion proteins described in claims 1-6; H2) contains an expression cassette containing the nucleic acid molecule described in H1); H3) A recombinant vector containing the nucleic acid molecule described in H1) or a recombinant vector containing the expression cassette described in H2); H4) Recombinant microorganisms containing the nucleic acid molecules described in H1), recombinant microorganisms containing the expression cassette described in H2), or recombinant microorganisms containing the recombinant vector described in H3).

8. The biomaterial according to claim 7, characterized in that: In H1), the nucleic acid molecule encoding the cytosine deaminase APOBEC3A is the DNA molecule shown at positions 2034-2630 of SEQ ID No. 1; The nucleic acid molecule encoding dCas12i3-5M is the DNA molecule shown at positions 2727-5870 of SEQ ID No. 1; The nucleic acid molecule encoding the uracil glycosylation inhibitor protein UGI is the DNA molecule shown at positions 5901-6149 of SEQ ID No. 1; The nucleic acid molecule encoding the DNA single-stranded binding protein Rad51 is the DNA molecule shown in positions 1-342 of SEQ ID No. 7; The nucleic acid molecule encoding the DNA double-strand binding protein HMG-D is the DNA molecule shown in positions 1-336 of SEQ ID No. 9; The nucleic acid molecule encoding the transcriptional activation domain VP64 is the DNA molecule shown in positions 1-150 of SEQ ID No. 5; The nucleic acid molecule encoding the adenine deaminase TadA8e is the DNA molecule shown at positions 2034-2531 of SEQ ID No.

3.

9. The use of the fusion protein of any one of claims 1-6 or the biomaterial of claim 7 or 8 in plant gene editing, or in the preparation of plant gene editing products.

10. The use of the fusion protein of any one of claims 1-6 or the biomaterial of claim 7 or 8 in the preparation of herbicide-resistant plants.