Optimization of the ultracompact Fanzor2 system and its application in plant genome editing

CN122563918APending Publication Date: 2026-08-14SANYA NATIONAL INSTITUTE OF SOUTHERN BREEDING CHINESE ACADEMY OF AGRICULTURAL SCIENCES +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

通过对ωRNA和enNlovFz2蛋白进行理性工程化改造,我们开发了enNlovFz2(K416R)/ωRNA-Rz系统。该系统在大豆和水稻中均表现出更高且稳定的基因组编辑活性,成功靶向了野生型版本难以编辑的位点。此外,基于该优化系统开发的新型胞嘧啶碱基编辑器,在植物中实现了高效的C-to-T碱基转换,最高编辑效率达到8.27%。这些发现为植物基因组编辑和碱基编辑建立了一个多功能平台,表明工程化的Fanzor系统可作为作物性状精准遗传修饰的有效工具。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This application discloses the optimization of the ultra-compact Fanzor2 system and its application in plant genome editing, belonging to the fields of protein and genetic engineering technology. This application developed the enNlovFz2(K416R) / ωRNA-Rz system through rational engineering of ωRNA and the enNlovFz2 protein. This system exhibited higher and more stable genome editing activity in both soybean and rice, successfully targeting sites that were difficult to edit in the wild-type version. Furthermore, a novel cytosine base editor developed based on this optimized system achieved highly efficient C-to-T base conversion in plants, with a maximum editing efficiency of 8.27%. These findings establish a multifunctional platform for plant genome editing and base editing, demonstrating that the engineered Fanzor system can serve as an effective tool for precise genetic modification of crop traits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of protein and genetic engineering technology, specifically relating to the optimization of the ultra-compact Fanzor2 system and its application in plant genome editing. Background Technology

[0002] The advent of CRISPR-Cas systems (including Cas9 and Cas12a) has enabled precise genetic modification of a wide range of organisms, from plants to humans, thus revolutionizing the field of genome engineering (Wang and Doudna, 2023). These systems are evolutionarily associated with the IS200 / 605 transposon-associated proteins IscB and TnpB (Altae-Tran et al., 2021), the latter two belonging to the obligate mobile element guided activity (OMEGA) system (Saito et al., 2023). TnpB is a prokaryotic RNA-guided DNA endonuclease that uses guide ωRNA to cleave target DNA and recognizes specific target neighbor motifs (TAMs). Recent studies have confirmed its genome editing activity in eukaryotic cells (Karmakar et al., 2024; Karvelis et al., 2021; Li et al., 2024; Lv et al., 2024; Thornton et al., 2025; Zhang et al., 2024). Fanzor, a eukaryotic homologue of TnpB, constitutes the core nuclease of the OMEGA system (Saito et al., 2023). As a programmable RNA-guided endonuclease, Fanzor is structurally and functionally similar to CRISPR-Cas12, but its molecular size is about half that of Cas12. This compact structure gives it unique advantages in the delivery of drugs and biotechnological applications (Jiang et al., 2023), and also shows great potential in plant virus-mediated genome editing (Ji et al., 2025).

[0003] Fanzor proteins are divided into two families: Fanzor1 is associated with various eukaryotic transposons and is found in fungi, protozoa, arthropods, plants, and giant viruses; Fanzor2 has higher homology with TnpB and is typically found in IS607-like transposons and double-stranded DNA viruses. Studies have confirmed that seven Fanzor proteins—SpuFz1 from Spizellomyces punctatus, KnFNuc from Klebsormidium nitens, NlovFz2 from Naegleria lovaniensis, MmeFz2 from Mercenaria mercenaria, ApmFNuc from Acanthamoeba polyphaga mimivirus, and DpFNuc from Dreissena polymorpha—are programmable RNA-guided nucleases with the potential for application in mammalian genome editing (Jiang et al., 2023; Saito et al., 2023; Xu et al., 2024). However, the performance of the Fanzor system in plants remains poorly reported, with preliminary studies showing limited editing activity. For example, SpuFz1 showed no editing activity in rice protoplasts (Zhang et al., 2024). Subsequently, Ji et al. (2025) systematically evaluated several candidate nucleases (SpuFz1, GtFz1, and MmeFz2) in rice and wheat using a single transcription unit (STU) strategy, finding that NlovFz2 (489 amino acids) was the most effective. This nuclease successfully induced mutations in protoplasts and regenerated rice plants. However, its editing efficiency was relatively limited: the mutation rate in protoplasts averaged 0.9% to 7.3%, and the frequency of biallelic or homozygous mutations in stable transgenic lines ranged from 1.2% to 26.5%. While these results provide crucial proof of concept for Fanzor-mediated plant genome editing, they also highlight the urgent need to further improve editing efficiency through protein engineering or guide RNA optimization.

[0004] With the support of AlphaFold3, researchers engineered an enhanced version of enNlovFz2. This variant recognizes extended TAMs (5'-NMYG-3'), exhibits approximately 11-fold improved editing efficiency, and enables in vivo therapeutic editing via a single recombinant adeno-associated virus (rAAV) delivery (Wei et al., 2025). Despite these advancements, the efficacy and optimization potential of enNlovFz2 in plant systems remain to be explored, representing a promising direction for future research.

[0005] References [1] Saito M, Xu P, Faure G, et al. Fanzor is a eukaryoticprogrammable RNA-guided endonuclease[J]. Nature, 2023, 620(7974): 660-668. [2]Ji Y, Sun Y, Zhou H, et al. Engineer the eukaryotic OMEGA Fanzorsystems for genome editing in plants[J]. Journal of Integrative PlantBiology, 2025. [3]Wei Y, Gao P, Pan D, et al. Engineering eukaryotic transposon-encoded Fanzor2 system for genome editing in mammals[J]. Nature ChemicalBiology, 2025: 1-10. [4]Zhang R, Tang X, He Y, et al. IsDge10 is a hypercompact TnpBnuclease that confers efficient genome editing in rice[J]. PlantCommunications, 2024: 101068. [5]Bin Moon S, Lee JM, Kang JG, et al. Highly efficient genomeediting by CRISPR-Cpf1 using CRISPR RNA with a uridinylate-rich 3′-overhang[J]. Nature Communications, 2018, 9(1): 3651. [6]Thornton BW, Weissman RF, Tran RV, et al. Latent activity in TnpBrevealed by mutational scanning[J]. bioRxiv, 2025: 2025.2002.2011.637750. [7]Luo J, Chia N, Qin Y, et al. STAGE: A compact and versatile TnpB-based genome editing toolkit for Streptomyces[J]. Proceedings of the NationalAcademy of Sciences, 2025, 122(17): e2509146122. [8]Yamatani H, Sato Y, Masuda Y, et al. NYC4, the rice ortholog ofArabidopsis THF1, is involved in the degradation of chlorophyll–proteincomplexes during leaf senescence[J]. The Plant Journal, 2013, 74(4): 652-662. [9]Bai S, Cao X, Hu L, et al. Engineering an optimized hypercompactCRISPR / Cas12j‐8 system for efficient genome editing in plants[J]. PlantBiotechnology Journal, 2025, 23(6): 1153-1164.

[10] Sun T, Liu Q, Chen X, et al. Hi-TOM 2.0: an improved platform forhigh-throughput mutation detection[J]. Science China Life Sciences, 2024: 1-3.

[11] Xie X, Ma X, Zhu Q, et al. CRISPR-GE: a convenient softwaretoolkit for CRISPR-based genome editing[J]. Molecular Plant, 2017, 10(9):1246-1249.

[12] Chen C, Wu Y, Li J, et al. TBtools-II: A “one for all, all for one” bioinformatics platform for biological big-data mining[J]. MolecularPlant, 2023, 16(11): 1733-1742.

[13] DeLano WL. Pymol: An open-source molecular graphics tool[J].CCP4 Newsletter on Protein Crystallography, 2002, 40: 82-92. Summary of the Invention The technical problem this application aims to solve is: how to obtain an RNA-guided nuclease with higher nucleic acid editing efficiency and a composition for nucleic acid editing, and how to use the obtained enhanced RNA-guided nuclease and composition for nucleic acid editing. To solve this technical problem, this application provides the following technical solution: In a first aspect, this application provides a protein selected from at least one of the following (A1) to (A4): (A1) The protein with an amino acid sequence as shown in SEQ ID NO:7, positions 10 to 497 (i.e., enNlovFz2(K416R)). Proteins with at least 70% identity to protein (A1) and possessing RNA-guided endonuclease function obtained by substituting, deleting and / or adding amino acid residues of the amino acid sequences shown in (A2) and (A1). (A3) A fusion protein comprising the protein described in (A1) or (A2) and other proteins or polypeptides; (A4) A conjugate comprising the protein described in (A1) or (A2) and a modified portion; In (A2), the replacement is D294A (i.e., the denNlovFz2(K416R) variant).

[0006] In this application, the protein can be used as an RNA-guided endonuclease for nucleic acid editing under the guidance of guide RNA.

[0007] In this application, the letter before the substitution number represents the amino acid residue of enNLovFz2(K416R) (corresponding to the amino acid residue at the numerical position in SEQ ID NO:15), and the letter after the number represents the mutant amino acid residue. For example, D294A indicates that the D at position 294 of SEQ ID NO:15 is replaced with A, while keeping the amino acid residues of SEQ ID NO:15 unchanged, to obtain a mutant protein, which is named the denNlovFz2(K416R) variant.

[0008] In some specific embodiments of this application, the amino acid sequence of the denNlovFz2(K416R) variant is shown as positions 184 to 671 of SEQ ID NO:9.

[0009] The protein has at least one of the following characteristics: The other proteins or polypeptides described in (A3) are selected from epitope tags, reporter genes, nuclear localization signals, deaminases, transcription activation domains, transcription repression domains, nuclease domains, reverse transcriptases, and any combination thereof; The modification portion described in (A4) is selected from the other proteins or polypeptides, detectable markers, and any combination thereof.

[0010] The detectable label is selected from enzyme labeling, radioactive isotope labeling, fluorescent labeling, biotin-avidin systems, and any combination thereof. The enzyme label is selected from horseradish peroxidase (HRP), alkaline phosphatase (AP), and any combination thereof. The radioactive isotope label is selected from... 35 S (labeled methionine) 3 H (labeled leucine) and 125 I (for iodinated proteins), etc. The fluorescent label is selected from FITC (fluorescein isothiocyanate), Alexa Fluor series dyes, Cy series (such as Cy3, Cy5), and any combination thereof. The biotin-avidin system, for example, utilizes the strong binding of biotin (small molecule) to streptavidin (protein) for detection or purification.

[0011] The fusion protein described in (A3) has the following characteristics as described in (i) or (ii): (i) The fusion protein contains an NLS sequence and / or an epitope tag; (ii) In the fusion protein, the additional protein or polypeptide is linked to the N-terminus or C-terminus of the protein via peptide bonds or linkers; And / or, the conjugate described in (A4) has the following characteristics as described in (i) or (ii): (i) The conjugate contains an NLS sequence and / or an epitope tag; (ii) In the conjugate, the modified portion is linked to the N-terminus or C-terminus of the protein via a peptide bond or a linker, or the modified portion is fused to the N-terminus or C-terminus of the protein.

[0012] The epitope tag (protein tag) refers to a polypeptide or protein fused with a target protein using in vitro DNA recombination technology for expression, detection, tracing, and / or purification of the target protein. The protein tag may be a Flag protein tag, His protein tag, MBP protein tag, HA protein tag, myc protein tag, GST protein tag, or 35 / or SUMO protein tag, etc.

[0013] Wherein, the fusion protein of A3) and / or the conjugate of A4) contain an NLS sequence.

[0014] The terms "nuclear localization signal," "nuclear localization sequence," or "NLS," used interchangeably in this article, refer to the amino acid sequence that "tags" a protein for transport into the nucleus via nuclear transport. Typically, this signal consists of a short sequence of one or more positively charged lysine or arginine residues exposed on the protein surface. Different nuclear localization proteins may share the same NLS. The NLS functions in contrast to nuclear export signals, which aim to expel the protein from the nucleus.

[0015] In some specific embodiments of this application, the fusion protein of A3) and / or the conjugate of A4) are linked with one or more NLS sequences.

[0016] In some specific embodiments of this application, the NLS sequence is selected from the SV40 nuclear localization signal (NLS) sequence, the nucleoplasmin nuclear localization signal sequence, and any combination thereof.

[0017] In some specific embodiments of this application, the SV40 nuclear localization signal (NLS) sequence is shown as SEQ ID NO:16 (same as SEQ ID NO:7, positions 3 to 9) (PKKKRKV).

[0018] In some specific embodiments of this application, the nucleoplasmin nucleus localization signal sequence is shown in SEQ ID NO:17 (same as bits 672 to 687 of SEQ ID NO:10) (KRPAATKKAGQAKKKK).

[0019] In some specific embodiments of this application, the additional protein or polypeptide is cytosine deaminase (mini-Sdd7) and uracil glycosylation inhibitor (UGI). The amino acid sequence of cytosine deaminase (mini-Sdd7) is shown in SEQ ID NO:18 (same as positions 9 to 183 of SEQ ID NO:10). The amino acid sequence of uracil glycosylation inhibitor (UGI) is shown in SEQ ID NO:19 (same as positions 695 to 777 of SEQ ID NO:10).

[0020] Specifically, the fusion protein described in (A3) is at least one of the following: (B1) A fusion protein (enNlovFz2(K416R)) with an amino acid sequence as shown in SEQ ID NO:15. (B2) A fusion protein with the amino acid sequence shown in SEQ ID NO:7 (SV40 NLS-enNlovFz2 (K416R) fusion protein); (B3) The fusion protein with the amino acid sequence shown in SEQ ID NO:10 (SV40 NLS-Cytidinedeaminases(mini Sdd7)-denNlovFz2(D294A)-Nucleoplasmin NLS-Uracil glycosylase inhibitor(UGI)-SV40NLS).

[0021] The second aspect: compositions and / or complexes for nucleic acid editing. This application also provides compositions and / or complexes for nucleic acid editing, said compositions and / or complexes comprising: (i) a protein component, wherein the protein component is at least one of the proteins described above; (ii) A nucleic acid component comprising, from 5' to 3', an ωRNA scaffold and a guide sequence capable of hybridizing with a target sequence, wherein the nucleotide sequence of the ωRNA scaffold is shown in SEQ ID NO:8.

[0022] In some specific embodiments of this application, the protein component is a protein with an amino acid sequence as shown in SEQ ID NO:7 or a protein with an amino acid sequence as shown in SEQ ID NO:10.

[0023] In some specific embodiments of this application, the guide sequence is attached to the 3' end of the ωRNA scaffold.

[0024] In some specific embodiments of this application, the guide sequence is the same as the target sequence. That is, the guide sequence and the target strand hybridize and complement each other.

[0025] In one specific implementation, the guide sequence is 16-17 nucleotides in length.

[0026] In some embodiments, the TAM of the nucleic acid component is 5'-NMYG-3'. N is A, T, C, or G; M is A or C; Y is T or C.

[0027] The TAM is located at the 5' end of the target sequence.

[0028] In this application, "TAM" has the same meaning as "PAM".

[0029] The third aspect: biomaterials related to the aforementioned proteins, compositions, and / or complexes. This application also provides biological materials, said biological materials being selected from at least one of the following: (C1) The DNA molecule encoding the above protein; (C2) An expression cassette containing the DNA molecule described in (C1); (C3) A recombinant vector containing the expression cassette described in (C2) and / or the above-mentioned nucleic acid components for transcription; (C4) Recombinant cells containing the expression cassette described in (C2) and / or the expression cassette that transcribes the above-described nucleic acid components.

[0030] Wherein, the DNA molecule described in (C1) is at least one of the following: (C1-1) The nucleotide sequence is as shown in the DNA molecule at positions 829 to 2295 of SEQ ID NO:6 (encoding enNlovFz2(K416R)). (C1-2) The nucleotide sequence is as shown in the DNA molecule at positions 808 to 2295 of SEQ ID NO:6 (encoding the SV40 NLS-enNlovFz2(K416R) fusion protein). (C1-3) The nucleotide sequence is as shown in the DNA molecule at positions 802 to 2295 of SEQ ID NO:6 (encoding the SV40 NLS-enNlovFz2(K416R) fusion protein). (C1-4) The nucleotide sequence is as shown in the DNA molecule at positions 1351 to 2814 of SEQ ID NO:9 (encoding denNlovFz2(K416R)). (C1-5) The nucleotide sequence is as shown in the DNA molecule at positions 805 to 3165 of SEQ ID NO:9 (encoding SV40 NLS-Cytidine deaminases (mini Sdd7)-denNlovFz2 (D294A)-Nucleoplasmin NLS-Uracilglycosylase inhibitor (UGI)-SV40NLS fusion protein). (C1-6) The nucleotide sequence is as shown in the DNA molecule at positions 802 to 3165 of SEQ ID NO:9 (encoding SV40 NLS-Cytidine deaminases (mini Sdd7)-denNlovFz2 (D294A)-Nucleoplasmin NLS-Uracilglycosylase inhibitor (UGI)-SV40NLS fusion protein).

[0031] Wherein, the expression box described in (C2) is at least one of the following: (C2-1) A DNA molecule with a nucleotide sequence as shown in SEQ ID NO:6; (C2-2) A DNA molecule with a nucleotide sequence as shown in SEQ ID NO:9; (C2-3) DNA molecules with nucleotide sequences as shown in SEQ ID NO:11.

[0032] In this application, the nucleic acid component or DNA molecule is an isolated nucleic acid molecule.

[0033] The term "separated" refers to a molecule that is at least partially separated from other molecules normally associated with it in its natural state. In one aspect, the term "separated" refers to a DNA molecule that is separated from nucleic acids normally located flanking the DNA molecule in its natural state. For example, a DNA molecule encoding a protein naturally present in bacteria would be a separated DNA molecule if it were not within the DNA of a bacterium that naturally contains a DNA molecule encoding that protein. Therefore, DNA molecules fused to or operatively linked to one or more other DNA molecules not associated with them in nature, such as as a result of recombinant DNA or plant transformation techniques, are considered separated herein. Even if these molecules are integrated into the chromosome of a host cell or exist together with other DNA molecules in a nucleic acid solution, they are considered separated.

[0034] The fourth aspect: products, methods, and industrial applications of nucleic acid editing. This application provides products for nucleic acid editing, said products containing the proteins, compositions and / or complexes described above, and / or biological materials.

[0035] This application also provides a method for nucleic acid editing, the method comprising the following steps: contacting the cells to be edited with the above-described proteins, compositions and / or complexes, biological materials and / or products to achieve nucleic acid editing.

[0036] This application also provides the use of the above-described proteins, compositions and / or complexes, and / or biological materials in the preparation of formulations for nucleic acid editing.

[0037] This application also provides the use of the above-described proteins, compositions and / or complexes, and / or biological materials in nucleic acid editing, wherein the nucleic acid editing is for non-disease treatment purposes.

[0038] The nucleic acid editing is gene editing, which includes gene knockout, base editing, altering the expression of gene products, repairing mutations, and / or inserting polynucleotides.

[0039] In some specific embodiments of this application, the proteins, conjugates, fusion proteins, isolated nucleic acid molecules (RNA), isolated nucleic acid molecule compositions (DNA), and vectors of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to: electroporation, lipid transfection, nuclear transfection, microinjection, acoustic wave effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.

[0040] In this application, the protein described in A1) or A2) above can be used alone as an RNA-directed endonuclease to bind with the RNA for nucleic acid editing.

[0041] In this application, the protein may also be complexed with other enzymes or co-expressed as a fusion protein, which then binds to the RNA for nucleic acid editing. Exemplarily, the effector protein described in A1) or A2) may be complexed with a deaminase or co-expressed as a fusion protein, which then binds to the RNA for base editing. The deaminase is selected from cytidine deaminase or adenosine deaminase. In some specific embodiments of this application, the protein is fused with cytosine deaminase and a glycosylation inhibitor for base editing (CtoT).

[0042] In this application, the nucleic acid editing or gene editing occurs in cells.

[0043] The cells can be in vitro cells or in vivo cells.

[0044] The cells are eukaryotic or prokaryotic cells. The eukaryotic cells are plant or animal cells.

[0045] The plant in question is either a monocotyledonous plant or a dicotyledonous plant.

[0046] In some specific embodiments of this application, the monocotyledonous plant is rice.

[0047] In some specific embodiments of this application, the dicotyledonous plant is soybean.

[0048] The beneficial technical effects achieved by this application are as follows: Through rational engineering of ωRNA and the enNlovFz2 protein, we developed the enNlovFz2(K416R) / ωRNA-Rz system. This system exhibited higher and more stable genome editing activity in both soybean and rice, successfully targeting sites that were difficult to edit in the wild-type version. Furthermore, a novel cytosine base editor developed based on this optimized system achieved highly efficient C-to-T base conversion in plants, with a maximum editing efficiency of 8.27%. These findings establish a multifunctional platform for plant genome editing and base editing, demonstrating that the engineered Fanzor system can serve as an effective tool for precise genetic modification of crop traits. Attached Figure Description

[0049] Figure 1 This section presents preliminary optimization and editing efficiency comparisons of the NlovFz2 system. A shows a schematic diagram of the vectors used for genome editing in soybean hairy roots by the NlovFz2, enNlovFz2, and enNlovFz2 / ωRNA-Rz systems; B compares the genome editing efficiency of the three systems at eight endogenous sites in soybean hairy roots; C shows the distribution of deletion mutation locations induced by enNlovFz2 / ωRNA-Rz at the GmSweet15a-T1 site; D shows the size distribution of deletion mutations induced by enNlovFz2 / ωRNA-Rz at the GmSweet15a-T1 site. Note: TAMs and spacers are indicated in red and blue, respectively; each data point represents a biological replicate of an independent experiment (n>3); data are expressed as mean ± standard deviation; ns, P>0.05; *, P<0.05; **, P<0.01; ***, P<0.001; ****, P<0.0001.

[0050] Figure 2 Multiple sequence alignments of enNlovFz2, ApmFz2, and IsDra2 TnpB.

[0051] Figure 3To rationally design and identify the key mutation K416R. In this study, A represents the structural alignment of enNlovFz2 with ApmFz2 and IsDra2 (the protein structures of ApmFz2, IsDra2, and enNlovFz2 are shown in green, blue, and yellow, respectively; the conserved residue I304 in the IsDra2 TnpB RuvC domain and its corresponding residue K416 in enNlovFz2 are highlighted in red); B represents the domain composition of NlovFz2 and the editing efficiency of each enNlovFz2 variant at the GmSweet15a-T1 site (UNK: unknown domain; REC: recognition domain; WED: wedge domain; RuvC: RuvC endonuclease domain; ZF: zinc finger element); and C represents the genome editing efficiency of enNlovFz2 (K416R) at the remaining 7 soybean target sites. Note: Each data point represents a biological replicate of an independent experiment (n>3); data are expressed as mean ± standard deviation; ns, P>0.05; *, P<0.05; **, P<0.01; ***, P<0.001; ****, P<0.0001.

[0052] Figure 4 The genome editing efficiency of enNlovFz2(K416R) at 18 non-CCG motif TAM sites.

[0053] Figure 5 This section describes the construction and activity validation of a Fanzor-based plant cytosine base editor. Figure A shows a schematic diagram of the vector for the denNlovFz2 (K416R) / ωRNA-Rz / CBE cytosine base editor; Figure B shows the editing window and efficiency of the denNlovFz2 (K416R) / ωRNA-Rz / CBE system at three sites; Figure C shows the base editing spectrum at the GmSweet15a-T1 site. Note: TAMs and spacers are indicated in red and blue, respectively; mutant bases are indicated by lowercase letters; each data point represents a biological replicate of an independent experiment (n>3); data are expressed as mean ± standard deviation.

[0054] Figure 6The editing performance of enNlovFz2 (K416R) / ωRNA-Rz in stable rice lines is shown. A is a schematic diagram of the enNlovFz2 (K416R) / ωRNA-Rz vector used for rice genome editing; B shows the editing efficiency of enNlovFz2 (K416R) / ωRNA-Rz at eight endogenous sites and the proportion of biallelic / homozygous mutants in rice T0 generation plants; C shows (i) the genotypes of different Osthf1 mutants (TAM and spacer regions are represented in red and blue, respectively; short lines indicate deletions; S indicates substitutions); and (ii) the chlorotic phenotype of Osthf1 mutants under dark-induced leaf senescence conditions. Detailed Implementation

[0055] Examples of resources describing many of the molecular biology-related terms used in this paper can be found in the following literature: Alberts et al., *Molecular Biology of The Cell*, 5th ed., Garland Science Publishing, Inc.: New York, 2007; Rieger et al., *Glossary of Genetics: Classical and Molecular*, 5th ed., Springer-Verlag: New York, 1991; King et al., *A Dictionary of Genetics*, 6th ed., Oxford University Press: New York, 2002; and Lewin, *GenesIX*, Oxford University Press: New York, 2007. Any references cited in this paper, including, for example, all patents, published patent applications, and non-patent publications, are incorporated herein by reference in their entirety.

[0056] Fanzor protein, a direct eukaryotic homolog of prokaryotic TnpB protein, represents a new class of ultracompact RNA-guided endonucleases with advantages in adeno-associated virus (AAV) delivery and plant virus-induced genome editing. However, its application in plants remains limited due to its relatively low editing activity. This application significantly improves its genome editing performance in plants by optimizing ωRNA expression and protein engineering. First, benchmark tests were conducted using soybean hairy roots to compare the original NlovFz2 / ωRNA system with the modified enNlovFz2 / ωRNA system. The results showed that enNlovFz2 / ωRNA successfully edited four sites that NlovFz2 / ωRNA could not edit, and improved the editing efficiency by 5.45-fold and 6.75-fold at two target sites, respectively. Further, by fusing the self-cleaving HDV ribozyme to ωRNA (enNlovFz2 / ωRNA-Rz), the editing efficiency at multiple target sites was improved by 1.79 to 63.02-fold, and successful editing of two target sites that enNlovFz2 / ωRNA could not edit was achieved. Through rational design and structure-guided protein engineering, we identified a key amino acid substitution (K416R) in the RuvC domain that significantly enhanced editing activity, and named the system enNlovFz2(K416R) / ωRNA-Rz. Furthermore, we developed, for the first time, a plant cytosine base editor based on enNlovFz2(K416R), demonstrating the potential for C-to-T base editing. In stable rice transformation lines, the enNlovFz2(K416R) / ωRNA-Rz system achieved targeted mutagenesis at eight endogenous loci, with editing efficiencies ranging from 17.6% to 93.3%, and the proportion of biallelic or homozygous lines was as high as 16.7% to 53.6%. In summary, the engineered enNlovFz2(K416R) / ωRNA-Rz system established in this application is an ultra-compact, multifunctional, and efficient plant genome editing platform that demonstrates application potential for crop improvement through conventional genetic transformation or virus-induced genome editing.

[0057] Fanzor (Fz), a eukaryotic homolog of the prokaryotic transposon encoding the TnpB protein, is a novel class of programmable RNA-guided DNA nucleases belonging to the OMEGA system. The ultracompact Fanzor system consists of an Fz protein of only 400-700 amino acids and an ωRNA of 75-200 nucleotides, offering unique advantages in adeno-associated virus (AAV) delivery and plant virus-mediated genome editing. However, the editing activity of the Fanzor system in plants is currently generally low, and no Fanzor-based plant base editors have been reported.

[0058] This application successfully constructed a significantly improved enNlovFz2 (K416R) / ωRNA-Rz genome editing system through a systematic optimization strategy: First, the editing activities of wild-type NlovFz2 and engineered enNlovFz2 were compared in soybean hairy roots, revealing that enNlovFz2 significantly improved editing efficiency and expanded the target range; by adding a self-cleaving hepatitis D virus (HDV) ribozyme to the 3' end of the ωRNA, the editing efficiency was further improved by 1.79-63.02 times, and editing of all 8 tested endogenous sites was achieved; furthermore, through AlphaFold3 structural alignment and rational design, the key K416R gain mutation in the RuvC domain was identified, further improving the editing efficiency by an average of 42%.

[0059] This system retains the enNlovFz2's ability to recognize extended 5'-NMYG-3' TAM and has been successfully applied to stable transformation in rice, achieving editing efficiencies of 17.6%-93.3% at eight endogenous sites, with T0 generation biallelic or homozygous mutants accounting for 16.7%-50.0%. Furthermore, this application is the first to construct a Fanzor-based plant cytosine base editor (CBE), achieving C-to-T base conversion in soybean, with the editing window located at +2 to +13 bp positions in the spacer region. The highly efficient, ultra-compact Fanzor2 system established in this application provides a new tool option for plant genome editing.

[0060] In this application, we demonstrate that the enNlovFz2 system significantly improves genome editing efficiency in plants compared to the wild-type version. Through rational engineering of ωRNA and the enNlovFz2 protein, we developed the enNlovFz2(K416R) / ωRNA-Rz system. This system exhibited stable genome editing activity in both soybean and rice, successfully targeting sites that were difficult to edit in the wild-type version. Furthermore, a novel cytosine base editor developed based on this optimized system achieved highly efficient C-to-T base conversion in plants, with a maximum editing efficiency of 8.27%. These findings establish a multifunctional platform for plant genome editing and base editing, demonstrating that the engineered Fanzor system can serve as an effective tool for precise genetic modification of crop traits.

[0061] The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.

[0062] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0063] Vector Construction: In the following examples, the soybean-introduced backbone vector was obtained by modifying the 35S:Ruby backbone (Addgene #160908). The modified vector was named r35S:Ruby, and its nucleotide sequence is shown in SEQ ID NO:20. Specifically, positions 5647 to 6324 of SEQ ID NO:20 are the 35S enhanced promoter; positions 6357 to 6362, 8228 to 8233, and 9021 to 9026 of SEQ ID NO:20 are XhoI restriction enzyme recognition sites; positions 9027 to 9201 of SEQ ID NO:20 are the CaMV 3'UTR sequence; and positions 117 to 124 of SEQ ID NO:20 are PmeI restriction enzyme recognition sites.

[0064] The plant codon-optimized sequence of the NlovFz2 gene was synthesized by Sangon Biotech (Shanghai) Co., Ltd. Specific mutations were introduced into NlovFz2 via overlap PCR. In soybean hairy roots and rice, the ωRNA is driven by the AtU6 and OsU3 promoters, respectively, while the NlovFz2 gene is driven by the CaMV35S and maize Ubiquitin promoters, respectively. All DNA fragments were integrated into a linearized r35S:Ruby vector using seamless cloning (TransGen Biotech), and the final vector was verified by Sanger sequencing. For the base editing vector, the coding sequences for the nuclear localization signal (NLS), cytosine deaminase (mini-Sdd7), linker peptide, NlovFz2 (D294A), and uracil glycosylation inhibitor (UGI) were synthesized after codon optimization, cloned into the r35S:Ruby backbone, and subsequently inserted with the corresponding ωRNA target sequences.

[0065] Soybean hairy root transformation and stable rice transformation: The soybean variety used in this application is Williams 82, and the rice variety is Zhonghua 11. Transgenic soybean hairy roots were induced using Agrobacterium rhizogenes strain K599. Stable genetic transformation of rice was mediated by Agrobacterium EHA105. Positive T0 generation plants were screened for further analysis by PCR detection of the hygromycin phosphotransferase (HPT) gene and the enNlovFz2 (K416R) gene.

[0066] Mutation analysis: Mutation frequencies in soybean hairy roots were analyzed using amplicon-based deep sequencing. After genomic DNA extraction, the target region was amplified by PCR, and the amplicon was sequenced using the Hi-TOM platform at the State Key Laboratory of Rice Biology, China National Rice Research Institute. At least 5000 reads were obtained from each sample for analysis. The mutation frequency was defined as the percentage of reads containing insertions or deletions (indels) within the target sequence and its flanking 30 bp regions. Editing results of stable T0 generation rice lines were determined by direct sequencing of PCR products (Sangon Biotech (Shanghai) Co., Ltd.), and sequencing data were analyzed using CRISPR-GE's DSDecodeM software.

[0067] Protein structure alignment and analysis: The protein structure of enNlovFz2 was predicted using AlphaFold 3 (https: / / alphafoldserver.com / ). Cryo-electron microscopy structures and model coordinates of IsDra2 TnpB and ApmFz2 were obtained from the Protein Database (PDB), accession numbers 8EXA and 9B0L, respectively. Multiple sequence alignment and phylogenetic tree construction were performed using TBtools v2.012. Protein structure visualization and plotting were performed using PyMOL.

[0068] Unless otherwise specified, the quantitative experiments in the following examples were performed in triplicate, and the results were averaged.

[0069] Example 1. Comparison of editing activities of NlovFz2 and enNlovFz2 in soybean hairy roots To systematically evaluate the editing performance of NlovFz2 and enNlovFz2 in plants, we constructed an all-inone expression vector: the plant codon-optimized NlovFz2 or enNlovFz2 was driven by the CaMV35S promoter, and its corresponding ωRNA was driven by the Arabidopsis U6 (AtU6) promoter (Figure 1A). We selected eight endogenous gene loci in soybean as targets, including GmBADH2 (3 loci), GmCCD4a (1 locus), GmSweet15a (1 locus), and GmFAD2-1A (3 loci). The vectors were introduced into soybean hairy roots via Agrobacterium rhizogenes K599-mediated transformation, the target regions were amplified by PCR, and the editing efficiency was analyzed using high-throughput sequencing.

[0070] Table 1. Detailed information on the eight target sequences in Example 1

[0071] The primer sequences (5'-3') used for PCR amplification of the target region are as follows: GmBADH2-T1-F: gagtacggtgtgctttgacttgatgtgattggt; GmBADH2-T1-R: ggatgctggatgggatagcgagcacgaacagag; GmBADH2-T2-F: gagtacggtgtgctctggcaatcagtatttgtc; GmBADH2-T2-R: ggatgctggatggtatgctgaatatttgcactg; GmBADH2-T3-F: gagtacggtgtgcggacctgtgctatgtgtgaa; GmBADH2-T3-R: ggatgctggatggggggataaaaaataacgatg; GmCCD4a-T1-F: gagtacggtgtgcGATCCCAAAACCAATCTTAG; GmCCD4a-T1-R: ggatgctggatggAGCTTTTTGAGTGGTGGTGG; GmSweet15a-T1-F: gagtacggtgtgcccaagtgtcattgtcatatt; GmSweet15a-T1-R: ggatgctggatggggagaaacccatgttgaaaa; GmFAD2-1A-T1-F: gagtacggtgtgcCCCTTATTTCTCATGGAAAA; GmFAD2-1A-T1-R: ggatgctggatggCCCTTCCTAGAGGGTTGTTT; GmFAD2-1A-T2-F: gagtacggtgtgcCCACTACCACCCTTATGCTC; GmFAD2-1A-T2-R: ggatgctggatggGCAAAGGCACCCCATAAACA; GmFAD2-1A-T3-F: gagtacggtgtgcCCACTACCACCCTTATGCTC; GmFAD2-1A-T3-R:ggatgctggatggGCAAAGGCACCCCATAAACA.

[0072] The construction of the expression vector for the enNlovFz2 / ωRNA genome editing system is illustrated as an example: The DNA fragment shown in positions 782 to 2295 of SEQ ID NO:1 is used to replace the fragment between the XhoI restriction site and the restriction site of the r35S:Ruby backbone vector. The DNA fragment shown in SEQ ID NO:2 is then inserted into the PmeI restriction site of the r35S:Ruby backbone vector, while keeping the other nucleotide sequences of the r35S:Ruby backbone vector unchanged. This yields the enNlovFz2 / ωRNA expression vector backbone for soybean. In SEQ ID NO:1, positions 1 to 781 represent the nucleotide sequence of the CaMV 35S promoter, positions 802 to 804 are the start codon, positions 808 to 828 are the coding gene for SV40 NLS, positions 829 to 2295 are the coding gene for enNloVFz2, and positions 2296 to 2486 are the CaMV 3'UTR sequence.

[0073] Then, the spacer sequences targeting the eight endogenous sites in Table 1 were inserted into the backbone of the enNlovFz2 / ωRNA expression vector between the ωRNA scaffold and Poly T in SEQ ID NO:2 (between positions 441 and 442), respectively, to obtain the recombinant vectors enNlovFz2 / ωRNA-GmBADH2-T1, enNlovFz2 / ωRNA-GmBADH2-T2, enNlovFz2 / ωRNA-GmBADH2-T3, enNlovFz2 / ωRNA-GmCCD4a-T1, enNlovFz2 / ωRNA-GmSweet15a-T1, enNlovFz2 / ωRNA-GmFAD2-1A-T1, enNlovFz2 / ωRNA-GmFAD2-1A-T2, and enNlovFz2 / ωRNA-GmFAD2-1A-T3.

[0074] The expression vector for the wild-type NlovFz2 / ωRNA genome editing system was constructed according to the same method as that for the enNlovFz2 / ωRNA genome editing system, except that the DNA fragment shown in SEQ ID NO:1 was replaced with the DNA fragment shown in SEQ ID NO:3, and the DNA fragment shown in SEQ ID NO:2 was replaced with the DNA fragment shown in SEQ ID NO:4. All other operations were the same.

[0075] The results showed that the wild-type NlovFz2 / ωRNA system detected editing activity at only 2 of the 8 sites (GmSweet15a-T1 and GmBADH2-T3), with an efficiency of only 1.07%-1.23% (Figure 1, B). In contrast, the enNlovFz2 / ωRNA system significantly improved editing activity: the efficiency at GmSweet15a-T1 and GmBADH2-T3 sites was increased by 5.45-fold and 6.75-fold, respectively, and it successfully edited 4 sites that NlovFz2 could not edit (GmBADH2-T1, GmBADH2-T2, GmFAD2-1A-T1, and GmFAD2-1A-T2), but no editing activity was detected at 2 sites (GmFAD2-1A-T3 and GmCCD4a-T1) (Figure 1, B).

[0076] Example 2. Adding HDV ribozyme to the 3' end of ωRNA significantly improves editing efficiency. To further enhance the editing performance of the enNlovFz2 / ωRNA system, we added a self-cleaving HDV ribozyme to the 3' end of the ωRNA, constructing the enNlovFz2 / ωRNA-Rz system (Figure 1A). The HDV ribozyme can generate a precise 3' end through self-cleavage, stabilizing the ωRNA structure and inhibiting the degradation of endogenous RNase.

[0077] The expression vector construction of the enNlovFz2 / ωRNA-Rz genome editing system followed the method described in Example 1, with the only difference being that the DNA fragment shown in SEQ ID NO:2 was replaced with the DNA fragment shown in SEQ ID NO:5, and the spacer sequences targeting the eight endogenous sites in Table 1 were inserted into the enNlovFz2 / ωRNA-Rz genome editing system. Z The remaining operations were the same between ωRNAscaffold and HDV ribozyme in the expression vector backbone (between positions 437 and 438 of SEQ ID NO:5). The vector was introduced into soybean hairy roots via Agrobacterium rhizogenes K599-mediated transformation, and the editing efficiency was analyzed using high-throughput sequencing.

[0078] High-throughput sequencing results showed that this modification significantly enhanced editing activity: it not only enabled editing of the remaining two unedited sites (GmFAD2-1A-T3 and GmCCD4a-T1), but also improved editing efficiency by 1.79-63.02 times at the remaining six sites (Figure 1, B). Analysis of the edited products showed that enNlovFz2 / ωRNA-Rz mainly produced deletion mutations of 1-30 bp, with the deletion sites mainly concentrated near the 14th nucleotide of the spacer region (with the proximal end of TAM as the first position) (Figure 1, C and D).

[0079] Example 3. Rational design to identify key mutation K416R to further enhance editing activity Rational design and protein engineering are effective strategies for enhancing the activity of genome editing tools. Although enNlovFz2 has low amino acid sequence homology with ApmFz2 and IsDra2 TnpB ( Figure 2 However, AlphaFold3 structure prediction shows that they have highly similar three-dimensional structures, with root mean square deviations (RMSD) of 4.15 and 3.82 from ApmFz2 and IsDra2, respectively. Figure 3 Based on structural alignment, we mapped the previously reported genome editing efficiency-enhancing mutations in ApmFz2 and IsDra2 to enNlovFz2, constructing 24 single-point mutants (Table 2). These mutations are distributed across different functional domains of enNlovFz2. Figure 3 (B) We constructed the enNlovFz2 / ωRNA-Rz single-point mutant system shown in Table 2 based on the enNlovFz2 / ωRNA-Rz system, and evaluated the editing efficiency of these mutants at the GmSweet15a-T1 site in soybean hairy roots.

[0080] The results showed that the editing efficiency of 23 mutants decreased to varying degrees, with only the K416R mutation significantly improving editing activity, with an average efficiency increase of 42%. Figure 3 (B). We named this optimized system the enNlovFz2(K416R) / ωRNA-Rz system.

[0081] To verify the broad spectrum of the K416R mutation, we tested the editing performance of this system at seven endogenous soybean sites other than GmSweet15a-T1. The structure of the enNlovFz2 (K416R) / ωRNA-Rz backbone vector in the enNlovFz2 (K416R) / ωRNA-Rz system is as follows: the fragment between the XhoI restriction sites of the r35S:Ruby backbone vector is replaced with the DNA fragment shown in positions 782 to 2295 of SEQ ID NO:6, and the DNA fragment shown in SEQ ID NO:5 is inserted into the PmeI restriction site of the r35S:Ruby backbone vector, keeping the other nucleotide sequences of the r35S:Ruby backbone vector unchanged. This yields the enNlovFz2 (K416R) / ωRNA-Rz expression vector backbone for soybean. In SEQ ID NO:6, positions 1 to 781 are the nucleotide sequence of the CaMV 35S promoter, positions 802 to 804 are the start codon, positions 808 to 828 are the coding gene for SV40 NLS, positions 829 to 2295 are the coding gene for enNloVFz2(K416R), and positions 2296 to 2486 are the CaMV 3'UTR sequence.

[0082] Then, the spacer sequences targeting the seven endogenous sites in Table 1 were inserted into the backbone of the enNlovFz2 (K416R) / ωRNA-Rz expression vector between the ωRNA scaffold and HDV ribozyme in SEQ ID NO:5 (between positions 437 and 438 of SEQ ID NO:5), respectively, to obtain the recombinant vectors enNlovFz2 (K416R) / ωRNA-Rz-GmBADH2-T1, enNlovFz2 (K416R) / ωRNA-Rz-GmBADH2-T2, enNlovFz2 (K416R) / ωRNA-Rz-GmBADH2-T3, enNlovFz2 (K416R) / ωRNA-RzGmCCD4a-T1, enNlovFz2 (K416R) / ωRNA-Rz-GmFAD2-1A-T1, and enNlovFz2 (K416R) / ωRNA-Rz-GmFAD2-1A-T2 and enNlovFz2 (K416R) / ωRNA-Rz-GmFAD2-1A-T3.

[0083] Recombinant vector enNlovFz2 (K416R) / ωRNA-Rz-GmBADH2-T1, enNlovFz2 (K416R) / ωRNA-Rz-GmBADH2-T2, enNlovFz2 (K416R) / ωRNA-Rz-GmBADH2-T3, enNlovFz2 (K416R) / ωRNA-RzGmCCD4a-T1, enNlovFz2 (K416R) / ωRNA-Rz-GmFAD2-1A-T1, enNlovFz2 (K416R) / ωRNA-Rz-GmFAD2-1A-T2 and enNlovFz2 Both (K416R) / ωRNA-Rz-GmFAD2-1A-T3 can express the protein and nucleic acid components of the enNlovFz2(K416R) / ωRNA-Rz system. The protein component is a fusion protein with the amino acid sequence shown in SEQ ID NO:7, where positions 3 to 9 of SEQ ID NO:7 are SV40 NLS sequences and positions 10 to 497 are amino acid sequences of enNlovFz2(K416R). The nucleic acid component is ωRNA targeting the above-mentioned target sites, wherein the ωRNA contains a ωRNA scaffold and a guide sequence capable of hybridizing with the target sequence from the 5' to 3' direction, and the nucleotide sequence of the ωRNA scaffold is shown in SEQ ID NO:8.

[0084] The vector was introduced into soybean hairy roots via Agrobacterium rhizogenes K599-mediated transformation, and the editing efficiency was analyzed using high-throughput sequencing.

[0085] The results showed that, except for GmFAD2-1A-T1, the editing efficiency of the other 6 sites was significantly higher than that of the wild-type NlovFz2 / ωRNA system, with an average improvement of 5%-110%. Figure 3 (C). We also observed that even within the same gene, the editing efficiency of different targets varied, indicating that this system, like most members of the Cas12 family, is target-dependent.

[0086] Table 2. NlovFz2 mutations correspond to gain-of-function mutations in IsDra2 TnpB and ApmFz2, respectively.

[0087] Example 4. Extension of TAM recognition range for enNlovFz2 (K416R) / ωRNA-Rz Previous studies have shown that engineered enNlovFz2 can recognize extended 5'-NMYG-3' TAM sequences in mammalian cells. To verify whether this characteristic is retained in plants, we selected three soybean genes—GmPDS, GmSMS6, and GmCKX3—and designed 18 target sites (6 ATG, 6 ACG, and 6 CTG) carrying 5'-NMYG-3' TAM for editing efficiency evaluation. The operation method was the same as the enNlovFz2 (K416R) / ωRNA-Rz system in Example 3, with the only difference being the spacer sequence introduced into the enNlovFz2 (K416R) / ωRNA-Rz expression vector backbone.

[0088] Table 3. Targets with expanded TAM recognition range for enNlovFz2 (K416R) / ωRNA-Rz

[0089] The PCR primer sequences (5' to 3') for amplifying the target region are as follows: ATGPAM-T1-F:gagtacggtgtgcGGATTGGCTGGTTTATCAACTGC; ATGPAM-T1-R:ggatgctggatggGGCGCACTAAGTGACAACTT; ATGPAM-T2-F:gagtacggtgtgcAGTTGGGGCTTACCCTAATG; ATGPAM-T2-R:ggatgctggatggTCGGGAAAATCAAATCGACT; ATGPAM-T3-F:gagtacggtgtgcTCCACGCGAAGAGTTCGTGT; ATGPAM-T3-R:ggatgctggatggTACCTGTGGGTAGTCTCCCTT; ATGPAM-T4-F:gagtacggtgtgcCCTCCAACTCACCCTATCAG; ATGPAM-T4-R:ggatgctggatggCTGTTTTGCAAGGCTGCATG; ATGPAM-T5-F:gagtacggtgtgcTAGTAACCATAACCCGTTTG; ATGPAM-T5-R:ggatgctggatggCCTTGGCCCCTCGCAGCTAT; ATGPAM-T6-F:gagtacggtgtgcTAGTAACCATAACCCGTTTG; ATGPAM-T6-R:ggatgctggatggCCTTGGCCCCTCGCAGCTAT; CTGPAM-T1-F:gagtacggtgtgcATATTTGGCTGATGCTGGGC; CTGPAM-T1-R:ggatgctggatggTTCTTATTTCATTAAACAGC; CTGPAM-T2-F:gagtacggtgtgcTACCACAGTATATACAACAT; CTGPAM-T2-R:ggatgctggatggAAAAAGAGTTAAACCCAAGA; CTGPAM-T3-F:gagtacggtgtgcTCCACGCGAAGAGTTCGTGT; CTGPAM-T3-R:ggatgctggatggTACCTGTGGTAGTCTCCCTT; CTGPAM-T4-F:gagtacggtgtgcGACATTGCTAACACCGAGCT; CTGPAM-T4-R:ggatgctggatggCTGTTTTGCAAGGCTGCATG; CTGPAM-T5-F:gagtacggtgtgcTCTGCTTTTACCAAAGACCA; CTGPAM-T5-R:ggatgctggatggAGTTATTAAAGAAGCTATTC; CTGPAM-T6-F:gagtacggtgtgcATTTGTACTTGACCGTGGGA; CTGPAM-T6-R:ggatgctggatggCAGTGATGACATCCATTTCA; ACGPAM-T1-F:gagtacggtgtgcGGATTGGCTGGTTTATCAACTGC; ACGPAM-T1-R:gagttggatgctggatggCAGCATGCTAAAATAATGAGC; ACGPAM-T2-F:gagtacggtgtgcAGTTGGGGCTTACCCTAATG; ACGPAM-T2-R:ggatgctggatggTCGGGAAAATCAAATCGACT; ACGPAM-T3-F:gagtacggtgtgcCCTCCATTGAGCAGAAGGAG; ACGPAM-T3-R:gagtacggtgtgcCCTGGCGCATCATCTCCTC; ACGPAM-T4-F:gagtacggtgtgcaactttgatattcactagGC; ACGPAM-T4-R:ggatgctggatggCTGCATGTCAGAGGTCCAGA; ACGPAM-T5-F:gagtacggtgtgcTAGTAACCATAACCCGTTTG; ACGPAM-T5-R:ggatgctggatggCCATGGCTTGTCCATGAGTG; ACGPAM-T6-F:gagtacggtgtgcATTTGTACTTGACCGTGGGA; ACGPAM-T6-R:ggatgctggatggCAGTGATGACATCCATTTCA.

[0090] The results showed that editing activity was detected at 4 out of 18 target sites. Figure 4 One target carrying CTG TAM showed high editing activity, with an average efficiency of 6.3% and a maximum of 17.9%; the other three targets carrying ATG or ACG TAM showed lower editing activity. These results indicate that the enNlovFz2 (K416R) / ωRNA-Rz system can indeed recognize the extended 5'-NMYG-3' TAM sequence in plant cells, further expanding its targeting range.

[0091] Example 5. Construction and Activity Verification of the First Fanzor-Based Plant Cytosine Base Editor Base editors enable precise single-base conversions without double-strand breaks, making them important tools for precise genome editing. Currently, no plant base editors based on Fanzor have been reported. To explore the application potential of Fanzor in base editing, we introduced the D294A mutation into the RuvC domain of enNlovFz2 (K416R) to obtain a catalytically inactivated denNlovFz2 (K416R) variant. Subsequently, we fused denNlovFz2 (K416R) with a mini-Sdd7 cytosine deaminase and a uracil glycosylase inhibitor (UGI) to construct the cytosine base editor denNlovFz2 (K416R) / ωRNA-Rz / CBE (…). Figure 5 The fusion protein (GmSweet15a-T1, GmBADH2-T1, and GmBADH2-T3) has a total length of 788 amino acids. We evaluated its base editing activity at three cytosine-rich sites in the spacer region: GmSweet15a-T1, GmBADH2-T1, and GmBADH2-T3.

[0092] The structure of the expression vector backbone of the cytosine base editor denNlovFz2 (K416R) / ωRNA-Rz / CBE is as follows: the fragment between the XhoI restriction sites of the r35S:Ruby backbone vector is replaced with the DNA fragment shown at positions 782 to 3168 of SEQ ID NO:9, and the DNA fragment shown in SEQ ID NO:5 is inserted at the PmeI restriction site of the r35S:Ruby backbone vector, while keeping the other nucleotide sequences of the r35S:Ruby backbone vector unchanged, thus obtaining the denNlovFz2 (K416R) / ωRNA-Rz / CBE expression vector backbone for soybean. In SEQ ID NO:9, positions 1 to 781 are the nucleotide sequence of the CaMV35S promoter, positions 802 to 804 are the start codon, positions 805 to 825 are the coding gene for SV40 NLS, positions 826 to 1350 are the coding gene for the deaminase mini Sdd7, positions 1351 to 2814 are the coding gene for denNlovFz2(K416R), positions 2815 to 2862 are the Nucleoplasmin NLS sequence, positions 2863 to 3144 are the coding gene for UGI, positions 3145 to 3165 are the coding gene for SV40 NLS, and positions 3169 to 3359 are the CaMV 3'UTR sequence.

[0093] Then, the spacer sequences targeting GmSweet15a-T1, GmBADH2-T1, and GmBADH2-T3 in Table 1 were inserted into the backbone of the denNlovFz2 (K416R) / ωRNA-Rz / CBE expression vector between ωRNAscaffold and HDV ribozyme in SEQ ID NO:5 (between positions 437 and 438 of SEQ ID NO:5), respectively, to obtain the recombinant vectors denNlovFz2 (K416R) / ωRNA-Rz / CBE-GmSweet15a-T1, denNlovFz2 (K416R) / ωRNA-Rz / CBE-GmBADH2-T1, and denNlovFz2 (K416R) / ωRNA-Rz / CBE-GmBADH2-T3.

[0094] The recombinant vectors denNlovFz2 (K416R) / ωRNA-Rz / CBE-GmSweet15a-T1, denNlovFz2(K416R) / ωRNA-Rz / CBE-GmBADH2-T1, and denNlovFz2 (K416R) / ωRNA-Rz / CBE-GmBADH2-T3 can express the protein and nucleic acid components of the denNlovFz2 (K416R) / ωRNA-Rz / CBE system. The protein component is a fusion protein with an amino acid sequence as shown in SEQ ID NO:10. In SEQ ID NO:10, positions 1 to 8 are the SV40 NLS sequence, positions 9 to 183 are the amino acid sequence of mini Sdd7 deaminase, positions 184 to 671 are the amino acid sequence of denNlovFz2 (D294A), positions 672 to 687 are the Nucleoplasmin NLS sequence, positions 695 to 777 are the amino acid sequence of uracil glycosylase inhibitor UGI, and positions 782 to 788 are the SV40 NLS sequence. The nucleic acid component is gRNA targeting the above targets respectively. The gRNA contains ωRNA scaffold and guide sequence capable of hybridizing with the target sequence from the 5' to 3' direction. The nucleotide sequence of ωRNA scaffold is shown in SEQ ID NO:8.

[0095] The results showed that the editing window of this CBE system was mainly located in the interval region from +2 to +13 bp, with an average editing efficiency of 0.94%-1.81%. The highest C-to-T editing efficiency of 8.27% was observed at the GmSweet15a-T1 site. Figure 5 (B and C). We also validated the results at other targets, demonstrating that further improvements in base editing efficiency are necessary.

[0096] Example 6. Efficient editing of enNlovFz2 (K416R) / ωRNA-Rz in stable transgenic rice lines To evaluate the editing performance of this system in monocotyledonous plants, we constructed an expression vector enNlovFz2 (K416R) / ωRNA-Rz suitable for rice. The structure of the rice expression vector enNlovFz2 (K416R) / ωRNA-Rz is as follows: SEQ ID NO:11 was cloned into the BamHI restriction site of the pCXUN vector using homologous recombination (Sun Y, Zhang X, Wu C, et al. Engineering herbicide-resistant rice plants through CRISPR / Cas9-mediated homologous recombination of acetolactate synthase[J].Molecular plant, 2016, 9(4): 628-631.), and the DNA fragment shown in SEQ ID NO:12 was cloned into the PmeI site of the above vector using homologous recombination back, keeping the other nucleotide sequences of the pCXUN backbone vector unchanged, thus obtaining the rice expression vector backbone enNlovFz2 (K416R) / ωRNA-Rz. The nucleotide sequence of the expression vector backbone for rice, enNlovFz2(K416R) / ωRNA-Rz, is shown in SEQ ID NO:21.

[0097] In SEQ ID NO:11, positions 1 to 1991 are the Maize Ubi promoter, positions 2018 to 2038 are the SV40NLS sequence, positions 2039 to 3505 are the coding gene for enNlovFz2(K416R), and positions 3530 to 3782 are the NOSterminator sequence. In SEQ ID NO:12, positions 1 to 437 are the OsU3 promoter sequence, positions 438 to 576 are the ωRNAscaffold sequence, positions 577 to 644 are the HDV ribozyme sequence, and positions 645 to 651 are the Poly T sequence.

[0098] enNlovFz2 (K416R) is driven by the maize ubiquitin (Ubi) promoter, and ωRNA-Rz is driven by the rice U3 (OsU3) promoter. Figure 6(A). We selected eight endogenous gene loci in rice for stable transformation, including Os03g0151800 (2 loci), OsNRT1.1B (1 locus), OsNYC1 (2 loci), Os11g0508600 (1 locus), OsDEP1 (1 locus), and OsTHF1 (1 locus). The recombinant vector can express both protein and nucleic acid components of the enNlovFz2 (K416R) / ωRNA-Rz system, which can be expressed in rice. The protein component is a fusion protein with an amino acid sequence as shown in SEQ ID NO:7, where positions 3 to 9 of SEQ ID NO:7 are SV40 NLS sequences and positions 10 to 497 are amino acid sequences of enNlovFz2 (K416R). The nucleic acid component is ωRNA targeting the above-mentioned target sites, wherein the ωRNA contains a ωRNA scaffold and a guide sequence capable of hybridizing with the target sequence from the 5' to 3' direction, and the nucleotide sequence of the ωRNA scaffold is shown in SEQ ID NO:7.

[0099] Transgenic seedlings were obtained by transforming rice callus tissue into a vector using Agrobacterium EHA105, PCR amplification of the target region, and high-throughput sequencing analysis of the editing efficiency.

[0100] Table 4 Targets in Rice

[0101] The primer sequences (5'-3') used for PCR amplification of the target region are as follows: Os03g0151800-T1-F:gagtacggtgtgcCACAATCACTTCTCCCCCCC; Os03g0151800-T1-R:ggatgctggatggCTACAAAACAGTAGACACAC; Os03g0151800-T2-F:gagtacggtgtgcGAGAGCCCTCCCTCCTC; Os03g0151800-T2-R:ggatgctggatggCTAGAAAGTAGCAAACATTC; OsNRT1.1B-T1-F:gagtacggtgtgcACCTCCTTCATGCTCTGCCT; OsNRT1.1B-T1-R:ggatgctggatggGAGCACCCCGAGCTGCGTCC; OsNYC1-T1-F:gagtacggtgtgcAAGCCGTAGGACGCCGGAGTGC; OsNYC1-T1-R:ggatgctggatggGCGGCGGCGACGACGAAGCCTC; OsNYC1-T2-F:gagtacggtgtgcAGTCTCCACGCCCGGCCATAC; OsNYC1-T2-R:ggatgctggatggGCCTCCCGATACGTGGCAAC; Os11g0508600-T1-F:gagtacggtgtgccctcattgatctcctcccac; Os11g0508600-T1-R:ggatgctggatggtaatcttgcatggttattta; OsDEP1-T1-F:gagtacggtgtgcggtcctcctcatcgcatcgc; OsDEP1-T1-R:ggatgctggatggagaaccacccctcgccgcct; OsDEP1-T2-F:gagtacggtgtgcGGCGGCCATATCTTCGCTTC; OsDEP1-T2-R:ggatgctggatggATTCATCTTTGTTTCTGCGA.

[0102] The results showed that enNlovFz2 (K416R) / ωRNA-Rz successfully induced targeted mutations at all 8 sites, with editing efficiencies ranging from 17.6% to 93.3%. Figure 6 (Middle B). Among them, the editing efficiency of the Os11g0508600-T1 site was the highest, with 28 out of 30 transgenic plants successfully edited, achieving an editing efficiency of 93.3%. Of the 28 edited plants, 15 were biallelic or homozygous mutations. Figure 6 (B) The rice THYLAKOID FORMATION1 gene (OsTHF1) is involved in regulating chlorophyll degradation during leaf senescence. We performed dark-induced leaf senescence phenotypic analysis on T0 generation OsTHF1 homozygous and biallelic knockout mutants. The results showed that the mutant leaves exhibited a distinct chlorophyll-holding phenotype, which contrasted sharply with the wild type. Figure 6 (C). This result further demonstrates the potential of this system for genetic improvement in monocotyledonous plants.

[0103] Example 7. Specificity analysis (off-target effect analysis) of the enNlovFz2 (K416R) / ωRNA-Rz system To evaluate the specificity of the enNlovFz2 (K416R) / ωRNA-Rz system, we selected the three most efficient editing sites in soybean and rice and analyzed their potential off-target sites. The results showed that no off-target editing activity was detected in the three soybean target sites (Table 5). However, in rice, the OsNRT1.1B-T1 site exhibited significant off-target activity: its two potential off-target sites had only one nucleotide mismatch, and off-target mutations were detected in 7 and 6 sequenced plants out of 15 tested plants, respectively; while off-target mutations were detected in 3 plants with two nucleotide mismatches (Table 5). These results suggest that careful target selection and off-target effect assessment are necessary in subsequent experimental designs.

[0104] Table 5. Analysis of potential off-target effects

[0105] Note: PAM sequences are identified by underscores; lowercase letters represent mismatched bases. The sequences in this application are as follows: SEQ ID NO:1, for soybean enNlovFz2 expression cassette CaMV 35S promoter-SV40 NLS-enNloVFz2-CaMV 3'UTR:

[0106] SEQ ID NO: 2, the wild-type ωRNA scaffold expression cassette AtU6 - wild type ωRNA scaffold - Poly T for soybean: TTTCCATTCGGAGTTTTTGTATCTTGTTTCATAGTTTGTCCCAGGATTAGAATGATTAGGCATCGAACCTTCAAGAATTTGATTGAATAAAACATCTTCATTCTTAAGATATGAAGATAATCTTCAAAAGGCCCCTGGGAATCTGAAAGAAGAGAAGCAGGCCCATTTATATGGGAAAGAACAATAGTATTTCTTATATAGGCCCATTTAAGTTGAAAACAATCTTCAAAAGTCCCACATCGCTTAGATAAGAAAACGAAGCTGAGTTTATATACAGCTAGAGTCGAAGTAGTGATTGGAGCACCTTGTGTGTTGGGTCTTCCCCACCTTGTGTGCGTTGGGTCCTTTCCCCTGGCTTTCACTCTTTGAGTGTTTGCTTGCAAGATCAGCACTTTGTTCGATTTGTTTGTTTTGTTCAAATTGTTGATTTAAATAAAGGCGCAGTCTTCTCAGAAGACGTTTTTTT.

[0107] SEQ ID NO: 3, the enNlovFz2 expression cassette for soybean CaMV 35S promoter - SV40 NLS - NloVFz2 - CaMV 3'UTR:

[0108] SEQ ID NO:4, for soybean engineered ωRNA scaffold expression cassette AtU6 - Engineered ωRNA scaffold - Poly T: TTTCCATTCGGAGTTTTTGTATCTTGTTTCATAGTTTGTCCCAGGATTAGAATGATTAGGCATCGAACCTTCAAGAATTTGATTGAATAAAACATCTTCATTCTTAAGATATGAAGATAATCTTCAAAAGGCCCCTGGGAATCTGAAAGAAGAGAAGCAGGCCCATTTATATGGGAAAGAACAATAGTATTTCTTATATAGGCCCATTTAAGTTGAAAACAATCTTCAAAAGTCCCACATCGCTTAGATAAGAAAACGAAGCTGAGTTTATATACAGCTAGAGTCGAAGTAGTGATTGGAGCACCTTGTGTGTTGGGTCTTCCCCACCTTGTGTGCGTTGGGTCCTTTCCCCTGGCTTTCACTCTTTGAGTGTTTGCTTGCAAGATCAGCACTTCGTTCGATGTTTGTTGTTCGAATTGTTGATTTAAATAAAGGCGTTTTTTT。

[0109] SEQ ID NO:5, for improved ωRNA scaffold expression cassette AtU6 - ωRNA scaffold - HDV ribozyme - Poly T in soybean: TTTCCATTCGGAGTTTTTGTATCTTGTTTCATAGTTTGTCCCAGGATTAGAATGATTAGGCATCGAACCTTCAAGAATTTGATTGAATAAAACATCTTCATTCTTAAGATATGAAGATAATCTTCAAAAGGCCCCTGGGAATCTGAAAGAAGAGAAGCAGGCCCATTTATATGGGAAAGAACAATAGTATTTCTTATATAGGCCCATTTAAGTTGAAAACAATCTTCAAAAGTCCCACATCGCTTAGATAAGAAAACGAAGCTGAGTTTATATACAGCTAGAGTCGAAGTAGTGATTGGAGCACCTTGTGTGTTGGGTCTTCCCCACCTTGTGTGCGTTGGGTCCTTTCCCCTGGCTTTCACTCTTTGAGTGTTTGCTTGCAAGATCAGCACTTCGTTCGATGTTTGTTGTTCGAATTGTTGATTTAAATAAAGGCGggccggcatggtcccagcctcctcgctggcgccggctgggcaacatgcttcggcatggcgaatgggacTTTTTTT。

[0110] SEQ ID NO:6, enNlovFz2 (K416R) expression cassette for soybean CaMV 35S promoter - SV40NLS - enNlovFz2 (K416R) - CaMV 3'UTR:

[0111] SEQ ID NO: 7, SV40 NLS-enNlovFz2 (K416R) fusion protein applicable to soybean: MAPKKKRKVEPTHPPTNPSLAHGIIPFWDEYSQQVSDKLWACSRDSFHEFNQYNNKGCTDGWFNFSQFTVIESQPVFDVPLNVHHSITENVAFDNSKKPPQLKKAKKGQKTPQKFQADKSMKIRLYPNEQERTTLNQWMGTARWIYNKCLEFTNKSKGVKKNKKNFRTFVVNNDNYQTENQWVVNTPYDVRDAAAIELLTAFNTNFEKKKAGTIDKFMIRFRRKKDRKDHFVLRCKHWKKKSGMYSFIRNIKSAEPLPEELQYDSIIIKNKLNHYYLCIPQVLDIRGENQAPRHSGQVVALDPGVRTFQTTFDLNGYSTKWGSGGAERIGRLCCAYDKLQSKWSQPEVRHCKRYKYKRAGRRIQQKIRNIVDDLHKKLCLWLCRNYQVILLPSFETQKMVKKLHRRINSKTARKMLTWSHYRFRQRLLHKAREHPWTHIYIVNEAYTSKTCSCCGHVYTVGSSEVFRCPSCGSIFDRDINGARNILLRFLTTHRISF*。

[0112] SEQ ID NO: 8, ωRNA scaffold (RNA): GAGCACCTTGTGTGTTGGGTCTTCCCCACCTTGTGTGCGTTGGGTCCTTTCCCCTGGCTTTCACTCTTTGAGTGTTTGCTTGCAAGATCAGCACTTCGTTCGATGTTTGTTGTTCGAATTGTTGATTTAAATAAAGGCG。

[0113] SEQ ID NO:9, soybean denNlovFz2(K416R) / ωRNA-Rz / CBE expression cassette, promoter-SV40NLS-Cytidine deaminases (mini Sdd7), -denNlovFz2 (D294A)-Nucleoplasmin NLS-Uracil glycosylase inhibitor (UGI)-SV40NLS-CaMV 3'UTR:

[0114] SEQ ID NO:10, protein fraction of denNlovFz2(K416R) / ωRNA-Rz / CBE in soybean, SV40 NLS-Cytidine deaminases (mini Sdd7), -denNlovFz2 (D294A)-Nucleoplasmin NLS-Uracil glycosylase inhibitor (UGI)-SV40NLS: MPKKKRKVEGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWSKSGSETPGTSESATPERPEPTHPPTNPSLAHGIIPFWDEYSQQVSDKLWACSRDSFHEFNQYNNKGCTDGWFNFSQFTVIESQPVFDVPLNVHHSITENVAFDNSKKPPQLKKAKKGQKTPQKFQADKSMKIRLYPNEQERTTLNQWMGTARWIYNKCLEFTNKSKGVKKNKKNFRTFVVNNDNYQTENQWVVNTPYDVRDAAAIELLTAFNTNFEKKKAGTIDKFMIRFRRKKDRKDHFVLRCKHWKKKSGMYSFIRNIKSAEPLPEELQYDSIIIKNKLNHYYLCIPQVLDIRGENQAPRHSGQVVALAPGVRTFQTTFDLNGYSTKWGSGGAERIGRLCCAYDKLQSKWSQPEVRHCKRYKYKRAGRRIQQKIRNIVDDLHKKLCLWLCRNYQVILLPSFETQKMVKKLHRRINSKTARKMLTWSHYRFRQRLLHKAREHPWTHIYIVNEAYTSKTCSCCGHVYTVGSSEVFRCPSCGSIFDRDINGARNILLRFLTTHRISFKRPAATKKAGQAKKKKTRDSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRKV。

[0115] SEQ ID NO:11, enNlovFz2 expression cassette for rice Maize Ubipromoter - SV40 NLS - enNlovFz2(K416R) - NOS terminator :

[0116] SEQ ID NO: 12, Improved ωRNA scaffold expression cassette OsU3-ωRNA scaffold-HDV ribozyme-Poly T for rice: gtaattcatccaggtctccaagttctaggattttcagaactgcaacttattttatcaaggaatctttaaacatacgaacagatcacttaaagttcttctgaagcaacttaaagttatcaggcatgcatggatcttggaggaatcagatgtgcagtcagggaccatagcacaagacaggcgtcttctactggtgctaccagcaaatgctggaagccgggaacactgggtacgttggaaaccacgtgatgtgaagaagtaagataaactgtaggagaaaagcatttcgtagtgggccatgaagcctttcaggacatgtattgcagtatgggccggcccattacgcaattggacgacaacaaagactagtattagtaccacctcggctatccacatagatcaaagctgatttaaaagagttgtgcagatgatccgtggcaGAGCACCTTGTGTGTTGGGTCTTCCCCACCTTGTGTGCGTTGGGTCCTTTCCCCTGGCTTTCACTCTTTGAGTGTTTGCTTGCAAGATCAGCACTTCGTTCGATGTTTGTTGTTCGAATTGTTGATTTAAATAAAGGCGggccggcatggtcccagcctcctcgctggcgccggctgggcaacatgcttcggcatggcgaatgggacTTTTTTT。

[0117] SEQ ID NO: 13, Amino acid sequence of wild-type NlovFz2: MEPTHPPTNPSLAHGIIPFWDEYSQQVSDKLWACSRDSFHEFNQYNNKGCTDGWFNFSQFTVIESQPVFDVPLNVHHSITENVAFDNSKKPPQLKKAKKGQKTPQKFQADKSMKIRLYPNEQERTTLNQWMGTARWIYNKCLEFTNKSKGVKKNKKNFRTFVVNNDNYQTENQWVVNTPYDVRDAAAIELLTAFNTNFEKKKAGTIDKFMIRFRRKKDRKDHFVLRCKHWKKKSGMYSFIRNIKSAEPLPEELQYDSIIIKNKLNHYYLCIPQVLDIRGENQAPQHSGQVVALDPGVRTFQTTFDLNGYSTKWGSGGAERIGRLCCAYDKLQSKWSQPEVRHCKRYKYKRAGRRIQQKIRNIVDDLHKKLCLWLCRNYQVILLPSFETQKMVKKLHRRINSKTARKMLTWSHYRFKQRLLHKAREHPWTHIYIVNEAYTSKTCSCCGHVYTVGSSEVFRCPSCGSIFDRDINGARNILLRFLTTHRISF*。

[0118] SEQ ID NO:14, Amino acid sequence of enNLovFz2: MEPTHPPTNPSLAHGIIPFWDEYSQQVSDKLWACSRDSFHEFNQYNNKGCTDGWFNFSQFTVIESQPVFDVPLNVHHSITENVAFDNSKKPPQLKKAKKGQKTPQKFQADKSMKIRLYPNEQERTTLNQWMGTARWIYNKCLEFTNKSKGVKKNKKNFRTFVVNNDNYQTENQWVVNTPYDVRDAAAIELLTAFNTNFEKKKAGTIDKFMIRFRRKKDRKDHFVLRCKHWKKKSGMYSFIRNIKSAEPLPEELQYDSIIIKNKLNHYYLCIPQVLDIRGENQAPRHSGQVVALDPGVRTFQTTFDLNGYSTKWGSGGAERIGRLCCAYDKLQSKWSQPEVRHCKRYKYKRAGRRIQQKIRNIVDDLHKKLCLWLCRNYQVILLPSFETQKMVKKLHRRINSKTARKMLTWSHYRFKQRLLHKAREHPWTHIYIVNEAYTSKTCSCCGHVYTVGSSEVFRCPSCGSIFDRDINGARNILLRFLTTHRISF*。

[0119] Amino acid sequence of SEQ ID NO:15, enNLovFz2(K416R): MEPTHPPTNPSLAHGIIPFWDEYSQQVSDKLWACSRDSFHEFNQYNNKGCTDGWFNFSQFTVIESQPVFDVPLNVHHSITENVAFDNSKKPPQLKKAKKGQKTPQKFQADKSMKIRLYPNEQERTTLNQWMGTARWIYNKCLEFTNKSKGVKKNKKNFRTFVVNNDNYQTENQWVVNTPYDVRDAAAIELLTAFNTNFEKKKAGTIDKFMIRFRRKKDRKDHFVLRCKHWKKKSGMYSFIRNIKSAEPLPEELQYDSIIIKNKLNHYYLCIPQVLDIRGENQAPRHSGQVVALDPGVRTFQTTFDLNGYSTKWGSGGAERIGRLCCAYDKLQSKWSQPEVRHCKRYKYKRAGRRIQQKIRNIVDDLHKKLCLWLCRNYQVILLPSFETQKMVKKLHRRINSKTARKMLTWSHYRFRQRLLHKAREHPWTHIYIVNEAYTSKTCSCCGHVYTVGSSEVFRCPSCGSIFDRDINGARNILLRFLTTHRISF*。

[0120] SEQ ID NO:16 (identical to positions 3 to 9 of SEQ ID NO:7), SV40 NLS: PKKKRKV.

[0121] SEQ ID NO:17 (identical to positions 672 to 687 of SEQ ID NO:10), nucleoplasmin nuclear localization signal: KRPAATKKAGQAKKKK.

[0122] SEQ ID NO:18 (identical to positions 9 to 183 of SEQ ID NO:10), cytosine deaminase (mini-Sdd7): EGGGPGAVPEGGDGPPAVPAEEVERLRGELPPPVVPGTGQKTHGRWIGPDGRVRAIVSGRDEDAALVHAQLAAKGIPDEPTRNSDVEQKLAAHMVANGIRHVTLVINHRPCRGFDDSCDTLVPIILPEGCTLTVHGQTDKGMRVRVRYTGGARPWWSKSGSETPGTSESATPERP.

[0123] SEQ ID NO:19 (same as positions 695 to 777 of SEQ ID NO:10), Uracil glycosylation enzyme inhibitor (UGI): TNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML.

[0124] The present application has been described in detail above. Those skilled in the art will recognize that the present application can be implemented in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments are given in this application, it should be understood that further modifications can be made to the present application. In summary, in accordance with the principles of this application, this application is intended to include any changes, uses, or improvements to the present application, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. A protein, characterized in that, The protein is selected from at least one of the following (A1) to (A4): (A1) Proteins with amino acid sequences such as positions 10 to 497 of SEQ ID NO:7; Proteins with at least 70% identity to protein A1 obtained by substituting, deleting and / or adding amino acid residues of the amino acid sequences shown in (A2) and (A1) and having RNA-guided endonuclease function. (A3) A fusion protein comprising the protein described in (A1) or (A2) and other proteins or polypeptides; (A4) A conjugate comprising the protein described in (A1) or (A2) and a modified portion; Wherein, the substitution described in A2) is D294A.

2. The protein according to claim 1, characterized in that, The other proteins or polypeptides described in (A3) are selected from epitope tags, reporter genes, nuclear localization signals, deaminases, transcription activation domains, transcription repression domains, nuclease domains, reverse transcriptases, and any combination thereof; And / or, the modification portion described in (A4) is selected from the other proteins or peptides, detectable markers, and any combination thereof.

3. The protein according to claim 2, characterized in that, (A3) The fusion protein has the characteristics described in (i) or (ii) below: (i) The fusion protein contains an NLS sequence and / or an epitope tag; (ii) In the fusion protein, the additional protein or polypeptide is linked to the N-terminus or C-terminus of the protein via peptide bonds or linkers; And / or, the conjugate described in (A4) has the following characteristics as described in (i) or (ii): (i) The conjugate contains an NLS sequence and / or an epitope tag; (ii) In the conjugate, the modified portion is connected to the N-terminus or C-terminus of the protein via a linker, or the modified portion is fused to the N-terminus or C-terminus of the protein.

4. The protein according to claim 2 or 3, characterized in that, (A3) The fusion protein is at least one of the following: (B1) A fusion protein with an amino acid sequence as shown in SEQ ID NO:15; (B2) A fusion protein with the amino acid sequence shown in SEQ ID NO:7; (B3) A fusion protein with an amino acid sequence as shown in SEQ ID NO:

10.

5. A composition and / or complex for nucleic acid editing, characterized in that, The composition or / and complex comprises: (i) a protein component, said protein component being a protein as described in at least one of claims 1 to 4; (ii) A nucleic acid component comprising, from 5' to 3', an ωRNA scaffold and a guide sequence capable of hybridizing with a target sequence, wherein the nucleotide sequence of the ωRNA scaffold is shown in SEQ ID NO:

8.

6. A biomaterial, characterized in that, The biomaterial is selected from at least one of the following: (C1) A DNA molecule encoding the protein of any one of claims 1 to 4; (C2) An expression cassette containing the DNA molecule described in (C1); (C3) A recombinant vector containing the expression cassette described in (C2) and / or the expression cassette that transcribes the nucleic acid components of claim 5; (C4) Recombinant cells containing the expression cassette described in (C2) and / or the expression cassette that transcribes the nucleic acid components described in claim 5.

7. The biomaterial according to claim 6, characterized in that, (C1) The DNA molecule is at least one of the following: (C1-1) DNA molecules with nucleotide sequences such as positions 829 to 2295 of SEQ ID NO:6; (C1-2) DNA molecules with nucleotide sequences such as positions 808 to 2295 of SEQ ID NO:6; (C1-3) DNA molecules with nucleotide sequences such as positions 802 to 2295 of SEQ ID NO:6; (C1-4) DNA molecules with nucleotide sequences such as positions 1351 to 2814 of SEQ ID NO:9; (C1-5) DNA molecules with nucleotide sequences such as positions 805 to 3165 of SEQ ID NO:9; (C1-6) DNA molecules with nucleotide sequences such as positions 802 to 3165 of SEQ ID NO:9; (C2) The expression box is shown in at least one of the following: (C2-1) A DNA molecule with a nucleotide sequence as shown in SEQ ID NO:6; (C2-2) A DNA molecule with a nucleotide sequence as shown in SEQ ID NO:9; (C2-3) DNA molecules with nucleotide sequences as shown in SEQ ID NO:

11.

8. A product for nucleic acid editing, said product comprising the protein of any one of claims 1 to 4, the composition or / and complex of claim 5, and / or the biomaterial of claim 6 or 7.

9. A method for nucleic acid editing, the method comprising the following steps: contacting a cell to be edited with a protein according to any one of claims 1 to 4, a composition or / and complex according to claim 5, a biological material according to claim 6 or 7, and / or a product according to claim 8 to achieve nucleic acid editing.

10. Uses described in at least one of the following: (D1) Use of the protein of any one of claims 1 to 4, the composition or / and complex of claim 5, and / or the biomaterial of claim 6 or 7 in the preparation of a formulation for nucleic acid editing; (D2) Use of the protein of any one of claims 1 to 4, the composition or / and complex of claim 5, and / or the biomaterial of claim 6 or 7 in nucleic acid editing.