A cas enzyme mutant and combination protein, nucleic acid molecule, recombinant vector, transgenic cell and application thereof

By performing point mutations, domain substitutions, and combinatorial mutations on the CasEDG-01 protein, the Cas enzyme mutant was optimized, solving the problem of low gene editing efficiency in rice and achieving highly efficient gene editing results, which is suitable for crop genetic improvement.

CN121653102BActive Publication Date: 2026-07-24EDGENE BIOTECHNOLOGY (WUHAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
EDGENE BIOTECHNOLOGY (WUHAN) CO LTD
Filing Date
2026-02-09
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

The existing CasEDG-01 protein has low gene editing efficiency in plants such as rice, which limits its application.

Method used

By performing point mutations, domain substitutions, and combinatorial mutations on the CasEDG-01 protein, Cas enzyme mutants were optimized. Combined with rice codon-optimized nucleic acid molecules and recombinant vectors, targeted editing of the host genome was achieved.

Benefits of technology

It significantly improves the editing efficiency of Cas enzyme mutants, up to 100%, which is 1.2-3.4 times higher than that of wild type, providing a new modification pathway that is suitable for crop genetic improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121653102B_ABST
    Figure CN121653102B_ABST
Patent Text Reader

Abstract

The application discloses a Cas enzyme mutant and a combined protein, a nucleic acid molecule, a recombinant vector, a transgenic cell and application thereof, relates to the field of gene editing technology, and the mutant is obtained based on CasEDG-01 protein modification, and the amino acid sequence of the CasEDG-01 protein is shown as SEQ ID NO. 1. The Cas enzyme mutant retains the PAM recognition preference (close to TTN PAM) of the Cas12a family, the editing efficiency is significantly improved, is up to 100%, is 1.2-3.4 times higher than that of the wild type, and an optimization strategy of point mutation, domain replacement and combined modification of the Cas protein is provided, a new path for the modification of low-activity Cas protein is provided, and the mutant shows high editing activity in rice, can be widely applied to crop genetic improvement, and has important agricultural application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing technology, specifically to a Cas enzyme mutant and its combinatorial protein, nucleic acid molecules, recombinant vectors, transgenic cells, and their applications. Background Technology

[0002] The CRISPR / Cas system is an adaptive immune system of bacteria and has been developed into a highly efficient gene editing tool. Among them, Cas9 and Cas12a (Cpf1) are commonly used enzymes in plant chromosome engineering, belonging to Type II and Type V, respectively. The Cpf1 protein has a bilobal structure, containing a catalytic domain (NUC) and a recognition domain (REC). The RuvC domain in the catalytic domain cleaves the non-complementary strand, the Nuc domain cleaves the complementary strand, and the WED and PI domains participate in PAM recognition.

[0003] Optimization of gene editing enzymes can be achieved through targeted screening, domain modification, point mutation, etc. The CasEDG-01 protein identified from Flavobacterium branchiophilum belongs to the Cas12a family, but its gene editing efficiency in plants such as rice is low, which limits its application. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a Cas enzyme mutant and its combinatorial protein, nucleic acid molecule, recombinant vector, transgenic cell, and their applications, thus solving the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a Cas enzyme mutant, wherein the mutant is obtained by modifying the CasEDG-01 protein, the amino acid sequence of which is shown in SEQ ID NO.1; the mutant is any one of the following:

[0006] Point mutants: at least one site mutation selected from Q162R; I878G; Y758A; S759R; Q892L; A821S; Q999F;

[0007] Domain substitution mutant: The RECⅠ or RECⅡ domain of CasEDG-01 is replaced with the corresponding domain of Lbcpf1, and the D156R mutation is introduced into the replaced RECⅠ domain. The RECⅠ domain of CasEDG-01 is located at positions 34-303, and the RECⅡ domain is located at positions 304-572. The RECⅠ domain of Lbcpf1 is located at positions 36-291, and the RECⅡ domain is located at positions 292-516.

[0008] The definitions of the RECⅠ and RECⅡ structural domains are as follows:

[0009] The RECⅠ domain of CasEDG-01 corresponds to positions 34-303 of its amino acid sequence SEQ ID NO.1;

[0010] The RECⅡ domain of CasEDG-01 corresponds to positions 304-572 of its amino acid sequence SEQ ID NO.1;

[0011] The amino acid sequence of Lbcpf1 is shown in SEQ ID NO.2;

[0012] The RECⅠ domain of Lbcpf1 corresponds to positions 36-291 of its amino acid sequence SEQ ID NO.2;

[0013] The RECⅡ domain of Lbcpf1 corresponds to positions 292-516 of its amino acid sequence SEQ ID NO.2;

[0014] Combinatorial mutants are also known as combinatorial proteins: they contain both point mutations and domain substitution mutants.

[0015] Furthermore, the point mutant is a combination of Q162R, Q162R&Q892L, Q162R&A821S, or Q162R&Q999F.

[0016] Furthermore, the domain substitution mutant is RECⅠ&D156R or RECⅠ&RECⅡ&D156R.

[0017] Furthermore, the combined mutants are RECⅠ&RECⅡ&D156R / A821S, RECⅠ&RECⅡ&D156R / Q999F, RECⅠ&RECⅡ&D156R / A821S / Q999F, RECⅠ&RECⅡ&D156R / Q892L / Q999F, or RECⅠ&RECⅡ&D156R / A821S / Q892L / Q999F.

[0018] A nucleic acid molecule encoding the Cas enzyme mutant, the sequence of which has been optimized with rice codons.

[0019] A recombinant vector comprising the aforementioned nucleic acid molecule, the recombinant vector further comprising an sgRNA expression cassette, a promoter, a nuclear localization signal coding sequence, and a selection marker gene.

[0020] Furthermore, the promoter includes the Ubi promoter, OsU6a promoter, OsU6b promoter, or OsU3 promoter; the nuclear localization signal is the BP nuclear localization signal; and the selection marker gene is a hygromycin resistance gene.

[0021] A transgenic cell comprising the Cas enzyme mutant, the nucleic acid molecule, or the recombinant vector; the cell is an Escherichia coli cell or a plant cell.

[0022] A gene editing method wherein the Cas enzyme mutant forms a ribonucleoprotein complex with crRNA targeting a target gene, or the Cas enzyme mutant and sgRNA are expressed in host cells via the recombinant vector to achieve targeted editing of the host genome.

[0023] An application of the Cas enzyme mutant, the nucleic acid molecule, the recombinant vector, or the transgenic cell in plant gene editing and crop genetic improvement; wherein the plant is rice.

[0024] This invention provides a Cas enzyme mutant and its combinatorial protein, nucleic acid molecule, recombinant vector, transgenic cell, and their applications, which have the following beneficial effects:

[0025] 1. This Cas enzyme mutant, along with its combinatorial proteins, nucleic acid molecules, recombinant vectors, transgenic cells, and their applications, retains the PAM recognition preference of the Cas12a family (close to TTNPAM) and exhibits significantly improved editing efficiency, reaching up to 100%, which is 1.2-3.4 times higher than the wild type. It also provides optimization strategies for Cas proteins through point mutations, domain substitutions, and combinatorial modifications, offering a new pathway for the modification of low-activity Cas proteins. Furthermore, the mutant demonstrates highly efficient editing activity in rice and can be widely applied to crop genetic improvement, possessing significant agricultural application value. Attached Figure Description

[0026] Figure 1 A three-dimensional schematic diagram of the locations of the 201 point mutation sites in CasEDG-01;

[0027] Figure 2 A map of the E. coli expression vector for CasEDG-01;

[0028] Figure 3 A map of the positive selection vector for CasEDG-01 Escherichia coli;

[0029] Figure 4 These are the first positive screening sequencing results for ATTA PAM;

[0030] Figure 5 These are the sequencing results from the second positive screening of ATTA PAM.

[0031] Figure 6 Plasmid map of 6N PAM target library;

[0032] Figure 7 A schematic diagram of the first part of the PAM motif for CasEDG-01 and its single-point mutants;

[0033] Figure 8 A schematic diagram of the second part of the PAM motif for CasEDG-01 and its single-point mutants;

[0034] Figure 9 A schematic diagram of the third part of the PAM motif for CasEDG-01 and its single-point mutants;

[0035] Figure 10 The first part of the diagram shows the PAM depletion of CasEDG-01 and its single-point mutants;

[0036] Figure 11 The second part of the figure shows the PAM depletion of CasEDG-01 and its single-point mutants;

[0037] Figure 12 The third part of the figure shows the PAM depletion of CasEDG-01 and its single-point mutants;

[0038] Figure 13 A schematic diagram of the vector structure used to determine the callus editing efficiency of CasEDG-01 and Lbcpf1 rice.

[0039] Figure 14 Vector map for determining the callus editing efficiency of CasEDG-01 rice;

[0040] Figure 15 A schematic diagram of the vector structure used to measure the callus editing efficiency of rice with CasEDG-01 domain replacement.

[0041] Figure 16 A schematic diagram of the vector structure used to determine the callus editing efficiency of rice with CasEDG-01 domain substitution and multi-point mutation.

[0042] Figure 17 The image shows a comparison of the three-dimensional structures of CasEDG-01 and its mutants. Detailed Implementation

[0043] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.

[0044] Unless otherwise specified, the instruments, reagents, and materials used in the following embodiments are all conventional instruments, reagents, and materials already available in the prior art and can be obtained through legitimate commercial channels. Unless otherwise specified, the experimental methods and detection methods used in the following embodiments are all conventional experimental methods and detection methods already available in the prior art.

[0045] The specific implementation scheme of this invention is as follows:

[0046] Method for constructing CasEDG-01 saturated mutant libraries:

[0047] To improve the coverage and effectiveness of the screening, a CasEDG-01 saturated mutant library was generated, containing all possible point mutations at the protein level across 201 sites. These 201 sites were obtained based on the 3D structure data of Lbcpf1 and Ascpf1 proteins. The 3D structure data of these two proteins and nucleic acid complexes were obtained from the PDB website (PDBID: 5XUT and 5B43, respectively), and their structures showed structural similarity to the structures predicted by CasEDG-01 using Alphafold3. Using PyMOL software, all amino acids within a 6 Å distance of Lbcpf1 and Ascpf1 proteins and nucleic acids were identified. In the 3D structure of protein 5B43, the amino acid of Ascpf1 within a 6 Å distance of the nucleic acid substrate was also identified. The amino acids were mapped to CasEDG-01 through homology alignment, yielding 235 amino acid sites. In the 5XUT protein three-dimensional structure, an amino acid from Lbcpf1 located 6 Å from the nucleic acid substrate was mapped to CasEDG-01 through homology alignment, yielding 244 amino acid sites. The intersection of the corresponding amino acids from these two proteins was used to obtain 201 candidate mutation sites for CasEDG-01. These sites include 12 sites: E18R, Q162R, N271G, D597H, Y758A, S759R, A821S, F825R, K831V, I878G, Q892L, and Q999F. The positions of these 201 sites in Ascpf1, Lbcpf1, and the intersection site in CasEDG-01 are shown below. Figure 1 As shown (red dots in the figure represent selected mutation sites);

[0048] A saturated mutant library was constructed from the 201 selected candidate sites: basic primers NNK (N and K are basic bases, where N is one of A, T, C, or G, and K is one of T or G) were designed at the sites requiring saturation mutations. Approximately 21 bp of wild-type sequence were added before and after these primers as forward primers, and reverse primers were designed 500-1000 bp downstream. The *CasEDG-01* *E. coli* expression vector is as follows: Figure 2 As shown, the main structure of the vector is modified based on pET28a. The nucleic acid sequence expressed by E. coli of CasEDG-01 is inserted into the downstream of the T7 promoter on the pET28a vector. The HIS tag near the T7 promoter end of the vector is retained, while the His tag sequence at the C end is removed.

[0049] Then, the 500bp-1000bp long fragment was amplified using the phanta enzyme Novizan. After amplification, 3μL of the product was taken for agarose gel electrophoresis to check for the presence of a 500-1000bp band. If no correct band was found, primers were redesigned. If a correct band was found, it was used directly for the second round of amplification. The second round of amplification used the first round of amplification products as primers to amplify the entire vector. In this amplification method, the 201 site saturation mutations were introduced into the vector.

[0050] The PCR amplification system for saturation mutation is shown in Table 1:

[0051] Table 1

[0052] Phanta enzyme 0.2μL 0.3μL Deoxyribonucleoside triphosphates (dNTPs) 0.2μL 0.3μL 10* Amplification Buffer 5μL 7.5μL template 5ng 80ng water Upto (to make up) 10μL Upto (to make up 15μL) Primers 0.2μL 1.2 μL of first-round amplification product

[0053] DpnI digestion, dialysis, transformation, and plasmid extraction:

[0054] The second-round amplification product was reacted at 37℃ for 3 hours, then at 80℃ for 20 minutes. The deionized product was then placed on a 0.025 μm Pore Size dialysis membrane and dialyzed with 1 / 3*TE buffer. After dialysis for 1 hour, the liquid was pipetted into a new PCR tube. This sample was then cryopreserved at -80℃ or directly transformed into DH10b electroporated competent cells. Using this method, a total of 201 single-point saturated mutant libraries were obtained. The concentration of the extracted plasmids was determined using nanodrop. These 201 single-point saturated mutant libraries were then mixed in one tube with the same plasmid mass (plasmid concentration controlled at 200 ng / μL). This mixed library is called a saturated mutant library.

[0055] Forward screening experiment of CasEDG-01 saturated mutant library:

[0056] The principle of forward selection experiments is to select plasmids by forward selection (selecting plasmid structures such as...) Figure 3 As shown in the figure, CasEDG-01 variants with enhanced editing activity were screened. The screening plasmid contained arabinose operons. When arabinose is present, it induces the expression of the toxic gene ccdB downstream of the operon, leading to bacterial death. However, the ccdB gene has PAM as an ATTA target, and the vector also contains an sgRNA expression element initiated by a T7 promoter (this sgRNA targets the target on the ccdB gene). If the CasEDG-01 variants expressed in the saturated mutant library have high editing efficiency, then such mutants are more likely to survive during screening. In subsequent sequencing, their frequency will be higher than that of the control without added arabinose.

[0057] The selection plasmid containing the ATTA target PAM was transformed into Rosetta (DE3) competent cells. Then, a single clone was selected for electroporation of competent cells, which will be referred to hereafter as ATTA-selected competent cells. The simplified procedure for electroporation of competent cells is as follows:

[0058] 1) Select Rosetta (DE3) single clones of screening plasmids containing ATTA target sites of PAM, inoculate them into test tubes containing 2 mL-5 mL of LB medium, and culture overnight at 37 °C with shaking at 180 rpm.

[0059] 2) On the second day, take 2 mL of the overnight culture and transfer it into a 1 L shake flask containing 200 mL of LB medium (1:100). Incubate at 37°C with vigorous shaking (220 rpm) until the OD600 is about 0.55. Then, place the bacterial culture in an ice-water mixture for 15 min to cool it down quickly.

[0060] 3) Centrifuge two pre-chilled 50mL centrifuge tubes at 2000g for 5 minutes at 4℃. Collect the bacterial cells and discard the supernatant;

[0061] 4) Gently resuspend the bacterial cells in pre-cooled water (pure double-distilled water, autoclaved). First, add a small amount of water to resuspend the bacterial cells, then add water to 40 mL. Centrifuge at 3000 g for 5 min at 4℃, discard the supernatant, and repeat this step 3 times.

[0062] 5) Centrifuge at 3000g for 5 minutes with 10% glycerol (autoclaved). Repeat step 2 times. After discarding the supernatant on the last time, suspend the cells in the glycerol residue at the bottom of the tube and dispense into EP tubes, 50μL / tube (generally resuspended at a ratio of 1:400-1:500).

[0063] 6) Quick-freeze in liquid N2 and store at -70℃ (the prepared competent cells can be used directly without quick-freezing);

[0064] Then, 1 μL of the saturated mutant library plasmid was electroporated into ATTA-selected competent cells. Immediately after transformation, 1 mL of LB medium was added, and the cells were incubated at 37°C for 1.5 h. IPTG was then added to a final concentration of 1 mM for 1 h of induction. The cells were then divided into two aliquots: one aliquot was plated on plates containing 2% (w / v) arabinose, 50 μg / mL kanapenem, and 100 μg / mL ampicillin as the experimental group; the other aliquot was plated on plates containing 50 μg / mL kanapenem and 100 μg / mL ampicillin. Using L-ampicillin plates as a control, all colonies on the plates were collected and plasmids were extracted. This constituted the first round of plasmid selection. The plasmids from the experimental group were then electroporated with two competent cells for the second round of selection. Similarly, the third, fourth, and fifth rounds of selection were performed. The experiment was conducted twice. In the first experiment, both the first and second rounds of selection were performed, and sequencing was performed on both. In the second experiment, only the first, third, and fifth rounds of selection were performed, and sequencing was performed on only the first, third, and fifth rounds. These selected plasmids were first processed using primers:

[0065] The forward primer sequence TCCTTCGGGCTTTGTTAG (SEQ ID NO.42) and the reverse primer sequence TTAGGAAGCAGCCCAGTAG (SEQ ID NO.43) amplified the entire protein sequence. The amplified product was 4.5 kb in size. This amplified product was then sent to a sample library for next-generation library construction and sequencing (second-generation library construction and sequencing were performed after sonication fragmentation, and finally, the mutation frequency of different mutation sites was statistically analyzed by bioinformatics analysis). The frequency of different mutants in each round of screening in the two experiments was finally obtained.

[0066] The Fold Rate values ​​for these 201 sites were calculated as follows: If the mutation frequency in the experimental group was less than that in the control group, the mutation frequency in the experimental group was divided by the mutation frequency in the control group; conversely, the mutation frequency in the control group was divided by the mutation frequency in the experimental group. The relative survival rate of each point mutation relative to the reference was calculated as the ratio of the normalized occurrence rate between the selected library and the input library (Fold Rate). Since cell survival indicates the DNA cleavage activity of each CasEDG-01 variant, any variant with a higher survival rate than the reference protein is a variant with enhanced activity at ATTAPAM.

[0067] Based on the sequencing results of the two sequencing sessions above (e.g.) Figure 4 and Figure 5As shown), the second sequencing results yielded candidate mutation sites for subsequent experiments (E18R, N271G, D597H, Y758A, S759R, A821S, F825R, K831V, I878G). The first sequencing results included Q892L and Q162R at the Q999F site, for a total of 12 sites, which were further validated. PAM depletion experiments were performed on E. coli to verify their PAM preference. At the same time, their depletion data at 24 sites in NTTN, NTCN, and TCTN were compared to compare their depletion capabilities under different PAMs.

[0068] PAM analysis and PAM depletion analysis of E. coli from the CasEDG-01 mutant:

[0069] First, 11 single-point mutants of the CasEDG-01 E. coli expression vector were constructed: E18R, N271G, D597H, Y758A, S759R, A821S, F825R, K831V, I878G, Q892L, and Q162R. The construction method of single-point mutants is similar to that of single-point saturated mutants. The primer design principles and amplification procedures are basically the same. The only difference is that NNK in the primers is replaced with the specific codons of these 11 point mutations.

[0070] These 11 successfully constructed mutants and wild-type CasEDG-01 were transformed into Rosetta (DE3) competent cells. One single clone of each mutant was selected and electroporated into competent cells. Then, a 6N target library plasmid (vector structure as shown in the target library plasmid) was used. Figure 6 (As shown) The plasmids were electroporated into 12 types of electroporation competent cells, 1 mL of antibiotic-free LB medium was added, and the cells were incubated at 37°C for 90 min. Then, they were plated on kanamycin, ampicillin, and IPTG plates and incubated overnight at 37°C. The next day, all colonies on the plates were washed off, and plasmids were extracted. These 12 extracted plasmid libraries, along with the original 6N target library (as a control), were then used for next-generation sequencing. Bioinformatics analysis of the sequencing results yielded their PAM motifs (the PAM motifs of these 12 plasmids are shown below). Figures 7 to 9 (as shown)

[0071] From these 12 motif diagrams, it can be seen that all 11 mutants are very similar to the wild-type PAM motif, close to a TTN PAM, indicating that these mutants do not change CasEDG-01's preference for recognizing PAM. At the same time, from the PAM motif diagrams, it can be seen that the fourth and fifth bases of the PAM of CasEDG-01 wild-type and mutants have the preference for both T and C bases. Therefore, we compared the specific PAM depletion levels of these 12 plasmids at the 24 PAM sites (NTTN, TTCN, TCTN) near the target site.

[0072] The method for calculating PAM exhaustion level is as follows: The frequency of the first four bases of the target sequence (256 PAMs) in all sequencing data is statistically analyzed. The frequency of these 256 PAMs in the control group (original 6N target library) is divided by the frequency of CasEDG-01 and its single-site mutants to obtain the PAM exhaustion level. A higher ratio indicates a greater decrease in the proportion of that PAM in the sequencing data compared to the control group, meaning a stronger substrate cleavage ability of the wild-type or mutant CasEDG-01 for that PAM. Then, the specific PAM exhaustion levels at 24 PAM sites (NTTN, TTCN, and TCTN) are plotted. The results are shown in the PAM exhaustion graph. For details, please refer to [link to relevant documentation]. Figures 10 to 12 ;

[0073] From the PAM exhaustion diagram Figures 10 to 12 From the data, it can be seen that Q162R, I878G, Y758A, S759R, and Q892L have better cutting performance than WT, A821S, K831V, and N271G. In PAM cutting efficiency, such as VTTV and TTCN, they are higher than wild-type.

[0074] Based on the results of PAM depletion, six single-point mutants, Q162R, I878G, Y758A, S759R, Q892L, and A821S, were selected for subsequent determination of rice callus editing efficiency. Since the locations of these six point mutations are all in the RECⅠ or WEDⅢ domains, when determining the rice callus editing efficiency, an additional point mutation, Q999F, located in the RuvCⅡ domain, obtained from the first screening, was also introduced.

[0075] The distribution of CasEDG-01 candidate mutation sites and their domains is shown in Table 2 below:

[0076] Table 2

[0077] Domain (structural domain) RECⅠ WEDIII WEDIII WEDIII WEDIII WEDIII RuvCⅡ

[0078] Determination of callus editing efficiency in rice using the CasEDG-01 mutant:

[0079] When designing vectors for measuring rice callus editing efficiency, single-point mutations, pairwise combination mutations, and three-point combination mutations were designed (when combining mutants, the main reference was the domain where the mutation site is located, and mutants from different domains were combined). The final designs included: Q162R; Q892L; Q162R / Q892L; S759R; I878G; Q162R / S759R; Q162R / I878G; S759R / I878G; Q162R / S Seventeen mutant vectors, including 759R / I878G, Y758A, A821S, Q162R / Y758A, Q162R / A821S, Q999F, Q162R / Q999F, Q162R / I878G / Q999F, and Q162R / S759R / Q999F, were used. Wild-type CasEDG-01 and Lbcpf1 were used as controls. The vector structure for determining the rice callus editing efficiency of CasEDG-01 is shown below. Figure 13 and Figure 14 As shown, the vector structure of Lbcpf1 is completely identical to that of CasEDG-01 except for changes in the protein expression sequence and DR sequence. The vector also contains two sgRNA expression frames targeting the rice EPFL9 and ROC5 genes, which are transcribed using promoters OsU6a and OsU3, respectively. Nuclear localization signals (BP) are introduced at both the N-terminus and C-terminus of the protein. The BP nuclear localization signals and the protein are linked at the N-terminus by a flexible peptide GS and at the C-terminus by a flexible peptide SG. The BP at the N-terminus and C-terminus are identical in amino acid sequence, with only slight differences in base sequence. Furthermore, the base sequences of CasEDG-01 wild type and its mutants, as well as the base sequences of Lbcpf1, have been optimized using rice codons. The optimized sequences are shown in SEQ ID NO.11 in the sequence listing.

[0080] After the vector was constructed, the callus editing efficiency of rice was determined. The detailed process of rice genetic transformation is as follows:

[0081] (1) Induction of rice callus:

[0082] Whole rice (grass rice "Zhonghua 11") seeds were selected, dehulled, and after disinfection, washing, and drying with filter paper, an appropriate amount of seeds were inoculated into NB medium and cultured under light at 28-30℃ for 8-10 days to induce callus tissue.

[0083] (2) Pre-culture of callus tissue:

[0084] Select dense, brightly colored, and granular callus tissues from the rice callus tissues induced in step (1) and transfer them into a pre-culture medium. Culture them under light at 28-30℃ for 3 days for transformation.

[0085] (3) Culture of Agrobacterium and Agrobacterium-mediated transformation:

[0086] The vector was transferred into Agrobacterium strain EHA105 by electroporation, positive bacteria were identified and cultured, the bacteria were shaken, and a certain amount of bacterial cells were collected and suspended in Agrobacterium saturation solution. The rice callus tissue cultured under light for 3 days in step (2) was soaked in the above Agrobacterium saturation solution and shaken on a shaker at 150 rpm and 28°C for 10 minutes.

[0087] (4) Co-culture of Agrobacterium and callus:

[0088] The rice callus tissue after step (3) was soaked was placed on sterile filter paper and dried. Then the callus tissue was transferred to a co-culture medium and cultured in the dark at 28°C for 3 days.

[0089] (5) Screening of resistant callus:

[0090] After washing the callus tissue co-cultured for 3 days as described in step (4) 4-5 times, wash the callus tissue again with carbenicillin solution, dry it with filter paper, and inoculate it on a selection medium containing hygromycin for screening. The culture conditions are 28°C and light culture for 20-22 days, with one subculture in between, until resistant callus tissue grows.

[0091] (6) Treatment of resistant callus:

[0092] The positive callus obtained after screening in step (5) was subjected to normal screening for 7 days. Genomic DNA was extracted from the callus and used for next-generation library construction, sequencing and analysis.

[0093] The final sequencing results are shown in Table 3. CasEDG-01-Q162R / A821S, CasEDG-01-Q162R / Q999F, and CasEDG-01-Q162R / Q892L have relatively high editing efficiency, and their editing efficiency on both OsEPFL9 and OsROC5 targets is higher than that of wild-type CasEDG-01.

[0094] Table 3: Statistical table of callus editing efficiency in rice with single-site and multi-site combined mutations of CasEDG-01.

[0095] Table 3

[0096] YX25A01 CasEDG-01-Q892L 48 9 / 48(18.7%) 23 / 48(47.9%) YX25A02 CasEDG-01-Q162RQ892L 48 34 / 48(70.8%) 42 / 48(87.5%) YX25A03 CasEDG-01-S759R 48 22 / 48(45.8%) 18 / 48(37.5%) YX25A04 CasEDG-01-I878G 48 9 / 48(18.7%) 26 / 48(54.1%) YX25A05 CasEDG-01-Q162RS759R 48 25 / 48(52.1%) 18 / 48(37.5%) YX25A06 CasEDG-01-Q162RI878G 48 28 / 48(58.3%) 30 / 48(62.5%) YX25A07 CasEDG-01-S759RI878G 48 4 / 48(8.3%) 2 / 48(4.1%) YX25A08 CasEDG-01-Q162RS759RI878G 48 18 / 48(37.5%) 6 / 48(12.5%) YX25A09 CasEDG-01-Y758A 48 14 / 48(29.1%) 19 / 48(39.5%) YX25A10 CasEDG-01-A821S 48 10 / 48(20.8%) 22 / 48(45.8%) YX25A11 CasEDG-01-Q162RY758A 48 20 / 48(41.6%) 27 / 48(56.2%) YX25A12 CasEDG-01-Q162RA821S 48 26 / 48(54.1%) 37 / 48(77.0%) YX25A13 CasEDG-01-Q999F 48 22 / 48(45.8%) 23 / 48(47.9%) YX25A14 CasEDG-01-Q162RQ999F 48 35 / 48(72.9%) 39 / 48(81.2%) YX25A15 CasEDG-01-Q162RI878GQ999F 48 30 / 48(62.5%) 18 / 48(37.5%) YX25A16 CasEDG-01-Q162RS759RQ999F 48 32 / 48(66.6%) 17 / 48(35.4%) YX25A17 CasEDG-01 48 11 / 48(22.9%) 27 / 48(56.2%) YX25A18 Lbcpf1 48 39 / 48(81.2%) 37 / 48(77.0%) YX25A19 CasEDG-01-Q162R 48 14 / 48(29.1%) 27 / 48(56.2%)

[0097] Simultaneously, the structure of the CasEDG-01 protein was modified based on its structure, for Type For type II gene-editing proteins, the most common reported domain substitutions are the WED and REC domains. Since Lbcpf1 and CasEDG-01 are both gene-editing enzymes of the Cas12a family, and Lbcpf1 shows better editing performance in plants, the WED and REC domains of CasEDG-01 were replaced with the corresponding domains of Lbcpf1. Information on these domains of CasEDG-01 was obtained through the Uniprot website (the parentheses indicate the amino acid site numbers): WEDⅠ domain (1-35), WEDⅡ domain (527-598), WEDⅢ domain (719-884), RECⅠ domain (36-320), and RECⅡ domain (321-526). Bioinformatics analysis was used to obtain the domain information for Lbcpf1 and CasEDG-01: Lbcpf1's WEDⅠ domain (1-35), WEDⅡ domain (517-586), W... The EDⅢ structural domain (678-808), RECⅠ structural domain (36-291), and RECⅡ structural domain (292-516) of CasEDG-01; the WEDⅠ structural domain (1-33), WEDⅡ structural domain (573-640), WEDⅢ structural domain (731-905), RECⅠ structural domain (34-303), and RECⅡ structural domain (304-572) of CasEDG-01, which will combine the WEDⅠ, WEDⅡ, WEDⅢ, and RECⅡ structural domains of CasEDG-01. The CⅠ and RECⅡ domains are respectively replaced with their corresponding domains in Lbcpf1, and variants with simultaneous replacement of RECⅠ and RECⅡ (referred to as RECⅠ&RECⅡ) are constructed. A highly efficient mutant of Lbcpf1 (D156R, located in the RECⅠ domain) is known. This invention incorporates this mutation into the existing RECⅠ domain replacement, for example, RECⅠ&D156R replacement or RECⅠ&RECⅡ&D156R replacement. The constructed vector structure is as follows: Figure 15 As shown in the vector structure map, the vector structure is similar to that of point mutations and combination mutations. The test target is the EPFL9 gene target used above. After constructing the domain replacement vector, the rice callus editing efficiency was determined. Forty-eight rice callus tissues were selected for each vector to test the editing efficiency. The final editing results are shown in Table 4.

[0098] Table 4: Statistics on the efficiency of rice callus editing with CasEDG-01 domain replacement.

[0099] Table 4

[0100] YX25A20 CasEDG-01 48 6 / 48(12.5%) YX25A21 CasEDG-01-RECⅠ 48 1 / 48(2.1%) YX25A22 CasEDG-01-RECⅠ&D156R 48 11 / 48(22.9%) YX25A23 CasEDG-01-RECⅡ 48 0 / 48(0%) YX25A24 CasEDG-01-RECⅠ&RECⅡ&D156R 48 29 / 48(60.4%) YX25A25 CasEDG-01-WEDⅠ 48 2 / 48(4.2%) YX25A26 CasEDG-01-WEDⅡ 48 0 / 48(0%) YX25A27 CasEDG-01-WEDⅢ 48 0 / 48(0%) YX25A28 Lbcpf1 48 42 / 48(87.5%)

[0101] The results in the table above show that:

[0102] The CasEDG-01-RECⅠ&D156R and CasEDG-01-RECⅠ&RECⅡ&D156R variants significantly improved editing efficiency at the EPFL9 site compared to wild-type CasEDG-01; therefore, the construction method of replacing only the RECⅠ domain and replacing both the RECⅠ and RECⅡ domains was finally selected for further testing.

[0103] Finally, the domain substitution and point mutation methods were combined for domain substitution, selecting the RECⅠ domain and RECⅠ&RECⅡ. For point mutations, Q162R, A821S, Q999F, and Q892L were selected. Furthermore, these four point mutations were combined into two-, three-, and four-mutations according to their domain positions, ultimately constructing RECⅠ&D156R / Q892L, RECⅠ&D156R / A821S, RECⅠ&D156R / Q892 / LQ999F, and RECⅠ&D156R / A821S / Q999F, RECⅠ&D156R / Q999F, RECⅠ&RECⅡ&D156R / Q892L, RECⅠ&RECⅡ&D156R / A821S, RECⅠ&RECⅡ&D156R / Q892L / Q999F, RECⅠ&RECⅡ&D156R / A821S / Q999F, RECⅠ&RECⅡ&D156R / Q999F and RECⅠ&RECⅡ&D156R / A821S / Q892 / LQ999F;

[0104] Simultaneous use of Q162R / A821S / Q999F, Q162R / Q892L / Q999F, and Q162R / Q892L / A821S / Q999F, these carriers are all based on Figure 16The vector diagram shown is used for construction (the diagram does not show the combined point mutations, only the structural changes and the point mutation D156R; other mutations can be obtained by mutating the corresponding amino acid sites based on this). The corresponding amino acid sequences of the above mutants are as follows: CasEDG-01-RECⅠ&D156R / Q892L (SEQ ID NO.44); CasEDG-01RECⅠ&D156R / A821S (SEQ ID NO.45); CasEDG-01-RECⅠ&D156R / Q892L / Q999F (SEQ ID NO.46); CasEDG-01-RECⅠ&D156R / A821S / Q999F (SEQ ID NO.47); CasEDG-01-RECⅠ&D156R / Q999F (SEQ ID NO.46); CasEDG-01-RECⅠ&D156R / A821S / Q999F (SEQ ID NO.47); CasEDG-01-RECⅠ&D156R / Q999F (SEQ ID NO.48). NO.48); CasEDG-01-RECⅠ&RECⅡ&D156R / Q892L (SEQ ID NO.49); CasEDG-01-RECⅠ&RECⅡ&D156R / A821S (SEQ ID NO.50); ​​CasEDG-01-RECⅠ&RECⅡ&D156R / Q892L / Q999F (SEQ ID NO.51); CasEDG-01-RECⅠ&RECⅡ&D156R / A821S / Q999F (SEQ ID NO.52); CasEDG-01-RECⅠ&RECⅡ&D156R / Q999F (SEQ ID NO.53); CasEDG-01-RECⅠ&RECⅡ&D156R / A821S / Q892L / Q999F (SEQ ID NO.54); CasEDG-01-Q162R / A821S / Q999F (SEQ ID NO.55); CasEDG-01-Q162R / Q892L / Q999F (SEQ ID NO.56); CasEDG-01-Q162R / Q892L / A821S / Q999F (SEQ ID NO.57); CasEDG-01-Q892L (SEQ ID NO.58); CasEDG-01-Q162R / Q892L (SEQ ID NO.59); CasEDG-01-S759R (SEQ ID NO.60); CasEDG-01-I878G (SEQ ID NO.61); CasEDG-01-Q162R / S759R (SEQ ID NO.62); CasEDG-01-Q162R / I878G (SEQ ID NO.63); CasEDG-01-S759R / I878G (SEQ ID NO.64); CasEDG-01-Q162R / S759R / I878G (SEQ ID NO.65); CasEDG-01-Y758A (SEQ ID NO.66); CasEDG-01-A821S (SEQ ID NO.67); CasEDG-01-Q162R / Y758A (SEQ ID NO.68); CasEDG-01-Q162R / A821S (SEQ ID NO.69); CasEDG-01-Q999F (SEQ ID NO.70); CasEDG-01-Q162R / Q999F (SEQ ID NO.71); CasEDG-01-Q162R / I878G / Q999F (SEQ ID NO.72); CasEDG-01-Q162R / S759R / Q999F (SEQ ID NO.73); CasEDG-01-Q162R (SEQ ID NO.74);.

[0105] Using wild-type CasEDG-01 and Lbcpf1 as controls, the number of target sites was increased to three compared to the previous vector map. After constructing the domain replacement vector, the rice callus editing efficiency was determined. Forty-eight rice callus tissues were selected for each vector to test the editing efficiency. The final editing results are shown in Table 5. The results show that the editing efficiency of CasEDG-01-RECⅠ&RECⅡ&D156R / A821S and CasEDG-01-RECⅠ&RECⅡ&D156R / Q999F (named ChiCasEDG-01) mutants was significantly improved compared with wild-type CasEDG-01. The editing efficiency of all three target sites exceeded 80%, with the highest editing efficiency reaching 100%.

[0106] The editing efficiency of rice callus editing vectors with multiple point mutations of CasEDG-01 domain substitution combinations is statistically analyzed as follows (Table 5):

[0107] Table 5

[0108] YX25A29 CasEDG-01-RECⅠ&D156RQ892L 48 38 / 48(79.2%) 36 / 48(75.0%) 41 / 48(85.4%) YX25A30 CasEDG-01-RECⅠ&D156RA821S 48 37 / 48(77.1%) 26 / 48(54.2%) 37 / 48(77.1%) YX25A31 CasEDG-01-RECⅠ&D156RQ892LQ999F 48 35 / 48(72.9%) 21 / 48(43.8%) 33 / 48(68.8%) YX25A32 CasEDG-01-RECⅠ&D156RA821SQ999F 48 39 / 48(81.2%) 22 / 48(45.8%) 31 / 48(64.6%) YX25A33 CasEDG-01-RECⅠ&D156RQ999F 48 41 / 48(85.4%) 30 / 48(62.5%) 40 / 48(83.3%) YX25A34 CasEDG-01-RECⅠ&RECⅡ&D156RQ892L 48 41 / 48(85.4%) 39 / 48(81.2%) 41 / 48(85.4%) YX25A35 CasEDG-01-RECⅠ&RECⅡ&D156RA821S 48 45 / 48(93.7%) 41 / 48(85.4%) 42 / 48(87.5%) YX25A36 CasEDG-01-RECⅠ&RECⅡ&D156RQ892LQ999F 48 37 / 48(77.1%) 33 / 48(68.8%) 34 / 48(70.8%) YX25A37 CasEDG-01-RECⅠ&RECⅡ&D156RA821SQ999F 48 48 / 48(100%) 35 / 48(72.9%) 37 / 48(77.1%) YX25A38 CasEDG-01-RECⅠ&RECⅡ&D156RQ999F 48 48 / 48(100%) 40 / 48(83.3%) 43 / 48(89.6%) YX25A39 CasEDG-01-RECⅠ&RECⅡ&D156RA821SQ892LQ999F 48 40 / 48(83.3%) 38 / 48(79.2%) 40 / 48(83.3%) YX25A40 Lbcpf1 48 41 / 48(85.4%) 41 / 48(85.4%) 41 / 48(85.4%) YX25A41 CasEDG-01 48 14 / 48(29.2%) 33 / 47(70.2%) 32 / 48(66.7%) YX25A42 CasEDG-01-Q162RA821SQ999F 48 37 / 48(77.1%) 43 / 48(89.6%) 38 / 48(79.2%) YX25A43 CasEDG-01-Q162RQ892LQ999F 48 41 / 48(85.4%) 44 / 47(93.6%) 47 / 48(97.9%) YX25A44 CasEDG-01-Q162RQ892LA821SQ999F 48 42 / 48(87.5%) 43 / 48(89.6%) 39 / 48(81.2%)

[0109] Three-dimensional structural comparison of the CasEDG-01 mutant:

[0110] The three-dimensional structures of Lbcpf1, CasEDG-01, and two chimeric site mutant proteins, RECⅠ&D156R / Q999F and RECⅠ&RECⅡ&D156R / Q999F (ChiCasEDG-01), were predicted using Alphafold3. The three-dimensional structures of Lbcpf1 and CasEDG-01, as well as their respective chimeric counterparts, were aligned using PyMOL, and their RMSD values ​​were calculated. The final results are as follows: Figure 17 As shown, the following table (Table 6) contains... Figure 17 Explanation of the 3D structural comparison diagrams of CasEDG-01 and its mutants:

[0111] Table 6

[0112]

[0113] From the perspective of RMSD values, whether considering overall structural similarity or core structural similarity, the three-dimensional structure of RECⅠ&RECⅡ&D156R / Q999F is close to the structures of CasEDG-01 and Lbcpf1, but there are significant differences; while the three-dimensional structural similarity of RECⅠ&D156R / Q999F with CasEDG-01 and Lbcpf1 is even lower than the similarity between CasEDG-01 and Lbcpf1.

[0114] in conclusion:

[0115] Combinatorial mutation tests on the structure and point mutations of CasEDG-01 showed that the ChiCasEDG-01 mutant improved the rice callus editing efficiency by 1.2-3.4 times compared with the wild-type CasEDG-01. In terms of innovative ideas for Cas protein optimization, chimeric mutants formed by key domain substitution and key domain site mutation between Cas proteins can significantly improve the editing activity of Cas proteins, providing a new optimization path for Cas proteins with low editing levels.

[0116] Therefore, the Cas enzyme mutant of this invention retains the PAM recognition preference of the Cas12a family (close to TTNPAM) and significantly improves the editing efficiency, up to 100%, which is 1.2-3.4 times higher than the wild type. It also provides a Cas protein optimization strategy of point mutation, domain substitution and combination modification, providing a new path for the modification of low-activity Cas proteins. Moreover, the mutant exhibits high-efficiency editing activity in rice and can be widely used in crop genetic improvement, which has important agricultural application value.

[0117] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described to better illustrate the principles and practical methods of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.

Claims

1. A Cas enzyme mutant, characterized in that: The mutant was obtained by modifying the CasEDG-01 protein, the amino acid sequence of which is shown in SEQ ID NO.1; the mutant is any one of the following: Specifically selected are RECⅠ&RECⅡ&D156R / Q999F, whose sequence is shown in SEQ ID NO.53; and CasEDG-01-RECⅠ&RECⅡ&D156R / A821S, whose sequence is shown in SEQ ID NO.

50.

2. A nucleic acid, characterized in that: It encodes the Cas enzyme mutant of claim 1, wherein the sequence of the nucleic acid is optimized with rice codons.

3. A recombinant vector, characterized in that: The recombinant vector comprises the nucleic acid of claim 2, and further comprises an sgRNA expression cassette, a promoter, a nuclear localization signal coding sequence, and a selection marker gene.

4. The recombinant vector according to claim 3, characterized in that: The promoter includes the Ubi promoter, OsU6a promoter, OsU6b promoter, or OsU3 promoter; the nuclear localization signal is the BP nuclear localization signal; and the selection marker gene is a hygromycin resistance gene.

5. A transgenic cell, characterized in that: It comprises the Cas enzyme mutant of claim 1, the nucleic acid of claim 2, or the recombinant vector of any one of claims 3-4; the cell is an Escherichia coli cell or a plant cell.

6. A gene editing method, characterized in that: The Cas enzyme mutant of claim 1 is combined with crRNA targeting the target gene to form a ribonucleoprotein complex, or the Cas enzyme mutant and sgRNA are expressed in host cells using any of the recombinant vectors of claims 3-4, thereby achieving targeted editing of the host genome.

7. An application, characterized in that: The application of the Cas enzyme mutant of claim 1, the nucleic acid of claim 2, the recombinant vector of any one of claims 3-4, or the transgenic cell of claim 5 in plant gene editing and crop genetic improvement.