Gene expression modulation system based on type i-f crispr / cas and application thereof

CN122609596APending Publication Date: 2026-08-21GERMPLASM INNOVATION GRAND SCIENCE CENTER OF WESTERN CHINA (CHONGQING) SCIENCE CITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610772422.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]然而关于是否能基于Type I-F CRISPR系统在植物中构建高效的转录调控系统仍是未知

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122609596A_ABST
    Figure CN122609596A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of genetic engineering, and particularly relates to a gene expression regulation system based on Type I-F CRISPR / Cas and application thereof. The technical problem to be solved by the application is to construct a Type I-F CRISPR system suitable for plant genome editing. The application constructs a Type I-F CRISPR expression frame suitable for plant genome editing, which comprises elements for expressing Cas5, Cas6, Cas8 and Cas7; the C-terminus of Cas6 and / or Cas7 is further fused with 2xTAD, 2xTAD-VP64 or TV. The application successfully develops a PaeCascade system, which can be applied to a plant gene regulation tool, further enriches the plant genome editing tool, and lays a foundation for popularization and application of the PaeCascade system in the field of plant genome editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of genetic engineering technology, specifically relating to a gene expression regulation system based on Type I-F CRISPR / Cas and its applications. Background Technology

[0002] The CRISPR-Cas system, a widely reported and highly efficient gene-editing tool, has traditionally been primarily used in gene editing, where its importance has grown significantly in recent years. Notably, current research has also preliminarily confirmed its potential applications and development space in gene expression regulation. Gene expression regulation specifically refers to the regulation and control of gene expression processes within organisms, encompassing multiple key stages at the gene, transcription, post-transcriptional, translational, and post-translational levels.

[0003] CRISPR-Cas systems are divided into two main categories. The first category consists of multiple subunits forming an effector complex to perform their function; the second category consists of a single effector protein (such as Cas9, Cas12a, etc.) to perform their function. The first category of CRISPR-Cas systems can be further divided into type I, type III, and type IV. Type I CRISPR systems consist of multiple effector proteins, requiring multiple subunits to assemble into a CRISPR-associated complex for antiviral defense (Cascade) to perform their function. Based on genetic characteristics such as the number and spatial arrangement of subunits, type I CRISPR systems can be further divided into eight subtypes (IA to IG and IU). The IF-type PaeCascade from Pseudomonas aeruginosa consists of a complex composed of four effector proteins (Cas5, Cas6, Cas8, and six Cas7s) to bind dsDNA. In the presence of Cas3, the PaeCascade complex can attract Cas3 to cleave target sequences within bacteria that meet the PAM (protospacer abject motif) 5'-CC-3' requirement, thereby achieving target gene editing. Previous literature reports the successful construction of a mammalian PaeCascade transcriptional regulatory system by optimizing the PaeCascade sequence and fusing the transcriptional activator VPR with PaeCascade.

[0004] However, it remains unknown whether an efficient transcriptional regulatory system can be constructed in plants based on the Type IF CRISPR system. Summary of the Invention

[0005] The technical problem to be solved by this invention is to construct a Type IF CRISPR system suitable for plant genome editing.

[0006] The technical solution of the present invention is a nucleic acid encoding Cas5, the sequence of which is shown in Seq ID No.1.

[0007] The present invention also provides a nucleic acid encoding Cas6, the sequence of which is shown in Seq ID No.2.

[0008] The present invention also provides a nucleic acid encoding Cas8, the sequence of which is shown in Seq ID No.3.

[0009] The present invention also provides a nucleic acid encoding Cas7, the sequence of which is shown in Seq ID No.4.

[0010] The present invention also constructs a Type IF CRISPR expression cassette suitable for plant genome editing, including elements expressing Cas5, Cas6, Cas8 and Cas7; the C-terminus of Cas6 and / or Cas7 is also fused with 2xTAD, 2xTAD-VP64 or TV.

[0011] Furthermore, the expression box also includes a Cas6 recognition sequence expression unit.

[0012] Specifically, the structure of the Cas6 recognition sequence expression unit is promoter-DR-LacZ-DR-terminator.

[0013] Specifically, the component structures representing Cas5, Cas6, Cas8, and Cas7 are one of the following:

[0014] a. Promoter - Cas5 - Cas8 - Terminator - Promoter - Cas6 - Cas7 - 2xTAD - Terminator;

[0015] b. Promoter - Cas5-Cas8-Terminator - Promoter - Cas6-Cas7-2xTAD-VP64-Terminator;

[0016] c. Starter - Cas5 - Cas8 - Terminator - Starter - Cas6 - Cas7 - TV - Terminator;

[0017] d. Starter - Cas5-Cas6-2xTAD - Terminator - Starter - Cas8-Cas7-2xTAD - Terminator.

[0018] Furthermore, the N-terminus and / or C-terminus of the Cas6 and / or Cas7 also incorporate a nuclear positioning signal.

[0019] Specifically, the nuclear positioning signal is SV40 NLS, BP NLS, or nucleoplasmin NLS.

[0020] The sequence of the nucleic acid encoding Cas6-2xTAD is shown in Seq ID No. 5.

[0021] The sequence of the nucleic acid encoding Cas7-2xTAD is shown in Seq ID No. 8.

[0022] Specifically, the sequence of the nucleic acid encoding Cas7-2xTAD-VP64 is shown in Seq ID No. 9.

[0023] Specifically, the sequence of the nucleic acid encoding Cas7-TV is shown in Seq ID No. 10.

[0024] Specifically, the promoter is p35S, pZmUbi1, pOsUbi, or pU6.

[0025] The terminator is pinII, AtHSP, NOS, poly T, or OCS.

[0026] Preferably, the expression box expresses the nucleic acid represented by any one of Seq ID No. 11 to 14.

[0027] The present invention also provides a plant genome editing system, vector, or host cell incorporating the expression framework.

[0028] Furthermore, the plant is a grass or a legume.

[0029] Specifically, the grasses mentioned are rice, wheat, barley, corn, or sorghum.

[0030] More preferably, the legume is soybean, peanut, broad bean, pea or adzuki bean.

[0031] This invention also provides the application of the aforementioned expression framework, plant genome editing system, vector, or host cell in plant genome editing.

[0032] Furthermore, the plant is a grass or a legume.

[0033] Specifically, the grasses mentioned are rice, wheat, barley, corn, or sorghum.

[0034] More preferably, the legume is soybean, peanut, broad bean, pea or adzuki bean.

[0035] The present invention also provides a method for gene editing of plant genomes, comprising the following steps: designing crRNA according to the target site, constructing the crRNA into a vector containing the editing system, and transforming the plant.

[0036] Furthermore, the plant is a grass or a legume.

[0037] Specifically, the grasses mentioned are rice, wheat, barley, corn, or sorghum.

[0038] More preferably, the legume is soybean, peanut, broad bean, pea or adzuki bean.

[0039] Specifically, the crRNA is 30-60 nt in length.

[0040] Preferably, the crRNA is 32-56 nt in length.

[0041] The beneficial effects of this invention are as follows: Based on the Type IF CRISPR system, this invention optimizes the PaeCascade sequence in a rice expression system, obtaining an optimized PaeCascade coding sequence. Furthermore, this invention fuses the optimized coding sequence with different transcription factors to construct expression vectors, resulting in a PaeCascade-based transcriptional regulation system. Genome-specific editing experiments were conducted using the constructed PaeCascade system targeting genes in rice. The results showed that the PaeCascade transcriptional activation system can regulate the transcriptional activation level at specific sites, achieving targeted genome editing, and the transcriptional activity is significantly higher than the control (dCas9-TV). This invention further designed crRNAs of different lengths for transformation at the same sites, and the results showed that the transcriptional activation effect was further enhanced with the lengthening of the crRNA. This demonstrates that the PaeCascade-based transcriptional activation system can regulate the transcriptional activation level at specific sites. Therefore, this invention successfully developed a PaeCascade-based system, which can be applied to plant gene regulation tools, further enriching plant genome editing tools and laying the foundation for the widespread application of the PaeCascade system in the field of plant genome editing. Attached Figure Description

[0042] Figure 1 Schematic diagrams of different transcriptional regulatory vector structures based on PaeCascade in this invention.

[0043] Figure 2 Comparison of different transcriptional regulatory vectors based on PaeCascade and dCas9-TV transcriptional activation of the GW7 site in rice.

[0044] Figure 3Comparison of the efficiency of different transcriptional regulatory vectors based on PaeCascade and dCas9-TV transcriptional activation of rice BBM1 and ER1 sites.

[0045] Figure 4 Comparison of activation efficiencies of different transcriptional regulatory vectors based on PaeCascade and dCas9-TV transcriptional activation of different crRNA lengths at BBM1 and ER1 sites. Detailed Implementation

[0046] To promote the application of PaeCascade in plant genome editing, this application first optimized the coding sequences of four PaeCascade effector proteins (Cas5, Cas6, Cas8, and Cas7) using rice and soybean as application targets. Furthermore, three transcriptional activation domains—2xTAD, 2xTAD-VP64, and TV—were selected and combined with Cas6 and Cas7 to construct different backbone vectors. Specifically, these are: pZmUbi-Cas5-Cas8 +p35S-Ca6-Cas7 +pOsU6-DR, pZmUbi-Cas5-Cas8+p35S-Cas6-Cas7-2xTAD / p35S-Cas6-Cas7-2xTAD-VP64 / p35S-Cas6-Cas7-TV +pOsU6-DR, pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV +p35S-Cas8-Cas7 +pOsU6-DR, and pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV+p35S-Cas8 -Cas7-2xTAD / p35S-Cas8 -Cas7-2xTAD-VP64 / p35S-Cas8 -Cas7-TV+pOsU6-DR.

[0047] The aforementioned backbone vector was combined with the known dCas9-TV backbone vector, a backbone vector without the transcriptional activation domain, and an empty vector to conduct single-site and two-site editing experiments (using rice as an example). The results showed that combining Cas7 with three transcriptional activation domains—2xTAD, 2xTAD-VP64, and TV—as well as simultaneously combining 2xTAD with Cas6, effectively improved the activation effect. This invention further designed crRNAs of different lengths for the same site and transformed them; the results showed that as the crRNA length increased, the transcriptional activation effect was further enhanced. This demonstrates that the PaeCascade-based transcriptional activation system can regulate the transcriptional activation level at specific sites.

[0048] Based on the aforementioned experimental results, a series of technical solutions of the present invention were obtained.

[0049] The technical solution of the present invention is a nucleic acid encoding Cas5, the sequence of which is shown in Seq ID No.1.

[0050] The present invention also provides a nucleic acid encoding Cas6, the sequence of which is shown in Seq ID No.2.

[0051] The present invention also provides a nucleic acid encoding Cas8, the sequence of which is shown in Seq ID No.3.

[0052] The present invention also provides a nucleotide encoding Cas7, the sequence of which is shown in Seq ID No.4.

[0053] The present invention also constructs a Type IF CRISPR expression cassette suitable for plant genome editing, including elements expressing Cas5, Cas6, Cas8 and Cas7; the C-terminus of Cas6 and / or Cas7 is also fused with 2xTAD, 2xTAD-VP64 or TV.

[0054] Furthermore, the expression box also includes a Cas6 recognition sequence expression unit.

[0055] Specifically, the structure of the Cas6 recognition sequence expression unit is promoter-DR-LacZ-DR-terminator.

[0056] Specifically, the component structures representing Cas5, Cas6, Cas8, and Cas7 are one of the following:

[0057] a. Promoter - Cas5 - Cas8 - Terminator - Promoter - Cas6 - Cas7 - 2xTAD - Terminator;

[0058] b. Promoter - Cas5-Cas8-Terminator - Promoter - Cas6-Cas7-2xTAD-VP64-Terminator;

[0059] c. Starter - Cas5 - Cas8 - Terminator - Starter - Cas6 - Cas7 - TV - Terminator;

[0060] d. Starter - Cas5-Cas6-2xTAD - Terminator - Starter - Cas8-Cas7-2xTAD - Terminator.

[0061] Furthermore, the N-terminus and / or C-terminus of the Cas6 and / or Cas7 also incorporate a nuclear positioning signal.

[0062] Specifically, the nuclear positioning signal is SV40 NLS, BP NLS, or nucleoplasmin NLS.

[0063] The sequence of the nucleic acid encoding Cas6-2xTAD is shown in Seq ID No. 5.

[0064] The sequence of the nucleic acid encoding Cas7-2xTAD is shown in Seq ID No. 8.

[0065] Specifically, the sequence of the nucleic acid encoding Cas7-2xTAD-VP64 is shown in Seq ID No. 9.

[0066] Specifically, the sequence of the nucleic acid encoding Cas7-TV is shown in Seq ID No. 10.

[0067] Specifically, the promoter is p35S, pZmUbi1, pOsUbi, or pU6.

[0068] The terminator is pinII, AtHSP, NOS, poly T, or OCS.

[0069] Preferably, the expression box expresses the nucleic acid represented by any one of Seq ID No. 11 to 14.

[0070] The present invention also provides a plant genome editing system, vector, or host cell incorporating the expression framework.

[0071] Furthermore, the plant is a grass or a legume.

[0072] Specifically, the grasses mentioned are rice, wheat, barley, corn, or sorghum.

[0073] More preferably, the legume is soybean, peanut, broad bean, pea or adzuki bean.

[0074] This invention also provides the application of the aforementioned expression framework, plant genome editing system, vector, or host cell in plant genome editing.

[0075] Furthermore, the plant is a grass or a legume.

[0076] Specifically, the grasses mentioned are rice, wheat, barley, corn, or sorghum.

[0077] More preferably, the legume is soybean, peanut, broad bean, pea or adzuki bean.

[0078] The present invention also provides a method for gene editing of plant genomes, comprising the following steps: designing crRNA according to the target site, constructing the crRNA into a vector containing the editing system, and transforming the plant.

[0079] Furthermore, the plant is a grass or a legume.

[0080] Specifically, the grasses mentioned are rice, wheat, barley, corn, or sorghum.

[0081] More preferably, the legume is soybean, peanut, broad bean, pea or adzuki bean.

[0082] Specifically, the crRNA is 30-60 nt in length.

[0083] Preferably, the crRNA is 32-56 nt in length.

[0084] The present invention will be further described below with reference to the embodiments. The following embodiments are intended to illustrate the present invention and not to further limit the present invention, and should not be used to limit the scope of protection of the present invention.

[0085] Example 1: Sequence Optimization of PaeCascade

[0086] In a 2020 article published in *Nature Communications* by Chen et al., titled "Repurposing type I–F CRISPR–Cas system as a transcriptional activation tool in human cell," they reported that PaeCascade can achieve transcriptional regulation in mammals. The amino acid sequences of four PaeCascade effector proteins (Cas5, Cas6, Cas8, and Cas7) were obtained based on literature records (see Addgene (pCsy_complex, 89232)). To promote PaeCascade expression in plants, the sequences (Cas5, Cas6, Cas8, and Cas7) were codon-optimized in rice and soybean samples at GenScript, with the specific sequences shown in Seq IDs 1–4. The corresponding gene fragments were then synthesized by Qingke Biotechnology.

[0087] Seq ID No. 1 Cas5 rice codon optimized nucleotide sequence:

[0088] TCTGTCACTGATCCTGAGGCACTGCTCTTACTACCTCGGTTATCCATACAAAATGCAAATGCCATCAGTTCGCCGCTAACATGGGGCTTCCCTTCACCTGGTGCTTTCACTGGATTTGTTCATGCTCTGCAGAGGCGCGTCGGGATATCGCTGGATATTGAGCTTGATGGTGTCGGCATCGTCTGCCACAGATTTGAAGCTCAGATATCTCAACCAGCTGGAAAAAGAACAAAGGTGTTCAACCTCACCAGGAATCCCCTCAACAGAGATGGGAGCACGGCGGCGATTGTTGAAGAAGGAAGGGCTCACCTGGAGGTTTCATTACTTTTAGGGGTGCATGGTGATGGTTTGGATGACCATCCTGCACAAGAAATTGCAAGACAAGTGCAGGAGCAAGCTGGCGCCATGAGGCTCGCCGGCGGCTCCATCCTCCCATGGTGCAATGAAAGGTTCCCGGCACCAAATGCTGAACTTTTGATGCTTGGTGGTTCTGATGAACAACGGAGAAAGAACCAGCGTCGTTTGACGCGTCGCTTGCTCCCCGGGTTTGCTCTTGTAAGTAGAGAAGCATTGTTGCAGCAGCATTTGGAGACATTGCGCACAACTCTGCCGGAAGCCACCACCCTCGACGCGCTCCTTGATCTTTGTAGAATCAACTTTGAGCCGCCGGCCACATCTTCAGAAGAGGAGGCGTCACCACCAGATGCTGCATGGCAGGTTAGGGACAAGCCTGGATGGCTGGTGCCAATTCCTGCTGGTTATAATGCACTTTCCCCACTCTATCTTCCCGGCGAGGTCAGGAACGCGCGCGACAGGGAAACACCATTGAGATTTGTGGAAAATCTTTTTGGATTAGGAGAGTGGCTTAGTCCCCACCGGGTTGCTGCTCTAAGTGACCTCCTATGGTATCATCATGCAGAGCCTGATAAAGGACTCTACAGATGGAGCACTCCCCGTTTCGTGGAGCACGCCATTGCT。

[0089] Seq ID No.2 Nucleotide sequence of codon-optimized Cas6 for rice:

[0090] GATCATTATTTGGATATACGGCTAAGGCCAGATCCAGAGTTCCCGCCGGCGCAGCTCATGTCTGTGCTGTTCGGGAAGCTGCACCAAGCACTTGTCGCACAAGGAGGCGACAGAATTGGTGTGTCATTTCCTGACCTTGATGAATCAAGAAGCAGATTGGGGGAGAGGCTTAGGATTCATGCTTCTGCTGATGATCTTAGGGCTTTGCTCGCCCGGCCGTGGCTTGAAGGACTCCGCGACCATCTTCAGTTTGGAGAGCCTGCTGTTGTTCCTCATCCTACACCATATCGCCAGGTGTCGAGAGTGCAGGCCAAGTCAAATCCTGAAAGGCTACGTCGCCGCTTGATGAGGAGGCATGACTTGAGTGAAGAGGAAGCAAGGAAGAGAATCCCAGATACAGTGGCACGGGCCCTGGACTTGCCATTTGTTACTTTACGAAGCCAAAGTACTGGCCAGCACTTCCGCCTCTTCATCAGGCATGGTCCCCTTCAAGTCACGGCGGAGGAGGGCGGGTTCACATGCTATGGCTTGTCCAAAGGTGGTTTTGTGCCATGGTTT。

[0091] Seq ID No.3 Nucleotide sequence of codon-optimized Cas8 for rice:

[0092]

[0093] Seq ID No. 4, Cas7 rice codon optimized nucleotide sequence:

[0094]

[0095] Based on the 2021 paper "CRISPR–Act3.0 for highly efficient multiplexed gene activation in plants" by Pan et al., the nucleotide sequences of three transcriptional activation domains—2xTAD (representing two TADs tandemly), 2xTAD-VP64, and TV—were obtained. 2xTAD, 2xTAD-VP64, and TV were fused to the C-terminus of Cas6 and Cas7, respectively, constructing six elements: Cas6-2xTAD (Seq ID No. 5), Cas6-2xTAD-VP64 (Seq ID No. 6), Cas6-TV (Seq ID No. 7), Cas7-2xTAD (Seq ID No. 8), Cas7-2xTAD-VP64 (Seq ID No. 9), and Cas7-TV (Seq ID No. 10).

[0096] Seq ID No. 5, Cas6-2xTAD nucleotide sequence; where positions 1-48 are SV40NLS, positions 49-606 are the Cas6 coding sequence, positions 607-654 are nucleoplasmin NLS, and positions 697-1017 are the 2xTAD coding sequence;

[0097]

[0098] Seq ID No. 6, Cas6-2xTAD-VP64 nucleotide sequence; where positions 1-48 are SV40NLS, positions 49-606 are the Cas6 coding sequence, positions 607-654 are nucleoplasmin NLS, and positions 697-1173 are the 2xTAD-VP64 coding sequence;

[0099]

[0100] Seq ID No. 7, Cas6-TV nucleotide sequence; where positions 1-48 are SV40NLS, positions 49-606 are the Cas6 coding sequence, positions 607-654 are the nucleoplasmin NLS, and positions 697-2040 are the TV coding sequence;

[0101]

[0102] Seq ID No. 8, Cas7-2xTAD nucleotide sequence; where positions 1-54 are SV40 NLS; positions 55-1077 are Cas7 coding sequence; positions 1078-1131 are BP NLS; and positions 1174-1494 are 2xTAD coding sequence.

[0103]

[0104] Seq ID No. 9, Cas7-2xTAD-VP64 nucleotide sequence; where positions 1-54 are SV40 NLS, positions 55-1077 are Cas7 coding sequence, positions 1078-1131 are BP NLS, and positions 1174-1650 are 2xTAD-VP64 coding sequence;

[0105]

[0106] Seq ID No. 10, Cas7-TV nucleotide sequence; where positions 1-54 are SV40 NLS, positions 55-1077 are Cas7 coding sequence, positions 1078-1131 are BP NLS, and positions 1174-2517 are TV coding sequence;

[0107]

[0108] Example 2: Construction of a backbone vector without transcription activation domain

[0109] The transcriptional activation domain-free backbone vector consists of three parts. The first part uses the maize promoter ZmUbi to promote the expression of Cas5 and Cas8, i.e., pZmUbi-Cas5-Cas8; the second part uses the CaMV 35S promoter to promote the expression of Cas6 and Cas7 proteins, i.e., p35S-Cas6-Cas7; and the third part uses the rice U6 promoter to promote the expression of DR (Cas6) and crRNA. DR (Cas6) refers to the direct repeat sequence (DR sequence) that can be recognized by Cas6, i.e., pOsU6-DR.

[0110] Step 1: Using the BsaI restriction site, Cas5 and Cas8 were ligated into the pTSWA plasmid via Golden Gate cloning to construct pZmUbi-Cas5-Cas8; Cas6 and Cas7 were ligated into the TSWB plasmid to construct p35S-Ca6-Cas7; the PaeCascade crRNA backbone (Direct Repeat, DR sequence) was derived from the endogenous CRISPR repeat sequence of Pseudomonas aeruginosa. The DR sequence, along with lacZα and its associated elements (crRNA replacement sequences), were inserted into the pTSWA vector using AscI and SacI to obtain pOsU6-DR.

[0111] Step 2: Using the AarI restriction site, pZmUbi-Cas5-Cas8, p35S-Ca6-Cas7, and pOsU6-DR were ligated into the pTRANS_210d backbone vector via Golden Gate cloning to obtain the backbone vector pZmUbi-Cas5-Cas8 + p35S-Ca6-Cas7 + pOsU6-DR. Figure 1 Group I in the middle.

[0112] Example 3: Construction of a backbone vector for expression of transcription activation domain fused with Cas7 protein

[0113] Step 1: The three transcriptional activation domains (2xTAD, 2xTAD-VP64, and TV) were fused with the Cas7 protein to obtain Cas7-2xTAD, Cas7-2xTAD-VP64, and Cas7-TV, respectively. Using the BsaI restriction site, Cas6 and Cas7-2xTAD / Cas7-2xTAD-VP64 / Cas7-TV were ligated into the TSWB plasmid via Golden Gate cloning to construct p35S-Ca6-Cas7-2xTAD / p35S-Ca6-Cas7-2xTAD-VP64 / p35S-Ca6-Cas7-TV.

[0114] Step 2: Using the AarI restriction site, pZmUbi-Cas5-Cas8 and pOsU6-DR from Example 2 were ligated into the pTRANS_210d backbone vector via Golden Gate cloning, respectively, along with p35S-Ca6-Cas7-2xTAD / p35S-Ca6-Cas7-2xTAD-VP64 / p35S-Ca6-Cas7-TV, to obtain a backbone vector expressing three transcriptional activation domains fused with Cas7: pZmUbi-Cas5-Cas8+p35S-Cas6-Cas7-2xTAD / p35S-Cas6-Cas7-2xTAD-VP64 / p35S-Cas6-Cas7-TV +pOsU6-DR ( Figure 1 Group III in the middle.

[0115] Example 4: Construction of a backbone vector for expression of transcription activation domain fused with Cas6 protein

[0116] Step 1: The three transcriptional activation domains (2xTAD, 2xTAD-VP64, and TV) were fused with the Cas6 protein to obtain Cas6-2xTAD, Cas6-2xTAD-VP64, and Cas6-TV, respectively. Using the BsaI restriction site, the maize ZmUbi promoter, Cas5, and Cas6-2xTAD / Cas6-2xTAD-VP64 / Cas6-TV were ligated into the TSWA plasmid via Golden Gate cloning to construct pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV. Using the BsaI restriction site, Cas8 and Cas7 were ligated into the pTSWB plasmid via Golden Gate cloning to construct p35S-Cas8-Cas7.

[0117] Step 2: Using the AarI restriction site, p35S-Cas8-Cas7 from Step 1 and pOsU6-DR from Example 2 were ligated into the pTRANS_210d backbone vector via Golden Gate cloning, respectively, to pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV, resulting in backbone vectors expressing three transcriptional activation domains fused with Cas6: pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV + p35S-Cas8-Cas7 + pOsU6-DR. Figure 1 Group in ).

[0118] Example 5: Construction of a backbone vector for expression of transcription activation domain fused with Cas6 and Cas7 proteins.

[0119] Construct plasmids pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV+p35S-Cas8-Cas7-2xTAD / p35S-Cas8-Cas7-2xTAD-VP64 / p35S-Cas8-Cas7-TV+pOsU6-DR.

[0120] Step 1: The three transcriptional activation domains (2xTAD, 2xTAD-VP64, and TV) were fused with Cas6 to obtain Cas6-2xTAD, Cas6-2xTAD-VP64, and Cas6-TV, respectively. Using the BsaI restriction site, the maize ZmUbi promoter, Cas5, and Cas6-2xTAD / Cas6-2xTAD-VP64 / Cas6-TV were linked to the pTSWA plasmid via Golden Gate cloning to construct pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV.

[0121] Step 2: The three transcriptional activation domains (2xTAD, 2xTAD-VP64, and TV) were fused with Cas7 to obtain Cas7-2xTAD, Cas7-2xTAD-VP64, and Cas7-TV, respectively. Using the BsaI restriction site, the 35S promoter, Cas8, and Cas7-2xTAD / Cas7-2xTAD-VP64 / Cas7-TV were ligated into the TSWB plasmid via Golden Gate cloning to construct p35S-Cas8-Cas7-2xTAD / p35S-Cas8-Cas7-2xTAD-VP64 / p35S-Cas8-Cas7-TV.

[0122] Step 3: Using the AarI restriction site, the unit with the same transcriptional activation domain from Step 1 and Step 2 was linked to pOsU6-DR in the pTRANS_210d backbone vector via Golden Gate cloning to obtain a backbone vector expressing three transcriptional activation domains fused with Cas6 and Cas7: pZmUbi-Cas5-Cas6-2xTAD / pZmUbi-Cas5-Cas6-2xTAD-VP64 / pZmUbi-Cas5-Cas6-TV+p35S-Cas8-Cas7-2xTAD / p35S-Cas8-Cas7-2xTAD-VP64 / p35S-Cas8-Cas7-TV+pOsU6-DR ( Figure 1 Group IV in the middle.

[0123] Example 6: Construction of the PaeCascade transcriptional activation backbone vector based on SunTag

[0124] The plasmids p35S-Cas7-10xGCN4+pZmUbi-Cas8-Cas5-Cas6+pOsUbi1-scFv-sfGFP-2xTAD+OsU6-DR, p35S-Cas7-10xGCN4+pZmUbi-Cas8-Cas5-Cas6+pOsUbi1-scFv-sfGFP-2xTAD-VP64+OsU6-DR and p35S-Cas7-10xGCN4+pZmUbi-Cas8-Cas5-Cas6+pOsUbi1-scFv-sfGFP-TV+OsU6-DR were constructed.

[0125] Step 1: Cas7 and 10xGCN4 were fused and expressed separately. Using the BsaI restriction site, the 35S promoter, Cas7 and 10xGCN4 were ligated into the pTSWA plasmid to construct p35S-Cas7-10xGCN4 via Golden Gate cloning.

[0126] Step 2: Cas8, Cas5 and Cas6 are fused and expressed. Using the BsaI restriction site, the maize ZmUbi promoter, Cas8, Cas5 and Cas6 are linked to the pTSWB plasmid to construct pZmUbi-Cas8-Cas5-Cas6 by Golden Gate cloning.

[0127] Step 3: The three transcriptional activation domains (2xTAD, 2xTAD-VP64, and TV) were fused with scFv-sfGFP to obtain scFv-sfGFP-2xTAD / scFv-sfGFP-2xTAD-VP64 / scFv-sfGFP-TV. Using the BsaI restriction site, the rice Ubi1 promoter was linked to the above three fusion proteins into the TSWB plasmid via Golden Gate cloning to construct pOsUbi1-scFv-sfGFP-2xTAD / pOsUbi1-scFv-sfGFP-2xTAD-VP64 / pOsUbi1-scFv-sfGFP-TV.

[0128] Step 4: Insert the DR sequence and lacZα and its associated elements (crRNA replacement sequence) into the pTSWD vector using SgsI and SacI to obtain pOsU6-DR.

[0129] Step 5: Using the AarI restriction site, p35S-Cas7-10xGCN4 from Step 1, pZmUbi-Cas8-Cas5-Cas6 from Step 2, pOsU6-DR obtained in Step 4, and pOsUbi1-scFv-sfGFP-2xTAD / pOsUbi1-scFv-sfGFP-2xTAD-VP64 / pOsUbi1-scFv-sfGFP-TV from Step 3 were ligated into the pTRANS_210d backbone vector via Golden Gate cloning to obtain three SunTag-based PaeCascade transcription activation backbone vectors. Figure 1 Group V in the middle.

[0130] Seq ID No. 11, nucleotide sequence of the Cas7-2xTAD expression unit; where positions 1-1998 are the ZmUbi promoter sequence, positions 2008-3084 are the Cas5 coding sequence, positions 3094-3150 are the P2A sequence, positions 3154-4560 are the Cas8 coding sequence, positions 4575-4824 are the HSP terminator, positions 4851-5683 are the 35S promoter, positions 5702-6358 are the Cas6 coding sequence, and positions 6368... Positions ~6424 are the P2A sequence, positions 6482~7504 are the Cas7 coding sequence, positions 7601~7921 are the 2xTAD coding sequence, positions 8211~8519 are the pinII terminator, positions 8542~8786 are the rice U6 promoter, positions 8787~8814 are the DR sequence, positions 8822~9380 are lazα and related elements, positions 9381~9408 are the DR sequence, and positions 9409~9415 are the transcription termination sequence.

[0131]

[0132] The nucleotide sequence of the Seq ID No. 12 Cas7-2xTAD-VP64 expression unit was obtained by replacing the Cas7-2xTAD sequence at positions 6482-7921 of the Seq ID No. 11 Cas7-2xTAD expression unit nucleotide sequence with the Seq ID No. 9 Cas7-2xTAD-VP64 sequence, while keeping the rest of the sequence unchanged. The structure of No. 12 is as follows: bits 1-1998 are the ZmUbi promoter sequence, bits 2008-3084 are the Cas5 encoding sequence, bits 3094-3150 are the P2A sequence, bits 3154-4560 are the Cas8 encoding sequence, bits 4575-4824 are the HSP terminator, bits 4851-5683 are the 35S promoter, bits 5702-6358 are the Cas6 encoding sequence, and bits 6368-6424 are the P2A sequence. Positions 6482–7504 are the Cas7 coding sequence, positions 7601–8077 are the 2xTAD-VP64 coding sequence, positions 8373–8681 are the pinII terminator, positions 8704–8948 are the rice U6 promoter, positions 8949–8976 are the DR sequence, positions 8984–9542 are lazα and its related elements, positions 9543–9570 are the DR sequence, and positions 9571–9577 are the transcription termination sequence.

[0133]

[0134]

[0135] Seq ID No. 14, nucleotide sequence of the Cas6-2xTAD+Cas7-2xTAD expression unit; where positions 1-1998 are the ZmUbi promoter sequence, positions 2008-3084 are the Cas5 coding sequence, positions 3094-3150 are the P2A sequence, positions 3154-3807 are the Cas6 coding sequence, positions 3850-4169 are the 2xTAD coding sequence, positions 4461-4710 are the HSP terminator, positions 4737-5569 are the 35S promoter, and positions 5591-6997 are... The Cas8 coding sequence has the following positions: 7007-7063 is the P2A sequence; 7067-8197 is the Cas7 coding sequence; 8240-8560 is the 2xTAD coding sequence; 8850-9158 is the pinII terminator; 9181-9425 is the rice U6 promoter; 9426-9453 is the DR sequence; 9461-10019 is lazα and its related elements; 10020-10047 is the DR sequence; and 10048-10054 is the transcription termination sequence.

[0136]

[0137] Example 7: Rice Single-Endogenous Site Activation Based on PaeCascade Transcriptional Activation System

[0138] Referring to the rice GW7 gene activation site used in the paper "CRISPR–Act3.0 for highly efficient multiplexed gene activation in plants" published by Pan et al. in Nature Plants in 2021, the last 32 bp of OsGW7 that satisfies the PAM 5'-CC-3' characteristic was searched as the target sequence (crRNA01-OsGW7: 5'-ATGTCCAGCTCCCGGCACCCACGGCCGAATGC-3') at 100-200 bp upstream of OsGW7. The target sequence was inserted into the above backbone vectors by BsaI digestion to construct crRNA target plasmids: Cas5-Cas8+Ca6-Cas7-crRNA01-OsGW7 (using the backbone vector in Example 2); Cas7-2xTAD-crRNA01-OsGW7, Cas7-2xTAD-VP64-crRNA01-OsGW7, Cas7-TV-crRNA01-OsGW7 (using the backbone vector in Example 3); Cas6-2xTAD-crRNA01-OsGW7, Cas6-2xTAD-VP64-crRNA01-OsGW7, Cas6-TV-crRNA01-OsGW7 (using the backbone vector in Example 4); Cas6-2xTAD +Cas7-2xTAD-crRNA01-OsGW7, Cas6-2xTAD-VP64+Cas7-2xTAD-VP64-crRNA01-OsGW7, Cas6-TV+Cas7-TV-crRNA01-OsGW7 (using the backbone vector in Example 5); Cas7-SunTag-2xTAD-crRNA01-OsGW7, Cas7-SunTag-2xTAD-VP64-crRNA01-OsGW7, Cas7-SunTag-TV-crRNA01-OsGW7 (using the backbone vector in Example 6).Using a PEG-mediated rice protoplast preparation and transformation system (Zhang et al., A highly efficient rice green tissue protoplast system for transient gene expression and studying light / chloroplast-related processes. Plant Methods, 2011, 7(1): 30), 40 μg of crRNA targeting plasmid was transiently transformed into rice protoplasts, and RNA was extracted. RT-qPCR was performed on the target gene GW7. It was found that except for Cas5-Cas8+Ca6-Cas7-crRNA01-OsGW7 (which does not contain a transcription activation domain), the other PaeCascade-based transcription activation vectors could significantly express the target gene OsGW7. Moreover, the expression vectors Cas7-2xTAD-crRNA01-OsGW7, Cas7-2xTAD-VP64-crRNA01-OsGW7, Cas7-TV-crRNA01-OsGW7, and Cas6-2xTAD+Cas7-2xTAD-crRNA01-OsGW7 showed better activation effects. Figure 2 The backbone vectors corresponding to Cas7-2xTAD-crRNA01-OsGW7, Cas7-2xTAD-VP64-crRNA01-OsGW7, Cas7-TV-crRNA01-OsGW7, and Cas6-2xTAD+Cas7-2xTAD-crRNA01-OsGW7 are pZmUbi-Cas5-Cas8+p35S-Cas6-Cas7-2xTAD / p35S-Cas6-Cas7-2xTAD-VP64 / p35S-Cas6-Cas7-TV +pOsU6-DR and pZmUbi-Cas5-Cas6-2xTAD+p35S-Cas8-Cas7-2xTAD +pOsU6-DR, respectively.

[0139] Figure 2CTRL is a backbone vector without assembly sites (i.e., pZmUbi-Cas5-Cas8+p35S-Ca6-Cas7+pOsU6-DR in Example 2); cascade is Cas5-Cas8+Ca6-Cas7-crRNA01-OsGW7; the dCas9-TV backbone vector is derived from the paper CRISPR–Act3.0 for highly efficient multiplexed gene activation in plants published by Pan et al. in Nature Plants in 2021. The target sequences (sgRNA-OsER1:5'-GTTCTACTCCTACCAACTCT-3'; sgRNA-OsBBM1:5'-CTAGTCTCAGCAAATAGCAG-3') were inserted into the backbone vector by BsaI restriction enzyme digestion.

[0140] Table 1. Backbone vectors corresponding to the expression vectors used in Examples 7 and 8.

[0141] .

[0142] Example 8: Simultaneous activation of two endogenous sites in rice mediated by the PaeCascade-based transcriptional activation system

[0143] Referring to the rice ER1 and BBM1 gene activation sites used in the 2021 paper "CRISPR–Act3.0 for highly efficient multiplexed gene activation in plants" published by Pan et al. in Nature Plants, the last 32 bp of OsER1 and OsBBM1 that meet the PAM 5'-CC-3' characteristics were searched as target sequences (crRNA01-OsER1: 5'-TCGACTTGCGACTCGAGTTCTACTCCTACCAA-3'; crRNA01-OsBBM1: 5'-TTTCTAGTCTCAGCAAATAGCAGAGGTAGAGA-3'). The target sequences were inserted into the above backbone vectors using BsaI digestion to construct the following crRNA target plasmids: Cas5-Cas8+Ca6-Cas7-crRNA01-OsER1-crRNA01-OsBBM1, Cas7-2xTAD-crRNA01-OsER1-crRNA01-OsBBM1, Cas7-2xTAD-VP64-crRNA01-OsER1-crRNA01-OsBBM1, Cas7-TV-crRNA01-OsER1-crRNA01-OsBBM1, Cas6-2xTAD-crRNA01-OsER1-crRNA01-OsBBM1, Cas6-2xTAD-VP64-crRNA01-OsER1-crRNA01-OsBBM1, C as6-TV-crRNA01-OsER1-crRNA01-OsBBM1, Cas6-2xTAD+Cas7-2xTAD-crRNA01-OsER1-crRNA01-OsBBM1, Cas6-2xTAD-VP64+Cas7-2xTAD-VP64-crRNA01-OsER1-crRNA01-OsBBM1, Cas6-TV+Cas7-TV-crRNA01-OsER1-crRNA01-OsBBM1, Cas7-SunTag-2xTAD-crRNA01-OsER1-crRNA01-OsBBM1, Cas7-SunTag-2xTAD-VP64-crRNA01-OsER1-crRNA01-OsBBM1, Cas7- SunTag-TV-crRNA01-OsER1-crRNA01-OsBBM1.Using a PEG-mediated rice protoplast preparation and transformation system, 40 μg of crRNA targeting plasmid was transiently transformed into rice protoplasts, and RNA was extracted. RT-qPCR was performed on the target genes ER1 and BBM1. It was found that, except for Cas5-Cas8+Ca6-Cas7-crRNA01-OsER1-crRNA01-OsBBM1 (this vector does not contain a transcription activation domain), all other PaeCascade-based transcription activation vectors could activate the expression of target genes ER1 and BBM1, consistent with the results of the GW7 single-site transcription activation experiment. The Cas7-2xTAD and Cas6-2xTAD+Cas7-2xTAD backbone vectors showed the best activation effects. Figure 3 ).

[0144] Example 9: Adjustable transcriptional activation mediated by the PaeCascade-based transcriptional activation system

[0145] The PaeCascade-based transcriptional activation system can induce different degrees of transcriptional activation by altering the length of crRNA. Based on the experimental results of Examples 7 and 8, four backbone vectors were selected: pZmUbi-Cas5-Cas8+p35S-Cas6-Cas7-2xTAD / p35S-Cas6-Cas7-2xTAD-VP64 / p35S-Cas6-Cas7-TV +pOsU6-DR and pZmUbi-Cas5-Cas6-2xTAD+p35S-Cas8-Cas7-2xTAD +pOsU6-DR. For the ER1 and BBM1 genes, 50nt and 56nt crRNA sequences were designed. Using the method described in Example 8, the 32nt crRNA was ligated into the PaeCascade-based transcriptional activation vector. Protoplast transformation revealed that the transcriptional activation effect further increased with the length of the crRNA. This demonstrates that the PaeCascade-based transcriptional activation system can regulate the transcriptional activation level at specific sites. Figure 4 ).

Claims

1. The nucleic acid encoding Cas5, the sequence of which is shown in Seq ID No. 1; or, the nucleic acid encoding Cas6, the sequence of which is shown in Seq ID No. 2; or, the nucleic acid encoding Cas8, the sequence of which is shown in Seq ID No. 3; or, the nucleic acid encoding Cas7, the sequence of which is shown in Seq ID No.

5.

2. A Type IF CRISPR expression cassette suitable for plant genome editing, characterized in that: It includes Cas5, Cas6, Cas8 and Cas7 elements; the C-terminus of Cas6 and / or Cas7 is also fused with 2xTAD, 2xTAD-VP64 or TV; the sequence of the nucleic acid encoding Cas5 is shown in Seq ID No.1; the sequence of the nucleic acid encoding Cas6 is shown in Seq ID No.2; the sequence of the nucleic acid encoding Cas8 is shown in Seq ID No.3; the sequence of the nucleic acid encoding Cas7 is shown in Seq ID No.

4.

3. The expression box according to claim 2, characterized in that: The expression box also includes a Cas6 recognition sequence expression unit; Preferably, the structure of the Cas6 recognition sequence expression unit is promoter-DR-LacZ-DR-terminator.

4. The expression box according to claim 2, characterized in that: The component structures representing Cas5, Cas6, Cas8, and Cas7 are one of the following: a. Promoter - Cas5 - Cas8 - Terminator - Promoter - Cas6 - Cas7 - 2xTAD - Terminator; b. Promoter - Cas5-Cas8-Terminator - Promoter - Cas6-Cas7-2xTAD-VP64-Terminator; c. Starter - Cas5 - Cas8 - Terminator - Starter - Cas6 - Cas7 - TV - Terminator; d. Starter - Cas5-Cas6-2xTAD - Terminator - Starter - Cas8-Cas7-2xTAD - Terminator.

5. The expression box according to claim 2, characterized in that: The N-terminus and / or C-terminus of the Cas6 and / or Cas7 also incorporate a nuclear positioning signal; Preferably, the nuclear positioning signal is SV40 NLS, BP NLS or nucleoplasmin NLS.

6. The expression box according to claim 5, characterized in that: One of the following: The sequence of the nucleic acid encoding Cas6-2xTAD is shown in Seq ID No. 5; The sequence of the nucleic acid encoding Cas7-2xTAD is shown in Seq ID No. 8; The sequence of the nucleic acid encoding Cas7-2xTAD-VP64 is shown in Seq ID No. 9; The sequence of the nucleic acid encoding Cas7-TV is shown in Seq ID No.

10.

7. The expression box according to any one of claims 3 to 6, characterized in that: The promoter is p35S, pZmUbi1, pOsUbi, or pU6; Alternatively, the terminator may be pinII, AtHSP, NOS, poly T, or T35S; Preferably, the expression box expresses the nucleic acid represented by any one of Seq ID No. 11 to 14.

8. A plant genome editing system, vector, or host cell comprising the expression framework described in any one of claims 1 to 7; Preferably, the plant is a grass or a legume; More preferably, the grass is rice, wheat, barley, corn or sorghum; More preferably, the legume is soybean, peanut, broad bean, pea or adzuki bean.

9. The expression framework according to any one of claims 1 to 7, the plant genome editing system of claim 8, the vector or host cell, and their application in plant genome editing.

10. A method for gene editing of a plant genome, characterized in that: The process includes the following steps: designing crRNA based on the target site, constructing the crRNA into a vector containing the plant genome editing system of claim 8, and transforming the plant; Preferably, the crRNA is 30-60 nt in length; More preferably, the crRNA is 32-56 nt long.