Editing guidance system and application thereof in plant genome editing

By introducing fusion proteins into the PE system, reducing RDR6 protein activity, and designing specific pegRNAs, the problem of low PE efficiency in Arabidopsis thaliana was solved, enabling efficient gene editing and the creation of transgenic mutants, thus improving the accuracy of gene function research.

CN120843508APending Publication Date: 2025-10-28BEIJING NORMAL UNIV AT ZHUHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510513498.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-17
Filing Date
2025-04-23
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing prime editors (PEs) are inefficient in dicotyledonous plants such as Arabidopsis thaliana, limiting their widespread application in plant gene editing.

Method used

A complete system is provided, including a fusion protein, a substance that reduces the content and/or activity of RDR6 protein, sgRNA and pegRNA. High-precision genome editing is performed by designing specific pegRNA, and combined with a screening marker protein for screening transgenic plants.

Benefits of technology

It significantly improves the editing efficiency of the PE system in Arabidopsis thaliana, enabling the creation of transgenic mutants whose targeted editing is heritable up to the T2 generation, and providing a more precise and powerful tool for gene function research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005371915110000091
    Figure BDA0005371915110000091
  • Figure HDA0005371915120000011
    Figure HDA0005371915120000011
  • Figure HDA0005371915120000012
    Figure HDA0005371915120000012
Patent Text Reader

Abstract

The invention discloses an editing guiding system and application thereof in plant genome editing. The guided editing system comprises a fusion protein, a substance for reducing the content and / or activity of RDR6 protein, sgRNA and pegRNA, the fusion protein is formed by fusing nuclease and reverse transcriptase; the sgRNA is used for generating a non-editing strand incision; the pegRNA comprises sgRNA ', a reverse transcription template sequence and a primer binding site sequence, and the sgRNA' is used for editing a chain target sequence in a targeted manner. The guiding editing system provided by the invention breaks through the application bottleneck of a PE system in arabidopsis thaliana, greatly improves the PE editing efficiency of an arabidopsis thaliana target, and provides a more accurate and powerful tool for arabidopsis thaliana gene function research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, specifically to a guided editing system and its application in plant genome editing. Background Art

[0002] Prime editors (PEs) are a novel, precise gene-editing technology developed in recent years based on the CRISPR / Cas9 system, showing broad application prospects in the field of plant gene editing. The PE system achieves high-precision genome editing by fusing a nuclease-deficient Cas9n (H840A) with reverse transcriptase (M-MLV RT) and designing specific pegRNAs (the 3' end of an elongated sgRNA containing primer-binding sequences and reverse transcription template sequences). Compared to the traditional CRISPR / Cas9 system, the PE system offers higher editing accuracy and significantly reduces off-target effects.

[0003] Currently, PE has achieved initial success in the application of monocotyledonous plants such as rice, wheat, and maize. In contrast, the efficiency of PE in dicotyledonous plants such as Arabidopsis thaliana remains very low, which greatly limits its widespread application. Summary of the Invention

[0004] One object of the present invention is to provide a complete system.

[0005] The complete system provided by this invention includes a fusion protein, a substance that reduces the content and / or activity of RDR6 protein, sgRNA, and pegRNA.

[0006] The fusion protein is formed by the fusion of a nuclease and a reverse transcriptase;

[0007] The sgRNA is used to generate non-editable strand cuts;

[0008] The pegRNA includes sgRNA', a reverse transcription template sequence, and a primer binding site sequence; the sgRNA' targets the editing strand target sequence.

[0009] In the above-described complete system, the reverse transcriptase is fused to the C-terminus of the nuclease in the fusion protein.

[0010] In the above-mentioned complete system, the nuclease is Cas9 nuclease.

[0011] Furthermore, the Cas9 nuclease is a Cas9 nicking enzyme. The Cas9 nicking enzyme can be any Cas9n or its variants known in the prior art, including bacterial Cas9n (such as SpCas9n, SaCas9n, SaCas9n-KKH, etc.), SpCas9 variant nicking enzymes that recognize different PAMs (such as xCas9n, Cas9n-NG, Cas9n-VQR, Cas9n-VRER, etc.), and Cas9 high-fidelity enzyme variant nicking enzymes (such as HypaCas9n, eSpCas9(1.1)n, Cas9-HF1n, etc.).

[0012] In some embodiments, the Cas9 nicking enzyme is Cas9n(H840A / R221K / N394K); the Cas9n(H840A / R221K / N394K) is either A1 or A2.

[0013] A1) The amino acid sequence is that of the protein shown in sequence 3;

[0014] A2) Proteins with the same function obtained by substituting and / or deleting and / or adding one or more amino acid residues of the amino acid sequence shown in Sequence 3.

[0015] Furthermore, the coding gene for Cas9n(H840A / R221K / N394K) is a1), a2), or a3):

[0016] a1) The DNA molecule shown at positions 4204-8304 of sequence 1;

[0017] a2) has 75% or more identity with the nucleotide sequence defined in a1) and encodes the DNA molecule of said Cas9n (H840A / R221K / N394K);

[0018] a3) hybridizes under stringent conditions with the nucleotide sequence defined by a1) or a2) and encodes the DNA molecule of said Cas9n(H840A / R221K / N394K).

[0019] In the above-described complete system, the reverse transcriptase is a truncated reverse transcriptase. The truncated reverse transcriptase is either B1 or B2.

[0020] B1) The amino acid sequence of the protein is shown in sequence 4;

[0021] B2) Proteins with the same function obtained by substituting and / or deleting and / or adding one or more amino acid residues of the amino acid sequence shown in Sequence 4.

[0022] The gene encoding the truncated reverse transcriptase is b1), b2), or b3.

[0023] b1) The DNA molecule shown in positions 8650-10191 of sequence 1;

[0024] b2) has 75% or more identity with the nucleotide sequence defined in b1) and encodes the DNA molecule of the truncated reverse transcriptase;

[0025] b3) Hybridizes under stringent conditions to the nucleotide sequence defined by b1) or b2) and encodes the DNA molecule of the truncated reverse transcriptase.

[0026] In the above-mentioned complete system, the substance that reduces the activity of RDR6 protein can be a protein, polypeptide, or small molecule compound that inhibits the function of RDR6 protein.

[0027] The substance that reduces RDR6 protein content may be a substance that inhibits RDR6 protein synthesis, promotes RDR6 protein degradation, or knocks down (reduces) or eliminates the RDR6 protein encoding gene.

[0028] Furthermore, the substance that knocks down (reduces) the RDR6 protein-coding gene can be any nucleic acid molecule that can inhibit (or interfere with) the expression of the RDR6 protein-coding gene, such as gRNA (e.g., sgRNA), siRNA, dsRNA, shRNA, miRNA, antisense RNA, etc.

[0029] The substance used to knock out the RDR6 protein-coding gene can be any substance that prevents the host cell from producing the functional protein product of this gene. Specific methods include removing all or part of the coding gene sequence, introducing mutations to prevent the production of the functional protein, removing or altering regulatory components (e.g., promoter editing) to prevent transcription of the coding gene sequence, or blocking translation by binding to mRNA. Typically, knockout is performed at the genomic DNA level, ensuring that the cell's offspring permanently carry the knockout.

[0030] Furthermore, the nucleic acid molecule that inhibits the expression of the RDR6 protein-coding gene is a miRNA that inhibits the expression of the RDR6 protein-coding gene.

[0031] In some embodiments, the nucleotide sequence of the miRNA that inhibits the expression of the RDR6 protein-encoding gene is shown in Sequence 7 or Sequence 8.

[0032] In the above-described system, the sgRNA is esgRNA, which consists of a target sequence (denoted as target sequence A) and an esgRNA backbone. The esgRNA backbone is an RNA molecule obtained by replacing the T with U in positions 11122-11207 of sequence 1. This esgRNA is used to generate non-editable strand cuts, and the non-editable strand cut sites can be arbitrarily selected.

[0033] In the above-mentioned complete system, the pegRNA includes, in sequence, esgRNA', reverse transcription template sequence (RT sequence) and primer binding site sequence (PBS sequence).

[0034] The esgRNA' consists of a target sequence (denoted as target sequence B) and an esgRNA' backbone. The esgRNA' backbone is an RNA molecule obtained by replacing T with U in positions 12306-12391 of sequence 1. Target sequence B and the aforementioned target sequence A are located on the two strands of the target DNA, respectively. They may be complementary or partially complementary, or they may be at a certain distance. This esgRNA' targets the editing strand target sequence, which can be any sequence in the genome sequence, preferably the target gene target sequence.

[0035] The RT sequence is the reverse complementary sequence of the three bases at the 3' end of the target sequence (target sequence B) and the subsequent continuous genomic sequence, in which a target mutation is introduced. This mutation serves as a reverse transcription template for reverse transcriptase, producing cDNA, which is then used as a repair template to repair genomic DNA. The RT sequence size can further be 8-34 bp.

[0036] The PBS sequence (primer binding site sequence) is the reverse complementary sequence (1≤n<17) of the target sequence (target sequence B) from the nth to the 17th base of the 5' end of the target sequence.

[0037] The design methods or principles of the RT sequence and the PBS sequence can refer to the design methods or principles related to the RT sequence and PBS sequence of pegRNA in prime editing (PE) technology that have been reported in the prior art.

[0038] In some embodiments, the pegRNA further includes a tevopreQ1 motif, wherein the pegRNA comprises, in sequence, an esgRNA', a reverse transcription template sequence (RT sequence), a primer binding site sequence (PBS sequence), and the tevopreQ1 motif. The nucleotide sequence of the tevopreQ1 motif is shown at positions 12425-12461 of Sequence 1.

[0039] In some embodiments, the pegRNA further includes tRNA and HDV sequences flanking it. Preferably, the 5' end of the pegRNA includes a tRNA sequence, and the 3' end includes an HDV sequence. The tRNA and HDV sequences are used together to cleave and generate pegRNA. The tRNA sequence is shown as positions 12209-12285 of Sequence 1; the HDV sequence is shown as positions 12462-12529 of Sequence 1.

[0040] The aforementioned complete system also includes screening marker proteins. These screening marker proteins are used to screen transgenic plants or transgenic plant cells, including but not limited to color screening marker proteins (such as red fluorescent protein, green fluorescent protein, anthocyanin-related protein, etc.) and resistance screening marker proteins (such as kanamycin resistance protein, herbicide phosphinic acid resistance protein, hygromycin resistance protein, methotrexate resistance protein, glyphosate resistance protein, etc.).

[0041] In some embodiments, the screening marker protein is the screening marker red fluorescent protein DsRed.

[0042] The screening marker, red fluorescent protein DsRed, is either C1 or C2.

[0043] C1) The amino acid sequence is the protein shown in sequence 2;

[0044] C2) A protein with the same function as the amino acid sequence shown in Sequence 2, by substitution and / or deletion and / or addition of one or more amino acid residues.

[0045] The gene encoding the screening marker, the red fluorescent protein DsRed, is c1), c2), or c3):

[0046] c1) A DNA molecule that is reverse complementary to positions 292-975 of sequence 1;

[0047] c2) has 75% or more identity with the nucleotide sequence defined by c1) and encodes the DNA molecule of the selection marker red fluorescent protein DsRed;

[0048] c3) hybridizes under stringent conditions with the nucleotide sequence defined by c1) or c2) and encodes the DNA molecule of the selection marker red fluorescent protein DsRed.

[0049] In any of the proteins described above, the substitution and / or deletion and / or addition of one or more amino acid residues may be as follows: substitution and / or deletion and / or addition of no more than 10 amino acid residues; substitution and / or deletion and / or addition of no more than 9 amino acid residues; substitution and / or deletion and / or addition of no more than 8 amino acid residues; substitution and / or deletion and / or addition of no more than 7 amino acid residues; substitution and / or deletion and / or addition of no more than 6 amino acid residues; substitution and / or deletion and / or addition of no more than 5 amino acid residues; substitution and / or deletion and / or addition of no more than 4 amino acid residues; substitution and / or deletion and / or addition of no more than 3 amino acid residues; substitution and / or deletion and / or addition of no more than 2 amino acid residues; or substitution and / or deletion and / or addition of no more than 1 amino acid residue.

[0050] Any of the proteins mentioned above can be synthesized artificially, or their encoding genes can be synthesized first and then expressed biologically.

[0051] Identity in any of the aforementioned coding genes refers to sequence similarity to natural nucleic acid sequences. This identity includes nucleotide sequences having 75% or higher, or 80% or higher, or 85% or higher, or 90% or higher, or 91% or higher, or 92% or higher, or 93% or higher, or 94% or higher, or 95% or higher, or 96% or higher, or 97% or higher, or 98% or higher, or 99% or higher identity with the nucleotide sequences of proteins composed of the amino acid sequences shown in Sequences 2, 3, and 4 of this invention. Identity can be evaluated visually or using computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.

[0052] The aforementioned complete system includes a fusion protein expression cassette, a miRNA expression cassette for inhibiting the expression of the RDR6 protein-coding gene, an esgRNA expression cassette, a pegRNA expression cassette, and a selection marker protein expression cassette.

[0053] Furthermore, the fusion protein expression cassette includes a promoter for driving the expression of the fusion protein.

[0054] The promoter used to drive the expression of the fusion protein is preferably the AtRPS5A promoter.

[0055] The nucleotide sequence of the AtRPS5A promoter is shown in positions 2299-3962 of Sequence 1.

[0056] The miRNA expression cassette that inhibits the expression of the RDR6 protein-encoding gene includes a promoter for driving the transcription of the miRNA.

[0057] The promoter used to drive the transcription of the miRNA is preferably the AtRPS5A promoter.

[0058] The esgRNA expression cassette includes a promoter for driving the transcription of the esgRNA.

[0059] The promoter used to drive the transcription of the esgRNA is preferably the AtU6 promoter.

[0060] The nucleotide sequence of the AtU6 promoter is shown in positions 10678-11101 of sequence 1.

[0061] The pegRNA expression cassette includes a promoter for driving the transcription of the pegRNA.

[0062] The promoter used to drive the transcription of the pegRNA is preferably the complex promoter CaMV 35S-CmYLCV-U6-26mini.

[0063] The nucleotide sequence of the complex promoter CaMV 35S-CmYLCV-U6-26 mini is shown in positions 11214-12208 of sequence 1.

[0064] The screening marker protein expression cassette includes a promoter for driving the expression of the screening marker protein.

[0065] The promoter used to drive the expression of the screening marker protein is preferably the CaMV 35S promoter.

[0066] The nucleotide sequence of the CaMV 35S promoter is shown as the reverse complementary sequence from position 1014 to 1691 in Sequence 1.

[0067] Furthermore, the fusion protein expression cassette includes a terminator for terminating the expression of the fusion protein.

[0068] The terminator used to terminate the expression of the fusion protein is preferably the tNOS terminator.

[0069] The nucleotide sequence of the tNOS terminator is shown in positions 10417-10671 of Sequence 1.

[0070] The miRNA expression cassette that inhibits the expression of the RDR6 protein-encoding gene includes a terminator for terminating the transcription of the miRNA.

[0071] The terminator used to terminate the transcription of the miRNA is preferably the tNOS terminator.

[0072] The esgRNA expression cassette includes a terminator for terminating the transcription of the esgRNA.

[0073] The terminator used to terminate the transcription of the esgRNA is preferably PolyT.

[0074] The nucleotide sequence of PolyT is shown in positions 11208-11213 of sequence 1.

[0075] The pegRNA expression cassette includes a terminator for terminating pegRNA transcription.

[0076] The terminator used to terminate the transcription of the pegRNA is preferably PolyT'.

[0077] The nucleotide sequence of PolyT' is shown in positions 12530-12537 of sequence 1.

[0078] The screening marker protein expression cassette includes a terminator for terminating the expression of the screening marker protein.

[0079] The terminator used to terminate the expression of the screening marker protein is the CaMVpolyA terminator.

[0080] The nucleotide sequence of the CaMVpolyA terminator is shown as the reverse complementary sequence from position 103 to 277 of sequence 1.

[0081] Furthermore, the complete system is expressed through a recombinant carrier. The various components or expression boxes included in the complete system can be expressed through the same carrier or through multiple carriers.

[0082] Another object of the present invention is to provide new uses for the above-described complete system.

[0083] This invention provides the application of the above-described complete system in any of the following S1)-S4):

[0084] S1) Editing of the genome sequence of an organism or its cells;

[0085] S2) Products for preparing edited genome sequences of organisms or biological cells;

[0086] S3) Improve the efficiency of editing genome sequences of organisms or biological cells;

[0087] S4) Prepare products that improve the efficiency of editing the genome sequence of organisms or biological cells.

[0088] Another object of the present invention is to provide the method described in any of the following T1)-T3):

[0089] T1) A method for editing genome sequences, comprising the following steps: causing an organism or biological cell to express the above-mentioned fusion protein, the above-mentioned sgRNA and the above-mentioned pegRNA, and reducing the content and / or activity of RDR6 protein in the organism or biological cell;

[0090] T2) A method for improving the efficiency of editing the genome sequence of an organism or biological cell, comprising the following steps: expressing the above-mentioned fusion protein, the above-mentioned sgRNA and the above-mentioned pegRNA in the organism or biological cell, and reducing the content and / or activity of RDR6 protein in the organism or biological cell;

[0091] T3) A method for preparing biological mutants includes the following steps: editing the genome sequence of an organism or biological cell according to the method described in T1) or T2) to obtain a biological mutant.

[0092] In any of the methods described above, the method of "expressing the above-mentioned fusion protein, sgRNA and pegRNA in an organism or biological cell, and reducing the content and / or activity of RDR6 protein in the organism or biological cell" involves introducing each element or expression cassette included in the above-mentioned complete system into a plant.

[0093] Furthermore, the various components or expression boxes included in the complete system can be introduced into the plant through the same carrier or through multiple carriers.

[0094] Furthermore, the entire system is introduced into plants via the same recombinant vector.

[0095] In some implementations, the recombinant vector is the PE-At-II-1 expression vector of the guided editing system described in the following embodiments.

[0096] In some implementations, the recombinant vector is the PE-At-II-2 expression vector of the guided editing system described in the following embodiments.

[0097] The amino acid sequence of any of the RDR6 proteins described above is shown in Sequence 9, and the gene sequence encoding the RDR6 protein is shown in Sequence 10.

[0098] Editing any of the aforementioned genome sequences includes base substitution, base insertion, and / or base deletion at the genome sequence editing site.

[0099] The bases include monobasic and / or polybasic bases.

[0100] The edit site may include one, two, or more. The site may be a neighboring site or a non-neighboring site.

[0101] In some implementations, the editing of the genome sequence is a base substitution of the genome sequence. The base substitution can be a single base substitution, the simultaneous substitution of several adjacent bases, or the simultaneous substitution of several non-adjacent bases.

[0102] Any of the organisms mentioned above is X1), X2), X3), or X4):

[0103] X1) Plants or animals;

[0104] X2) Monocotyledonous or dicotyledonous plants;

[0105] X3) Cruciferous plants;

[0106] X4) Arabidopsis thaliana (e.g., Col-0);

[0107] The biological cells described above are Y1, Y2, Y3, or Y4:

[0108] Y1) Plant cells or animal cells;

[0109] Y2) Monocotyledonous plant cells or dicotyledonous plant cells;

[0110] Y3) Cruciferous plant cells;

[0111] Y4) Arabidopsis thaliana cells.

[0112] To overcome the application bottleneck of the PE system in Arabidopsis thaliana, this invention provides a guided editing system, PE-At-II. Compared to the guided editing system PE-At-I, PE-At-II significantly improves the editing efficiency of Arabidopsis targets, and PE-At-II-I is more suitable for creating transgenic mutants whose target editing is heritable up to the T2 generation. This invention provides a more precise and powerful tool for Arabidopsis gene function research. Attached Figure Description

[0113] Figure 1 A schematic diagram of the structure of the expression carrier for different guided editing systems.

[0114] Figure 2 The results are PCR identification results of T1 generation transgenic plants.

[0115] Figure 3 The results are PCR identification of non-transgenic plants in the T2 generation segregating population.

[0116] Figure 4 To analyze the editing efficiency of expression vectors with different guided editing systems in T1 generation transgenic plants. Detailed Implementation

[0117] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0118] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0119] The amino acid sequence of the RDR6 protein in the following examples is shown in Sequence 9, and its encoding gene sequence is shown in Sequence 10.

[0120] Example 1: Design of main components and expression carriers of different guided editing systems

[0121] I. Design of Main Components in Different Guided Editing Systems

[0122] The guided editing system PE-At-I includes a fusion protein, esgRNA, pegRNA, and a selection marker protein; the fusion protein includes a Cas9 nicking enzyme (such as Cas9n(H840A / R221K / N394K)) and a truncated reverse transcriptase M-MLV (ΔRNase H) (the amino acid sequence of the truncated reverse transcriptase M-MLV is shown in Sequence 4); the esgRNA is used to generate a non-editing strand cut; the pegRNA includes, in sequence, esgRNA', an RT sequence, a PBS sequence, and a tevopreQ1 motif, wherein the esgRNA' targets the editing strand target sequence.

[0123] The guided editing system PE-At-II includes a fusion protein, amiR-RDR6, esgRNA, pegRNA, and a selection marker protein. The fusion protein includes a Cas9 nicking enzyme (such as Cas9n(H840A / R221K / N394K)) and a truncated reverse transcriptase M-MLV (ΔRNase H) (the amino acid sequence of this truncated reverse transcriptase M-MLV is shown in Sequence 4). amiR-RDR6 is a small RNA (sRNA) that can inhibit the expression of the RDR6 protein-encoding gene. The esgRNA is used to generate a non-editing strand cut. The pegRNA includes, in sequence, esgRNA', an RT sequence, a PBS sequence, and a tevopreQ1 motif, wherein the esgRNA' targets the editing strand target sequence.

[0124] II. Design of Expression Carriers for Different Guided Editing Systems

[0125] The PE-At-I-1 guided editing system's expression vectors include a Cas9n(H840A / R221K / N394K)&ΔRNase H expression cassette, an esgRNA expression cassette, a pegRNA expression cassette, and a selection marker protein expression cassette. In the Cas9n(H840A / R221K / N394K)&ΔRNase H expression cassette, ΔRNase H is fused to the C-terminus of Cas9n(H840A / R221K / N394K), and this cassette is driven by the AtRPS5A promoter. The esgRNA expression cassette is driven by the AtU6 promoter. The pegRNA expression cassette is driven by the complex promoter CaMV 35S-CmYLCV-U6-26mini. The selection marker protein expression cassette is driven by the CaMV 35S promoter.

[0126] The PE-At-I-2 guided editing system's expression vectors include a Cas9n(H840A / R221K / N394K)&ΔRNase H expression cassette, an esgRNA expression cassette, a pegRNA expression cassette, and a selection marker protein expression cassette. In the Cas9n(H840A / R221K / N394K)&ΔRNase H expression cassette, ΔRNase H is fused to the C-terminus of Cas9n(H840A / R221K / N394K), and this cassette is driven by the CaMV 35S promoter. The esgRNA expression cassette is driven by the AtU6 promoter. The pegRNA expression cassette is driven by the complex promoter CaMV 35S-CmYLCV-U6-26mini. The selection marker protein expression cassette is driven by the CaMV 35S promoter.

[0127] The expression vectors of the PE-At-II-1 guided editing system include the Cas9n(H840A / R221K / N394K)&ΔRNaseH expression cassette, the amiR-RDR6 expression cassette, the esgRNA expression cassette, the pegRNA expression cassette, and the selection marker protein expression cassette. In the Cas9n(H840A / R221K / N394K)&ΔRNase H expression cassette, ΔRNase H is fused to the C-terminus of Cas9n(H840A / R221K / N394K), and this expression cassette is driven by the AtRPS5A promoter. The amiR-RDR6 expression cassette is driven by the AtRPS5A promoter. The esgRNA expression cassette is driven by the AtU6 promoter. The pegRNA expression cassette is driven by the complex promoter CaMV 35S-CmYLCV-U6-26 mini. The selection marker protein expression cassette is driven by the CaMV 35S promoter.

[0128] The PE-At-II-2 guided editing system's expression vectors include the Cas9n(H840A / R221K / N394K)&ΔRNaseH expression cassette, the amiR-RDR6 expression cassette, the esgRNA expression cassette, the pegRNA expression cassette, and the selection marker protein expression cassette. In the Cas9n(H840A / R221K / N394K)&ΔRNase H expression cassette, ΔRNase H is fused to the C-terminus of Cas9n(H840A / R221K / N394K), and this expression cassette is driven by the CaMV 35S promoter. The amiR-RDR6 expression cassette is driven by the CaMV 35S promoter. The esgRNA expression cassette is driven by the AtU6 promoter. The pegRNA expression cassette is driven by the complex promoter CaMV 35S-CmYLCV-U6-26 mini. The selection marker protein expression cassette is driven by the CaMV 35S promoter.

[0129] The structural diagrams of the expression carriers of the above-mentioned guided editing systems are as follows: Figure 1 As shown.

[0130] Example 2: Construction of expression vectors with different guided editing systems and comparison of their efficiency in base editing of the Arabidopsis genome.

[0131] I. Construction of Expression Carriers for Different Guided Editing Systems

[0132] This invention uses AtGL1 as a target and artificially constructs the following recombinant vectors (each vector is a circular plasmid):

[0133] The nucleotide sequence of the PE-At-I-1 expression vector, a guided editing system, consists of Sequence 1 and Sequence 11. In Sequence 1, positions 103-277 are the reverse complementary sequence of the CaMVpolyA terminator; positions 292-975 are the reverse complementary sequence of the gene encoding the selection marker red fluorescent protein DsRed (encoding the DsRed protein shown in Sequence 2); positions 1014-1691 are the reverse complementary sequence of the CaMV 35S promoter; positions 2299-3962 are the nucleotide sequence of the AtRPS5A promoter; positions 4204-8304 are the gene encoding the Cas9n (H840A / R221K / N394K) protein (encoding the Cas9n (H840A / R221K / N394K) protein shown in Sequence 3); and positions 8650-10191 are the gene encoding the truncated reverse transcriptase M-MLV (ΔRNaseH) (encoding the ΔRNaseH protein shown in Sequence 4). The sequence consists of: H protein, positions 10417-10671 being the nucleotide sequence of the tNOS terminator; positions 10678-11101 being the nucleotide sequence of the AtU6 promoter; positions 11102-11121 being the esgRNA target sequence that generates a non-coding strand cleavage; positions 11122-11207 being the esgRNA backbone sequence that generates a non-coding strand cleavage; positions 11208-11213 being the nucleotide sequence of PolyT; and positions 11214-12208 being the complex promoter CaMV 35S-CmYLCV-U6-26. The nucleotide sequence of mini is as follows: positions 12209-12285 are the tRNA sequence, positions 12286-12305 are the esgRNA target sequence corresponding to pegRNA, positions 12306-12391 are the esgRNA backbone sequence corresponding to pegRNA, positions 12392-12405 are the RT sequence on pegRNA, positions 12406-12416 are the PBS sequence on pegRNA, positions 12425-12461 are the tevopreQ1 motif sequence on pegRNA, positions 12462-12529 are the nucleotide sequence of HDV (HDV is used in conjunction with tRNA and can be used to cleave and generate pegRNA), and positions 12530-12537 are the nucleotide sequence of Poly T'.

[0134] The PE-At-I-2 expression vector of the guide editing system is obtained by replacing the nucleotide sequence of the AtRPS5A promoter in the PE-At-I-1 expression vector of the guide editing system with the nucleotide sequence of the CaMV 35S promoter (the nucleotide sequence of the CaMV 35S promoter is shown in Sequence 5), while keeping the other sequences of the PE-At-I-1 expression vector unchanged.

[0135] The PE-At-II-1 expression vector is obtained by inserting the amiR-RDR6 expression cassette between positions 10672 and 10673 of the PE-At-I-1 expression vector while keeping the other sequences of the PE-At-I-1 expression vector unchanged. The nucleotide sequence of the amiR-RDR6 expression cassette is shown in Sequence 6, where positions 1-1664 are the AtRPS5A promoter sequence, positions 1680-2091 are the amiR-RDR6 gene sequence, and positions 2115-2367 are the tNOS terminator sequence. The amiR-RDR6 expression cassette can express a miRNA that represses RDR6 gene expression; the precursor sequence of this miRNA is shown in Sequence 7, and the mature sequence is shown in Sequence 8.

[0136] The PE-At-II-2 expression vector of the guide editing system is obtained by replacing all nucleotide sequences of the AtRPS5A promoter in the PE-At-II-1 expression vector with the nucleotide sequence of the CaMV 35S promoter, while keeping other sequences of the PE-At-II-1 expression vector unchanged.

[0137] II. Genetic transformation and identification of Arabidopsis thaliana

[0138] To test the editing efficiency of different guided editing systems, the guided editing system expression vectors PE-At-I-1, PE-At-I-2, PE-At-II-1, and PE-At-II-2 constructed in step one were transformed into wild-type Arabidopsis thaliana Col-0 according to the method described in the literature "Clough SJ & Bent AF (1998) Floral dip: a simplified method for Agrobacterium-mediated transformation of Arabidopsis thaliana. Plant J 16(6):735-743.". Approximately 100 T0 generation transgenic Arabidopsis seeds with red fluorescence were selected from each vector based on the reporter gene DsRed. The T0 generation transgenic Arabidopsis seeds were germinated on MS medium. After 7 days of growth, the seedlings were transplanted into nutrient soil and allowed to grow for another 14 days to obtain T1 generation transgenic Arabidopsis plants. DNA was extracted from leaves of T1 generation transgenic Arabidopsis plants. PCR was used to detect the DsRed and Cas9 genes (the amplification products of DsRed-F and DsRed-R primers were 631 bp in size, annealed at 55℃; the amplification products of Cas9-F and Cas9-R primers were 600 bp in size, annealed at 55℃). Positive T1 generation transgenic Arabidopsis plants were obtained. Electrophoretic detection results of the amplification products from some positive T1 generation transgenic Arabidopsis plants are shown below. Figure 2 As shown in the figure. Transformation with the PE-At-I-1 vector yielded 59 T1 generation transgenic positive plants, transformation with the PE-At-I-2 vector yielded 54 T1 generation transgenic positive plants, transformation with the PE-At-II-1 vector yielded 60 T1 generation transgenic positive plants, and transformation with the PE-At-II-2 vector yielded 38 T1 generation transgenic positive plants. The primer sequences are as follows:

[0139] DsRed-F: 5'-TCCGAGAACGTCATCACCGAG-3';

[0140] DsRed-R: 5'-ACTGCTCCACGATGGTGTAG-3';

[0141] Cas9-F: 5'-GCAACACCGACCGCCACT-3';

[0142] Cas9-R: 5'-CGTTCTTCTCTCACCAGGGA-3'.

[0143] III. Analysis of Editing Efficiency of Different Guided Editing Systems

[0144] Using DNA from T1 generation transgenic Arabidopsis thaliana positive plants obtained in step two as templates, the target editing site in the T1 generation transgenic Arabidopsis thaliana positive plants was amplified using PCR-specific primers (F: TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGTGATTCGTTGATAGGGCTAA; R: GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGATGTACCTATTGCCGAGGAGCT), yielding PCR products. The PCR products were then subjected to next-generation sequencing (NGS), with 1000 reads sequenced per single plant. The T1 generation PE editing efficiency was detected and calculated. T1 generation PE editing efficiency = (number of reads with target mutation / total number of reads) × 100% (target mutation refers to the change from TGAAATTGCCTTTG to TGAAATT). T C A TT A G, where the underlined bases are the target mutant bases). T1 generation NGS sequencing results showed that the PE-At-II guided editing system significantly improved PE efficiency compared to the PE-At-I system. Specifically, the average increase in efficiency for PE-At-II-1 was 11.4 times compared to PE-At-I-1; and the average increase in efficiency for PE-At-II-2 was 26.9 times compared to PE-At-I-2. Figure 4 (Table 1).

[0145] T1 generation plants with high PE editing efficiency were selected for seed collection. Approximately 50 non-transgenic seeds (without red light) were selected from the T2 generation population and germinated on MS medium. After 7 days of growth, the seedlings were transplanted into nutrient soil. After another 14 days of growth, T2 generation non-transgenic Arabidopsis plants were obtained. DNA was extracted from the leaves of the T2 generation non-transgenic Arabidopsis plants, and PCR detection of the DsRed and Cas9 genes was performed to obtain the T2 generation non-transgenic population. Electrophoretic detection results of amplification products from some non-transgenic plants are shown below. Figure 3 As shown. Using DNA from the T2 generation of non-transgenic population as a template, the target editing site in the T2 generation of non-transgenic population was amplified using PCR-specific primers (F: CGAAAACCCATCATAAGTTC; R: AACTTAACCGGCCAAATCTT) to obtain PCR products. The PCR products were then subjected to Sanger sequencing to detect and calculate the PE efficiency. PE efficiency (%) = {(heterozygous mutants + homozygous mutants) / total number of sequenced strains} × 100% (the mutation is the same as the target mutation mentioned above).

[0146] T2 generation Sanger sequencing results showed that PE-edited materials without transgenic components could be identified in the progeny of the AtRPS5A promoter-based guided editing system PE-At-II-1, while no PE-edited materials without transgenic components could be obtained in the progeny of the guided editing systems PE-At-I and PE-At-II-2 (Table 1). This indicates that the guided editing system PE-At-II of the present invention not only significantly improves the PE efficiency in Arabidopsis thaliana, but also, compared with PE-At-II-2, PE-At-II-1 is more suitable for creating transgenic-free mutants whose targeted editing is heritable to the T2 generation.

[0147] Table 1

[0148]

[0149] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.

Claims

1. A complete system comprising a fusion protein, a substance that reduces the content and / or activity of RDR6 protein, sgRNA, and pegRNA; The fusion protein is formed by the fusion of a nuclease and a reverse transcriptase; The sgRNA is used to generate non-editable strand cuts; The pegRNA includes sgRNA', a reverse transcription template sequence, and a primer binding site sequence; the sgRNA' targets the editing strand target sequence.

2. The complete system according to claim 1, characterized in that: In the fusion protein, the reverse transcriptase is fused to the C-terminus of the nuclease.

3. The complete system according to claim 1 or 2, characterized in that: The nuclease is a Cas9 nuclease; Alternatively, the Cas9 nuclease may be a Cas9 nicking enzyme; Alternatively, the Cas9 nicking enzyme is Cas9n (H840A / R221K / N394K); the Cas9n (H840A / R221K / N394K) is either A1 or A2. A1) The amino acid sequence is that of the protein shown in sequence 3; A2) Proteins with the same function obtained by substituting and / or deleting and / or adding one or more amino acid residues of the amino acid sequence shown in Sequence 3. Alternatively, the reverse transcriptase is a truncated reverse transcriptase; the truncated reverse transcriptase is B1) or B2): B1) The amino acid sequence of the protein is shown in sequence 4; B2) Proteins with the same function obtained by substituting and / or deleting and / or adding one or more amino acid residues as shown in Sequence 4.

4. The complete system according to any one of claims 1-3, characterized in that: The substance that reduces the content and / or activity of RDR6 protein is a nucleic acid molecule that inhibits the expression of the RDR6 protein encoding gene; Alternatively, the nucleic acid molecule that inhibits the expression of the RDR6 protein-coding gene is a miRNA that inhibits the expression of the RDR6 protein-coding gene; Alternatively, the nucleotide sequence of the miRNA that inhibits the expression of the RDR6 protein-coding gene is shown in Sequence 7 or Sequence 8.

5. The complete system according to any one of claims 1-4, characterized in that: The complete system also includes screening marker proteins.

6. The complete system according to any one of claims 1-5, characterized in that: The complete system includes a fusion protein expression cassette, a miRNA expression cassette that inhibits the expression of the RDR6 protein-coding gene, an esgRNA expression cassette, a pegRNA expression cassette, and a selection marker protein expression cassette. Alternatively, the fusion protein expression cassette may include a promoter for driving the expression of the fusion protein; Alternatively, the promoter used to drive the expression of the fusion protein is the AtRPS5A promoter; Alternatively, the miRNA expression cassette that inhibits the expression of the RDR6 protein-coding gene includes a promoter for driving the transcription of the miRNA; Alternatively, the promoter used to drive the transcription of the miRNA is the AtRPS5A promoter; Alternatively, the nucleotide sequence of the AtRPS5A promoter is shown as positions 2299-3962 of sequence 1; Alternatively, the esgRNA expression cassette may include a promoter for driving the transcription of the esgRNA; Alternatively, the promoter used to drive the transcription of the esgRNA is the AtU6 promoter; Alternatively, the nucleotide sequence of the AtU6 promoter is shown as positions 10678-11101 of sequence 1; Alternatively, the pegRNA expression cassette may include a promoter for driving the transcription of the pegRNA; Alternatively, the promoter used to drive the transcription of the pegRNA is the complex promoter CaMV 35S-CmYLCV-U6-26mini; Alternatively, the nucleotide sequence of the complex promoter CaMV 35S-CmYLCV-U6-26 mini is shown in positions 11214-12208 of sequence 1; Alternatively, the screening marker protein expression cassette may include a promoter for driving the expression of the screening marker protein; Alternatively, the promoter used to drive the expression of the screening marker protein is the CaMV 35S promoter; Alternatively, the nucleotide sequence of the CaMV 35S promoter is shown as the reverse complementary sequence at positions 1014-1691 of Sequence 1.

7. The application of the complete system according to any one of claims 1-6 in any of the following S1)-S4): S1) Editing of the genome sequence of an organism or its cells; S2) Products for preparing edited genome sequences of organisms or biological cells; S3) Improve the efficiency of editing genome sequences of organisms or biological cells; S4) Prepare products that improve the efficiency of editing the genome sequence of organisms or biological cells.

8. The method described in any of T1)-T3): T1) A method for editing a genome sequence, comprising the following steps: causing an organism or biological cell to express the fusion protein, the sgRNA, and the pegRNA as described in claim 1, and reducing the content and / or activity of the RDR6 protein in the organism or biological cell; T2) A method for improving the efficiency of editing the genome sequence of an organism or biological cell, comprising the following steps: expressing the fusion protein, sgRNA and pegRNA of claim 1 in the organism or biological cell, and reducing the content and / or activity of RDR6 protein in the organism or biological cell; T3) A method for preparing biological mutants includes the following steps: editing the genome sequence of an organism or biological cell according to the method described in T1) or T2) to obtain a biological mutant.

9. The complete system according to any one of claims 1-6, the application according to claim 7, or the method according to claim 8, characterized in that: The editing of the genome sequence includes base substitution, base insertion, and / or base deletion at the genome sequence editing site; Alternatively, the bases may include monobasic and / or polybasic bases; Alternatively, the edit sites may include one, two, or more.

10. The complete system according to any one of claims 1-6, the application according to claim 7, or the method according to claim 8, characterized in that: The organism is X1), X2), X3), or X4. X1) Plants or animals; X2) Monocotyledonous or dicotyledonous plants; X3) Cruciferous plants; X4) Arabidopsis thaliana; The biological cell is Y1, Y2, Y3, or Y4. Y1) Plant cells or animal cells; Y2) Monocotyledonous plant cells or dicotyledonous plant cells; Y3) Cruciferous plant cells; Y4) Arabidopsis thaliana cells.