Method for inserting exogenous sequence in genome at fixed point

By combining the dual pegRNA and ePPE system with the Cre/Lox or FLP/FRT system, the problem of low insertion efficiency of large exogenous sequences in higher plant cells was solved, achieving efficient site-directed insertion.

CN121826033APending Publication Date: 2026-04-10INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-05-25
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve efficient site-specific insertion of large exogenous sequences into higher plant cells, particularly due to the stability and efficiency issues of reverse transcriptase, resulting in low insertion efficiency.

Method used

By employing a dual pegRNA strategy and an enhanced plant guided editing system (ePPE) combined with the Cre/Lox system or FLP/FRT system from the tyrosine recombinase family, large-fragment exogenous sequences are inserted site-specifically using the SSA repair pathway by providing donor DNA with recombination sites.

Benefits of technology

This technology enables efficient site-specific insertion of short fragments and precise insertion of large exogenous sequences in plant cells, improving insertion efficiency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention belongs to the field of gene engineering. Specifically, the invention relates to a method for inserting an exogenous sequence in a genome at a fixed point. Specifically, on the basis of a guided editing system (PE), two adjacent pegRNAs with partially overlapped sequences on a reverse transcription template are used, and efficient and accurate exogenous sequence fixed-point insertion is realized in a genome, especially a plant genome. The system is further coupled with a recombinase system such as Cre / Lox or FLP / FRT and the like, and large-fragment exogenous sequence fixed-point insertion is realized in genomes, especially plant genomes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 202310599213.X, filed on May 25, 2023, entitled "A method for site-specific insertion of exogenous sequences into the genome".

[0002] This application claims priority to Chinese patent application No. 202210580767.0 filed on May 25, 2022 and Chinese patent application No. 202310363943.X filed on April 6, 2023. Technical Field

[0003] This invention belongs to the field of genetic engineering. Specifically, this invention relates to a method for site-specific insertion of exogenous sequences into the genome. Specifically, this invention is based on a guided editing system (PE), using two adjacent pegRNAs with partially overlapping sequences on a reverse transcription template to achieve efficient and precise site-specific insertion of exogenous sequences into the genome, particularly the plant genome. Furthermore, by coupling this system with recombinase systems such as Cre / Lox or FLP / FRT, site-specific insertion of large exogenous sequences into the genome, particularly the plant genome, is achieved. Background of the Invention The rapid development of DNA sequencing technology has propelled the life sciences into the genomic era. Technologies such as GWAS have greatly advanced genetics, particularly in plants, where the functions of many key genes have been elucidated. This presents a significant opportunity for molecular crop breeding. Traditional crop breeding methods, such as hybridization and backcrossing, are no longer sufficient to support the rapid growth demands of crop breeding due to their time and labor-intensive nature. Therefore, the development of novel molecular breeding technologies is becoming increasingly important.

[0004] Transgenic technology, with its ability to rapidly and efficiently acquire desirable traits, has been rapidly applied to plant molecular breeding. However, due to its introduction of foreign genes, it is subject to strict regulation. In contrast, genome editing technology can precisely modify functional genes at specific sites without introducing foreign genes, thus achieving desirable traits more quickly and efficiently. Currently, plant genome editing tools mainly fall into three categories: zinc finger nucleases (ZFNs); transcription activator-like effector nucleases (TALENs); and clustered regularly spaced short palindromic repeats and their associated proteins (CRISPR / Cas). Among these, the CRISPR / Cas system is the simplest and most efficient, and has made significant contributions to genetic research and plant molecular breeding in recent years.

[0005] The widely used CRISPR / Cas system consists of a single-stranded guide RNA (sgRNA) and a site-specific nuclease, Cas9. The sgRNA targets specific locations on the genomic DNA via base pairing. Cas9 and sgRNA form a ribonucleoprotein complex (RNP) within the cell. Simultaneously, Cas9 undergoes a conformational change. A domain on Cas9 (PAM-interaction domain, PI domain) continuously interacts with motifs (NGG, PAM) at various locations on the genome until it finds a site that can pair complementary with the sgRNA. At this point, the RNP complex interacts with the DNA to form a new complex. Cas9 unwinds the DNA double helix to form an R-loop, and the conformation changes again. The RuvC and HNH nuclease active domains on Cas9 are activated, respectively cleaving the non-target and target strands, resulting in a DNA double-strand break (DSB). This DNA double-strand break triggers the cell's endogenous DNA repair mechanism, typically through the most frequent non-homologous end joining (NHEJ). NHEJ is an error-prone repair pathway, which may introduce random insertions or deletions (indels) near the DSB during repair, leading to abnormal gene expression. If a foreign DNA segment (donor) with homologous arms flanking the DSB is provided during DSB generation, the cellular endogenous repair mechanism may use this donor as a template for homologous recombination repair (HR). HR is a precise repair pathway that can introduce arbitrary point mutations, insertions, and deletions into the genome. However, this repair pathway occurs very infrequently in higher organisms, especially plant cells, and therefore has not been widely adopted. Subsequently, a CRISPR-based base editor (BE) was developed. The BE system utilizes a RuvC-inactivated Cas9 (nCas9-D10A) coupled with a deaminase (cytosine deaminase or adenine deaminase, corresponding to CBE and ABE, respectively). When the RNP complex binds to DNA to form an R-Loop, the deaminase deaminizes cytosine (C) or adenine (A) on the non-target strand to form uracil (U) or hypoxanthine (I). Intracellular repair mechanisms recognize uracil as thymine (T) and hypoxanthine as guanine (G). At this point, nCas9 cleaves the target strand, thereby promoting the base excision repair pathway (BER) to complete the repair of CUT or AIG. The BE system can perform efficient and precise point mutations without relying on DSB generation and the HR pathway, and therefore has been rapidly and widely adopted. Meanwhile, the CRISPR system, represented by BE, has seen rapid development in its genome editing toolkit, which couples other effector factors, including targeted activation, inhibition, and epigenetic modification by coupling transcriptional activation, repressor, or epigenetic modification factors.

[0006] Despite the rapid development of the CRISPR molecular toolkit, from simple gene knockout to precise base editing, and further to transcriptional activation, repression, and epigenetic modification, targeted and precise insertion of DNA fragments has remained a challenge in higher plant cells. Traditional strategies for targeted insertion rely on the generation of DNA fragments from gene sites (DSBs). When a donor DNA segment without genomic homologous sequences is provided, this donor may be inserted near the DSB via the NHEJ repair pathway after DSB generation. However, this process is highly imprecise and inefficient due to issues such as the donor delivery method. When a donor DNA segment containing genomic homologous sequences is provided, the target fragment within this donor may be inserted into the target site via the HR repair pathway after DSB generation. However, this process is extremely inefficient, and almost impossible to achieve in higher plant cells.

[0007] Due to the low efficiency of recombinant recombinases (HRs), site-specific large DNA integration can be accomplished using site-specific recombinases (SSRs). SSRs specifically recognize and bind to a DNA sequence (recombination site, RS) and form a synaptic complex. A strand exchange process occurs between the two synaptic complexes, completing DNA recombination. This process is mediated by the attack of tyrosine or serine residues at the active site of the SSR on the RS phosphate backbone, leading to DNA cleavage. After cleavage, a covalent intermediate is formed, and a strand exchange reaction occurs between the two RSs. This process does not require high-energy cofactors and does not rely on endogenous DNA repair pathways, making it highly efficient. Based on differences in the residues at the active site of SSRs, they can be divided into tyrosine recombinase families and serine recombinase families. These mainly originate from bacteriophages, bacteria, and fungi, and perform biological functions such as cleavage, inversion, integration, and transposition. Common tyrosine recombinases include E. coli bacteriophage λ integrase, P1 bacteriophage Cre recombinase, and yeast FLP recombinase. They all utilize a conserved tyrosine residue to attack one strand of the RS backbone, exposing a 5' phosphate group and a 3' hydroxyl group. The 5' phosphate group and 3' hydroxyl group of the two RSs then bind separately, achieving chain exchange. Simultaneously, the recombinase bound to the RS undergoes conformational change, attacking the other strand and achieving chain exchange through the same pathway, thus completing the recombination process. Common serine recombinases include Tn3 transposase, Salmonella recombinase Hin, Streptomyces bacteriophage ФC31 integrase, and Mycobacterium bacteriophage Bxb1 integrase. Their recombination process is similar to that of tyrosine recombinases, but they utilize a serine residue to simultaneously attack both strands of the RS backbone, achieving simultaneous exchange of both strands of the two RSs, thus completing the recombination process. SSRs have a wide range of applications: in vitro, they are mainly used as a molecular cloning tool, and their high efficiency in DNA recombination makes the in vitro cloning of large and multi-fragment molecules very simple; in prokaryotic cells, they can be used as a gene or staining engineering tool to perform large DNA deletion, inversion, translocation, or integration; in eukaryotic cells of higher organisms, they are mainly used as a tool for deleting transgenic marker genes, but the current difficulty in site-specific knock-in of SSRs makes site-specific integration of large DNA fragments very difficult.

[0008] Recently, guided editing systems (PEs) capable of arbitrary base mutations, short DNA insertions, and deletions have been developed and rapidly adopted for widespread use in plant and animal genome editing due to their powerful capabilities and independence from DSBs. PEs utilize an HNH domain-inactivated Cas9 (nCas9-H840A) coupled with a reverse transcriptase (MLV). Simultaneously, a reverse transcription template sequence (RT) and a reverse transcriptase primer binding site (PBS) are sequentially introduced at the 3' end of the sgRNA. The RT carries the target mutant sequence and sequences flanking the mutant sequence that are homologous to the genome; this sgRNA is called pegRNA. After nCas9 cleaves the non-target strand, PBS binds to its 5' end, serving as the initiation primer for reverse transcriptase. The reverse transcriptase then extends to the 3' end of the RT, reverse transcribing the RT sequence into DNA, forming a 3' overhang with the mutant sequence. After repair by endogenous DNA within the cell, this mutant sequence can potentially be introduced into the genome, thereby completing any type of genome editing within a certain length.

[0009] Guided editing systems remain inefficient in higher plant cells, insufficient for efficient insertion, and the length of inserted fragments is highly limited. Three main reasons are speculated: first, the repair pathways utilized by guided editing systems in higher plants occur less frequently, leading to lower editing efficiency; second, RT competes with genomic homologous sequences for binding to genomic DNA, hindering reverse transcription; and third, reverse transcriptases or pegRNAs are easily degraded or have insufficient reverse transcription capabilities. There is still a need in this field for systems and methods to achieve efficient insertion of exogenous nucleotide sequences, especially large exogenous nucleotide fragments, into plant genomes. Invention Summary To avoid the first two reasons for the low efficiency of PE in higher plants, the inventors first designed a dual pegRNA strategy. The two pegRNAs target and bind to the two strands of genomic DNA respectively, with a certain distance between them (approximately 20bp-60bp). The RT of both pegRNAs contains only the required insertion sequence, and the 3' ends have partially overlapping sequences. After reverse transcription is completed, the two newly synthesized DNA strands bind to each other due to the overlapping sequences and anneal. Insertion is completed through a DNA repair pathway different from the original PE system (according to some results of this application, this repair pathway may be SSA, a repair pathway with a relatively high frequency of occurrence in plants).

[0010] Recently, an enhanced plant guided editing system (ePPE) constructed by fusing a retroviral nucleocapsid protein (NC) and deleting the RNaseH active domain of the reverse transcriptase MLV has significantly improved the efficiency of the plant guided editing system by enhancing reverse transcription capacity or reverse transcriptase stability. Furthermore, adding a secondary structure tevopre (epegRNA) to the 3' end of pegRNA can also enhance reverse transcription capacity or pegRNA stability and thus improve the efficiency of PE.

[0011] To further improve insertion efficiency, the inventors simultaneously used the aforementioned ePPE system and epigRNA, thereby achieving highly efficient short-fragment site-directed insertion in plant somatic cells. The DNA integration capabilities of the Cre / Lox and FLP / FRT systems from the tyrosine recombinase family, and the ФC31 and Bxb1 recombinase systems from the serine family, in rice somatic cells were evaluated. The Cre / Lox and FLP / FRT systems were found to be more effective. Therefore, they were combined with the aforementioned highly efficient insertion system, and by providing an additional donor containing the desired insertion gene (RS), site-directed insertion of large exogenous nucleotide sequences was achieved in a one-step process. Brief description of the attached diagram Figure 1. Testing the efficiency of five constructs in rice protoplasts for inserting double pegRNA into Lox66 or FRT1.

[0012] Figure 2. Testing the efficiency of inserting PPE+pegRNA, ePPE+pegRNA, PPE+epegRNA, and ePPE+epegRNA into RS.

[0013] Figure 3 The relationship between insertion length (30bp-100bp) and the distance between the two pegRNAs (PAM distance 20bp-80bp) and the length of overlap between the two RTs (10bp-50bp) was evaluated using the ePPE+epegRNA combination to assess the effect on insertion efficiency.

[0014] Figure 4. Efficiency of fixed-point insertion on NG PAM when Cas9 is replaced with SpG-Cas9 or SpRY-Cas9.

[0015] Figure 5. Effect of different promoters on pegRNA insertion efficiency.

[0016] Figure 6. The effect of using a 37-degree temperature treatment method (6B) and a system using MS2-MCP to recruit MLVs (6C) on long fragment insertion efficiency.

[0017] Figure 7. A) Schematic diagram of the GFP reporter system; B) Evaluation of the editing effect of eight recombinases, and the corresponding recombinase site sequences. The microscopic images show the recombinases corresponding to rice protoplast transformation or non-transformation; C) Verification of recombinase editing effect using a fluorescent reporter system; D) Schematic diagram of the construct for evaluating the DNA integration ability of recombinases using a fluorescent reporter system; F) Schematic diagram of the construct for one-step large fragment insertion of recombinase combined with ePPE.

[0018] Figure 8 The insertion efficiency of the PrimeROOT.v1 system was detected by ddPCR.

[0019] Figure 9. A) Percentage of GFP-positive plant protoplast cells detected by flow cytometry, reflecting the efficiency of "one-step" large fragment insertion using different combinations of recombinases; B) GFP insertion efficiency of rice protoplast OsALS determined by ddPCR.

[0020] Figure 10. The editing efficiency of different base editing systems was demonstrated using fluorescence microscopy and flow cytometry.

[0021] Figure 11. Detection of the insertion percentage of different donors into four endogenous sites using ddPCR.

[0022] Figure 12. Percentage of insertions at six endogenous sites in maize using ddPCR with the PrimeROOT.v2C-Cre system.

[0023] Figure 13. Percentage of large fragment insertions in different gene editing systems detected by ddPCR.

[0024] Figure 14. Comparison of the precision editing efficiency of PrimeROOT.v2C-Cre and NHEJ using base sequencing results.

[0025] Figure 15. A) Schematic diagram of inserting the Act1 promoter into the OsHPPD site using PrimeROOT.v2C-Cre; B) Screening of pegRNA pairs; C) Insertion efficiency.

[0026] Figure 16. GSH sites obtained by high-throughput sequencing and insertion efficiency of recombination sites in GSH1 detected by high-throughput sequencing.

[0027] Figure 17. Schematic diagram of PrimeROOT.v3 and the efficiency of precise insertion through PrimeROOT.v3.

[0028] Figure 18. Efficiency and sequencing results of precise insertion in human HEK293 cells using the PrimeROOT system. Invention Details I. Definition In this invention, unless otherwise stated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are all widely used terms and routine procedures in their respective fields. For example, the standard recombinant DNA and molecular cloning techniques used in this invention are well known to those skilled in the art and are described more fully in the following literature: Sambrook, J., Fritsch, EF, and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). Meanwhile, to better understand this invention, definitions and explanations of relevant terms are provided below.

[0029] As used herein, the term “and / or” covers all combinations of items connected by the term and should be regarded as if each combination had been listed separately herein. For example, “A and / or B” covers “A,” “A and B,” and “B.” For example, “A, B, and / or C” covers “A,” “B,” “C,” “A and B,” “A and C,” “B and C,” and “A and B and C.”

[0030] When the term "comprising" is used herein to describe a protein or nucleic acid sequence, the protein or nucleic acid may consist of the stated sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, while still possessing the activities described in this invention. Furthermore, those skilled in the art will understand that the methionine encoded by the start codon at the N-terminus of a polypeptide may be retained in certain practical situations (e.g., when expressed in a specific expression system) without substantially affecting the polypeptide's function. Therefore, when describing a specific polypeptide amino acid sequence in this specification and claims, although it may not contain the methionine encoded by the start codon at the N-terminus, the sequence containing that methionine is still included, and correspondingly, its encoding nucleotide sequence may also contain the start codon; and vice versa.

[0031] As used herein, a “genome editing system” refers to a combination of components required for genome editing within a cell. The individual components of such a system, such as a guide editing fusion protein or its expression construct, pegRNA or its expression construct, donor construct, etc., may exist independently or in any combination as a composition.

[0032] The term "genome," as used in this article, encompasses not only chromosomal DNA located in the cell nucleus but also organelle DNA located in subcellular components of the cell, such as mitochondria and plastids.

[0033] The term "genetically modified plant" as used in this article refers to a plant whose genome contains inserted exogenous polynucleotides. For example, exogenous polynucleotides can be stably integrated into the plant genome and inherited across generations.

[0034] In relation to a sequence, “exogenous” means a sequence that originates from a foreign species, or, if from the same species, a sequence whose composition and / or loci have been significantly altered from its natural form through deliberate human intervention.

[0035] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” are used interchangeably and refer to single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “D” for A, T, or G, “I” for inosine, and “N” for any nucleotide. Although nucleotide sequences may be represented as DNA sequences (containing T) herein, when referring to RNA, those skilled in the art can readily determine the corresponding RNA sequence (i.e., replacing T with U).

[0036] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably in this invention to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.

[0037] As used in this invention, "expression construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to generate mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein.

[0038] The "expression construct" of the present invention may be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, may be a translatable RNA (such as mRNA), for example, RNA transcribed in vitro.

[0039] The "expression construct" of the present invention may contain regulatory sequences and nucleotide sequences of interest from different sources, or regulatory sequences and nucleotide sequences of interest from the same source but arranged in a manner different from those normally found in nature.

[0040] A "promoter" refers to a nucleic acid fragment that controls the transcription of another nucleic acid fragment. In some embodiments of the present invention, a promoter is a promoter capable of controlling gene transcription in a cell, regardless of whether it originates from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmental regulatory promoter, or an inducible promoter.

[0041] Examples of promoters include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. When used for plants, promoters may be the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, the maize U3 promoter, or the rice actin promoter.

[0042] "Introducing" nucleic acid molecules (e.g., plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the organism's cells with the nucleic acid or protein, enabling the nucleic acid or protein to perform its function within the cell. The term "transformation" as used in this invention includes stable transformation and transient transformation. "Stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome, resulting in the stable inheritance of the foreign gene. Once stable transformation occurs, the foreign nucleic acid sequence is stably integrated into the genome of the organism and its genome in any subsequent generations. "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell to perform its function without the stable inheritance of the foreign gene. In transient transformation, the foreign nucleic acid sequence does not integrate into the genome.

[0043] "Temperament" refers to the physiological, morphological, biochemical, or physical characteristics of a cell or organism.

[0044] "Agronomic traits" specifically refer to measurable parameters of crop plants, including but not limited to: leaf greenness, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, total nitrogen content of plants, nitrogen content of fruits, nitrogen content of seeds, nitrogen content of plant vegetative tissues, total free amino acid content of plants, free amino acid content of fruits, free amino acid content of seeds, free amino acid content of plant vegetative tissues, total protein content of plants, protein content of fruits, protein content of seeds, protein content of plant vegetative tissues, herbicide resistance and drought resistance, nitrogen uptake, root lodging, harvest index, stem lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt tolerance, and tiller number, etc.

[0045] II. Genome editing systems for site-specific modifications in the genome of organisms, such as site-specific insertion of exogenous nucleotide sequences. In one aspect, the present invention relates to a genome editing system for site-specific modification, such as site-specific insertion of exogenous nucleotide sequences, in the genome of an organism, comprising: i) a) An expression construct containing a CRISPR nuclease and / or a nucleotide sequence encoding the CRISPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase, or b) A guide-editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide-editing fusion protein, wherein the guide-editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase; ii) An expression construct containing a first pegRNA and / or a nucleotide sequence encoding the first pegRNA, and iii) An expression construct containing a second pegRNA and / or a nucleotide sequence encoding the second pegRNA. The first pegRNA contains, from 5' to 3', a first guide sequence, a first scaffold sequence, a first reverse transcription template (RT) sequence, and a first primer binding site (PBS) sequence. The second pegRNA contains, from 5' to 3', a second guide sequence, a first scaffold sequence, a second reverse transcription template (RT) sequence, and a second primer binding site (PBS) sequence. The first pegRNA targets a first target sequence on the sense strand of the organism's genomic DNA, and the second pegRNA targets a second target sequence on the antisense strand of the organism's genomic DNA. In some embodiments, the organism is a plant.

[0046] As used herein, a "target sequence" refers to a sequence of approximately 20 nucleotides in length in the genome characterized by a PAM (pre-intermediate sequence adjacent motif) sequence flanking the 5' or 3' region. Typically, the PAM is necessary for the recognition of the target sequence by the complex formed by the CRISPR nuclease or its variants with the guide RNA. For example, for Cas9 nuclease and its variants, the target sequence is adjacent to the PAM at the 3' end, such as 5'-NGG-3'. Based on the presence of the PAM, those skilled in the art can readily identify target sequences in the genome that can be used for targeting. Moreover, depending on the location of the PAM, the target sequence can be located on any strand of the genomic DNA molecule; the strand containing the target sequence is called the target strand. For Cas9 or its derivatives, such as Cas9 nickase, the target sequence is preferably 20 nucleotides long. The PAM sequence may vary depending on the different CRISPR nucleases or their different variants.

[0047] In some embodiments, the pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a target sequence in the genome, resulting in a nick on the target strand (e.g., within the target sequence).

[0048] In some embodiments, the PAM between the first target sequence and the second target sequence is spaced approximately 1 to approximately 300 bp, for example, 10 bp to approximately 100 bp, or for example, approximately 20 bp to approximately 60 bp. In some embodiments, the PAM between the first target sequence and the second target sequence may be spaced approximately 10 bp, approximately 20 bp, approximately 30 bp, approximately 40 bp, approximately 50 bp, approximately 60 bp, approximately 70 bp, approximately 80 bp, approximately 100 bp, approximately 150 bp, or approximately 300 bp.

[0049] In some embodiments, the CRISPR nuclease is a Cas9 nuclease, for example, derived from Streptococcus pyogenes (Streptococcus pyogenes). S. pyogenes SpCas9. An exemplary wild-type SpCas9 contains the amino acid sequence shown in SEQ ID NO:1.

[0050] In some embodiments, the CRISPR nuclease is a CRISPR nickase. The CRISPR nickase in the fusion protein is capable of forming a nick within the target sequence on the target strand of the genomic DNA. In some embodiments, the CRISPR nickase is a Cas9 nickase.

[0051] In some embodiments, the Cas9 nickase is derived from Streptococcus pyogenes (Streptococcus pyogenes). S. pyogenesThe Cas9 cleavage enzyme comprises, relative to wild-type SpCas9, at least the amino acid substitution H840A. In some embodiments, the Cas9 cleavage enzyme comprises the amino acid sequence shown in SEQ ID NO:2. In some embodiments, the Cas9 cleavage enzyme in the fusion protein is capable of forming a cleavage between the -3 nucleotide (the first nucleotide at the 5' end of the PAM sequence is +1) and the -4 nucleotide of the target sequence.

[0052] In some embodiments, the Cas9 nuclease, such as a nickase, is a Cas9 nuclease or nickase variant capable of recognizing an altered PAM sequence. Many Cas9 nickase variants capable of recognizing altered PAM sequences are known in the art. In some embodiments, the Cas9 nuclease, such as a nickase, is a Cas9 variant that recognizes the PAM sequence 5'-NG-3'. In some embodiments, the Cas9 nickase variant recognizing the PAM sequence 5'-NG-3' comprises, relative to wild-type Cas9, the following amino acid substitutions: H840A, D1135L, S1136W, G1218K, E1219Q, R1335Q, T1337R, wherein the amino acid numbers refer to SEQ ID NO:1. In some embodiments, the Cas9 nickase variant (SpG-Cas9 nickase) comprises the amino acid sequence shown in SEQ ID NO:42. In some embodiments, the Cas9 nickase variant recognizing the PAM sequence 5'-NG-3' comprises, relative to wild-type Cas9, the following amino acid substitutions: H840A, A61R, L1111R, D1135L, S1136W, G1218K, E1219Q, N1317R, A1322R, R1333P, R1335Q, and T1337R, wherein the amino acid numbers are referenced to SEQ ID NO:1. In some embodiments, the Cas9 nickase variant (SpRY-Cas9 nickase) comprises the amino acid sequence shown in SEQ ID NO:43.

[0053] The nicks formed by the Cas9 nuclease described in this invention, such as the nicking enzyme, can lead to the formation of a free single strand with a 3' end (3' free single strand) and a free single strand with a 5' end (5' free single strand).

[0054] In some implementations, the CRISPR nucleases, such as Cas9 nickase and the reverse transcriptase, in the fusion protein are linked by a linker.

[0055] In some embodiments, the reverse transcriptase of the present invention may be derived from different sources. In some embodiments, the reverse transcriptase is a viral reverse transcriptase. For example, in some embodiments, the reverse transcriptase is M-MLV reverse transcriptase or a functional variant thereof. An exemplary wild-type M-MLV reverse transcriptase sequence is shown in SEQ ID NO:3.

[0056] In some embodiments, the reverse transcriptase is, for example, M-MLV reverse transcriptase or a functional variant thereof. (a) A mutation at positions 155, 156, 200 and / or 524, for example, a mutation selected from any one or a combination of F155Y, F155V, F156Y, D524N, N200C, the amino acid positions of which refer to SEQ ID NO:3; (b) The connection sequence is missing; and / or (c) The RNase H domain is mutated or deleted.

[0057] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, comprises a mutation selected from D524N, the amino acid position of which refers to SEQ ID NO:3.

[0058] In some preferred embodiments, the RNase H domain of the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is missing.

[0059] In some embodiments, the connection sequence comprises an amino acid sequence as shown in SEQ ID NO:4.

[0060] In some embodiments, the RNase H domain comprises an amino acid sequence as shown in SEQ ID NO:5.

[0061] In some embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, comprises the sequence of any one of SEQ ID NO:9-15, preferably comprising the amino acid sequence shown in SEQ ID NO:14.

[0062] In some embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused at its N-terminus or C-terminus directly or via a linker to a nucleocapsid protein (NC), a hydrolase (PR), or an integrase (IN). The nucleocapsid protein (NC), hydrolase (PR), or integrase (IN) is, for example, derived from M-MLV.

[0063] In some embodiments, the nucleocapsid protein (NC) comprises an amino acid sequence as shown in SEQ ID NO:6.

[0064] In some embodiments, the hydrolase (PR) comprises an amino acid sequence as shown in SEQ ID NO:7.

[0065] In some embodiments, the integrase (IN) comprises an amino acid sequence as shown in SEQ ID NO:8.

[0066] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused at the N-terminus to a nucleocapsid protein (NC) directly or via a linker.

[0067] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused at its C-terminus to a nucleocapsid protein (NC) directly or via a linker.

[0068] In some embodiments, the reverse transcriptase may also be fused via a linker or directly to an RNA aptamer-binding protein sequence (e.g., the MCP protein sequence). Thereby, the reverse transcriptase can be recruited to a CRISPR nuclease through the interaction of the RNA aptamer-binding protein sequence (e.g., the MCP protein sequence) and one or more RNA aptamer sequences (e.g., the MS2 sequence) present on the pegRNA. In this case, it is not necessary to fuse the CRISPR nuclease to the reverse transcriptase. An exemplary MCP protein comprises the amino acid sequence of SEQ ID NO:44.

[0069] As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids without secondary or higher structures. For example, the linker can be a flexible linker, such as GGGGS, GS, GAP, (GGGGS) x 3, GGS, and (GGS) x 7. For example, it can be the linker shown in SEQ ID NO: 16.

[0070] In some embodiments, the CRISPR nuclease, such as a CRISPR nickase, in the fusion protein is located at the N-terminus of the reverse transcriptase. In some embodiments, the CRISPR nuclease, such as a CRISPR nickase, in the fusion protein is located at the C-terminus of the reverse transcriptase.

[0071] In some embodiments of the present invention, the CRISPR nuclease, reverse transcriptase, recombinase, or fusion protein of the present invention may further comprise one or more nuclear localization sequences (NLS). Generally, one or more NLS in the CRISPR nuclease, reverse transcriptase, or fusion protein should have sufficient strength to drive the accumulation of the CRISPR nuclease, reverse transcriptase, or fusion protein in the nucleus of the cell to achieve its base-editing function. Generally, the strength of nuclear localization activity is determined by the number and location of the NLS in the CRISPR nuclease, reverse transcriptase, or fusion protein, the use of one or more specific NLS, or a combination of these factors.

[0072] In some preferred embodiments, the fusion protein comprises, from N-terminus to C-terminus, a CRISPR nuclease such as a nickase, the nucleocapsid protein (NC), and the reverse transcriptase, linked by or without a linker. In some preferred embodiments, the fusion protein comprises, from N-terminus to C-terminus, a nuclear localization sequence-the CRISPR nuclease such as a nickase-linker-the nucleocapsid protein (NC)-nuclear localization sequence-linker-the reverse transcriptase-nuclear localization sequence.

[0073] In some preferred embodiments, the fusion protein comprises the amino acid sequence (ePPE) shown in SEQ ID NO:19.

[0074] In some embodiments, the fusion protein comprises a nuclease moiety and a reverse transcriptase moiety. The nuclease moiety comprises a CRISPR nuclease, such as a CRISPR nickase, and one or more NLSs. The reverse transcriptase moiety comprises an RNA aptamer-binding protein sequence (e.g., an MCP protein sequence), the reverse transcriptase, one or more NLSs, and optionally the nucleocapsid protein (NC). The nuclease moiety and the reverse transcriptase moiety are linked by a self-cleaving peptide. When the fusion protein is translated in vivo, separate nuclease moiety polypeptides and reverse transcriptase moiety polypeptides are formed. The reverse transcriptase moiety is recruited to the nuclease moiety through the interaction of the RNA aptamer-binding protein sequence (e.g., the MCP protein sequence) and one or more RNA aptamer sequences (e.g., the MS2 sequence) present on the pegRNA. An exemplary MCP protein comprises the amino acid sequence of SEQ ID NO:44.

[0075] In some embodiments, the pegRNA of the present invention further comprises one or more RNA aptamer sequences (e.g., MS2 sequences). Exemplary one or more MS2 sequences are shown in SEQ ID NO:45. In some embodiments, the one or more RNA aptamer sequences (e.g., MS2 sequences) are located at the 3' end of the pegRNA. In some embodiments, the one or more RNA aptamer sequences (e.g., MS2 sequences) are located in the middle of the pegRNA, for example, between the scaffold sequence and the RT sequence. The one or more RNA aptamer sequences (e.g., MS2 sequences) can be used to recruit reverse transcriptase containing an RNA aptamer-binding protein sequence (e.g., an MCP protein sequence) to the CRISPR nuclease-pegRNA complex.

[0076] The guide sequence (also called seed sequence or spacer sequence) in the pegRNA of the present invention is configured to have sufficient sequence identity (preferably 100% identity) with the target sequence, thereby enabling it to bind to the complementary strand of the target sequence through base pairing and achieve sequence-specific targeting.

[0077] For example, the guide sequence in the first pegRNA may have sufficient sequence identity (preferably 100% identity) with the first target sequence, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the first target sequence; the guide sequence in the second pegRNA may have sufficient sequence identity (preferably 100% identity) with the second target sequence on the opposite strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the second target sequence, thereby the two pegRNAs result in nicks on different strands of the genomic DNA.

[0078] Various scaffold sequences of gRNAs suitable for CRISPR-based genome editing (e.g., Cas9) are known in the art and can be used in the pegRNAs of this invention. In some specific embodiments, the scaffold sequence of the gRNA is shown in SEQ ID NO:17.

[0079] In some embodiments, the primer-binding sequence is configured to be complementary to at least a portion of the target sequence (preferably perfectly paired with at least a portion of the target sequence). Preferably, the primer-binding sequence is complementary to at least a portion of the 3' free single strand in the DNA strand containing the target sequence due to a nick (preferably perfectly paired with at least a portion of the 3' free single strand), particularly complementary to the nucleotide sequence at the 3' end of the 3' free single strand (preferably perfectly paired). When the 3' free single strand of the strand binds to the primer-binding sequence through base pairing, the 3' free single strand can act as a primer, using the reverse transcription template (RT) sequence immediately adjacent to the primer-binding sequence as a template, to perform reverse transcription under the action of reverse transcriptase in the fusion protein, extending the DNA sequence corresponding to the reverse transcription template (RT) sequence.

[0080] The primer-binding sequence depends on the length of the free single strand formed by the CRISPR nicking enzyme in the target sequence; however, it should have a minimum length to ensure specific binding. In some embodiments, the primer-binding sequence can be 4-20 nucleotides long, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0081] In some embodiments, the primer-binding sequence is configured to have a Tm (denaturation temperature) not exceeding approximately 52°C. In some embodiments, the Tm (denaturation temperature) of the primer-binding sequence is approximately 18°C-52°C, preferably approximately 24°C-36°C, more preferably approximately 28°C-32°C, and even more preferably approximately 30°C.

[0082] Methods for calculating the Tm of a nucleic acid sequence are well known in the art; for example, it can be calculated using the Oligo Analysis Tool online analysis tool. An exemplary formula is Tm = N. G:C *4+N A:T *2, where N G:C It is the number of G and C bases in the sequence, N A:T It refers to the number of A and T bases in the sequence. A suitable Tm can be obtained by selecting an appropriate PBS length. Alternatively, a PBS sequence with a suitable Tm can be obtained by selecting an appropriate target sequence.

[0083] In some embodiments, the RT template sequence can be any sequence. Through reverse transcription, its sequence information can be integrated into the DNA strand containing the target sequence (i.e., the strand containing the target sequence PAM), and then, through cellular DNA repair, a DNA double strand containing the RT template sequence information is formed. In some embodiments, the RT template sequence contains desired modifications. For example, the desired modifications include substitution, deletion, and / or addition of one or more nucleotides. In some embodiments, the RT template sequence is configured to correspond to a sequence downstream of the target sequence nick (e.g., complementary to at least a portion of the sequence downstream of the target sequence nick), but contains desired modifications. The desired modifications include substitution, deletion, and / or addition of one or more nucleotides.

[0084] In some implementations, the two pegRNAs are configured to introduce the same desired modification. For example, one pegRNA is configured to introduce A-G substitutions at the sense strand, while the other pegRNA is configured to introduce T-C substitutions at the corresponding position on the antisense strand. As another example, one pegRNA is configured to introduce a two-nucleotide deletion at the sense strand, and the other pegRNA is configured to similarly introduce a two-nucleotide deletion at the corresponding position on the antisense strand. Other types of modifications can be deduced similarly. The same desired modification can be achieved by designing suitable RT template sequences that target two different strands of the pegRNA.

[0085] In some embodiments, the RT sequence is configured to generate a foreign nucleotide sequence or a portion thereof of the genome to be inserted into after reverse transcription using it as a template, or to generate a complementary sequence to a foreign nucleotide sequence or a portion thereof of the genome of the organism to be inserted, such as a plant. In some embodiments, the RT sequence does not contain a genomic sequence near the target sequence or a complementary sequence to a genomic sequence near the target sequence. In some embodiments, the RT sequence does not contain sequence information other than the foreign nucleotide sequence to be inserted.

[0086] In some embodiments, the first RT sequence and the second RT sequence are used to insert a first exogenous nucleotide sequence, for example, by inserting the first exogenous nucleotide sequence between the first target sequence and the second target sequence (such as between a cleavage of the first target sequence and a cleavage of the second target sequence).

[0087] In some embodiments, the first RT sequence of the first pegRNA is configured to generate a first fragment of a first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; the second RT sequence of the second pegRNA is configured to be a complementary sequence to generate a second fragment of a first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template.

[0088] In some embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted at least partially overlap. In some embodiments, the first and second fragments overlap by at least about 10 bp to about 50 bp, for example, at least about 10 bp, about 15 bp, about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, or about 50 bp. In some embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted completely overlap.

[0089] In some embodiments, the length of the first exogenous nucleotide sequence to be inserted is approximately 1 bp to approximately 700 bp, for example approximately 10 bp, approximately 20 bp, approximately 30 bp, approximately 40 bp, approximately 50 bp, approximately 60 bp, approximately 70 bp, approximately 80 bp, approximately 90 bp, approximately 100 bp, approximately 150 bp, approximately 200 bp, approximately 250 bp, approximately 300 bp, approximately 350 bp, approximately 400 bp, approximately 450 bp, approximately 500 bp, approximately 600 bp, approximately 700 bp, or any value in between.

[0090] In some embodiments, the pegRNA further includes a tevopre sequence at the 3' end of the PBS. The design of the tevopre sequence can be found in James W. Nelson et al., Engineered pegRNAs improve prime editing efficiency. 2022, Nature Biotech. Volume 40, pages 402–410. An exemplary tevopre sequence is shown in SEQ ID NO:20.

[0091] In some embodiments, the pegRNA also includes a polyA sequence at its 3' end. The polyA sequence, for example, contains approximately 10-30 consecutive adenosine nucleotides (A).

[0092] In some implementations, the pegRNA includes a guide sequence, a scaffold sequence, a reverse transcription template (RT) sequence, a primer binding site (PBS) sequence, a tevopre sequence, and a polyA sequence from 5' to 3'.

[0093] In some embodiments, the pegRNA can be precisely processed using a self-processing system. In some specific embodiments, the 5' end of the pegRNA is linked to a first ribozyme or tRNA, which is designed to cleave the fusion at the 5' end of the pegRNA; and / or the 3' end of the pegRNA is linked to a second ribozyme or tRNA, which is designed to cleave the fusion at the 3' end of the pegRNA. The design of the first or second ribozyme or tRNA is within the capabilities of those skilled in the art. For example, see Gao et al., JIPB, Apr, 2014; Vol 56, Issue 4, 343-349. Methods for precisely processing gRNA can be found, for example, in WO2018 / 149418.

[0094] In some implementations, the first pegRNA and the second pegRNA are transcribed by different promoters. For example, the first pegRNA is expressed by the OsU3 promoter, and the second pegRNA is expressed by the TaU3 promoter.

[0095] In some embodiments, the pegRNA is transcribed by a type II promoter, meaning that in an expression construct containing the nucleotide sequence encoding the pegRNA, the coding nucleotide sequence of the pegRNA is operatively linked to a type II promoter. In some specific embodiments, the type II promoter is a GS promoter. An exemplary GS promoter sequence is shown in SEQ ID NO:21.

[0096] In some embodiments, the first target sequence, the second target sequence, and / or the desired modification, such as the first exogenous nucleotide sequence, are associated with an organism, such as a plant trait, such as an agronomic trait, whereby the insertion of the desired modification, such as the first exogenous nucleotide sequence, results in the organism, such as a plant, having altered (preferably improved) traits, such as agronomic traits, relative to a wild-type organism, such as a plant.

[0097] In some implementations, the first exogenous nucleotide sequence contains one or more recombinase recognition sites (RS).

[0098] In some embodiments, the recombinase is a recombinase from the tyrosine recombinase family or the serine recombinase family, preferably a recombinase from the tyrosine recombinase family. Exemplary tyrosine recombinases include, but are not limited to, E. coli bacteriophage λ integrase, P1 bacteriophage Cre recombinase (cyclization recombinase), and yeast FLP recombinase (flippase recombinase). Exemplary serine recombinases include, but are not limited to, Tn3 transposase, Salmonella recombinase Hin, Streptomyces bacteriophage ФC31 integrase, and Mycobacterium bacteriophage Bxb1 integrase. Different recombinases and their corresponding recombinase recognition sites (RS) are known in the art, and those skilled in the art can select them as needed.

[0099] In some embodiments, the recombinase is a Dre recombinase. An exemplary Dre recombinase comprises the amino acid sequence of SEQ ID NO: 56. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, rox (SEQ ID NO: 57, 58).

[0100] In some embodiments, the recombinase is a ФC31 integrase. An exemplary ФC31 integrase comprises the amino acid sequence of SEQ ID NO:22. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, aTTP (SEQ ID NO:38) and / or aTTB (SEQ ID NO:39).

[0101] In some preferred embodiments, the recombinase is a Bxb1 integrase. An exemplary Bxb1 integrase comprises the amino acid sequence of SEQ ID NO:23. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, aGTP (SEQ ID NO:40) and / or aGTB (SEQ ID NO:41).

[0102] In some preferred embodiments, the recombinase is a Cre recombinase. An exemplary Cre recombinase comprises the amino acid sequence of SEQ ID NO:24. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, loxP (SEQ ID NO:26), Lox2272 (SEQ ID NO:29), Lox71 (SEQ ID NO:27), Lox66 (SEQ ID NO:28), or variants thereof, and any combination thereof.

[0103] In some preferred embodiments, the recombinase is an FLP recombinase. An exemplary FLP recombinase comprises the amino acid sequence of SEQ ID NO:25. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, FRT1 (SEQ ID NO:30), FRT6 (SEQ ID NO:31), or variants thereof, and any combination thereof. In some embodiments, the one or more recombinase recognition sites (RS) are variants of FRT1, for example, comprising the sequence described in one of SEQ ID NO:32-37.

[0104] In some embodiments, the recombinase is a B2 recombinase. An exemplary B2 recombinase comprises the amino acid sequence of SEQ ID NO:50. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, the nucleotide sequence shown in SEQ ID NO:53.

[0105] In some embodiments, the recombinase is a KD recombinase. An exemplary KD recombinase comprises the amino acid sequence of SEQ ID NO:51. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, the nucleotide sequence shown in SEQ ID NO:54.

[0106] In some embodiments, the recombinase is a pSR1 recombinase. An exemplary pSR1 recombinase comprises the amino acid sequence of SEQ ID NO:52. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, the nucleotide sequence shown in SEQ ID NO:55.

[0107] Based on one or more recombinase recognition sites (RS) in a first exogenous nucleotide sequence inserted into the genome, by providing a donor containing RS and a second exogenous nucleotide sequence, the second exogenous nucleotide sequence can be inserted into the genome of an organism such as a plant via recombination using the corresponding recombinase. The recombinase can be a separately expressed recombinase or included in the guide editing fusion protein. Those skilled in the art can select a suitable combination of RS located at the first exogenous polynucleotide already inserted into the genome and RS located at the donor to insert the second exogenous nucleotide sequence into the genome via recombination.

[0108] Therefore, in some implementations, the genome editing system further includes: iv) a recombinase and / or an expression construct containing a nucleotide sequence encoding the recombinase, and v) A donor construct containing one or more recombinase recognition sites (RS) and a second exogenous polynucleotide sequence to be inserted into the plant genome.

[0109] In some preferred embodiments, the recombinase is contained within the guided editing fusion protein. In some embodiments, the recombinase is located at the N-terminus of the guided editing fusion protein relative to the CRISPR nuclease and reverse transcriptase. In some embodiments, the recombinase is located at the C-terminus of the guided editing fusion protein relative to the CRISPR nuclease and reverse transcriptase.

[0110] The second exogenous polynucleotide sequence can be of any length. It can range from 1 bp to approximately 10 kb or longer. Preferably, the second exogenous polynucleotide is a long fragment, such as at least 300 bp, at least 500 bp, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, or longer. In some embodiments, the second exogenous polynucleotide can be a full-length gene.

[0111] In some embodiments, the second exogenous nucleotide sequence is associated with an organism such as a plant trait or agronomic trait, whereby the insertion of the second exogenous nucleotide sequence results in the organism such as a plant having altered (preferably improved) traits, such as agronomic traits, relative to a wild-type organism such as a plant.

[0112] Different components of the genome editing system of the present invention, such as the coding sequences of CRISPR nuclease, reverse transcriptase, guide editing fusion protein, pegRNA and / or recombinase, and the second exogenous polynucleotide sequence, may be located in the same construct in different combinations, or in different constructs respectively.

[0113] The genome editing system of this invention can be used to perform site-specific modifications, such as site-specific insertion of exogenous nucleotide sequences, in organisms that can be non-human animals, humans, or plants, preferably plants. Suitable plants include monocotyledonous and dicotyledonous plants, for example, crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0114] In order to achieve effective expression in organisms such as plants, in some embodiments of the present invention, the nucleotide sequence encoding the fusion protein is codon-optimized for the organism, such as the plant species, whose genome is to be modified.

[0115] Codon optimization refers to the modification of nucleic acid sequences to enhance expression in host cells of interest by replacing at least one codon of the natural sequence with codons that are used more frequently or most frequently in the gene in the host cell (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons while maintaining the natural amino acid sequence). Different species exhibit specific preferences for certain codons of specific amino acids. Codon preference (differences in codon use between organisms) is often associated with the translation efficiency of messenger RNA (mRNA), which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be customized to achieve optimal gene expression in a given organism based on codon optimization. Codon utilization tables are readily available, for example, in the Codon Usage Database (“Codon Usage Database”) available at www.kazusa.orjp / codon / , and these tables can be adapted in various ways. See Nakamura. Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).

[0116] III. Methods for site-specific modification of the plant genome, such as site-specific insertion of exogenous nucleotide sequences. On the other hand, the present invention provides a method for site-directed modification of a plant genome, comprising introducing the genome editing system of the present invention into at least one of said plants. The site-directed modification includes substitution, deletion, and / or addition of one or more nucleotides. For example, the site-directed modification includes the site-directed insertion of a foreign nucleotide sequence.

[0117] On the other hand, the present invention provides a method for producing genetically modified plants, said genetically modified plants comprising site-directed modifications, said method comprising introducing the genome editing system of the present invention into at least one said plant. The site-directed modifications include substitutions, deletions, and / or additions of one or more nucleotides. For example, the site-directed modifications include the site-directed insertion of a foreign nucleotide sequence.

[0118] In some embodiments, the method further includes screening plants from the at least one plant for plants with desired site-directed modifications, such as site-directed exogenous nucleotide sequence insertions.

[0119] In the method of this invention, the genome editing system can be introduced into plants using various methods well known to those skilled in the art. Methods for introducing the genome editing system of this invention into plants include, but are not limited to: gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the genome editing system is introduced into plants via transient transformation.

[0120] In some embodiments, the components of the genomic editing system are introduced into the plant simultaneously. In some embodiments, the components of the genomic editing system are introduced into the plant separately or sequentially.

[0121] In some implementations, the method includes the following steps: 1) Transform components i)-iv) of the genome editing system into isolated plant cells or tissues to obtain plant cells or tissues with a first exogenous nucleotide sequence containing recognition sites (RS) of one or more recombinases; 2) Component v) of the genome editing system is converted into the plant cells or tissues obtained in step 1), thereby obtaining plant cells or tissues containing the inserted second exogenous polynucleotide sequence; and 3) Regenerate a complete plant from the plant cells or tissues obtained in step 2).

[0122] In some embodiments, the exogenous nucleotide sequence is inserted into a safe harbor site in the plant genome, the safe harbor site being located in the plant genome. 1) At least 5kb away from the protein-coding region; 2) At least 30kb away from the miRNA coding region; 3) At least 20kb away from the lncRNA coding region; 4) At least 20kb away from the tRNA coding region; 5) At least 5 kb away from the promoter and / or enhancer; 6) The distance from the LTR repeat should be at least 20kb; 7) At least 200 bp away from non-LTR repeats; and 8) At least 10 kb away from the centromere.

[0123] In some implementations, the plant is rice, and the safe harbor sites are selected from the sites shown in Tables 1 and 2.

[0124] In some embodiments, the introduction includes converting the genome editing system of the present invention into isolated plant cells or tissues, and then regenerating the converted plant cells or tissues into complete plants. Preferably, no selectants targeting the selection genes carried on the expression vector are used during tissue culture.

[0125] In other embodiments, the genome editing system of the present invention can be transformed into specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young spikelets, or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate through tissue culture.

[0126] In some embodiments of the present invention, in vitro expressed proteins and / or in vitro transcribed RNA molecules (e.g., the expression construct is an in vitro transcribed RNA molecule) and / or donor DNA molecules are directly transformed into the plant.

[0127] In some embodiments, the method further includes treating (e.g., culturing) plant cells, tissues, or whole plants that have been introduced into the genome editing system at an elevated temperature (relative to the temperature of conventional culture, such as room temperature), said elevated temperature being, for example, 37°C.

[0128] In some embodiments of the invention, the site-directed modification, such as a site-directed insertion of a foreign nucleotide sequence and / or the target sequence is associated with plant traits such as agronomic traits, thereby the site-directed modification, such as a site-directed insertion, results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.

[0129] In some embodiments, the method further includes the step of screening plants with desired site modifications, such as site insertions, and / or desired traits, such as agronomic traits.

[0130] In some embodiments of the invention, the method further includes obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring have the desired modification (e.g., site-directed exogenous polynucleotide insertion) and / or desired traits such as agronomic traits.

[0131] In another aspect, the present invention also provides genetically modified plants or their offspring or portions thereof, wherein the plants are obtained by the methods described above. Preferably, the genetically modified plants or their offspring have desired genetic modifications (such as site-directed exogenous polynucleotide insertion) and / or desired traits such as agronomic traits.

[0132] In another aspect, the present invention also provides a plant breeding method, comprising crossing a genetically modified first plant obtained by the method described above with a second plant that does not contain the modification, thereby introducing the modification (such as site-directed exogenous polynucleotide insertion) into the second plant. Preferably, the genetically modified first plant and the second plant have desired traits, such as agronomic traits.

[0133] The genome editing system of the present invention can be used to perform site-specific modifications, such as site-specific insertion of exogenous nucleotide sequences, on plants including monocotyledonous and dicotyledonous plants. For example, the plants are crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0134] In another aspect, the present invention provides a method for producing genetically modified plants, said genetically modified plants comprising site-directed insertion of a foreign nucleotide sequence, said method comprising inserting the foreign nucleotide sequence into a safe harbor site in the plant genome. 1) At least 5kb away from the protein-coding region; 2) At least 30kb away from the miRNA coding region; 3) At least 20kb away from the lncRNA coding region; 4) At least 20kb away from the tRNA coding region; 5) At least 5 kb away from the promoter and / or enhancer; 6) The distance from the LTR repeat should be at least 20kb; 7) At least 200 bp away from non-LTR repeats; and 8) At least 10 kb away from the centromere.

[0135] In some implementations, the plant is rice, and the safe harbor sites are selected from the sites shown in Tables 1 and 2.

[0136] IV. Methods for site-specific modifications in the genomes of humans or non-human animals, such as site-specific insertion of exogenous nucleotide sequences. On the other hand, the present invention provides a method for site-directed modification of the genome of a human or non-human animal, comprising introducing the genome editing system of the present invention into at least one said human or non-human animal cell. The site-directed modification includes substitution, deletion, and / or addition of one or more nucleotides. For example, the site-directed modification includes the site-directed insertion of a foreign nucleotide sequence.

[0137] On the other hand, this invention provides the use of site-specific modification of human or non-human animal genomes for in vivo and in vitro gene therapy, enabling the deletion, addition, upregulation, downregulation, inactivation, activation, or mutation correction of disease-related genes, thereby achieving disease prevention and / or treatment. For example, the target nucleic acid region described in this invention can be located within the protein-coding region of a disease-related gene, or, for example, within a gene expression regulatory region such as a promoter region or enhancer region, thereby enabling modification of the function or expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes modification of the disease-related gene itself (e.g., protein-coding region), as well as modification of its expression regulatory regions (e.g., promoters, enhancers, introns, etc.).

[0138] On the other hand, the present invention provides a method for generating genetically modified human or non-human animal somatic cells, said genetically modified somatic cells comprising site-specific modifications, said method comprising introducing the genome editing system of the present invention into at least one said human or animal somatic cell. The site-specific modifications include substitution, deletion, and / or addition of one or more nucleotides. For example, the site-specific modifications include site-specific insertion of a foreign nucleotide sequence.

[0139] Therefore, the present invention also provides a method for treating a disease in a subject in need, comprising delivering an effective amount of the genome editing system of the present invention to the subject to modify a gene associated with the disease. The present invention also provides the use of the genome editing system in the preparation of a pharmaceutical composition for treating a disease in a subject in need, wherein the genome editing system is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject in need, comprising the genome editing system of the present invention and optionally a pharmaceutically acceptable vector, wherein the genome editing system is used to modify a gene associated with the disease. In some embodiments, the subject is a human being.

[0140] V. Reagent Kit The present invention also includes a kit for use with the methods of the present invention, the kit comprising at least components of the genome editing system of the present invention. The kit may also contain reagents for introducing said genome editing system into an organism or somatic cells. The kit generally includes a label indicating the intended use and / or method of use of the kit contents. Terminology labels include any written or documented material provided on or with the kit or otherwise accompanied by the kit.

[0141] Example Example 1. Design of a novel genome editing system 1.1. Filtering by the Guided Editor (PE) System Guided editing (PE) is a precise genome editing technique that can produce base changes and short DNA insertions and deletions without forming DSBs (Dual Substances in Generic DNA). It is widely used across species, such as humans, mice, rice, wheat, and maize. To develop novel genome editing systems, this embodiment first screens the efficiency of reported PE systems for endogenous target editing using a dual pegRNA strategy. Five PE system constructs were compared: PPE, Art-PPE (Cas9 5' fused with the mouse exonuclease Artemis), PPE-NCV1, and ePPE (Zong, Y., Liu, Y., Xue, C. et al. An engineered prime editor with enhanced editing efficiency in plants). Nat Biotechnol 40, 1394–1402 (2022).) and the ePPE-wtCas9 (replacing H840A-Cas9 in ePPE with wtCas9) construct, using a dual pegRNA strategy, were used to evaluate the efficiency of inserting two recombinase recognition sites (RS) of Lox66 (34 bp) and / or FRT1 (48 bp) at the endogenous target site. A schematic diagram of the vector construction strategy is shown below. Figure 1A .

[0142] Rice protoplasts were selected as the model cells. The above constructs were transformed into rice protoplasts using PEG transformation. The efficiency of five guided editing system constructs in inserting five pairs of pegRNAs into Lox66 or FRT1 at the endogenous site of the rice protoplast was tested. Next-generation sequencing results are shown below. Figure 1B .

[0143] The results showed that when using the dual pegRNA strategy for site-specific insertion, ePPE was the most efficient, improving the accuracy of insertion by 10-50 times compared to PPE.

[0144] 1.2. Guide RNA screening We further compared ordinary pegRNAs with tevoPre-containing eegRNAs that have been reported to improve PE efficiency (Nelson, JW, Randolph, PB, Shen, SP et al. Engineered pegRNAs improve prime editing efficiency). Nat BiotechnolEditing efficiency under different guided editing systems was tested in 40, 402–410 (2022). Four combinations were tested: PPE+pegRNA, ePPE+pegRNA, PPE+epegRNA, and ePPE+epegRNA. Vector construction was as follows: Figure 2A As shown.

[0145] The efficiency of the above four combinations in site-directed insertion of 8 pairs of pegRNA / epegRNA into RS was also tested using rice protoplasts. The next-generation sequencing results are as follows: Figure 2B .

[0146] The results showed that the dual-ePPE strategy mediated by ePPE+epegRNA (hereinafter referred to as "dual-ePPE") achieved the highest site-specific insertion efficiency, with some sites reaching over 50%, representing a more than 100-fold improvement compared to the ordinary PPE+pegRNA combination. It also demonstrated high efficiency at most inefficient target sites. Notably, while the editing tool improved the efficiency at the target site, it also increased the probability of inaccurate editing or insertion / deletion at other sites. While dual-ePPE significantly improved the accuracy of editing at the target site compared to other combinations, it did not significantly alter the insertion / deletion efficiency at other sites.

[0147] 1.3. Dual-ePPE Insertion The relationship between insert length (30bp-100bp) and the distance between the two pegRNAs (PAM distance 20bp-80bp) was further evaluated using dual-ePPE, as well as the impact of the overlap length between the two RTs (10bp-50bp) on insertion efficiency. Rice protoplast assays were used, and the next-generation sequencing results are as follows: Figure 3 As shown.

[0148] The results show that there is no obvious linear relationship between the insertion length and the distance between pegRNAs. Efficiency is higher when the insertion length is greater than the distance between pegRNAs, and high site-specific insertion efficiency is achieved when the overlap length between two RTs is between 10bp and 50bp. These results demonstrate that the ePPE+epegRNA system of this invention can achieve efficient and site-specific insertion of tag sequences such as Flag and Tag sequences.

[0149] Example 2. Optimization of the dual-ePPE system In order to further verify and optimize the effect of the dual-ePPE system of the present invention under different usage environments, and thus obtain the preferred technical solution, this embodiment verifies and analyzes the possible improvements of each component of the dual-ePPE system.

[0150] 2.1. CRISPR system effector proteins This embodiment designs SpG-Cas9 with an NGN PAM recognition sequence and SpRY-Cas9 variants that are almost unrestricted by PAM sequences (Christie KA, Guo JA, Silverstein RA, Doll RM, Mabuchi M, Stutzman HE, Lin J, Ma L, Walton RT, Pinello L, Robb GB, Kleinstiver BP. Precise DNAcleavage using CRISPR-SpRYgests. Nat Biotechnol. 2023 Mar;41(3):409-416.) into dual-ePPE and evaluates their insertion efficiency in NGN-containing PAMs, aiming to expand the targeting range of dual-ePPE. Vector construction is as follows. Figure 4A As shown.

[0151] The efficiency of site-specific insertion under three PAM combinations (NGA, NGC, and NGT) was tested using rice protoplasts. Next-generation sequencing results are as follows: Figure 4B As shown.

[0152] The results show that SpG-ePPE and SpRY-ePPE also have high efficiency in site-directed insertion of PAMs for NGA, NGC, and NGT, thus verifying that the dual-ePPE system can be applied to various CRISPR system effector proteins and can effectively perform their functions. These results demonstrate that the dual-ePPE of this invention can effectively achieve RS sequence insertion in plants.

[0153] 2.2. RT Sequence Synonymous Mutations Reports have shown that introducing synonymous mutations (SM) into RT sequences can improve recombination editing efficiency (Xu, W., Yang, Y., Yang, B.). et al. A design optimized prime editor with expanded scope and capability in plants. Nat. Plants 8, 45–52 (2022). This embodiment tested the editing efficiency of the system on RT for SM processing. The carrier was constructed as follows... Figure 4A As shown.

[0154] The efficiency of point protrusion at four target sites using two different RT methods was tested using rice protoplasts. Next-generation sequencing results are as follows: Figure 4C As shown in the figure. The results indicate that when there is a uniform mismatch between the RT sequence and the genomic sequence, the efficiency of point mutation can be greatly improved (4-20 times).

[0155] 2.3. Promoters that drive epigRNA expression Further investigation was conducted on the efficiency of inserting longer fragments (150bp-300bp) using the aforementioned system, and the editing efficiency of guide RNA expression using the U3 promoter and the compound type II promoter (pGS promoter) was tested. The pGS-epegRNA vector was constructed as follows: Figure 5A As shown.

[0156] The efficiency of site-directed insertion of fragments of different lengths using the U3 promoter and pGS promoter was compared in rice protoplasts. The ddPCR results are shown below. Figure 5B C. The results showed that there was no significant difference between the U3 promoter and the pGS promoter in performing small fragment insertion. Figure 5B When the insertion length reaches 150bp or more, the pGS promoter drives epigRNA more efficiently than the U3 promoter. Using the pGS promoter to express epigRNA can improve the efficiency of large fragment site-directed insertion by 2-5 times, and precise insertion can still be achieved when the insertion fragment length reaches 700bp.

[0157] 2.4. The impact of MS2-MCP and variable temperature processing on editing efficiency Further improvements were made to the efficiency of long fragment insertion by using an MS2-MCP system to recruit MLVs and a 37°C temperature treatment method. Vector construction was carried out as follows: Figure 6A As shown.

[0158] The efficiency of large fragment site-specific insertion using rice protoplasts was tested for both recruitment methods. The efficiency of 37℃ treatment (TT, normal culture 12h → 37℃ culture 12h → normal culture 24h) was also tested to see if it improved efficiency. The ddPCR results are as follows: Figure 6B As shown.

[0159] The results showed that using a temperature of 37℃ could improve the insertion efficiency of large fragments by approximately 1.2-5 times. Figure 6B Using the MS2-MCP system to recruit MLVs can improve the insertion efficiency of large fragments by approximately 2-4 times. Figure 6C ).

[0160] Example 3. Achieving large-fragment DNA insertion in plants without double-strand breaks using the PrimeROOT system. This embodiment uses dual-ePPE combined with recombinase as a prime editing-mediated Recombination of Opportune Targets (PrimeROOT) system and verifies its ability to insert DNA fragments in plants.

[0161] 3.1. Construction of the Fluorescence Reporting System To verify the DNA recombination capabilities of various recombinases in plant base editing, the inventors first constructed a fluorescent reporter system to characterize the DNA recombination efficiency of commonly used site-specific recombinases in rice protoplasts. This reporter system divides GFP into two domains: the N-terminal (GFP-N) and the C-terminal (GFP-C), each encoding a separate plasmid. Figure 7A Each of the two plasmids carries a recombinase site. See the schematic diagram of plasmid construction. Figure 7A Following recombinase expression and recombination, GFP-N and GFP-C are linked through an intron linker, enabling GFP expression in protoplasts. The activity of the recombinase in protoplasts can then be characterized by observation of GFP fluorescence using fluorescence microscopy and detection of GFP fluorescence using flow cytometry.

[0162] 3.2. PrimeROOT Construction for Detection The inventors constructed independent fluorescent reporter systems for six different tyrosine recombinases and two serine recombinases (all recombinases were codon-optimized and can be expressed in rice). GFP fluorescence microscopy observation results ( Figure 7B ) and flow cytometry search results ( Figure 7C The results show that the Cre and FLP recombinase system produces the strongest fluorescence and can be used as the best recombinase system for verifying and optimizing the effectiveness of the technical solution of the present invention.

[0163] In another set of parallel experiments, the inventors constructed fluorescent reporter systems targeting the Cre / Lox system and FLP / FRT system of the tyrosine recombinase family, as well as the ФC31 and Bxb1 recombinases of the serine family. The vector construction is as follows: Figure 7D As shown.

[0164] The above report system was transformed into rice protoplasts, and then observed under a fluorescence microscope and detected by flow cytometry. The results are as follows: Figure 7E As shown.

[0165] The results showed that the Cre / Lox and FLP / FRT systems had stronger DNA integration capabilities. Therefore, they were combined with the aforementioned site-directed insertion systems to transfer all components (ePPE, two epigRNAs, recombinase, and the gene to be inserted with the recombination site) into rice cells in a "one-step" process, achieving large-fragment site-directed insertion at the gene level. Figure 7F As shown. The above "one-step" reagent containing dual-ePPE, recombinase, and the gene to be inserted with a recombination site is named PrimeROOT.v1, and is named PrimeROOT.v1-Cre and PrimeROOT.v1-FLP respectively, depending on whether the recombinase is a Cre / Lox system or an FLP / FRT system.

[0166] 3.3. Verification of PrimeROOT.v1's ability to insert large fragments To verify the insertion capability of PrimeROOT.v1 into large DNA fragments, the inventors tested the integration efficiency of PrimeROOT.v1-Cre and PrimeROOT.v1-FLP into four endogenous sites of GFP (720 kp) in rice protoplasts using ddPCR. The experimental results are as follows: Figure 8 The results showed that both PrimeROOT methods achieved precise and targeted large fragment insertion at all four sites.

[0167] 3.4. System Optimization Because FRT1 contains short repetitive sequences, some FRT1 mutants have been reported to promote the efficiency of FLP recombinase (Bruckner, RC & Cox, MM Specific Contacts between the Flp Protein of the Yeast 2-Micron Plasmid and Its Recombination Site. Journal of Biological Chemistry 261, 1798-1807 (1986).; Senecoff, JF, Rossmeissl, PJ & Cox, MM DNA recognition by the FLP recombinase of the yeast 2-micron plasmid. A mutational analysis of the FLP binding site. J Mol Biol 201, 405-421 (1988).). Further optimization of the editing system is needed to obtain a better technical solution. The inventors artificially designed multiple FRT1 mutants (F1m1, F1m2, and F1m3) and two truncated FRT1 (tFRT1) sequence mutants (tF1m2 and tF1m3). During integration using PrimeROOT, ddPCR was used to evaluate the efficiency of one-step large fragment insertion at the endogenous target site, considering recombinase fusion and the presence or absence of the above recombinases, as well as the FRT variants. After inserting GFP into the rice endogenous gene using a one-step method, protoplast cells were luminescent. The ddPCR results are shown in Figure 9; the combination of FRT1 mutants exhibited higher mutation efficiency compared to the wild type.

[0168] 3.5. PrimeROOT System Optimization Based on PrimeROOT.v1, the inventors further optimized it to obtain a better technical solution. In this solution, the inventors fused ePPE from the PrimeROOT composition with a recombinase, and created two structural schemes based on different fusion sites. See example sequences. Figure 10A : In scheme 1, the recombinase was ligated to the N-terminus of the ePPE system via SV40 NLS and a 32-amino acid flexible linker, named PrimeROOT.v2N; in scheme 2, the recombinase was ligated to the C-terminus of the ePPE system via the same pathway, named PrimeROOT.v2C. Fluorescence microscopy and flow cytometry results showed that the PrimeROOT.v2N and PrimeROOT.v2C systems had higher GFP insertion efficiency at the four endogenous sites compared to PrimeROOT.v1 (Figure 10).

[0169] 3.6. Verification of PrimeROOT.v2's ability to insert large fragments To verify the insertion capability of PrimeROOT.v2 into large DNA fragments, the inventors constructed vector constructs containing any one or a combination of three genes (pigmR, OsMYB30, and OsHPPD), with donor lengths of 1.4 kb, 4.9 kb, 7.7 kb, and 11.1 kb, respectively. The vector construction is as follows: Figure 11A The inventors used ddPCR to detect the insertion efficiency of the four donors at four endogenous sites, and found that as the donor length gradually increased, precise and targeted large fragment insertion was achieved, and the editing efficiency did not decrease significantly. Figure 11B ).

[0170] Example 4. Achieving large-fragment DNA insertion without double-strand breaks in maize seeds using the PrimeROOT system. In addition to rice protoplasts, the inventors also evaluated the editing efficiency of dual-ePPE and its PrimeROOT in maize protoplasts.

[0171] The inventors first tested the precise RS insertion editing efficiency of dual-ePPE at six endogenous gene loci in maize protoplasts, and the experimental results showed that it could achieve an editing efficiency of up to 40%. Figure 12A ).

[0172] The inventors then tested PrimeROOT.v2C-Cre's editing efficiency on large GFP DNA fragments, and the experimental results showed that it achieved a GFP sequence editing efficiency of up to 4% at endogenous sites. Figure 12B ).

[0173] The experimental results are similar to the editing efficiency in rice, indicating that the dual-ePPE of the present invention and the PrimeROOT system composed therefrom have broad and universal application prospects in plant synthetic biology and gene editing engineering, and can precisely insert the required DNA sequence without introducing the donor backbone sequence.

[0174] Example 5. Editing capabilities of the PrimeROOT and CRISPR-mediated NHEJ system The CRISPR-mediated NHEJ system is currently the only reported system capable of targeted large-fragment insertion in plants (Li, J. et al. Gene replacements and insertions in rice by intron targeting using CRISPR-Cas9. Nature Plants 2 (2016).; Dong, OXO et al. Marker-free carotenoid-enriched rice generated through targeted gene insertion using CRISPR-Cas9. Nature Communications 11 (2020).). This example uses PrimeROOT.v2C-Cre as an example to compare the performance of PrimeROOT and the CRISPR-mediated NHEJ system in inserting GFP (720 bp), Act1 promoter (Act1P, 1.4 kb), and Act1P-... pigmR Gene cassette (4.9 kb) and Act1P- pigmR The targeted insertion capability of the -Act1P- OsMYB30 gene cassette (7.7 kb) was demonstrated. Results showed that both systems exhibited similar insertion efficiencies for GFP and Act1P insertions. However, for longer donor insertions, the PrimeROOT.v2C-Cre system demonstrated an average efficiency 2-4 times higher than the NHEJ system (see schematic diagram of the construct). Figure 13A See the editing efficiency chart. Figure 13B ).

[0175] Regarding editing accuracy, the inventors observed that Act1P events inserted using the PrimeROOT.v2C-Cre system showed clear Sanger sequencing results, but mixed peaks appeared in the results inserted using NHEJ. Figure 14A (The underscore indicates inaccurate insertion). This demonstrates that the PrimeROOT system offers superior editing precision compared to the traditional CRISPR-mediated NHEJ system.

[0176] Subsequently, the inventors cloned the edited insertion events from protoplasts into bacteria and sequenced the connections between the endogenous genome and the individual clone insertion fragments. When the inventors randomly selected 20 clones from the Act1P insertion samples processed by PrimeROOT and NHEJ, they found that all 20 insertions generated by PrimeROOT contained the exact inserted sequence as expected, while all 20 insertions generated by NHEJ contained random DNA base insertions and deletions / deletions at their ligation sites. Figure 14A B).

[0177] Next, the inventors used PrimeROOT and CRISPR-mediated NHEJ to combine Act1P and Act1P- pigmR Sequence insertion sites in rice callus genomic loci ( Figure 14C After transfer and induction of callus tissue, the inventors analyzed 95 callus clones from each treatment to compare editing efficiency and accuracy. PrimeROOT generated 2 accurate Act1P insertions and 2 accurate Act1P- pigmR Insertion, while NHEJ generated 3 inaccurate Act1P insertions and 1 inaccurate Act1P-. pigmR insert( Figure 14C The underscore indicates imprecise insertion. Figure 14D These results demonstrate that PrimeROOT is an effective editing tool for creating large, targeted, and precise DNA insertions, compared to the NHEJ system, which heavily relies on double-strand DNA breaks as intermediates.

[0178] Example 6: Precise and targeted insertion of the actin promoter using the PrimeROOT tool Many desirable agronomic traits are quantitative traits, dependent on the upregulation or downregulation of certain genes, or on tissue-specific expression. This embodiment utilizes the PrimeROOT system to precisely insert advantageous promoters upstream of target genes, thereby enabling the application of the PrimeROOT tool in plant trait improvement.

[0179] Specifically, the inventors used PrimeROOT.v2C-Cre to knock a strong promoter into the 5'UTR region of OsHPPD. Figure 15A The inventors first designed 16 pairs of pegRNAs in the 5' UTR and compared their RS insertion editing efficiency in rice protoplasts, determining that the optimal RS insertion frequency for the pegRNA pairs was 30%. Figure 15BNext, the inventors used PrimeROOT.v2C-Cre and the pegRNA to bombard rice Actin1 promoter (Act1P) particles into rice callus tissue. The inventors identified the edited plants by amplifying the connection between the genome and the inserted donor sequence, and assessed the insertion accuracy using Sanger sequencing. A total of 12 precise Act1P insertion events (2.4%) were detected in 507 regenerated rice plants. Figure 15C These results indicate that PrimeROOT can serve as an effective genome insertion tool to introduce novel genetic regulatory elements into plant genomes for breeding.

[0180] Example 7: Precise gene insertion in the GSH region To ensure the safe insertion of transgenes into the plant genome, the inventors predicted genomic safe harbor (GSH) regions throughout the Kitaake rice genome. This was based on previous research methods on GSH (Aznauryan, E. et al. Discovery and validation of human genomic safe harbor sites for gene and celltherapies). Cell Rep Methods 2, 100154 (2022). ; Sadelain, M., Papapetrou, EP&Bushman, FD Safe harbors for the integration of new DNA in the humangenome. Nat Rev Cancer 12, 51-58 (2011). The inventors used various algorithms to identify regions that are at a certain distance from certain elements (such as gene coding regions, small RNAs, miRNAs, lncRNAs, tRNAs, promoters, enhancers, LTRs, etc.). In this way, the inventors generated a new set of GSH regions, consisting of 30 regions, totaling 40 kb. Figure 16A The complete GSH regions of Kitaake are shown in Table 1. In addition, the inventors identified the GSH regions of 33 rice genomes, and their inter-genome mappings are shown in Table 2.

[0181] The inventors selected GSH1 (kitaake, Chr1:7660637-7661671) as the proof-of-concept region and designed four pairs of pegRNAs for inserting RS into this region (Table 3). When comparing RS insertion efficiencies using dual-ePPE in GSH1, the highest RS insertion efficiency was >40%. Figure 16B The inventors then tested 4.9kb of ActP1P- pigmR The donor box was inserted into the GSH1 region. Gel electrophoresis and Sanger sequencing results showed that 19 Act1- cells were identified in 744 regenerated plantlets. pigmR Insertion events (2.6%). Importantly, all 19 ligations produced amplification products of the same size, and sequencing showed that these were the results of exact insertion events, with the ends of the donor boxes perfectly matching the predictions.

[0182] Example 8: Methods for transferring from PrimeROOT and donor To investigate the insertion efficiency of PrimeROOT and donor components during plant editing, the inventors used Lox66 and the FRT mutant F1m2 as landing sites to test the recovery efficiency of whole-plant editing by sequentially transforming PrimeROOT and donor components into rice callus (the sequentially transformed system is referred to as PrimeROOT.v3). The inventors first evaluated dual-ePPE-mediated RS insertion into rice callus and achieved an editing efficiency as high as 84.7%. Figure 17A In the first round of transformation, the inventors transformed callus tissue with PrimeROOT reagent (without donor) using Agrobacterium. After one month of hygromycin selection, the inventors enriched callus tissue containing the desired RS insertions. These callus tissues were then used as substrates for a second round of transformation, containing donor vectors delivered by particle bombardment or Agrobacterium. After G418 selection and regeneration, the inventors examined the regenerated plants and measured the editing frequency of the desired insertion events. Figure 17B The inventors found that the Cre-Lox66 and FLP-F1m2 sites achieved editing efficiencies of 7.1% and 8.3% respectively for precise insertion of Act1P into the OsHPPD 5'UTR, representing improvements of 3-fold and 3.5-fold compared to one-step transformation. When evaluating the editing efficiency of Act1P-pigmR for precise insertion into GSH1, the inventors obtained 4.2% efficiency for the Cre-Lox66 site and 6.3% efficiency for the FLP-F1m2 site, representing improvements of 1.6-fold and 2.4-fold compared to integrated plant transformation. When the inventors delivered the product to the donor via Agrobacterium-mediated transformation, they obtained Act1P- pigmR The efficiency of precise insertion events into the GSH1 site was 3.9%. These results indicate that PrimeROOT.v3 can be used with different delivery methods and further improves the efficiency of precise targeted gene insertion in plants.

[0183] Example 9: Testing of large fragment insertion of PrimeROOT into human cells To test the functionality of PrimeROOT in human cells, the inventors first replaced the promoters of PrimeROOT.V2N-Cre and PrimeROOT.V2C-Cre with the commonly used human cell expression promoter, CMV. They designed pegRNAs in four regions: hAAVS1, hACTB, hCCR5, and hLMNB1, and constructed these pegRNAs into the hU6 expression vector. Subsequently, the plasmids, along with a GFP-containing donor plasmid, were transformed into the HEK293 cell line via plasmid transformation. Cell DNA was extracted after 72 hours, and ddPCR was performed to detect the efficiency. Figure 18A Simultaneously, junction PCR was performed for first-generation sequencing, revealing that site-specific integration of GFP into the genome is completely precise and predictable. Figure 18B This example demonstrates that the PrimeROOT system can precisely target gene insertion in human cells.

[0184] Table 1: Summary of GSH Regions

[0185] Table 2: Cross-mapping GSH regions of 33 rice genomes

[0186] Table 3 Information on the designed pegRNAs used to validate GSH1

[0187] Sequence information Wild-type SpCas9 amino acid sequence (SEQ ID NO:1) >nCas9(H840A) amino acid sequence (SEQ ID NO:2) >Wild-type M-MLV-RT amino acid sequence (SEQ ID NO.3) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP >M-MLV-RT-connection (SEQ ID NO.4) DQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDT >RT-RNase H (SEQ ID NO.5) PDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLL >NC(SEQ ID NO.6) ATVVSGQKQDRQGGERRRSQLDRDQCAYCKEKGHWAKDCPKKPRGPRGPRPQTSLL >PR(SEQ ID NO.7) TLDDQGGQGQEPPPEPRITLKVGGQPVTFLVDTGAQHSVLTQNPGPLSDKSAWVQGATGGKRYRWTTDRKVHLATGKVTHSFLHVPDCPYPLLGRDLLTKLKAQIHFEGSGAQVMGPMGQPLQVL >IN(SEQ ID NO.8) ENSSPYTSEHFHYTVTDIKDLTKLGAIYDKTKKYWVYQGKPVMPDQFTFELLDFLHQLTHLSFSKMKALLERSHSPYYMLNRDRTLKNITETCKACAQVNASKSAVKQGTRVRGHRPGTHWEIDFTEIKPGLYGYKYLLVFIDTFSGWIEAFPTKKETAKVVTKKLLEEIFPRFGMPQVLGTDNGPAFVSKVSQTVADLLGIDWKLHCAYRPQSSGQVERMNRTIKETLTKLTLATGSRDWVLLLPLALYRARNTPGPHGLTPYEILYGAPPPLVNFPDPDMTRVTNSPSLQAHLQALYLVQHEVWRPLAAAYQEQLDRPVVPHPYRVGDTVWVRRHQTKNLEPRWKGPYTVLLTTPTALKVDGIAAWIHAAHVKAADPGGGPSSRLTWRVQRSQNPLKIRLTREAP >M-MLV-RT-F155Y(SEQ ID NO.9) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAYFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP >M-MLV-RT-F155V(SEQ ID NO.10) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAVFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP >M-MLV-RT-F156Y (SEQ ID NO.11) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFYCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP >M-MLV-RT-D524N (SEQ ID NO.12) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTNGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP >M-MLV-RT-N200C (SEQ ID NO.13) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAYFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP >M-MLV-RT-ΔRNase H (SEQ ID NO.14) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPL >M-MLV-RT-ΔRNase H-ΔConnection (SEQ ID NO.15) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGP >Linker sequence (SEQ ID NO:16) SGGSSGGSSGSETPGTSESATPESSGGSSGGS >gRNA scaffold (SEQ ID NO:17) guuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugc >PPE (SEQ ID NO:18) TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TGGMGLAVRQAPLIIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPV QDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLP QGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQV KYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKA YQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLT KDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD ILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAE GKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRM ADQAARKAAITETPDTSTLLIENSSP SGGSPKKKRKV >ePPE(SEQ ID NO:19) >tevopre (SEQ ID NO:20) CGCGGTTCTATCTTAGTTACGCGTTAAACCAACTAGAA >pGS (SEQ ID NO:21) Atggagtcaaagattcaaatagaggacctaacagaactcgccgtaaagactggcgaacagttcatacagagtctcttacgactcaatgacaagaagaaaatcttcgtcaacatggtggagcacgacacacttgtctactccaaaaatatcaaagatacagtctcagaagaccaaagggcaattgagacttttcaacaaagggtaatatccggaaacctcctcggattccattgcccagctatctgtcactttattgtgaagatagtggaaaaggaaggtggctcctacaaatgccatcattgcgataaaggaaaggccatcgttgaagatgcctctgccgacagtggtcccaaagatggacccccacccacgaggagcatcgtggaaaaagaagacgttccaaccacgtcttcaaagcaagtggattgatgtgattggcagacatactgtcccacaaatgaagatggaatctgtaaaagaaaacgcgtgaaataatgcgtctgacaaaggttaggtcggctgcctttaatcaataccaaagtggtccctaccacgatggaaaaactgtgcagtcggtttggctttttctgacgaacaaataagattcgtggccgacaggtgggggtccaccatgtgaaggcatcttcagactccaataatggagcaatgacgtaagggcttacgaaataagtaagggtagtttgggaaatgtccactcacccgtcagtctataaatacttagcccctccctcattgttaagggagcaaaatctcagagagatagtcctagagagagaaagagagcaagtagcctagaagtagtcaaggcggcgaagtattcaggcacgtggccaggaagaagaaaagccaagacgacgaaaacaggtaagagctaagcatctagataagttgaaaacaatcttcaaaagtcccacatcgcttagataagaaaacgaagctgagtttatatacagctagagtcgaagtagtgatt >ФC31 (SEQ ID NO:22) MDTYAGAYDRQSRERENSSAASPATQRSANEDKAADLQREVERDGGRFRFVGHFSEAPGTSAFGTAERPEFERILNECRAGRLNMIIVYDVSRFSRLKVMDAIPIVSELLALGVTIVSTQEGVFRQGNVMDLIHLIMRLDASHKESSLKSAKILDTKNLQRELGGYVGGKAPYGFELVSETKEITRNGRMVNVVINKLAHSTTPLTGPFEFEPDVIRWWWREIKTHKHLPFKPGSQAAIHPGSITGLCKRMDADAVPTRGETIGKKTASSAWDPATVMRILRDPRIAGFAAEVIYKKKPDGTPTTKIEGYRIQRDPITLRPVELDCGPIIEPAEWYELQAWLDGRGRGKGLSRGQAILSAMDKLYCECGAVMTSKRGEESIKDSYRCRRRKVVDPSAPGQHEGTCNVSMAALDKFVAERIFNKIRHAEGDEETLALLWEAARRFGKLTEAPEKSGERANLVAERADALNALEELYEDRAAGAYDGPVGRKHFRKQQAALTLRQQGAEERLAELEAAEAPKLPLDQWFPEDADADPTGPKSWWGRASVDDKRVFVGLFVDKIVVTKSTTGRGQGTPIEKRASITWAKPPTDDDEDDAQDGTEDVAA >Bxb1(SEQ ID NO:23) MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLN GKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLLRGSVVERLHTGMSEF >Cre (SEQ ID NO:24) SNLLTVHQNLPALPVDATSDEVRKNLMDMFDRQAFSEHTWKMLLSVCRSWAAWCKLNNRKWFPAEPEDVRDYLLYLQARGLAVKTIQQHLGQLNMLHRRSGLPRPSDSNAVSLVMRRIRKENVDAGERAKQALAFERTDFDQVRSLMENSDRCQDIRNLAFLGIAYNTLLRIAEIARIRVKDISRTDGGRMLIHIGRTKTLVSTAGVEKALSLGVTKLVERWISVSGVADDPNNYLFCRVRKNGVAAPSATSQLSTRALEGIFEATHRLIYGAKDDSGQRYLAWSGHSARVGAARDMARAGVSIPEIMQAGGWTNVNVIVMNYIRNLDSETGAMVRLLEDGD >FLP (SEQ ID NO:25) MPQFDILCKTPPKVLVRQFVERFERPSGEKIALCAAELTYLCWMITHNGTAIKRATFMSYNTIISNSLSFDIVNKSLQFKYKTQKATILEASLKKLIPAWEFTIIPYYGQKHQSDITDIVSSLQLQFESSEEADKGNSHSKKMLKALLSEGESIWEITEKILNSFEYTSRFTKTKTLYQFLFLATFINCGRFSDIKNVDPKSFKLVQNKYLGVIIQCLVTETKTSVSRHIYFFSARGRIDPLVYLDEFLRNSEPVLKRVNRTGNSSSNKQEYQLLKDNLVRSYNKALKKNAPYSIFAIKNGPKSHIGRHLMTSFLSMKGLTELTNVVGNWSDKRASAVARTTYTHQITAIPDHYFALVSRYYAYDPISKEMIALKDETNPIEEWQHIEQLKGSAEGSIRYPAWNGIISQEVLDYLSSYINRRI >LoxP (SEQ ID NO:26) ATAACTTCGTATAGCATACATTATACGAAGTTAT Lox66 (SEQ ID NO:27) ATAACTTCGTATAGCATACATTATACGAACGGTA Lox71 (SEQ ID NO:28) taccgTTCGTATAGCATACATTATACGAAGTTAT Lox2272 (SEQ ID NO:29) Ataacttcgtataggatactttatacgaagttat FRT1 (SEQ ID NO:30) GAAGTTCCTATTCCGAAGTTCCTATTCTCTAGAAAGTATAGGAACTTC FRT6 (SEQ ID NO:31) Gaagttcctattccgaagttcctattcttcaaaaagtataggaacttc FRT1m1 (SEQ ID NO:32) GgAGgTCtTATTtCGAAGTTCCTATTCTCTAGAAAGTATAGGAACTTC FRT1m2 (SEQ ID NO:33) GAAGTTCCTATTCCGgAGgTCtTATTtTCTAGAAAGTATAGGAACTTC FRT1m3 (SEQ ID NO:34) GgAGgTCtTATTtCGAAGTTCCTATTCTCTAGAAAGTATAaGAcCTcC mFRT1 (SEQ ID NO:35) GAAGTTCCTATTCTCTAGAAAGTATAGGAACTTC mFRT1m1 (SEQ ID NO:36) GgAGgTCtTATTtTCTAGAAAGTATAGGAACTTC mFRT1m2 (SEQ ID NO:37) GAAGTTCCTATTCTCTAGAAAGTATAaGAcCTcC aTTP (SEQ ID NO:38) gtagtgccccaactggggtaacctttgagttctctcagttgggggcgtag aTTB (SEQ ID NO:39) cggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccac aGTP (SEQ ID NO:40) GGTTTGTCTGGTCAACCACCGCGGTCTCAGTGGTGTACGGTACAAACC aGTB (SEQ ID NO:41) GGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGG SpG-nCas9 (SEQ ID NO:42) SpRY-nCas9:(SEQ ID NO:43) MCP (SEQ ID NO:44) ASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY 2×MS2 (SEQ ID NO:45) GGGAGCACATGAGGATCACCCATGTGCCACGAGCGACATGAGGATCACCCATGTCGCTCGTGTTCCC PrimeROOT.v2N-Cre (SEQ ID NO:46) PrimeROOT.v2N-FLP (SEQ ID NO:47) PrimeROOT.v2C-Cre (SEQ ID NO:48) PrimeROOT.v2C-FLP (SEQ ID NO:49) Amino acid sequence of B2 recombinase (SEQ ID NO:50) MSEFSELVRILPLDQVAEIKRILSRGDPIPLQRLASLLTMVILTVNMSKKRKSSPIKLSTFTKYRRNVAKSLYYDMSSKTVFFEYHLKNTQDLQEGLEQAIAPYNFVVKVHKKPIDWQKQLSSVHERKAGHRSILSNNVGAEISKLAETKDSTWSFIERTMDLIEARTRQPTTRVAYRFLLQLTFMNCCRANDLKNADPSTFQIIADPHLGRILRAFVPETKTSIERFIYFFPCKGRCDPLLALDSYLLWVGPVPKTQTTDEETQYDYQLLQDTLLISYDRFIAKESKENIFKIPNGPKAHLGRHLMASYLGNNSLKSEATLYGNWSVERQEGVSKMADSRYMHTVKKSPPSYLFAFLSGYYKKSNQGEYVLAETLYNPLDYDKTLPITTNEKLICRRYGKNAKVIPKDALLYLYTYAQQKRKQLADPNEQNRLFSSESPAHPFLTPQSTGSSTPLTWTAPKTLSTGLMTPGEE Amino acid sequence of KD recombinase (SEQ ID NO:51) MSTFAEAAHLTPHQCANEINEILESDTFNINAKEIRNKLASLFSILTMQSLSIRREMKINTYRSYKSAIGKSLSFDKDDKIIKFTVRLRKTESLQKDIESALPSYKVVVSPFKNQEVSLFDRYEETHKYDASMVGLQFTNILSKEKDIWKIVSRIACFFDQSCVTTTKRAEYRLLLLGAVGNCCRYSDLKNLDPRTFEIYNNSFLGPIVRATVTETKSRTERYVNFYPVNGDCDLLISLYDYLRVCSPIEKTVSSNRPTNQTHQFLPESLARTFSRFLTQHVDEPVFKIWNGPKSHFGRHLMATFLSRSEKGKYVSSLGNWAGDREIQSAVARSHYSHGSVTVDDRVFAFISGFYKEAPLGSEIYVLKDPSNKPLSREELLEEEGNSLGSPPLSPPSSPRLVAQSFSAHPSLQLFEQWHGIISDEVLQFIAEYRRKHELRSQRTVVA Amino acid sequence of pSR1 recombinase (SEQ ID NO:52) MQLTKDTEISTINRQMSDFSELSQILPLHQISKIKDILENENPLPKEKLASHLTMIILMANLASQKRKDVPVKRSTFLKYQRSISKTLQYDSSTKTVSFEYHLKDPSKLIKGLEDVVSPYRFVVGVHEKPDDVMSHLSAVHMRKEAGRKRDLGNKINDEITKIAETQETIWGFVGKTMDLIEARTTRPTTKAAYNLLLQATFMNCCRADDLKNTDIKTFEVIPDKHLGRMLRAFVPETKTGTRFVYFFPCKGRCDPLLALDSYLQWTDPIPKTRTTDEDARYDYQLLRNSLLGSYDGFISKQSDESIFKIPNGPKAHLGRHVTASYLSNNEMDKEATLYGNWSAAREEGVSRVAKARYMHTIEKSPPSYLFAFLSGFYNITAERACELVDPNSNPCEQDKNIPMISDIETLMARYGKNAEIIPMDVLVFLSSYARFKNNEGKEYKLQARSSRGVPDFPDNGRTALYNALTAAHVKRRKISIVVGRSIDTS B2 recombinase recognition site (SEQ ID NO:53) GAGTTTCATTAAGGAATAACTAATTCССТAATGAAAACTC KD recombinase recognition site (SEQ ID NO:54) AAACGATATCAGACATTTGTCTGATAATGCTTCATTATCAGACAAATGTCTGATATCGTTT pSR1 recombinase recognition site (SEQ ID NO:55) TTGATGAAAGAATAACGTATTCTTTCATCAA Dre recombinase amino acid sequence (SEQ ID NO:56) MSELIISGSSGGFLRNIGKEYQEAAENFMRFMNDQGAYAPNTLRDLRLVFHSWARWCHARQLAWFPISPEMAREYFLQLHDADLASTTIDKHYAMLNMLLSHCGLPPLSDDKSVSLAMRRIRREAATEKGERTGQAIPLRWDDLKLLDVLLSRSERLVDLRNRAFLFVAYN TLMRMSEISRIRVGDLDQTGDTVTLHISHTKTITTAAGLDKVLSRRTTAVLNDWLDVSGLREHPDAVLFPPIHRSNKARITTTPLTAPAMEKIFSDAWVLLNKRDATPNKGRYRTWTGHSARVGAAIDMAEKQVSMVEIMQEGTWKKPETLMRYLRRGGVSVGANSRLMDS Dre recombinase recognition site rox (SEQ ID NO:57) TAACTTTAAATAATGCCAATTATTTAAAGTTA Dre recombinase recognition site rox (SEQ ID NO:58) TAACTTTAAATAATGTCCATTATTTAAAGTTA Figure 14A PrimeROOT.v2-Cre inserts S20-T2 (SEQ ID NO:59) GATCCTGTGCAATTTGAAAGGAACCCTGACGAGATTCCGTGGGCTGAATAACTTCGTATAGCATACATTATACGAAGTTATTCGAGGTCATTCATAT Figure 14A PrimeROOT.v2-Cre insertion at S20-T4 (SEQ ID NO:60) ATACCTTCAAGTGAGCAGCAGCCTTCTCCTTGTCAGTGAAGACACTACCGTTCGTATAATGTATGCTATACGAACGGTAGGTCTACCTAC Figure 14A NHEJ insertion at S20-T2 (SEQ ID NO:61) GATCCTGTGCAATTTGAAAGGAACCCTGACGAGATTCCGTGGGCTGAGGGTGGGCTTGGCTTTGTTTTCGGTCTCCGCCCCCCCGGGCGTTTTTATG Figure 14A NHEJ insertion at S20-T4 (SEQ ID NO:62) CAGTGAAGACACCGGTGGACTCCACGACATACTCAGCACCAGCCGGGTGGGCGGGACCTCTTCTACCTACAAAAAAGCTCCGCACGA Figure 14A NHEJ insertion at S20-T2 -80bp (SEQ ID NO:63) CGTGGAACTGATGTTT / / GA… / / … / / AAGGTGGTATA Figure 14A NHEJ insertion at S20-T2 +35bp (SEQ ID NO:64) CGTGGAACTGATGTTT / / ATGGCTGGGCTTGGCCTTGAATTCGAGCTCGGTACCCTCGA / / TCAGTTAAAAGGTGGTATA Figure 14A NHEJ insertion at S20-T2 +34 / -1bp (SEQ ID NO:65) CGTGGAACTGATGTTT / / GCTGGGCTTGGCCTTGAATTCGAGCTCGGTACCC-TCGAGG / / TCAGTTAAAAGGTGGTATA Figure 14A NHEJ insertion at S20-T2 +44 / -79bp (SEQ ID NO:66) CGTGGAACTGATGTTTCAGTA / / CGGA… / / …AAAAGAGTTG / / AAGGTGGTATA Figure 14A NHEJ insertion S20-T2 +62 / -4bp (SEQ ID NO:67) CGTGGAACTGATGTTT / / GAGGGAGAGGCGGTG / / GCTCGCTGCGCTC…GGTCGTTCA / / AGTTAAAAGGTGGTATA Figure 14A NHEJ insertion S20-T2 +289 / -1bp (SEQ ID NO:68) CGTGGAACTGATGTTT / / TGCGTTTCTGGGTGAG / / TCGAGCTCGGTACCC-TCGAGGTC / / CAGTTAAAAGGTGGTATA Figure 14A NHEJ insertion S20-T2 +271 / -7bp (SEQ ID NO:69) CGTGGAACTGATGTTT / / TGGTCGTTCGCTCCAAGC / / CTCGAGGTCAT…TCGAGGTC / / TAAAAGGTGGTATA Figure 14A NHEJ insertion S20-T2 +151bp (SEQ ID NO:70) CGTGGAACTGATGTTT / / GAGTTTTCGTTCCACTGACT / / TAATTCGAGCTCGGTACCCCG / / AGTTAAAAGGTGGTATA Figure 14A NHEJ insertion S20-T2 +72bp (SEQ ID NO:71) CGTGGAACTGATGTTT / / TGAGGTAAGATTACCTGGTC / / GAATTCGAGCTCGGTACCCTC / / AGTTAAAAGGTGGTATA Figure 14A NHEJ insertion S20-T2 +183 / -92bp (SEQ ID NO:72) CGTGG / / AA-A / / ATCC / / AAAT… / / … / / AAGGTGGTATA Figure 14A NHEJ insertion S20-T4 +15bp (1) (SEQ ID NO:73) CTTTTCATGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGCCGGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +12bp (SEQ ID NO:74) CTTTTCATGATTTGTGACAAATGCAGCCT / AGAAGAGGTACCGGCCAGCCTGGTGCTGAGTATGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +109 / -2bp (SEQ ID NO:75) CCTTTCATGATTTGTGACAAATGC / / GAAGAGGTAT / / CTGCGTTA--GCCGGTGCTGAGTATGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +15bp (2) (SEQ ID NO:76) CTTTTCACGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGCCGGCTGGTGCTGAGTGTGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +172 / -5bp (SEQ ID NO:77) CTTTTCATGATTTGTGACAAATG / / GAAGAGGTAC / / AGACCCCGT-----TGGTGCTGAGTATGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +15bp (3) (SEQ ID NO:78) CTTTTCACGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGCCGGCTGGTGCTGAGTGTGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +47 / -2bp (SEQ ID NO:79) CTTTTCATGATTTGTGACAAATGC / / GAAGAGGTACC / / AGCTCGG--GGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +143 / -2bp (SEQ ID NO:80) CTTTTCATGATTTGTGGCAAATGC / / GAAGAGGTACC / / CCAGCCT--GGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +15bp (4) (SEQ ID NO:81) CTTTTCATGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGTCGGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG Figure 14A NHEJ insertion S20-T4 +15bp (5) (SEQ ID NO:82) CTTTTCATGATTTGTGACAAATGCATC / AAAGAGGTACCGGCCAGCCGGCTGGCGCTGAGTATGTCGAGGAGTCCACCGG Figure 14C PrimeROOT.v2C-Cre-Act1P or PrimeROOT.v2C-Act1P-pigmT insertion (SEQ ID NO:83) GAAGCATCTGTCTGTCCАCТССССАCТCGTATAACТТCGTATAGCATACATTATACGAA Figure 14C NHEJ strategy-ActP insertion ins 36bp (SEQ ID NO:84) CTCCCACTCGTTGCACGGGCTTGGC / / TTCGAGCTCGGTACCCTCGAGGTCATTCATA Figure 14C NHEJ strategy-ActP insertion ins 12 / del 1bp (SEQ ID NO:85) ATGTCTGTCCАCТCCCCАCTCGAGCTCGGTACCC-TCGAGGTCATTCATATGCTTGAGA Figure 14C NHEJ strategy-ActP insertion ins 131bp (SEQ ID NO:86) CTCCCCACTCGTTTTTCCGAAGGTAAC / / CGAGCTCGGTACCCTTCGAGGTCATTCATA Figure 14C NHEJ strategy - ActP - pigmR insertion ins 38 / del 11bp (SEQ ID NO:87) CTCCCCACTGCTTGGC / / CCTCCCTGC-----------CATTCATATGCTTGAGA Figure 18B GFP - N - terminal - hLMNB1 - F (SEQ ID NO:88) ACGGCATGGACGAGCTGTACAAGTAATTTTTTTTACCGTTCGTATAGCATACATTATACGAACGGTAAGCCCCACGCGCCTGTCGCGGCTCCAGGAGAAGGAGGAGCTGCGCGAGCTCAATGAC Figure 18B GFP - C - terminal - hLMNB1 - R (SEQ ID NO:89) CCCGGTGAACAGCTCCTCGCCCTTGCTCACCATATAACTTCGTATAATGTATGCTATACGAAGTTATCGGGCGGCGGAGACAGCGGGGCGGCGAGGCCGCGAGCGGGACCGTGATAAGGAG。

Claims

1. A recombinase recognition site for inserting a foreign nucleotide sequence into the genome, characterized in that, The recombinase recognition site is the FRT1 variant, which is selected from the nucleotide sequences shown in SEQ ID NO:32-37.

2. The recombinase recognition site according to claim 1, characterized in that, The recombinase that binds to the recognition site is an FLP recombinase.

3. A site-specific recombination system, characterized in that, The site-specific recombination system includes: I) Any of the following combinations: 1) A nucleic acid molecule or its expression construct containing a wild-type FRT1 recognition site, and at least one nucleic acid molecule or its expression construct containing a recombinase recognition site as described in claim 1 or 2. 2) At least two nucleic acid molecules or their expression constructs containing the recombinase recognition site as described in claim 1 or 2; and II) FLP recombinase, a fusion protein containing FLP recombinase, or an expression construct encoding said FLP recombinase / fusion protein. The nucleotide sequence of the wild-type FRT1 recognition site is SEQ ID NO:

30.

4. The site-specific recombination system according to claim 3, characterized in that, The site-specific recombination system further includes: III) A genome editing system that introduces a wild-type FRT1 recognition site or a recombinase recognition site as described in claim 1 or 2 into the cell genome.

5. The site-specific recombination system according to claim 4, characterized in that, The genome editing system is selected from a wide range of genome editing systems mediated by nucleases, zinc finger proteins, TALEN, CRISPR-Cas, topoisomerases, or recombinases.

6. A method for inserting a foreign nucleotide sequence into a genome, characterized in that, The insertion system used is the PrimeROOT system, which includes: i) A donor construct comprising a wild-type FRT1 recognition site or a recombinase recognition site as described in claim 1 or 2, and a second exogenous nucleotide sequence to be inserted into the plant genome; ii) Any of the following components: a) An expression construct containing a CRISPR nuclease and / or a nucleotide sequence encoding the CRISPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase. b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein is composed of a CRISPR nuclease linked with a reverse transcriptase; iii) An expression construct containing a first pegRNA and / or a nucleotide sequence encoding the first pegRNA; iv) A second pegRNA and / or an expression construct containing a nucleotide sequence encoding the second pegRNA; and v) An FLP recombinase and / or an expression construct containing a nucleotide sequence encoding the FLP recombinase; The first pegRNA contains, from the 5' end to the 3' end, a first guide sequence, a first scaffold sequence, a first reverse transcription template (RT) sequence, and a first primer binding site (PBS) sequence, and the first pegRNA targets the first target sequence on the sense strand of the plant genomic DNA. The second pegRNA contains, from the 5' end to the 3' end, a second guide sequence, a first scaffold sequence, a second reverse transcription template (RT) sequence, and a second primer binding site (PBS) sequence. The second pegRNA targets the second target sequence on the antisense strand of the plant genomic DNA. The first RT sequence and the second RT sequence are used to insert a first exogenous nucleotide sequence into the plant genome, the first exogenous nucleotide sequence containing one or more wild-type FRT1 recognition sites or recombinase recognition sites as described in claim 1 or 2; The pegRNA can form a complex with the CRISPR nuclease or the guided editing fusion protein and target the CRISPR nuclease or the guided editing fusion protein to a target sequence in the plant genome, thereby causing a cut in the target strand within the target sequence. The CRISPR nuclease is a Cas9 nickase, the Cas9 nickase being selected from the amino acid sequence of SEQ ID NO:2, 42-43, and the reverse transcriptase is M-MLV reverse transcriptase or an optimized variant thereof.

7. The method according to claim 6, characterized in that, The FLP recombinase is contained in the guided editing fusion protein described in ii)-b), wherein the guided editing fusion protein containing the FLP recombinase consists of the amino acid sequence shown in SEQ ID NO:47 or 49.

8. The method according to any one of claims 6-7, characterized in that, The second exogenous nucleotide sequence is 1 bp to 11.1 kb in length.

9. The method according to claim 8, characterized in that, The first target sequence, the second target sequence, the first exogenous nucleotide sequence, and / or the second exogenous nucleotide sequence are associated with plant traits. After the first exogenous nucleotide sequence and / or the second exogenous nucleotide sequence are inserted into the plant genome, the plant has altered plant traits relative to the wild type plant.

10. The method according to claim 9, characterized in that, The plants include monocotyledonous and dicotyledonous plants, such as crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

Citation Information

Patent Citations

  • Genome editing system and method

    WO2018149418A1