A method for site-specific insertion of exogenous sequences into a genome
By combining the dual pegRNA and ePPE system with the Cre/Lox or FLP/FRT system, the problem of low insertion efficiency of large exogenous sequences in higher plant cells has been solved, achieving efficient insertion of exogenous sequences, especially significantly improving insertion efficiency in rice and maize.
Patent Information
- Application Number
- CN202310599213.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-06
- Filing Date
- 2023-05-25
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2043-05-25
AI Technical Summary
Existing technologies struggle to achieve efficient site-specific insertion of large exogenous sequences into higher plant cells, particularly due to the stability and efficiency issues of reverse transcriptase, resulting in low insertion efficiency.
By employing a dual pegRNA strategy and an enhanced plant guided editing system (ePPE) combined with the Cre/Lox system or FLP/FRT system from the tyrosine recombinase family, large-fragment exogenous sequences are inserted site-specifically using the SSA repair pathway by providing donor DNA with recombination sites.
This technology enables efficient site-directed insertion of short and large exogenous sequences into plant cells, improving insertion efficiency and enhancing the stability and insertion length of reverse transcriptase.
Smart Images

Figure BDA0004248229970000251 
Figure BDA0004248229970000261 
Figure BDA0004248229970000262
Abstract
Description
[0001] This application claims priority to Chinese Patent Application No. 202210580767.0, filed May 25, 2022, and Chinese Patent Application No. 202310363943.X, filed April 6, 2023. TECHNICAL FIELD
[0002] The present application belongs to the field of genetic engineering. Specifically, the present application relates to a method for site-specific insertion of exogenous sequences in the genome. Specifically, the present application is based on the prime editing system (PE) using two adjacent pegRNAs with partially overlapping sequences on the reverse transcription template to achieve efficient and precise site-specific insertion of exogenous sequences in the genome, particularly in plant genomes. Further coupling this system with recombinase systems such as Cre / Lox or FLP / FRT, etc. to achieve site-specific insertion of large fragments of exogenous sequences in the genome, particularly in plant genomes. BACKGROUND
[0003] The rapid development of DNA sequencing technology has enabled the life science field to rapidly enter the genomic era. The emergence of technologies represented by GWAS has greatly promoted the development of genetics, especially the resolution of numerous key gene functions in plants, which is a great opportunity for the development of molecular crop breeding. Traditional crop breeding methods represented by hybridization and backcrossing have been insufficient to support the rapid growth of crop breeding due to time and labor consumption, etc. Therefore, the development of new molecular breeding techniques is increasingly important.
[0004] Transgenic technology has been rapidly applied in plant molecular breeding due to its ability to quickly and efficiently obtain excellent traits. However, it is strictly regulated due to the introduction of exogenous genes. In contrast, genome editing technology can perform site-specific and precise modification of functional genes without introducing exogenous genes, thereby more quickly and efficiently obtaining excellent traits. Current plant genome editing tools mainly include three categories: zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and clustered regularly interspaced short palindromic repeats and their associated proteins (CRISPR / Cas). Among them, the CRISPR / Cas system is the most convenient and efficient, and has made a significant contribution to genetic research and plant molecular breeding in recent years.
[0005] The widely used CRISPR / Cas system includes a single guide RNA (sgRNA) and a site-specific nuclease Cas9. sgRNA targets the specific genomic DNA through base pairing, and Cas9 forms a ribonucleoprotein complex (RNP) with sgRNA in cells. At the same time, the conformation of Cas9 changes, and a domain on Cas9 (PAM-interaction domain, PI domain) continuously interacts with the motif NGG (PAM) at each position on the genome until it finds a position that can base pair with sgRNA. At this time, the RNP complex interacts with DNA to form a new complex. Cas9 unwinds the DNA double helix to form an R-loop, and its RuvC and HNH nuclease active domains are activated to cut the non-target strand and target strand, respectively, to produce a DNA double-strand break (DSB). At this time, the DNA double-strand break will trigger the cell's endogenous DNA repair mechanism, which will generally be repaired by the highest frequency of non-homologous end joining (NHEJ). NHEJ is an error-prone repair pathway, so some base insertions or deletions (Indels) may be randomly introduced near the DSB during the repair process, resulting in the inability of the gene to be normally expressed. In the process of generating DSB, if a piece of exogenous DNA (donor) with homologous arms on both sides of the DSB is provided, the cell's endogenous repair mechanism may use the donor as a template for homologous recombination repair (HR). HR is a precise repair pathway that can introduce any point mutations, fragment insertions, and deletions on the genome. However, this repair pathway occurs at a very low frequency in higher organism cells, especially in plant cells, and therefore has not been widely used. Subsequently, the base editor (BE) based on the CRISPR system was developed. The BE system uses a RuvC domain-inactivated Cas9 (nCas9-D10A) coupled with a deaminase (cytosine deaminase or adenine deaminase, corresponding to CBE and ABE, respectively). When the RNP complex binds to DNA to form an R-loop, the deaminase deaminates cytosine (C) or adenine (A) on the non-target strand to form uracil (U) or hypoxanthine (I), respectively. The cell's repair mechanism recognizes uracil as thymine (T) and hypoxanthine as guanine (G). At this time, nCas9 cuts the target strand, thereby promoting the cell to undergo the base excision repair pathway (BER) to repair C-U-T or A-I-G. The BE system does not need to rely on the generation of DSB and the HR pathway to complete efficient and precise point mutations, so it has been rapidly and widely used. The toolbox of CRISPR system coupled with other effector factors represented by BE has also been rapidly developed, including coupling with transcriptional activators, inhibitors, or epigenetic modification factors for targeted activation, inhibition, and epigenetic modification.
[0006] Although the CRISPR molecular toolbox has rapidly evolved from simple gene knockout to precise base editing, to transcriptional activation, repression, and epigenetic modification, targeted precise insertion of DNA fragments in higher plant cells has been difficult to achieve. Traditional strategies for achieving targeted insertion rely on the generation of DSBs, when an additional donor DNA that does not contain genomic homologous sequences is provided, the donor is likely to be inserted near the DSB through the NHEJ repair pathway after the DSB is generated, but this process is very imprecise, and its efficiency is also low due to donor provision and other problems; when an additional donor DNA containing genomic homologous sequences is provided, the target fragment in the donor is likely to be inserted into the target site through the HR repair pathway after the DSB is generated, but the efficiency of this process is extremely low, and it is almost impossible to achieve in higher plant cells.
[0007] Because of the low efficiency of HR, site-specific recombination (SSR) can be used to achieve the integration of large DNA fragments. SSR can specifically recognize and bind to a certain DNA sequence (recombination site, RS) and form a complex, and the strand exchange process between two complexes can occur and complete the recombination of DNA. This process is mediated by the attack of the tyrosine or serine residue in the active center of SSR on the phosphate backbone of RS, resulting in the cleavage of DNA. After cleavage, a covalent intermediate is formed and a strand exchange reaction between two RSs occurs. This process does not require the participation of high-energy cofactors and does not rely on endogenous DNA repair pathways in cells, and is more efficient. SSR enzymes can be divided into tyrosine recombinase family and serine recombinase family according to the difference of their active center residues. They are mainly derived from bacteriophages, bacteria and fungi, and play biological functions such as excision, inversion, integration and transposition. Common tyrosine recombinases include E. coli bacteriophage lambda integrase, P1 bacteriophage Cre recombinase and yeast FLP recombinase, which use a conserved tyrosine residue to attack one strand of the RS backbone, exposing the 5' phosphate group and 3' hydroxyl group. At this time, the 5' phosphate group and 3' hydroxyl group of two RSs are combined to achieve strand exchange, and the recombinase bound to the RS is changed in structure, attacking the other strand and achieving strand exchange through the same way, thereby completing the recombination process. Common serine recombinases include Tn3 transposase, Salmonella recombinase Hin, Streptomyces phage ΦC31 integrase and Mycobacterium phage Bxb1 integrase, which have similar recombination processes as tyrosine recombinases. The difference is that they use a serine residue to attack both strands of the RS backbone, achieving simultaneous exchange of two strands of two RSs, thereby completing the recombination process. SSR has a wide range of applications: in vitro, it is mainly used as a molecular cloning tool. The high efficiency of DNA intermolecular recombination makes it very simple to perform in vitro molecular cloning of large fragments and multiple fragments; in prokaryotic cells, it can be used as a gene or chromosomal engineering tool to delete, invert, translocate or integrate large DNA fragments; in higher eukaryotic cells, it is mainly used as a tool for deleting transgenic marker genes. However, due to the difficulty of site-specific knock-in of RS, it is very difficult to achieve site-specific integration of large DNA fragments.
[0008] Recently, a prime editing system (PE) that can achieve arbitrary base mutations, short DNA insertions and deletions has been developed, and it has been widely used in animals and plants for genome editing due to its powerful and DSB-independent function. The prime editing system uses a Cas9 with an inactive HNH domain (nCas9-H840A) coupled with a reverse transcriptase (MLV), and a reverse transcription template sequence (RT) and a reverse transcriptase primer binding site (PBS) are introduced in sequence at the 3' end of the sgRNA. The RT has a desired mutation sequence and homologous sequences on both sides of the mutation sequence, and this sgRNA is referred to as a pegRNA. After nCas9 cuts the non-target strand, PBS binds to the 5' end to serve as the starting primer for reverse transcriptase, which then extends to the 3' end of the RT sequence, reverse transcribing the RT sequence into DNA to form a 3' overhang with a mutation sequence. After endogenous DNA repair in the cell, the mutation sequence can be introduced into the genome, thereby completing genome editing of a certain length and any type.
[0009] The efficiency of the prime editing system in higher plant cells is still too low to achieve efficient insertion, and the length of the inserted fragment is very limited. The main reasons are speculated to be threefold. First, the repair pathway used by the prime editing system in higher plants occurs at a low frequency, resulting in low final editing efficiency. Second, the RT and the homologous sequence of the genome competitively bind to the genomic DNA, hindering the reverse transcription process. Third, the reverse transcriptase or pegRNA is easily degraded or lacks sufficient reverse transcription capacity. There is still a need in the art for systems and methods that can achieve efficient insertion of exogenous nucleotide sequences, particularly large exogenous nucleotide sequences, in plant genomes. SUMMARY
[0010] To avoid the first two reasons for the low efficiency of the above PE in higher plants, the inventors first designed a double-pegRNA strategy, in which two pegRNAs target and bind to two strands of genomic DNA with a certain distance (about 20 bp to about 60 bp) between PAMs, and the RT of each pegRNA contains only the desired insertion sequence and has a partial overlapping sequence at the 3' end. After reverse transcription, the two newly synthesized DNA strands are combined and annealed due to the overlapping sequence, and the insertion is completed through a different DNA repair pathway from the original PE system (according to some results of the present application, this repair pathway is likely to be SSA, a repair pathway that occurs at a high frequency in plants).
[0011] Recently, an enhanced version of the plant prime editing system (ePPE) was established by fusing the retroviral nucleocapsid protein (NC) and deleting the RNaseH active domain of the reverse transcriptase MLV, which can enhance the reverse transcription ability or enhance the stability of reverse transcriptase, thereby greatly improving the efficiency of the plant prime editing system. In addition, by adding a secondary structure tevopre at the 3' end of the pegRNA (epegRNA), it can also enhance the reverse transcription ability or enhance the stability of pegRNA and improve the efficiency of PE.
[0012] To further improve the insertion efficiency, the inventors used the ePPE system and epegRNA described above at the same time, thereby achieving efficient short fragment site-specific insertion in plant somatic cells. At the same time, the DNA integration ability of the Cre / Lox system in the tyrosine recombinase family, the FLP / FRT system, etc. and the ФC31, Bxb1 recombinase system in the serine family in rice somatic cells were evaluated, and it was found that the Cre / Lox system and the FLP / FRT system were better, so they were combined with the above-mentioned efficient insertion system, and by additionally providing a donor with a RS required insertion gene, the site-specific insertion of large fragment exogenous nucleotide sequence was realized by one-step method. SUMMARY
[0013] Figure 1. Test the efficiency of five constructs using double pegRNA to insert Lox66 or FRT1 in rice protoplasts.
[0014] Figure 2. Test the efficiency of inserting RS using PPE+pegRNA, ePPE+pegRNA, PPE+epegRNA, ePPE+epegRNA.
[0015] Figure 3 . Use ePPE+epegRNA combination to evaluate the relationship between insertion length (30bp-100bp) and distance between two pegRNAs (PAM distance 20bp-80bp) and the length of overlap between two RTs (10bp-50bp) on the insertion efficiency.
[0016] Figure 4. The efficiency of site-specific insertion on NG PAM when Cas9 is replaced by SpG-Cas9 or SpRY-Cas9.
[0017] Figure 5. The effect of pegRNA driven by different promoters on insertion efficiency.
[0018] Figure 6. The effect of 37-degree temperature treatment (6B) and the system of recruiting MLV using MS2-MCP (6C) on long fragment insertion efficiency.
[0019] Figure 7. A) Schematic of GFP reporter system; B) Evaluation of 8 recombinases editing efficiency and corresponding recombinase site sequence. Microscope images are of rice protoplasts transformed with or without corresponding recombinases; C) Evaluation of recombinase editing efficiency using fluorescent reporter system; D) Schematic of construct used to evaluate DNA integration ability of recombinases using fluorescent reporter system; F) Schematic of construct used for one-step large fragment insertion with recombinases combined with ePPE.
[0020] Figure 8 . Evaluation of insertion efficiency of PrimeROOT.vl system by ddPCR.
[0021] Figure 9. A) Flow cytometry detection of percentage of GFP positive plant protoplast cells, reflecting efficiency of one-step large fragment insertion using different combinations of recombinases; B) ddPCR determination of GFP insertion efficiency of OsALS in rice protoplasts.
[0022] Figure 10. Evaluation of editing efficiency of different base editing systems using fluorescent microscope and flow cytometry.
[0023] Figure 11. ddPCR detection of percentage of insertion of different donors into four endogenous sites.
[0024] Figure 12. ddPCR detection of percentage of insertion of six endogenous sites in maize using PrimeROOT.v2C-Cre system.
[0025] Figure 13. ddPCR detection of percentage of insertion of different gene editing systems in large fragment insertion.
[0026] Figure 14. Comparison of precise editing efficiency of PrimeROOT.v2C-Cre and NHEJ using base sequencing results.
[0027] Figure 15. A) Schematic of insertion of Actl promoter into OsHPPD site using PrimeROOT.v2C-Cre; B) Screening of pegRNA pairs; C) Insertion efficiency.
[0028] Figure 16. High-throughput sequencing of GSH site and high-throughput sequencing detection of insertion efficiency of inserted recombinase site in GSHl.
[0029] Figure 17. Schematic of PrimeROOT.v3 and efficiency of precise insertion by PrimeROOT.v3.
[0030] Figure 18. Efficiency and sequencing results of precise insertion in human HEK293 cells using PrimeROOT system. DETAILED DESCRIPTION
[0031] I. DEFINITIONS
[0032] In the present application, the scientific and technical terms used herein have the meanings commonly understood by one of ordinary skill in the art, unless otherwise indicated. Also, the terms and phrases used in the context of protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology, and molecular genetics have the same meaning as commonly understood by one of ordinary skill in the art in the field of the relevant art. For example, standard recombinant DNA and molecular cloning techniques used in the present application are well known in the art and are described more fully in Sambrook, J., Fritsch, E. F. and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter "Sambrook"). Also, the definitions and explanations of the relevant terms are provided below for better understanding of the present application.
[0033] As used herein, the term "and / or" encompasses all combinations of the items linked by the term. For example, "A and / or B" covers "A", "B", and "A and B". For example, "A, B, and / or C" covers "A", "B", "C", "A and B", "A and C", "B and C", and "A and B and C".
[0034] The word "comprise" as used herein, when used in relation to a protein or nucleic acid sequence, means that the protein or nucleic acid can consist of the sequence, or can have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in the present application. Furthermore, it is clear to a person skilled in the art that the methionine encoded by the start codon at the N-terminus of a polypeptide is in some practical cases (e.g. when expressed in a particular expression system) retained, but does not materially affect the function of the polypeptide. Therefore, when describing a specific polypeptide amino acid sequence in the specification and claims of the present application, although it can not comprise the methionine encoded by the start codon at the N-terminus, it is also intended to encompass sequences comprising the methionine, and correspondingly, the encoding nucleotide sequence can comprise the start codon; vice versa.
[0035] As used herein, "genome editing system" refers to a combination of components required for genome editing of a genome in a cell. Each component of the system, such as a guide-editing fusion protein or an expression construct thereof, a pegRNA or an expression construct thereof, a donor construct, etc. can exist independently, or can exist in any combination as a composition.
[0036] "Genome" as used herein encompasses not only chromosomal DNA present in the nucleus of a cell, but also organelle DNA present in subcellular components of a cell, such as mitochondria, plastids.
[0037] "Genetically modified plant" as described herein means a plant comprising an inserted exogenous polynucleotide within its genome. The exogenous polynucleotide can be stably integrated into the genome of the plant and inherited through successive generations, for example.
[0038] "Exogenous" with respect to a sequence means a sequence from a foreign species, or if from the same species, a sequence that has been significantly altered from its natural form by deliberate human intervention in its constitution and / or locus.
[0039] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and are single- or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single-letter designation: "A" is adenine or deoxyadenosine (RNA or DNA, respectively), "C" is cytosine or deoxycytosine, "G" is guanine or deoxyguanosine, "U" is uridine, "T" is deoxythymidine, "R" is purine (A or G), "Y" is pyrimidine (C or T), "K" is G or T, "H" is A or C or T, "D" is A, T or G, "I" is inosine, and "N" is any nucleotide. While nucleotide sequences herein can be represented in DNA sequence (containing T), the corresponding RNA sequence (i.e., with U in place of T) can be readily determined by one of skill in the art when RNA is referred to.
[0040] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" can also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.
[0041] "Expression construct" as used herein refers to a vector, such as a recombinant vector, suitable for expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to transcription of the nucleotide sequence (e.g., to produce mRNA or functional RNA) and / or translation of the RNA into a precursor or mature protein.
[0042] An "expression construct" of the application can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, can be a translatable RNA (such as an mRNA), for example an in vitro transcribed RNA.
[0043] An "expression construct" of the application can comprise regulatory sequences and nucleotide sequences of interest of different origin, or regulatory sequences and nucleotide sequences of interest of the same origin but arranged in a manner different from that which normally occurs in nature.
[0044] A "promoter" is a nucleic acid fragment that is capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the application, a promoter is a promoter that is capable of controlling the transcription of a gene in a cell, whether or not it is derived from that cell. A promoter can be a constitutive promoter or a tissue-specific promoter or a developmentally-regulated promoter or an inducible promoter.
[0045] Examples of promoters include, but are not limited to, a pol I, pol II, or pol III promoter. When used in plants, the promoter can be a cauliflower mosaic virus 35S promoter, a maize Ubi-1 promoter, a wheat U6 promoter, a rice U3 promoter, a maize U3 promoter, a rice actin promoter.
[0046] To "introduce" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, an RNA, etc.) or a protein into an organism means to transform a cell of the organism with the nucleic acid or protein such that the nucleic acid or protein is able to function in the cell. As used herein, "transform" includes both stable transformation and transient transformation. "Stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome resulting in stable inheritance of the foreign gene. Once stably transformed, the foreign nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof. "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell to perform a function without stable inheritance of the foreign gene. In transient transformation, the foreign nucleic acid sequence is not integrated into the genome.
[0047] A "trait" refers to a physiological, morphological, biochemical, or physical characteristic of a cell or organism.
[0048] "Agronomic traits" refer in particular to measurable index parameters of crop plants, including but not limited to: leaf green, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, total plant nitrogen content, fruit nitrogen content, seed nitrogen content, nitrogen content of vegetative plant tissue, total plant free amino acid content, fruit free amino acid content, seed free amino acid content, free amino acid content of vegetative plant tissue, total plant protein content, fruit protein content, seed protein content, protein content of vegetative plant tissue, herbicide resistance, drought resistance, nitrogen uptake, root lodging, harvest index, stalk lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt resistance, and tiller number.
[0049] II. Genome editing system for site-directed modification, e.g. site-directed insertion of exogenous nucleotide sequences, in a genome of an organism
[0050] In one aspect, the present application relates to a genome editing system for site-directed modification, e.g. site-directed insertion of exogenous nucleotide sequences, in a genome of an organism, comprising:
[0051] i) a) a CRISPR nuclease and / or an expression construct containing a nucleotide sequence encoding said CRISPR nuclease, and a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding said reverse transcriptase, or
[0052] b) a prime editing fusion protein and / or an expression construct containing a nucleotide sequence encoding said prime editing fusion protein, wherein said prime editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase;
[0053] ii) a first pegRNA and / or an expression construct containing a nucleotide sequence encoding said first pegRNA, and
[0054] iii) a second pegRNA and / or an expression construct containing a nucleotide sequence encoding said second pegRNA,
[0055] wherein said first pegRNA comprises from 5' to 3' direction a first guide sequence, a first scaffold sequence, a first reverse transcription template (RT) sequence, and a first primer binding site (PBS) sequence,
[0056] wherein said second pegRNA comprises from 5' to 3' direction a second guide sequence, a first scaffold sequence, a second reverse transcription template (RT) sequence, and a second primer binding site (PBS) sequence,
[0057] wherein the first pegRNA targets a first target sequence on the sense strand of the organism’s genomic DNA and the second pegRNA targets a second target sequence on the anti-sense strand of the organism’s genomic DNA. In some embodiments, the organism is a plant.
[0058] As used herein, a “target sequence” refers to a sequence of about 20 nucleotides in length in a genome characterized by a PAM (protospacer adjacent motif) sequence flanking either the 5’ or 3’ end. Generally, a PAM is necessary for the complex of a CRISPR nuclease or variant thereof and a guide RNA to recognize a target sequence. For example, for Cas9 nucleases and variants thereof, the target sequence is immediately adjacent to the PAM at the 3’ end, e.g., 5’-NGG-3’. Based on the presence of a PAM, one of skill in the art can readily determine target sequences in a genome that can be targeted. Depending on the location of the PAM, the target sequence can be on either strand of the genomic DNA molecule, the strand on which the target sequence is located is referred to as the target strand. For Cas9 or derivatives thereof, e.g., Cas9 nickases, the target sequence is preferably 20 nucleotides. Depending on different CRISPR nucleases or different variants thereof, the PAM sequence can vary.
[0059] In some embodiments, the pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a target sequence in the genome, resulting in a nick on the target strand (e.g., within the target sequence).
[0060] In some embodiments, the PAMs of the first target sequence and the second target sequence are separated by about 1 to about 300 bp, e.g., 10 bp to about 100 bp, e.g., about 20 bp to about 60 bp. In some embodiments, the PAMs of the first target sequence and the second target sequence can be separated by about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 100 bp, about 150 bp, about 300 bp.
[0061] In some embodiments, the CRISPR nuclease is a Cas9 nuclease, e.g., SpCas9 derived from S. pyogenes. An exemplary wild-type SpCas9 comprises the amino acid sequence set forth in SEQ ID NO: 1.
[0062] In some embodiments, the CRISPR nuclease is a CRISPR nickase. The CRISPR nickase in the fusion protein is capable of forming a nick within the target sequence on the target strand (the strand on which the target sequence is located) of the genomic DNA. In some embodiments, the CRISPR nickase is a Cas9 nickase.
[0063] In some embodiments, the Cas9 nickase is derived from SpCas9 of S. pyogenes and comprises at least the amino acid substitution H840A relative to wild-type SpCas9. In some embodiments, the Cas9 nickase comprises the amino acid sequence set forth in SEQ ID NO: 2. In some embodiments, the Cas9 nickase in the fusion protein is capable of forming a nick between the -3 position nucleotide of the PAM (the first nucleotide 5' of the PAM sequence is the +1 position) and the -4 position nucleotide of the PAM of a target sequence.
[0064] In some embodiments, the Cas9 nuclease such as a nickase is a Cas9 nuclease or nickase variant capable of recognizing an altered PAM sequence. There are many Cas9 nickase variants capable of recognizing an altered PAM sequence known in the art. In some embodiments, the Cas9 nuclease such as a nickase is a Cas9 variant recognizing a PAM sequence 5'-NG-3'. In some embodiments, the Cas9 nickase variant recognizing a PAM sequence 5'-NG-3' comprises the following amino acid substitutions H840A, D1135L, S1136W, G1218K, E1219Q, R1335Q, T1337R relative to wild-type Cas9, wherein the amino acid numbering is with reference to SEQ ID NO: 1. In some embodiments, the Cas9 nickase variant (SpG-Cas9 nickase) comprises the amino acid sequence set forth in SEQ ID NO: 42. In some embodiments, the Cas9 nickase variant recognizing a PAM sequence 5'-NG-3' comprises the following amino acid substitutions H840A, A61R, L1111R, D1135L, S1136W, G1218K, E1219Q, N1317R, A1322R, R1333P, R1335Q, T1337R relative to wild-type Cas9, wherein the amino acid numbering is with reference to SEQ ID NO: 1. In some embodiments, the Cas9 nickase variant (SpRY-Cas9 nickase) comprises the amino acid sequence set forth in SEQ ID NO: 43.
[0065] The nick formed by the Cas9 nuclease such as a nickase of the present application can result in the target strand forming a free single strand with a 3' end (3' free single strand) and a free single strand with a 5' end (5' free single strand).
[0066] In some embodiments, the CRISPR nuclease such as a Cas9 nickase and the reverse transcriptase in the prime editing fusion protein are connected by a linker.
[0067] In some embodiments, the reverse transcriptase in the present application can be derived from different sources. In some embodiments, the reverse transcriptase is a reverse transcriptase derived from a virus. For example, in some embodiments, the reverse transcriptase is an M-MLV reverse transcriptase or a functional variant thereof. An exemplary wild-type M-MLV reverse transcriptase sequence is set forth in SEQ ID NO: 3.
[0068] In some embodiments, the reverse transcriptase, e.g. M-MLV reverse transcriptase or a functional variant thereof
[0069] (a) comprises a mutation at position 155, 156, 200 and / or 524, e.g. comprises a mutation selected from any one of F155Y, F155V, F156Y, D524N, N200C or a combination thereof, the amino acid positions referring to SEQ ID NO: 3;
[0070] (b) the connection sequence is deleted; and / or
[0071] (c) the RNase H domain is mutated or deleted.
[0072] In some preferred embodiments, the reverse transcriptase, e.g. M-MLV reverse transcriptase or a functional variant thereof comprises a mutation selected from D524N, the amino acid position referring to SEQ ID NO: 3.
[0073] In some preferred embodiments, the RNase H domain of the reverse transcriptase, e.g. M-MLV reverse transcriptase or a functional variant thereof is deleted.
[0074] In some embodiments, the connection sequence comprises an amino acid sequence as set forth in SEQ ID NO: 4.
[0075] In some embodiments, the RNase H domain comprises an amino acid sequence as set forth in SEQ ID NO: 5.
[0076] In some embodiments, the reverse transcriptase, e.g. M-MLV reverse transcriptase or a functional variant thereof comprises a sequence as set forth in any one of SEQ ID NOs: 9-15, preferably an amino acid sequence as set forth in SEQ ID NO: 14.
[0077] In some embodiments, the reverse transcriptase, e.g. M-MLV reverse transcriptase or a functional variant thereof is fused at the N- or C-terminus directly or via a linker to a nucleocapsid protein (NC), a protease (PR) or an integrase (IN), e.g. from M-MLV.
[0078] In some embodiments, the nucleocapsid protein (NC) comprises an amino acid sequence as set forth in SEQ ID NO: 6.
[0079] In some embodiments, the hydrolytic enzyme (PR) comprises an amino acid sequence as set forth in SEQ ID NO: 7.
[0080] In some embodiments, the integrase (IN) comprises an amino acid sequence as set forth in SEQ ID NO: 8.
[0081] In some preferred embodiments, the reverse transcriptase, e.g., M-MLV reverse transcriptase or a functional variant thereof, is fused at the N-terminus directly or via a linker to a nucleocapsid protein (NC).
[0082] In some preferred embodiments, the reverse transcriptase, e.g., M-MLV reverse transcriptase or a functional variant thereof, is fused at the C-terminus directly or via a linker to a nucleocapsid protein (NC).
[0083] In some embodiments, the reverse transcriptase can also be fused via a linker or directly to an RNA aptamer binding protein sequence, e.g., a MCP protein sequence. Thereby, the reverse transcriptase can be recruited to the CRISPR nuclease via the interaction of the RNA aptamer binding protein sequence, e.g., a MCP protein sequence, and one or more RNA aptamer sequences, e.g., MS2 sequences, present on the pegRNA. In this case, the CRISPR nuclease does not need to be fused to the reverse transcriptase. An exemplary MCP protein comprises the amino acid sequence of SEQ ID NO: 44.
[0084] As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids, which does not have a secondary structure above. For example, the linker can be a flexible linker, such as GGGGS, GS, GAP, (GGGGS)x3, GGS, and (GGS)x7, etc. For example, it can be the linker set forth in SEQ ID NO: 16.
[0085] In some embodiments, the CRISPR nuclease, e.g., CRISPR nickase, in the fusion protein is located at the N-terminus of the reverse transcriptase. In some embodiments, the CRISPR nuclease, e.g., CRISPR nickase, in the fusion protein is located at the C-terminus of the reverse transcriptase.
[0086] In some embodiments of the application, the CRISPR nuclease, reverse transcriptase, recombinase or fusion protein of the application can further comprise one or more nuclear localization sequences (NLS). Generally, the one or more NLS in the CRISPR nuclease, reverse transcriptase or fusion protein should be of sufficient strength to drive accumulation of the CRISPR nuclease, reverse transcriptase or fusion protein in the nucleus of a cell in an amount that enables it to perform its base editing function. Generally, the strength of nuclear localization activity is determined by the number, location, specific NLS(es) used, or a combination of these factors, of NLS in the CRISPR nuclease, reverse transcriptase or fusion protein.
[0087] In some preferred embodiments, the fusion protein comprises, in the N-terminal to C-terminal direction, the CRISPR nuclease such as a nickase, the nucleocapsid protein (NC) and the reverse transcriptase, connected by or without a linker. In some preferred embodiments, the fusion protein comprises, in the N-terminal to C-terminal direction, a nuclear localization sequence - the CRISPR nuclease such as a nickase - a linker - the nucleocapsid protein (NC) - a nuclear localization sequence - a linker - the reverse transcriptase - a nuclear localization sequence.
[0088] In some preferred embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 19 (ePPE).
[0089] In some embodiments, the fusion protein comprises a nuclease portion comprising the CRISPR nuclease such as a CRISPR nickase and one or more NLS, and a reverse transcriptase portion comprising an RNA aptamer binding protein sequence (e.g., a MCP protein sequence), the reverse transcriptase, one or more NLS and optionally the nucleocapsid protein (NC), wherein the nuclease portion and reverse transcriptase portion are connected by a self-cleaving peptide. When the fusion protein is translated in vivo, a separate nuclease portion polypeptide and reverse transcriptase portion polypeptide are formed, the reverse transcriptase portion recruited to the nuclease portion by the interaction of the RNA aptamer binding protein sequence (e.g., a MCP protein sequence) and one or more RNA aptamer sequences (e.g., MS2 sequences) present on the pegRNA. An exemplary MCP protein comprises the amino acid sequence of SEQ ID NO: 44.
[0090] In some embodiments, the pegRNA of the application further comprises one or more RNA aptamer sequences (e.g., MS2 sequences). Exemplary one or more MS2 sequences are set forth in SEQ ID NO: 45. In some embodiments, the one or more RNA aptamer sequences (e.g., MS2 sequences) are located at the 3’ end of the pegRNA. In some embodiments, the one or more RNA aptamer sequences (e.g., MS2 sequences) are located in the middle of the pegRNA, e.g., between the scaffold sequence and the RT sequence. The one or more RNA aptamer sequences (e.g., MS2 sequences) can be used to recruit a reverse transcriptase comprising a RNA aptamer binding protein sequence (e.g., MCP protein sequence) to the CRISPR nuclease-pegRNA complex.
[0091] The guide sequence (also referred to as seed sequence or spacer sequence) in the pegRNA of the application is arranged to have sufficient sequence identity (preferably 100% identity) to the target sequence so as to be able to bind to the complementary strand of the target sequence through base pairing, enabling sequence-specific targeting.
[0092] For example, the guide sequence in the first pegRNA can have sufficient sequence identity (preferably 100% identity) to the first target sequence, which, in complex with a CRISPR nuclease such as a nickase, results in a nick in the first target sequence; the guide sequence in the second pegRNA can have sufficient sequence identity (preferably 100% identity) to the second target sequence on the opposite strand, which, in complex with a CRISPR nuclease such as a nickase, results in a nick in the second target sequence, whereby both pegRNAs result in nicks on different strands of the genomic DNA.
[0093] A variety of scaffold sequences for gRNAs suitable for use in CRISPR nuclease (e.g., Cas9)-based genome editing are known in the art, which can be used in the pegRNAs of the application. In some particular embodiments, the scaffold sequence of the gRNA is set forth in SEQ ID NO: 17.
[0094] In some embodiments, the primer binding sequence is arranged to be complementary to at least a portion of the target sequence, preferably to be perfectly paired to at least a portion of the target sequence, preferably to be complementary to, preferably to be perfectly paired to, at least a portion of the 3' overhang single strand resulting from the nick in the DNA strand in which the target sequence is located. When the 3' overhang single strand of the strand is bound to the primer binding sequence by base pairing, the 3' overhang single strand can serve as a primer to perform reverse transcription on the reverse transcription template (RT) sequence immediately adjacent to the primer binding sequence as a template under the action of the reverse transcriptase in the fusion protein, to extend a DNA sequence corresponding to the reverse transcription template (RT) sequence.
[0095] The primer binding sequence depends on the length of the overhang single strand formed in the target sequence by the CRISPR nickase used, however, it should have a minimum length to ensure specific binding. In some embodiments, the primer binding sequence can have a length of 4-20 nucleotides, for example, a length of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides.
[0096] In some embodiments, the primer binding sequence is arranged to have a Tm (melting temperature) of no more than about 52°C. In some embodiments, the primer binding sequence has a Tm (melting temperature) of about 18°C-52°C, preferably about 24°C-36°C, more preferably about 28°C-32°C, more preferably about 30°C.
[0097] Methods for calculating the Tm of a nucleic acid sequence are well known in the art, for example, the Oligo Analysis Tool online analysis tool can be used. An exemplary calculation formula is Tm = N G:C *4 + N A:T *2, wherein N G:C is the number of G and C bases in the sequence, and N A:T is the number of A and T bases in the sequence. A suitable Tm can be obtained by selecting a suitable length of PBS. Alternatively, a PBS sequence with a suitable Tm can be obtained by selecting a suitable target sequence.
[0098] In some embodiments, the RT template sequence can be any sequence. Through the reverse transcription described above, its sequence information can be integrated into the DNA strand where the target sequence is located (i.e., the strand containing the PAM of the target sequence), and through the DNA repair of the cell, a DNA double strand containing the sequence information of the RT template sequence is formed. In some embodiments, the RT template sequence contains a desired modification. For example, the desired modification includes substitution, deletion, and / or addition of one or more nucleotides. In some embodiments, the RT template sequence is configured to correspond to (e.g., complementary to at least a portion of) the sequence downstream of the target sequence cut, but contains a desired modification. The desired modification includes substitution, deletion, and / or addition of one or more nucleotides.
[0099] In some embodiments, the two pegRNAs are configured to introduce the same desired modification. For example, one pegRNA is configured to introduce a substitution of A to G in the sense strand, and the other pegRNA is configured to correspondingly introduce a substitution of T to C in the corresponding position of the anti-sense strand. For another example, one pegRNA is configured to introduce a deletion of two nucleotides in the sense strand, and the other pegRNA is configured to correspondingly introduce a deletion of two nucleotides in the corresponding position of the anti-sense strand. Other types of modifications can be similarly introduced. The pegRNAs targeting two different strands respectively can be configured to introduce the same desired modification by designing appropriate RT template sequences.
[0100] In some embodiments, the RT sequence is configured to generate, upon reverse transcription using it as a template, an exogenous nucleotide sequence or a portion thereof to be inserted into the genome, or generate a complement of the exogenous nucleotide sequence or a portion thereof to be inserted into the genome of an organism such as a plant. In some embodiments, the RT sequence does not contain the genomic sequence near the target sequence or the complement of the genomic sequence near the target sequence. In some embodiments, the RT sequence does not contain sequence information other than the exogenous nucleotide sequence to be inserted.
[0101] In some embodiments, the first RT sequence and the second RT sequence are used to insert a first exogenous nucleotide sequence, for example, between the first target sequence and the second target sequence (e.g., between the cut of the first target sequence and the cut of the second target sequence).
[0102] In some embodiments, the first RT sequence of the first pegRNA is configured to generate, upon reverse transcription using it as a template, a first fragment of a first exogenous nucleotide sequence to be inserted into the genome; and the second RT sequence of the second pegRNA is configured to generate, upon reverse transcription using it as a template, a complement of a second fragment of the first exogenous nucleotide sequence to be inserted into the genome.
[0103] In some embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted at least partially overlap. In some embodiments, the first and second fragments overlap by at least about 10 bp to about 50 bp, such as at least about 10 bp, about 15 bp, about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp. In some embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted completely overlap.
[0104] In some embodiments, the first exogenous nucleotide sequence to be inserted is about 1 bp to about 700 bp in length, such as about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp, about 600 bp, about 700 bp, or any value therebetween.
[0105] In some embodiments, the pegRNA further comprises a tevopre sequence at the 3’ end of the PBS. The design of the tevopre sequence can refer to the literature James W. Nelson et al., Engineered pegRNAs improve prime editing efficiency. 2022, Nature Biotech. volume 40, pages 402-410. An exemplary tevopre sequence is shown in SEQ ID NO: 20.
[0106] In some embodiments, the pegRNA further comprises a polyA sequence at the 3’ end. The polyA sequence, for example, comprises about 10-30 consecutive adenosine (A) nucleotides.
[0107] In some embodiments, the pegRNA comprises, from 5’ to 3’ direction, a guide sequence, a scaffold sequence, a reverse transcription template (RT) sequence, a primer binding site (PBS) sequence, a tevopre sequence, and a polyA sequence.
[0108] In some embodiments, the pegRNA can be precisely processed using a self-processing system for its sequence. In some specific embodiments, the 5’ end of the pegRNA is linked to a first ribozyme or tRNA designed to cleave the fusion at the 5’ end of the pegRNA; and / or the 3’ end of the pegRNA is linked to a second ribozyme or tRNA designed to cleave the fusion at the 3’ end of the pegRNA. The design of the first or second ribozyme or tRNA is within the capability of one skilled in the art. For example, see Gao et al., JIPB, Apr, 2014; Vol 56, Issue 4, 343-349. Methods for precisely processing gRNA can be found, for example, in WO 2018 / 149418.
[0109] In some embodiments, the first pegRNA and the second pegRNA are driven for transcription by different promoters. For example, the first pegRNA is driven for expression by OsU3 promoter, and the second pegRNA is driven for expression by TaU3 promoter.
[0110] In some embodiments, the pegRNA is driven for transcription by a type II promoter, i.e. in an expression construct comprising a nucleotide sequence encoding the pegRNA, the coding nucleotide sequence of the pegRNA is operably linked to a type II promoter. In some specific embodiments, the type II promoter is a GS promoter. An exemplary sequence of a GS promoter is set forth in SEQ ID NO: 21.
[0111] In some embodiments, wherein the first target sequence, the second target sequence, and / or the desired modification, such as the insertion of the first exogenous nucleotide sequence, is associated with a trait, such as an agronomic trait, of an organism, such as a plant, whereby the desired modification, such as the insertion of the first exogenous nucleotide sequence, results in the organism, such as the plant, having an altered, preferably improved, trait, e.g. agronomic trait, relative to a wild type organism, such as a plant.
[0112] In some embodiments, the first exogenous nucleotide sequence comprises one or more recombinase recognition sites (RS).
[0113] In some embodiments, the recombinase is a recombinase of the tyrosine recombinase family or a recombinase of the serine recombinase family, preferably a recombinase of the tyrosine recombinase family. Exemplary tyrosine recombinases include, but are not limited to, bacteriophage lambda integrase, P1 phage Cre recombinase (cyclization recombinase), yeast FLP recombinase (flippase recombinase). Exemplary serine recombinases include, but are not limited to, Tn3 transposase, Salmonella recombinase Hin, Streptomyces phage ΦC31 integrase, and Mycobacterium phage Bxbl integrase. Different recombinases and their corresponding recombinase recognition sites (RS) are known in the art, and a person skilled in the art can select as needed.
[0114] In some embodiments, the recombinase is a Dre recombinase. Exemplary Dre recombinases comprise the amino acid sequence of SEQ ID NO: 56. Correspondingly, the one or more recombinase recognition sites (RS) include, but are not limited to, rox (SEQ ID NO: 57, 58).
[0115] In some embodiments, the recombinase is a ΦC31 integrase. Exemplary ΦC31 integrases comprise the amino acid sequence of SEQ ID NO: 22. Correspondingly, the one or more recombinase recognition sites (RS) include, but are not limited to, aTTP (SEQ ID NO: 38) and / or aTTB (SEQ ID NO: 39).
[0116] In some preferred embodiments, the recombinase is a Bxbl integrase. Exemplary Bxbl integrases comprise the amino acid sequence of SEQ ID NO: 23. Correspondingly, the one or more recombinase recognition sites (RS) include, but are not limited to, aGTP (SEQ ID NO: 40) and / or aGTB (SEQ ID NO: 41).
[0117] In some preferred embodiments, the recombinase is a Cre recombinase. Exemplary Cre recombinases comprise the amino acid sequence of SEQ ID NO: 24. Correspondingly, the one or more recombinase recognition sites (RS) include, but are not limited to, loxP (SEQ ID NO: 26), Lox2272 (SEQ ID NO: 29), Lox71 (SEQ ID NO: 27), Lox66 (SEQ ID NO: 28), or variants thereof, and any combination thereof.
[0118] In some preferred embodiments, the recombinase is an FLP recombinase. An exemplary FLP recombinase comprises the amino acid sequence of SEQ ID NO:25. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, FRT1 (SEQ ID NO:30), FRT6 (SEQ ID NO:31), or variants thereof, and any combination thereof. In some embodiments, the one or more recombinase recognition sites (RS) are variants of FRT1, for example, comprising the sequence described in one of SEQ ID NO:32-37.
[0119] In some embodiments, the recombinase is a B2 recombinase. An exemplary B2 recombinase comprises the amino acid sequence of SEQ ID NO:50. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, the nucleotide sequence shown in SEQ ID NO:53.
[0120] In some embodiments, the recombinase is a KD recombinase. An exemplary KD recombinase comprises the amino acid sequence of SEQ ID NO:51. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, the nucleotide sequence shown in SEQ ID NO:54.
[0121] In some embodiments, the recombinase is a pSR1 recombinase. An exemplary pSR1 recombinase comprises the amino acid sequence of SEQ ID NO:52. Accordingly, the one or more recombinase recognition sites (RS) include, but are not limited to, the nucleotide sequence shown in SEQ ID NO:55.
[0122] Based on one or more recombinase recognition sites (RS) in a first exogenous nucleotide sequence inserted into the genome, by providing a donor containing RS and a second exogenous nucleotide sequence, the second exogenous nucleotide sequence can be inserted into the genome of an organism such as a plant via recombination using the corresponding recombinase. The recombinase can be a separately expressed recombinase or included in the guide editing fusion protein. Those skilled in the art can select a suitable combination of RS located at the first exogenous polynucleotide already inserted into the genome and RS located at the donor to insert the second exogenous nucleotide sequence into the genome via recombination.
[0123] Therefore, in some implementations, the genome editing system further includes:
[0124] iv) a recombinase and / or an expression construct containing a nucleotide sequence encoding the recombinase, and
[0125] v) A donor construct containing one or more recombinase recognition sites (RS) and a second exogenous polynucleotide sequence to be inserted into the plant genome.
[0126] In some preferred embodiments, the recombinase is comprised in the prime editing fusion protein. In some embodiments, the recombinase is located N-terminal to the CRISPR nuclease and reverse transcriptase in the prime editing fusion protein. In some embodiments, the recombinase is located C-terminal to the CRISPR nuclease and reverse transcriptase in the prime editing fusion protein.
[0127] The second exogenous polynucleotide sequence can be of any length. The second exogenous polynucleotide sequence can be 1 bp to about 10 kb or longer. Preferably, the second exogenous polynucleotide is a long fragment, e.g., at least 300 bp, at least 500 bp, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb or longer. In some embodiments, the second exogenous polynucleotide can be a full-length gene.
[0128] In some embodiments, wherein the second exogenous nucleotide sequence is associated with a trait of an organism, e.g., a plant trait, e.g., an agronomic trait, whereby insertion of the second exogenous nucleotide sequence results in the organism, e.g., the plant, having an altered, preferably improved, trait, e.g., an agronomic trait, relative to a wild-type organism, e.g., a wild-type plant.
[0129] The different components of the genome editing system of the application, e.g., the coding sequences for the CRISPR nuclease, the reverse transcriptase, the prime editing fusion protein, the pegRNA and / or the recombinase, and the second exogenous polynucleotide sequence, can be located in the same construct, or in different constructs, respectively, in different combinations.
[0130] The organism that can be subjected to site-directed modification, e.g., site-directed insertion of an exogenous nucleotide sequence, by the genome editing system of the application can be a non-human animal, a human, or a plant, preferably a plant. Suitable plants include monocotyledonous and dicotyledonous plants, e.g., the plant is a crop plant, including but not limited to wheat, rice, maize, soybean, sunflower, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.
[0131] In order to obtain efficient expression in an organism, e.g., a plant, in some embodiments of the application, the nucleotide sequence encoding the fusion protein is codon-optimized for the organism, e.g., the plant species, whose genome is to be modified.
[0132] Codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) with a codon that is more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. Different species exhibit particular biases toward certain codons for a particular amino acid. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which is believed to be dependent on the properties of the codons being translated and the availability of the particular transfer RNA (tRNA) molecules. The predominance of selected tRNAs within a cell generally reflects the frequency with which codons are used in protein synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the Codon Usage Database available at www.kazusa.orjp / codon / , and these tables can be adapted in different ways. See, Nakamura Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).
[0133] III. Methods of making site-directed modifications in a plant genome, e.g., site-directed insertion of foreign nucleotide sequences
[0134] In another aspect, the application provides a method of making a genetically modified plant comprising a site-directed modification, the method comprising introducing into at least one of the plants a genome editing system of the application. The site-directed modification comprises substitution, deletion, and / or addition of one or more nucleotides. For example, the site-directed modification comprises site-directed insertion of a foreign nucleotide sequence.
[0135] In another aspect, the application provides a method of making a genetically modified plant comprising a site-directed modification, the method comprising introducing into at least one of the plants a genome editing system of the application. The site-directed modification comprises substitution, deletion, and / or addition of one or more nucleotides. For example, the site-directed modification comprises site-directed insertion of a foreign nucleotide sequence.
[0136] In some embodiments, the method further comprises screening the at least one plant for a desired site-directed modification, e.g., site-directed insertion of a foreign nucleotide sequence.
[0137] In the methods of the present application, the genome editing system can be introduced into the plant by various methods well known to the person skilled in the art. Methods that can be used to introduce the genome editing system of the present application into a plant include, but are not limited to, biolistics, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway and ovary injection. Preferably, the genome editing system is introduced into the plant by transient transformation.
[0138] In some embodiments, the components of the genome editing system are introduced into the plant simultaneously. In some embodiments, the components of the genome editing system are introduced into the plant separately or sequentially.
[0139] In some embodiments, the method comprises the following steps:
[0140] 1 ) transforming components i) - iv) of the genome editing system into an isolated plant cell or tissue, obtaining a plant cell or tissue into which a first exogenous nucleotide sequence comprising one or more recognition sites (RS) of a recombinase is inserted;
[0141] 2) transforming component v) of the genome editing system into the plant cell or tissue obtained in step 1 ), thereby obtaining a plant cell or tissue comprising a second exogenous polynucleotide sequence inserted; and
[0142] 3) regenerating a whole plant from the plant cell or tissue obtained in step 2).
[0143] In some embodiments, wherein the exogenous nucleotide sequence is inserted into a safe harbor site in the plant genome, the safe harbor site is in the plant genome
[0144] 1 ) at least 5 kb away from a protein coding region;
[0145] 2) at least 30 kb away from a miRNA coding region;
[0146] 3) at least 20 kb away from a IncRNA coding region;
[0147] 4) at least 20 kb away from a tRNA coding region;
[0148] 5) at least 5 kb away from a promoter and / or enhancer;
[0149] 6) at least 20 kb away from a LTR repeat;
[0150] 7) at least 200 bp away from a non-LTR repeat; and
[0151] 8) at least 10 kb away from a centromere.
[0152] In some embodiments, wherein the plant is rice, and the safe harbor site is selected from the sites shown in Tables 1, 2.
[0153] In some embodiments, the introducing comprises transforming the genome editing system of the application into an isolated plant cell or tissue, and then regenerating the transformed plant cell or tissue into a whole plant. Preferably, no selection agent directed against the selection gene carried on the expression vector is used during tissue culture.
[0154] In other embodiments, the genome editing system of the application can be transformed into a specific location on a whole plant, such as a leaf, a stem tip, a pollen tube, an ear shoot, or a hypocotyl. This is particularly suitable for transformation of plants that are difficult to regenerate via tissue culture.
[0155] In some embodiments of the application, an in vitro expressed protein and / or an in vitro transcribed RNA molecule (e.g., the expression construct is an in vitro transcribed RNA molecule) and / or a donor DNA molecule is directly transformed into the plant.
[0156] In some embodiments, the method further comprises treating (e.g., culturing) the plant cell, tissue, or whole plant into which the genome editing system has been introduced at an elevated temperature (relative to the temperature at which the culture is normally maintained, such as room temperature), such as 37°C.
[0157] In some embodiments of the application, wherein the site-directed modification, such as site-directed insertion, of an exogenous nucleotide sequence and / or the target sequence is associated with a plant trait, such as an agronomic trait, whereby the site-directed modification, such as site-directed insertion, results in the plant having an altered (preferably improved) trait, such as an agronomic trait, relative to a wild type plant.
[0158] In some embodiments, the method further comprises the step of screening the plants for the desired site-directed modification, such as site-directed insertion, and / or the desired trait, such as an agronomic trait.
[0159] In some embodiments of the application, the method further comprises obtaining progeny of the genetically modified plant. Preferably, the genetically modified plant or its progeny has the desired modification (such as site-directed exogenous polynucleotide insertion) and / or the desired trait, such as an agronomic trait.
[0160] In another aspect, the application also provides a genetically modified plant or its progeny or a part thereof, wherein the plant is obtained by the above-described method of the application. Preferably, the genetically modified plant or its progeny has the desired genetic modification (such as site-directed exogenous polynucleotide insertion) and / or the desired trait, such as an agronomic trait.
[0161] In another aspect, the present application also provides a method for breeding a plant, comprising crossing a first genetically modified plant obtained by the above-mentioned method of the present application with a second plant not containing the modification, thereby introducing the modification (e.g. site-directed exogenous polynucleotide insertion) into the second plant. Preferably, the first genetically modified plant and the second plant have desirable traits such as agronomic traits.
[0162] Plants that can be subjected to site-directed modification such as site-directed insertion of exogenous nucleotide sequences by the genome editing system of the present application include monocotyledonous and dicotyledonous plants, for example, the plants are crop plants including but not limited to wheat, rice, maize, soybean, sunflower, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava and potato.
[0163] In another aspect, the present application provides a method for producing a genetically modified plant comprising a site-directed inserted exogenous nucleotide sequence, the method comprising inserting an exogenous nucleotide sequence into a safe harbor site in a plant genome, the safe harbor site being at least
[0164] 1) at least 5 kb away from a protein-coding region;
[0165] 2) at least 30 kb away from a miRNA-coding region;
[0166] 3) at least 20 kb away from a IncRNA-coding region;
[0167] 4) at least 20 kb away from a tRNA-coding region;
[0168] 5) at least 5 kb away from a promoter and / or enhancer;
[0169] 6) at least 20 kb away from a LTR repeat;
[0170] 7) at least 200 bp away from a non-LTR repeat; and
[0171] 8) at least 10 kb away from a centromere.
[0172] In some embodiments, wherein the plant is rice, and the safe harbor site is selected from the sites shown in Tables 1, 2.
[0173] IV. A method for making a genetically modified plant comprising a site-directed inserted exogenous nucleotide sequence
[0174] In another aspect, the present application provides a method of site-specifically modifying a human or non-human animal genome, comprising introducing the genome editing system of the present application into at least one of the human or non-human animal cells. The site-specific modification includes substitution, deletion and / or addition of one or more nucleotides. For example, the site-specific modification includes site-specific insertion of an exogenous nucleotide sequence.
[0175] In another aspect, the present application provides a use of the genome editing system of the present application for site-specifically modifying a human or non-human animal genome in vivo or ex vivo, which can achieve deletion, addition, up-regulation, down-regulation, inactivation, activation or mutation correction of a disease-associated gene, thereby achieving prevention and / or treatment of a disease. For example, the target nucleic acid region of the present application can be located within a protein-coding region of a disease-associated gene, or for example, can be located in a gene expression regulatory region such as a promoter region or an enhancer region, thereby achieving modification of the function of the disease-associated gene or modification of the expression of the disease-associated gene. Thus, the modification of a disease-associated gene described herein includes modification of the disease-associated gene itself (e.g., a protein-coding region), and also includes modification of the expression regulatory region (e.g., a promoter, an enhancer, an intron, etc.) thereof.
[0176] In another aspect, the present application provides a method of producing a genetically modified human or non-human animal somatic cell comprising a site-specific modification, the method comprising introducing the genome editing system of the present application into at least one of the human or animal somatic cells. The site-specific modification includes substitution, deletion and / or addition of one or more nucleotides. For example, the site-specific modification includes site-specific insertion of an exogenous nucleotide sequence.
[0177] Thus, the present application also provides a method of treating a disease in a subject in need thereof, comprising delivering to the subject an effective amount of the genome editing system of the present application to modify a gene associated with the disease. The present application also provides use of the genome editing system of the present application in the manufacture of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the genome editing system is used to modify a gene associated with the disease. The present application also provides a pharmaceutical composition for treating a disease in a subject in need thereof, comprising the genome editing system of the present application, and optionally a pharmaceutically acceptable carrier, wherein the genome editing system is used to modify a gene associated with the disease. In some embodiments, the subject is a human.
[0178] V. Kits
[0179] The present application also includes kits for use in the methods of the present application, which kits comprise at least components of the genome editing system of the present application. The kits can also comprise reagents for introducing the genome editing system into an organism or organism cell. The kits generally include a label indicating intended use and / or method of use of the contents of the kit. The term label includes any written or recorded material that is provided with the kit or is otherwise available to one of ordinary skill in the art to which the kit relates to to practice the present application. Examples
[0180] Example 1. Design of novel genome editing system
[0181] 1.1. Prime editing (PE) system screening
[0182] Prime editing (PE) is a precise genome editing technology that can generate base changes and short DNA insertions and deletions without forming DSBs, and is widely used across species such as humans, mice, rice, wheat, corn, etc. To develop a novel genome editing system, this embodiment first screened the efficiency of the reported PE system in using the double pegRNA strategy to edit endogenous target points. Five PE system constructs were compared, namely PPE, Art-PPE (5' end of Cas9 fused with mouse exonuclease Artemis), PPE-NCV1, ePPE (Zong, Y., Liu, Y., Xue, C. et al. An engineered prime editor with enhanced editing efficiency in plants. Nat Biotechnol 40, 1394-1402 (2022).) and ePPE-wtCas9 (replace H840A-Cas9 in ePPE with wtCas9) constructs, in the efficiency of using the double pegRNA strategy to insert Lox66 (34 bp in length) and / or FRT1 (48 bp in length) two recombinase recognition sites (RS) at endogenous target points. The schematic diagram of the construction strategy of the vector is as follows Figure 1A .
[0183] Rice protoplasts were selected as model cells. The above constructs were transformed into rice protoplasts by PEG transformation method, and the efficiency of five prime editing system constructs in inserting Lox66 or FRT1 at rice protoplast endogenous sites using 5 pairs of pegRNA was tested, and the second generation sequencing results are as follows Figure 1B .
[0184] The results show that when using the double pegRNA strategy for site-specific insertion, the efficiency of using ePPE is the highest, which can improve the precise insertion efficiency by 10-50 times compared with PPE.
[0185] 1.2. Guide RNA screening
[0186] Further comparison of the general pegRNA and the reported epegRNA containing tevoPre that can improve PE efficiency (Nelson, J. W., Randolph, P. B., Shen, S. P. et al. Engineered pegRNAs improve prime editing efficiency. Nat Biotechnol 40, 402-410 (2022).) in different prime editing systems. Four combinations were tested respectively, including PPE+pegRNA, ePPE+pegRNA, PPE+epegRNA, ePPE+epegRNA. Vector construction is shown as Figure 2A .
[0187] Similarly, the above four combinations were tested for the efficiency of inserting RS at 8 pairs of pegRNA / epegRNA using rice protoplasts, and the results of second-generation sequencing are shown as Figure 2B .
[0188] The results show that the combination of ePPE+epegRNA (hereinafter referred to as "dual-ePPE") has the highest efficiency in the double pegRNA strategy mediated site-specific insertion, and the highest efficiency of some sites can reach more than 50%, which is more than 100 times higher than the general PPE+pegRNA combination, and has higher efficiency in most low-efficiency target sites. It is worth noting that while the editing tool improves the editing efficiency of the target site, the probability of non-precise editing or insertion or deletion at other sites also increases. While dual-ePPE has a significant improvement in the precise editing efficiency of the target site compared to other combinations, it does not significantly change the insertion or deletion efficiency at other sites.
[0189] 1.3. dual-ePPE insertion
[0190] Further evaluation of the relationship between insertion length (30bp-100bp) and distance between two pegRNAs (PAM distance 20bp-80bp) and the effect of overlap length between two RTs (10bp-50bp) on insertion efficiency using dual-ePPE. Rice protoplasts were used for testing, and the results of second-generation sequencing are shown as Figure 3 .
[0191] The results show that there is no obvious linear relationship between the insertion length and the distance between pegRNA, and the efficiency is higher when the insertion length is greater than the distance between pegRNA, and the overlap length between the two RTs is between 10bp-50bp. The above results show that the ePPE+epegRNA system of the application can meet the efficient and site-specific insertion of Flag, Tag and other tag sequences.
[0192] Example 2. Optimization of dual-ePPE system
[0193] In order to further verify and optimize the effect of the dual-ePPE system of the application in different use environments, thereby obtaining the preferred technical scheme, the possible improvements of each component of the dual-ePPE system are verified and analyzed in this embodiment.
[0194] 2.1. CRISPR system effector protein
[0195] In this embodiment, SpG-Cas9 with recognition sequence NGN PAM and SpRY-Cas9 variant almost not limited by PAM sequence (Christie KA, Guo JA, Silverstein RA, Doll RM, Mabuchi M, Stutzman HE, Lin J, Ma L, Walton RT, Pinello L, Robb GB, Kleinstiver BP. Precise DNA cleavage using CRISPR-SpRYgests. Nat Biotechnol. 2023 Mar; 41(3): 409-416.) are designed into dual-ePPE, and the insertion efficiency of PAM containing NGN is evaluated, in order to expand the targeting range of dual-ePPE. The vector construction is as shown in Figure 4A .
[0196] The efficiency of site-specific insertion under the combination of NGA, NGC, and NGT PAM was tested using rice protoplast. The second-generation sequencing results are as shown in Figure 4B .
[0197] The results show that using SpG-ePPE and SpRY-ePPE has high efficiency for site-specific insertion under PAM as NGA, NGC, and NGT, thereby verifying that the dual-ePPE system can be suitable for various CRISPR system effector proteins and can effectively exert its function. The results show that the dual-ePPE of the application can effectively realize the insertion of RS sequence in plants.
[0198] 2.2. RT sequence synonymous mutation
[0199] It has been reported that introducing synonymous mutations (SM) on the RT sequence can improve the efficiency of recombineering editing (Xu, W., Yang, Y., Yang, B. et al. A design optimized prime editor with expanded scope and capability in plants. Nat. Plants 8, 45-52 (2022)). This example tests the editing efficiency of the system in the processing of SM on the RT. The vector construction is as shown in Figure 4A .
[0200] The efficiency of point mutation using two RTs at four target sites was tested using rice protoplasts. The results of second-generation sequencing are as shown in Figure 4C . The results show that when the RT sequence has uniform mismatches with the genomic sequence, the efficiency of point mutation can be greatly improved (4-20 times).
[0201] 2.3. Promoters driving epegRNA expression
[0202] Further research was conducted on the efficiency of inserting longer fragments (150bp-300bp) using the above system, and the editing efficiency of U3 promoters and composite type II promoters (pGS promoters) expressing guide RNA was tested. The pGS-epegRNA vector construction is as shown in Figure 5A .
[0203] The efficiency of using U3 promoters and pGS promoters to insert different lengths of fragments at specific sites was compared in rice protoplasts. The results of ddPCR detection are as shown in Figure 5B , C. The results show that there is no significant difference between U3 promoters and pGS promoters in small fragment insertion Figure 5B . When the insertion length is more than 150bp, the efficiency of pGS promoters driving epegRNA is higher than that of U3 promoters. Using pGS promoters to express epegRNA can improve the efficiency of large fragment insertion by 2-5 times, and when the length of the inserted fragment reaches 700bp, precise insertion can still be achieved.
[0204] 2.4. Influence of MS2-MCP and temperature treatment on editing efficiency
[0205] Further use of the MS2-MCP recruiting MLV system and 37°C temperature treatment to further improve the efficiency of long fragment insertion, the vector construction is as shown in Figure 6A .
[0206] The efficiency of large fragment site-directed insertion using the above two recruitment forms was tested using rice protoplasts, and the efficiency of 37°C treatment (TT, normal culture for 12 h→ 37°C culture for 12 h→ normal culture for 24 h) was also tested. The ddPCR results are shown in Figure 6B .
[0207] The results show that using 37°C temperature treatment can improve the large fragment insertion efficiency by about 1.2-5 times Figure 6B , and using the MS2-MCP system to recruit MLV can improve the large fragment insertion efficiency by about 2-4 times Figure 6C .
[0208] Example 3. Large fragment DNA insertion in plants using the PrimeROOT system
[0209] In this embodiment, dual-ePPE is combined with recombinase as a Prime editing-mediated Recombination Of Opportune Targets (hereinafter referred to as PrimeROOT) mediated precision targeted DNA recombination system, and its DNA fragment insertion function in plants is verified.
[0210] 3.1. Construction of fluorescent reporter system
[0211] To verify the DNA recombination ability of various recombinases in plant base editing, the inventors first constructed a fluorescent reporter system to characterize the DNA recombination efficiency of commonly used site-specific recombinases in rice protoplasts. The reporter system divides GFP into N-terminal (GFP-N) and C-terminal (GFP-C) two domains, which are encoded on two separate plasmids Figure 7A ), each of which carries a recombinase site. The plasmid construction is shown in Figure 7A . After expression of the recombinase and recombination, GFP-N and GFP-C are linked by an intron linker, thereby allowing GFP to be expressed in protoplasts. Further, GFP fluorescence can be observed by fluorescence microscopy and detected by flow cytometry to characterize the activity of recombinases in protoplasts.
[0212] 3.2. Construction of PrimeROOT for detection
[0213] The inventors constructed independent fluorescent reporter systems for 6 different tyrosine recombinases and 2 serine recombinases (all recombinases were codon-optimized and can be expressed in rice). The results of GFP fluorescence microscopy observation Figure 7B ) and flow cytometry retrieval Figure 7C) The Cre and FLP recombinase system can produce the strongest fluorescence, and can be used as the best recombinase system for verifying and optimizing the effectiveness of the technical solutions of the present application.
[0214] In another set of parallel experiments, the inventors constructed a fluorescence reporter system for the Cre / Lox system of the tyrosine recombinase family, the FLP / FRT system, and the ФC31 and Bxb1 recombinases in the serine family, and the vector construction is as shown in Figure 7D .
[0215] The above reporter system was transformed into rice protoplasts, and fluorescence microscopy observation and flow cytometry detection were performed, and the results are shown in Figure 7E .
[0216] The results show that the DNA integration ability of the Cre / Lox system and the FLP / FRT system is stronger. Therefore, they are combined with the above-mentioned site-specific insertion system, and all components (ePPE, two epegRNAs, recombinase, and genes to be inserted with recombination sites) are transferred into rice cells by "one-step method" to realize the site-specific insertion of large fragments at the gene level, as shown in Figure 7F . The "one-step method" composition containing dual-ePPE, recombinase, and genes to be inserted with recombination sites is named PrimeROOT.v1, and according to the recombinase, it is named PrimeROOT.v1-Cre and PrimeROOT.v1-FLP for the Cre / Lox system or the FLP / FRT system, respectively.
[0217] 3.3. Verification of large fragment insertion ability of PrimeROOT.v1
[0218] To verify the insertion ability of PrimeROOT.v1 for large fragment DNA molecules, the inventors tested the integration efficiency of PrimeROOT.v1-Cre and PrimeROOT.v1-FLP for GFP (720kp) at four endogenous sites in rice protoplasts by ddPCR, and the experimental results are shown in Figure 8 . The results show that both PrimeROOTs achieve precise and targeted large fragment insertion at the four sites.
[0219] 3.4. Optimization of recombinase system
[0220] Because of the short repeat sequence in FRT1, it has been reported that some FRT1 mutants have a promoting effect on the efficiency of FLP recombinase (Bruckner, R. C. & Cox, M. M. Specific Contacts between the Flp Protein of the Yeast 2-Micron Plasmid and Its Recombination Site. Journal of Biological Chemistry 261, 1798-1807 (1986).; Senecoff, J. F., Rossmeissl, P. J. & Cox, M. M. DNA recognition by the FLP recombinase of the yeast 2mu plasmid. A mutational analysis of the FLP binding site. J Mol Biol 201, 405-421 (1988).). In order to further optimize the editing system to obtain a more preferred technical solution, the inventors designed a plurality of FRT1 mutants (F1m1, F1m2 and F1m3) and two truncated FRT1 (tFRT1) sequence mutant (tF1m2 and tF1m3). When using PrimeROOT for integration, the efficiency of one-step large fragment insertion on the endogenous target site was evaluated by the method of ddPCR, and the protoplast cells were made to emit light after inserting GFP into the endogenous gene of rice using one-step method. The ddPCR results are shown in Figure 9, and the combination of FRT1 mutants has higher mutation efficiency than the wild type.
[0221] 3.5. Optimization of PrimeROOT system
[0222] On the basis of PrimeROOT.v1, the inventors further optimized it to obtain a more preferred technical solution. In this technical solution, the inventors fused the ePPE of PrimeROOT composition species with the recombinase, and according to the fusion site, two structural schemes were created, and the example sequences are shown in Figure 10A :
[0223] Scheme 1 connects the recombinase to the N-terminus of the ePPE system through an SV40 NLS and a 32 amino acid flexible linker, named PrimeROOT.v2N; Scheme 2 connects the recombinase to the C-terminus of the ePPE system through the same way, named PrimeROOT.v2C. Fluorescence microscopy observation and flow cytometry detection results show that PrimeROOT.v2N and PrimeROOT.v2C systems have higher GFP insertion efficiency at four endogenous sites than PrimeROOT.v1 (Figure 10).
[0224] 3.6. Verification of large fragment insertion ability of PrimeROOT.v2
[0225] To verify the large fragment DNA insertion ability of PrimeROOT.v2, the inventors constructed a vector construct containing any one or combination of three genes (pigmR, OsMYB30 and OsHPPD), with donor lengths of 1.4 kb, 4.9 kb, 7.7 kb and 11.1 kb, respectively, and the vector construction is as follows Figure 11A . The inventors detected the insertion efficiency of the four donors at the four endogenous sites by ddPCR, and found that as the donor length gradually increased, precise and targeted large fragment insertion was achieved, and the editing efficiency did not decrease significantly Figure 11B ).
[0226] Example 4. Large fragment DNA insertion without double-strand breaks in maize using the PrimeROOT system
[0227] In addition to rice protoplasts, the inventors also evaluated the editing efficiency of dual-ePPE and its PrimeROOT in corn protoplasts.
[0228] The inventors first tested the precise RS insertion editing efficiency of dual-ePPE at six endogenous gene sites in corn protoplasts, and the experimental results showed that it could achieve an editing efficiency as high as 40% Figure 12A ).
[0229] Subsequently, the inventors tested the editing efficiency of PrimeROOT.v2C-Cre for large fragment DNA of GFP, and the experimental results showed that it achieved a GFP sequence editing efficiency as high as 4% at the endogenous site Figure 12B ).
[0230] The experimental results are similar to the editing efficiency in rice, indicating that the dual-ePPE and the PrimeROOT system composed of it of the present application have a wide and universal application prospect in plant synthetic biology and gene editing engineering, and can precisely insert the required DNA sequence without introducing the donor backbone sequence.
[0231] Example 5. Editing capacity of PrimeROOT vs. CRISPR-mediated NHEJ system
[0232] CRISPR-mediated NHEJ system is the only system reported to date that can perform targeted large fragment insertion in plants (Li, J. et al. Gene replacements and insertions in rice by intron targeting using CRISPR-Cas9. Nature Plants 2 (2016).; Dong, O. X. O. et al. Marker-free carotenoid-enriched rice generated through targeted gene insertion using CRISPR-Cas9. Nature Communications 11 (2020).). In this example, we compared the editing capacity of PrimeROOT.v2C-Cre with CRISPR-mediated NHEJ system in performing targeted insertion of GFP (720 bp), Act1 promoter (Act1P, 1.4 kb), Act1P-pigmR gene cassette (4.9 kb) and Act1P-pigmR-Act1P-OsMYB30 gene cassette (7.7 kb). The results showed that for the insertion of GFP, Act1P, both systems have similar insertion efficiency. However, for longer donor insertion, the average efficiency of PrimeROOT.v2C-Cre system is 2-4 times of NHEJ system (schematic of constructs see Figure 13A , graph of editing efficiency see Figure 13B ).
[0233] As for editing precision, we observed that the Act1P events inserted by PrimeROOT.v2C-Cre system showed clear Sanger sequencing results, but the results of insertion using NHEJ showed mixed peaks Figure 14A , underlined indicates imprecise insertion). This indicates that PrimeROOT system has superior editing precision compared to traditional CRISPR-mediated NHEJ system.
[0234] Subsequently, the inventors cloned edited insertions from protoplasts into bacteria and sequenced junctions between endogenous genomes and individual cloned insertions. When the inventors randomly selected 20 clones from PrimeROOT and NHEJ treated Act1P insertion samples, the inventors found that all 20 PrimeROOT generated insertions contained precise insertions as expected, while all 20 NHEJ generated insertions contained random DNA base insertions and deletions / duplications at their junctions Figure 14A
[0235] Next, the inventors used PrimeROOT and CRISPR-mediated NHEJ to insert Act1P and Act1P-pigmR sequences into genomic sites in rice calli Figure 14C ). After transformation and induction of calli, the inventors analyzed 95 callus clones from each treatment to compare editing efficiency and precision. PrimeROOT generated 2 precise Act1P insertions and 2 precise Act1P-pigmR insertions, while NHEJ generated 3 imprecise Act1P insertions and 1 imprecise Act1P-pigmR insertion Figure 14C imprecise insertions are denoted by underlined text, Figure 14D ). These results demonstrate that PrimeROOT is an effective editing tool that can be used to create large, targeted precise DNA insertions compared to NHEJ systems that rely heavily on double-stranded DNA breaks as intermediates.
[0236] Example 6: Precise, targeted insertion of actin promoters using the PrimeROOT tool
[0237] Many desirable agronomic traits are quantitative traits that depend on up- or down-regulation of certain specific genes, or on tissue-specific expression. This example utilizes the PrimeROOT system to precisely insert a beneficial promoter upstream of a gene of interest, thereby demonstrating the application of the PrimeROOT tool for plant trait improvement.
[0238] Specifically, the inventors used PrimeROOT.v2C-Cre to knock in a strong promoter into the 5’ UTR region of OsHPPD Figure 15A ). The inventors first designed 16 pairs of pegRNAs in the 5’ UTR and compared their RS insertion editing efficiency in rice protoplasts, determining the optimal pegRNA pair to have a 30% RS insertion frequency Figure 15B ). Next, the inventors utilized PrimeROOT.v2C-Cre and the pegRNA to perform particle bombardment insertion of the rice Actin1 promoter (Act1P) into rice calli. The inventors identified edited plants by amplifying the junction between the genome and the inserted donor sequence and assessed insertion precision by Sanger sequencing. A total of 12 precise Act1P insertion events were detected in 507 regenerated rice plants (2.4) Figure 15C ). These results demonstrate that PrimeROOT can serve as an effective genome insertion tool to introduce new genetic regulatory elements into plant genomes for breeding.
[0239] Example 7: Precise insertion of genes in GSH regions
[0240] To ensure that transgenes can be safely inserted into the plant genome, the inventors predicted the genomic safe harbor (GSH) regions throughout the Kitaake rice genome. Based on previous methods for GSH discovery and validation (Aznauryan, E. et al. Discovery and validation of human genomic safe harbor sites for gene and cell therapies. Cell Rep Methods 2, 100154 (2022); Sadelain, M., Papapetrou, E. P. & Bushman, F. D. Safe harbors for the integration of new DNA in the human genome. Nat Rev Cancer 12, 51-58 (2011)), the inventors used multiple algorithms to identify regions with certain distance from some elements such as gene coding regions, small RNA, miRNA, IncRNA, tRNA, promoters, enhancers, LTRs, etc. In this way, the inventors generated a new set of GSH regions, consisting of 30 regions, totaling 40 kb Figure 16A ). The Kitaake GSH regions are shown in Table 1. In addition, the inventors identified 33 GSHs in the rice genome, and the GSH regions that map to each other are shown in Table 2.
[0241] The inventors selected GSH1 (kitaake, Chr1:7660637-7661671) as a proof-of-concept region and designed 4 pairs of pegRNAs for RS insertion in this region (Table 3). When comparing the RS insertion efficiency using dual-ePPE in GSH1, the highest RS insertion efficiency was >40% Figure 16B). The inventors then tested the insertion of the 4.9 kb ActPlp- pigmR donor cassette into the GSH1 region. Gel electrophoresis and Sanger sequencing results showed that 19 Actl-pigmR insertion events were identified in 744 regenerated plants (2.6%). Importantly, all 19 junctions produced the same size amplification product and were shown by sequencing to be the result of precise insertion events, with the ends of the donor cassette perfectly matching the predictions.
[0242] Example 8: Method of PrimeROOT and donor delivery
[0243] To test the insertion efficiency of the PrimeROOT and donor components in the plant editing process, the inventors used Lox66 and FRT mutant Flm2 as landing sites to test the recovery efficiency of the overall edited plants from sequential transformation of PrimeROOT and donor components into rice calli (the sequential transformation system is referred to as PrimeROOT.v3). The inventors first evaluated dual-ePPE-mediated RS insertion into rice calli and achieved an editing efficiency of up to 84.7% ( Figure 17A ). In the first round of transformation, the inventors transformed PrimeROOT reagents (without donor) into calli via Agrobacterium, and after 1 month of hygromycin selection, the inventors enriched calli containing the desired RS insertion. These calli were then used as substrates for the second round of transformation, which contained donor vectors delivered via particle bombardment or Agrobacterium. After G418 selection and regeneration, the inventors examined the regenerated plants and measured the editing frequency of the desired insertion events ( Figure 17B ). The inventors found that the editing efficiency of Cre-Lox66 and FLP-Flm2 site-specific insertion of OsHPPD 5'UTR into ActlP was 7.1% and 8.3%, respectively, which was 3-fold and 3.5-fold higher than the efficiency of one-step transformation; when evaluating the editing efficiency of ActlP-pigmR site-specific insertion into GSH1, the inventors obtained an efficiency of 4.2% for the Cre-Lox66 site and 6.3% for the FLP-Flm2 site, which was 1.6-fold and 2.4-fold higher than the one-step plant transformation. When the inventors delivered the donor via Agrobacterium transformation, the inventors obtained an efficiency of 3.9% for the precise insertion event of ActlP-pigmR insertion into the GSH1 site. These results show that PrimeROOT.v3 can be performed using different delivery methods and further improves the efficiency of precise targeted gene insertion in plants.
[0244] Example 9: Testing of PrimeROOT for large fragment insertion in human cells
[0245] To test whether PrimeROOT works in human cells, the inventors first replaced the promoters of PrimeROOT.V2N-Cre, PrimeROOT.V2C-Cre with the commonly used expression promoter CMV promoter in human cells. The inventors designed pegRNA in four regions of hAAVS1, hACTB, hCCR5, hLMNB1, respectively, and constructed the pegRNA on the expression vector of hU6, and then transformed the above plasmids and donor plasmids containing GFP into HEK293 cell lines by the method of plasmid transformation. After 72 hours, the cell DNA was extracted, and then the efficiency was detected by ddPCR ( Figure 18A ), and at the same time, junction PCR was performed for first-generation sequencing detection, and it was found that the site-specific integration of GFP on the genome was completely accurate and predictable ( Figure 18B ). This example shows that the PrimeROOT system has the effect of accurate targeted gene insertion in human cells.
[0246] Table 1: GSH region summary
[0247]
[0248]
[0249] Table 2: 33 rice genome inter-mapped GSH regions
[0250]
[0251]
[0252] Table 3 GSH1 verification of designed pegRNA information
[0253]
[0254] Sequence information
[0255] > Wild-type SpCas9 amino acid sequence (SEQ ID NO: 1)
[0256]
[0257] nCas9(H840A) amino acid sequence (SEQ ID NO: 2)
[0258]
[0259] >Wild type M-MLV-RT amino acid sequence (SEQ ID NO. 3)
[0260] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP
[0261] >M-MLV-RT-connection (SEQ ID NO. 4)
[0262] DQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDT>RT-RNase H (SEQ ID NO. 5)
[0263] PDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLL
[0264] >NC (SEQ ID NO. 6)
[0265] ATVVSGQKQDRQGGERRRSQLDRDQCAYCKEKGHWAKDCPKKPRGPRGPRPQTSLL
[0266] >PR (SEQ ID NO. 7)
[0267] TLDDQGGQGQEPPPEPRITLKVGGQPVTFLVDTGAQHSVLTQNPGPLSDKSAWVQGATGGKRYRWTTDRKVHLATGKVTHSFLHVPDCPYPLLGRDLLTKLKAQIHFEGSGAQVMGPMGQPLQVL
[0268] >IN (SEQ ID NO. 8)
[0269] ENSSPYTSEHFHYTVTDIKDLTKLGAIYDKTKKYWVYQGKPVMPDQFTFELLDFLHQLTHLSFSKMKALLERSHSPYYMLNRDRTLKNITETCKACAQVNASKSAVKQGTRVRGHRPGTHWEIDFTEIKPGLYGYKYLLVFIDTFSGWIEAFPTKKETAKVVTKKLLEEIFPRFGMPQVLGTDNGPAFVSKVSQTVADLLGIDWKLHCAYRPQSSGQVERMNRTIKETLTKLTLATGSRDWVLLLPLALYRARNTPGPHGLTPYEILYGAPPPLVNFPDPDMTRVTNSPSLQAHLQALYLVQHEVWRPLAAAYQEQLDRPVVPHPYRVGDTVWVRRHQTKNLEPRWKGPYTVLLTTPTALKVDGIAAWIHAAHVKAADPGGGPSSRLTWRVQRSQNPLKIRLTREAP
[0270] M-MLV-RT-F155Y (SEQ ID NO. 9)
[0271] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAYFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP
[0272] M-MLV-RT-F155V (SEQ ID NO. 10)
[0273] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAVFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP
[0274] M-MLV-RT-F156Y (SEQ ID NO. 11)
[0275] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFYCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP
[0276] M-MLV-RT-D524N (SEQ ID NO. 12)
[0277] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTNGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP
[0278] > M-MLV-RT-N200C (SEQ ID NO. 13)
[0279] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAYFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP
[0280] > M-MLV-RT-ΔRNase H (SEQ ID NO. 14)
[0281] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPL
[0282] > M-MLV-RT-ΔRNase H-ΔConnection (SEQ ID NO. 15)
[0283] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGP
[0284] > Linker sequence (SEQ ID NO: 16)
[0285] SGGSSGGSSGSETPGTSESATPESSGGSSGGS
[0286] >gRNA scaffold (SEQ ID NO: 17)
[0287] guuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugc
[0288] >PPE (SEQ ID NO: 18)
[0289] TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TGGMGLAVRQAPLIIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPV QDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLP QGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQV KYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKA YQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLT KDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD ILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAE GKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRM ADQAARKAAITETPDTSTLLIENSSP SGGSPKKKRKV
[0290] >ePPE (SEQ ID NO: 19)
[0291]
[0292] tevopre (SEQ ID NO: 20)
[0293] CGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAA
[0294] pGS (SEQ ID NO: 21)
[0295] Atggagtcaaagattcaaatagaggacctaacagaactcgccgtaaagactggcgaacagttcatacagagtctcttacgactcaatgacaagaagaaaatcttcgtcaacatggtggagcacgacacacttgtctactccaaaaatatcaaagatacagtctcagaagaccaaagggcaattgagacttttcaacaaagggtaatatccggaaacctcctcggattccattgcccagctatctgtcactttattgtgaagatagtggaaaaggaaggtggctcctacaaatgccatcattgcgataaaggaaaggccatcgttgaagatgcctctgccgacagtggtcccaaagatggacccccacccacgaggagcatcgtggaaaaagaagacgttccaaccacgtcttcaaagcaagtggattgatgtgattggcagacatactgtcccacaaatgaagatggaatctgtaaaagaaaacgcgtgaaataatgcgtctgacaaaggttaggtcggctgcctttaatcaataccaaagtggtccctaccacgatggaaaaactgtgcagtcggtttggctttttctgacgaacaaataagattcgtggccgacaggtgggggtccaccatgtgaaggcatcttcagactccaataatggagcaatgacgtaagggcttacgaaataagtaagggtagtttgggaaatgtccactcacccgtcagtctataaatacttagcccctccctcattgttaagggagcaaaatctcagagagatagtcctagagagagaaagagagcaagtagcctagaagtagtcaaggcggcgaagtattcaggcacgtggccaggaagaagaaaagccaagacgacgaaaacaggtaagagctaagcatctagataagttgaaaacaatcttcaaaagtcccacatcgcttagataagaaaacgaagctgagtttatatacagctagagtcgaagtagtgatt
[0296] > ΦC31 (SEQ ID NO: 22)
[0297] MDTYAGAYDRQSRERENSSAASPATQRSANEDKAADLQREVERDGGRFRFVGHFSEAPGTSAFGTAERPEFERILNECRAGRLNMIIVYDVSRFSRLKVMDAIPIVSELLALGVTIVSTQEGVFRQGNVMDLIHLIMRLDASHKESSLKSAKILDTKNLQRELGGYVGGKAPYGFELVSETKEITRNGRMVNVVINKLAHSTTPLTGPFEFEPDVIRWWWREIKTHKHLPFKPGSQAAIHPGSITGLCKRMDADAVPTRGETIGKKTASSAWDPATVMRILRDPRIAGFAAEVIYKKKPDGTPTTKIEGYRIQRDPITLRPVELDCGPIIEPAEWYELQAWLDGRGRGKGLSRGQAILSAMDKLYCECGAVMTSKRGEESIKDSYRCRRRKVVDPSAPGQHEGTCNVSMAALDKFVAERIFNKIRHAEGDEETLALLWEAARRFGKLTEAPEKSGERANLVAERADALNALEELYEDRAAGAYDGPVGRKHFRKQQAALTLRQQGAEERLAELEAAEAPKLPLDQWFPEDADADPTGPKSWWGRASVDDKRVFVGLFVDKIVVTKSTTGRGQGTPIEKRASITWAKPPTDDDEDDAQDGTEDVAA
[0298] > Bxbl (SEQ ID NO: 23)
[0299] MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMSEF
[0300] > Cre (SEQ ID NO: 24)
[0301] SNLLTVHQNLPALPVDATSDEVRKNLMDMFRDRQAFSEHTWKMLLSVCRSWAAWCKLNNRKWFPAEPEDVRDYLLYLQARGLAVKTIQQHLGQLNMLHRRSGLPRPSDSNAVSLVMRRIRKENVDAGERAKQALAFERTDFDQVRSLMENSDRCQDIRNLAFLGIAYNTLLRIAEIARIRVKDISRTDGGRMLIHIGRTKTLVSTAGVEKALSLGVTKLVERWISVSGVADDPNNYLFCRVRKNGVAAPSATSQLSTRALEGIFEATHRLIYGAKDDSGQRYLAWSGHSARVGAARDMARAGVSIPEIMQAGGWTNVNIVMNYIRNLDSETGAMVRLLEDGD
[0302] > FLP (SEQ ID NO: 25)
[0303] MPQFDILCKTPPKVLVRQFVERFERPSGEKIALCAAELTYLCWMITHNGTAIKRATFMSYNTIISNSLSFDIVNKSLQFKYKTQKATILEASLKKLIPAWEFTIIPYYGQKHQSDITDIVSSLQLQFESSEEADKGNSHSKKMLKALLSEGESIWEITEKILNSFEYTSRFTKTKTLYQFLFLATFINCGRFSDIKNVDPKSFKLVQNKYLGVIIQCLVTETKTSVSRHIYFFSARGRIDPLVYLDEFLRNSEPVLKRVNRTGNSSSNKQEYQLLKDNLVRSYNKALKKNAPYSIFAIKNGPKSHIGRHLMTSFLSMKGLTELTNVVGNWSDKRASAVARTTYTHQITAIPDHYFALVSRYYAYDPISKEMIALKDETNPIEEWQHIEQLKGSAEGSIRYPAWNGIISQEVLDYLSSYINRRI
[0304] LoxP (SEQ ID NO: 26)
[0305] ATAACTTCGTATAGCATACATTATACGAAGTTAT
[0306] Lox66 (SEQ ID NO: 27)
[0307] ATAACTTCGTATAGCATACATTATACGAACGGTA
[0308] Lox71 (SEQ ID NO: 28)
[0309] taccgTTCGTATAGCATACATTATACGAAGTTAT
[0310] Lox2272 (SEQ ID NO: 29)
[0311] Ataacttcgtataggatactttatacgaagttat
[0312] FRT1 (SEQ ID NO: 30)
[0313] GAAGTTCCTATTCCGAAGTTCCTATTCTCTAGAAAGTATAGGAACTTC
[0314] FRT6 (SEQ ID NO: 31)
[0315] Gaagttcctattccgaagttcctattcttcaaaaagtataggaacttc
[0316] FRT1m1 (SEQ ID NO: 32)
[0317] GgAGgTCtTATTtCGAAGTTCCTATTCTCTAGAAAGTATAGGAACTTC
[0318] FRT1m2 (SEQ ID NO: 33)
[0319] GAAGTTCCTATTCCGgAGgTCtTATTtTCTAGAAAGTATAGGAACTTC
[0320] FRT1m3 (SEQ ID NO: 34)
[0321] GgAGgTCtTATTtCGAAGTTCCTATTCTCTAGAAAGTATAaGAcCTcC
[0322] mFRT1 (SEQ ID NO: 35)
[0323] GAAGTTCCTATTCTCTAGAAAGTATAGGAACTTC
[0324] mFRT1m1 (SEQ ID NO: 36)
[0325] GgAGgTCtTATTtTCTAGAAAGTATAGGAACTTC
[0326] mFRT1m2 (SEQ ID NO: 37)
[0327] GAAGTTCCTATTCTCTAGAAAGTATAaGAcCTcC
[0328] aTTP (SEQ ID NO: 38)
[0329] gtagtgccccaactggggtaacctttgagttctctcagttgggggcgtag
[0330] aTTB (SEQ ID NO: 39)
[0331] cggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccac
[0332] aGTP(SEQ ID NO:40)
[0333] GGTTTGTCTGGTCAACCACCGCGGTCTCAGTGGTGTACGGTACAAACC
[0334] aGTB(SEQ ID NO:41)
[0335] GGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGG
[0336] SpG-nCas9(SEQ ID NO:42)
[0337]
[0338] SpRY-nCas9: (SEQ ID NO: 43)
[0339]
[0340] MCP (SEQ ID NO: 44)
[0341] ASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEV PKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY
[0342] 2xMS2 (SEQ ID NO: 45)
[0343] GGGAGCACATGAGGATCACCCATGTGCCACGAGCGACATGAGGATCACCCATGTCGCTCGTGTTCCC
[0344] PrimeROOT.v2N-Cre (SEQ ID NO: 46)
[0345]
[0346] PrimeROOT.v2N-FLP (SEQ ID NO: 47)
[0347]
[0348] PrimeROOT.v2C-Cre (SEQ ID NO: 48)
[0349]
[0350] PrimeROOT.v2C-FLP (SEQ ID NO: 49)
[0351]
[0352] B2 recombinase amino acid sequence (SEQ ID NO: 50)
[0353] MSEFSELVRILPLDQVAEIKRILSRGDPIPLQRLASLLTMVILTVNMSKKRKSSPIKLSTFTKYRRNVAKSLYYDMSSKTVFFEYHLKNTQDLQEGLEQAIAPYNFVVKVHKKPIDWQKQLSSVHERKAGHRSILSNNVGAEISKLAETKDSTWSFIERTMDLIEARTRQPTTRVAYRFLLQLTFMNCCRANDLKNADPSTFQIIADPHLGRILRAFVPETKTSIERFIYFFPCKGRCDPLLALDSYLLWVGPVPKTQTTDEETQYDYQLLQDTLLISYDRFIAKESKENIFKIPNGPKAHLGRHLMASYLGNNSLKSEATLYGNWSVERQEGVSKMADSRYMHTVKKSPPSYLFAFLSGYYKKSNQGEYVLAETLYNPLDYDKTLPITTNEKLICRRYGKNAKVIPKDALLYLYTYAQQKRKQLADPNEQNRLFSSESPAHPFLTPQSTGSSTPLTWTAPKTLSTGLMTPGEE
[0354] KD recombinase amino acid sequence (SEQ ID NO: 51)
[0355] MSTFAEAAHLTPHQCANEINEILESDTFNINAKEIRNKLASLFSILTMQSLSIRREMKINTYRSYKSAIGKSLSFDKDDKIIKFTVRLRKTESLQKDIESALPSYKVVVSPFKNQEVSLFDRYEETHKYDASMVGLQFTNILSKEKDIWKIVSRIACFFDQSCVTTTKRAEYRLLLLGAVGNCCRYSDLKNLDPRTFEIYNNSFLGPIVRATVTETKSRTERYVNFYPVNGDCDLLISLYDYLRVCSPIEKTVSSNRPTNQTHQFLPESLARTFSRFLTQHVDEPVFKIWNGPKSHFGRHLMATFLSRSEKGKYVSSLGNWAGDREIQSAVARSHYSHGSVTVDDRVFAFISGFYKEAPLGSEIYVLKDPSNKPLSREELLEEEGNSLGSPPLSPPSSPRLVAQSFSAHPSLQLFEQWHGIISDEVLQFIAEYRRKHELRSQRTVVA
[0356] pSR1 recombinase amino acid sequence (SEQ ID NO: 52)
[0357] MQLTKDTEISTINRQMSDFSELSQILPLHQISKIKDILENENPLPKEKLASHLTMIILMANLASQ KRKDVPVKRSTFLKYQRSISKTLQYDSSTKTVSFEYHLKDPSKLIKGLEDVVSPYRFVVGVHE KPDDVMSHLSAVHMRKEAGRKRDLGNKINDEITKIAETQETIWGFVGKTMDLIEARTTRPTT KAAYNLLLQATFMNCCRADDLKNTDIKTFEVIPDKHLGRMLRAFVPETKTGTRFVYFFPCKG RCDPLLALDSYLQWTDPIPKTRTTDEDARYDYQLLRNSLLGSYDGFISKQSDESIFKIPNGPKA HLGRHVTASYLSNNEMDKEATLYGNWSAAREEGVSRVAKARYMHTIEKSPPSYLFAFLSGFY NITAERACELVDPNSNPCEQDKNIPMISDIETLMARYGKNAEIIPMDVLVFLSSYARFKNNE GKEYKLQARSSRGVPDFPDNGRTALYNALTAAHVKRRKISIVVGRSIDTS
[0358] B2 recombinase recognition site (SEQ ID NO: 53)
[0359] GAGTTTCATTAAGGAATAACTAATTCATTAAACTC
[0360] KD recombinase recognition site (SEQ ID NO: 54)
[0361] AAACGATATCAGACATTTGTCTGATAATGCTTCATTATCAGACAAATGTCTGATATCGTTT
[0362] pSR1 recombinase recognition site (SEQ ID NO: 55)
[0363] TTGATGAAAGAATAACGTATTCTTTCATCAA
[0364] Dre recombinase amino acid sequence (SEQ ID NO: 56)
[0365] MSELIISGSSGGFLRNIGKEYQEAAENFMRFMNDQGAYAPNTLRDLRLVFHSWARWCHARQLAWFPISPEMAREYFLQLHDADLASTTIDKHYAMLNMLLSHCGLPPLSDDKSVSLAMRRIRREAATEKGERTGQAIPLRWDDLKLLDVLLSRSERLVDLRNRAFLFVAYNTLMRMSEISRIRVGDLDQTGDTVTLHISHTKTITTAAGLDKVLSRRTTAVLNDWLDVSGLREHPDAVLFPPIHRSNKARITTTPLTAPAMEKIFSDAWVLLNKRDATPNKGRYRTWTGHSARVGAAIDMAEKQVSMVEIMQEGTWKKPETLMRYLRRGGVSVGANSRLMDS
[0366] Dre recombinase recognition site rox (SEQ ID NO: 57)
[0367] TAACTTTAAATAATGCCAATTATTTAAAGTTA
[0368] Dre recombinase recognition site rox (SEQ ID NO: 58)
[0369] TAACTTTAAATAATGTCCATTATTTAAAGTTA
[0370] Figure 14A PrimeROOT.v2-Cre insertion S20-T2 (SEQ ID NO: 59)
[0371] GATCCTGTGCAATTTGAAAGGAACCCTGACGAGATTCCGTGGGCTGAATAACTTCGTATAGCATACATTATACGAAGTTATTCGAGGTCATTCATAT
[0372] Figure 14A PrimeROOT.v2-Cre insertion S20-T4 (SEQ ID NO: 60)
[0373] ATACCTTCAAGTGAGCAGCAGCCTTCTCCTTGTCAGTGAAGACACTACCGTTCGTATAATGTATGCTATACGAACGGTAGGTCTACCTAC
[0374] Figure 14ANHEJ insertion S20-T2 (SEQ ID NO: 61)
[0375] GATCCTGTGCAATTTGAAAGGAACCCTGACGAGATTCCGTGGGCTGAGGGTGGGCTTGGCTTTGTTTTCGGTCTCCGCCCCCCCGGGCGTTTTTATG
[0376] Figure 14A NHEJ insertion S20-T4 (SEQ ID NO: 62)
[0377] CAGTGAAGACACCGGTGGACTCCACGACATACTCAGCACCAGCCGGGTGGGCGGGACCTCTTCTACCTACAAAAAAGCTCCGCACGA
[0378] Figure 14A NHEJ insertion S20-T2-80bp (SEQ ID NO: 63)
[0379] CGTGGAACTGATGTTT / / GA… / / … / / AAGGTGGTATA
[0380] Figure 14A NHEJ insertion S20-T2+35bp (SEQ ID NO: 64)
[0381] CGTGGAACTGATGTTT / / ATGGCTGGGCTTGGCCTTGAATTCGAGCTCGGTACCCTCGA / / TCAGTTAAAAGGTGGTATA
[0382] Figure 14A NHEJ insertion S20-T2+34 / -1bp (SEQ ID NO: 65)
[0383] CGTGGAACTGATGTTT / / GCTGGGCTTGGCCTTGAATTCGAGCTCGGTACCC-TCGAGG / / TCAGTTAAAAGGTGGTATA
[0384] Figure 14A NHEJ insertion S20-T2+44 / -79bp (SEQ ID NO: 66)
[0385] CGTGGAACTGATGTTTCAGTA / / CGGA… / / …AAAAGAGTTG / / AAGGTGGTATA
[0386] Figure 14A NHEJ insertion S20-T2+62 / -4bp (SEQ ID NO: 67)
[0387] CGTGGAACTGATGTTT / / GAGGGAGAGGCGGTG / / GCTCGCTGCGCTC...GGTCGTTCA / / AGTTAAAAGGTGGTATA
[0388] Figure 14A NHEJ insertion S20-T2+289 / -1bp (SEQ ID NO: 68)
[0389] CGTGGAACTGATGTTT / / TGCGTTTCTGGGTGAG / / TCGAGCTCGGTACCC-TCGAGGTC / / CAGTTAAAAGGTGGTATA
[0390] Figure 14A NHEJ insertion S20-T2+271 / -7bp (SEQ ID NO: 69)
[0391] CGTGGAACTGATGTTT / / TGGTCGTTCGCTCCAAGC / / CTCGAGGTCAT...TCGAGGTC / / TAAAAGGTGGTATA
[0392] Figure 14A NHEJ insertion S20-T2+151bp (SEQ ID NO: 70)
[0393] CGTGGAACTGATGTTT / / GAGTTTTCGTTCCACTGACT / / TAATTCGAGCTCGGTACCCTC / / AGTTAAAAGGTGGTATA
[0394] Figure 14A NHEJ insertion S20-T2+72bp (SEQ ID NO: 71)
[0395] CGTGGAACTGATGTTT / / TGAGGTAAGATTACCTGGTC / / GAATTCGAGCTCGGTACCCTC / / AGTTAAAAGGTGGTATA
[0396] Figure 14A NHEJ insertion S20-T2+183 / -92bp (SEQ ID NO: 72)
[0397] CGTGG / / AA-A / / ATCC / / AAAT... / / ... / / AAGGTGGTATA
[0398] Figure 14A NHEJ insertion S20-T4+ 15 bp (1) (SEQ ID NO: 73)
[0399] CTTTTCATGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGCCGGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG
[0400] Figure 14A NHEJ insertion S20-T4+ 12 bp (SEQ ID NO: 74)
[0401] CTTTTCATGATTTGTGACAAATGCAGCCT / AGAAGAGGTACCGGCCAGCCTGGTGCTGAGTATGTCGTGGAGTCCACCGG
[0402] Figure 14A NHEJ insertion S20-T4+ 109 / -2 bp (SEQ ID NO: 75)
[0403] CCTTTCATGATTTGTGACAAATGC / / GAAGAGGTAT / / CTGCGTTA--GCCGGTGCTGAGTATGTCGTGGAGTCCACCGG
[0404] Figure 14A NHEJ insertion S20-T4+ 15 bp (2) (SEQ ID NO: 76)
[0405] CTTTTCACGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGCCGGCTGGTGCTGAGTGTGTCGTGGAGTCCACCGG
[0406] Figure 14A NHEJ insertion S20-T4+ 172 / -5 bp (SEQ ID NO: 77)
[0407] CTTTTCATGATTTGTGACAAATG / / GAAGAGGTAC / / AGACCCCGT-----TGGTGCTGAGTATGTCGTGGAGTCCACCGG
[0408] Figure 14ANHEJ insertion S20-T4+15bp (3) (SEQ ID NO: 78)
[0409] CTTTTCACGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGCCGGCTGGTGCTGAGTGTGTCGTGGAGTCCACCGG
[0410] Figure 14A NHEJ insertion S20-T4+47 / -2bp (SEQ ID NO: 79)
[0411] CTTTTCATGATTTGTGACAAATGC / / GAAGAGGTACC / / AGCTCGG--GGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG
[0412] Figure 14A NHEJ insertion S20-T4+143 / -2bp (SEQ ID NO: 80)
[0413] CTTTTCATGATTTGTGGCAAATGC / / GAAGAGGTACC / / CCAGCCT--GGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG
[0414] Figure 14A NHEJ insertion S20-T4+15bp (4) (SEQ ID NO: 81)
[0415] CTTTTCATGATTTGTGACAAATGCAGC / / GAAGAGGTACCGGCCAGTCGGCTGGTGCTGAGTATGTCGTGGAGTCCACCGG
[0416] Figure 14A NHEJ insertion S20-T4+15bp (5) (SEQ ID NO: 82)
[0417] CTTTTCATGATTTGTGACAAATGCATC / AAAGAGGTACCGGCCAGCCGGCTGGCGCTGAGTATGTCGAGGAGTCCACCGG
[0418] Figure 14C PrimeROOT.v2C-Cre-Act1P or PrimeROOT.v2C-Act1P-pigmT insertion (SEQ ID NO: 83)
[0419] GAA GCA TCT GTC TGT CCAC TCC CCC ACT CGT TTT TCC GAA GTA AC / / CGA GCT CGG TAC CCT TCG AGG TCA TTC ATA
[0420] Figure 14C NHEJ strategy - ActP insertion ins 36bp (SEQ ID NO: 84)
[0421] CTC CCA CTC GTT GCA GGG CTT GGC / / TTC GAG CTC GGT ACC CTC GAG GTC ATTC AT A
[0422] Figure 14C NHEJ strategy - ActP insertion ins 12 / del 1bp (SEQ ID NO: 85)
[0423] ATG TCT GTC CCAC TCC CCC ACT CGA GCT CGG TAC CCT CCG AGG TCA TTC AT A
[0424] Figure 14C NHEJ strategy - ActP insertion ins 131bp (SEQ ID NO: 86)
[0425] CTC CCC ACT CGT TTT TCC GAA GTA AC / / CGA GCT CGG TAC CCT TCG AGG TCA TTC ATA
[0426] Figure 14C NHEJ strategy - ActP-pigmR insertion ins 38 / del 11bp (SEQ ID NO: 87)
[0427] CTC CCC ACT GCT TGG C / / CCT CCC TGC-----------CAT TCA TAT GCT TGA GA
[0428] Figure 18B GFP-N-terminal-hLMNB1-F (SEQ ID NO: 88)
[0429] ACG GCA TGG ACG AGC TGT ACA AGT AAT TTT TTT TAC CGT TCG TAT AGC AT ACAA TAT ACG AAC GTA AGC CCC ACG CGC CTG TCG CGG CTC CAG GAG AAG GAG GAG CTG CGC GAG CTC AAT GAC
[0430] Figure 18BGFP-C-term-hLMNBl-R (SEQ ID NO: 89)
[0431] CCCGGTGAACAGCTCCTCGCCCTTGCTCACCATATAACTTCGTATAATGTATGCTATACGAAGTTATCGGGCGGCGGAGACAGCGGGGCGGCGAGGCCGCGAGCGGGACCGTGATAAGGAG
Claims
1. A genome editing system for inserting an exogenous nucleotide sequence in a plant genome, comprising: i) a) a CRISPR nuclease and / or an expression construct containing a nucleotide sequence encoding the CRISPR nuclease, and a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the reverse transcriptase, or b) a prime editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the prime editing fusion protein, wherein the prime editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase; ii) a first pegRNA and / or an expression construct containing a nucleotide sequence encoding the first pegRNA, iii) a second pegRNA and / or an expression construct containing a nucleotide sequence encoding the second pegRNA, iv) a recombinase and / or an expression construct containing a nucleotide sequence encoding the recombinase, and v) a donor construct comprising one or more recognition sites (RS) of the recombinase and a second exogenous nucleotide sequence to be inserted into the genome; wherein the first pegRNA comprises, in the 5' to 3' direction, a first guide sequence, a first scaffold sequence, a first reverse transcription template (RT) sequence, and a first primer binding site (PBS) sequence, wherein the second pegRNA comprises, in the 5' to 3' direction, a second guide sequence, a first scaffold sequence, a second reverse transcription template (RT) sequence, and a second primer binding site (PBS) sequence, wherein the first pegRNA targets a first target sequence on the sense strand of the plant genomic DNA and the second pegRNA targets a second target sequence on the anti-sense strand of the plant genomic DNA, wherein the first RT sequence and the second RT sequence are used to insert a first exogenous nucleotide sequence comprising one or more recognition sites (RS) of the recombinase; wherein the recombinase is a Cre recombinase or a FLP recombinase.
2. The genome editing system of claim 1, wherein the pegRNA is capable of forming a complex with the CRISPR nuclease or fusion protein and targeting the CRISPR nuclease or fusion protein to a target sequence in the genome, resulting in a nick within the target sequence on the target strand.
3. The genome editing system of claim 1, wherein the first target sequence and the second target sequence are separated by 20-80 bp.
4. The genome editing system of claim 3, wherein the first target sequence and the second target sequence are separated by 20-60 bp.
5. The genome editing system of any one of claims 1-4, wherein the CRISPR nuclease is a Cas9 nuclease.
6. The genome editing system of any one of claims 1-4, wherein the CRISPR nuclease is a CRISPR nickase.
7. The genome editing system of claim 5, wherein the CRISPR nuclease is a Cas9 nickase.
8. The genome editing system of claim 7, wherein the Cas9 nickase consists of an amino acid sequence selected from the group consisting of SEQ ID NO: 2 and 42-43.
9. The genome editing system of any one of claims 1-4, wherein the CRISPR nuclease and the reverse transcriptase are linked by a linker.
10. The genome editing system of any one of claims 1-4, wherein the reverse transcriptase is an M-MLV reverse transcriptase.
11. The genome editing system of any one of claims 1-4, wherein the reverse transcriptase is missing the RNase H domain.
12. The genome editing system of any one of claims 1-4, wherein the reverse transcriptase is fused to a nucleocapsid protein (NC) at the N-terminus or C-terminus, directly or through a linker.
13. The genome editing system of claim 12, wherein the nucleocapsid protein (NC) is an amino acid sequence as set forth in SEQ ID NO:
6.
14. The genome editing system of any one of claims 1-4, wherein the reverse transcriptase is fused to an RNA aptamer binding protein sequence, either through a linker or directly, and the pegRNA comprises one or more RNA aptamer sequences.
15. The genome editing system of claim 14, wherein the RNA aptamer binding protein sequence is an MCP protein sequence.
16. The genome editing system of claim 14, wherein the RNA aptamer sequence is an MS2 sequence.
17. The genome editing system of any one of claims 1-4, wherein the CRISPR nuclease in i)-b) is fused to the reverse transcriptase through a self-cleaving peptide.
18. The genome editing system of any one of claims 1-4, wherein the CRISPR nuclease in i)-b) is fused to the N-terminus of the reverse transcriptase.
19. The genome editing system of any one of claims 1-4, wherein the fusion protein in i)-b) consists of an amino acid sequence as set forth in SEQ ID NO:
19.
20. The genome editing system of any one of claims 1-4, wherein the guide sequence in the first pegRNA has sequence identity to a first target sequence on the sense strand, which complex with the CRISPR nuclease results in a nick in the first target sequence; and the guide sequence in the second pegRNA has sequence identity to a second target sequence on the anti-sense strand, which complex with the CRISPR nuclease results in a nick in the second target sequence.
21. The genome editing system of any one of claims 1-4, wherein the scaffold sequence of the gRNA is set forth in SEQ ID NO:
17.
22. The genome editing system of any one of claims 1-4, wherein the primer binding site sequence is disposed to be complementary to at least a portion of the target sequence, the primer binding site sequence being complementary to at least a portion of a 3’ overhang single strand resulting from the nick in the DNA strand in which the target sequence resides.
23. The genome editing system of any one of claims 1-4, wherein the RT sequence is configured to generate, upon reverse transcription with it as a template, a first foreign nucleotide sequence or a portion thereof to be inserted into the genome, or to generate a complement of the first foreign nucleotide sequence or the portion thereof to be inserted into the genome of the plant.
24. The genome editing system of any one of claims 1-4, wherein the first RT sequence of the first pegRNA is configured to generate, upon reverse transcription with it as a template, a first fragment of a first foreign nucleotide sequence to be inserted into the genome; and the second RT sequence of the second pegRNA is configured to generate, upon reverse transcription with it as a template, a complement of a second fragment of the first foreign nucleotide sequence to be inserted into the genome.
25. The genome editing system of claim 24, wherein the first and second fragments of the first foreign nucleotide sequence to be inserted at least partially overlap.
26. The genome editing system of claim 25, wherein the first and second fragments overlap by at least 10 bp to 50 bp.
27. The genome editing system of any one of claims 1-4, wherein the pegRNA further comprises a tevopre sequence at the 3’ end of the PBS.
28. The genome editing system of any one of claims 1-4, wherein the pegRNA further comprises a polyA sequence at the 3’ end.
29. The genome editing system of any one of claims 1-4, wherein the first foreign nucleotide sequence to be inserted is 1 bp to 700 bp in length.
30. The genome editing system of any one of claims 1-4, wherein the 5’ end of the pegRNA is linked to a first ribozyme or tRNA designed to cleave the fusion of the first ribozyme or tRNA to the pegRNA at the 5’ end of the pegRNA; and / or the 3’ end of the pegRNA is linked to a second ribozyme or tRNA designed to cleave the fusion of the second ribozyme or tRNA to the pegRNA at the 3’ end of the pegRNA.
31. The genome editing system of any one of claims 1-4, wherein the pegRNA is driven for transcription by a type II promoter.
32. The genome editing system of claim 31, wherein the type II promoter is a GS promoter.
33. The genome editing system of any one of claims 1-4, wherein the one or more recombinase recognition sites (RS) are selected from the group consisting of loxP, Lox2272, Lox71, Lox66, and any combination thereof.
34. The genome editing system of any one of claims 1-4, wherein the one or more recombinase recognition sites (RS) are selected from the group consisting of FRT1, FRT3, FRT5, FRT6, or a FRT1 variant consisting of one of SEQ ID NOs: 32-37, and any combination thereof.
35. The genome editing system of any one of claims 1-4, wherein the recombinase is comprised in the prime editing fusion protein of i)-b). The recombinase is located at the N- or C-terminus of the fusion protein, either directly or via a linker to the other part of the fusion protein.
36. The genome editing system of any one of claims 1-4, wherein the guide-editing fusion protein consists of the amino acid sequence set forth in any one of SEQ ID NOs: 46-49.
37. The genome editing system of any one of claims 1-4, wherein the second exogenous nucleotide sequence can be between 1 bp and 11.1 kb.
38. The genome editing system of any one of claims 1-4, wherein the first target sequence, the second target sequence, the first exogenous nucleotide sequence, and / or the second exogenous nucleotide sequence is associated with a plant trait, whereby insertion of the first and / or second exogenous nucleotide sequence results in the plant having an altered plant trait relative to a wild-type plant.
39. The genome editing system of any one of claims 1-4, wherein the plant trait is an agronomic trait.
40. The genome editing system of any one of claims 1-4, wherein the plant comprises a monocot and a dicot.
41. The genome editing system of claim 40, wherein the plant is a crop plant.
42. The genome editing system of claim 40, wherein the plant is selected from the group consisting of wheat, rice, maize, soybean, sunflower, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.
43. A method of producing a genetically modified plant, wherein the genetically modified plant comprises a site-specifically inserted exogenous nucleotide sequence, the method comprising introducing the genome editing system of any one of claims 1-42 into at least one of the plants.
44. The method of claim 43, wherein the method further comprises screening the at least one plant for a plant having a desired insertion of the exogenous nucleotide sequence.
45. The method of claim 43, wherein the genome editing system is introduced into the plant by a method selected from the group consisting of biolistics, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway, and ovary injection.
46. The method of any one of claims 43-45, wherein the introducing comprises transforming the genome editing system into an isolated plant cell or tissue, and then regenerating the transformed plant cell or tissue into a whole plant.
47. The method of any one of claims 43-45, wherein the introducing comprises transforming the genome editing system into a specific location on a whole plant.
48. The method of claim 47, wherein the specific location is selected from the group consisting of a leaf, a stem tip, a pollen tube, an ear shoot, or a hypocotyl.
49. The method of any one of claims 43-45, further comprising treating a plant cell, tissue, or whole plant that has been introduced with the genome editing system at an elevated temperature.
50. The method of claim 49, wherein the elevated temperature is 37 °C.
51. The method of any one of claims 43-45, wherein the components of the genome editing system are introduced into the plant simultaneously.
52. The method of any one of claims 43-45, wherein comprising introducing the genome editing system of any one of claims 1-42 into at least one of the plants, and comprising the steps of: 1) transforming components i)-iv) of the genome editing system to an isolated plant cell or tissue to obtain a plant cell or tissue with an inserted first foreign nucleotide sequence comprising one or more recognition sites (RS) for a recombinase; 2) transforming component v) of the genome editing system to the plant cell or tissue obtained in step 1) to thereby obtain a plant cell or tissue comprising an inserted second foreign nucleotide sequence; and 3) regenerating a whole plant from the plant cell or tissue obtained in step 2).
53. The method of any one of claims 43-45, wherein the foreign nucleotide sequence is inserted into a safe harbor site in the plant genome, the safe harbor site being in the plant genome 1) at least 5 kb from a protein coding region; 2) at least 30 kb from a miRNA coding region; 3) at least 20 kb from a IncRNA coding region; 4) at least 20 kb from a tRNA coding region; 5) at least 5 kb from a promoter and / or enhancer; 6) at least 20 kb from an LTR repeat; 7) at least 200 bp from a non-LTR repeat; and 8) at least 10 kb from a centromere.
54. The method of claim 53, wherein the plant is rice, and the safe harbor site is selected from the sites shown in Table 1 or Table 2.
Citation Information
Patent Citations
Genome editing system and method
WO2018149418A1
Novel gene editing system and related vector and method
CN113549648A
Methods and compositions for prime editing nucleotide sequences
US11447770B1
Prime editing guide RNAS, compositions thereof, and methods of using the same
US20230357766A1
Method for targeted modification of sequence of plant genome
WO2021082830A1