Efficient and precise plant gene knockout method
By inserting stop codon clusters into the plant genome using guided editing technology and a dual pegRNA strategy, the problems of high off-target efficiency and low efficiency in existing technologies have been solved, enabling highly efficient and precise gene knockout and multi-type genome editing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
- Filing Date
- 2025-11-07
- Publication Date
- 2026-05-15
AI Technical Summary
Existing gene knockout technologies suffer from high off-target efficiency, inaccurate protein translation termination positions, and low efficiency, making it difficult to achieve efficient and precise gene knockout.
We designed DNA sequences containing stop codon clusters (SCCs) using prime editing (PE) combined with a dual pegRNA strategy. The SCCs were then inserted into the plant genome using a Twin PE strategy, and CRISPR nuclease and reverse transcriptase were used to achieve efficient gene knockout with low off-target effects.
It achieves efficient and precise gene knockout, and can perform multiple types of genome editing simultaneously, such as gene knockout, base substitution, small fragment insertion and deletion, improving editing efficiency and flexibility.
Smart Images

Figure CN2025133387_15052026_PF_FP_ABST
Abstract
Description
A highly efficient and precise method for plant gene knockout Technical Field
[0001] This invention belongs to the field of genetic engineering. Specifically, this invention includes the design and efficient insertion method of protein translation stop codon cluster sequences to achieve efficient and precise knockout of plant genes. Utilizing the plant guided editing system (PE) to reverse transcribe DNA sequences and combining the ability of double pegRNAs to efficiently insert heterologous sequences, a DNA sequence containing a stop codon cluster (SCC) is efficiently inserted into the target gene to achieve efficient and precise knockout of plant genes. Based on this, multiple different pegRNAs are combined to construct an all-in-one system for plant genome modification, meaning that only one guided editing PE protein is needed to achieve multi-gene, multi-type genome modifications such as gene knockout, base substitution, small fragment insertion, deletion, substitution, and long fragment deletion and replication. Background Technology
[0002] Genomic variation is the foundation for creating new plant varieties. Plant breeding has a long history. From ancient times to the present, humans have been selecting low-frequency natural variations to domesticate plants. With the development of life science technologies, humans have been able to autonomously create genomic variations to improve plants through techniques such as hybridization and artificial mutagenesis. However, the genomic variations obtained through these techniques are random and inefficient. The development of targeted genome editing technologies, especially the simple and efficient CRISPR / Cas system, has enabled precise and efficient modification of plant genomes, greatly accelerating the plant breeding process.
[0003] CRISPR / Cas is a defense system used by bacteria and archaea to resist viral infection. It uses guide RNA (gRNA) to guide Cas nucleases to target and cleave specific sequences. Based on this principle, the CRISPR / Cas9 system has been developed, capable of targeted cleavage of the eukaryotic genome, resulting in small insertion or deletion mutations. Furthermore, by fusing the Cas9 protein with other effector factors, such as deaminases and reverse transcriptases, a series of genome editing tools, including base editors and guide editors, have been developed. These rich editing tools can not only perform conventional gene knockout but also perform precise base substitutions, small fragment replacements, insertions, or deletions in the genome. CRISPR / Cas technology is continuously developing towards greater precision, efficiency, and versatility.
[0004] Gene knockout is a genetic engineering technique that selectively inactivates a specific gene. Gene knockout has important applications in biological research, breeding, and disease treatment. The conventional gene knockout method uses gRNA to guide the Cas9 nuclease to cut the double strand of the target genome. The Cas9 protein contains two domains, HNH and RuvC, which cut the target strand and the non-target strand of the gRNA, respectively, resulting in a double-strand break (DSB) in the target genome. Subsequently, the cell repairs the DSB by non-homologous end joining (NHEJ), which may introduce a small number of base insertion or deletion (indel) mutations at the cut. If the indel is not a multiple of 3, it will cause a frameshift mutation in the target gene to produce a premature stop codon, causing the target protein to lose its function and achieving the knockout of the target gene. Cas9-mediated gene knockout is simple to operate, but it has a series of shortcomings. For example: (1) High off-target efficiency. When there is a region in the genome that is similar to the target sequence, it is very likely to cause non-target DNA cutting, resulting in unexpected and non-specific genetic modifications. (2) Inaccurate protein translation termination position. Since Cas9-mediated gene knockout generates indels through cellular NHEJ repair, non-3-fold indels cause frameshift mutations in the target protein until a stop codon is found in the downstream sequence. The sequence content and size of indels are uncertain, making the termination position of protein translation uncertain and difficult to achieve precise termination. (3) Gene knockout efficiency is lower than genome mutation efficiency. Based on the generation of indels, a certain percentage of 3N-fold indel mutations are generated, only adding or deleting a few amino acid residues in the wild-type protein, failing to generate a premature stop codon, resulting in actual gene knockout efficiency being lower than that of indel mutations.
[0005] In addition, by using a cytosine base editor, specific codons can be targeted, including the forward strand CGA (Arg,R), CAG (Gln,Q), and CAA (Gln,Q) codons and the reverse strand TGG (Trp,W) codons, causing cytosine (C) to be converted to thymine (T), generating an in situ premature stop codon, resulting in the knockout of the target gene. This strategy is known as CRISPR-STOP or iSTOP and CRISPR-BETS. CRISPR-STOP uses nickase Cas9 (nCas9) for gene knockout, which has the advantage of low off-target effects. However, the target design for CRISPR-STOP is difficult, and a suitable protospacer must meet several conditions: (1) it needs to be in the first half of the CDS; (2) it needs to be NGG PAM; (3) it needs to contain the CGA (Arg,R), CAG (Gln,Q), CAA (Gln,Q), and TGG (Trp,W) codons within the active window (1–17-nt) of the cytosine base editing. This results in a relatively small number of suitable protospacers. Furthermore, the position of the codon to be edited within the base editing window significantly impacts gene knockout efficiency. Summary of the Invention
[0006] The problem the invention aims to solve
[0007] Prime editing (PE) utilizes reverse transcriptase to provide exogenous integration fragments, enabling genomic modifications such as arbitrary base substitutions, small fragment insertions, deletions, and replacements. Key components of a prime editing guide RNA (pegRNA) include nCas9, reverse transcriptase, and pegRNA. The pegRNA consists of a protospacer, gRNA scaffold, RT template, and prime binding site (PBS). By programming the RT template, various types of editing at target sites can be achieved. Because PE editing requires multiple steps, including gRNA guidance, PBS binding, and the binding of the reverse-transcribed template to the genome, PE has a low off-target rate. Although the early efficiency of prime editing guide RNA was low, the efficiency of PE has been significantly improved by optimizing the reverse transcriptase protein sequence, the nCas9 protein sequence, the pegRNA structure, and the implementation strategy. Previous reports have shown that the PAM-in dual pegRNA strategy can efficiently insert recombinase recognition sites with editing efficiency of over 50% (Anzalone et al, Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing, Nature Biotechnology 2022; Sun et al., Precise integration of large DNA sequences in plant genomes using PrimeRoot editors, Nature biotechnology 2023).
[0008] In this study, the inventors designed a series of DNA sequences containing stop codon clusters (SCCs) and used a twin prime editing (Twin PE) strategy to insert the SCCs into the coding sequence (CDS) of genes, achieving efficient, low-off-target, and precise colonization of stop codons in plants to knock out target genes (see Figure 1a). The core technology of this study is named TwinPE-mediated gene KnockOut editor (TKO). Furthermore, combining the base substitution, small fragment insertion, deletion, and replacement functions of the prime editing technology, the inventors established an all-in-one multi-gene, multi-type genome editing TRIM (TKO editor-enabled gene Rupture and development of Integrated Multi-type genome modification systems) platform. Through this platform, efficient multi-type genome editing, including gene knockout, base substitution, small fragment insertion, deletion, and replacement, can be achieved with just one PE protein. This comprehensive editing platform will bring greater convenience and efficiency to the field of genome editing, providing researchers with more possibilities and flexibility.
[0009] Solution for solving the problem
[0010] [1]. A guide editing RNA (pegRNA) comprising a first pegRNA and a second pegRNA, wherein the first pegRNA or the second pegRNA includes a guide sequence, a scaffold sequence, a reverse transcription template (RT) sequence, and a primer binding site (PBS) sequence.
[0011] The scaffold sequence of the first pegRNA can be complexed with CRISPR nuclease and create a nick in the first target sequence of the sense strand of the target double-stranded DNA sequence, and the reverse transcription template (RT) sequence of the first pegRNA sequence has a stop codon cluster (SCC) sequence.
[0012] The SCC contains at least one, at least two, or at least three stop codons;
[0013] Optionally, the stop codon is selected from TAA, TAG, and TGA, preferably from TAA and TAG;
[0014] The reverse transcription template (RT) sequence of the second pegRNA is partially or completely complementary to the RT sequence of the first pegRNA, preferably with at least 13 bp to about 40 bp overlap.
[0015] The second pegRNA can complex with a CRISPR nuclease and create a nick in the second target sequence of the antisense strand of the target double-stranded DNA sequence.
[0016] [2]. The guide editing RNA according to [1], wherein the sequence length of the SCC is 13 to 38 nt.
[0017] [3]. The guide editing RNA according to [1] or [2], wherein the SCC comprises at least one of the following features (a) to (b):
[0018] (a) The RNA secondary structure of the SCC includes a stem-loop structure;
[0019] (b) The Gibbs free energy of the RNA secondary structure of the SCC is in the range of -0.2 to -0.05 kcal / mol / base.
[0020] [4]. The guide editing RNA according to any one of [1] to [3], wherein the 3' end of the SCC is a base C.
[0021] [5]. The guide editing RNA according to any one of [1] to [4], wherein the SCC comprises a sequence as shown in any one of SEQ ID NO:39-72.
[0022] [6]. The guide editing RNA according to any one of [1] to [5], wherein the interval between the cuts of the first target sequence and the second target sequence is not less than 20 bp, preferably not less than 30 bp, and more preferably 30 bp to 100 bp.
[0023] [7]. The guide editing RNA according to any one of [1] to [6], wherein the scaffold sequence of the first pegRNA or the second pegRNA is as shown in SEQ ID NO:7.
[0024] [8] A genome editing system comprising:
[0025] i)a) An expression construct containing a CRISPR nuclease and / or a nucleotide sequence encoding the CRISPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase, or
[0026] b) A guide-editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide-editing fusion protein, wherein the guide-editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase; and,
[0027] ii) at least one pegRNA and / or an expression construct containing a nucleotide sequence encoding the at least one pegRNA, wherein the at least one pegRNA comprises at least one, at least two, at least three, at least four, at least five, at least six or more pegRNAs as described in any one of [1] to [7], and the pegRNA as described in any one of [1] to [7] consists of a first pegRNA and a second pegRNA.
[0028] [9]. According to the genome editing system described in [8], wherein the fusion protein in i)-b) further comprises a recombinase, preferably Cre, Bxb1, or phiC31 recombinase, more preferably Cre.
[0029]
[0010] . According to the genome editing system described in [8] or [9], wherein the at least one pegRNA is transcribed by a complex promoter comprising a 35S enhancer, a CmYLCV promoter and / or a U6 promoter;
[0030] Optionally, the U6 promoter is a truncated variant of the U6 promoter.
[0031]
[0011] . The genome editing system according to any one of [8] to
[0010] , wherein,
[0032] The first pegRNA and / or the second pegRNA in the at least one pegRNA are both in different expression constructs, or
[0033] At least two first pegRNAs, at least two second pegRNAs, or at least one first pegRNA and at least one second pegRNA are in the same expression construct, or
[0034] The first and second pegRNAs in the at least one pegRNA are both in the same expression construct.
[0035]
[0012] . The genome editing system according to any one of [8] to
[0011] , wherein the first pegRNA and / or the second pegRNA further comprises a tevopre sequence at the 3' end of the PBS sequence; and / or
[0036] The first pegRNA and / or the second pegRNA also contain a polyT sequence at the 3' end.
[0037]
[0013] . A genome editing system according to any one of [8] to
[0012] , wherein the 5' end of a first pegRNA and / or a second pegRNA is linked to a first ribozyme or tRNA, the first ribozyme or tRNA being designed to cleave the first pegRNA and / or the second pegRNA at the 5' end of the first pegRNA and / or the second pegRNA, or to cleave a fusion of the first pegRNA and / or the second pegRNA, a tevopre sequence, and / or a polyT sequence; and / or the 3' end of the first pegRNA and / or the second pegRNA is linked to a second ribozyme or tRNA, the second ribozyme or tRNA being designed to cleave the first pegRNA and / or the second pegRNA at the 3' end of the first pegRNA and / or the second pegRNA, or to cleave a fusion of the first pegRNA and / or the second pegRNA, a tevopre sequence, and / or a polyT sequence;
[0038] Optionally, the tRNA is selected from at least one of tRNAGly, tRNAOsAsp, and tRNAZmIle;
[0039] Optionally, the ribozyme includes the HDV ribozyme.
[0040]
[0014] . According to the genome editing system described in
[0013] , wherein when the first pegRNA and the second pegRNA in the at least one pegRNA are both in the same expression construct, and the total number of the first pegRNA and the second pegRNA in the at least one pegRNA is 6 or less, the tRNA is tRNAGly; or,
[0041] When the first and second pegRNAs in the at least one pegRNA are both in the same expression construct, and the total number of the first and second pegRNAs in the at least one pegRNA is greater than 6, the tRNA is selected from at least two of tRNAGly, tRNAOsAsp and tRNAZmIle, preferably three.
[0042]
[0015] . A genome editing system according to any one of [8] to
[0014] , wherein the first target sequence and the second target sequence in the pegRNA are associated with plant traits such as agronomic traits, thereby causing the plant to have altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant by the gene editing system.
[0043]
[0016] . A genome editing system according to any one of [8] to
[0015] , wherein the at least one pegRNA is capable of forming a complex with the CRISPR nuclease or fusion protein and targeting the CRISPR nuclease or fusion protein to a target sequence in the genome, resulting in a cut in the target strand within the target sequence.
[0044]
[0017] . A genome editing system according to any one of [8] to
[0016] , wherein the CRISPR nuclease is a Cas9 nuclease or a variant thereof.
[0045] 18. A genome editing system according to any one of [8] to
[0017] , wherein the CRISPR nuclease is a CRISPR nickase, such as a Cas9 nickase or a variant thereof, such as the Cas9 nickase or a variant thereof comprising a sequence selected from the sequence shown in SEQ ID NO: 1 or 2;
[0046] Optionally, the CRISPR nuclease, such as Cas9 nickase, and the reverse transcriptase are linked by a adapter.
[0047]
[0019] . A genome editing system according to any one of [8] to
[0018] , wherein the reverse transcriptase is M-MLV reverse transcriptase or a functional variant thereof;
[0048] Optionally, the RNase H domain of the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is deleted, and it contains the sequence shown in SEQ ID NO:5.
[0049]
[0020] . A genome editing system according to any one of [8] to
[0019] , wherein a reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused to a nucleocapsid protein (NC) directly or via a linker at the N-terminus or C-terminus.
[0050]
[0021] . According to the genome editing system of
[0020] , wherein the nucleocapsid protein (NC) comprises the amino acid sequence shown in SEQ ID NO:6.
[0051]
[0022] . The genome editing system according to any one of [8] to
[0021] , wherein the CRISPR nuclease, such as the CRISPR nickase, described in i)-b) is fused to the N-terminus of the reverse transcriptase;
[0052] Optionally, the fusion protein described in i)-b) comprises the amino acid sequence shown in SEQ ID NO:11.
[0053]
[0023] . A genome editing system according to any one of [8] to
[0022] , wherein the genome is derived from a microorganism, animal or plant;
[0054] The microorganisms include bacteria and fungi;
[0055] The animals include mammals and poultry;
[0056] The plants include monocotyledons and dicotyledons;
[0057] Optionally, the mammals include humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats;
[0058] Optionally, the poultry includes chickens, ducks, and geese;
[0059] Optionally, the monocotyledonous plants include rice, corn, wheat, sorghum, and barley, and the dicotyledonous plants include soybean, peanut, Arabidopsis thaliana, cotton, and rapeseed.
[0060]
[0024] . The application of the genome editing system described in any of [8] to
[0023] in the following (A) or (B):
[0061] (A) Editing of the genome sequence of an organism or its cells;
[0062] (B) Prepare products that modify the genome sequence of an organism or biological cell.
[0063]
[0025] . A kit comprising the genome editing system described in any one of [8] to
[0024] .
[0064]
[0026] . A method for producing genetically modified cells or organisms, the method comprising introducing a genome editing system as described in any one of [8] to
[0024] into at least one of the cells or organisms, thereby resulting in modification of the genome sequence of the at least one cell or organism, such as the modification comprising substitution, deletion and / or addition of one or more nucleotides.
[0065]
[0027] . According to the method of
[0026] , wherein the method further includes screening organisms having desired exogenous nucleotide sequence insertions from the at least one organism.
[0066]
[0028] . The method according to
[0026] or
[0027] , wherein the components of the genome editing system are simultaneously introduced into the organism.
[0067] The effects of the invention
[0068] This invention also screened out stop codon cluster sequences with high editing efficiency, which can achieve efficient knockout of target genes.
[0069] The genome editing method provided by this invention can precisely edit and modify target genes in various ways, including gene knockout, base substitution, insertion, deletion and replacement of small fragments.
[0070] Furthermore, the method provided by this invention can edit multiple genes using only one guide editing protein, thus possessing excellent multi-gene editing capabilities. Attached Figure Description
[0071] Figure 1 shows the selection of SCCs that can be efficiently inserted to establish a TKO.
[0072] Figure 2 shows a schematic diagram of the gRNA expression plasmid structure, SCC secondary structure, and Gibbs free energy.
[0073] Figure 3 shows the expression forms of (pe)gRNA, Cas9, and ePPEplus, a schematic diagram of TKO paired target design, and a comparison of the efficiency of Cas9 and TKO gene knockout in rice, wheat, and maize.
[0074] Figure 4 shows the selection of more SCCs that can be efficiently inserted and the selection of orthogonal SCCs.
[0075] Figure 5 shows a comparison of the efficiency of Cas9 and orthogonal TKO editors in multi-gene knockout in rice.
[0076] Figure 6 shows the structural diagrams of the multi-gene editing vectors OsAssembly1, OsAssembly2, OsAssembly3, ZmAssembly1, and TaAssembly1.
[0077] Figure 7 shows the All-In-One multi-gene, multi-type TRIM genome modification strategy tested in rice. Detailed Implementation
[0078] Various exemplary embodiments, features, and aspects of the present invention will be described in detail below. The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0079] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art should understand that the present invention can be practiced without certain specific details. In other instances, methods, means, apparatus, and steps well known to those skilled in the art have not been described in detail in order to highlight the spirit of the present invention.
[0080] Unless otherwise stated, all units used in this specification are international standard units, and all numerical values and ranges appearing in this invention should be understood to include systematic errors that are unavoidable in industrial production.
[0081] In this specification, the word "may" has two meanings: to perform a certain process and not to perform a certain process.
[0082] In this specification, references to "some specific / preferred embodiments," "other specific / preferred embodiments," "implementation," etc., refer to specific elements (e.g., features, structures, properties, and / or characteristics) related to that embodiment, which are included in at least one of the embodiments described herein and may or may not be present in other embodiments. Furthermore, it should be understood that these elements may be combined in any suitable manner in various embodiments.
[0083] In this specification, "optional" and "optionally" mean that the events or circumstances described below may or may not occur, and the description includes both cases where the events or circumstances occur and cases where the events or circumstances do not occur.
[0084] In this specification, the range of values referred to as "value A to value B" refers to the range including the endpoint values A and B.
[0085] In this specification, a "genome editing system" refers to a combination of components required for genome editing within cells. The individual components of such a system, such as a guide editing fusion protein or its expression construct, pegRNA or its expression construct, donor construct, etc., may exist independently or in any combination as a composition.
[0086] In this specification, the term "genome," as used herein, encompasses not only chromosomal DNA present in the cell nucleus but also organelle DNA present in subcellular components of the cell, such as mitochondria and plastids.
[0087] In this specification, "genetically modified plant" means a plant whose genome contains inserted exogenous polynucleotides. For example, exogenous polynucleotides can be stably integrated into the plant's genome and inherited across generations.
[0088] In relation to a sequence, “exogenous” means a sequence that originates from a foreign species, or, if from the same species, a sequence whose composition and / or loci have been significantly altered from its natural form through deliberate human intervention.
[0089] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” are used interchangeably and refer to single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “D” for A, T, or G, “I” for inosine, and “N” for any nucleotide. Although nucleotide sequences may be represented as DNA sequences (containing T) herein, when referring to RNA, those skilled in the art can readily determine the corresponding RNA sequence (i.e., replacing T with U).
[0090] In this specification, the terms "polypeptide," "peptide," and "protein" are used interchangeably to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.
[0091] As used in this invention, "expression construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to generate mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein.
[0092] The "expression construct" of the present invention may be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, may be a translatable RNA (such as mRNA), for example, RNA transcribed in vitro.
[0093] The "expression construct" of the present invention may contain regulatory sequences and nucleotide sequences of interest from different sources, or regulatory sequences and nucleotide sequences of interest from the same source but arranged in a manner different from those normally found in nature.
[0094] In this specification, "promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the invention, a promoter is a promoter capable of controlling gene transcription in a cell, regardless of whether it originates from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmental regulatory promoter, or an inducible promoter. Examples of promoters include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. When used in plants, the promoter can be the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, the maize U3 promoter, or the rice actin promoter.
[0095] "Introducing" nucleic acid molecules (e.g., plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the organism's cells with the nucleic acid or protein, enabling the nucleic acid or protein to perform its function within the cell. The term "transformation" as used in this invention includes stable transformation and transient transformation. "Stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome, resulting in the stable inheritance of the foreign gene. Once stable transformation occurs, the foreign nucleic acid sequence is stably integrated into the genome of the organism and its genome in any subsequent generations. "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell to perform its function without the stable inheritance of the foreign gene. In transient transformation, the foreign nucleic acid sequence does not integrate into the genome.
[0096] "Temperament" refers to the physiological, morphological, biochemical, or physical characteristics of a cell or organism.
[0097] "Agronomic traits" specifically refer to measurable parameters of crop plants, including but not limited to: leaf greenness, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, total nitrogen content of plants, nitrogen content of fruits, nitrogen content of seeds, nitrogen content of plant vegetative tissues, total free amino acid content of plants, free amino acid content of fruits, free amino acid content of seeds, free amino acid content of plant vegetative tissues, total protein content of plants, protein content of fruits, protein content of seeds, protein content of plant vegetative tissues, herbicide resistance and drought resistance, nitrogen uptake, root lodging, harvest index, stem lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt tolerance, and tiller number, etc.
[0098] In this specification, InDel (Insertion and Deletion) refers to small insertion or deletion variations occurring in the genome, typically between 1 and 50 bp in length. These variations are relatively common in the human genome and can significantly impact gene function. InDel variations may occur within gene coding regions, leading to frameshift mutations, or affect gene expression regulatory regions, thereby influencing gene expression and function.
[0099] In this specification, "PAM-in" refers to the PAM (Protospacer Adjacent Motif) sequence recognized by the CRISPR / Cas system in gene editing technology. A PAM is a short nucleotide sequence recognized by the CRISPR system, typically located upstream of the target gene sequence. Different Cas proteins recognize different PAM sequences; for example, the commonly used Cas9 protein usually recognizes the NGG sequence as its PAM.
[0100] In this specification, "orthogonality" means that SCCs can work independently and in parallel without cross-reaction or interference between them.
[0101] The technical solution of the present invention will be described in detail below:
[0102] The dual pegRNA strategy (i.e., TwinPE) of the present invention comprises two pegRNAs that target and bind to two strands of genomic DNA respectively, with a certain distance between the two cuts (approximately 20bp to approximately 100bp, for example 30bp to 50bp). The RT of both pegRNAs contains only the desired insertion sequence, and the 3' ends have partially overlapping sequences. After reverse transcription is completed, the two newly synthesized DNA strands bind to each other due to the overlapping sequences and anneal, and the insertion is completed through a DNA repair pathway different from the original PE system.
[0103] Meanwhile, by fusing retroviral nucleocapsid protein (NC), deleting the RNaseH active domain of reverse transcriptase MLV, and modifying the Cas9 protein, an enhanced plant guided editing system (ePPEplus) was established. This system can enhance reverse transcription ability or reverse transcriptase stability, and simultaneously modify epigRNA, thereby greatly improving the efficiency of the plant guided editing system.
[0104] Based on this, the present invention designs a series of sequences (SCCs) containing stop codon clusters, and then inserts these SCCs into the CDS of the target gene to achieve the knockout of the target gene.
[0105] I. Guided editing of RNA
[0106] This invention provides a guide editing RNA (pegRNA), which in some embodiments consists of a first pegRNA and a second pegRNA, wherein the first pegRNA or the second pegRNA includes a guide sequence, a scaffold sequence, a reverse transcription template (RT) sequence, and a primer binding site (PBS) sequence.
[0107] In some embodiments, the scaffold sequence of the first pegRNA can be complexed with a CRISPR nuclease and create a nick in the first target sequence of the sense strand of the target double-stranded DNA sequence, and the reverse transcription template (RT) sequence of the first pegRNA sequence has a stop codon cluster (SCC) sequence.
[0108] In some implementations, the SCC contains at least one, at least two, or at least three stop codons.
[0109] In some alternative implementations, the stop codon is selected from TAA, TAG, and TGA, preferably from TAA and TAG.
[0110] In some embodiments, the reverse transcription template (RT) sequence of the second pegRNA has partial or complete complement to the RT sequence of the first pegRNA, preferably with at least 13 bp to about 40 bp overlap.
[0111] The guide sequence (also called seed sequence or spacer sequence) in the pegRNA of the present invention is configured to have sufficient sequence identity (preferably 100% identity) with the target sequence, thereby enabling it to bind to the complementary strand of the target sequence through base pairing and achieve sequence-specific targeting.
[0112] For example, the guide sequence in the first pegRNA may have sufficient sequence identity (preferably 100% identity) with the first target sequence, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the first target sequence; the guide sequence in the second pegRNA may have sufficient sequence identity (preferably 100% identity) with the second target sequence on the opposite strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the second target sequence, thereby the two pegRNAs result in nicks on different strands of the genomic DNA.
[0113] In some embodiments, the primer-binding sequence in the pegRNA is configured to be complementary to at least a portion of the target sequence (preferably perfectly paired with at least a portion of the target sequence). Preferably, the primer-binding sequence is complementary to at least a portion of the 3' free single strand in the DNA strand containing the target sequence due to a nick (preferably perfectly paired with at least a portion of the 3' free single strand), particularly complementary to the nucleotide sequence at the 3' end of the 3' free single strand (preferably perfectly paired). When the 3' free single strand of the strand binds to the primer-binding sequence through base pairing, the 3' free single strand can act as a primer, using the reverse transcription template (RT) sequence immediately adjacent to the primer-binding sequence as a template, to perform reverse transcription under the action of reverse transcriptase in the fusion protein, extending the DNA sequence corresponding to the reverse transcription template (RT) sequence.
[0114] The primer-binding sequence depends on the length of the free single strand formed by the CRISPR nicking enzyme in the target sequence; however, it should have a minimum length to ensure specific binding. In some embodiments, the primer-binding sequence can be 4-20 nucleotides long, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.
[0115] In some embodiments, the RT template sequence can be any sequence. Through reverse transcription, its sequence information can be integrated into the DNA strand containing the target sequence (i.e., the strand containing the target sequence PAM), and then, through cellular DNA repair, a DNA double strand containing the RT template sequence information is formed. In some embodiments, the RT template sequence contains desired modifications. For example, the desired modifications include substitution, deletion, and / or addition of one or more nucleotides. In some embodiments, the RT template sequence is configured to correspond to a sequence downstream of the target sequence nick (e.g., complementary to at least a portion of the sequence downstream of the target sequence nick), but contains desired modifications. The desired modifications include substitution, deletion, and / or addition of one or more nucleotides.
[0116] In some implementations, the two pegRNAs are configured to introduce the same desired modification. For example, one pegRNA is configured to introduce A-G substitutions at the sense strand, while the other pegRNA is configured to introduce T-C substitutions at the corresponding position on the antisense strand. As another example, one pegRNA is configured to introduce a two-nucleotide deletion at the sense strand, and the other pegRNA is configured to similarly introduce a two-nucleotide deletion at the corresponding position on the antisense strand. Other types of modifications can be deduced similarly. The same desired modification can be achieved by designing suitable RT template sequences that target two different strands of the pegRNA.
[0117] In some embodiments, the RT sequence is configured to generate a foreign nucleotide sequence or a portion thereof of the genome to be inserted into after reverse transcription using it as a template, or to generate a complementary sequence to a foreign nucleotide sequence or a portion thereof of the genome of the organism to be inserted, such as a plant. In some embodiments, the RT sequence does not contain a genome sequence near the target sequence or a complementary sequence to a genome sequence near the target sequence. In some embodiments, the RT sequence does not contain sequence information other than the foreign nucleotide sequence to be inserted. In some exemplary embodiments of the present invention, the foreign nucleotide sequence includes a stop codon cluster (SCC) sequence. The stop codon cluster contains at least one, at least two, or at least three stop codons.
[0118] In some implementations, the stop codon is selected from TAA, TAG, and TGA, preferably TAA and TAG.
[0119] In some preferred embodiments, the 3' end of the SCC is a base C.
[0120] In some preferred embodiments, the SCC is 13 to 38 nt in length and has a secondary structure similar to that of recombinase recognition sites (e.g., attB, attP, lox66, lox71) (stem-loop structure, local double helix structure and unpaired regions formed within the RNA molecule by base pairing); and, optionally, the Gibbs free energy of the SCC is in the range of -0.2 to -0.05 kcal / mol / base.
[0121] In some implementations, the SCC comprises sequences as shown in SEQ ID NO:39-72.
[0122] In some specific implementations, the RT sequence is configured to generate a foreign nucleotide sequence or a portion thereof of the genome to be inserted after reverse transcription using it as a template, or to generate a complementary sequence of the foreign nucleotide sequence or a portion thereof of the genome to be inserted.
[0123] In some embodiments, the first RT sequence of the first pegRNA is configured to generate a first fragment of a first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; the second RT sequence of the second pegRNA is configured to be a complementary sequence to generate a second fragment of a exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template.
[0124] In some embodiments, the first and second fragments of the exogenous nucleotide sequence to be inserted at least partially overlap. In some embodiments, the first and second fragments overlap by at least about 10 bp to about 50 bp, for example, at least about 10 bp, about 15 bp, about 20 bp, about 25 bp, about 30 bp, about 35 bp, or about 40 bp. In some specific embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted completely overlap, or are guaranteed to have a complete overlap of 13–40 bp, 20–30 bp, or 25–31 bp.
[0125] In some embodiments, the length of the foreign nucleotide sequence to be inserted is approximately 1 bp to approximately 40 bp, such as approximately 10 bp, approximately 20 bp, approximately 30 bp, approximately 40 bp, or any value in between.
[0126] As used herein, a "target sequence" refers to a sequence of approximately 20 nucleotides in length in the genome characterized by a PAM (pre-intermediate sequence adjacent motif) sequence flanking the 5' or 3' region. Typically, the PAM is necessary for the recognition of the target sequence by the complex formed by the CRISPR nuclease or its variants with the guide RNA. For example, for Cas9 nuclease and its variants, the target sequence is adjacent to the PAM at the 3' end, such as 5'-NGG-3'. Based on the presence of the PAM, those skilled in the art can readily identify target sequences in the genome that can be used for targeting. Moreover, depending on the location of the PAM, the target sequence can be located on any strand of the genomic DNA molecule; the strand containing the target sequence is called the target strand. For Cas9 or its derivatives, such as Cas9 nickase, the target sequence is preferably 20 nucleotides long. The PAM sequence may vary depending on the different CRISPR nucleases or their different variants.
[0127] In some embodiments, the pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a target sequence in the genome, resulting in a nick on the target strand (e.g., within the target sequence).
[0128] In some embodiments, the PAM between the first target sequence and the second target sequence is spaced approximately 1 to approximately 300 bp, for example, 10 bp to approximately 100 bp, for example, approximately 20 bp to approximately 100 bp, approximately 20 bp to approximately 80 bp, approximately 20 bp to approximately 60 bp, or approximately 30 bp to approximately 50 bp. In some embodiments, the PAM between the first target sequence and the second target sequence may be spaced approximately 10 bp, approximately 20 bp, approximately 30 bp, approximately 40 bp, approximately 50 bp, approximately 60 bp, approximately 70 bp, approximately 80 bp, approximately 100 bp, approximately 150 bp, or approximately 300 bp.
[0129] II. Genome editing systems used for modifying the genome of organisms
[0130] This invention provides a genome editing system comprising:
[0131] i)a) An expression construct containing a CRISPR nuclease and / or a nucleotide sequence encoding the CRISPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase, or
[0132] b) A guide-editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide-editing fusion protein, wherein the guide-editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase; and,
[0133] ii) An expression construct containing at least one pegRNA and / or a nucleotide sequence encoding the at least one pegRNA, wherein the at least one pegRNA comprises at least one, at least two, at least three, at least four, at least five, at least six or more pegRNAs as described in <Guided Editing Guide RNA>, and wherein, as described in <Guided Editing Guide RNA>, the pegRNA is composed of a first pegRNA and a second pegRNA.
[0134] In some preferred embodiments, the fusion protein comprises the amino acid sequence shown in SEQ ID NO:11 (ePPEplus). In some alternative embodiments, the fusion protein in i)-b) further comprises a recombinase, preferably a Cre, Bxb1, or phiC31 recombinase, more preferably a Cre recombinase.
[0135] In some preferred embodiments, the least one pegRNA is transcribed by a complex promoter comprising a 35S enhancer, a CmYLCV promoter, and / or a U6 promoter;
[0136] Optionally, the U6 promoter is a truncated variant of the U6 promoter.
[0137] The exemplary 35S enhancer contains the sequence shown in SEQ ID NO:13, the CmYLCV promoter contains the sequence shown in SEQ ID NO:14, and the truncated U6 promoter contains the sequence shown in SEQ ID NO:15.
[0138] The first pegRNA and / or the second pegRNA further contain a tevopre sequence at the 3' end of the PBS sequence; and / or
[0139] The first pegRNA and / or the second pegRNA also contain a polyT sequence at the 3' end.
[0140] In some embodiments, the 5' end of the first pegRNA and / or the second pegRNA is linked to a first ribozyme or tRNA, the first ribozyme or tRNA being designed to cleave the first pegRNA and / or the second pegRNA at the 5' end, or to cleave a fusion of the first pegRNA and / or the second pegRNA, a tevopre sequence, and / or a polyT sequence; and / or the 3' end of the first pegRNA and / or the second pegRNA is linked to a second ribozyme or tRNA, the second ribozyme or tRNA being designed to cleave the first pegRNA and / or the second pegRNA at the 3' end, or to cleave a fusion of the first pegRNA and / or the second pegRNA, a tevopre sequence, and / or a polyT sequence.
[0141] In some implementations, the design of the tevopre (i.e., tevopreQ1) sequence can be found in James W. Nelson et al., Engineered pegRNAs improve prime editing efficiency. 2022, Nature Biotech. Volume 40, pages 402–410. An exemplary tevopre sequence is shown in SEQ ID NO:12.
[0142] In some embodiments, the polyT sequence comprises, for example, approximately 10-30 consecutive thymine (T) molecules.
[0143] In some optional embodiments, the tRNA is selected from at least one of tRNAGly, tRNAOsAsp, and tRNAZmIle;
[0144] In some optional embodiments, the ribozyme includes the HDV ribozyme.
[0145] In some specific embodiments, the 5' end of the first pegRNA and / or the second pegRNA is linked to a first ribozyme or tRNA, the first ribozyme or tRNA being designed to cleave the fusion at the 5' end of the first pegRNA and / or the second pegRNA; and / or the 3' end of the first pegRNA and / or the second pegRNA is linked to a second ribozyme or tRNA, the second ribozyme or tRNA being designed to cleave the fusion at the 3' end of the first pegRNA and / or the second pegRNA. The design of the first or second ribozyme or tRNA is within the capabilities of those skilled in the art. For example, see Gaoetal., JIPB, Apr, 2014; Vol. 56, Issue 4, 343-349. Methods for precisely processing gRNA can be found, for example, in WO2018 / 149418. In some exemplary embodiments, the tRNA includes tRNAGly (tGly for short), tRNAOsAsp (tOsAsp for short), and tRNAZmIle (tZmIle for short); the first or second ribozyme includes HDV (Hepatitis Delta Virus), which is used as a self-cutting RNA structure that can cleave RNA molecules at specific sequences.
[0146] In some specific embodiments, the at least one pegRNA comprises at least one, at least two, at least three, at least four, at least five, at least six, or more first and second pegRNAs, and / or
[0147] The first pegRNA and / or the second pegRNA in the at least one pegRNA are both in different expression constructs, or
[0148] At least two first pegRNAs, at least two second pegRNAs, or at least one first pegRNA and at least one second pegRNA are in the same expression construct, or
[0149] The first and second pegRNAs in the at least one pegRNA are both in the same expression construct.
[0150] In some specific embodiments, when the first and second pegRNAs of the at least one pegRNA are both in the same expression construct, and the total number of the first and second pegRNAs of the at least one pegRNA is 6 or less, the tRNA is tRNAGly; or,
[0151] When the first and second pegRNAs of the at least one pegRNA are both in the same expression construct, and the total number of the first and second pegRNAs of the at least one pegRNA is greater than 6, the tRNA is selected from at least two of tRNAGly, tRNAOsAsp and tRNAZmIle, preferably three.
[0152] Various scaffold sequences for gRNAs suitable for CRISPR-based genome editing (e.g., Cas9) are known in the art and can be used in the pegRNAs of this invention. In some specific embodiments, the scaffold sequence of the first or second pegRNA is shown in SEQ ID NO:7.
[0153] In some implementations, the design of the tevopre (i.e., tevopreQ1) sequence can be found in James W. Nelson et al., Engineered pegRNAs improve prime editing efficiency. 2022, Nature Biotech. Volume 40, pages 402–410. An exemplary tevopre sequence is shown in SEQ ID NO:12.
[0154] In some embodiments, the polyT sequence comprises, for example, approximately 10-30 consecutive thymine (T) molecules.
[0155] In some embodiments, the CRISPR nuclease is a Cas9 nuclease, such as SpCas9 derived from Streptococcus pyogenes. An exemplary wild-type SpCas9 contains the amino acid sequence shown in SEQ ID NO:1.
[0156] In some embodiments, the CRISPR nuclease is a CRISPR nickase. The CRISPR nickase in the fusion protein is capable of forming a nick within the target sequence on the target strand of the genomic DNA. In some embodiments, the CRISPR nickase is a Cas9 nickase.
[0157] In some embodiments, the Cas9 nickase in the fusion protein is capable of forming a nick between the -3 nucleotide (the first nucleotide at the 5' end of the PAM sequence is the +1 position) and the -4 nucleotide of the target sequence PAM.
[0158] In some embodiments, the Cas9 nuclease, such as a nickase, is a Cas9 nuclease or nickase variant capable of recognizing an altered PAM sequence. Many Cas9 nickase variants capable of recognizing altered PAM sequences are known in the art. In some embodiments, the Cas9 nuclease, such as a nickase, is a Cas9 variant that recognizes the PAM sequence 5'-NGG-3'. In some embodiments, the Cas9 nickase variant recognizing the PAM sequence 5'-NGG-3' is derived from the wild-type Cas9 (SpCas9) of *S. pyogenes*, and contains, relative to wild-type Cas9, the following amino acid substitutions: H840A, R221K, N394K, comprising the amino acid sequence shown in SEQ ID NO:2.
[0159] The nicks formed by the Cas9 nuclease described in this invention, such as the nicking enzyme, can lead to the formation of a free single strand with a 3' end (3' free single strand) and a free single strand with a 5' end (5' free single strand).
[0160] In some implementations, the CRISPR nucleases, such as Cas9 nickase and the reverse transcriptase, in the fusion protein are linked by a linker.
[0161] In some embodiments, the reverse transcriptase is a viral reverse transcriptase, such as M-MLV reverse transcriptase derived from Moloney murine leukemia virus. Further, the RNase H domain of the M-MLV reverse transcriptase is mutated or deleted. In some embodiments, the M-MLV reverse transcriptase with a mutated or deleted RNase H domain comprises the sequence shown in SEQ ID NO:5.
[0162] In some embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused to the nucleocapsid protein (NC) directly or via a linker at the N-terminus or C-terminus.
[0163] In some embodiments, the nucleocapsid protein (NC) comprises an amino acid sequence as shown in SEQ ID NO:6.
[0164] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused at the N-terminus to a nucleocapsid protein (NC) directly or via a linker.
[0165] As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids without secondary or higher structures. For example, the linker can be a flexible linker, such as SGGS, SGSETPGTSESATPES, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, etc.
[0166] In some embodiments, the CRISPR nuclease, such as a CRISPR nickase, in the fusion protein is located at the N-terminus of the reverse transcriptase. In some embodiments, the CRISPR nuclease, such as a CRISPR nickase, in the fusion protein is located at the C-terminus of the reverse transcriptase.
[0167] In some embodiments of the present invention, the CRISPR nuclease, reverse transcriptase, recombinase, or fusion protein of the present invention may further comprise one or more nuclear localization sequences (NLS). Generally, one or more NLS in the CRISPR nuclease, reverse transcriptase, or fusion protein should have sufficient strength to drive the accumulation of the CRISPR nuclease, reverse transcriptase, or fusion protein in the nucleus of the cell to achieve its base-editing function. Generally, the strength of nuclear localization activity is determined by the number, location, one or more specific NLS used, or a combination of these factors in the CRISPR nuclease, reverse transcriptase, or fusion protein. In some exemplary embodiments, the NLS comprises sequences as shown in any of SEQ ID NO: 73–76.
[0168] In some preferred embodiments, the fusion protein comprises, from N-terminus to C-terminus, a CRISPR nuclease such as a nickase, the nucleocapsid protein (NC), and the reverse transcriptase, linked by or without a linker. In some preferred embodiments, the fusion protein comprises, from N-terminus to C-terminus, a nuclear localization sequence-the CRISPR nuclease such as a nickase-linker-the nucleocapsid protein (NC)-linker-the reverse transcriptase-nuclear localization sequence.
[0169] In some embodiments, the fusion protein comprises a nuclease portion and a reverse transcriptase portion, the nuclease portion comprising the CRISPR nuclease such as CRISPR nickase and one or more NLS, and the reverse transcriptase portion comprising an RNA aptamer-binding protein sequence (e.g., an MCP protein sequence), the reverse transcriptase, one or more NLS, and optionally the nucleocapsid protein (NC).
[0170] In some embodiments, the first target sequence, the second target sequence, and / or the desired modification, such as the first exogenous nucleotide sequence, are associated with an organism, such as a plant trait, such as an agronomic trait, whereby the insertion of the desired modification, such as the first exogenous nucleotide sequence, results in the organism, such as a plant, having altered (preferably improved) traits, such as agronomic traits, relative to a wild-type organism, such as a plant.
[0171] The genome editing system of this invention can be used to perform site-specific modifications, such as site-specific insertion of exogenous nucleotide sequences, in organisms that can be non-human animals, humans, or plants, preferably plants. Suitable plants include monocotyledonous and dicotyledonous plants, for example, crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.
[0172] In order to achieve effective expression in organisms such as plants, in some embodiments of the present invention, the nucleotide sequence encoding the fusion protein is codon-optimized for the organism, such as the plant species, whose genome is to be modified.
[0173] Codon optimization refers to the modification of nucleic acid sequences to enhance expression in host cells of interest by replacing at least one codon of the natural sequence with codons that are used more frequently or most frequently in the gene in the host cell (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons while maintaining the natural amino acid sequence). Different species exhibit specific preferences for certain codons of specific amino acids. Codon preference (the difference in codon use between organisms) is often associated with the translation efficiency of messenger RNA (mRNA), which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be customized to achieve optimal gene expression in a given organism based on codon optimization. Codon utilization tables are readily available, for example, from the Codon Usage database (“Codon Usage”) available at www.kazusa.orjp / codon / . In the database, these tables can be adapted in different ways. See Nakamura Y. et al., “Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).
[0174] III. Applications of Genome Editing Systems
[0175] On the one hand, the present invention provides the application of the genome editing system of the present invention in the following (A) or (B):
[0176] (A) Modification of the genome sequence of an organism or its cells;
[0177] (B) Prepare products that modify the genome sequence of an organism or biological cell.
[0178] In some specific embodiments, the cells are microorganisms such as bacteria and fungi; animals, including mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats, and poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, with monocots such as rice, corn, wheat, sorghum, and barley, and dicots such as soybeans, peanuts, Arabidopsis thaliana, rapeseed, and cotton. In some preferred embodiments, the cells are derived from humans.
[0179] On the other hand, this invention provides the use of the genome editing system in in vivo and in vitro gene therapy, enabling the deletion, addition, upregulation, downregulation, inactivation, activation, or mutation correction of disease-related genes, thereby achieving disease prevention and / or treatment. For example, the target nucleic acid region described in this invention can be located within the protein-coding region of a disease-related gene, or, for example, within a gene expression regulatory region such as a promoter region or enhancer region, thereby enabling modification of the function or expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes modification of the disease-related gene itself (e.g., protein-coding region), as well as modification of its expression regulatory regions (e.g., promoters, enhancers, introns, etc.).
[0180] IV. Methods for modifying target sequences in the genome of cells or organisms; methods for generating genetically modified cells; genetically modified organisms.
[0181] On one hand, the present invention provides a method for producing genetically modified cells, the method comprising introducing the gene editing system of the present invention into at least one cell, thereby inserting one or more exogenous nucleotide sequences, resulting in modification of the genome sequence of the at least one cell. The modification includes substitution, deletion, and / or addition of one or more nucleotides. For example, the modification includes one or more substitutions selected from the following: C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or includes the deletion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide deletions; and / or includes the insertion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide insertions.
[0182] In another aspect, the present invention also provides genetically modified organisms comprising genetically modified cells or their progeny cells produced by the method of the present invention.
[0183] In this invention, the modification can be located anywhere in the genome, such as within a functional gene like a protein-coding gene, or in a gene expression regulatory region such as a promoter or enhancer region, thereby achieving modification of gene function or gene expression. The modification in the cell genome sequence can be detected using T7EI, PCR / RE, or sequencing methods.
[0184] In this invention, the gene editing system can be introduced into cells using various methods well known to those skilled in the art. For example, methods for introducing the gene editing system of this invention into cells include, but are not limited to: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, and other viruses), gene gun method, PEG-mediated protoplast transformation, and Agrobacterium-mediated transformation. Cells that can be gene-edited using the methods of this invention can be derived from, for example, microorganisms such as bacteria and fungi; animals, including mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, with monocots such as rice, corn, wheat, sorghum, and barley, and dicots such as soybeans, peanuts, Arabidopsis, rapeseed, and cotton. The organism may include animals, including mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots. Monocots include rice, corn, wheat, sorghum, and barley, while dicots include soybeans, peanuts, Arabidopsis thaliana, rapeseed, and cotton. In some preferred embodiments, the cells are derived from humans.
[0185] In some embodiments, the method of the present invention is performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.
[0186] In other embodiments, the method of the present invention can also be performed in vivo. For example, the cells are cells within an organism, and the system of the present invention can be introduced into the cells in vivo via, for example, a viral or Agrobacterium-mediated method.
[0187] In another aspect, the present invention also provides a method for treating a disease in a subject in need, comprising delivering an effective amount of the genome editing system of the present invention to the subject to modify a gene associated with the disease. The present invention also provides the use of the genome editing system in the preparation of a pharmaceutical composition for treating a disease in a subject in need, wherein the genome editing system is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject in need, comprising the genome editing system of the present invention and optionally a pharmaceutically acceptable vector, wherein the genome editing system is used to modify a gene associated with the disease. In some embodiments, the subject is a human being.
[0188] V. Reagent Kit
[0189] The present invention also includes a kit for use with the methods of the present invention, the kit comprising at least components of the genome editing system of the present invention. The kit may also contain reagents for introducing said genome editing system into an organism or somatic cells. The kit generally includes a label indicating the intended use and / or method of use of the kit contents. Terminology labels include any written or documented material provided on or with the kit or otherwise accompanied by the kit.
[0190] Example
[0191] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0192] Materials and Methods
[0193] 1. Plasmid construction
[0194] (1) Target design of the TKO editor: The inventors first compared the genomic DNA and coding sequence (CDS) of the target gene, selecting common coding sequences in the first half of all reliable transcripts. Specific targets were quickly obtained using webservers (rice and maize, http: / / skl.scau.edu.cn / targetdesign / ; wheat, http: / / www.rgenome.net / cas-designer / ), and paired targets with PAM-in intervals of 20-50 bp were selected. For Cas9 targets, one or two targets identical to those in the TKO were selected for a more reasonable control.
[0195] (2) Construction of epigRNA expression plasmid: First, PCR was performed to obtain DNA fragments containing protospacer, gRNA scaffold, RT template and PBS sequence. These fragments and tevopreQ1-HDV fragment were inserted into the p35C-epegRNA-ccdB vector (which contains one 35S enhancer, one CmYLCV promoter, one truncated U6 promoter, tGly, HSP18.2t, and the p35C-epegRNA-ccdB vector sequence is shown in SEQ ID NO:16; the structure of the constructed p35C-epegRNA plasmid is shown in Figure 2a).
[0196] (3) Construction of gRNA expression plasmid:
[0197] By replacing elements such as the promoter, tGly, and HDV, an empty expression vector of p35C-gRNA-ccdB was first constructed (containing one 35S enhancer, one type II CmYLCV promoter, one truncated U6 promoter, an HDV fragment, and HSP18.2t). Then, the protospacer sequence obtained after primer annealing was ligated using T4 ligase (https: / / www.neb.cn / ) into pOsU3-gRNA (targets for rice and maize), pTaU6-gRNA (targets for wheat), and p35C-gRNA-ccdB vectors, respectively (the constructed p35C-gRNA vector structure is shown in Figure 2a).
[0198] (4) Construction of multi-gene editing expression plasmids: Specific fragments containing a single epigRNA and corresponding IIS restriction enzyme sites (BsaI or BsmBI) and protective bases were amplified from the corresponding plasmids by PCR. Subsequently, these fragments were digested with the corresponding IIS enzymes to obtain DNA fragments with 5' overhangs. Then, multiple fragments were inserted into the p35C-epegRNA vector using T4 ligase.
[0199] 2. Protoplast transfection
[0200] The inventors used rice Kitaake, winter wheat KN199 and maize B73 to prepare protoplasts to construct a transient expression system.
[0201] The Cas9 and ePPEplus plasmids were co-transfected into protoplasts along with their corresponding gRNAs and epigRNAs. It is noteworthy that, due to the significant differences in molecular weight between the expression plasmids for different proteins and epigRNAs, the inventors adjusted the transformation amounts according to molar mass to ensure the same molar quantity (5 μg of plasmid expressing ePPEplus and 5 μg of plasmid expressing epigRNA). Each experimental treatment was performed in triplicate. The transfected protoplasts were cultured at 28°C for 48 h, after which they were collected for genome extraction.
[0202] 3. DNA extraction and amplicon library construction
[0203] DNA was extracted from protoplasts using the CTAB method. For identical or adjacent editing sites, the same primers were used to amplify the target fragment in the first round of PCR. Then, using the first round PCR product diluted 10-fold as a template and barcode-containing forward and reverse primers, a second round of PCR was performed to amplify the fragment containing the target site, with a size between 100-270 bp. Equal volumes of the library were then mixed and subjected to high-throughput sequencing.
[0204] 4. High-throughput amplicon sequencing and data analysis
[0205] Amplicon library sequencing was performed using the Novaseq 2×150bp platform (https: / / www.genewiz.com.cn / ). The Cas9-mediated indel generation efficiency is the ratio of reads containing insertions or deletions to the total number of reads. The desired precise editing efficiency is the ratio of reads containing precisely edited sequences to the total number of reads. For Cas9 and TKO-mediated gene knockout efficiency, a different algorithm is used. Specifically, each read is first aligned with a reference sequence and a stop codon scan is performed, classifying it into seven types: WT, SNP, SNP-STOP, Indel-3N+1, Indel-3N+2, Indel-3N, and Indel-3N-STOP. Reads containing SNP-STOP, Indel-3N+1, Indel-3N+2, and Indel-3N-STOP are defined as gene knockout reads, and their ratio to the total number of reads represents the gene knockout efficiency.
[0206] 5. Statistical Analysis
[0207] Data analysis and graphing were performed using Graphpad Prism 9.0.0 software. All bar charts and scatter plots represent the mean ± standard deviation. Significance analysis was performed using Graphpad Prism 9.0.0 software with t-tests, one-way ANOVA, and multiple comparisons. * indicates P-value less than 0.05; ** indicates P-value less than 0.01; *** indicates P-value less than 0.001; **** indicates P-value less than 0.0001; ns, not significant, indicates P-value greater than 0.05.
[0208] Example 1: Screening for efficient SCC insertions to establish TKO
[0209] In eukaryotic protein synthesis, commonly used stop codons are TAA, TAG, and TGA. Among them, TAA and TAG have high protein termination fidelity, while TGA has low fidelity and is prone to readthrough. Therefore, in this invention, the inventors chose TAA and TAG as stop codons for the SCC sequence. In addition, when using TwinPE (see Lin, Q., Jin, S., Zong, Y. et al. High-efficiency prime editing with optimized, paired pegRNAs in plants. Nat Biotechnol 39, 923–927 (2021), which is incorporated herein by reference) to insert SCC, the distance between the nicks of the two ephedra RNAs is approximately 30–50 nt, but the actual distance is not fixed. Regardless of whether the left nick is located at codon 3N, 3N+1, or 3N+2, to ensure that stop codons can be read in the SCC, the inventors included three stop codons in the SCC and filled the spaces between them with bases to adjust the GC content and length of the SCC. Furthermore, research showed that the first extension base at the 3' end of the gRNAscaffold is preferably G, so the inventors added a C base to the 3' end of the DNAFlap in the SCC. Additionally, research indicates that the recombinase recognition site has high insertion efficiency, which may be related to its secondary structure. [1] Therefore, the filling DNA sequence was simultaneously adjusted to obtain a secondary structure similar to the recombinase recognition site. Based on the above principles, 18 SCC sequences (SCC1-18, SEQ ID NO:39-56) were designed, with lengths between 13 and 33 nt (see Figures 1a and 2b).
[0210] Next, using rice protoplasts as a transient expression system, the precise insertion efficiency of 18 SCC sequences at the OsWx target site was tested. Specifically, the TwinPE method involved transfecting a plasmid expressing ePPEplus and two plasmids expressing epigRNA (p35C-epegRNA) into rice protoplasts. For example, the epigRNA sequences in the two epigRNA-expressing plasmids are shown in SEQ ID NO:19-28, respectively. The ePPEplus plasmid and the two p35C-epegRNA plasmids were then simultaneously introduced into rice protoplasts, thereby inserting the corresponding SCC sequences into the OsWx gene (the forward target sequence is CCGGAATCCTGGAAGCCGAC, and the reverse target sequence is ATGAGCTCCTCGGCGTAGTA).
[0211] The results showed that SCC1, 2, 6, 8, and 11 had high insertion efficiencies (see Figure 1b). Simultaneously, the structure of the SCCs was analyzed, and their Gibbs free energy was measured. The SCC structures are shown in Figure 2b. SCC1–4 were too short to detect their secondary structures. The remaining SCCs were similar to the recombinase recognition site, exhibiting certain stem-loop secondary structures. Two RT templates with average Gibbs free energies maintained between -0.2 and -0.05 kcal / mol / base showed high integration efficiency (Figure 1c). Subsequently, the inventors tested the integration efficiency of these five highly efficient inserted SCC sequences at six other rice target sites. The results showed that SCC11 had the highest insertion efficiency (see Figure 1d). Therefore, in this invention, the inventors used the TwinPE method to insert SCC11 into the target gene CDS to achieve efficient and precise knockout of the target gene. This method was named TKO.
[0212] Furthermore, by testing the effect of different nick distances on SCC insertion efficiency, the inventors found that the insertion efficiency of SCC was highest when the distance between the nicks at the target point was greater than 30 bp (see Figure 1e).
[0213] Example 2: Comparison of the efficiency of Cas9 and TKO gene knockout in rice and maize
[0214] Typically, in Cas9-mediated gene knockout, a type II promoter is used to express the Cas9 protein, while a type III promoter is used to express the gRNA. In this invention, the expression of epigRNA in the TKO editor uses a complex promoter, comprising a type II CmYLCV promoter, a truncated U6 promoter, and a 35S enhancer, referred to as the p35C promoter. To fairly compare gene knockout efficiency at different gene target sites, the inventors selected seven rice genes and compared the gene knockout efficiency of Cas9 and TKO in rice protoplasts (see Figure 3a), while gRNA expression was driven by type III (OsU3 or TaU6) and type II promoters, respectively (see Figure 2a). In each gene, Cas9 selected the left and right targets of the dual pegRNA (see Table 1). Experimental results showed that even when using p35C to achieve high levels of gRNA expression, TKO exhibited significantly higher gene knockout efficiency compared to Cas9, with a knockout efficiency 2.0–2.3 times that of Cas9 (see Figure 3b). Furthermore, TKO achieved precise gene knockout, while Cas9 caused approximately 30 amino acids to be added to the C-terminus of the protein through frameshift mutations (see Figure 3c). Additionally, Cas9 produced approximately 16% of 3n indels; this type of mutation alters only a few amino acids and is mostly considered ineffective knockout, while TKO produced almost no such mutations (see Figure 3d). Genetic transformation in rice callus tissue further confirmed that TKO-mediated gene knockout was both efficient and precise (see Figures 3e and 3f, where in Figure 3f, Chi represents chimeras, Bi / Bp represents biallelic mutants / byproducts, Het represents heterozygous mutants, and Ho represents homozygous mutants).
[0215] Table 1
[0216] Furthermore, the inventors tested the TKO system in maize. In nine maize genes (target sequences are shown in Table 1), the TKO-mediated gene knockout efficiency was significantly higher than Cas9, ranging from 3.0 to 8.5 times (see Figure 3g), and also more precise (see Figures 3h and 3i). These results demonstrate that the TKO method achieves gene knockout in rice and maize with higher efficiency and precision.
[0217] Example 3: Comparison of gene knockout efficiency of Cas9 and TKO in wheat
[0218] In the aforementioned embodiments, TKO has demonstrated significant efficiency advantages in rice and maize, and reduced 3n indel mutations. The inventors further tested the efficiency of TKO gene knockout in the polyploid crop wheat. In wheat, the inventors selected eight genes (TaSD1, TaQ, TaKRN2, TaGASR7, TaGW2, TaMLO, TaMTL, and TaNP1) and compared the efficiency of Cas9 and TKO gene knockout in wheat protoplasts. The results showed that, similar to the results in rice and maize, the knockout efficiency of TKO-mediated single-copy genes in wheat was significantly higher than that of Cas9, ranging from 2.0 to 10.6 times (see Figure 3j), and more precise (see Figure 3k). Similarly, TKO almost completely eliminated 3n indel types (see Figure 3l). Genetic transformation of wheat embryos showed that TKO was significantly more efficient at simultaneously knocking out all three copies of TaKRN2 than Cas9, by 5.2 times (see Figure 3m), mainly due to the elimination of the 3n indel type (see Figure 3n). This indicates that the TKO method has significant advantages in knocking out polyploid genes such as wheat, providing a powerful tool and method for wheat gene editing research.
[0219] Table 2
[0220] Example 4: Developing orthogonal TKOs to achieve efficient multi-gene knockout
[0221] Multiple gene knockout is of great significance in biological research and biobreeding applications. To avoid the overlap of dangling ends generated by reverse transcription from different target sites, which could lead to large-fragment deletions and other side effects, the inventors screened for orthogonal SCCs (see Figure 4a). In addition to the previously screened SCC6, SCC8, and SCC11, which were capable of efficient insertion, the inventors designed 16 more SCCs (SCC19–34) based on previous SCC design principles and the Gibbs free energy rule of the secondary structure of RT templates. The insertion efficiency of the 19 SCCs was tested using three rice genes (OsINV3L, OsINV3R, and OsEPSPS) (Figure 4b), and the three least efficient SCCs (SCC30, SCC32, and SCC33) were removed. Furthermore, through homology alignment, four highly similar SCCs were removed, leaving 12 SCCs (Figure 4c). Furthermore, the rice OsKRN2 gene and its neighboring genes, located only 7 kb apart, were selected to test the deletion efficiency of the intermediate fragment caused by different SCC combinations (without adding on-target paired pegRNAs). Two SCCs were removed based on the deletion efficiency of the intermediate fragment (see Figures 4d and 4e). Finally, 10 highly orthogonal SCCs were retained, and two sets of on-target pegRNAs were added for orthogonal testing. The results confirmed the extremely high degree of non-interference of orthogonal SCCs, with an average crosstalk frequency reduction of 91.3%. In most cases, the large-scale deletion frequency was below 0.1% (see Figures 4f to 4h), showing that TKO has good orthogonality when inserting different SCCs. Through SCC orthogonality optimization, at least 10 genes can be knocked out simultaneously using TKO.
[0222] To improve the efficiency of TKO-mediated multi-gene knockout, the inventors first optimized the expression mode of multiple pegRNAs: they designed three All pegRNA-In-One (APIO) multi-pegRNA expression forms to express six pegRNAs. The specific design is as follows:
[0223] peg1 and peg2 were designed to insert the recombinase recognition site m3pR at the OsINV3L target site, peg3 and peg4 were designed to insert the recombinase recognition site lox71 in OsINV3R, and peg5 and peg6 were designed to replace the three important amino acids TAP of OsEPSPS with IVS (exon variant sites).
[0224] Regarding expression formats, MPP1 uses only tRNAGly to process pegRNA, while MPP2 uses both tRNAGly and HDV to process pegRNA. To reduce sequence repetition in the multi-pegRNA expression vector, the inventors screened the tRNAs with the highest copy numbers in rice and maize, namely tRNAOsAsp and tRNAZmIle, and constructed the MPP3 expression format by combining it with tRNAGly and HDV (see Figure 5a), which was then tested in rice protoplasts.
[0225] Table 3
[0226] Experimental results showed that TKO exhibited high editing efficiency across different expression formats. The MPP2 expression format, specifically the tRNAGly and HDV-binded form, demonstrated the highest efficiency in multi-gene editing (see Figure 5b, where "Mixed single" indicates a mixture of plasmids expressing six single pegRNAs). Therefore, when the number of expressed pegRNAs is ≤6, the tRNAGly and HDV-binded expression format is preferred; when the number of pegRNAs is >6, to reduce sequence repetition in the vector, the pegRNA processing formats of tRNAGly, tRNAOsAsp, and tRNAZmIle bound to HDV are preferred. These optimization measures help improve the efficiency of TKO-mediated multi-gene knockout and provide important technical support for subsequent multi-gene editing research.
[0227] Next, the inventors constructed All gRNA In One (AGIO) and APIO vectors to test the effectiveness of TKO in multi-gene knockout. They designed expression plasmids that simultaneously knocked out 2, 3, and 4 genes, respectively, and compared the effects of TKO, the combination of Cas9 and pOsU3-gRNA, and the combination of Cas9 and p35C-gRNA in multi-gene knockout.
[0228] Experimental results showed that, compared to Cas9, TKO exhibited significantly higher efficiency when simultaneously knocking out 2, 3, and 4 genes (see Figures 5c and 5d). Genetic transformation of rice callus revealed that using orthogonal TKO to simultaneously knock out 4 genes was more efficient than the Cas9 system (see Figures 5e and 5f). This indicates that TKO has significant advantages in multi-gene knockout, providing an efficient tool and method for multi-gene editing research.
[0229] Example 5: Establishing a high-efficiency multi-gene, multi-type genome editing TRIM platform
[0230] Multi-gene, multi-type genome modification plays a crucial role in biological research and biobreeding applications. By simultaneously editing multiple genes using different methods, such as knockout, insertion, and substitution, precise regulation and improvement of multiple target traits can be achieved. PE (Progeny Genome Optimizer) possesses the functions of base substitution, small fragment insertion, deletion, and substitution, while current gene knockout primarily relies on Cas9-mediated methods. When multi-gene, multi-type editing is required, iterative editing using different editors or editing different plant bodies followed by hybridization to aggregate mutated alleles is necessary. The TKO (Transgenic Koine) method eliminates the Cas9 dependency of gene knockout, enabling multi-gene, multi-type gene editing—covering gene knockout, base substitution, small fragment insertion, deletion, and substitution—to be achieved with only a single PE protein, i.e., an all-in-one platform, named TRIM (TKO editor-enabled gene Rupture and development of Integrated Multi-type genome modification systems).
[0231] Table 4
[0232] The inventors constructed three APIO multi-gene editing vectors, OsAssembly1, OsAssembly2, and OsAssembly3 (see Figure 6), and used the TRIM1 platform to perform multi-site, multi-type genome modifications (see Figure 7a). Experimental results showed that, using only the ePPEplus protein, efficient knockout of OsKRN2, OsGn1a, OsGW2, OsGS3, and OsSD1 genes, creation of uORF in the OsDLT gene, amino acid substitution in the OsEPSPS gene, deletion of the OsRDD1 miRNA recognition site, insertion of a small fragment in OsINV8L, insertion of an HSE response element in the OsGIF1 promoter, disruption of the miR156 recognition site in OsIPA1 through multiple base substitutions, and insertion of an enhancer in the OsDREB1C promoter (see Figure 7b). Genetic transformation of OsAssembly2 resulted in up to 23.1% of regenerated T0 rice plants exhibiting OsKRN2 knockout, while homozygous mutations occurred at the other three sites (see Fig. 7c, Fig. 7d). In maize and wheat, TRIM1 mediated similar results in multi-site, multi-type genomic modifications (see Fig. 7e, Fig. 7f).
[0233] To expand editing capabilities, the inventors developed a second TRIM platform, TRIM2, using a protein fused with PE and Cre recombinase. TRIM2 amplifies the editing capabilities of TRIM1 for gene knockout, single base substitution and small fragment substitution, insertion, substitution, deletion, replication, and inversion. It allows for chromosome-level manipulation of DNA fragments at the kb and even Mb levels using the Cre-lox system (see Figure 7g). The inventors also utilized the LoxAR2 Cre recombinase recognition site at the OsINV8 site in OsAssembly2 and OsAssembly3 as a donor, enabling site-specific insertion of fragments up to 4.9 kb in the TRIM2 system, in addition to gene knockout, single base substitution, and small fragment editing (see Figure 7h).
[0234] The TRIM1 and TRIM2 platforms provide efficient tools and methods for multi-gene, multi-type editing, which will drive the development of biological research and bio-breeding applications.
[0235] SEQ ID NO:1, Wild-type spCas9 amino acid sequence:
[0236] SEQ ID NO:2, nSpCas9 (H840A, R221K, N394K) amino acid sequence:
[0237] SEQ ID NO:5, reverse transcriptase M-MLV-RT amino acid sequence:
[0238] SEQ ID NO:6, nucleocapsid protein (NC) sequence:
[0239] SEQ ID NO:7, pegRNA backbone sequence:
[0240] SEQ ID NO:8, Connector sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGS
[0241] SEQ ID NO:9, Connector sequence: SGSETPGTSESATPES
[0242] SEQ ID NO:10, Connector Sequence: SGGS
[0243] SEQ ID NO:11, ePPEplus sequence:
[0244] SEQ ID NO:12, tevopreQ1 sequence:
[0245] SEQ ID NO:13, 35S enhancer subsequence:
[0246] SEQ ID NO:14, CmYLCV promoter
[0247] SEQ ID NO:15, truncated U6 starter
[0248] SEQ ID NO:16, p35C-epegRNA-ccdB vector sequence:
[0249] en35S is represented by an underscore; CmYLCV by a double underscore; U6 by a wavy line; ccdB by a dashed underscore; and Terminator by a dotted underscore.
[0250] SEQ ID NO:18, MPP1-3 backbone sequences:
[0251] en35S is represented by an underscore; CmYLCV is represented by a double underscore; U6 is represented by a wavy line; Terminator is represented by a dotted underscore.
[0252] SEQ ID NO:19, the pegRNA sequence used for inserting SCC1 (forward direction, taking the OsWx gene as an example):
[0253] SEQ ID NO:20, the pegRNA sequence used for inserting SCC1 (reverse, taking the OsWx gene as an example):
[0254] SEQ ID NO:21, the pegRNA sequence used for inserting SCC2 (forward direction, taking the OsWx gene as an example):
[0255] SEQ ID NO:22, the pegRNA sequence used for inserting SCC2 (in reverse, taking the OsWx gene as an example):
[0256] SEQ ID NO:23, the pegRNA sequence used for inserting SCC6 (forward direction, taking the OsWx gene as an example):
[0257] SEQ ID NO:24, the pegRNA sequence used for inserting SCC6 (in reverse, taking the OsWx gene as an example):
[0258] SEQ ID NO:25, the pegRNA sequence used for inserting SCC8 (forward direction, taking the OsWx gene as an example):
[0259] SEQ ID NO:26, the pegRNA sequence used for inserting SCC8 (in reverse, taking the OsWx gene as an example):
[0260] SEQ ID NO:27, the pegRNA sequence used for inserting SCC11 (forward direction, taking the OsWx gene as an example):
[0261] SEQ ID NO:28, the pegRNA sequence used for inserting SCC11 (in reverse, taking the OsWx gene as an example):
[0262] SEQ ID NO:29, peg1 sequence:
[0263] SEQ ID NO:30, peg2 sequence:
[0264] SEQ ID NO:31, peg3 sequence:
[0265] SEQ ID NO:32, peg4 sequence:
[0266] SEQ ID NO:33, peg5 sequence:
[0267] SEQ ID NO:34, peg6 sequence:
[0268] SEQ ID NO:35, tRNAGly sequence:
[0269] SEQ ID NO:36, tRNAOsAsp sequence:
[0270] SEQ ID NO:37, tRNAZmIle sequence:
[0271] SEQ ID NO:38, HDV sequence:
[0272] SEQ ID NO:39, SCC1 sequence: GTAACTAGCTAAC
[0273] SEQ ID NO:40, SCC2 sequence: GTAACTAGGTAAC
[0274] SEQ ID NO:41, SCC3 sequence: GTAACCTAGCGTAAC
[0275] SEQ ID NO:42, SCC4 sequence: GTAACGTAGCCTAAC
[0276] SEQ ID NO:43, SCC5 sequence: GTAACGGCTAGCTAACGCC
[0277] SEQ ID NO:44, SCC6 sequence: GTAACCGGTAGGTCCTAAC
[0278] SEQ ID NO:45, SCC7 sequence: GTAACGGCTAGCCGTTAAC
[0279] SEQ ID NO:46, SCC8 sequence: GTAACTGCACGTAGTCACTAAC
[0280] SEQ ID NO:47, SCC9 sequence: GTAACGCCTACTAGGCGTTAAC
[0281] SEQ ID NO:48, SCC10 sequence: GTAACTCGTAGTAGTGCCTAAC
[0282] SEQ ID NO:49, SCC11 sequence: GTAACAGTCGCCTAGCATCGAGCTAAC
[0283] SEQ ID NO:50, SCC12 sequence: GTAAAGCTCGCCTAGCATCGAGCTAAC
[0284] SEQ ID NO:51, SCC13 sequence: GTAAAGCTCGCCTAGCAGCGAGCTAAC
[0285] SEQ ID NO:52, SCC14 sequence: TAGCAGTCGCTTAACATCGAGTTAGCTAC
[0286] SEQ ID NO:53, SCC15 sequence: GTAAGTCTGGTACATAGGGTACGGACGTAAC
[0287] SEQ ID NO:54, SCC16 sequence: GTAAGTCTGGTACATAGGGTACGGACTTAAC
[0288] SEQ ID NO:55, SCC17 sequence: GTAAGTCTGGTACATAGGACACGGACTTAAC
[0289] SEQ ID NO:56, SCC18 sequence: TAAGTCTGCACGTACATAGGACACGGACTTAAC
[0290] SEQ ID NO:57, SCC19 sequence: GTAGCTAACGCGTAAGCAC
[0291] SEQ ID NO:58, SCC20 sequence: GTAGGGTAACATGCTAGAC
[0292] SEQ ID NO:59, SCC21 sequence: GTAGGCTAAGCCAGTAGCGACC
[0293] SEQ ID NO:60, SCC22 sequence: GTAACGCACTAGGCGTGTAAGCTAC
[0294] SEQ ID NO:61, SCC23 sequence: GTAGCCTAGCGACGTAAGACGCGAC
[0295] SEQ ID NO:62, SCC24 sequence: GTAGAAGGTAACGCTTCGTAGGGAC
[0296] SEQ ID NO:63, SCC25 sequence: GTCCTAAGGTAAGAGCCTAGCAGCC
[0297] SEQ ID NO:64, SCC26 sequence: TAGCTGCCCGTAAGAGTCCGTAGGCGAC
[0298] SEQ ID NO:65, SCC27 sequence: TAGCCTCGTAACCAGTGCGTAACAGCTC
[0299] SEQ ID NO:66, SCC28 sequence: TAGCCTCGTAACCAGTGCGTAACAGTCC
[0300] SEQ ID NO: 67, SCC29 sequence: TAGCGCTCAGGTAAGGAGTACGTAGGCAC
[0301] SEQ ID NO:68, SCC30 sequence: TAGGCGTGTAAGCGAGCAGTCGTAAGAGC
[0302] SEQ ID NO: 69, SCC31 sequence: TAGCAGCGTGATAACGAGATAGCCGCGTC
[0303] SEQ ID NO:70, SCC32 sequence: TAGGACGCTCGTAGGCATGTAACCGTCGGCC
[0304] SEQ ID NO:71, SCC33 sequence: TAGCTCGCTAATCGCCGACGTCTAGAGCACC
[0305] SEQ ID NO:72, SCC34 sequence: TAAGTCACCCACGTACATAGGACGTCAGCAGACTTAAC
[0306] SEQ ID NO:73, NLS SV40 Sequence: PKKKRKV
[0307] SEQ ID NO:74, NLS c-Myc Sequence: PAAKRVKLD
[0308] SEQ ID NO:75, bpNLS SV40 Sequence: KRTADGSEFESPKKKRKV
[0309] SEQ ID NO:76, vbpNLS SV40 Sequence: KRTADSQHSTPPKTKRKV
[0310] References:
[0311] [1]Koeppel,J.,Weller,J.,Peets,E.M.et al.Prediction of prime editing insertion efficiencies using sequence features and DNArepair determinants.Nat Biotechnol 41,1446–1456(2023)。
Claims
1. A guide editing RNA (pegRNA) comprising a first pegRNA and a second pegRNA, wherein the first pegRNA or the second pegRNA includes a guide sequence, a scaffold sequence, a reverse transcription template (RT) sequence, and a primer binding site (PBS) sequence. The scaffold sequence of the first pegRNA can be complexed with CRISPR nuclease and create a nick in the first target sequence of the sense strand of the target double-stranded DNA sequence, and the reverse transcription template (RT) sequence of the first pegRNA sequence has a stop codon cluster (SCC) sequence. The SCC contains at least one, at least two, or at least three stop codons; Optionally, the stop codon is selected from TAA, TAG, and TGA, preferably from TAA and TAG; The reverse transcription template (RT) sequence of the second pegRNA is partially or completely complementary to the RT sequence of the first pegRNA, preferably with at least 13 bp to about 40 bp overlap. The second pegRNA can complex with a CRISPR nuclease and create a nick in the second target sequence of the antisense strand of the target double-stranded DNA sequence.
2. The guide editing RNA according to claim 1, wherein, The sequence length of the SCC is 13–38 nt.
3. The guide editing RNA according to claim 1 or 2, wherein, The SCC includes at least one of the following features (a) to (b): (a) The RNA secondary structure of the SCC includes a stem-loop structure; (b) The Gibbs free energy of the RNA secondary structure of the SCC is in the range of -0.2 to -0.05 kcal / mol / base.
4. The guide editing RNA according to any one of claims 1 to 3, wherein, The 3' end of the SCC is a base C.
5. The guide editing RNA according to any one of claims 1 to 4, wherein, The SCC comprises a sequence as shown in any one of SEQ ID NO:39-72.
6. The guide editing RNA according to any one of claims 1 to 5, wherein the interval between the cuts of the first target sequence and the second target sequence is not less than 20 bp, preferably not less than 30 bp, and more preferably 30 bp to 100 bp.
7. The guide editing RNA according to any one of claims 1 to 6, wherein, The scaffold sequence of the first or second pegRNA is shown in SEQ ID NO:
7.
8. A genome editing system comprising: i)a) An expression construct containing a CRISPR nuclease and / or a nucleotide sequence encoding the CRISPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase, or b) A guide-editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide-editing fusion protein, wherein the guide-editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase; and, ii) An expression construct containing at least one pegRNA and / or a nucleotide sequence encoding the at least one pegRNA, wherein the at least one pegRNA comprises at least one, at least two, at least three, at least four, at least five, at least six or more pegRNAs as described in any one of claims 1 to 7, and wherein the pegRNA as described in any one of claims 1 to 7 is composed of a first pegRNA and a second pegRNA.
9. The genome editing system according to claim 8, wherein, The fusion protein described in i)-b) further comprises a recombinase, preferably Cre, Bxb1, or phiC31 recombinase, and more preferably Cre.
10. The genome editing system according to claim 8 or 9, wherein, The at least one pegRNA is transcribed by a complex promoter comprising a 35S enhancer, a CmYLCV promoter, and / or a U6 promoter; Optionally, the U6 promoter is a truncated variant of the U6 promoter.
11. The genome editing system according to any one of claims 8 to 10, wherein, The first pegRNA and / or the second pegRNA in the at least one pegRNA are both in different expression constructs, or At least two first pegRNAs, at least two second pegRNAs, or at least one first pegRNA and at least one second pegRNA are in the same expression construct, or The first and second pegRNAs in the at least one pegRNA are both in the same expression construct.
12. The genome editing system according to any one of claims 8 to 11, wherein, The first pegRNA and / or the second pegRNA also contain a tevopre sequence at the 3' end of the PBS sequence; and / or The first pegRNA and / or the second pegRNA also contain a polyT sequence at the 3' end.
13. The genome editing system according to any one of claims 8 to 12, wherein, The 5' end of the first pegRNA and / or the second pegRNA is linked to a first ribozyme or tRNA, the first ribozyme or tRNA being designed to cleave the first pegRNA and / or the second pegRNA at the 5' end, or to cleave a fusion of the first pegRNA and / or the second pegRNA, a tevopre sequence, and / or a polyT sequence; and / or the 3' end of the first pegRNA and / or the second pegRNA is linked to a second ribozyme or tRNA, the second ribozyme or tRNA being designed to cleave the first pegRNA and / or the second pegRNA at the 3' end, or to cleave a fusion of the first pegRNA and / or the second pegRNA, a tevopre sequence, and / or a polyT sequence; Optionally, the tRNA is selected from at least one of tRNAGly, tRNAOsAsp, and tRNAZmIle; Optionally, the ribozyme includes the HDV ribozyme.
14. The genome editing system according to claim 13, wherein, When the first and second pegRNAs of the at least one pegRNA are both in the same expression construct, and the total number of the first and second pegRNAs of the at least one pegRNA is 6 or less, the tRNA is tRNAGly; or, When the first and second pegRNAs in the at least one pegRNA are both in the same expression construct, and the total number of the first and second pegRNAs in the at least one pegRNA is greater than 6, the tRNA is selected from at least two of tRNAGly, tRNAOsAsp and tRNAZmIle, preferably three.
15. The genome editing system according to any one of claims 8 to 14, wherein, The first and second target sequences in the pegRNA are associated with plant traits such as agronomic traits, thereby causing the plant to have altered (preferably improved) traits, such as agronomic traits, relative to the wild type plant via a gene editing system.
16. The genome editing system according to any one of claims 8 to 15, wherein, The at least one pegRNA is capable of forming a complex with the CRISPR nuclease or fusion protein and targeting the CRISPR nuclease or fusion protein to a target sequence in the genome, resulting in a cut on the target strand within the target sequence.
17. The genome editing system according to any one of claims 8 to 16, wherein, The CRISPR nuclease is the Cas9 nuclease or a variant thereof.
18. The genome editing system according to any one of claims 8 to 17, wherein, The CRISPR nuclease is a CRISPR nicking enzyme, such as the Cas9 nicking enzyme or a variant thereof, for example, the Cas9 nicking enzyme or a variant thereof contains a sequence selected from the sequence shown in SEQ ID NO:1 or 2; Optionally, the CRISPR nuclease, such as Cas9 nickase, and the reverse transcriptase are linked by a adapter.
19. The genome editing system according to any one of claims 8 to 18, wherein the reverse transcriptase is M-MLV reverse transcriptase or a functional variant thereof; Optionally, the RNase H domain of the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is deleted, and it contains the sequence shown in SEQ ID NO:
5.
20. The genome editing system according to any one of claims 8 to 19, wherein, Reverse transcriptases such as M-MLV reverse transcriptase or their functional variants are fused to nucleocapsid proteins (NC) directly or via linkers at the N-terminus or C-terminus.
21. The genome editing system according to claim 20, wherein, The nucleocapsid protein (NC) contains the amino acid sequence shown in SEQ ID NO:
6.
22. The genome editing system according to any one of claims 8 to 21, wherein, The CRISPR nuclease described in i)-b) such as the CRISPR nickase is fused to the N-terminus of the reverse transcriptase; Optionally, the fusion protein described in i)-b) comprises the amino acid sequence shown in SEQ ID NO:
11.
23. The genome editing system according to any one of claims 8 to 22, wherein the genome is derived from microorganisms, animals, or plants; The microorganisms include bacteria and fungi; The animals include mammals and poultry; The plants include monocotyledons and dicotyledons; Optionally, the mammals include humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; Optionally, the poultry includes chickens, ducks, and geese; Optionally, the monocotyledonous plants include rice, corn, wheat, sorghum, and barley, and the dicotyledonous plants include soybean, peanut, Arabidopsis thaliana, cotton, and rapeseed.
24. The use of the genome editing system according to any one of claims 8 to 23 in either (A) or (B): (A) Editing of the genome sequence of an organism or its cells; (B) Prepare products that modify the genome sequence of an organism or biological cell.
25. A kit comprising the genome editing system according to any one of claims 8 to 24.
26. A method for producing a genetically modified cell or organism, the method comprising introducing a genome editing system as described in any one of claims 8 to 24 into at least one of the cells or organisms, thereby resulting in modification of the genome sequence of the at least one cell or organism, such as the modification comprising substitution, deletion and / or addition of one or more nucleotides.
27. The method according to claim 26, wherein, The method further includes screening organisms with desired exogenous nucleotide sequence insertions from the at least one organism.
28. The method according to claim 26 or 27, wherein, The components of the genome editing system are simultaneously introduced into the organism.