Cre-lox system, genome editing system comprising same, and use thereof

By optimizing the Cre-Lox system and designing a PAM-free guided editing system, the problems of insufficient editing efficiency and reversible recombination activity in the existing Cre-Lox system have been solved, enabling efficient and precise manipulation of large DNA fragments in rice, maize, and wheat, and expanding the types and scope of editing.

WO2026012437A1PCT designated stage Publication Date: 2026-01-15INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/107933
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2025-07-10
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing genome editing technologies struggle to achieve efficient and precise manipulation of large DNA fragments, especially in rice, maize, and wheat. The Cre-Lox system suffers from insufficient editing efficiency and issues with reversible recombination activity and residual recombination sites, limiting its application scope and editing types.

Method used

A high-throughput screening platform for recombination sites was developed. Cre proteins were optimized using artificial intelligence and combined with different types of Lox-pegRNAs to improve the recombination efficiency of the Cre-Lox system. A PAM-free guided editing system was designed to achieve targeted and precise deletion, replacement, inversion, and translocation of large DNA fragments. This was further optimized into the RePCE system to achieve completely precise and traceless manipulation of large DNA fragments.

Benefits of technology

It significantly improved the integration efficiency of large DNA fragments in rice, maize, and wheat, enabling efficient and precise genome editing. It solved the problems of editing efficiency and reversible recombination activity in the Cre-Lox system, and expanded the genetic manipulation capabilities and target range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025107933_15012026_PF_FP_ABST
    Figure CN2025107933_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A Cre-Lox system, a genome editing system comprising same, and a use thereof. The Cre-Lox system comprises a recombinase and a recognition site for the recombinase. The recombinase is a mutant of a Cre recombinase, and the recognition site is a mutant of a Lox site. The genome editing system comprises the mutant of the Cre recombinase in the Cre-Lox system. The Cre-Lox system has high recombination activity and reduced reversible recombination activity. The genome editing system improves the DNA integration efficiency, implements targeted deletion, replacement, inversion and translocation of DNA, and can operate without PAM restrictions, achieving completely precise and scarless manipulation of large DNA fragments.
Need to check novelty before this filing date? Find Prior Art

Description

The Cre-Lox system, genome editing systems incorporating the Cre-Lox system, and their applications

[0001] Priority and related applications

[0002] This invention claims priority to Chinese Patent Application No. 202410931217.8, filed on July 11, 2024, entitled "Cre-Lox System, Genome Editing System Containing Cre-Lox System and Application Thereto", the entire contents of which, including the appendices, are incorporated herein by reference. Technical Field

[0003] This invention belongs to the field of genetic engineering. Specifically, this invention relates to the Cre-Lox system, a genome editing system incorporating the Cre-Lox system, and its applications. Background Technology

[0004] Rapid population growth, frequent extreme weather events, and shrinking arable land pose comprehensive challenges to current food security, creating an urgent need to cultivate crops with higher yields and superior quality. However, conventional methods based on backcrossing and mutation breeding fall far short of the demands in terms of both speed and quality of germplasm development. Genome editing technology provides an efficient and feasible solution for the development of rapid molecular breeding techniques. By manipulating target genes genetically, including DNA deletion, insertion, and base substitution, genetic information can be directly altered to rapidly obtain new germplasm with the desired superior traits. For a long time, researchers have focused on simply knocking in or out single genes and altering one or a few bases to improve crop traits. However, in recent years, with the deepening of genomics research, numerous studies have shown that genomic structural variations, including large-segment DNA insertions, deletions, inversions, and translocations, have a significant impact on phenotypic shaping. Compared with single-gene variations, structural variations can cause large-scale perturbations in cis-regulatory regions or directly change gene copy numbers, thereby causing quantitative changes in gene expression. In agricultural production, structural variation has been proven to have a significant impact on key agronomic traits, including flowering time, fruit size, and stress resistance. For example, a 13Mb inversion, Inv4m, in maize is closely related to adaptability to high-altitude stress; the stably heritable wheat-rye 1RS.1BL translocation line carries multiple disease resistance genes, planthopper resistance genes, and genes related to high yield and adaptability, significantly improving wheat's disease resistance and yield, playing a crucial role in ensuring wheat production and food security in my country and the world. Furthermore, certain chromosomal rearrangement events can greatly hinder the breeding process. For instance, some inversions reduce the frequency of recombination and exchange between homologous chromosomes, thus significantly reducing breeding efficiency; some reverse translocations can cause male or female semi-sterility in plants, increasing the difficulty of breeding. Therefore, the development of large-scale DNA and even chromosome rearrangement technologies will promote the research and development of new breeding technologies, such as precise repair of certain inversions and translocations to restore gene exchange frequencies or fertility; precise creation of inversions or translocations to break genetic linkage groups and manipulate genetic associations between genes to achieve genetic linkage between beneficial traits, thereby greatly improving the speed and effectiveness of breeding; and fixing certain gene flows through precise chromosome rearrangement to fix superior traits. Furthermore, currently, artificial chromosome synthesis is limited to bacteria and yeast. The synthesis of eukaryotic chromosomes is more difficult and limited by biotransformation, making it impossible to deliver synthetic chromosomes to eukaryotic cells. The development and application of precise chromosome rearrangement tools will provide new ideas for the reconstruction of artificial chromosomes in vivo. In summary, the development of large-scale DNA and even chromosome precise genetic manipulation technologies can greatly promote the development of molecular agricultural breeding and plant synthetic biology, representing the next important direction for genome editing breeding and an effective genetic manipulation tool for achieving food security.

[0005] The latest genome editing technologies are primarily based on the CRISPR-Cas9 system, which originates from prokaryotes. This system uses guide RNA (gRNA) components to target genomic DNA, and the Cas9 components perform sequence-specific cleavage after recognizing the PAM (Protospacer Adjacent Motif) sequence. Using an engineered CRISPR-Cas9 system as a base editor, a base editor is developed by fusing a deaminase component; a prime editor is developed by fusing a reverse transcriptase (M-MLV) component and combining it with engineered gRNA (prime editing guide RNA, pegRNA). Base editors can achieve precise replacement of single bases, while prime editors can achieve precise editing of small DNA fragments in any form. Traditional methods for manipulating large DNA fragments rely on the generation of DNA double-strand breaks (DSBs). DSBs, under the action of homologous end joining (NHEJ), the main DNA repair mechanism in cells, can achieve large DNA fragment editing. However, due to the randomness and error-proneness of the NHEJ repair method, the resulting mutations are uncontrollable, leading to random and imprecise DNA insertions and deletions (Indels). This imprecision is unsuitable for certain applications and can also result in unexpected large-segment DNA deletions. Furthermore, the generation of DSBs often damages cells. Therefore, developing new tools capable of efficient and precise manipulation of large DNA segments remains one of the bottlenecks in the field of genome editing.

[0006] Numerous studies both domestically and internationally have focused on improving the scale of manipulation in plant genome editing, developing various methods including the NHEJ (Non-Homologous End Joining) strategy, the HDR (Homology-Directed Repair) strategy, and guided editing strategies. However, these methods all suffer from problems such as inaccurate editing, low editing efficiency, small editing scale, and insufficient editing types, failing to meet the demands. Site-specific recombinase systems do not rely on endogenous DNA repair mechanisms or high-energy cofactors; they require only a single component to complete recombination between DNA molecules, thus being considered to have the potential for precise editing of large DNA fragments (including insertion, deletion, inversion, and translocation). However, because these systems can only recognize recombination sites (RS) of specific sequences, they lack programmability and are therefore not widely used. In the past two years, researchers have coupled site-specific recombinase systems with guided editing technology, enabling precise manipulation of large DNA fragments that do not rely on DSB by giving them programmability. Specifically, the recognition site (recombination site) of the recombinase is first introduced into the target genome through a dual-guided editor. The recombination site can recruit site-specific recombinase and insert donor DNA into the target site, achieving precise integration of large DNA fragments. These include PASSIGE, developed in human cells, which achieves highly efficient and precise insertion of up to 5.6 kb of DNA without DSB generation by combining the efficient dual-pegRNA-based guided editing system Twin PE with the serine recombinase Bxb1. However, this method integrates the entire plasmid containing the exogenous donor DNA and the vector backbone, and leaves two recombination sites on the genome after editing. The PrimeRoot system in plants uses the efficient dual-pegRNA-based guided editing system dual-ePPE combined with the tyrosine recombinase Cre, also achieving precise insertion of up to 11.1 kb of DNA without DSB generation. This method can also achieve precise insertion without exogenous donor backbone DNA (WO / 2023 / 227050 and Sun, Chao et al. “Precise integration of large DNA sequences in plant genomes using PrimeRoot editors.” Nature biotechnology). vol.42,2(2024):316-327.doi:10.1038 / s41587-023-01769-w). These two systems, using the same principle, have achieved precise manipulation of large DNA fragments in plants and animals without relying on DSB, and have demonstrated the superiority of such systems in their ability to manipulate large DNA fragments compared to other genetic engineering methods.However, the PrimeRoot system in plants still suffers from limitations in editing efficiency and scale, as well as in the variety of editing types. Furthermore, it is constrained by the limitations of PAM (Polymer Algorithm) and the presence of residual recombination sites on the genome after editing. Specifically, the Cre-Lox system suffers from insufficient reversible recombination activity and recombination capacity, resulting in insufficient editing efficiency and preventing its application to other editing types at larger scales. Highly efficient guided editing systems based on dual pegRNAs require two target sites spaced a certain distance (typically 20-100 bp), and the PAM on the genome must have an NGG target and a CCN target, thus significantly limiting the system's targetable range on the genome. The inherent site-specific recombination characteristics of this system leave two recombination site sequences on the genome after editing, further hindering its applicability in many scenarios. Therefore, optimizing and improving the editing efficiency and scale of the PrimeRoot system, expanding its genetic manipulation capabilities and target range, and achieving efficient, precise, and traceless manipulation of any type of large DNA fragment will lay a solid technical foundation for molecular agricultural breeding and plant synthetic biology. Summary of the Invention

[0007] The problem the invention aims to solve

[0008] To address the aforementioned issues, this invention develops a high-throughput screening platform for recombination sites, obtaining combinations of recombination site variants that maintain high recombination activity while significantly reducing reversible recombination activity. Artificial intelligence methods are used to optimize the Cre protein, greatly improving its recombination efficiency. An integrated and optimized Cre-Lox system is developed, and a PCE system is created by designing different types of Lox-pegRNAs, significantly improving the integration efficiency of large DNA fragments in rice, maize, and wheat. This system is further used to achieve targeted and precise deletion, replacement, inversion, and translocation manipulation of large DNA fragments. A guided editing system without PAM limitations, capable of efficient and precise insertion of short DNA fragments, is developed based on a dual-pegRNA strategy and applied to PCE. Furthermore, the RePCE system is developed based on this, achieving completely precise and traceless manipulation of large DNA fragments.

[0009] Solution for solving the problem

[0010] In a first aspect of the invention, a Cre-Lox system is provided, wherein the Cre-Lox system includes a recombinase and a recognition site (RS) of the recombinase;

[0011] The recombinase is a mutant of Cre recombinase, and the mutant of Cre recombinase is selected from any one of the following groups (a1)-(a3):

[0012] (a1) The amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has R72V, R72I, R72L, Q94M, R101L, D141Q, R146I, R159L, R159M, R159V, R159W, K211I, R243D, T268I, A296I, G313F, G313M, G313W, S257P, Q35D, E39 Mutations of P, H40E, W63P, K86E, Q94A, N111P, M117L, Q156R, R173P, I197V, I225L, S226E, R241P, H269S, I272L, A275P, K276P, Q281G, M299L, T316R, D329P, T332A or any combination thereof;

[0013] (a2) is a polypeptide having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the amino acid sequences shown in (a1);

[0014] (a3) A polypeptide with an amino acid sequence as shown in (a1) or (a2) having one or more amino acids added or deleted at at least one end of the N-terminus and C-terminus.

[0015] In some embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the mutations shown in 1) to 69) below:

[0016] 66)D141Q, H40E, A275P, K276P, Q281G; 1) R72V; 2) R72I; 3) R72L; 4) Q94M; 5) R101L; 6) D141Q; 7) R146I; 8) R159L; 9) R159M; 10) R159V; 11 )R159W; 12) K211I; 13) R243D; 14) T268I; 15) A296I; 16) G313F; 17) G313M; 18) G313W; 19) S257P; 20) Q35D; 21) E39P; 22) H40E; 23) W63P; 2 4)K86E; 25) Q94A; 26) N111P; 27) M117L; 28) Q156R; 29) R173P; 30) I197V; 31) I225L; 32) S226E; 33) R241P; 34) H269S; 35) I272L; 36) A275 P; 37) K276P; 38) Q281G; 39) M299L; 40) T316R; 41) D329P; 42) T332A; 43) D141Q, R243D; 44) D141Q, K276P; 45) D141Q, G313F; 46) R243D, K2 76P; 47) Q94M, D141Q; 48) Q94M, R243D; 49) Q94M, K276P; 50) R159M, R243D; 51) R159M, K276P; 52) G313F, K276P; 53) E39P, H40E; 54) S226 E, R241P; 55) I272L, A275P; 56) K276P, Q281G; 57) M299L, T332A; 58) D141Q, R243D, K276P; 59) I225L, S226E, R241P; 60) I272L, A275P, K2 76P; 61) A275P, K276P, Q281G; 62) Q156R, S226E, R241P; 63) Q156R, I225L, S226E, R241P; 64) I272L, A275P, K276P, Q281G; 65) H40E, A275 P, K276P, Q281G; 67) D141Q, I272L, A275P, K276P, Q281G; 68) R243D, I272L, A275P, K276P, Q281G; 69) R243D, H40E, A275P, K276P, Q281G.

[0017] In some preferred embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations:

[0018] 6)D141Q; 13)R243D; 22)H40E; 33)R241P; 34)H269S; 35)I272L; 36)A275P; 37)K276P; 38)Q281G; 42)T332A.

[0019] In some preferred embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations:

[0020] 66)D141Q, H40E, A275P, K276P, Q281G; 43) D141Q, R243D; 44) D141Q, K276P; 45)D141Q, G313F; 46) R243D, K276P; 47) Q94M, D141Q; 48) Q94M, R243D; 49) Q 94M, K276P; 50) R159M, R243D; 51) R159M, K276P; 52) G313F, K276P; 53) E39P , H40E; 55) I272L, A275P; 56) K276P, Q281G; 57) M299L, T332A; 58) D141Q, R24 3D, K276P; 59) I225L, S226E, R241P; 60) I272L, A275P, K276P; 61) A275P, K2 76P, Q281G; 63) Q156R, I225L, S226E, R241P; 64) I272L, A275P, K276P, Q281G ;65)H40E, A275P, K276P, Q281G; 67)D141Q, I272L, A275P, K276P, Q281G; 68 )R243D, I272L, A275P, K276P, Q281G; 69) R243D, H40E, A275P, K276P, Q281G.

[0021] In some preferred embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations:

[0022] 66)D141Q, H40E, A275P, K276P, Q281G; 67)D141Q, I272L, A275P, K276P, Q 281G; 65) H40E, A275P, K276P, Q281G; 46) R243D, K276P; 44) D141Q, K276P.

[0023] In some embodiments, the recognition site (RS) of the recombinase is selected from any one of the following or any combination thereof:

[0024] (a) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:9, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:9;

[0025] (b) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:10, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:10;

[0026] (c) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:11, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:11;

[0027] (d) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:12, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:12;

[0028] (e) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:13, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:13;

[0029] (f) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:14, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:14;

[0030] (g) Contains a nucleotide sequence as shown in SEQ ID NO:15, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:15;

[0031] (h) Contains a nucleotide sequence as shown in SEQ ID NO:16, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:16;

[0032] (i) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:17, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:17;

[0033] (j) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:3, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:3;

[0034] (k) Contains a nucleotide sequence as shown in SEQ ID NO:4, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:4;

[0035] (l) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:5, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:5;

[0036] (m) contains a nucleotide sequence as shown in SEQ ID NO:6, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:6;

[0037] (n) contains a nucleotide sequence as shown in SEQ ID NO:7, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:7;

[0038] (o) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:8, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:8;

[0039] (p) Contains a nucleotide sequence as shown in SEQ ID NO:20, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:20;

[0040] In some preferred embodiments, the recognition site (RS) of the recombinase is selected from the following combinations:

[0041] (L) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:13;

[0042] (A) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:11;

[0043] (B) Contains the nucleotide sequence shown in SEQ ID NO:3 and the nucleotide sequence shown in SEQ ID NO:6;

[0044] (C) Contains the nucleotide sequence shown in SEQ ID NO:4 and the nucleotide sequence shown in SEQ ID NO:7;

[0045] (D) Contains the nucleotide sequence shown in SEQ ID NO:5 and the nucleotide sequence shown in SEQ ID NO:8;

[0046] (E) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:12;

[0047] (F) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:13;

[0048] (G) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:14;

[0049] (H) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:15;

[0050] (I) Containing a nucleotide sequence as shown in SEQ ID NO:9 and a nucleotide sequence as shown in SEQ ID NO:16;

[0051] (J) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:17;

[0052] (K) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:12;

[0053] (M) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:14;

[0054] (N) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:15;

[0055] (O) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:16;

[0056] (P) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:17;

[0057] More preferably, the recognition site (RS) of the recombinase is selected from the following combinations:

[0058] (L) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:13;

[0059] (B) Contains the nucleotide sequence shown in SEQ ID NO:3 and the nucleotide sequence shown in SEQ ID NO:6;

[0060] (C) Contains the nucleotide sequence shown in SEQ ID NO:4 and the nucleotide sequence shown in SEQ ID NO:7;

[0061] (D) Contains the nucleotide sequence shown in SEQ ID NO:5 and the nucleotide sequence shown in SEQ ID NO:8;

[0062] (G) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:14;

[0063] (H) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:15;

[0064] (I) Containing a nucleotide sequence as shown in SEQ ID NO:9 and a nucleotide sequence as shown in SEQ ID NO:16;

[0065] (J) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:17;

[0066] (K) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:12;

[0067] (M) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:14;

[0068] (N) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:15;

[0069] (O) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:16;

[0070] (P) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:17.

[0071] In a second aspect of the present invention, a genome editing system is provided, wherein the genome editing system comprises:

[0072] i)a) An expression construct containing a CRISPR nuclease and / or a nucleotide sequence encoding the CRISPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase, or

[0073] b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase;

[0074] ii) an expression construct containing a first pegRNA and / or a nucleotide sequence encoding the first pegRNA, and

[0075] iii) An expression construct containing a second pegRNA and / or a nucleotide sequence encoding the second pegRNA.

[0076] The first pegRNA contains, from 5' to 3', a first guide sequence, a first scaffold sequence, a first reverse transcription template (RT) sequence, and a first primer binding site (PBS) sequence.

[0077] The second pegRNA contains, from 5' to 3', a second guide sequence, a first scaffold sequence, a second reverse transcription template (RT) sequence, and a second primer binding site (PBS) sequence.

[0078] The first pegRNA targets a first target sequence on the sense strand of the organism's genomic DNA, and the second pegRNA targets a second target sequence on the antisense strand of the organism's genomic DNA.

[0079] The first RT sequence and the second RT sequence are used to insert a first exogenous nucleotide sequence.

[0080] The first exogenous nucleotide sequence contains one or more recombinase recognition sites (RS); and

[0081] iv) Recombinase and / or expression constructs containing a nucleotide sequence encoding said recombinase,

[0082] The recombinase described therein includes a mutant of the Cre recombinase in the Cre-Lox system as described in the first aspect of the invention.

[0083] The recognition site (RS) of the recombinase includes any one or any combination of the recognition sites (RS) of the recombinase in the Cre-Lox system as described in the first aspect of the invention.

[0084] In some implementations, the genome editing system further includes:

[0085] v) A donor construct comprising one or more recognition sites (RS) of the recombinase and a second exogenous nucleotide sequence in the genome of the organism to be inserted;

[0086] The recognition site (RS) of the recombinase includes any one or any combination of the recognition sites (RS) of the recombinase in the Cre-Lox system as described in the first aspect of the invention.

[0087] In some implementations, the organism is a plant.

[0088] In some implementations, the CRISPR nuclease is a Cas9 nickase or a variant thereof.

[0089] In some preferred embodiments, the Cas9 nickase or a variant thereof comprises an amino acid sequence selected from SEQ ID NO:26 and 41-44.

[0090] In some implementations, the CRISPR nuclease and the reverse transcriptase are linked by a adapter.

[0091] In some embodiments, the reverse transcriptase is M-MLV reverse transcriptase or a functional variant thereof.

[0092] In some preferred embodiments, the M-MLV reverse transcriptase or a functional variant thereof comprises an amino acid sequence as shown in SEQ ID NO:31.

[0093] In some embodiments, the reverse transcriptase is fused to the nucleocapsid protein (NC) directly or via a linker at the N-terminus or C-terminus.

[0094] In some embodiments, the nucleocapsid protein (NC) comprises an amino acid sequence as shown in SEQ ID NO:28.

[0095] In some implementations, the CRISPR nuclease described in i)-b) is fused to the N-terminus of the reverse transcriptase.

[0096] In some embodiments, the recombinase is contained in the guide editing fusion protein described in i)-b).

[0097] In some optional embodiments, the recombinase is located at the N-terminus or C-terminus of the guide editing fusion protein, and is directly or via a linker linked to the guide editing fusion protein.

[0098] In some preferred embodiments, the fusion protein comprising the recombinase and the guided editing fusion protein comprises the amino acid sequence shown in SEQ ID NO:45 or an amino acid sequence having 85%, 90%, or 95% identity with it.

[0099] In some implementations, the genome editing system further includes:

[0100] vi) an expression construct containing a third pegRNA and / or a nucleotide sequence encoding said third pegRNA, and

[0101] vii) An expression construct containing a fourth pegRNA and / or a nucleotide sequence encoding the fourth pegRNA.

[0102] The third pegRNA, from 5' to 3', comprises a third guide sequence, a first scaffold sequence, a third reverse transcription template (RT) sequence, and a third primer binding site (PBS) sequence.

[0103] The fourth pegRNA, from 5' to 3', comprises a fourth guide sequence, a first scaffold sequence, a fourth reverse transcription template (RT) sequence, and a fourth primer binding site (PBS) sequence.

[0104] The third pegRNA targets a third target sequence on the sense strand of the organism's genomic DNA, which is introduced into the recombinase recognition site (RS) via the genome editing system. The fourth pegRNA targets a fourth target sequence on the antisense strand of the organism's genomic DNA, which is introduced into the recombinase recognition site (RS) via the genome editing system.

[0105] The third RT sequence and the fourth RT sequence are used to insert a third exogenous nucleotide sequence to delete the recognition site (RS) of recombinase introduced into the genome of an organism by the genome editing system.

[0106] In some embodiments, the pegRNA is capable of forming a complex with the CRISPR nuclease, the guided editing fusion protein, or a fusion protein containing a recombinase and the guided editing fusion protein, and targeting the CRISPR nuclease, the guided editing fusion protein, or a fusion protein containing a recombinase and the guided editing fusion protein to a target sequence in the genome, resulting in a cut in the target strand within the target sequence.

[0107] In some implementations, the spacing between the PAMs of the first target sequence and the second target sequence, or between the PAMs of the third target sequence and the fourth target sequence, is approximately 20 bp to approximately 60 bp.

[0108] In some embodiments, the guide sequence in the first pegRNA has sufficient sequence identity (preferably 100%) with the first target sequence on the sense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the first target sequence; the guide sequence in the second pegRNA has sufficient sequence identity (preferably 100%) with the second target sequence on the antisense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the second target sequence; and / or

[0109] The guide sequence in the third pegRNA has sufficient sequence identity (preferably 100%) with the third target sequence on the sense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the third target sequence; the guide sequence in the fourth pegRNA has sufficient sequence identity (preferably 100%) with the fourth target sequence on the antisense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the fourth target sequence.

[0110] In some implementations, the primer binding site sequence is configured to be complementary to at least a portion of the target sequence, preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the DNA strand containing the target sequence caused by a nick.

[0111] In some embodiments, the pegRNA scaffold sequence is shown in SEQ ID NO:37.

[0112] In some embodiments, the first RT sequence and the second RT sequence are configured to generate a first exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted after reverse transcription using them as templates, or to generate a complementary sequence of the first exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted; and / or

[0113] The third RT sequence and the fourth RT sequence are configured to generate a third exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted after reverse transcription using them as templates, or to generate a complementary sequence of the third exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted.

[0114] In some embodiments, the first RT sequence of the first pegRNA is configured to generate a first fragment of a first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; the second RT sequence of the second pegRNA is configured to be the complementary sequence of a second fragment of the first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; and / or

[0115] The third RT sequence of the third pegRNA is configured as the third fragment of the third exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; the fourth RT sequence of the fourth pegRNA is configured as the complementary sequence of the fourth fragment of the third exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template.

[0116] In some embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted into the genome of the organism at least partially overlap; and / or

[0117] The third and fourth segments of the third exogenous nucleotide sequence to be inserted into the genome of the organism to be inserted at least partially overlap.

[0118] In some implementations, the first and second fragments have at least about 10 bp to about 30 bp overlap; and / or

[0119] The third and fourth segments overlap by at least approximately 10 bp to approximately 30 bp.

[0120] In some implementations, the pegRNA also contains a tevopre sequence at the 3' end of the PBS.

[0121] In some embodiments, the 5' end of the pegRNA is linked to a first ribozyme or tRNA, the first ribozyme or tRNA being designed to cleave the fusion at the 5' end of the pegRNA; and / or the 3' end of the pegRNA is linked to a second ribozyme or tRNA, the second ribozyme or tRNA being designed to cleave the fusion at the 3' end of the pegRNA.

[0122] In some implementations, the pegRNA is transcribed by a type II promoter, optionally a GS promoter.

[0123] In some implementations, the second exogenous nucleotide sequence may be 100 bp to approximately 18 kb or longer.

[0124] In some embodiments, the recognition site (RS) in the first exogenous nucleotide sequence is flanked by a PAM that can be recognized by the CRISPR nuclease or the guide editing fusion protein; and / or the second exogenous nucleotide sequence is flanked by a PAM that can be recognized by the CRISPR nuclease or the guide editing fusion protein.

[0125] In some embodiments, the second exogenous nucleotide sequence is associated with plant traits such as agronomic traits, thereby the insertion of the second exogenous nucleotide sequence results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.

[0126] In some implementations, the plants include monocotyledonous and dicotyledonous plants, such as crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0127] In a third aspect of the invention, a method for producing genetically modified plants is provided, wherein the genetically modified plants comprise at least one modification of site-directed insertion of a foreign nucleotide sequence, deletion, substitution, inversion, or translocation of at least a portion of the genome, the method comprising introducing the Cre-Lox system described in the first aspect of the invention or the genome editing system described in the second aspect of the invention into at least one plant.

[0128] In some embodiments, the method further includes screening plants with desired modifications from the at least one plant.

[0129] In some implementations, the genome editing system is introduced into plants via a method selected from: gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method.

[0130] In some implementations, the introduction includes converting the genome editing system into isolated plant cells or tissues, and then regenerating the converted plant cells or tissues into a complete plant.

[0131] In some implementations, the importation includes converting the genome editing system to a specific part of the whole plant, such as a leaf, shoot tip, pollen tube, young spike, or hypocotyl.

[0132] In some implementations, components of the genome editing system are simultaneously introduced into the plant.

[0133] In some implementations, the method includes the following steps:

[0134] 1) Transform components i)-iv) of the genome editing system into isolated plant cells or tissues to obtain plant cells or tissues containing the first exogenous nucleotide sequence with a recognition site (RS) of one or more recombinases inserted;

[0135] Optionally, 2) the component v) of the genome editing system is converted into the plant cells or tissues obtained in step 1), thereby obtaining plant cells or tissues containing the inserted second exogenous nucleotide sequence;

[0136] Optionally, 3) converting components vi)-vii) of the genome editing system into the plant cells or tissues obtained in step 1) or step 2); and

[0137] 4) Regenerate complete plants from plant cells or tissues obtained in step 1), step 2), or step 3).

[0138] In a fourth aspect of the invention, the use of the Cre-Lox system described in the first aspect of the invention or the genome editing system described in the second aspect of the invention in the preparation of pharmaceutical compositions for treating diseases in subjects of need is provided.

[0139] In a fifth aspect of the invention, a pharmaceutical composition is provided for treating a disease in a subject of need, wherein the pharmaceutical composition comprises the Cre-Lox system described in the first aspect of the invention or the genome editing system described in the second aspect of the invention, and optionally, a pharmaceutically acceptable carrier.

[0140] In a sixth aspect of the invention, a method for site-directed modification of the genome of a human or non-human animal is provided, wherein the method comprises introducing the Cre-Lox system described in the first aspect of the invention or the genome editing system described in the second aspect of the invention into at least one human or non-human animal cell; optionally, the method is a method for non-diagnostic and non-therapeutic purposes.

[0141] The effects of the invention

[0142] To avoid reversible recombination activity at Lox sites in the PrimeRoot system, this invention develops a rapid optimization and modification platform for recombination sites. Specifically, a high-throughput sequencing-based method for identifying base preferences at recombination sites is developed. A fluorescence reporter system is then used for rapid screening of recombination site variants, resulting in combinations of recombination site variants that significantly reduce reversible recombination activity while maintaining high recombination activity. An AI-assisted protein optimization method is used to design Cre spikes, and through testing and spike combination, Cre variants with significantly improved recombination efficiency are obtained. Integrating the above optimizations, a PCE system is developed, significantly improving the integration efficiency of large DNA fragments in rice, maize, and wheat. This system is further used to achieve targeted and precise deletion, replacement, inversion, and translocation manipulation of large DNA fragments. Based on a dual-pegRNA guided editing system capable of efficient and precise insertion of short DNA fragments, and utilizing the ePPEplus guided editing system, currently the most efficient in plants, Cas9 is replaced with SpRY-Cas9, and enhancements to Cas9 are introduced. Point mutations in nick activity yielded the SpRY-ePPEplus system. The system's efficiency in inserting Lox sites into random PAMs (NAN, NTN, NCN, NGN) was tested, showing high Lox site insertion efficiency in almost all PAMs. This system was further applied to the PCE system. Leveraging its PAM-free nature, additional pegRNAs (Re-pegRNAs) were designed to re-edit the Lox sites remaining on the genome after editing, developing the RePCE system. This enables completely precise and seamless manipulation of large DNA fragments. Attached Figure Description

[0143] Figure 1: Schematic diagram of rapid modification process of recombination sites.

[0144] Figure 2: Schematic diagram of the principle and method of constructing recombinant site libraries.

[0145] Figures 3A and 3B: Heatmaps of Lox-dm next-generation sequencing results.

[0146] Figure 4: Schematic diagram of Cre's base preference for Lox sites (weblogo).

[0147] Figures 5A and 5B: Schematic diagrams of recombination reporting systems and reversible reporting systems for recombination site variants.

[0148] Figure 6: Results of RS variant reporting system identified by flow cytometry.

[0149] Figure 7: Schematic diagram of Cre point mutation test vector construction.

[0150] Figures 8A to 8C: Results of Cre variants identified by flow cytometry using the reporter system.

[0151] Figure 9: Insertion efficiency of different endogenous targets in rice mediated by the PrimeRoot system.

[0152] Figure 10: Schematic diagram of PCE construction.

[0153] Figures 11 and 12: Comparison of insertion efficiency of PrimeRoot.v2 and PCE at different endogenous target sites in rice, maize and wheat.

[0154] Figure 13: PCE-mediated inversion efficiency of large DNA fragments at different endogenous target sites in the rice genome.

[0155] Figure 14: Comparison of NHEJ strategy and PCE in rice genome inversion efficiency.

[0156] Figure 15: Comparison of inversion efficiency of PrimeRoot.v2 and PCE in rice genome.

[0157] Figure 16: Efficiency of PCE-mediated large-fragment DNA replacement at endogenous targets in rice.

[0158] Figure 17: First-generation sequencing of PCR products from PCE-mediated large DNA deletion events in the rice genome.

[0159] Figure 18: PCE-mediated large-fragment DNA deletion efficiency in the rice genome.

[0160] Figure 19: First-generation sequencing of PCR products at the junction of PCE-mediated rice genome chromosome translocation events.

[0161] Figure 20: Efficiency of PCE-mediated chromosome translocation in rice genome.

[0162] Figure 21: Gel electrophoresis identification of large-fragment DNA inversion variants in rice plants mediated by PCE.

[0163] Figure 22: Statistical analysis of large DNA inversion efficiency in rice plants mediated by PCE and PrimeRoot systems.

[0164] Figure 23: Gel electrophoresis identification of T1 generation rice plants with large DNA inversion mediated by PCE.

[0165] Figure 24: Schematic diagram of the construction of PAMless test PE vectors mediated by different Cas9 variants.

[0166] Figure 25: Efficiency heatmap of Cas9 variant-mediated guide editors in different PAMs of the rice genome.

[0167] Figure 26: A schematic comparison between the PrimeRoot system with PAM restrictions and Lox residue after editing and the RePCE system without PAM restrictions and capable of seamless editing.

[0168] Figure 27: Schematic diagram of RePCE construction.

[0169] Figure 28: Efficiency of RePCE-mediated randomization of 4 PAM sites in the rice genome.

[0170] Figure 29: Schematic diagram of the construction and operation of the RePCE-based fluorescence reporting system SPR.

[0171] Figure 30: The proportion of all large DNA editing events in the rice genome that were achieved through RePCE without leaving a trace at five endogenous target sites.

[0172] Figure 31: Schematic diagram of seamless editing in the RePCE system based on the rG strategy.

[0173] Figure 32: Efficiency of large-fragment DNA editing without scarring in the RePCE system mediated by the rG strategy.

[0174] Figure 33: The proportion of all large-fragment DNA editing events in the RePCE system mediated by the rG strategy.

[0175] Figure 34: Efficiency of large DNA fragment insertion into human cells mediated by the RePCE system.

[0176] Figure 35: PCE system-mediated efficiency of large DNA insertion in human cells.

[0177] Figure 36: The proportion of all large DNA insertion events in human cells mediated by the RePCE system without any visible insertions.

[0178] Figure 37: Efficiency of chromosome editing without scarring in human cells mediated by the RePCE system.

[0179] Figure 38: The proportion of chromosome-free editing in human cells mediated by the RePCE system.

[0180] Figure 39: First-generation sequencing of PCR products at the junction of PCE and RePCE-mediated human cell chromosome editing events. Detailed Implementation

[0181] Various exemplary embodiments, features, and aspects of the present invention will be described in detail below. The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0182] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art should understand that the present invention can be practiced without certain specific details. In other instances, methods, means, apparatus, and steps well known to those skilled in the art have not been described in detail in order to highlight the spirit of the present invention.

[0183] Unless otherwise stated, all units used in this specification are international standard units, and all numerical values ​​and ranges appearing in this invention should be understood to include systematic errors that are unavoidable in industrial production.

[0184] [Terminology Definition]

[0185] In this specification, the word "may" has two meanings: to perform a certain process and not to perform a certain process.

[0186] In this specification, references to "some specific / preferred embodiments," "other specific / preferred embodiments," "implementation," etc., refer to specific elements (e.g., features, structures, properties, and / or characteristics) related to that embodiment, which are included in at least one of the embodiments described herein and may or may not be present in other embodiments. Furthermore, it should be understood that these elements may be combined in any suitable manner in various embodiments.

[0187] In this specification, the range of values ​​referred to as "value A to value B" refers to the range including the endpoint values ​​A and B.

[0188] Unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are all widely used terms and routine procedures in their respective fields. For example, the standard recombinant DNA and molecular cloning techniques used in this invention are well known to those skilled in the art and are described more fully in the following literature: Sambrook, J., Fritsch, EF, and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). Meanwhile, to better understand this invention, definitions and explanations of relevant terms are provided below.

[0189] In this specification, the term "and / or" covers all combinations of items connected by the term and should be regarded as if each combination had been listed separately herein. For example, "A and / or B" covers "A", "A and B", and "B". For example, "A, B and / or C" covers "A", "B", "C", "A and B", "A and C", "B and C", and "A and B and C".

[0190] In this specification, the term "comprising" is used to describe a protein or nucleic acid sequence. The protein or nucleic acid may consist of the stated sequence, or it may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, while still possessing the activities described in this invention. Furthermore, those skilled in the art will understand that the methionine encoded by the start codon at the N-terminus of a polypeptide may be retained in certain practical situations (e.g., when expressed in a specific expression system) without substantially affecting the polypeptide's function. Therefore, when describing a specific polypeptide amino acid sequence in this application and claims, although it may not contain the methionine encoded by the start codon at the N-terminus, the sequence containing that methionine is still included. Correspondingly, its encoding nucleotide sequence may also contain the start codon; and vice versa.

[0191] In this specification, a "genome editing system" refers to a combination of components required for genome editing within cells. The individual components of such a system, such as a guide editing fusion protein or its expression construct, pegRNA or its expression construct, donor construct, etc., may exist independently or in any combination as a composition.

[0192] In this specification, the term "genome," as used herein, encompasses not only chromosomal DNA present in the cell nucleus but also organelle DNA present in subcellular components of the cell, such as mitochondria and plastids.

[0193] In this specification, "genetically modified plant" means a plant whose genome contains inserted exogenous polynucleotides. For example, the exogenous polynucleotides can be stably integrated into the plant's genome and inherited for successive generations. "Genetically modified plant" also means a plant whose genome has been artificially (e.g., by using the gene editing system provided in this invention) deleted, replaced, inverted, and / or translocated.

[0194] In this specification, the term “exogenous” for the purposes of referring to a sequence means a sequence derived from a foreign species, or, if derived from the same species, a sequence whose composition and / or loci have been significantly altered from its natural form through deliberate human intervention.

[0195] In this specification, the terms "polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and refer to single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their individual letter names as follows: "A" for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), "C" for cytidine or deoxycytidine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purine (A or G), "Y" for pyrimidine (C or T), "K" for G or T, "H" for A, C, or T, "D" for A, T, or G, "I" for inosine, and "N" for any nucleotide. Although nucleotide sequences may be represented as DNA sequences (containing T) herein, when RNA is referred to, those skilled in the art can readily determine the corresponding RNA sequence (i.e., replacing T with U).

[0196] In this specification, the terms "polypeptide," "peptide," and "protein" are used interchangeably to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.

[0197] As used in this invention, the term "amino acid" can include natural amino acids, non-natural amino acids, amino acid analogs, and all their D and L stereoisomers. The amino acids and their abbreviations and English abbreviations in this invention are shown below:

[0198] Histidine (His, H); Serine (S); Glutamic acid (Glu, E); Glutamine (Gln, Q); Glycine (Gly, G); Threonine (Thr, T); Phenylalanine (Phe, F); Aspartic acid (Asp, D); Tyrosine (Tyr, Y); Leucine (Leu, L); Isoleucine (Ile, I); Arginine (Arg, R); Alanine (Ala, A); Valine (Val, V); Tryptophan (Trp, W); Methionine (Met, M); Asparagine (Asn, N); Cysteine ​​(Cys, C); Lysine (Lys, K); Proline (Pro, P).

[0199] As used herein, the term "wild-type" refers to an object that can be found in nature. For example, a polypeptide or polynucleotide sequence that exists in an organism, can be isolated from a natural source, and has not been intentionally modified by humans in a laboratory is naturally occurring. As used herein, "naturally occurring" and "wild-type" are synonyms.

[0200] As used herein, the term "mutant" refers to a polynucleotide or polypeptide that contains alterations (i.e., substitutions, insertions, and / or deletions) at one or more (e.g., several) positions relative to the "wild type" or "comparative" polynucleotide or polypeptide, wherein substitution refers to replacing a nucleotide or amino acid occupying a position with a different nucleotide or amino acid. Deletion refers to removing a nucleotide or amino acid occupying a position. Insertion refers to adding a nucleotide or amino acid adjacent to and immediately following the nucleotide or amino acid occupying the position.

[0201] As used in this invention, the term "mutated amino acid" includes "a single or more amino acids that have been substituted, repeated, deleted, or added." In this invention, the term "mutation" refers to an alteration of the amino acid sequence. In one specific embodiment, the term "mutation" refers to "substitution."

[0202] As used herein, the terms “corresponding to” or “relative to” have the meaning commonly understood by one of ordinary skill in the art. Specifically, “corresponding to” means that, after homology or sequence identity alignment, one sequence corresponds to a specified position in another sequence. Therefore, for example, regarding “corresponding to the 150th amino acid residue of the amino acid sequence shown in Sequence 1,” if a 6×His tag is added to one end of the amino acid sequence shown in Sequence 1, then the 150th position in the resulting mutant corresponding to the 156th position of the amino acid sequence shown in Sequence 1 could be the 156th position.

[0203] In this specification, "expression construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to generate mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein.

[0204] In this specification, "expression construct" can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (such as mRNA), for example, RNA transcribed in vitro.

[0205] In this specification, an "expression construct" may contain regulatory sequences and nucleotide sequences of interest from different sources, or regulatory sequences and nucleotide sequences of interest from the same source but arranged in a manner different from those normally found in nature.

[0206] In this specification, "promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the invention, a promoter is a promoter capable of controlling gene transcription in a cell, regardless of whether it originates from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmental regulatory promoter, or an inducible promoter.

[0207] In this specification, examples of promoters include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. When used for plants, promoters may be the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, the maize U3 promoter, or the rice actin promoter.

[0208] In this specification, "introducing" nucleic acid molecules (e.g., plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism means transforming the organism's cells with the nucleic acid or protein, enabling the nucleic acid or protein to perform its function within the cell. The term "transformation" as used in this invention includes stable transformation and transient transformation. "Stable transformation" refers to the introduction of a foreign nucleotide sequence into the genome, resulting in the stable inheritance of the foreign gene. Once stable transformation occurs, the foreign nucleic acid sequence is stably integrated into the genome of the organism and its genome in any subsequent generations. "Transient transformation" refers to the introduction of nucleic acid molecules or proteins into cells to perform their function without the stable inheritance of the foreign gene. In transient transformation, the foreign nucleic acid sequence does not integrate into the genome.

[0209] In this specification, "characteristics" refers to the physiological, morphological, biochemical, or physical features of a cell or organism.

[0210] In this specification, "agronomic traits" specifically refers to measurable parameters of crop plants, including but not limited to: leaf greenness, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, total nitrogen content of plants, nitrogen content of fruits, nitrogen content of seeds, nitrogen content of plant vegetative tissues, total free amino acid content of plants, free amino acid content of fruits, free amino acid content of seeds, free amino acid content of plant vegetative tissues, total protein content of plants, protein content of fruits, protein content of seeds, protein content of plant vegetative tissues, herbicide resistance and drought resistance, nitrogen uptake, root lodging, harvest index, stem lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt tolerance, and tiller number, etc.

[0211] In this specification, the term "forward" means, for example, a DNA sequence positioned at a 5' to 3' orientation, with its orientation the same as its upstream promoter sequence.

[0212] In this specification, the term "opposite" means, for example, a DNA sequence placed in a 3' to 5' orientation that is opposite to its upstream promoter sequence.

[0213] In this specification, the terms "safe harbor site" or "safe harbor locus (SHL)" are well known in the art. An SHL is a genomic locus where a gene or other genetic element can be safely inserted and expressed without altering the physiological state of the cell. An SHL is further described as a genomic location where a new gene or genetic element can be introduced without interfering with the expression or regulation of adjacent genes.

[0214] [Detailed Description of the Invention]

[0215] This invention includes the optimization and modification of the Lox site used in the PrimeRoot system disclosed in WO / 2023 / 227050 (which is incorporated herein by reference) and the point mutation optimization of Cre recombinase. The optimized and integrated PCE system is used to perform targeted and precise deletion, replacement, inversion, translocation and other editing of large fragments of plant genome DNA. Furthermore, a PAM-free RePCE system for plants has been developed, and a traceless large fragment DNA precision manipulation system has been developed based on this.

[0216] Specifically, the present invention first developed a method for rapidly modifying and optimizing the recombination site sequence based on high-throughput sequencing and a fluorescence reporting system, and obtained a combination of Lox site variants that ensured high recombination activity and significantly reduced reversible recombination activity; point mutation optimization of Cre was performed based on an artificial intelligence-assisted protein optimization method, significantly improving its recombination efficiency; the optimized components were integrated to obtain an optimized version of PCE, significantly improving the precise insertion efficiency of large fragment DNA in rice, maize, and wheat; further, various manipulation types such as targeted precise deletion, replacement, inversion, and translocation of large fragment DNA were performed using PCE; an efficient and precise recombination site insertion system without PAM restriction was developed based on SpRY and applied to the PCE system; the residual Lox sites on the genome after large fragment DNA editing were re-edited using its PAM-free property to obtain the RePCE system, achieving completely precise and traceless manipulation of large fragment DNA.

[0217] <Cre-Lox system>

[0218] In some aspects of the present invention, a Cre-Lox system is provided, wherein the Cre-Lox system comprises a recombinase and a recognition site (RS) of the recombinase, and the recombinase is a mutant of Cre recombinase.

[0219] In some embodiments, the mutant of Cre recombinase is selected from any one of the groups consisting of the following (a1)-(a3):

[0220] (a1) The amino acid sequence of the mutant of Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO: 2, and has mutations of R72V, R72I, R72L, Q94M, R101L, D141Q, R146I, R159L, R159M, R159V, R159W, K211I, R243D, T268I, A296I, G313F, G313M, G313W, S257P, Q35D, E39P, H40E, W63P, K86E, Q94A, N111P, M117L, Q156R, R173P, I197V, I225L, S226E, R241P, H269S, I272L, A275P, K276P, Q281G, M299L, T316R, D329P, T332A or any combination thereof;

[0221] (a2) A polypeptide having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, most preferably at least 99% sequence identity with the amino acid sequence shown in (a1);

[0222] (a3) A polypeptide with an amino acid sequence as shown in (a1) or (a2) having one or more amino acids added or deleted at at least one end of the N-terminus and C-terminus.

[0223] In some embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the mutations shown in 1) to 69) below:

[0224] 66)D141Q, H40E, A275P, K276P, Q281G; 1) R72V; 2) R72I; 3) R72L; 4) Q94M; 5) R101L; 6) D141Q; 7) R146I; 8) R159L; 9) R159M; 10) R159V; 11 )R159W; 12) K211I; 13) R243D; 14) T268I; 15) A296I; 16) G313F; 17) G313M; 18) G313W; 19) S257P; 20) Q35D; 21) E39P; 22) H40E; 23) W63P; 2 4)K86E; 25) Q94A; 26) N111P; 27) M117L; 28) Q156R; 29) R173P; 30) I197V; 31) I225L; 32) S226E; 33) R241P; 34) H269S; 35) I272L; 36) A275 P; 37) K276P; 38) Q281G; 39) M299L; 40) T316R; 41) D329P; 42) T332A; 43) D141Q, R243D; 44) D141Q, K276P; 45) D141Q, G313F; 46) R243D, K2 76P; 47) Q94M, D141Q; 48) Q94M, R243D; 49) Q94M, K276P; 50) R159M, R243D; 51) R159M, K276P; 52) G313F, K276P; 53) E39P, H40E; 54) S226 E, R241P; 55) I272L, A275P; 56) K276P, Q281G; 57) M299L, T332A; 58) D141Q, R243D, K276P; 59) I225L, S226E, R241P; 60) I272L, A275P, K2 76P; 61) A275P, K276P, Q281G; 62) Q156R, S226E, R241P; 63) Q156R, I225L, S226E, R241P; 64) I272L, A275P, K276P, Q281G; 65) H40E, A275 P, K276P, Q281G; 67) D141Q, I272L, A275P, K276P, Q281G; 68) R243D, I272L, A275P, K276P, Q281G; 69) R243D, H40E, A275P, K276P, Q281G.

[0225] In some preferred embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations:

[0226] 6)D141Q; 13)R243D; 22)H40E; 33)R241P; 34)H269S; 35)I272L; 36)A275P; 37)K276P; 38)Q281G; 42)T332A.

[0227] In some preferred embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations:

[0228] 66)D141Q, H40E, A275P, K276P, Q281G; 43) D141Q, R243D; 44) D141Q, K276P; 45)D141Q, G313F; 46) R243D, K276P; 47) Q94M, D141Q; 48) Q94M, R243D; 49) Q 94M, K276P; 50) R159M, R243D; 51) R159M, K276P; 52) G313F, K276P; 53) E39P , H40E; 55) I272L, A275P; 56) K276P, Q281G; 57) M299L, T332A; 58) D141Q, R24 3D, K276P; 59) I225L, S226E, R241P; 60) I272L, A275P, K276P; 61) A275P, K2 76P, Q281G; 63) Q156R, I225L, S226E, R241P; 64) I272L, A275P, K276P, Q281G ;65)H40E, A275P, K276P, Q281G; 67)D141Q, I272L, A275P, K276P, Q281G; 68 )R243D, I272L, A275P, K276P, Q281G; 69) R243D, H40E, A275P, K276P, Q281G.

[0229] In some preferred embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations:

[0230] 66)D141Q, H40E, A275P, K276P, Q281G; 67)D141Q, I272L, A275P, K276P, Q 281G; 65) H40E, A275P, K276P, Q281G; 46) R243D, K276P; 44) D141Q, K276P.

[0231] In some further preferred embodiments, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has the following mutations:

[0232] 66)D141Q, H40E, A275P, K276P, Q281G.

[0233] In some embodiments, the recognition site (RS) of the recombinase is selected from any one of the following or any combination thereof:

[0234] (a) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:9, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:9;

[0235] (b) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:10, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:10;

[0236] (c) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:11, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:11;

[0237] (d) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:12, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:12;

[0238] (e) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:13, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:13;

[0239] (f) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:14, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:14;

[0240] (g) Contains a nucleotide sequence as shown in SEQ ID NO:15, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:15;

[0241] (h) Contains a nucleotide sequence as shown in SEQ ID NO:16, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:16;

[0242] (i) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:17, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:17;

[0243] (j) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:3, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:3;

[0244] (k) Contains a nucleotide sequence as shown in SEQ ID NO:4, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:4;

[0245] (l) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:5, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:5;

[0246] (m) contains a nucleotide sequence as shown in SEQ ID NO:6, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:6;

[0247] (n) contains a nucleotide sequence as shown in SEQ ID NO:7, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:7;

[0248] (o) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:8, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:8;

[0249] (p) contains a nucleotide sequence as shown in SEQ ID NO:20, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:20.

[0250] In some preferred embodiments, the recognition site (RS) of the recombinase is selected from the following combinations:

[0251] (L) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:13;

[0252] (A) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:11;

[0253] (B) Contains the nucleotide sequence shown in SEQ ID NO:3 and the nucleotide sequence shown in SEQ ID NO:6;

[0254] (C) Contains the nucleotide sequence shown in SEQ ID NO:4 and the nucleotide sequence shown in SEQ ID NO:7;

[0255] (D) Contains the nucleotide sequence shown in SEQ ID NO:5 and the nucleotide sequence shown in SEQ ID NO:8;

[0256] (E) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:12;

[0257] (F) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:13;

[0258] (G) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:14;

[0259] (H) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:15;

[0260] (I) Containing a nucleotide sequence as shown in SEQ ID NO:9 and a nucleotide sequence as shown in SEQ ID NO:16;

[0261] (J) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:17;

[0262] (K) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:12;

[0263] (M) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:14;

[0264] (N) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:15;

[0265] (O) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:16;

[0266] (P) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:17.

[0267] In some preferred embodiments, the recognition site (RS) of the recombinase is selected from the following combinations:

[0268] (L) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:13;

[0269] (B) Contains the nucleotide sequence shown in SEQ ID NO:3 and the nucleotide sequence shown in SEQ ID NO:6;

[0270] (C) Contains the nucleotide sequence shown in SEQ ID NO:4 and the nucleotide sequence shown in SEQ ID NO:7;

[0271] (D) Contains the nucleotide sequence shown in SEQ ID NO:5 and the nucleotide sequence shown in SEQ ID NO:8;

[0272] (G) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:14;

[0273] (H) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:15;

[0274] (I) Containing a nucleotide sequence as shown in SEQ ID NO:9 and a nucleotide sequence as shown in SEQ ID NO:16;

[0275] (J) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:17;

[0276] (K) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:12;

[0277] (M) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:14;

[0278] (N) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:15;

[0279] (O) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:16;

[0280] (P) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:17.

[0281] In some further preferred embodiments, the recognition site (RS) of the recombinase is selected from the following combinations:

[0282] (L) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:13.

[0283] <Genome Editing Systems>

[0284] Some aspects of the present invention provide a genome editing system, wherein the genome editing system comprises:

[0285] i) a) A CRISPR nuclease and / or an expression construct containing a nucleotide sequence encoding the CRISPR nuclease, and a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the reverse transcriptase, or

[0286] b) A prime editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the prime editing fusion protein, wherein the prime editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase;

[0287] ii) A first pegRNA and / or an expression construct containing a nucleotide sequence encoding the first pegRNA, and

[0288] iii) A second pegRNA and / or an expression construct containing a nucleotide sequence encoding the second pegRNA,

[0289] wherein the first pegRNA comprises a first guide sequence, a first scaffold sequence, a first reverse transcription template (RT) sequence, and a first primer binding site (PBS) sequence in the 5' to 3' direction,

[0290] wherein the second pegRNA comprises a second guide sequence, a first scaffold sequence, a second reverse transcription template (RT) sequence, and a second primer binding site (PBS) sequence in the 5' to 3' direction,

[0291] wherein the first pegRNA targets a first target sequence on the sense strand of the genomic DNA of an organism, and the second pegRNA targets a second target sequence on the antisense strand of the genomic DNA of the organism,

[0292] wherein the first RT sequence and the second RT sequence are used for inserting a first exogenous nucleotide sequence,

[0293] wherein the first exogenous nucleotide sequence comprises one or more recognition sites (RS) of a recombinase; and

[0294] iv) A recombinase and / or an expression construct containing a nucleotide sequence encoding the recombinase,

[0295] wherein the recombinase comprises a mutant of the Cre recombinase as described in the <Cre-Lox system>,

[0296] wherein the recognition site (RS) of the recombinase comprises any one or any combination of the recognition sites (RS) of the recombinase as described in the <Cre-Lox system>.

[0297] In some embodiments, the genome editing system further comprises:

[0298] v) A donor construct comprising one or more recognition sites (RS) of the recombinant enzyme and a second exogenous nucleotide sequence to be inserted into the genome of the organism.

[0299] In some embodiments, the recognition site (RS) of the recombinant enzyme comprises any one or any combination of the recognition sites (RS) of the recombinant enzymes described in the <Cre-Lox system>.

[0300] In some embodiments, the organism is a plant or animal cell.

[0301] In this specification, a "target sequence" refers to a sequence of approximately 20 nucleotides in the genome characterized by a PAM (protospacer adjacent motif) sequence flanking the 5' or 3' end. Generally, PAM is required for the complex formed by a CRISPR nuclease or its variant and a guide RNA to recognize the target sequence. For example, for Cas9 nuclease and its variants, its target sequence is adjacent to the PAM at the 3' end, such as 5'-NGG-3'. Based on the presence of PAM, those skilled in the art can easily determine the target sequences available for targeting in the genome. Moreover, depending on the position of PAM, the target sequence can be located on either strand of the genomic DNA molecule, and the strand where the target sequence is located is called the target strand. For Cas9 or its derivatives such as Cas9 nickase, the target sequence is preferably 20 nucleotides. Depending on different CRISPR nucleases or their different variants, the PAM sequence will vary.

[0302] In some embodiments, the pegRNA is capable of forming a complex with a fusion protein (a prime editing fusion protein or a fusion protein of a prime editing fusion protein and a recombinant enzyme) and targeting the fusion protein to a target sequence in the genome, resulting in a nick on the target strand (e.g., within the target sequence).

[0303] In some alternative embodiments, the interval between the PAMs of the first target sequence and the second target sequence is approximately 20 bp - approximately 60 bp.

[0304] (CRISPR nuclease)

[0305] In some embodiments, the CRISPR nuclease is Cas9 nuclease, such as SpCas9 derived from Streptococcus pyogenes (S.pyogenes). An exemplary wild-type SpCas9 comprises the amino acid sequence shown in SEQ ID NO:1 in WO / 2023 / 227050.

[0306] In some embodiments, the CRISPR nuclease is a CRISPR nickase. The CRISPR nickase in the fusion protein is capable of forming a nick within the target sequence on the target strand of the genomic DNA. In some embodiments, the CRISPR nickase is a Cas9 nickase.

[0307] In some embodiments, the Cas9 nickase is derived from *S. pyogenes* SpCas9 and contains at least the amino acid substitution H840A relative to wild-type SpCas9. In some embodiments, the Cas9 nickase comprises the amino acid sequence shown in SEQ ID NO:2 of WO / 2023 / 227050. In some embodiments, the Cas9 nickase in the fusion protein is capable of forming a nick between the -3 nucleotide (the first nucleotide at the 5' end of the PAM sequence is +1) and the -4 nucleotide of the target sequence.

[0308] In some embodiments, the Cas9 nickase is derived from SpCas9 of *S. pyogenes* and contains at least the amino acid substitutions R211K, N394K, and H840A relative to wild-type SpCas9. In some preferred embodiments, the Cas9 nickase comprises the amino acid sequence shown in SEQ ID NO:26.

[0309] In some embodiments, the Cas9 nuclease, such as a nickase, is a Cas9 nuclease or nickase variant capable of recognizing an altered PAM sequence. Many Cas9 nickase variants capable of recognizing altered PAM sequences are known in the art. In some embodiments, the Cas9 nuclease, such as a nickase, is a Cas9 variant that recognizes the PAM sequence 5'-NG-3'. In some embodiments, the Cas9 nickase variant recognizing the PAM sequence 5'-NG-3' comprises, relative to wild-type Cas9, the following amino acid substitutions: H840A, D1135L, S1136W, G1218K, E1219Q, R1335Q, T1337R, wherein the amino acid numbers refer to SEQ ID NO:1 in WO / 2023 / 227050. In some embodiments, the Cas9 nickase variant (SpG-Cas9 nickase) comprises the amino acid sequence shown in SEQ ID NO:42 in WO / 2023 / 227050. In some embodiments, the Cas9 nickase variant recognizing the PAM sequence 5'-NG-3' comprises, relative to wild-type Cas9, the following amino acid substitutions: H840A, A61R, L1111R, D1135L, S1136W, G1218K, E1219Q, N1317R, A1322R, R1333P, R1335Q, and T1337R, wherein the amino acid numbers refer to SEQ ID NO:1 in WO / 2023 / 227050. In some embodiments, the Cas9 nickase variant (SpRY-Cas9 nickase) comprises the amino acid sequence shown in SEQ ID NO:43 in WO / 2023 / 227050.

[0310] In some preferred embodiments, the Cas9 nuclease, such as a nickase, is a PAM-widened or PAM-free Cas9 variant, such as SpRY-Cas9(R211K,N394K,H840A)(SEQ ID NO:40), BCA(R221K,H838A)(SEQ ID NO:41), FCA(R221K,H838A)(SEQ ID NO:42), and Cas9(R221K,N394K,H840,FCA*)(SEQ ID NO:43).

[0311] In some preferred embodiments, the sequence of the Cas9 nuclease is shown in SEQ ID NO:40.

[0312] The nicks formed by the Cas9 nuclease described in this invention, such as the nicking enzyme, can lead to the formation of a free single strand with a 3' end (3' free single strand) and a free single strand with a 5' end (5' free single strand).

[0313] In some implementations, the CRISPR nucleases, such as Cas9 nickase and the reverse transcriptase, in the fusion protein are linked by a linker.

[0314] (Reverse transcriptase)

[0315] In some embodiments, the reverse transcriptase of the present invention may be derived from different sources. In some embodiments, the reverse transcriptase is a viral reverse transcriptase. For example, in some embodiments, the reverse transcriptase is M-MLV reverse transcriptase or a functional variant thereof. An exemplary wild-type M-MLV reverse transcriptase sequence is shown in SEQ ID NO:3 of WO / 2023 / 227050.

[0316] In some embodiments, the reverse transcriptase is, for example, M-MLV reverse transcriptase or a functional variant thereof:

[0317] (a) A mutation comprising positions 155, 156, 200, 223 and / or 524, for example, a mutation comprising any one or a combination of F155Y, F155V, F156Y, D524N, N200C, V223A, the amino acid positions being referenced to SEQ ID NO:3 in WO / 2023 / 227050;

[0318] (b) The connection sequence is missing; and / or

[0319] (c) The RNase H domain is mutated or missing.

[0320] In some alternative embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, comprises a mutation selected from D524N, the amino acid position of which refers to SEQ ID NO:3 in WO / 2023 / 227050.

[0321] In some alternative embodiments, the RNase H domain of the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is missing.

[0322] In some alternative embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, comprises a mutation selected from V223A, the amino acid position of which refers to SEQ ID NO:3 in WO / 2023 / 227050.

[0323] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, comprises a mutation selected from V223A, the amino acid position of which refers to SEQ ID NO:3, and the RNase H domain is missing.

[0324] In some embodiments, the connection sequence comprises an amino acid sequence as shown in SEQ ID NO:4 of WO / 2023 / 227050.

[0325] In some embodiments, the RNase H domain comprises an amino acid sequence as shown in SEQ ID NO:5 of WO / 2023 / 227050.

[0326] In some embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, comprises the sequence of any one of SEQ ID NO:9-15 in WO / 2023 / 227050.

[0327] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, comprises the amino acid sequence shown in SEQ ID NO:31.

[0328] In some embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused at its N-terminus or C-terminus directly or via a linker to a nucleocapsid protein (NC), a hydrolase (PR), or an integrase (IN). The nucleocapsid protein (NC), hydrolase (PR), or integrase (IN) is, for example, derived from M-MLV.

[0329] In some embodiments, the nucleocapsid protein (NC) comprises the amino acid sequence shown in SEQ ID NO:6 of WO / 2023 / 227050.

[0330] In some embodiments, the hydrolase (PR) comprises an amino acid sequence as shown in SEQ ID NO:7 of WO / 2023 / 227050.

[0331] In some embodiments, the integrase (IN) comprises the amino acid sequence shown in SEQ ID NO:8 of WO / 2023 / 227050.

[0332] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused at the N-terminus to a nucleocapsid protein (NC) directly or via a linker.

[0333] In some preferred embodiments, the reverse transcriptase, such as M-MLV reverse transcriptase or a functional variant thereof, is fused at its C-terminus to a nucleocapsid protein (NC) directly or via a linker.

[0334] In some embodiments, the reverse transcriptase may also be fused via a linker or directly to an RNA aptamer-binding protein sequence (e.g., the MCP protein sequence). Thereby, the reverse transcriptase can be recruited to a CRISPR nuclease through the interaction of the RNA aptamer-binding protein sequence (e.g., the MCP protein sequence) and one or more RNA aptamer sequences (e.g., the MS2 sequence) present on the pegRNA. In this case, it is not necessary to fuse the CRISPR nuclease to the reverse transcriptase. An exemplary MCP protein comprises the amino acid sequence of SEQ ID NO:44 in WO / 2023 / 227050.

[0335] As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids without secondary or higher structures. For example, the linker can be a flexible linker, such as GGGGS, GS, GAP, (GGGGS)x3, GGS, and (GGS)x7. For example, it can be the linker shown in SEQ ID NO:27 or the linker shown in SEQ ID NO:30.

[0336] In some embodiments, the CRISPR nuclease, such as a CRISPR nickase, in the fusion protein is located at the N-terminus of the reverse transcriptase. In some embodiments, the CRISPR nuclease, such as a CRISPR nickase, in the fusion protein is located at the C-terminus of the reverse transcriptase.

[0337] In some embodiments, the recombinase may be a separately expressed recombinase or may be included in the guide editing fusion protein.

[0338] In some preferred embodiments, the recombinase is contained within the guided editing fusion protein. In some embodiments, the recombinase is located at the N-terminus of the guided editing fusion protein relative to the CRISPR nuclease and reverse transcriptase. In some embodiments, the recombinase is located at the C-terminus of the guided editing fusion protein relative to the CRISPR nuclease and reverse transcriptase. That is, the guided editing fusion protein and the recombinase further form a fusion protein.

[0339] In some embodiments of the present invention, the CRISPR nuclease, reverse transcriptase, recombinase, or fusion protein (guided editing fusion protein or a fusion protein of guided editing fusion protein and recombinase) of the present invention may further comprise one or more nuclear localization sequences (NLS). Generally, one or more NLS of the CRISPR nuclease, reverse transcriptase, recombinase, or fusion protein should have sufficient strength to drive the accumulation of the CRISPR nuclease, reverse transcriptase, recombinase, or fusion protein in the nucleus of the cell to an amount sufficient to achieve its base editing function. Generally, the strength of nuclear localization activity is determined by the number, location, one or more specific NLS used, or a combination of these factors in the CRISPR nuclease, reverse transcriptase, recombinase, or fusion protein.

[0340] In some preferred embodiments, the fusion protein guiding the editing of the fusion protein and the recombinase comprises, from the N-terminus to the C-terminus, the CRISPR nuclease such as a nickase, the nucleocapsid protein (NC), the reverse transcriptase, and the recombinase, linked by or without a linker. In some preferred embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, a nuclear localization sequence-the CRISPR nuclease such as a nickase-linker-the nucleocapsid protein (NC)-nuclear localization sequence-linker-the reverse transcriptase-nuclear localization sequence-linker-the recombinase-nuclear localization sequence.

[0341] In some preferred embodiments, the fusion protein of the guide editing fusion protein and the recombinase comprises the amino acid sequence shown in SEQ ID NO:45.

[0342] (pegRNA)

[0343] The guide sequence (also called seed sequence or spacer sequence) in the pegRNA of the present invention is configured to have sufficient sequence identity (preferably 100% identity) with the target sequence, thereby enabling it to bind to the complementary strand of the target sequence through base pairing and achieve sequence-specific targeting.

[0344] For example, the guide sequence in the first pegRNA may have sufficient sequence identity (preferably 100% identity) with the first target sequence, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the first target sequence; the guide sequence in the second pegRNA may have sufficient sequence identity (preferably 100% identity) with the second target sequence on the opposite strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the second target sequence, thereby the two pegRNAs result in nicks on different strands of the genomic DNA.

[0345] Various scaffold sequences for gRNAs suitable for CRISPR-based genome editing (e.g., Cas9) are known in the art and can be used in the pegRNAs of this invention. In some specific embodiments, the scaffold sequence of the gRNA is shown in SEQ ID NO:37.

[0346] In some embodiments, the primer binding site (PBS) sequence is configured to be complementary to at least a portion of the target sequence (preferably perfectly paired with at least a portion of the target sequence). Preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the DNA strand containing the target sequence due to cleavage (preferably perfectly paired with at least a portion of the 3' free single strand), particularly complementary to the nucleotide sequence at the 3' end of the 3' free single strand (preferably perfectly paired). When the 3' free single strand of the strand binds to the primer binding site sequence through base pairing, the 3' free single strand can act as a primer, using the reverse transcription template (RT) sequence adjacent to the primer binding site sequence as a template, to perform reverse transcription under the action of reverse transcriptase in the fusion protein, extending the DNA sequence corresponding to the reverse transcription template (RT) sequence.

[0347] The primer binding site sequence depends on the length of the free single strand formed by the CRISPR nickase in the target sequence; however, it should have a minimum length to ensure specific binding. In some embodiments, the primer binding site sequence can be 4-20 nucleotides long, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0348] In some embodiments, the primer binding site sequence is configured to have a Tm (unwinding temperature) not exceeding approximately 52°C. In some embodiments, the Tm (unwinding temperature) of the primer binding site sequence is approximately 18°C-52°C, preferably approximately 24°C-36°C, and more preferably approximately 28°C-32°C.

[0349] Methods for calculating the Tm of a nucleic acid sequence are well known in the art; for example, it can be calculated using the Oligo Analysis Tool online analysis tool. An exemplary formula is Tm = N. G:C *4+N A:T *2, where N G:C It is the number of G and C bases in the sequence, N A:T It refers to the number of A and T bases in the sequence. A suitable Tm can be obtained by selecting an appropriate PBS length. Alternatively, a PBS sequence with a suitable Tm can be obtained by selecting an appropriate target sequence.

[0350] In some embodiments, the RT template sequence can be any sequence. Through reverse transcription, its sequence information can be integrated into the DNA strand containing the target sequence (i.e., the strand containing the target sequence PAM), and then, through cellular DNA repair, a DNA double strand containing the RT template sequence information is formed. In some embodiments, the RT template sequence contains desired modifications. For example, the desired modifications include substitution, deletion, and / or addition of one or more nucleotides. In some embodiments, the RT template sequence is configured to correspond to a sequence downstream of the target sequence nick (e.g., complementary to at least a portion of the sequence downstream of the target sequence nick), but contains desired modifications. The desired modifications include substitution, deletion, and / or addition of one or more nucleotides.

[0351] In some implementations, the two pegRNAs are configured to introduce the same desired modification. For example, one pegRNA is configured to introduce A-G substitutions at the sense strand, while the other pegRNA is configured to introduce T-C substitutions at the corresponding position on the antisense strand. As another example, one pegRNA is configured to introduce a two-nucleotide deletion at the sense strand, and the other pegRNA is configured to similarly introduce a two-nucleotide deletion at the corresponding position on the antisense strand. Other types of modifications can be deduced similarly. The same desired modification can be achieved by designing suitable RT template sequences that target two different strands of the pegRNA.

[0352] In some embodiments, the RT sequence is configured to generate a foreign nucleotide sequence or a portion thereof of the genome to be inserted into after reverse transcription using it as a template, or to generate a complementary sequence to a foreign nucleotide sequence or a portion thereof of the genome of the organism to be inserted, such as a plant. In some embodiments, the RT sequence does not contain a genomic sequence near the target sequence or a complementary sequence to a genomic sequence near the target sequence. In some embodiments, the RT sequence does not contain sequence information other than the foreign nucleotide sequence to be inserted.

[0353] In some embodiments, the first RT sequence and the second RT sequence are used to insert a first exogenous nucleotide sequence, for example, by inserting the first exogenous nucleotide sequence between the first target sequence and the second target sequence (such as between a cleavage of the first target sequence and a cleavage of the second target sequence).

[0354] In some embodiments, the first RT sequence of the first pegRNA is configured to generate a first fragment of a first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; the second RT sequence of the second pegRNA is configured to be a complementary sequence to generate a second fragment of a first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template.

[0355] In some embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted at least partially overlap. In some embodiments, the first and second fragments overlap by at least about 10 bp to about 30 bp, for example, at least about 10 bp, about 15 bp, about 20 bp, about 25 bp, or about 30 bp. In some embodiments, the first and second fragments of the first exogenous nucleotide sequence to be inserted completely overlap.

[0356] In some embodiments, the pegRNA further includes a tevopre sequence at the 3' end of the PBS. The design of the tevopre sequence can be found in James W. Nelson et al., Engineered pegRNAs improve prime editing efficiency. 2022, Nature Biotech. Volume 40, pages 402-410. An exemplary tevopre sequence is shown in SEQ ID NO:34.

[0357] In some implementations, the pegRNA includes a guide sequence, a scaffold sequence, a reverse transcription template (RT) sequence, a primer binding site (PBS) sequence, and a tevopre sequence from 5' to 3'.

[0358] In some embodiments, the pegRNA can be precisely processed using a self-processing system. In some specific embodiments, the 5' end of the pegRNA is linked to a first ribozyme or tRNA, which is designed to cleave the fusion (i.e., a fusion containing a guide sequence, scaffold sequence, reverse transcription template (RT) sequence, primer binding site (PBS) sequence, and tevopre sequence) at the 5' end of the pegRNA; and / or the 3' end of the pegRNA is linked to a second ribozyme or tRNA, which is designed to cleave the fusion at the 3' end of the pegRNA. The design of the first or second ribozyme or tRNA is within the capabilities of those skilled in the art. For example, see Gao et al., JIPB, Apr, 2014; Vol 56, Issue 4, 343-349. Methods for precisely processing gRNA can be found, for example, in WO 2018 / 149418.

[0359] In some embodiments, the pegRNA is transcribed by a type II promoter, meaning that in an expression construct containing the nucleotide sequence encoding the pegRNA, the coding nucleotide sequence of the pegRNA is operatively linked to a type II promoter. In some specific embodiments, the type II promoter is a GS promoter. An exemplary GS promoter sequence is shown in SEQ ID NO:32.

[0360] In some embodiments, the first target sequence, the second target sequence, and / or the desired modification, such as the first exogenous nucleotide sequence, are associated with an organism, such as a plant trait, such as an agronomic trait, whereby the insertion of the desired modification, such as the first exogenous nucleotide sequence, results in the organism, such as a plant, having altered (preferably improved) traits, such as agronomic traits, relative to a wild-type organism, such as a plant.

[0361] In some embodiments, the first exogenous nucleotide sequence contains one or more recombinase recognition sites (RS).

[0362] In some embodiments, the first exogenous nucleotide sequence further comprises a PAM sequence recognized by the CRISPR nuclease or the guide editing fusion protein. When the CRISPR nuclease used (or the CRISPR nuclease in the guide editing fusion protein) is not PAM-free, the PAM sequence is recognized by the CRISPR nuclease and eliminates the residual recombinase recognition site (RS) after gene editing. In some specific embodiments, the PAM sequence is located flanking the recombinase recognition site (RS) in the first exogenous nucleotide sequence. In some preferred embodiments, the PAM sequence is designed with approximately 3 protective bases upstream, for example, for 5'-NGG-3', adding a protective base results in 5'-NNNNGG-3', and correspondingly, for 3'-CCN-5', adding a protective base results in 3'-CCNNNN-5'.

[0363] In some embodiments, the donor construct further includes a PAM sequence recognized by the CRISPR nuclease or the guided editing fusion protein. In some specific embodiments, the PAM is located flanking a second exogenous nucleotide sequence. In some preferred embodiments, the PAM sequence is designed with approximately 3 protective bases upstream, for example, 5'-NGG-3' becomes 5'-NNNNGG-3' after adding a protective base, and correspondingly, 3'-CCN-5' becomes 3'-CCNNNN-5' after adding a protective base. When the CRISPR nuclease used (or the CRISPR nuclease in the guided editing fusion protein) is not PAM-free, this PAM sequence is recognized by the CRISPR nuclease and eliminates the residual recombinase recognition site (RS) after gene editing.

[0364] In some exemplary embodiments, the genome editing system provided by this invention can achieve large-scale targeted and precise DNA inversion in the genome. For example, using genome editing systems with different first exogenous nucleotide sequences containing recombinase recognition sites, positive recombinase recognition sites (e.g., positive LoxAR2) and negative recombinase recognition sites (e.g., negative Lox71) are inserted upstream and downstream of a certain region of the genome, respectively. Under the action of the recombinase in the genome editing system, the inversion of that region is achieved.

[0365] In some exemplary embodiments, the genome editing system provided by this invention can be used to precisely target deletions at the chromosome level. For example, using genome editing systems that contain recombinase recognition sites in different first exogenous nucleotide sequences, recombinase recognition sites (e.g., LoxAR2 and Lox71 in the same direction) are inserted upstream and downstream of a region of the genome, respectively. Under the action of the recombinase in the genome editing system, the deletion of that region is achieved.

[0366] In some exemplary embodiments, the genome editing system provided by this invention can be used to precisely target translocations at the chromosome level. Specifically, for example, using a genome editing system containing recombinase recognition sites in different first exogenous nucleotide sequences, one recombinase recognition site (e.g., LoxAR2) can be introduced into one chromosome, and another recombinase recognition site (e.g., Lox71) can be introduced into another chromosome, resulting in chromosomal translocation under the action of the recombinase in the genome editing system.

[0367] In some exemplary embodiments, the genome editing system provided by the present invention can achieve targeted and precise insertion of large-scale DNA into the genome. Specifically, a genome editing system in which a recombinase recognition site is included in the first exogenous nucleotide sequence is used to insert a recombinase recognition site (such as LoxAR2) into a certain region of the genome, and a donor construct is provided. The donor construct contains another recombinase recognition site (such as Lox71 site) on both sides of the second exogenous nucleotide sequence (i.e., the DNA fragment to be inserted) in the genome of the organism to be inserted, so as to achieve precise insertion.

[0368] In some exemplary embodiments, the genome editing system provided by the present invention can achieve targeted and precise replacement of large-scale DNA in the genome. Specifically, different genome editing systems in which a recombinase recognition site is included in the first exogenous nucleotide sequence are used to insert orthogonal recombinase recognition sites with different central sequences (such as orthogonal LoxAR2 with different central sequences; for the design of orthogonal recombinase recognition sites with different central sequences, refer to Lee, G, and I Saito. “Role of nucleotide sequences of loxP spacer region in Cre-mediated recombination.” Gene vol. 216, 1 (1998): 55-65. doi:10.1016 / s0378-1119(98)00325-4) upstream and downstream of a certain region of the genome respectively, and a donor construct is provided. The donor construct contains another recombinase recognition site with the corresponding central sequence (such as orthogonal Lox71 site with the corresponding central sequence) on both sides of the second exogenous nucleotide sequence (i.e., the DNA fragment to be replaced) in the genome of the organism to be inserted, so as to achieve precise replacement.

[0369] Based on one or more recombinase recognition sites (RS) inserted into the first exogenous nucleotide sequence in the genome, by providing a donor containing RS and the second exogenous nucleotide sequence, the second exogenous nucleotide sequence can be inserted into the genome of an organism such as a plant by recombination using the corresponding recombinase. The recombinase can be a recombinase expressed alone or can be included in the guide editing fusion protein. Those skilled in the art can select a suitable combination of RS located in the RS of the first exogenous polynucleotide inserted into the genome and RS located in the donor to insert the second exogenous nucleotide sequence into the genome by recombination.

[0370] In some preferred embodiments, the combination of the RS is the combination of the RS described in the <Cre-Lox system>.

[0371] The second exogenous nucleotide sequence can be of any length. The second exogenous nucleotide sequence can be 100 bp to approximately 18 kb or longer. Preferably, the second exogenous polynucleotide is a long fragment, such as at least 300 bp, at least 500 bp, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, at least 13 kb, at least 15 kb, at least 18 kb or longer. In some embodiments, the second exogenous polynucleotide can be a full-length gene.

[0372] In some embodiments, the second exogenous nucleotide sequence is associated with an organism such as a plant trait or agronomic trait, whereby the insertion of the second exogenous nucleotide sequence results in the organism such as a plant having altered (preferably improved) traits, such as agronomic traits, relative to a wild-type organism such as a plant.

[0373] Different components of the genome editing system of the present invention, such as the coding sequences of CRISPR nuclease, reverse transcriptase, guide editing fusion protein, pegRNA and / or recombinase, and the second exogenous nucleotide sequence, may be located in the same construct in different combinations, or in different constructs respectively.

[0374] The genome editing system of this invention can be used to perform site-specific modifications, such as site-specific insertion of exogenous nucleotide sequences, in organisms that can be non-human animals, humans, or plants, preferably plants. Suitable plants include monocotyledonous and dicotyledonous plants, for example, crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0375] In order to achieve effective expression in organisms such as plants, in some embodiments of the present invention, the nucleotide sequence encoding the fusion protein is codon-optimized for the organism, such as the plant species, whose genome is to be modified.

[0376] Codon optimization refers to the modification of nucleic acid sequences to enhance expression in host cells of interest by replacing at least one codon of the natural sequence with codons that are used more frequently or most frequently in the gene in the host cell (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons while maintaining the natural amino acid sequence). Different species exhibit specific preferences for certain codons of specific amino acids. Codon preference (differences in codon use between organisms) is often associated with the translation efficiency of messenger RNA (mRNA), which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be customized to achieve optimal gene expression in a given organism based on codon optimization. Codon utilization tables are readily available, for example, in the Codon Usage Database (“Codon Usage Database”) available at www.kazusa.orjp / codon / , and these tables can be adapted in various ways. See Nakamura. Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).

[0377] (Re-pegRNA)

[0378] In some implementations, the genome editing system further includes Re-pegRNA. Specifically, the genome editing system further includes:

[0379] vi) an expression construct containing a third pegRNA and / or a nucleotide sequence encoding said third pegRNA, and

[0380] vii) An expression construct containing a fourth pegRNA and / or a nucleotide sequence encoding the fourth pegRNA.

[0381] The third pegRNA, from 5' to 3', comprises a third guide sequence, a first scaffold sequence, a third reverse transcription template (RT) sequence, and a third primer binding site (PBS) sequence.

[0382] The fourth pegRNA, from 5' to 3', comprises a fourth guide sequence, a first scaffold sequence, a fourth reverse transcription template (RT) sequence, and a fourth primer binding site (PBS) sequence.

[0383] The third pegRNA targets a third target sequence on the sense strand of the organism's genomic DNA, which is introduced into the recombinase recognition site (RS) via the genome editing system. The fourth pegRNA targets a fourth target sequence on the antisense strand of the organism's genomic DNA, which is introduced into the recombinase recognition site (RS) via the genome editing system.

[0384] The third RT sequence and the fourth RT sequence are used to insert a third exogenous nucleotide sequence to delete the recognition site (RS) of recombinase introduced into the genome of an organism by the genome editing system.

[0385] The third and fourth pegRNAs are designed in accordance with the design method of the first and second pegRNAs described in the preceding (pegRNA) section (i.e., the pegRNA design method disclosed in WO / 2023 / 227050).

[0386] By designing suitable RT template sequences, the recognition sites (RS) of recombinases introduced into the genome through the genome editing system can be deleted. For example, only the residual recognition sites (RS) of recombinases after genome editing can be deleted, thereby altering the genome sequence after genome editing.

[0387] In some exemplary embodiments, a third pegRNA and a fourth pegRNA can be designed flanking a residual recombinase recognition site (RS). The RT template can be the original genomic sequence or other sequences, thereby using the third pegRNA and the fourth pegRNA to eliminate the residual recombinase recognition site (RS).

[0388] In some embodiments, when there is no PAM restriction for the CRISPR nuclease (or the CRISPR nuclease in the guided editing fusion protein), the PAM sequence can be designed such that the third guide sequence and fourth guide sequence of the third and fourth pegRNAs are located upstream or downstream of the recognition site (RS) of the recombinase introduced by the genome editing system. For example, it is located 1–20 bp upstream, preferably 1–10 bp, of the recognition site (RS) of the recombinase introduced by the genome editing system. For example, it is located 1–20 bp downstream, preferably 1–10 bp, of the recognition site (RS) of the recombinase introduced by the genome editing system.

[0389] In some alternative implementations, the PAM sequence may also be located in the recognition site (RS) of the recombinase introduced by the genome editing system, for example, from the last 1-10 bp.

[0390] In some embodiments, the third or fourth pegRNA may be constructed into a separate expression construct. In other embodiments, the third or fourth pegRNA sequence may be constructed into the same vector as the first or second pegRNA. For example, the first and third pegRNAs may be constructed into the same expression construct; the second and fourth pegRNAs may be constructed into the same expression construct.

[0391] <Methods for Site-Specific Modification of Plant Genomes>

[0392] On the other hand, the present invention provides a method for site-directed modification of a plant genome, comprising introducing the Cre-Lox system or genome editing system of the present invention into at least one of said plants. The site-directed modification includes substitution, deletion, and / or addition of one or more nucleotides. For example, the site-directed modification includes the site-directed insertion of a foreign nucleotide sequence. In some embodiments, the site-directed modification further includes deletion, substitution, inversion, and translocation of at least a portion of the genome.

[0393] On the other hand, the present invention provides a method for producing genetically modified plants, said genetically modified plants comprising site-directed modifications, said method comprising introducing the Cre-Lox system or genome editing system of the present invention into at least one of said plants. The site-directed modifications include substitution, deletion, and / or addition of one or more nucleotides. For example, the site-directed modifications include the site-directed insertion of a foreign nucleotide sequence. In some embodiments, the site-directed modifications further include deletion, substitution, inversion, and translocation of at least a portion of the genome.

[0394] In some embodiments, the method further includes screening plants with desired site-specific modifications from the at least one plant.

[0395] In the method of this invention, the Cre-Lox system or genome editing system can be introduced into plants using various methods well known to those skilled in the art. Methods for introducing the Cre-Lox system or genome editing system of this invention into plants include, but are not limited to: gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the Cre-Lox system or genome editing system is introduced into plants via transient transformation.

[0396] In some embodiments, the components of the genomic editing system are introduced into the plant simultaneously. In some embodiments, the components of the genomic editing system are introduced into the plant separately or sequentially.

[0397] In some implementations, the method includes the following steps:

[0398] 1) Transform components i)-iv) of the Cre-Lox system or genome editing system into isolated plant cells or tissues to obtain plant cells or tissues with a first exogenous nucleotide sequence containing recognition sites (RS) of one or more recombinases;

[0399] Optionally, 2) the component v) of the Cre-Lox system or genome editing system is converted into the plant cells or tissues obtained in step 1), thereby obtaining plant cells or tissues containing the inserted second exogenous nucleotide sequence;

[0400] Optionally, 3) converting components vi)-vii) of the Cre-Lox system or genome editing system into the plant cells or tissues obtained in step 1) or step 2); and

[0401] 3) Regenerate complete plants from plant cells or tissues obtained in step 1), step 2), or step 3).

[0402] In some embodiments, the exogenous nucleotide sequence is inserted into a safe harbor site in the plant genome, the safe harbor site being located in the plant genome.

[0403] 1) At least 5 kb away from the protein-coding region;

[0404] 2) At least 30kb away from the miRNA coding region;

[0405] 3) At least 20 kb away from the lncRNA coding region;

[0406] 4) At least 20kb away from the tRNA coding region;

[0407] 5) At least 5 kb away from the promoter and / or enhancer;

[0408] 6) The distance from the LTR repeat should be at least 20kb;

[0409] 7) Distance from non-LTR repeats must be at least 200 bp; and

[0410] 8) The distance from the centromere should be at least 10 kb.

[0411] In some embodiments, the introduction includes transforming the Cre-Lox system or genome editing system of the present invention into isolated plant cells or tissues, and then regenerating the transformed plant cells or tissues into complete plants. Preferably, no selectants targeting the selection genes carried on the expression vector are used during tissue culture.

[0412] In other embodiments, the Cre-Lox system or genome editing system of the present invention can be transformed into specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young spikelets, or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate through tissue culture.

[0413] In some embodiments of the present invention, in vitro expressed proteins and / or in vitro transcribed RNA molecules (e.g., the expression construct is an in vitro transcribed RNA molecule) and / or donor DNA molecules are directly transformed into the plant.

[0414] In some embodiments of the invention, the site-directed modification, such as a site-directed insertion of a foreign nucleotide sequence and / or the target sequence is associated with plant traits such as agronomic traits, thereby the site-directed modification, such as a site-directed insertion, results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.

[0415] In some embodiments, the method further includes the step of screening plants with desired site modifications, such as site insertions, and / or desired traits, such as agronomic traits.

[0416] In some embodiments of the invention, the method further includes obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring have the desired modification (e.g., site-directed exogenous polynucleotide insertion) and / or desired traits such as agronomic traits.

[0417] In another aspect, the present invention also provides genetically modified plants or their offspring or portions thereof, wherein the plants are obtained by the methods described above. Preferably, the genetically modified plants or their offspring have desired genetic modifications (such as site-directed exogenous polynucleotide insertion) and / or desired traits such as agronomic traits.

[0418] In another aspect, the present invention also provides a plant breeding method, comprising crossing a genetically modified first plant obtained by the method described above with a second plant that does not contain the modification, thereby introducing the modification (such as site-directed exogenous polynucleotide insertion) into the second plant. Preferably, the genetically modified first plant and the second plant have desired traits, such as agronomic traits.

[0419] Plants that can be site-directed modified, such as those with site-directed insertion of exogenous nucleotide sequences, using the Cre-Lox system or genome editing system of the present invention include monocotyledonous and dicotyledonous plants, for example, crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0420] In another aspect, the present invention provides a method for producing genetically modified plants, said genetically modified plants comprising site-directed insertion of a foreign nucleotide sequence, said method comprising inserting the foreign nucleotide sequence into a safe harbor site in the plant genome.

[0421] 1) At least 5 kb away from the protein-coding region;

[0422] 2) At least 30kb away from the miRNA coding region;

[0423] 3) At least 20 kb away from the lncRNA coding region;

[0424] 4) At least 20kb away from the tRNA coding region;

[0425] 5) At least 5 kb away from the promoter and / or enhancer;

[0426] 6) The distance from the LTR repeat should be at least 20kb;

[0427] 7) Distance from non-LTR repeats must be at least 200 bp; and

[0428] 8) The distance from the centromere should be at least 10 kb.

[0429] <Methods for Site-Specific and Scarless Modification of Human or Non-Human Animal Genomes>

[0430] On the other hand, the present invention provides a method for site-specific, traceless modification of the genome of a human or non-human animal, comprising introducing the Cre-Lox system or genome editing system of the present invention into at least one of the human or non-human animal cells. The site-specific modification includes the site-specific insertion of a large exogenous nucleotide sequence, and the insertion is precise and traceless, excluding vector backbone and recombination site sequence remnants. In some embodiments, the site-specific modification further includes large deletions, substitutions, inversions, and translocations of at least a portion of the genome, and the modification is precise and traceless, excluding recombination site sequence remnants.

[0431] On the other hand, this invention provides a method for site-specific, traceless modification of human or non-human animal genomes for in vivo and in vitro gene therapy. This allows for the deletion, addition, upregulation, downregulation, inactivation, activation, or mutation correction of disease-related genes, thereby achieving disease prevention and / or treatment. For example, the target nucleic acid region described in this invention can be located within the protein-coding region of a disease-related gene, or, for example, within a gene expression regulatory region such as a promoter or enhancer region, thereby enabling modification of the function or expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes modification of the disease-related gene itself (e.g., protein-coding region) and modification of its expression regulatory regions (e.g., promoters, enhancers, introns, etc.).

[0432] Therefore, the present invention also provides a method for treating a disease in a subject in need, comprising delivering an effective amount of the Cre-Lox system or genome editing system of the present invention to the subject to modify a gene associated with the disease. The present invention also provides the use of the Cre-Lox system or genome editing system in the preparation of a pharmaceutical composition for treating a disease in a subject in need, wherein the Cre-Lox system or genome editing system is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject in need, comprising the Cre-Lox system or genome editing system of the present invention, and optionally a pharmaceutically acceptable vector, wherein the Cre-Lox system or genome editing system is used to modify a gene associated with the disease. In some embodiments, the subject is a human being.

[0433] Example

[0434] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.

[0435] Example 1: Developing a rapid optimization platform for recombination sites to obtain combinations of recombination site variants with high recombination activity and significantly reduced reversible recombination activity.

[0436] 1. This embodiment envisions a strategy for rapidly optimizing and modifying recombination sites. Specifically, primers are used to introduce saturation mutations of the Lox sequence into two plasmids containing Lox sites, which are then transformed into protoplast cells along with Cre expression plasmids. The recombinant plasmids are then subjected to PCR amplification and deep sequencing to obtain the base preference of Cre recombinase for the Lox site. Furthermore, by designing symmetrical (reverse repeat) and asymmetrical Lox variants flanking the Lox site, a fluorescent reporter system is used to rapidly screen combinations of recombination site variants with high recombination activity and low reversible recombination activity. The specific process is shown in Figure 1.

[0437] 2. According to the process shown in Figure 1, this embodiment first constructs plasmids P1 and P2. Plasmid P1 contains a Lox site and other sequences. Using primers, a saturation mutation is introduced into the LA part (the first half of the central sequence, i.e., the 5' end) of the Lox site (LoxP, whose nucleotide sequence is shown in SEQ ID NO:1). The resulting Lox site variant is called Lox-smLA. Similarly, a saturation mutation is introduced into the RA part (the second half of the central sequence) of the other Lox site in plasmid P2. The resulting Lox site variant is called Lox-smRA. Plasmids P1 and P2, together with the Cre expression plasmid (whose amino acid sequence is shown in SEQ ID NO:2), are transformed into protoplast cells. After recombination, a plasmid containing both LA and RA double-mutated Lox sequences is obtained, called Lox-dsm, and a plasmid containing LoxP is obtained, as shown in Figure 2.

[0438] PCR amplification and deep sequencing of the double-mutant Lox site on the Lox-dsm plasmid can yield the different base abundances at each position on the Lox site, as shown in Figures 3A and 3B.

[0439] Further analysis revealed the base preference of Cre recombinase for the Lox site, as shown in Figure 4.

[0440] 3. The base preference results of the Lox site showed that the bases in the LA part were more conserved, while the bases in the RA part were more loose. Therefore, this example is based on two LA variants, including Lox71, which had the highest efficiency in the previous study, and LoxPF and LoxPR designed according to base preference. Corresponding symmetric (RA and LA reverse repeat) variant combinations were designed, including Scm1 (LoxPF+LoxPSR), Scm2 (Sm1Lox66+Sm1Lox71), Scm3 (Sm2Lox66+Sm2Lox71), Scm4 (Sm3Lox66+Sm3Lox71), and asymmetric (RA and LA misplacement mutation) variant combinations, including Acm1-6 (LoxPF+LoxAR1-6), Acm7-12 (Lox71+LoxAR1-6), and the LoxPF+LoxPR combination (as a control). See Table 1 below for details. The corresponding recombinant active fluorescent reporter system and reversible recombinant active fluorescent reporter system were constructed. Specifically, taking the LoxPF and LoxPR combination as an example, the recombinant reporter system is constructed by building LoxPF and LoxPR onto plasmids containing GFP-N and GFP-C, respectively (see Sun, Chao et al. "Precise integration of large DNA sequences in plant genomes using PrimeRoot editors." Nature biotechnology vol.42,2(2024):316-327.doi:10.1038 / s41587-023-01769-w). The reversible reporter system is constructed by building the double-mutant LoxPF / PR sequence and the wild-type LoxP sequence onto plasmids containing GFP-N and GFP-C, respectively, as shown in Figures 5A and 5B.

[0441] Table 1:

[0442] The constructed recombination site variant combinations were characterized using a reporter system to assess recombination activity and reversible recombination activity. The results were compared with those of the Bxb1 system (see Thomson, James G et al. “The Bxb1 recombination system demonstrates heritable transmission of site-specific excision in Arabidopsis.” BMC biotechnology vol.12 9.21 Mar.2012, doi:10.1186 / 1472-6750-12-9) and the previously demonstrated Lox71 / Lox66 (see Albert, H et al. “Site-specific integration of DNA into wild-type and mutant lox sites placed in the plant genome.” The Plant journal: for cell and molecular biology vol.7,4(1995):649-59. doi:10.1046 / j.1365-313x.1995.7040649.x) and Lox71 / LoxKR3 (see Araki, Kimi et al.). The results of the comparisons between the Lox71 / LoxTC9 combination (see Liu, Bo et al. "Comparative analysis of right element mutant lox sites on recombination efficiency in embryonic stem cells." BMC biotechnology vol.10 29.31 Mar.2010, doi:10.1186 / 1472-6750-10-29) and the Lox71 / LoxTC9 combination (see Liu, Bo et al. "Large-scale multiplexed mosaic CRISPR perturbation in the whole organism." Cell vol.185,16(2022):3008-3024.e16. doi:10.1016 / j.cell.2022.06.039) and flow cytometry analysis are shown in Figure 6.

[0443] The above results indicate that rationally designed Lox variant combinations based on Lox site base preferences can significantly reduce reversible recombination activity while maintaining high recombination efficiency. The best-performing variant combination is Acm8, i.e., Lox71 + LoxAR2, which has a recombination efficiency comparable to Lox71 / Lox66, but with reversible recombination activity reduced by more than tenfold, approaching that of the Bxb1 system. Furthermore, LoxPF / LoxPR exhibits significantly higher recombination activity than Lox71 / Lox66, while the Scm2, Scm3, and Scm4 combinations, as well as the Acm3–12 combinations, all show significantly lower reversible recombination activity than Lox71 / LoxTC9.

[0444] Example 2: Using an artificial intelligence-assisted protein optimization method, Cre point mutations were designed to significantly improve their recombination activity.

[0445] 1. In this embodiment, ePPEplus is used to replace the ePPE component in the PrimeRoot system (see PrimeROOT.v2C as described in WO / 2023 / 227050). A point mutation of Cre is designed using an artificial intelligence-assisted protein optimization method. The construct is shown in Figure 7, and specifically includes: c-Myc NLS—BP-SV40 NLS—Cas9(R211K,N394K,H840A)—XTEN linker—NC—SV40 NLS—32aa linker—MLV-ΔRNase(V223A)—vBP-SV40 NLS—SV40 NLS—32aa linker—Cre mutant—c-Myc NLS—BP-SV40 NLS. The specific sequences of each element are shown in the sequence information section.

[0446] The Cre mutants designed and detected in this embodiment are shown in Table 2 below:

[0447] Table 2:

[0448] The amino acid positions in Table 2 above refer to SEQ ID NO:2.

[0449] Using the previously developed PrimeRoot one-step fluorescence reporter system (i.e., all-in-one, AR system, see Sun, Chao et al. "Precise integration of large DNA sequences in plant genomes using PrimeRoot editors." Nature biotechnology vol.42,2(2024):316-327. doi:10.1038 / s41587-023-01769-w), the Lox66 sequence to be inserted in the AR system was changed to the LoxAR2 sequence for rapid identification of Cre variants. A total of 19 variants were identified in the first batch, among which two variants, D141Q and R243D, showed significantly improved efficiency, which was 1.8 and 2.3 times that of wild-type Cre, respectively. The results are shown in Figure 8A.

[0450] 2. In this embodiment, a second batch of 23 Cre point mutants were further tested. Among them, eight variants, namely H40E, R241P, H269S, I272L, A275P, K276P, Q281G, and T332A, showed significant improvements in efficiency, which were 1.7 to 2.3 times that of the wild-type variants, as shown in Figure 8B.

[0451] 3. The efficient variants obtained in steps 1 and 2 were further combined and tested to obtain a variant combination with an efficiency up to 3.5 times that of the wild type, as shown in Figure 8C.

[0452] 4. Further, the Cre variants containing the highly efficient variant combinations cm2, cm4, cm23, cm25, and cm24 were subjected to large-fragment DNA efficiency tests at rice endogenous gene target sites. A total of 6 target sites were tested, and the efficiency of inserting GFP (720 bp) and Act1-GFP (2.4 Kb) was tested respectively. The efficiency was detected by ddPCR. The results showed that the Cre mutation combination optimized at the endogenous target sites could achieve an efficiency of up to 3.2 times that of the wild type. The results are shown in Figure 9.

[0453] The above results indicate that AI-assisted protein optimization methods can significantly improve the recombination efficiency of Cre proteins, thereby greatly enhancing the integration efficiency of large DNA fragments.

[0454] Example 3: Integrating optimized components of the Cre-Lox system to construct a PCE system

[0455] 1. In this embodiment, the optimized components Lox variant and Cre variant of Cre-Lox obtained in Examples 1 and 2 are integrated into the PrimeRoot system disclosed in WO / 2023 / 227050. Specifically, two epigRNAs guiding the editing system are expressed using the compound type II promoter GS (pGS) to insert into the recombination site LoxAR2 (the design principles of the two epigRNAs are as described in the previous (pegRNA) section, and are also the same as the design principles of the first and second pegRNAs in WO / 2023 / 227050, hereinafter referred to as the dual-pegRNA design principle), forming a PCE version. The construct is shown in Figure 10, wherein the vector for expressing the fusion protein specifically includes:

[0456] ubi-P promoter—c-Myc NLS—BP-SV40 NLS—Cas9(R211K,N394K,H840A)—XTEN linker—NC—SV40 NLS—32aa linker—MLV-ΔRNase(V223A)—vBP-SV40 NLS—SV40 NLS—32aa linker—Cre mutant (cm24)—c-Myc NLS—BP-SV40 NLS;

[0457] The vectors for epegRNA specifically include:

[0458] pGS promoter—tRNA—Lox-pegRNA—tevopre—HDV—HSPT;

[0459] For the specific sequence of each component, please refer to the sequence information section.

[0460] The ability of PCE and PrimeRoot.v2C-Cre (hereinafter referred to as PrimeRoot.v2) as described in WO / 2023 / 227050 to insert a 18.8Kb ultra-large DNA fragment was tested in rice. The insertion efficiency was tested by ddPCR at six endogenous target sites. The results showed that the efficiency of PCE in integrating the 18.8Kb large DNA fragment was up to 4.6 times that of PrimeRoot.v2, as shown in Figure 11.

[0461] 2. In this embodiment, the working ability of PCE was further tested in corn and wheat, respectively. The ability to insert 720bp and 18.8Kb at 6 endogenous target sites was tested with PrimeRoot.v2. The ddPCR results showed that the efficiency in corn was up to 3.2 times that of PrimeRoot.v2, and the efficiency of inserting 18.8Kb in wheat was up to 11 times that of PrimeRoot.v2. The results are shown in Figure 12.

[0462] The results above demonstrate the significant leap in editing efficiency and scale of the optimized PCE compared to PrimeRoot.v2 disclosed in WO / 2023 / 227050, proving its potential for precise manipulation of larger-scale DNA in the genome.

[0463] Example 4: Precise manipulation of large DNA fragments in plants using a PCE system, including targeted deletion, substitution, inversion, and translocation.

[0464] 1. Using PCE for targeted and precise inversion of large-scale DNA in the genome.

[0465] In this embodiment, three scales—10Kb, 100Kb, and 500Kb—were selected to test targeted inversions in four rice genome regions. Specifically, two pairs of pegRNAs—one forward-directed LoxAR2 and the other reverse-directed Lox71—were designed at both ends of each region. RT represented sequences containing LoxAR2 and Lox71, respectively. The spacer and PBS were adjusted according to the genomic targeting sequence. The entire region was inverted using a PCE system containing these two pairs of pegRNAs. ddPCR was used to detect the efficiency; the 10Kb inversion efficiency was high, and the 500Kb inversion also showed some efficiency. The results are shown in Figure 13.

[0466] Previous studies have shown that the generation of two DSBs on the genome can also mediate the inversion of the intermediate DNA sequence through the NHEJ repair pathway (Schmidt, Carla et al. "Efficient induction of heritable inversions in plant genomes using the CRISPR / Cas system." The Plant journal: for cell and molecular biology vol.98,4(2019):577-589. doi:10.1111 / tpj.14322). Therefore, the method in that paper was adopted, and Cas9 target sites were designed at the same positions of the inversion regions targeted by the PCE system in step 1. The results were compared with the PCE system. ddPCR identification showed that the inversion efficiency of PCE was up to 17.6 times that of the NHEJ strategy, as shown in Figure 14.

[0467] Further, target sites were designed before and after the promoters of some functional genes, as well as before and after promoters of stronger or weaker expression nearby, to test inversion efficiency. Specifically, four genes were selected as targets, and five regions (named 27TB1, 35Kb, 54Kb, 180Kb, and 315Kb, respectively) were selected upstream or downstream of them at distances of 27Kb, 35Kb, 54Kb, 180Kb, and 315Kb for inversion testing. The results were compared with PrimeRoot.v2, and the results are shown in Figure 15.

[0468] 2. Use PCE for targeted and precise replacement of DNA at the kb level.

[0469] By inserting two orthogonal LoxAR2 sites with different central sequences on either side of a region of the genome at a certain distance (see Lee, G, and I Saito. "Role of nucleotide sequences of loxP spacer region in Cre-mediated recombination." Gene vol. 216, 1(1998): 55-65. doi: 10.1016 / s0378-1119(98)00325-4), and providing exogenous donors with orthogonal Lox71 sites containing the corresponding central sequences on both sides of the DNA fragment to be replaced, the PrimeRoot system can be used to accurately replace the DNA in this region with the desired replacement DNA fragment. Target sites with intervals of 1.5Kb, 2.5Kb, and 5Kb were designed in the two regions, and the middle sequence was replaced with the provided DNA sequence of 2.4Kb or 4.9Kb. ddPCR detection showed that highly efficient and accurate replacement of up to 5Kb of DNA could be achieved, as shown in Figure 16.

[0470] 3. Use PCE for targeted and precise deletion and translocation at the chromosome level.

[0471] Precise deletion of intermediate sequences can be mediated by designing target sites on both sides of a large DNA fragment in the genome and inserting unidirectional recombination sites. For example, in this embodiment, pegRNA was designed and inserted into LoxAR2 and Lox71 on both sides of the centromere of rice chromosomes 2 and 9, respectively. Precise deletion of up to 4Mb was detected by PCE. Precise deletion events can be detected by junction PCR. The first-generation sequencing results are shown in Figure 17. There are no overlapping peaks, indicating that it is a completely precise deletion.

[0472] Further ddPCR was used to detect efficiency, and the results are shown in Figure 18.

[0473] Furthermore, by designing target sites on two different chromosomes and inserting LoxAR2 and Lox71 respectively, precise translocations between chromosomes could be achieved under the action of PCE. Specifically, pegRNAs were designed and inserted into the long arms of rice chromosomes 1, 4, 6, 7, 8, and 9, as well as the broken arms of chromosomes 2 and 7, respectively, to generate six translocation events. After transformation into protoplasts, precise translocation events were detected by junction PCR and first-generation sequencing, as shown in Figure 19. Similarly, ddPCR was used to quantify the translocation events, and the results are shown in Figure 20.

[0474] 4. Creation of Rice Variants with Precise Inversion of Large DNA Fragments Using PCE. Using Lox-pegRNAs with a 23bp reverse transcription template, we transformed rice plants and analyzed the regenerated plants by ligation site PCR and Sanger sequencing, as shown in Figure 21. The PCE system achieved precise inversion rates of 26.2% and 20.1% in inv-35W and inv-315H, respectively, significantly higher than the 3.7% and 2.8% of the PrimeRoot system, as shown in Figure 22. This indicates that the evolutionarily modified, nearly irreversible Lox site is highly effective in creating single-plant mutants. Analysis of the T1 generation plants showed that approximately one-quarter of the progeny were identified as homozygous inversion mutants by ligation site PCR targeting the inversion mutation and the wild-type OsHIS1 site. Further T-DNA primer analysis confirmed that some homozygous mutants did not contain exogenous transgenes, as shown in Figure 23. These results demonstrate the potential of the PCE system in achieving targeted large-scale DNA inversions, reversing natural genomic inversions, and enhancing desirable plant traits through gene expression modification.

[0475] Example 5: Development of a guided editing system without PAM limitations, enabling efficient and precise insertion of recombination sites.

[0476] To address the PAM-free limitation and enable efficient and precise insertion of recombination sites, this embodiment first requires the development of a PAM-free guided editing system capable of efficiently and precisely inserting recombination sites. The chosen Cas9 variant SpRY, which significantly broadens the PAM (see Walton, Russell T et al. “Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants.” Science (New York, NY) vol. 368, 6488(2020): 290-296. doi:10.1126 / science.aba8853), and two anCas proteins, BCA and FCA, reconstructed during Cas9 evolution (see Alonso-Lerma, Borja et al. “Evolution of CRISPR-associated endonucleases as inferred from resurrected proteins.” Nature microbiology). vol.8,1(2023):77-90.doi:10.1038 / s41564-022-01265-y), in vitro experiments have demonstrated that both have the potential to achieve PAM-free processing. Furthermore, their research showed that although their double-stranded DNA cleavage activity is not as high as Cas9, they have strong nick activity, thus potentially making them suitable as a targeting module in guided editing systems. Secondly, this study found four major PAM-related mutations in both FCA and Cas9; therefore, this embodiment will also introduce these four mutations (K1107D, S1109T, W1126L, K1334) into Cas9. M) To detect whether it can broaden its target range, guide editing systems SpRY-ePPEplus, BCA-ePPEplus, FCA-ePPEplus, and FCACas9-ePPEplus were constructed based on ePPEplus, which introduced R221K and N394K, BCA, FCA, and Cas9-FCA point mutations, respectively. These systems were compared with ePPEplus and the guide editing system ePPE(SpRY) constructed based on SpRY-nCas9, which can broaden PAM, as disclosed in WO / 2023 / 227050. The specific vector construction strategy is shown in Figure 24.

[0477] The above six constructs were transformed into rice protoplasts using PEG transformation. The efficiency of LoxAR2 insertion into 16 randomly designed PAMs at four positions (NAA, NAT, NAC, NAG, NTA, NTT, NTC, NTG, NCA, NCT, NCC, NCG, NGA, NGT, NGC, and NGG, where N represents A, T, G, or C) was tested. The design principle was the same as that of dual-pegRNA. The results showed that SpRY-ePPEplus had a high Lox site insertion efficiency (up to 20%) in almost all PAMs (except NTT, which is a PAM that comes with the target sequence on the plasmid and may be affected by plasmid self-cleavage). The next-generation sequencing results are shown in Figure 25.

[0478] The above results demonstrate that SpRY-ePPEplus can perform highly efficient and precise insertion of dual-pegRNA-mediated recombination sites with almost no PAM restrictions.

[0479] Example 6: Development of a RePCE system that is not limited by PAM and can perform precise manipulation of large DNA fragments without leaving a trace.

[0480] 1. To achieve seamless large-fragment DNA manipulation, this embodiment envisions using the PAM-free PrimeRoot system constructed in Example 5. Leveraging its inherently efficient guided editing capabilities, pegRNAs (Re-pegRNAs) are designed to re-edit residual Lox sites on the edited genome according to the dual-pegRNA design principle, thereby achieving precise and seamless large-fragment DNA manipulation. Specifically, two pegRNAs are designed flanking each residual Lox site. The RT template can be the original genomic sequence or other sequences, thus using the dual-pegRNA strategy to eliminate residual Lox sites, as illustrated in Figure 26.

[0481] First, based on SpRY-ePPEplus and the above concept, the large-fragment DNA manipulation system RePCE without PAM restrictions was constructed. The vector construction is shown in Figure 27.

[0482] The ddPCR method was used to test its large DNA fragment insertion efficiency outside of NGG PAM. Four different PAMs were randomly tested, including NGA, NAT, NTA, and NGC. The results are shown in Figure 28.

[0483] The above results demonstrate the effectiveness of RePCE in large-fragment DNA editing on different PAMs, proving its potential to be PAM-free.

[0484] 2. Further based on the AR system described above, a reporter system SPR using a traceless large-fragment DNA insertion strategy was constructed. Specifically, the GFP expression frame was divided into three parts: G, F, and P. The GFP-N part in the AR reporter system was deleted, G and P were placed on either side of the target site, and the NGG target site was replaced with a non-NGG target site. The F part replaced the GFP-C part, and another Lox71 sequence was added to the plasmid to complete the vector-backbone-free F part recombination. At the same time, a Re-pegRNA was designed to repair the Lox sequence on the recombinated GFP to complete the traceless editing. The SPR luminescence after transformation of protoplasts is shown in Figure 29.

[0485] Fluorescence identification proved the effectiveness of the RePCE system in performing scarless editing. Therefore, the proportion of scarless editing when performing precise inversion and insertion of large DNA fragments at endogenous targets was further tested. Re-pegRNAs were designed for two large DNA fragment inversion events and three large DNA fragment insertion events, respectively. The results were transformed into rice protoplasts and the scarless editing ratio was detected by amplicon next-generation sequencing. The results are shown in Figure 30.

[0486] This embodiment first tested the efficiency of PAM Re-pegRNA at Lox sites positions 4-6 and 29-31, or positions 3-5 and 30-32 (with the first base at the 5' end of the Lox site as position 1) for scarless editing of Lox sites. The results showed that the efficiency of C-terminal re-editing of the Lox71 / LoxAR2 double mutant sequence was as high as over 90%, but the efficiency of N-terminal re-editing of the LoxP sequence was extremely low. It is speculated that the Cre protein (cm24) binds to the wild-type LoxP sequence, causing steric hindrance and hindering the occurrence of guided editing. Therefore, the efficiency of PAM Re-pegRNA at LoxP sites positions -1 to -10 and 35 to 45 (upstream or downstream of LoxP) for scarless editing of LoxP sites was further tested. The results showed that the efficiency was also as high as over 90% (Figure 30), which proved the effectiveness of the system in scarless editing of endogenous targets.

[0487] 3. Develop a high-efficiency RePCE system mediated by SpCas9 based on the "rG" strategy.

[0488] To improve the efficiency of scarless editing, we proposed an "rG" strategy (i.e., redesigning Lox-pegRNAs and donors using GG motifs). By artificially introducing N3NGG (NNNNGG) and CCNN3 (CCNNNN) motifs into the RS and donor sequences, we can design Re-pegRNAs suitable for secondary editing of SpCas9 (SEQ ID NO:26; not SpRY-Cas9, i.e., using the recombinase and ePPEplus guided editing of the fusion protein shown in Figure 10) using the NGG and CCN PAM sequences of the Lox sites after editing. The results are shown in Figure 31. We tested this strategy at four target sites: IS-S20 and IS-8L (GFP insertion sites) and inv-35W and inv-315H (large DNA inversion sites). Protoplast transformation experiments revealed that the RePCE system mediated by the rG strategy, as assessed by ddPCR analysis (using probe-GFP / inv targeting GFP or the inversion junction), showed a two-fold increase in editing efficiency compared to SpRY-Cas9-mediated RePCE. Notably, the efficiency significantly decreased when using a probe targeting LoxP (probe-LoxP), indicating a substantial improvement in scarless editing efficiency, as shown in Figure 32. We further validated the effectiveness of the rG strategy by deep sequencing of the 5' and 3' ends of the edited events. The results showed that the scarless editing efficiency mediated by Re-pegRNA under the rG strategy approached 100%, as shown in Figure 33. These results demonstrate that the RePCE system is a highly efficient tool for scarless editing of plant chromosomes, providing a significant technological breakthrough for precise and scarless large-scale DNA manipulation.

[0489] 4. RePCE-mediated precise and scarless chromosome editing in human cells. Next, we evaluated the ability of the RePCE system to achieve scarless large-scale genomic alterations (related to genetic diseases) in human cells. Since existing Bxb1-based systems produce large DNA insertions that retain the donor vector backbone and RS sequence, we first tested the efficiency of RePCE in achieving scarless large DNA insertions in human cells. We constructed a PCE system driven by the CMV promoter (using the construct shown in Figure 10, replacing the Ubi promoter with the CMV promoter), selecting four target sites at AAVS1, ACTB, LMNB1, and HEK3, and constructed Lox-pegRNAs and Re-pegRNAs driven by the hU6 promoter (construct shown in Figure 10, replacing the pGS promoter with hU6). After transfection into HEK293T cells, ddPCR (using probe-LoxP and probe-GFP probes) was used to detect the efficiency of targeted GFP insertion. When Lox-pegRNAs and Re-pegRNAs were transfected simultaneously, the insertion efficiency detected by probe-GFP reached as high as 18.6%, while the detection result of probe-LoxP was close to zero, as shown in Figure 34. To verify the effectiveness of probe-LoxP, we performed ddPCR detection after transfecting only Lox-pegRNAs, obtaining an efficiency comparable to probe-GFP, as shown in Figure 35. This indicates that when using Re-pegRNAs, most insertion events eliminated the LoxP sequence through secondary editing. Deep sequencing at the insertion junction showed that the secondary editing efficiency of the Lox sequence exceeded 90%, as shown in Figure 36. We further explored the potential of the RePCE system in constructing large-scale structural variations related to genetic diseases without scarring. For this purpose, two typical disease models were selected: Ewing sarcoma (ES) caused by translocation of chromosomes 11 and 22 leading to EWSR1-FLI1 fusion, and non-small cell lung cancer (NSCLC) caused by 12Mb inversion leading to ALK-EML4 fusion on chromosome 2. We designed Lox-pegRNAs to insert LoxAR2 into EWSR1 and ALK, Lox71 into FLI1, and rLox71 into EML4, as well as Re-pegRNAs for secondary editing of Lox sites, aiming to achieve seamless gene fusion of exon regions. After transfection, ddPCR analysis showed chromosomal translocation and inversion efficiencies of 1.9% and 1.2%, respectively, as shown in Figure 37. Deep sequencing revealed a seamless editing efficiency of up to 90% at the 5' and 3' ends of the edited regions, as shown in Figure 38. Ligation PCR and Sanger sequencing of the monoclonal PCR products confirmed the precise and seamless chromosomal translocation and inversion events mediated by the PCE system, as shown in Figure 39.These results highlight the potential of the RePCE system to reverse large-scale DNA structural variations at the chromosomal level through precise, scarless editing.

[0490] Sequence information

[0491] LoxP (SEQ ID NO:1)

[0492] Cre(SEQ ID NO:2)

[0493] The recombination sites or recombination site variant sequences used or constructed in this invention are as follows:

[0494] Figure 7 shows the nucleotide / amino acid sequences of the elements involved in the construct: nucleotide sequence of the ubi-P promoter (SEQ ID NO:23)

[0495] The amino acid sequence of c-Myc NLS (SEQ ID NO:24)

[0496] Amino acid sequence of BP-SV40 NLS (SEQ ID NO:25)

[0497] The amino acid sequence of Cas9 (R211K, N394K, H840A) (SEQ ID NO:26)

[0498] The amino acid sequence of the XTEN linker (SEQ ID NO:27)

[0499] The amino acid sequence of NC (SEQ ID NO:28)

[0500] The amino acid sequence of SV40 NLS (SEQ ID NO:29)

[0501] The amino acid sequence of vBP-SV40NLS (SEQ ID NO:47)

[0502] The amino acid sequence of the 32aa linker (SEQ ID NO:30)

[0503] The amino acid sequence of MLV-ΔRNase (V223A) (SEQ ID NO:31)

[0504] ePPEplus guides the editing of the fusion protein's amino acid sequence (SEQ ID NO:46):

[0505] Figure 10 shows the nucleotide / amino acid sequences of the elements involved in the construct:

[0506] The nucleotide sequence of the pGS promoter (SEQ ID NO:32)

[0507] The nucleotide sequence of tRNA (SEQ ID NO:33)

[0508] The nucleotide sequence of tevopre (SEQ ID NO:34)

[0509] The nucleotide sequence of HDV (SEQ ID NO:35)

[0510] The nucleotide sequence of HSPT (SEQ ID NO:36)

[0511] The nucleotide sequence of the pegRNA scaffold (SEQ ID NO:37)

[0512] The sequence of Lox-pegRNA (taking LoxAR2 as an example):

[0513] LoxAR2-pegRNA-1 (SEQ ID NO:38)

[0514] LoxAR2-pegRNA-2 (SEQ ID NO:39)

[0515] Note: In the above Lox-pegRNA sequences, a single underline indicates the selectable target spacer sequence, a double underline indicates the selectable target PBS sequence, a dotted underline indicates the LoxAR2 sequence on both pegRNAs, and N represents A, T, G, or C.

[0516] The amino acid sequence of the Cas protein used in Example 5:

[0517] The amino acid sequence of Cas9 (R211K, N394K, H840A) in ePPEplus is shown in SEQ ID NO:26.

[0518] The amino acid sequence of SpRY-Cas9(H840A) in ePPE(SpRY) is disclosed in WO / 2023 / 227050.

[0519] SpRY-Cas9 (R211K, N394K, H840A) (SEQ ID NO:40) in SpRY-ePPEplus:

[0520] The amino acid sequence of BCA (R221K, H838A) in BCA-ePPEplus (SEQ ID NO:41)

[0521] The amino acid sequence of FCA (R221K, H838A) in FCA-ePPEplus (SEQ ID NO:42)

[0522] The Cas9 (R221K, N394K, H840, FCA*) (SEQ ID NO:43) in FCACas9-ePPEplus:

[0523] SpRY-ePPEplus guides the editing of the amino acid sequence of the fusion protein (SEQ ID NO:44):

[0524] The double-underlined portion represents SpRY-Cas9(R211K,N394K,H840A). The amino acid sequences of the BCA-ePPEplus, FCA-ePPEplus, and FCACas9-ePPEplus guided editing fusion proteins can be obtained by replacing the SpRY-Cas9(R211K,N394K,H840A) portion in the SpRY-ePPEplus guided editing fusion protein with the corresponding Cas9 protein sequences disclosed above.

[0525] The amino acid sequence (SEQ ID NO:45) of the recombinase and the fusion protein SpRY-ePPEplus-cm24, which is the guide editing fusion protein:

[0526] The single underlined part is cm24; the double underlined part is SpRY-Cas9(R211K,N394K,H840A).

[0527] The amino acid sequences of other recombinases and guide editing fusion proteins can be obtained by replacing mutants of the guide editing fusion protein (or the Cas protein therein) and / or Cre recombinase.

[0528] The sequences involved in Figure 2:

[0529] Lox-smLA(SEQ ID NO:48):NNNNNNNNNNNNNGCATACATTATACGAAGTTAT

[0530] Lox-smRA(SEQ ID NO:49):ATAACTTCGTATAGCATACATNNNNNNNNNNNNNN

[0531] Lox-dsm:NNNNNNNNNNNNNGCATACATNNNNNNNNNNNNN

[0532] LoxP(SEQ ID NO:1):ATAACTTCGTATAGCATACATTATACGAAGTTAT

[0533] The sequence involved in Figure 17 (SEQ ID NO:50):

[0534] The sequence involved in Figure 19 (SEQ ID NO:51):

[0535] The sequences involved in Figure 39:

[0536] From top to bottom:

[0537] The CMV promoter sequence used in the PCE system in the human cell (SEQ ID NO:60):

[0538] The hU6 promoter sequence (SEQ ID NO:61) used in the PCE system in the human cell:

[0539] It should be noted that although the technical solution of the present invention has been described with specific examples, those skilled in the art will understand that the present invention should not be limited thereto.

[0540] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A Cre-Lox system, wherein, The Cre-Lox system includes a recombinase and a recognition site (RS) for the recombinase; The recombinase is a mutant of Cre recombinase, and the mutant of Cre recombinase is selected from any one of the following groups (a1)-(a3): (a1) The amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has R72V, R72I, R72L, Q94M, R101L, D141Q, R146I, R159L, R159M, R159V, R159W, K211I, R243D, T268I, A296I, G313F, G313M, G313W, S257P, Q35D, E39 Mutations of P, H40E, W63P, K86E, Q94A, N111P, M117L, Q156R, R173P, I197V, I225L, S226E, R241P, H269S, I272L, A275P, K276P, Q281G, M299L, T316R, D329P, T332A or any combination thereof; (a2) is a polypeptide having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the amino acid sequences shown in (a1); (a3) A polypeptide with an amino acid sequence as shown in (a1) or (a2) having one or more amino acids added or deleted at at least one end of the N-terminus and C-terminus.

2. The Cre-Lox system according to claim 1, wherein, The amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the mutations shown in 1) to 69) below: 66)D141Q, H40E, A275P, K276P, Q281G; 1) R72V; 2) R72I; 3) R72L; 4) Q94M; 5) R101L; 6)D141Q; 7)R146I; 8) R159L; 9)R159M; 10) R159V; 11)R159W; 12)K211I; 13)R243D; 14)T268I; 15)A296I; 16)G313F; 17)G313M; 18)G313W; 19)S257P; 20)Q35D; 21)E39P; 22)H40E; 23)W63P; 24)K86E; 25)Q94A; 26)N111P; 27)M117L; 28)Q156R; 29)R173P; 30)I197V; 31)I225L; 32)S226E; 33)R241P; 34)H269S; 35)I272L; 36) A275P; 37) K276P; 38)Q281G; 39)M299L; 40)T316R; 41)D329P; 42)T332A; 43)D141Q, R243D; 44)D141Q, K276P; 45)D141Q, G313F; 46) R243D, K276P; 47)Q94M, D141Q; 48) Q94M, R243D; 49) Q94M, K276P; 50)R159M, R243D; 51) R159M, K276P; 52) G313F, K276P; 53)E39P, H40E; 54)S226E, R241P; 55) I272L, A275P; 56) K276P, Q281G; 57) M299L, T332A; 58)D141Q, R243D, K276P; 59) I225L, S226E, R241P; 60) I272L, A275P, K276P; 61) A275P, K276P, Q281G; 62)Q156R, S226E, R241P; 63)Q156R, I225L, S226E, R241P; 64)I272L, A275P, K276P, Q281G; 65)H40E, A275P, K276P, Q281G; 67)D141Q, I272L, A275P, K276P, Q281G; 68)R243D, I272L, A275P, K276P, Q281G; 69)R243D, H40E, A275P, K276P, Q281G; Preferably, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations: 6)D141Q; 13)R243D; 22)H40E; 33)R241P; 34)H269S; 35)I272L; 36) A275P; 37) K276P; 38)Q281G; 42)T332A; Preferably, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations: 66)D141Q, H40E, A275P, K276P, Q281G; 43)D141Q, R243D; 44)D141Q, K276P; 45)D141Q, G313F; 46) R243D, K276P; 47)Q94M, D141Q; 48) Q94M, R243D; 49) Q94M, K276P; 50)R159M, R243D; 51) R159M, K276P; 52) G313F, K276P; 53)E39P, H40E; 55) I272L, A275P; 56) K276P, Q281G; 57) M299L, T332A; 58)D141Q, R243D, K276P; 59) I225L, S226E, R241P; 60) I272L, A275P, K276P; 61) A275P, K276P, Q281G; 63)Q156R, I225L, S226E, R241P; 64)I272L, A275P, K276P, Q281G; 65)H40E, A275P, K276P, Q281G; 67)D141Q, I272L, A275P, K276P, Q281G; 68)R243D, I272L, A275P, K276P, Q281G; 69)R243D, H40E, A275P, K276P, Q281G; More preferably, the amino acid sequence of the mutant Cre recombinase corresponds to the amino acid sequence shown in SEQ ID NO:2, and has any of the following mutations: 66)D141Q, H40E, A275P, K276P, Q281G; 67)D141Q, I272L, A275P, K276P, Q281G; 65)H40E, A275P, K276P, Q281G; 46) R243D, K276P; 44)D141Q、K276P.

3. The Cre-Lox system according to claim 1 or 2, wherein, The recognition site (RS) of the recombinase is selected from any one of the following or any combination thereof: (a) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:9, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:9; (b) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:10, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:10; (c) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:11, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:11; (d) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:12, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:12; (e) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:13, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:13; (f) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:14, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:14; (g) Contains a nucleotide sequence as shown in SEQ ID NO:15, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:15; (h) Contains a nucleotide sequence as shown in SEQ ID NO:16, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:16; (i) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:17, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:17; (j) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:3, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:3; (k) Contains a nucleotide sequence as shown in SEQ ID NO:4, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:4; (l) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:5, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:5; (m) contains a nucleotide sequence as shown in SEQ ID NO:6, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:6; (n) contains a nucleotide sequence as shown in SEQ ID NO:7, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:7; (o) A nucleotide sequence comprising the nucleotide sequence shown in SEQ ID NO:8, or having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:8; (p) Contains a nucleotide sequence as shown in SEQ ID NO:20, or a nucleotide sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in SEQ ID NO:20; Preferably, the recognition site (RS) of the recombinase is selected from the following combinations: (L) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:13; (A) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:11; (B) Contains the nucleotide sequence shown in SEQ ID NO:3 and the nucleotide sequence shown in SEQ ID NO:6; (C) Contains the nucleotide sequence shown in SEQ ID NO:4 and the nucleotide sequence shown in SEQ ID NO:7; (D) Contains the nucleotide sequence shown in SEQ ID NO:5 and the nucleotide sequence shown in SEQ ID NO:8; (E) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:12; (F) Contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:13; (G) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:14; (H) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:15; (I) Containing a nucleotide sequence as shown in SEQ ID NO:9 and a nucleotide sequence as shown in SEQ ID NO:16; (J) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:17; (K) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:12; (M) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:14; (N) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:15; (O) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:16; (P) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:17; More preferably, the recognition site (RS) of the recombinase is selected from the following combinations: (L) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:13; (B) Contains the nucleotide sequence shown in SEQ ID NO:3 and the nucleotide sequence shown in SEQ ID NO:6; (C) Contains the nucleotide sequence shown in SEQ ID NO:4 and the nucleotide sequence shown in SEQ ID NO:7; (D) Contains the nucleotide sequence shown in SEQ ID NO:5 and the nucleotide sequence shown in SEQ ID NO:8; (G) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:14; (H) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:15; (I) Containing a nucleotide sequence as shown in SEQ ID NO:9 and a nucleotide sequence as shown in SEQ ID NO:16; (J) contains the nucleotide sequence shown in SEQ ID NO:9 and the nucleotide sequence shown in SEQ ID NO:17; (K) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:12; (M) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:14; (N) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:15; (O) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:16; (P) contains the nucleotide sequence shown in SEQ ID NO:20 and the nucleotide sequence shown in SEQ ID NO:

17.

4. A genome editing system, wherein the genome editing system comprises: i)a) An expression construct containing a CRISPR nuclease and / or a nucleotide sequence encoding the CRISPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase, or b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein comprises a CRISPR nuclease and a reverse transcriptase; ii) an expression construct containing a first pegRNA and / or a nucleotide sequence encoding the first pegRNA, and iii) An expression construct containing a second pegRNA and / or a nucleotide sequence encoding the second pegRNA. The first pegRNA contains, from 5' to 3', a first guide sequence, a first scaffold sequence, a first reverse transcription template (RT) sequence, and a first primer binding site (PBS) sequence. The second pegRNA contains, from 5' to 3', a second guide sequence, a first scaffold sequence, a second reverse transcription template (RT) sequence, and a second primer binding site (PBS) sequence. The first pegRNA targets a first target sequence on the sense strand of the organism's genomic DNA, and the second pegRNA targets a second target sequence on the antisense strand of the organism's genomic DNA. The first RT sequence and the second RT sequence are used to insert a first exogenous nucleotide sequence. The first exogenous nucleotide sequence contains one or more recombinase recognition sites (RS); and iv) Recombinase and / or expression constructs containing a nucleotide sequence encoding said recombinase, The recombinase therein comprises a mutant of the Cre recombinase in the Cre-Lox system as described in any one of claims 1-3. The recognition site (RS) of the recombinase includes any one or any combination of the recognition sites (RS) of the recombinase in the Cre-Lox system as described in any one of claims 1-3.

5. The genome editing system of claim 4, wherein the genome editing system further comprises: v) A donor construct comprising one or more recognition sites (RS) of the recombinase and a second exogenous nucleotide sequence in the genome of the organism to be inserted; The recognition site (RS) of the recombinase includes any one or any combination of the recognition sites (RS) of the recombinase in the Cre-Lox system as described in any one of claims 1-3.

6. The genome editing system according to claim 4 or 5, wherein the organism is a plant or an animal.

7. The genome editing system according to any one of claims 4-6, wherein the CRISPR nuclease is a Cas9 nickase or a variant thereof; Preferably, the Cas9 nickase or a variant thereof comprises an amino acid sequence selected from SEQ ID NO:26 and 41-44.

8. The genome editing system according to any one of claims 4-7, wherein the CRISPR nuclease and the reverse transcriptase are linked by a adapter.

9. The genome editing system according to any one of claims 4-8, wherein the reverse transcriptase is M-MLV reverse transcriptase or a functional variant thereof; Preferably, the M-MLV reverse transcriptase or a functional variant thereof comprises an amino acid sequence as shown in SEQ ID NO:

31.

10. The genome editing system according to any one of claims 4-9, wherein the reverse transcriptase is fused to the nucleocapsid protein (NC) directly or via a linker at the N-terminus or C-terminus.

11. The genome editing system of claim 10, wherein the nucleocapsid protein (NC) comprises the amino acid sequence shown in SEQ ID NO:

28.

12. The genome editing system according to any one of claims 4-11, wherein the CRISPR nuclease described in i)-b) is fused to the N-terminus of the reverse transcriptase.

13. The genome editing system according to any one of claims 4-12, wherein the recombinase is contained in the guide editing fusion protein described in i)-b); Optionally, the recombinase is located at the N-terminus or C-terminus of the guide editing fusion protein, and is directly or through a linker linked to the guide editing fusion protein; Preferably, the fusion protein comprising the recombinase and the guided editing fusion protein comprises the amino acid sequence shown in SEQ ID NO:45 or an amino acid sequence having 85%, 90%, or 95% identity with it.

14. The genome editing system according to any one of claims 4-13, wherein the genome editing system further comprises: vi) an expression construct containing a third pegRNA and / or a nucleotide sequence encoding said third pegRNA, and vii) An expression construct containing a fourth pegRNA and / or a nucleotide sequence encoding the fourth pegRNA. The third pegRNA, from 5' to 3', comprises a third guide sequence, a first scaffold sequence, a third reverse transcription template (RT) sequence, and a third primer binding site (PBS) sequence. The fourth pegRNA, from 5' to 3', comprises a fourth guide sequence, a first scaffold sequence, a fourth reverse transcription template (RT) sequence, and a fourth primer binding site (PBS) sequence. The third pegRNA targets a third target sequence on the sense strand of the organism's genomic DNA, which is introduced into the recombinase recognition site (RS) via the genome editing system. The fourth pegRNA targets a fourth target sequence on the antisense strand of the organism's genomic DNA, which is introduced into the recombinase recognition site (RS) via the genome editing system. The third RT sequence and the fourth RT sequence are used to insert a third exogenous nucleotide sequence to delete the recognition site (RS) of recombinase introduced into the genome of an organism by the genome editing system.

15. The genome editing system according to any one of claims 4-14, wherein the pegRNA is capable of forming a complex with the CRISPR nuclease, the guide editing fusion protein, or a fusion protein comprising a recombinase and the guide editing fusion protein, and targeting the CRISPR nuclease, the guide editing fusion protein, or the fusion protein comprising a recombinase and the guide editing fusion protein to a target sequence in the genome, resulting in a cut in the target strand within the target sequence.

16. The genome editing system according to any one of claims 4-15, wherein the interval between the PAMs of the first target sequence and the second target sequence or the interval between the PAMs of the third target sequence and the fourth target sequence is approximately 20 bp to approximately 60 bp.

17. The genome editing system according to any one of claims 4-16, wherein the guide sequence in the first pegRNA has sufficient sequence identity (preferably 100% identity) with the first target sequence on the sense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the first target sequence; the guide sequence in the second pegRNA has sufficient sequence identity (preferably 100% identity) with the second target sequence on the antisense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the second target sequence; and / or The guide sequence in the third pegRNA has sufficient sequence identity (preferably 100%) with the third target sequence on the sense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the third target sequence; the guide sequence in the fourth pegRNA has sufficient sequence identity (preferably 100%) with the fourth target sequence on the antisense strand, and its complex with a CRISPR nuclease such as a nicking enzyme results in a nick in the fourth target sequence.

18. The genome editing system according to any one of claims 4-17, wherein the primer binding site sequence is configured to be complementary to at least a portion of the target sequence, preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the DNA strand containing the target sequence caused by a nick.

19. The genome editing system according to any one of claims 4-18, wherein the pegRNA scaffold sequence is shown in SEQ ID NO:

37.

20. The genome editing system according to any one of claims 4-19, wherein, The first RT sequence and the second RT sequence are configured to generate, after reverse transcription using them as templates, a first exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted, or a complementary sequence of the first exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted; and / or The third RT sequence and the fourth RT sequence are configured to generate a third exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted after reverse transcription using them as templates, or to generate a complementary sequence of the third exogenous nucleotide sequence or a portion thereof of the genome of the organism to be inserted.

21. The genome editing system according to any one of claims 4-20, wherein, The first RT sequence of the first pegRNA is configured to generate a first fragment of the first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; the second RT sequence of the second pegRNA is configured to be the complementary sequence of a second fragment of the first exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; and / or The third RT sequence of the third pegRNA is configured as the third fragment of the third exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template; the fourth RT sequence of the fourth pegRNA is configured as the complementary sequence of the fourth fragment of the third exogenous nucleotide sequence to be inserted into the genome after reverse transcription using it as a template.

22. The genome editing system of claim 21, wherein the first fragment and the second fragment of the first exogenous nucleotide sequence of the organism to be inserted into the genome at least partially overlap; and / or The third and fourth segments of the third exogenous nucleotide sequence to be inserted into the genome of the organism to be inserted at least partially overlap.

23. The genome editing system of claim 22, wherein the first and second fragments have at least about 10 bp to about 30 bp overlap; and / or The third and fourth segments overlap by at least approximately 10 bp to approximately 30 bp.

24. The genome editing system according to any one of claims 4-23, wherein the pegRNA further comprises a tevopre sequence at the 3' end of the PBS.

25. The genome editing system according to any one of claims 4-24, wherein the 5' end of the pegRNA is linked to a first ribozyme or tRNA, the first ribozyme or tRNA being programmed to cleave the fusion at the 5' end of the pegRNA; and / or the 3' end of the pegRNA is linked to a second ribozyme or tRNA, the second ribozyme or tRNA being programmed to cleave the fusion at the 3' end of the pegRNA.

26. The genome editing system according to any one of claims 4-25, wherein the pegRNA is transcribed by a type II promoter, optionally, the type II promoter is a GS promoter.

27. The genome editing system according to any one of claims 4-26, wherein the second exogenous nucleotide sequence may be 100 bp to about 18 kb or longer.

28. The genome editing system according to any one of claims 14-27, wherein the recognition site (RS) in the first exogenous nucleotide sequence is flanked by a PAM that can be recognized by the CRISPR nuclease or the guide editing fusion protein; and / or the second exogenous nucleotide sequence is flanked by a PAM that can be recognized by the CRISPR nuclease or the guide editing fusion protein.

29. The genome editing system according to any one of claims 4-28, wherein the second exogenous nucleotide sequence is associated with a plant trait such as an agronomic trait, and the insertion of the second exogenous nucleotide sequence results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.

30. The genome editing system of any one of claims 4-29, wherein the plant includes monocotyledonous and dicotyledonous plants, for example, the plant is a crop plant, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava and potato.

31. A method for producing a genetically modified plant, wherein the genetically modified plant comprises at least one modification of a site-directed inserted exogenous nucleotide sequence, deletion, substitution, inversion, or translocation of at least a portion of the genome, the method comprising introducing the Cre-Lox system of any one of claims 1-3 or the genome editing system of any one of claims 4-30 into at least one plant.

32. The method of claim 31, wherein the method further comprises screening plants having the desired modification from the at least one plant.

33. The method according to claim 31 or 32, wherein the genome editing system is introduced into plants by a method selected from: gene gun method, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method.

34. The method according to any one of claims 31-33, wherein the introduction comprises converting the genome editing system into isolated plant cells or tissues, and then regenerating the converted plant cells or tissues into a complete plant.

35. The method according to any one of claims 31-34, wherein the introduction comprises converting the genome editing system to a specific site on a whole plant, such as a leaf, shoot tip, pollen tube, young spike, or hypocotyl.

36. The method according to any one of claims 31-35, wherein components of the genome editing system are simultaneously introduced into the plant.

37. The method according to any one of claims 31-36, comprising the following steps: 1) Transform components i)-iv) of the genome editing system into isolated plant cells or tissues to obtain plant cells or tissues containing the first exogenous nucleotide sequence with a recognition site (RS) of one or more recombinases inserted; Optionally, 2) the component v) of the genome editing system is converted into the plant cells or tissues obtained in step 1), thereby obtaining plant cells or tissues containing the inserted second exogenous nucleotide sequence; Optionally, 3) converting components vi)-vii) of the genome editing system into the plant cells or tissues obtained in step 1) or step 2); and 4) Regenerate complete plants from plant cells or tissues obtained in step 1), step 2), or step 3).

38. Use of the Cre-Lox system of any one of claims 1-3 or the genome editing system of any one of claims 4-30 in the preparation of a pharmaceutical composition for treating diseases in subjects of need.

39. A pharmaceutical composition for treating a disease in a subject of need, wherein, The pharmaceutical composition comprises the Cre-Lox system of any one of claims 1-3 or the genome editing system of any one of claims 4-30, and optionally, a pharmaceutically acceptable vector.

40. A method for site-specific modification of the genome of a human or non-human animal, wherein, The method includes introducing the Cre-Lox system of any one of claims 1-3 or the genome editing system of any one of claims 4-30 into at least one human or non-human animal cell; Optionally, the method is for non-diagnostic and non-therapeutic purposes.

Citation Information

Patent Citations

  • Conditional expression of wild-type and mutant porcine IκBα gene DNA fragments

    CN102286475A

  • Rasamsonia transformants

    CN104838003A

  • Method for inserting exogenous sequence in genome at fixed point

    CN117126876A

  • Site-specific recombinase for efficient and specific genome editing

    CN118234856A

  • SINGLE pegRNA-MEDIATED LARGE INSERTIONS

    WO2023212594A2