Bxb1 recombinase with improved activity and application thereof

By optimizing the amino acid sequence of Bxb1 recombinase and constructing a genome editing system, the problems of low efficiency and poor specificity of recombinase were solved, achieving efficient genome editing, expanding application scenarios, and making it suitable for the treatment of animal and plant diseases and trait improvement.

CN121065133APending Publication Date: 2025-12-05BEIJING QI BIODESIGN BIOTECHNOLOGY CO LTD

Patent Information

Application Number
CN202511093123.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2025-08-05
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing recombinases have low recombination efficiency and poor specificity in genome editing, and their application scenarios are limited, making it difficult to meet the needs of animal and plant disease treatment and trait improvement.

Method used

By optimizing the amino acid sequence of Bxb1 recombinase and introducing specific amino acid substitutions, its recombination activity and efficiency were improved, and a highly active Bxb1 mutant recombinase was developed. Combined with CRISPR effector proteins and reverse transcriptase, a genome editing system was constructed.

Benefits of technology

It improves the integration and flipping efficiency of genome editing, expands the application scenarios of recombinases, enhances the precision and flexibility of gene editing tools, and is suitable for gene modification in a variety of organisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121065133A_ABST
    Figure CN121065133A_ABST
Patent Text Reader

Abstract

The invention discloses Bxb1 recombinase with improved activity and application of the Bxb1 recombinase. The invention provides a recombinase, and the amino acid sequence of the recombinase comprises one or more amino acid substitutions relative to the amino acid sequence shown in SEQ ID NO: 1. The recombinase provided by the invention can act on large-fragment DNA in a genome, so that the activity of a recombinase genome editing system is improved, and the application scene and selection of recombinase genome editing are enriched.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority and Related Applications

[0002] The present application claims priority to Chinese Patent Application No. 202411068225.0, filed August 5, 2024, entitled “Bxb1 Recombinase with Improved Activity and Applications Thereof,” and Chinese Patent Application No. 202510237975.4, filed February 28, 2025, entitled “Bxb1 Recombinase with Improved Activity and Applications Thereof.” The entire contents of the above-referenced patent applications are incorporated herein by reference in their entirety. TECHNICAL FIELD

[0003] The present application belongs to the field of genetic engineering. Specifically, the present application relates to a Bxb1 recombinase with improved activity and applications thereof. More specifically, the present application provides a Bxb1 mutant recombinase that can act on a genome, as well as a genome editing system and a fusion protein based thereon, a method for genome editing of an organism or cells of an organism using the recombinase, and a genetically modified organism and its offspring produced by the method. BACKGROUND

[0004] Site-specific recombinases catalyze the specific recombination of fragments between two specific DNA sequences, mediate DNA fragment integration, excision or inversion, and play a key role in the life cycle of many microorganisms, including bacteria and bacteriophages.

[0005] Due to the characteristics of not introducing DNA double-strand breaks in the process of editing the genome and being able to edit large fragments of chromosomes, recombinases can be used as an ideal gene editing tool for large fragment insertion, deletion and inversion of DNA, for the treatment of diseases and the improvement of traits in animals and plants.

[0006] Currently, the recombinases in the prior art have problems such as low recombination efficiency and poor specificity. For example, the recombination activity of tyrosine recombinase is reversible, resulting in low editing efficiency, and there are few choices of tyrosine family recombinases, the cargo fragments that can be effectively and accurately edited are small, and in addition, tyrosine recombinase Cre has cytotoxicity after overexpression. Large serine recombinases (LSR) have the characteristics of irreversible recombination, making them a potentially valuable genome editing tool. However, LSR still has problems of low editing efficiency and single variety. So far, only a few LSRs have been discovered, including Bxb1 and PhiC31 recombinases, but their recombination efficiency in animal and plant cells is very limited. Therefore, finding recombinases with high activity, high recombination efficiency and wide application scenarios is of great significance for expanding the existing DNA large fragment editing system and developing a library of gene editing tools that can precisely manipulate target DNA sequences. SUMMARY

[0007] Problem to be solved by the invention

[0008] Site-specific recombinases can directly modify the genome at the target sequence, or work together with other genome editing tools to achieve precise modification of the genome. Finding recombinases with high catalytic activity, high recombination efficiency, and wide application scenarios can expand existing gene editing techniques and have great research potential for the treatment of animal and plant diseases and trait improvement.

[0009] Solution for solving the problem

[0010] The present application provides a recombinase in the first aspect, wherein the amino acid sequence of the recombinase comprises one or more amino acid substitutions relative to the amino acid sequence shown in SEQ ID NO: 1, wherein the one or more amino acid substitutions include substitutions at positions 2, 3, 4, 5, 6, 7, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 23, 24, 25, 26, 27, 28, 29, 31, 32, 34, 35, 36, 37, 38, 39, 40, 41, 42, 44, 45, 46, 49, 50, 51, 52, 53, 54, 55, 58, 59, 60, 62, 63, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 77, 78, 79, 83, 84, 86, 87, 88, 89, 90, 92, 93, 94, 95, 96, 97, 98, 100, 101, 102, 103, 104, 105, 106, 107, 108, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 122, 123, 124, 125, 126, 128, 129, 132, 134, 135, 137, 138, 139, 142, 143, 144, 145, 146, 147, 148, 149, 150, or 151 of the amino acid sequence shown in SEQ ID NO: 1, and any combination thereof.

[0011] In some embodiments, the recombinase has improved recombination activity relative to the recombinase as shown in SEQ ID NO: 1.

[0012] In some embodiments, the amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions selected from the following: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C, S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V 37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, E42L, E42R, D 45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, D51C, D51S, P52L, P5 2M, P52V, F53H, F53L, F53M, F53R, F53W, D54C, R55A, R55Q, R55W, R55Y, P5 9D, P59G, N60M, A62H, A62N, R63C, R63Q, A66W, F67S, E68D, E68P, E68Q, E6 8R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I7 5F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84 N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89 S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S , A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V , K101Y, K102C, K102H, L103M, L103Q, V104C, V104L, V105E, V105Q, S106N,A107V, T108A, T108G, T108Y, A110C, A110L, A110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, or F118S.

[0013] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, R2Q, A3G, L4M, V129A, I123T, or V76C, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0014] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution T84N, relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more of the amino acid substitutions selected from H146C, A144S, P71M, P71S, Q92C, D51S, E42L, E42R, D51C, E69M, D54C, A66W, V104L.

[0015] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, P71M, Q92C, E42L, E69M, D54C, A66W.

[0016] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of T84N and P71M relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, Q92C.

[0017] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of T84N and P71S relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, Q92C.

[0018] In some embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in any one of (a) to (s) relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0019] (a) T84N

[0020] (b) T84N and E69M;

[0021] (c) T84N and Q92C;

[0022] (d) T84N and E42L;

[0023] (e) T84N and A66W;

[0024] (f) T84N and D54C;

[0025] (g) T84N and H146C;

[0026] (h) T84N and P71M;

[0027] (i) T84N and P71S;

[0028] (j) T84N and E42R;

[0029] (k) T84N and V104L;

[0030] (l) T84N and D51C;

[0031] (m) T84N and A144S;

[0032] (n) T84N and D51S;

[0033] (o) T84N, P71M and H146C;

[0034] (p) T84N, P71M and Q92C;

[0035] (q) T84N, P71S and H146C;

[0036] (r) T84N, P71S and Q92C.

[0037] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of I87T relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from Y78R, A144T, H146C, H146Q, F118M, M132L, A144G.

[0038] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a1) to (h1) relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0039] (a1) I87T;

[0040] (b1) I87T and A144G;

[0041] (c1) I87T and Y78R;

[0042] (d1) I87T and A144T;

[0043] (e1) I87T and H146C;

[0044] (f1) I87T and H146Q;

[0045] (g1) I87T and F118M;

[0046] (h1) I87T and M132L.

[0047] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, Y78R, H146C, I87T, Q92C.

[0048] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a2) to (f2) relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0049] (a2) A144S;

[0050] (b2) A144S and V94G;

[0051] (c2) A144S and Y78R;

[0052] (d2) A144S and H146C;

[0053] (e2) A144S and I87T;

[0054] (f2) A144S and Q92C.

[0055] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from H146C, K101E.

[0056] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in any one of (a3) to (c3) relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0057] (a3) L93A;

[0058] (b3) L93A and H146C;

[0059] (c3) L93A and K101E.

[0060] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, H146C.

[0061] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in any one of (a4) to (c4) relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0062] (a4) Q92C;

[0063] (b4) Q92C and V94G;

[0064] (c4) Q92C and H146C.

[0065] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, Y78R, H146C.

[0066] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in (a5) to (d5) below relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0067] (a5) M132L;

[0068] (b5) M132L and V94G;

[0069] (c5) M132L and Y78R;

[0070] (d5) M132L and H146C.

[0071] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from P59D, Y78R, V94G.

[0072] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in (a6) to (d6) below relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0073] (a6) N148A;

[0074] (b6) N148A and P59D;

[0075] (c6) N148A and Y78R;

[0076] (d6) N148A and V94G.

[0077] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of E42W relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from T84N, P59D, V94G, K102H, Y78R.

[0078] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in (a7) to (f7) below relative to the amino acid sequence set forth in SEQ ID NO: 1:

[0079] (a7) E42W;

[0080] (b7) E42W and T84N;

[0081] (c7) E42W and P59D;

[0082] (d7) E42W and V94G;

[0083] (e7) E42W and K102H;

[0084] (f7) E42W and Y78R.

[0085] In some embodiments, the amino acid sequence of the recombinant enzyme comprises any one or more of the amino acid substitutions selected from A3G, V129A, V76C, L4M, or I123T relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0086] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution of R2Q, A3G, or L4M relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0087] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution of V76C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0088] The second aspect of the application provides a genome editing system, wherein the genome editing system comprises:

[0089] 1) a recombinase recognition site sequence (RS) comprising an attP or attB site;

[0090] 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase, the recombinase comprising the recombinase of the first aspect of the application.

[0091] In some embodiments, the recombinase recognition site sequence (RS) is naturally occurring or engineered or optimized.

[0092] In some embodiments, one or more of the recombinase recognition site sequence (RS) is inserted into a desired location in the genome by a recombinase recognition site integration unit.

[0093] In some embodiments, the recombinase recognition site integration unit comprises a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or the functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.

[0094] In some embodiments, the CRISPR effector protein functional variant is a CRISPR nuclease with full / partial loss of cleavage activity, preferably, the CRISPR effector protein functional variant is a CRISPR nickase, for example, Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase, or TraC nickase.

[0095] In some embodiments, the guide RNA comprises a scaffold sequence and a primer binding sequence, and a recombination enzyme recognition site sequence (RS) integration template sequence.

[0096] In some embodiments, the recombination enzyme recognition site integration unit further comprises a reverse transcriptase and / or an expression construct comprising a nucleotide sequence encoding the reverse transcriptase.

[0097] In some embodiments, the CRISPR nickase of the recombination enzyme recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with the CRISPR nickase and targets a desired location in the genome, wherein the CRISPR nickase makes a nick on a strand of the genome, and the reverse transcriptase incorporates the RS integration template sequence in the guide RNA into the nick site, thereby inserting at least one recombination enzyme recognition site sequence (RS) recognizable by the recombination enzyme at the desired location of the genome.

[0098] In some embodiments, the reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii Marathonase RT (Marathon RT).

[0099] In some embodiments, the recombination enzyme recognition site integration unit is not covalently linked to a recombination enzyme; or, the recombination enzyme recognition site integration unit is covalently linked to a recombination enzyme.

[0100] In some embodiments, the genome editing system further comprises:

[0101] 3) a donor of an exogenous nucleotide sequence to be inserted into the genome, optionally, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer.

[0102] The third aspect of the present application provides a fusion protein comprising the genome editing system of the second aspect of the present application, wherein the CRISPR nickase is linked to a reverse transcriptase at the C-terminus, and the reverse transcriptase is linked to a recombination enzyme via a linker.

[0103] The fourth aspect of the application provides a polynucleotide comprising a nucleotide sequence encoding the recombinase of the first aspect of the application, the genome editing system of the second aspect of the application, or the fusion protein of the third aspect of the application. Optionally, the polynucleotide is RNA, e.g., mRNA.

[0104] The fifth aspect of the application provides an expression construct comprising the polynucleotide of the fourth aspect of the application.

[0105] The sixth aspect of the application provides a cell comprising the recombinase of the first aspect of the application, the genome editing system of the second aspect of the application, the fusion protein of the third aspect of the application, the polynucleotide of the fourth aspect of the application, or the expression construct of the fifth aspect of the application.

[0106] The seventh aspect of the application provides a kit comprising the recombinase of the first aspect of the application, the genome editing system of the second aspect of the application, the fusion protein of the third aspect of the application, the polynucleotide of the fourth aspect of the application, the expression construct of the fifth aspect of the application, or the cell of the sixth aspect of the application.

[0107] The eighth aspect of the application provides a method of performing gene editing in an organism or a cell of an organism, wherein the method introduces into the organism or the cell of the organism the recombinase of the first aspect of the application, the polynucleotide of the fourth aspect of the application, the expression construct of the fifth aspect of the application, the genome editing system of the second aspect of the application, or the fusion protein of the third aspect of the application.

[0108] In some embodiments, components 1), 2), and optionally 3) of the genome editing system are introduced into the organism or the cell of the organism simultaneously.

[0109] In some embodiments, components 1), 2), or optionally 3) of the genome editing system are introduced into the organism or the cell of the organism stepwise, respectively.

[0110] In some embodiments, the component 1) is inserted into the donor construct of the genomic or exogenous nucleotide sequence in the same direction or in the opposite direction.

[0111] In some embodiments, the method comprises recombining DNA of the genome of the organism or the cell of the organism; optionally, the recombining DNA of the genome of the organism or the cell of the organism comprises deleting DNA, inverting DNA in the genome, and / or integrating exogenous DNA into the genome.

[0112] In some embodiments, the genome editing system is introduced into the cell by a method selected from the group consisting of calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, viral infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, or other viruses), biolistics, N-acetylgalactosamine (GalNAc)-mediated, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.

[0113] In some embodiments, the organism or organism cell is from a mammal such as a human, a mouse, a rat, a monkey, a dog, a pig, a sheep, a cow, a cat; poultry such as a chicken, a duck, a goose; a plant, preferably a crop plant, for example, wheat, rice, corn, soybean, sunflower, leafy greens, lettuce, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.

[0114] Effects of the invention

[0115] The present inventors, through a large number of experiments and explorations, screened and optimized Bxbl recombinase, and unexpectedly obtained a mutant Bxbl recombinase with high recombination activity, the integration efficiency and inversion efficiency of the mutant Bxbl were higher than those of wild-type Bxbl, and the mutant Bxbl could be used for large-fragment DNA insertion and translocation in the genome. This discovery helps to improve the editing activity of the genome editing system based on the recombinase, and the various mutant Bxbl enriches the application scenarios and selection of the genome editing based on the recombinase. BRIEF DESCRIPTION OF DRAWINGS

[0116] Figure 1 Figure 1 is a schematic diagram of an expression construct and a screening method of Bxbl mutant screening system 1 (evolutionary screening system 1).

[0117] Figure 2 Figure 2 is a schematic diagram of an expression construct of Bxbl mutant screening system 2 (evolutionary screening system 2).

[0118] Figure 3 Figure 3 is a schematic diagram of an expression construct of recombinase activity verification system 1.

[0119] Figure 4 Figure 4 is a schematic diagram of a method of inserting RS in the genome by double pegRNA of recombinase activity verification system 2.

[0120] Figure 5 Figure 5 is a schematic diagram of a recombinase expression construct of recombinase activity verification system 2.

[0121] Figure 6 Figure 6 is a schematic diagram of a detection method of recombinase activity verification system 2.

[0122] Figure 7, Recombinase activity verification system 2 in rice cells Recombination vector 1 recombination activity verification results.

[0123] Figure 8 , Recombinase activity verification system 2 in rice cells Recombination vector 2 recombination activity verification results.

[0124] Figure 9 , Recombinase activity verification system 2 in human cells Recombination activity verification results.

[0125] Figure 10 , Recombinase activity verification system 1 in tobacco Bxbl multi-mutant recombination activity verification results.

[0126] Figure 11 , Recombinase activity verification system 1 in rice Bxbl multi-mutant recombination activity verification results. DETAILED DESCRIPTION

[0127] I. DEFINITIONS

[0128] In the present application, unless otherwise indicated, the scientific and technical terms used herein have the meanings that would be generally understood by one of ordinary skill in the art. Also, the terms related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology, and the like, and the laboratory procedures steps used herein are the terms and procedures commonly used in the corresponding fields. For example, the standard recombinant DNA and molecular cloning techniques used in the present application are well known in the art and are described more fully in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter "Sambrook"). Also, for better understanding of the present application, the definitions and explanations of the related terms are provided below.

[0129] As used herein, the term "and / or" encompasses all combinations of the items linked by the term. It will be understood that each combination is individually contemplated, even though the combinations are not individually listed. For example, "A and / or B" covers "A," "B," "A and B." For example, "A, B, and / or C" covers "A," "B," "C," "A and B," "A and C," "B and C," and "A and B and C."

[0130] The term "comprising" as used herein to describe a protein or nucleic acid sequence means that the protein or nucleic acid can consist of the sequence, or can have additional amino acids or nucleotides at either or both ends of the protein or nucleic acid, but still have the activity described herein. Furthermore, it is clear to one skilled in the art that the methionine encoded by the start codon at the N-terminus of a polypeptide can in some instances (e.g., when expressed in certain expression systems) be retained, but does not materially affect the function of the polypeptide. Thus, the specification and claims herein, in describing specific polypeptide amino acid sequences, encompass sequences that include the methionine encoded by the start codon at the N-terminus, even though they can not include it, and correspondingly, the nucleotide sequences that encode them can include the start codon, and vice versa.

[0131] "Genome" as used herein encompasses not only chromosomal DNA present in the nucleus of a cell, but also organelle DNA present in subcellular components of a cell, such as mitochondria, plastids.

[0132] "Organism" includes any organism suitable for genome editing, preferably a eukaryote. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry such as chickens, ducks, geese; plants including monocots and dicots, for example rice, maize, wheat, sorghum, barley, soybean, peanut, Arabidopsis, etc.

[0133] "Foreign" means a sequence from a foreign species, or, if from the same species, a sequence that has been significantly altered by deliberate human intervention from its natural form.

[0134] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and are single- or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single letter designation: "A" for adenosine or deoxyadenosine (RNA or DNA, respectively), "C" for cytosine or deoxycytosine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "I" for inosine, and "X" for any nucleotide.

[0135] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" can also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.

[0136] As used herein, the term "amino acid" can include natural amino acids, non-natural amino acids, amino acid analogs, and all their D and L stereoisomers. Amino acids and abbreviations and English names in the present application are shown below:

[0137] Histidine (His, H); Serine (Ser, S); Glutamic acid (Glu, E); Glutamine (Gln, Q); Glycine (Gly, G); Threonine (Thr, T); Phenylalanine (Phe, F); Aspartic acid (Asp, D); Tyrosine (Tyr, Y); Leucine (Leu, L); Isoleucine (lie, I); Arginine (Arg, R); Alanine (Ala, A); Valine (Val, V); Tryptophan (Trp, W); Methionine (Met, M); Asparagine (Asn, N); Cysteine (Cys, C); Lysine (Lys, K); Proline (Pro, P).

[0138] As used herein, the term "wild type" refers to an object that can be found in nature. For example, a polypeptide or polynucleotide sequence that exists in an organism, can be isolated from a source in nature and has not been intentionally modified by humans in a laboratory is naturally occurring. As used herein, "naturally occurring" and "wild type" are synonymous.

[0139] As used herein, the term "mutant" refers to a polynucleotide or polypeptide that contains an alteration (i.e., a substitution, an insertion, and / or a deletion) at one or more (e.g., several) positions relative to a "wild type," or "compared" polynucleotide or polypeptide, wherein a substitution refers to replacing a nucleotide or amino acid occupying a position with a different nucleotide or amino acid. A deletion refers to removing a nucleotide or amino acid occupying a position. An insertion refers to adding a nucleotide or amino acid after nucleotides or amino acids that abut and immediately follow the nucleotide or amino acid occupying a position.

[0140] As used herein, an "altered" amino acid includes a substitution, a duplication, a deletion, or an addition of one or more amino acids. In the present application, the term "mutation" refers to an alteration in an amino acid sequence. In a particular embodiment, the term "mutation" refers to a "substitution."

[0141] The term "corresponding to" or "relative to" as used herein has the meaning generally understood by one of ordinary skill in the art. In particular, "corresponding to" means that after homology or sequence identity alignment of two sequences, a position in one sequence corresponds to a specified position in the other sequence. Thus, for example, in the context of "the amino acid residue corresponding to position 150 of the amino acid sequence shown in SEQ ID NO: 1," if a 6xHis tag is added to one end of the amino acid sequence shown in SEQ ID NO: 1, then the corresponding position to position 150 of the amino acid sequence shown in SEQ ID NO: 1 in the resulting mutant can be position 156.

[0142] Sequence "identity" has the meaning commonly understood in the art and can be calculated as the percentage of nucleotide or amino acid residues in the two sequences that are the same. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule. (See, e.g., Computational Molecular Biology, Lesk, A.M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A.M., and Griffin, H.G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). Although there are a number of methods for measuring sequence identity between two polynucleotides or polypeptides, the term "identity" is well known to one of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).

[0143] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art and can generally be made without altering the biological activity of the resulting molecule. In general, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a polypeptide will not substantially alter biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).

[0144] As used herein, "expression construct" or "construct" refers to a vector, such as a recombinant vector, suitable for expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to transcription of the nucleotide sequence (e.g., to produce mRNA or functional RNA) and / or translation of the RNA into a precursor or mature protein.

[0145] An "expression construct" of the application can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (e.g., mRNA).

[0146] An "expression construct" of the application can comprise regulatory sequences and a nucleotide sequence of interest of different origin, or regulatory sequences and a nucleotide sequence of interest of the same origin arranged in a manner different from that normally existing in nature.

[0147] As used herein, the term "recombinase" and the like refers to a site-specific enzyme that mediates recombination of DNA between recombinase recognition sequences. During the reaction, an attack on the DNA phosphate backbone is launched by a tyrosine or serine located in the catalytic active center of the recombinase, causing a DNA strand break. Recombinases can be divided into two different families: serine recombinases (e.g., resolvases and invertases) and tyrosine recombinases (e.g., integrases). Some examples of serine recombinases include, but are not limited to, Bxbl, Ilin, Gin, Tn3, beta-six, CinH, ParA, gamma delta, Bxbl, TP901, TGl, Rl, R2, R3, R4, R5, MRll, Al 18, U153, and gp29.

[0148] The term "recombination" refers to excision, integration, inversion, or exchange (e.g., translocation) of a DNA fragment between recombinase recognition sequences. In some embodiments, the recombinase recombination activity boost includes an increase in integration activity or inversion activity of the recombinase.

[0149] As used herein, the term "large serine recombinase," "LSR" is a classification of serine recombinases, which are integrases carried by bacteriophages that integrate large fragments of DNA sequences into the genome of a bacterium by recognizing specific sequences on the bacterial genome and the phage DNA fragment, without the need for any cellular cofactors and without the generation of harmful DNA double-strand breaks (DSBs). Many serine recombinases function to resolve transposition intermediates or to regulate gene expression by inverting regulatory sequences. During the catalytic reaction of a serine recombinase, a complex is formed from DNA and the recombinase, then a serine in the recombinase domain attacks the DNA phosphate backbone causing the DNA to form a nick with a 3'-OH double-strand break end and a 5'-phosphoserine covalently linked to the DNA, the complex is flipped, the double-stranded DNA is re-ligated, and a recombination product is formed. Most serine recombinases have a 150 amino acid catalytic domain at their amino terminus, followed by a small HTH (Helix-Turn-Helix)-DNA binding domain. Large serine recombinases have a similar amino-terminal catalytic domain, but have a larger carboxy-terminal region that varies in size from 300 residues in the bacteriophage R4 and A118 recombinases to 550 residues in the TnpX transposase, and are referred to as large serine recombinases because of the larger carboxy-terminal region.

[0150] As used herein, the terms "attP," "attB" are the phage attachment site (attP) and the bacterial attachment site (attB), respectively, and are collectively referred to herein as "recombinase recognition site sequences" or "recombination sites." The original function of integrases is to recombine short sequences of DNA between the phage attachment site (attP) and the bacterial attachment site (attB), so the recognition site for the target DNA in the gene editing process is called "attB" and the recognition site for the DNA sequence to be edited is called "attP."

[0151] As used herein, the terms "recombinase recognition site sequence," "recombination site," "recognition site," "RS site," "RS sequence," "recognition sequence," "RS" generally refer to a nucleic acid (e.g., DNA) sequence that is recognized (e.g., can be bound by) a recombinase polypeptide, which can be naturally occurring, or engineered or optimized.

[0152] II. Recombinases

[0153] In one aspect, the present application provides a recombinase that is capable of recombining DNA between a recombinase recognition site sequence (RS). In some embodiments, the recombinase is from a bacteriophage. In some embodiments, the recombinase is from a mycolicibacterium smegmatis bacteriophage. In some embodiments, the recombinase is a Bxbl recombinase.

[0154] In some embodiments, the recombinase comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to the amino acid sequence set forth in SEQ ID NO: 1.

[0155] In some embodiments, the amino acid sequence of the recombinase comprises one or more amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1, wherein the one or more amino acid substitutions comprise a substitution at position 2, 3, 4, 5, 6, 7, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 23, 24, 25, 26, 27, 28, 29, 31, 32, 34, 35, 36, 37, 38, 39, 40, 41, 42, 44, 45, 46, 49, 50, 51, 52, 53, 54, 55, 58, 59, 60, 62, 63, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 77, 78, 79, 83, 84, 86, 87, 88, 89, 90, 92, 93, 94, 95, 96, 97, 98, 100, 101, 102, 103, 104, 105, 106, 107, 108, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 122, 123, 124, 125, 126, 128, 129, 132, 134, 135, 137, 138, 139, 142, 143, 144, 145, 146, 147, 148, 149, 150, or 151 of the amino acid sequence set forth in SEQ ID NO: 1, and any combination thereof.

[0156] In some embodiments, the recombinase has increased recombination activity relative to a recombinase as set forth in SEQ ID NO: 1.

[0157] In some embodiments, the amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions selected from the following: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C. S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V 37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, E42L, E42R, D 45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, D51C, D51S, P52L, P5 2M, P52V, F53H, F53L, F53M, F53R, F53W, D54C, R55A, R55Q, R55W, R55Y, P5 9D, P59G, N60M, A62H, A62N, R63C, R63Q, A66W, F67S, E68D, E68P, E68Q, E6 8R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I7 5F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84 N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89 S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S , A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V , K101Y, K102C, K102H, L103M, L103Q, V104C, V104L, V105E, V105Q, S106N,A107V, T108A, T108G, T108Y, A110C, A110L, A110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, or F118S.

[0158] In some embodiments, the amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, and 59 amino acids. 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 1 16, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, or 319 amino acid substitutions.

[0159] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C, S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, D45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, P52L, P52M, P52V, F53H, F53L, F53M, F53R, F53W, R55A, R55Q, R55W, R55Y, P59D, P59G, N60M, A62H, A62N, R63C, R63Q, F67S, E68D, E68P, E68Q, E68R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I75F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S, A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, L103M, L103Q, V104C, V105E, V105Q, S106N, A107V, T108A, T108G, T108Y, A110C, A110L, A110V, A112C, A112G, A112L, A112M, A112S, A112V, A112Y, A113C, A113G, A113L, A113M, A113S, A113V, A113Y, A114C, A114G, A114L, A114M, A114S, A114V, A114Y, A115C, A115G, A115L, A115M, A115S, A115V, A115Y, A116C, A116G, A116L, A116M, A116S, A116V, A116Y, A117C, A117G, A117L, A117M, A117S, A117V, A117Y, A118C, A118G, A118L, A118M, A118S, A118V, A118Y, A119C, A119G, A119L, A119M, A119S, A119V, A119Y, A120C, A120G, A120L, A120M, A120S, A120V, A120Y, A121C, A121G, A121L, A121M, A121S, A121V, A121Y, A122C, A122G, A122L, A122M, A122S, A122V, A122Y, A123C, A123G, A123L, A123M, A123S, A123V, A123Y, A124C, A124G, A124L, A124M, A124S, A124V, A124Y, A125C, A125G, A125L, A125M, A125S, A125V, A125Y, A126C, A126G, A126L, A126M, A126S, A126V, A126Y, A127C, A127G, A127L, A127M, A127S, A127V, A127Y, A128C, A128G, A128L, A128M, A128S, A128V, A128Y, A129C, A129G, A129L, A129M, A129S, A129V, A129Y, A130C, A130G, A130L, A130M, A130S, A130V, A130Y, A131C, A131G, A131L, A131M, A131S, A131V, A131Y, A132C, A132G, A132L, A132M, A132S, A132V, A132Y, A133C, A133G, A133L, A133M, A133S, A133V, A133Y, A134C, A134G, A134L, A134M, A134S, A134V, A134Y, A135C, A135G, A135L, A135M, A135S, A135V, A135Y, A136C, A136G, A136L, A136M, A136S, A136V, A136Y, A137C, A137G, A137L, A137M, A137S, A137V, A137Y, A138C, A138G, A138L, A138M, A138S, A138V, A138Y, A139C, A139G, A139L, A139M, A139S, A139V, A139Y, A140C, A140G, A140L, A140M, A140S, A140V, A140Y, A141C, A141G, A141L, A141M, A141S, A141V, A141Y, A142C, A142G, A142L, A142M, A142S, A142V, A142Y, A143C, A143G, A143L, A143M, A143S, A143V, A143Y, A144C, A144G, A144L, A144M, A144S, A144V, A144Y, A145C, A145G, A145L, A145M, A145S, A145V, A145Y, A146C, A146G, A146L, A146M, A146S, A146V, A146Y, A147C, A147G, A147L, A147M, A147S, A147V, A147Y, A148C, A148G, A148L, A148M, A148S, A148V, A148Y, A149C, A149G, A149L, A149M, A149S, A149V, A149Y, A150C, A150G, A150L, A150M, A150S, A150V, A150Y, A151C, A151G, A151L, A151M, A151S, A151V, A151Y, A152C, A152G, A152L, A152M, A152S, A152V, A152Y, A153C, A153G, A153L, A153M, A153S, A153V, A153Y, A154C, A154G, A154L, A154M, A154S, A154V, A154Y, A155C, A155G, A155L, A155M, A155S, A155V, A155Y, A156C, A156G, A156L, A156M, A156S, A156V, A156Y, A157C, A157G, AA110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T.

[0160] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0161] In some embodiments, the amino acid sequence of the recombinant enzyme has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, or 131 amino acid substitution relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0162] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0163] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from H146C, A144S, P71M, P71S, Q92C, D51S, E42L, E42R, D51C, E69M, D54C, A66W, V104L.

[0164] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from H146C, P71M, Q92C, E42L, E69M, D54C, A66W.

[0165] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and E69M relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0166] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0167] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and E42L relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0168] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and A66W relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0169] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and D54C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0170] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0171] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and P71M relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0172] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N, P71M, and any one or more amino acid substitutions selected from H146C, Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0173] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N, P71M, and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0174] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N, P71M, and Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0175] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and P71S relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0176] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N, P71S amino acid substitutions, and any one or more amino acid substitutions selected from H146C, Q92C, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0177] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N, P71S, and H146C amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0178] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N, P71S, and Q92C amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0179] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and E42R amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0180] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and V104L amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0181] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and D51C amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0182] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and A144S amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0183] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and D51S amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0184] In some embodiments, the amino acid sequence of the recombinant enzyme comprises I87T amino acid substitution, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0185] In some embodiments, the amino acid sequence of the recombinant enzyme comprises I87T amino acid substitution, and any one or more amino acid substitutions selected from Y78R, A144T, H146C, H146Q, F118M, M132L, A144G, relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0186] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and A144G relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0187] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0188] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and A144T relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0189] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0190] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and H146Q relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0191] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and F118M relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0192] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and M132L relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0193] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution A144S relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0194] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution A144S relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more of the amino acid substitutions selected from V94G, Y78R, H146C, I87T, Q92C.

[0195] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions A144S and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0196] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions A144S and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0197] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0198] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S and I87T relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0199] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S and Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0200] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0201] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A relative to the amino acid sequence set forth in SEQ ID NO: 1, and one or more amino acid substitutions selected from H146C, K101E.

[0202] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0203] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A and K101E relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0204] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0205] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1, and one or more amino acid substitutions selected from V94G, H146C.

[0206] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0207] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0208] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0209] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and any one or more amino acid substitutions selected from V94G, Y78R, H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0210] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0211] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0212] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0213] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0214] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and any one or more amino acid substitutions selected from P59D, Y78R, V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0215] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and P59D relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0216] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0217] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0218] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of E42W relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0219] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of E42W relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from T84N, P59D, V94G, K102H, Y78R.

[0220] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and T84N relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0221] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and P59D relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0222] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0223] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and K102H relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0224] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0225] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from A3G, V5A, I7T, I7V, I87T, I87V, T108A, V129A, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A41V, A62Y, P71S, P71I, H89Y, A119Y, F118I, A120R, F118S, or H146Y relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0226] In some embodiments, the amino acid sequence of the recombinase comprises one or more amino acid substitutions of R2Q, A3G, V129A, V76C, L4M, or I123T relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0227] In some embodiments, the amino acid sequence of the recombinase comprises an amino acid substitution of R2Q, A3G, or L4M relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0228] In some embodiments, the amino acid sequence of the recombinase comprises an amino acid substitution of V76C relative to the amino acid sequence set forth in SEQ ID NO: 1.

[0229] In some embodiments, the recombinase recognition site sequence (RS) of the recombinase provided herein comprises an attP site (SEQ ID NO: 2), an attB site (SEQ ID NO: 3), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the nucleotide sequence set forth in SEQ ID NO: 2-3.

[0230] III. Genome editing system

[0231] In one aspect, the present disclosure provides a genome editing system, wherein the genome editing system comprises:

[0232] 1) a recombinase recognition site sequence (RS) comprising an attP or attB site (hereinafter also referred to as Component 1);

[0233] 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase (hereinafter also referred to as Component 2), wherein the recombinase comprises any one or more of the recombinases described above.

[0234] In some embodiments, the recombinase recognition site sequence (RS) is naturally occurring, or engineered or optimized.

[0235] In some embodiments, one or more recombinase recognition site sequence (RS) is inserted into a desired location in the genome by a recombinase recognition site integration unit.

[0236] A person skilled in the art can select a suitable gene editing tool to insert the RS into a desired location in the genome.

[0237] In some embodiments, the recombinase recognition site integration unit comprises: a sequence-specific DNA cleavage protein and a RS integration template, and / or an expression construct encoding the sequence-specific DNA cleavage protein and the RS integration template. In some embodiments, the recombinase recognition site integration unit specifically recognizes a desired location in a genome and cleaves DNA, followed by insertion of a recombinase recognition site sequence into the desired location in the genome by an endogenous or exogenous repair mechanism of the cell, mediated by the RS integration template.

[0238] In some embodiments, the RS integration template is a donor DNA template or a reverse transcription RNA template, wherein the donor DNA template comprises a desired recombinase recognition site sequence or a complement thereof, and the reverse transcription RNA template comprises a transcription sequence of a desired recombinase recognition site sequence or a complement thereof. In some embodiments, the RS integration template comprises a DNA synthetic template encoding a RS recognizable by a recombinase.

[0239] In some embodiments, the sequence-specific DNA cleavage protein is selected from one or more of a meganuclease (MGN), a zinc-finger nuclease (ZFN), a transcription-activator-like effector nuclease (TALEN), and a CRISPR effector protein.

[0240] In some embodiments, the recombinase recognition site integration unit comprises: a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or the functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.

[0241] In some embodiments, the CRISPR effector protein or the functional variant thereof comprises a CRISPR nuclease and a functional variant thereof.

[0242] As used herein, the term "CRISPR effector protein" generally refers to a nuclease or a functional variant thereof that is present in a naturally occurring CRISPR system. The term encompasses any effector protein based on a CRISPR system that is capable of achieving sequence-specific targeting within a cell.

[0243] As used herein, the term "CRISPR nuclease" can be derived from a Cas9 nuclease, including a Cas9 nuclease or a functional variant thereof. The Cas9 nuclease can be a Cas9 nuclease from different species, such as spCas9 from S. pyogenes or SaCas9 derived from S. aureus. "Cas9 nuclease" and "Cas9" are used interchangeably herein to refer to an RNA-guided nuclease including a Cas9 protein or a fragment thereof (e.g., a protein comprising the active DNA cleavage domain of Cas9 and / or the gRNA binding domain of Cas9). Cas9 is a component of the CRISPR / Cas (clustered regularly interspaced short palindromic repeats and associated systems) genome editing system that can target and cleave a DNA target sequence under the guidance of a guide RNA to form a DNA double-strand break (DSB).

[0244] A "CRISPR nuclease" can also be derived from a Cpf1 nuclease, including a Cpf1 nuclease or a functional variant thereof. The Cpf1 nuclease can be a Cpf1 nuclease from different species, such as Cpf1 nuclease from Francisella novicida U112, Acidaminococcus sp. BV3L6 and Lachnospiraceae bacterium ND2006.

[0245] As used herein, a "functional variant" with respect to a CRISPR nuclease means that it retains at least the ability of sequence-specific targeting mediated by a guide RNA. Preferably, the functional variant is a nuclease-inactive variant, i.e., it lacks the double-stranded nucleic acid cleavage activity. However, a CRISPR nuclease that lacks the double-stranded nucleic acid cleavage activity also encompasses a nickase, which forms a nick in a double-stranded nucleic acid molecule, but does not completely cut the double-stranded nucleic acid.

[0246] In some preferred embodiments of the present application, the CRISPR effector protein of the present application has nickase activity (or is referred to as a CRISPR nickase). In some embodiments, the functional variant recognizes a different PAM (protospacer adjacent motif) sequence relative to the wild-type nuclease.

[0247] As used herein, the term “CRISPR nickase” can be derived from a Cas9 of different species, e.g., from S. pyogenes Cas9 (SpCas9), or from S. aureus Cas9 (SaCas9). Simultaneous mutation of the HNH nuclease subdomain and the RuvC subdomain of Cas9, e.g., comprising mutations D10A and H840A, renders Cas9 nuclease-inactive, a nuclease-inactive Cas9 (dCas9). Mutation of only one of the subdomains can render Cas9 to have nickase activity, i.e., a Cas9 nickase (nCas9), e.g., nCas9 with only mutation D10A (or referred to as Cas9-D10A), and nCas9 with only mutation H840A (or referred to as Cas9-H840A). In some embodiments, the CRISPR nickase can be derived from a Cas12 of different species, e.g., Cas12a, Cas12b, Cas12i. In some embodiments, the CRISPR nickase is based on TraC protein that only retains DNA single-strand cleavage activity among transposon and CRISPR-Cas12 intermediate TraC effector protein (wherein TraC nuclease is disclosed in PCT / CN2023 / 097783 (publication number WO / 2023 / 232109), which is incorporated herein by reference).

[0248] In some embodiments, the CRISPR nuclease functional variant is a CRISPR nickase, e.g., Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase, and / or TraC nickase (see patents CN202310646033.2, CN202411698901.2).

[0249] In some embodiments, the recombinase recognition site integration unit further comprises a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the reverse transcriptase. In these embodiments, the recombinase recognition site integration unit can be based on prime editors and iterations thereof, which have demonstrated insertion, deletion, and base conversion with modest editing efficiency (as disclosed in WO 2020 / 191234 Al, incorporated by reference herein), as well as plant prime editors (PPEs), disclosed in Lin, Qiupeng et al. “Prime genome editing in rice and wheat.” Nature biotechnology vol. 38, 5 (2020): 582-585. doi:10.1038 / s41587-020-0455-x; Lin, Qiupeng et al. “High-efficiency prime editing with optimized, paired pegRNAs in plants.” Nature biotechnology vol. 39, 8 (2021): 923-927. doi:10.1038 / s41587-021-00868-w, and the like; dual-ePPEs, disclosed in Sun, Chao et al. “Precise integration of large DNA sequences in plant genomes using PrimeRoot editors.” Nature biotechnology vol. 42, 2 (2024): 316-327. doi:10.1038 / s41587-023-01769-w; TwinPEs, disclosed in Anzalone, Andrew V et al. “Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing.” Nature biotechnology vol. 40, 5 (2022): 731-740. doi:10.1038 / s41587-021-01133-w, and the like, incorporated by reference herein.

[0250] In some embodiments, the recombinase recognition site integration unit further comprises all the technologies in the prior art that can insert specific DNA sequences into DNA fragments at a specific site, including but not limited to: technologies based on CRISPR and transposon systems, such as homologous directed repair (HDR); INTEGRATE system, which is disclosed in Strecker, Jonathan et al. “RNA-guided DNA insertion with CRISPR-associated transposases.” Science (New York, N.Y.) vol. 365, 6448 (2019): 48-53. doi:10.1126 / science.aax9181, and the like; CRISPR-associated transposase (CAST) technology, which is disclosed in Lampe, George D et al. “Targeted DNA integration in human cells without double-strand breaks using CRISPR RNA-guided transposases.” bioRxiv: the preprint server for biology 2023.03.17.533036. 18 Mar. 2023, doi:10.1101 / 2023.03.17.533036. Preprint. and the like, which are incorporated herein by reference.

[0251] As used herein, the term “reverse transcriptase” is an enzyme that directs the synthesis of deoxyribonucleotide triphosphates into complementary DNA (cDNA) using RNA as a template.

[0252] In some embodiments, the reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii mature enzyme RT (Marathon RT).

[0253] In some embodiments, the guide RNA comprises a scaffold sequence, an RS integration template sequence, and a primer binding sequence, the RS integration template sequence being connected to the primer binding sequence.

[0254] As used herein, the terms "guide RNA" and "gRNA" are used interchangeably to refer to an RNA molecule capable of forming a complex with a CRISPR effector protein and capable of targeting the complex to a target sequence due to a certain identity with the target sequence. The guide RNA targets the target sequence through base pairing between the complementary strands of the target sequence. The guide RNA functions in conjunction with the CRISPR effector protein is collectively referred to as "scaffold sequence", for example, the scaffold sequence adopted by Cas9 nuclease or its functional variants is usually composed of crRNA and tracrRNA that form a complex in part complementarily. "Primer binding sequence" and "PBS" are used interchangeably to refer to a sequence of at least 10, at least 15 or at least 20 consecutive nucleotides of the guide RNA complementary to the target sequence. In some embodiments, the recombinase recognition site sequence is linked to the primer binding sequence. In some embodiments, the "guide RNA" is a single guide RNA (sgRNA), a guide RNA of prime editor (pegRNA). It is within the ability of those skilled in the art to design a suitable guide RNA based on the CRISPR nuclease used and the target sequence to be edited.

[0255] In some embodiments, the CRISPR nickase of the recombinase recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with the CRISPR nickase and targets a desired location in the genome, wherein the CRISPR nickase makes a cut in the strand of the genome, and the reverse transcriptase incorporates the RS integration template sequence in the guide RNA into the cut site, thereby inserting at least one RS recognizable by the recombinase at the desired location of the genome.

[0256] In some embodiments, the recombinase site (RS) is selected from an attP site, an attB site. In some embodiments, the attP site comprises a nucleotide sequence set forth in SEQ ID NO: 2, and the attB site comprises a nucleotide sequence set forth in SEQ ID NO: 3.

[0257] In some embodiments, a RS can be inserted into the genome or a fragment thereof of a cell using a nuclease, a gRNA, and / or an integrase. In some embodiments, the RS is carried on a guide RNA. The guide RNA can target any site known in the art. The complementary RS can be operably linked to a gene or nucleic acid sequence of interest of an exogenous DNA or RNA. In some embodiments, one RS is added to the target genome. In some embodiments, more than one RS is added to the target genome.

[0258] In some embodiments, the recombinase recognition site integration unit is not covalently linked to a recombinase.

[0259] In some embodiments, the recombinase recognition site integration unit is covalently linked to the recombinase.

[0260] In some embodiments, the genome editing system further comprises:

[0261] 3) a donor comprising an exogenous nucleotide sequence to be inserted into the genome (hereinafter also referred to as component 3).

[0262] In some embodiments, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer, such as at least 50 bp, at least 100 bp, at least 300 bp, at least 500 bp, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, at least 20 kb, at least 50 kb.

[0263] In some embodiments, the exogenous nucleotide sequence further comprises one or more RSs that can be recognized by the recombinase.

[0264] In some embodiments, the exogenous nucleotide sequence is flanked by one or two RSs.

[0265] IV. Fusion protein

[0266] In another aspect of the present application, a fusion protein / chimeric protein comprising the genome editing system of any one of the above is provided, wherein the sequence-specific DNA cleavage protein, the reverse transcriptase and the recombinase are directly linked or linked via a linker.

[0267] In some specific embodiments, the sequence-specific DNA cleavage protein, such as a CRISPR nickase, is linked at the C-terminus to the reverse transcriptase, which is linked via a linker to the recombinase.

[0268] As used herein, the term "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids, without secondary structure above. For example, the linker can be a flexible linker.

[0269] V. Polynucleotide

[0270] In one aspect, the present application provides a polynucleotide comprising a nucleotide sequence encoding the aforementioned recombinase, the aforementioned genome editing system, or the aforementioned fusion protein.

[0271] The polynucleotide of the present application can be in the form of DNA or RNA. The DNA form includes cDNA, genomic DNA, or artificially synthesized DNA. The DNA can be single-stranded or double-stranded. The DNA can be a coding strand or a non-coding strand. The RNA form includes mRNA.

[0272] The polynucleotide encoding the recombinase of the present application includes: a coding sequence encoding only the recombinase; a coding sequence of the recombinase and various additional coding sequences; a coding sequence of the recombinase (and optional additional coding sequences) and non-coding sequences.

[0273] Six, expression construct

[0274] In one aspect, the present application provides an expression construct comprising the aforementioned polynucleotide.

[0275] In some embodiments, the expression construct is selected from the group consisting of viral, bacterial, yeast, plant, mammalian cell expression constructs.

[0276] Seven, cell

[0277] In one aspect, the present application provides a cell comprising the aforementioned recombinase, the aforementioned genome editing system, the aforementioned fusion protein, the aforementioned polynucleotide, or the aforementioned expression construct.

[0278] Eight, kit

[0279] In one aspect, the present application provides a kit comprising the aforementioned recombinase, the aforementioned polynucleotide, the aforementioned genome editing system, the aforementioned fusion protein, the aforementioned expression construct, or the aforementioned cell.

[0280] Nine, method of genome editing

[0281] In another aspect of the present application, a method of genome editing in an organism or an organism cell is provided, selected from: i) introducing into the organism or the organism cell the aforementioned recombinase and a donor comprising an exogenous nucleotide sequence to be inserted into the genome, wherein the target genome of the organism or the organism cell contains a recombinase recognition site sequence (RS) corresponding to the recombinase; or ii) introducing into the organism or the organism cell the aforementioned genome editing system. In some embodiments, the method of genome editing can result in one or more nucleotide substitutions, or one or more nucleotide insertions. In some embodiments, the length of the substitution of the plurality of nucleotides is at least 2 bp, at least 10 bp, at least 100 bp, at least 1 kbp, at least 10 kbp, at least 20 kbp, at least 50 kbp.

[0282] In some embodiments, components 1), 2) and optional component 3) of the genome editing system are introduced into the organism or organism cell simultaneously.

[0283] In some embodiments, components 1), 2) or optional component 3) of the genome editing system are introduced into the organism or organism cell separately in steps.

[0284] In some embodiments, component 2) of the genome editing system is introduced into the organism or organism cell separately.

[0285] The skilled person can select the RS insertion into the genome and / or the foreign nucleotide sequence, and the direction of insertion into the genome or foreign nucleotide sequence, according to the purpose of DNA recombination, such as deletion, inversion, integration, etc.

[0286] In some embodiments, component 1) is inserted into the donor construct of the genome or foreign nucleotide sequence in the same direction or in the opposite direction.

[0287] In some embodiments, the method comprises recombining DNA of the genome of the organism or organism cell.

[0288] In some specific embodiments, the recombining DNA of the genome of the organism or organism cell comprises deleting DNA, inverting DNA in the genome, and / or integrating foreign DNA into the genome.

[0289] In the method of the present application, the meganuclease or genome editing system can be introduced into the cell by various methods well known to the skilled person.

[0290] In some embodiments, the method of introducing the recombinase or genome editing system into the organism or organism cell is selected from the group consisting of: calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), biolistics, N-acetylgalactosamine (GalNAc)-mediated, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.

[0291] In some embodiments, the organism or organism cell is from a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; a plant, preferably a crop plant, such as wheat, rice, corn, soybean, sunflower, leafy greens, lettuce, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, kiwi, tobacco, cassava, potato.

[0292] In the present application, the target nucleic acid region to be edited can be located at any position of the genome, for example, at a safe harbor of the genome, within a functional gene such as a protein-coding gene, or for example, can be located at a gene expression regulatory region such as a promoter region or an enhancer region, so as to achieve modification of the function of the gene or modification of the gene expression. In some embodiments, the desired nucleotide sequence substitution results in a desired modification of the function of the gene or modification of the gene expression.

[0293] EMBODIMENTS

[0294] The embodiments of the present application will be described in detail below with reference to the examples, but those skilled in the art will understand that the following examples are only for illustration of the present application and should not be regarded as limiting the scope of the present application. If no specific conditions are specified in the examples, the conventional conditions or the conditions recommended by the manufacturer are used. If no manufacturer of the reagent or instrument used is specified, it is a conventional product that can be obtained by purchase on the market.

[0295] Example 1 Screening of Bxbl single mutants with improved recombination activity

[0296] The inventors designed two parallel screening systems to screen the screening library of Bxbl single mutants constructed by single-point saturation mutation, and carried out multi-dimensional coverage based on the screening strategy of functional complementation to avoid missing screening.

[0297] Firstly, the inventors designed an evolutionary screening system 1 in tobacco leaves Figure 1), the expression of Bxbl is initiated by CaMV 35S promoter and terminated by NOS terminator, a screening library is constructed by introducing single point saturation mutation in the region of 150 amino acids at the N-terminus of Bxbl using NNK degenerate primer, the N-terminus of Bxbl is equally divided into 15 fragments, each of which contains single point saturation mutation of 10 amino acids, the two recognition sites (RS) of attB and attP are opposite in direction, a fragment of about 1 kb is reversely inserted between the two sites, after the construction of the 15 libraries, Agrobacterium is transformed respectively to prepare Agrobacterium library. Then Agrobacterium is used to infect tobacco leaves, and tobacco leaves are taken after two days, and genome is extracted. PCR is performed using genome as template, two pairs of primers are used for PCR respectively, the first pair of primers is F1+R1 (SEQ ID NO: 11, SEQ ID NO: 13), and product 1 is obtained as background control; the second pair of primers is F2+R2 (SEQ ID NO: 12, SEQ ID NO: 14), which is used to amplify the fragment containing Bxbl variant with sufficient activity which occurs inversion, as product 2. Among them, the higher the proportion in product 2 is, the higher the activity of the corresponding Bxbl variant is. Subsequently, product 1 and 2 are used as templates to continue PCR reaction, the same pair of primers is used to enrich the short fragment containing Bxbl variant in product 1 (using primer F1+R1) and product 2 (using primer F2+R2) respectively, and then barcode label is added for second-generation sequencing, the proportion of different variants in product 1 and product 2 is calculated, and the enrichment degree of active Bxbl variant in product 2 relative to product 1 is further calculated, and the higher the enrichment degree is, the higher the activity of the Bxbl variant is. Each library is biologically repeated three times, and the Bxbl variant which is enriched in three biological repeats is screened, and the enrichment degree is the average value of three biological repeats.

[0298] The screening results are shown in Table 1, and the mutants with average enrichment degree of more than 10% are listed in the table, which proves that these screened mutants have recombinase activity and can be used as potential Bxbl active mutants.

[0299] Table 1 Bxbl mutants with potential recombinase activity

[0300]

[0301] Secondly, the inventors designed an evolution screening system 2 in tobacco leaves Figure 2The screening system 2 consists of two vectors, the first binary vector is the receptor vector (Vector 1, SEQ ID NO: 24), which expresses a truncated firefly luciferase nF-LUC initiated by the cassava vein mosaic virus CSVMV promoter (P-CSVMV) and terminated by the tobacco mosaic virus-derived CaMV 35S terminator (T-CaMV) at the 3' end, and an intron sequence and the recognition sequence attB of Bxbl. The second binary vector is the donor vector (Vector 2), which contains a sea pansy luciferase (R-LUC) expression frame; another recognition sequence attP of Bxbl, an intron and the C-terminal of firefly luciferase (cF-LUC); and a Bxbl expression frame.

[0302] Firstly, wild type Bxbl (WT_Bxbl) and inactivated Bxbl (dead_Bxbl) were used to detect whether the evolution screening system works in tobacco leaves. Vector 1 and Vector 2 (WT_Bxbl or dead_Bxbl) were transformed into Agrobacterium, respectively, and co-infected tobacco leaves. The sequence of Vector 1 is shown in SEQ ID NO: 24, the sequence of Vector 2 (WT_Bxbl) is shown in SEQ ID NO: 25, and the sequence of Vector 2 (dead_Bxbl) is shown in SEQ ID NO: 26. Figure 2 The results are shown in FIG. 2a and FIG. 2b. The WT_Bxbl sequence in the sequence of FIG. 2a and SEQ ID NO: 25 was replaced by dead_Bxbl. After 48 hours, the infected leaves were detected for luminescence signals, and the results are shown in FIG. 2a and FIG. 2b. Figure 2 As shown in FIG. 2b, the leaves expressing WT_Bxbl detected obvious firefly luciferase signals, while the leaves expressing dead_Bxbl had almost no signals, indicating that the evolution screening system 2 can be used to screen Bxbl variants with higher integration activity in tobacco leaves.

[0303] The screening library was constructed by introducing single-point saturation mutations into the catalytic region of Bxb1 (containing 150 amino acids) using NNK degenerate primers. The N-terminal of Bxb1 was divided into 15 fragments, each containing a single-point saturation mutation of 10 amino acids. After constructing the 15 libraries, Agrobacterium was transformed, and Agrobacterium libraries were prepared. Then, Agrobacterium was used to infect tobacco leaves, and tobacco leaves were collected after two days. Genomic DNA was extracted from the tobacco leaves. PCR was performed using genomic DNA as a template. Two pairs of primers were used for PCR. The first pair of primers was P1+P3 (SEQ ID NO: 26, SEQ ID NO: 28), and product 1 was obtained as a background control. The second pair of primers was P2+P3 (SEQ ID NO: 27, SEQ ID NO: 28), and product 2 was obtained by amplifying the fragment containing the Bxb1 variant with sufficient activity. The Bxb1 variant with high activity theoretically has a high proportion in product 2. Then, product 1 and product 2 were used as templates for further PCR. The same pair of primers was used to enrich the short fragment containing the Bxb1 variant in product 1 (using primers F1+R1) and product 2 (using primers F2+R2). After adding a barcode tag, second-generation sequencing was performed. The proportion of different variants in product 1 and product 2 was calculated. The enrichment degree of the active Bxb1 variant in product 2 relative to product 1 was further calculated. The higher the enrichment degree, the higher the activity of the Bxb1 variant. Each library was biologically repeated three times. The Bxb1 variant that was enriched in three biological repeats was selected, and the enrichment degree was calculated as the average of three biological repeats.

[0304] Meanwhile, the evolution screening system plasmid was transformed into rice and corn protoplasts using the PEG (polyethylene glycol) method. After 48 hours of transformation, the protoplasts were collected, and genomic DNA was extracted. The samples were prepared using the above primers and methods, and second-generation sequencing analysis was performed. Three biological repeats were performed.

[0305] The screening results are shown in Table 2. The mutants that were enriched in three biological repeats and at least in two species were used as the Bxb1 mutants with high activity.

[0306] Table 2: Candidate Bxb1 single mutants with improved integration activity

[0307]

[0308] In summary, the screening results of the two parallel screening systems are summarized, and the final screened candidate recombinant activity improved Bxbl mutant sites are: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C, S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, D45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, P52L, P52M, P52V, F53H, F53L, F53M, F53R, F53W, R55A, R55Q, R55W, R55Y, P59D, P59G, N60M, A62H, A62N, R63C, R63Q, F67S, E68D, E68P, E68Q, E68R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I75F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S, A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, L103M, L103Q, V104C, V105E, V105Q, S106N, A107V, T108A, T108G, T108Y, A110C, A110L, A110M, A110Q,A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, F118S.

[0309] Example 2: Recombination activity verification of Bxbl single mutants

[0310] In molecular biology, Bxb1 recombination activity can be divided into three types according to the spatial arrangement of its recognition sites (attB / attP) and the reaction results: integration, deletion and inversion. The integration type recombination refers to the specific recombination between the attP site on the foreign DNA and the host genome attB site (same direction arrangement) mediated by Bxb1 recombinase, which inserts the foreign DNA into the host genome at a specific site and generates a new hybrid site attL and attR. The deletion type recombination refers to the recognition of two same direction arranged target sites (attB / attP) on the same DNA molecule by the enzyme, which loops out the middle segment and generates hybrid sites attL and attR, respectively. The inversion type recombination refers to the catalysis of the 180° inversion of the middle DNA segment with the same direction of the target site (attB and reverse attP, etc.) on the same DNA molecule. The above three types of activity promotion can be counted as the promotion of the recombination activity of the recombinase. For this, the inventors constructed two sets of verification systems to verify the recombination editing activity of the candidate Bxb1 single mutant sites in Example 1.

[0311] Firstly, the inventors constructed a recombinase activity verification system 1 for recombination activity characterization. As shown in Figure 3 , the tobacco mosaic virus CaMV 35S promoter (P-CaMV 35S) initiates the expression of the recombinase (e.g. Bxb1 or its variant), the T-NOS terminator terminates the expression of the recombinase (Bxb1 or its variant), the attB and attP sequences are arranged in opposite directions, and the reverse CaMV 35S promoter sequence (P-CaMV 35S) initiates the expression of firefly luciferase, a segment from the Agrobacterium tumefaciens 3' untranslated region sequence (T-ORF1 3'UTR) terminates the expression of firefly luciferase, the cassava vein mosaic virus CSVMV promoter (P-CSVMV) initiates the expression of Renilla luciferase, and the CaMV 35S terminator (T-CaMV 35S) terminates the expression of Renilla luciferase. Only after the active recombinase (Bxb1 or its variant) is expressed, the sequence of the CaMV 35S promoter in the att site is inverted to start the expression of firefly luciferase. The activity of firefly luciferase is detected by a microplate reader, representing the activity of Bxb1 and its variants. The activity of the constantly expressed Renilla luciferase is used as an internal reference to correct the expression of firefly luciferase. The expression construct of the recombinase activity quantitative system 1 is shown in Figure 3 . The luciferase detection kit used is Transdetect Double-Luciferase Reporter Assay Kit (Transgen, FR201).

[0312] The data of three biological repeats of fold activity improvement in tobacco leaf relative to wild type Bxbl (Bxbl-WT; SEQ ID NO: 1) were averaged, the activity results were normalized, Bxbl-WT activity as 1, Bxbl mutant activity as fold of 1 were compared. The results are shown in Table 3, in which are the preferred Bxbl single mutants, whose editing activity is 1.01-8.86 fold of Bxbl-WT, indicating that the mutation of these preferred sites has significantly improved the recombination activity of Bxbl recombinase.

[0313] Table 3 Verification results of verification system 1 of recombinase activity

[0314]

[0315] Therefore, by the recombination enzyme activity verification system 1, the inventors verified a batch of single mutation sites with improved recombination activity, which are: R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q.

[0316] Secondly, the inventors constructed a recombination enzyme activity verification system 2 to perform supplementary verification of editing activity. A recognition site of the recombination enzyme was inserted into a safe harbor site of the genome by using a Prime Editor (PE) and a double pegRNA, and the principle of inserting the RS is as follows Figure 4The other recognition site of the recombinase is placed near the donor large fragment DNA, and when the recombinase is expressed, the sites on the inserted genome and the sites on the donor are respectively recognized to form a tetramer, activate the recombinase activity, and thus integrate the donor large fragment DNA into the genome. In order to facilitate detection of integration efficiency, a R primer on the genome is pre-placed on the donor large fragment DNA vector.

[0317] The endogenous recombination activity of the candidate Bxbl mutants was verified in rice and human cells respectively, and the schematic diagram of the expression construct of the recombinase is shown in Figure 5 Figure 5 a is used for rice cells, Figure 5 b is used for human cells). The selected site in rice is a safe harbor position GSH1 on the genome (kitaake, Chr1:7660637-7661671), the integration vector 1 is an 8.5 kb donor vector (SEQ ID NO: 19), the PE tool used is ePPE (SEQ ID NO: 16), and the double pegRNA is shown as SEQ ID NO: 17 and SEQ ID NO: 18; the integration vector 2 is a 5.7 kb donor vector (SEQ ID NO: 35), the PE tool used is ePPE (SEQ ID NO: 16), and the double pegRNA is shown as SEQ ID NO: 36 and SEQ ID NO: 37; the TRAC site is selected in human cells HEK293T, and the integration activity of the 5.6 kb donor vector (SEQ ID NO: 23) of the variant is verified, the PE tool used is PE2 (SEQ ID NO: 20), and the double pegRNA is shown as SEQ ID NO: 21 and SEQ ID NO: 22.

[0318] The activity detection method is shown in Figure 6 ​As shown, a pair of F and R primers amplify genomic DNA, if recombination occurs, then three sequences can be amplified, one is the wild type sequence (genome amplicon), the second is the sequence after PE work and install attB (attB installed amplicon), and the third is the sequence after the work of recombinase and the insertion of donor into the genome (insertion of donor). After second-generation sequencing, the number of three kinds of amplicons is counted respectively by using the specific marker sequences of the three sequences, wherein the number of reads of the second and third divided by the total number of reads = the efficiency of PE, and the number of reads of the third divided by the number of reads of the second and third = the recombination efficiency of the recombinase. The F and R primers used for amplification in rice cells are SEQ ID NO: 33 and SEQ ID NO: 34, respectively; and the F and R primers used for amplification in human cells are SEQ ID NO: 30 and SEQ ID NO: 31, respectively.

[0319] Table 4: Correspondence between variant name and mutation type

[0320]

[0321]

[0322] The recombination activity detection results of the recombinase activity verification system 2 are shown in Figures 7-9 As shown in the figure, PE represents the efficiency of inserting the recognition site of Bxb1 at the genomic safe harbor site, and IN represents the recombination efficiency of the Bxb1 recombinase. The correspondence between the variant name and the mutation type is shown in Table 4.

[0323] The recombination activity verification results of integrating vector 1 in rice cells are shown in Figure 7 It is found through verification that the recombination efficiency of V27, V39, V53, and V54, a total of 4 variants, is significantly improved compared with the wild type Bxb1, and the corresponding mutation types are A3G, L4M, V129A, and I123T. The recombination activity of the above-mentioned variants is improved by 1.3-1.7 times of WT_Bxb1.

[0324] The recombination activity verification results of integrating vector 2 in rice cells are shown in Figure 8 V27 and V39 still maintain high recombination efficiency, and the recombination efficiency of V1 variant is improved by 1.69 times compared with the wild type Bxb1 activity, even slightly higher than the recombination efficiency of V27 and V39, proving that R2Q is also an ideal mutation type for efficiency improvement.

[0325] The recombination activity verification results in human cells are shown in Figure 9As shown, hV27 and hV39 have a slight increase compared with WT Bxbl, and the corresponding mutation types are A3G and L4M; hV30 has a significant increase compared with WT Bxbl, and the corresponding mutation type is V76C, and the recombination activity reaches 1.9 times that of WT_Bxbl.

[0326] Therefore, through the recombination enzyme activity verification system 2, the inventors verified a batch of single-mutation sites with increased recombination activity: R2Q, A3G, L4M, V129A, I123T and V76C.

[0327] Based on the above, the single mutation sites with improved recombination activity verified by the two sets of verification systems were summarized, and the Bxbl mutation sites with significantly improved recombination activity were finally obtained as follows: R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, R2Q, A3G, L4M, V129A, I123T and V76C.

[0328] Experimental Example 3: Verification of the activity of Bxbl multi-mutants with improved recombination activity in tobacco

[0329] The highly active Bxb1 single mutation sites from Experiments 1 and 2 were combined to obtain a series of Bxb1 mutant libraries with multiple mutation sites. The recombinase activity verification system 1 from Experiment 1 was used to identify the activity of these Bxb1 mutant libraries with multiple mutation sites. The verification method was the same as in Experiment 2: the average of three biological replicates of activity enhancement relative to WT_Bxb1 obtained from tobacco leaves was taken, and the activity results were normalized, with WT_Bxb1 activity as 1 and Bxb1 mutant activity as a multiple of 1 for comparison. Table 5 lists the combinations of mutation sites and the activity verification results. The efficiency enhancement of double mutants ranged from 1.72 to 5.1, and the efficiency enhancement of triple mutants ranged from 3.35 to 4.48. Among these multi-mutant mutants with enhanced recombination activity, the multi-mutant recombination activity enhancement was more significant when using T84N, L93A, I87T, or A144S as the baseline mutants. Figure 10 As shown, especially T84N, its recombinant activity with dual-protrusion combinations of H146C, P71M, E42L, E69M, D54C or A66W is significantly improved compared to the single-protrusion activity of T84N, and the activity improvement is even more obvious compared to WT_Bxb1.

[0330] In summary, the results above all indicate that both double and triple mutants of Bxb1 have the effect of further enhancing recombination activity.

[0331] Table 5. Bxb1 N mutants with enhanced editing activity and activity verification results.

[0332]

[0333] Experiment Example 4: Validation of the activity of recombinant Bxb1 multiple mutants with enhanced recombinant activity in rice protoplasts

[0334] The recombinase activity verification system 1 was introduced into rice protoplast cells. The results showed that the enhanced activity of the Bxb1 multimutant recombinase could also undergo efficient DNA recombination in rice cells. Figure 11 As shown, the activity of WT_Bxb1 was taken as 1, and the recombination activity of Bxb1 mutant was taken as a multiple of 1 for comparison. The results showed that in rice cells, the recombination activity of double mutants composed of T84N as the baseline mutant and Q92C, H146C, P71M, E42L, E69M or A66W, and the recombination activity of double mutants composed of I87T as the baseline mutant and A144G were significantly improved. Among them, the combination of E69M+T84N had the highest activity, which was 11.69 times that of WT_Bxb1.

[0335] It should be noted that although the technical solution of the present invention has been described with specific examples, those skilled in the art will understand that the present invention should not be limited thereto.

[0336] The foregoing description of various embodiments of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the various embodiments of the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It was described above, and will be apparent from this disclosure and various embodiments, that each of the foregoing embodiments can be implemented alone or in various combination with one another. It will be apparent to those ordinarily skilled in the art that many modifications and variations are possible in light of the above teachings. It is intended that the various embodiments disclosed in this detailed description be considered in all their novel respects, alone and in various combination and permutation. The various embodiments described herein are chosen and described in order to best explain the principles of the various embodiments and their best mode of operation. The use of certain terms in various places of this specification is intended to be interpreted not as a limitation but rather meant to educate the reader on the principles of the various embodiments.

[0337] Sequences referred to in the specification:

[0338] >SEQ ID NO: 1 BxBl

[0339] MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMSEF

[0340] >SEQ ID NO: 2 attP

[0341] GGTTTGTCTGGTCAACCACCGCGGWCTCAGTGGTGTACGGTACAAACC

[0342] Where: W is A (typically for integration) or T (typically for inversion).

[0343] >SEQ ID NO: 3 attB

[0344] GGCCGGCTTGTCGACGACGGCGGWCTCCGTCGTCAGGATCATCCGG

[0345] wherein: W is A (typically for integration) or T (typically for inversion).

[0346] >SEQ ID NO:4 CaMV 35S

[0347] tgagactttt caacaaaggg taatatccgg aaacctcctc ggattccatt gcccagctat ctgtcacttt attgtgaaga tggtggaaaa ggaaggtggc tcctacaaat gccatcattg cgataaagga aaggccatcg ttgaagatgc ctctgccgac agtggtccca aagatggacc cccacccacg aggagcatcg tggaaaaaga agacgttcca accacgtctt caaagcaagt ggattgatgt gatatctcca ctgacgtaag ggatgacgcacaatcccactatccttcgcaagacccttcctctatataaggaagttcatttcatttggagagaaca

[0348] >SEQ ID NO:5 T-NOS

[0349] gatcgttcaaacatttggcaataaagtttcttaagattgaatcctgttgccggtcttgcgatgattatcatataatttctgttgaattacgttaagcatgtaataattaacatgtaatgcatgacgttatttatgagatgggtttttatgattagagtcccgcaattatacatttaatacgcgatagaaaacaaaatatagcgcgcaaactaggataaattatcgcgcgcggtgtcatctatgttactagatc

[0350] >SEQ ID NO:6 Firefly luciferase

[0351]

[0352] >SEQ ID NO: 7 T-At ORF1 3'UTR

[0353] AGATCGGCGGCAATAGCTTCTTAGCGCCATCCCGGGTTGATCCTATCTGTGTTGAAATAGTTGCGGTGGGCAAGGCTCTCTTTCAGAAAGACAGGCGGCCAAAGGAACCCAAGGTGAGGTGGGCTATGGCTCTCAGTTCCTTGTGGAAGCGCTTGGTCTAAGGTGCAGAGGTGTTAGCGGGATGAAGCAAAAGTGTCCGATTGTAACAAGATATGTTGATCCTACGTAAGGATATTAAAGTATGTATTCATCACTAATATAATCAGTGTATTCCAATATGTACTACGATTTCCAATGTCTTTATTGTCGCCGTATGTAATCGGCGTCACAAAATAATCCCCGGTGACTTTCTTTTAATCCAGGATGAAATAATATGTTATTATAATTTTTGCGATTTGGTCCGTTATAGGAATTGAAGTGTGCTTGCGGTCGCCACCACTCCCATTTCATAATTTTACATGTATTTGAAAAATAAAAATTTATGGTATTCAATTTAAACACGTATACTTGTAAAGAATGATATCTTGAAAGAAATATAGTTTAAATATTTATTGATAAAATAACAAGTCAGGTATTATAGTCCAAGCAAAAACATAAATTTATTGATGCAAGTTTAAATTCAGAAATATTTCAATAACTGATTATATCAGCTGGTACATTGCCGTAGATGAAAGACTGAGTGCGATATTATGGTGTAATACATA

[0354] >SEQ ID NO: 8 P-CSVMV

[0355] CCAGAAGGTAATTATCCAAGATGTAGCATCAAGAATCCAATGTTTACGGGAAAAACTATGGAAGTATTATGTAAGCTCAGCAAGAAGCAGATCAATATGCGGCACATATGCAACCTATGTTCAAAAATGAAGAATGTACAGATACAAGATCCTATACTGCCAGAATACGAAGAAGAATACGTAGAAATTGAAAAAGAAGAACCAGGCGAAGAAAAGAATCTTGAAGACGTAAGCACTGACGACAACAATGAAAAGAAGAAGATAAGGTCGGTGATTGTGAAAGAGACATAGAGGACACATGTAAGGTGGAAAATGTAAGGGCGGAAAGTAACCTTATCACAAAGGAATCTTATCCCCCACTACTTATCCTTTTATATTTTTCCGTGTCATTTTTGCCCTTGAGTTTTCCTATATAAGGAACCAAGTTCGGCATTTGTGAAAACAAGAAAAAATTTGGTGTAAGCTATTTTCTTTGAAGTACTGAGGATACAACTTCAGAGAAATTTGTAAGTTTGTA

[0356] >SEQ ID NO: 9 Renilla luciferase

[0357] atgacttcgaaagtttatgatccagaacaaaggaaacggatgataactggtccgcagtggtgggccagatgtaaacaaatgaatgttcttgattcatttattaattattatgattcagaaaaacatgcagaaaatgctgttatttttttacatggtaacgcggcctcttcttatttatggcgacatgttgtgccacatattgagccagtagcgcggtgtattataccagaccttattggtatgggcaaatcaggcaaatctggtaatggttcttataggttacttgatcattacaaatatcttactgcatggtttgaacttcttaatttaccaaagaagatcatttttgtcggccatgattggggtgcttgtttggcatttcattatagctatgagcatcaagataagatcaaagcaatagttcacgctgaaagtgtagtagatgtgattgaatcatgggatgaatggcctgatattgaagaagatattgcgttgatcaaatctgaagaaggagaaaaaatggttttggagaataacttcttcgtggaaaccatgttgccatcaaaaatcatgagaaagttagaaccagaagaatttgcagcatatcttgaaccattcaaagagaaaggtgaagttcgtcgtccaacattatcatggcctcgtgaaatcccgttagtaaaaggtggtaaacctgacgttgtacaaattgttaggaattataatgcttatctacgtgcaagtgatgatttaccaaaaatgtttattgaatcggacccaggattcttttccaatgctattgttgaaggtgccaagaagtttcctaatactgaatttgtcaaagtaaaaggtcttcatttttcgcaagaagatgcacctgatgaaatgggaaaatatatcaaatcgttcgttgagcgagttctcaaaaatgaacaataa

[0358] >SEQ ID NO: 10 T-CaMV35S

[0359] tttctccata aataatgtgt gagtagtttc ccgataaggg aaattagggt tcttataggg ttcgctcatg tgttgagcat ataagaaacc cttagtatgt atttgtattt gtaaaatact tctatcaata aaatttctaa ttcctaaaac caaaatccag tactaaaatc cagatc

[0360] >SEQ ID NO: 11 F1

[0361] tacaattaca accATGggaC

[0362] >SEQ ID NO: 12 F2

[0363] TCCTCTATAT AAGGAAGTTC

[0364] >SEQ ID NO: 13 R1

[0365] cgcttagaca acttaataac acat

[0366] >SEQ ID NO: 14 R2

[0367] GTTGAAGTTCTGGTTGAA

[0368] >SEQ ID NO: 15 dead_Bxbl

[0369] MRALVVIRLARVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKCRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMSEF

[0370] >SEQ ID NO: 16 ePPE vector, for rice

[0371]

[0372]

[0373] >SEQ ID NO: 18 epegRNA2, for rice

[0374]

[0375] >SEQ ID NO: 19 8.5 kb donor

[0376]

[0377] >SEQ ID NO: 20 PE2 vector, human cells

[0378]

[0379] >SEQ ID NO:21 pegRNA1, human cell

[0380]

[0381] >SEQ ID NO: 22 pegRNA2, human cell

[0382]

[0383] >SEQ ID NO:23 5.6kb donor

[0384]

[0385] >SEQ ID NO:24 Vector1

[0386]

[0387] >SEQ ID NO: 25 Vector2 (WT_Bxbl)

[0388]

[0389] >SEQ ID NO: 26 Figure 2 P1

[0390] ttgtgccagagtccttcgat

[0391] >SEQ ID NO: 27 Figure 2 P2

[0392] GTAATACATATAAATGCCCGGGCGT

[0393] >SEQ ID NO: 28 Figure 2 P3

[0394] GTCAACTCTGGTAGGCAGGT

[0395] >SEQ ID NO: 29 WT_Bxbl expression vector for use in human cells

[0396]

[0397] >SEQ ID NO: 30 Specific binding sequence of F primer used for detection of efficiency in human cells

[0398] ACTTGCCAGCCCCACAGAG

[0399] >SEQ ID NO: 31 R primer sequence inserted in donor vector used in human cells

[0400] TCCAGTGACAAGTCTGTCTGC

[0401] >SEQ ID NO: 32 WT_Bxbl expression vector used in rice

[0402]

[0403] >SEQ ID NO:33 Specific binding sequence of F primer used for detection efficiency in rice

[0404] CTCATGTGCATGGAAGCATC

[0405] >SEQ ID NO:34 R primer sequence inserted on the donor vector used in rice

[0406] CTTGCTCCTTCTCACCGTG

[0407] >SEQ ID NO:35 5.7 kb donor, for rice

[0408]

[0409] >SEQ ID NO: 36 epegRNA3, for rice

[0410]

[0411] >SEQ ID NO: 37 epegRNA4, for rice

[0412]

Claims

1. A recombinase, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions, wherein the one or more amino acid substitutions include positions 2, 3, 4, 5, 6, 7, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 23, 24, 25, 26, 27, 28, 29, 31, 32, 34, 35, 36, 37, 38, 39, 40, 41, 42, 44, 45, 46, 49, 50, 51, 52, 53, 54, 55, 58, 59, 60, 62, 63, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 77, 78, 79, 83, 84, 86, 87, 88 of the amino acid sequence shown in SEQ ID NO:

1. The substitution of positions 89, 90, 92, 93, 94, 95, 96, 97, 98, 100, 101, 102, 103, 104, 105, 106, 107, 108, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 122, 123, 124, 125, 126, 128, 129, 132, 134, 135, 137, 138, 139, 142, 143, 144, 145, 146, 147, 148, 149, 150, or 151, and any combination thereof.

2. The recombinase according to claim 1, wherein, The recombinase has enhanced recombinant activity compared to the recombinase shown in SEQ ID NO:

1.

3. The recombinase according to claim 1 or 2, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions selected from the following: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C, S2 5D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32 C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V37M , V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, E42L, E42R, D45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, D51C, D51S, P52L, P52M, P5 2V, F53H, F53L, F53M, F53R, F53W, D54C, R55A, R55Q, R55W, R55Y, P59D, P59 G, N60M, A62H, A62N, R63C, R63Q, A66W, F67S, E68D, E68P, E68Q, E68R, E69M , E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I75F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84N, S86A, S8 6E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89S, Q92C, Q92 E. Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S, A97I, A97S , A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102 C. K102H, L103M, L103Q, V104C, V104L, V105E, V105Q, S106N, A107V, T108A,T108G, T108Y, A110C, A110L, A110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, ​​R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, or F118S.

4. The recombinase according to any one of claims 1-3, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions selected from the following: R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A 62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I 87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135 V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R15 0F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, R2Q, A3G, L4M, V129A, I123T or V76C.

5. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of T84N; and optionally, includes any one or more amino acid substitutions selected from H146C, A144S, P71M, P71S, Q92C, D51S, E42L, E42R, D51C, E69M, D54C, A66W, and V104L.

6. The recombinase according to claim 5, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of T84N, and any one or more amino acid substitutions selected from H146C, P71M, Q92C, E42L, E69M, D54C, and A66W.

7. The recombinase according to claim 5, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes amino acid substitutions of T84N and P71M, and any one or more amino acid substitutions selected from H146C and Q92C.

8. The recombinase according to claim 5, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes amino acid substitutions of T84N and P71S, and any one or more amino acid substitutions selected from H146C and Q92C.

9. The recombinase according to any one of claims 5-8, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes at least one of the following amino acid substitutions (a) to (s): (a)T84N; (b) T84N and E69M; (c)T84N and Q92C; (d)T84N and E42L; (e)T84N and A66W; (f)T84N and D54C; (g)T84N and H146C; (h)T84N and P71M; (i)T84N and P71S; (j)T84N and E42R; (k)T84N and V104L; (l)T84N and D51C; (m)T84N and A144S; (n)T84N and D51S; (o)T84N, P71M and H146C; (p)T84N, P71M and Q92C; (q)T84N, P71S and H146C; (r)T84N, P71S and Q92C.

10. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of I87T; and optionally, includes any one or more amino acid substitutions selected from Y78R, A144T, H146C, H146Q, F118M, M132L, and A144G. Preferably, the amino acid sequence of the recombinase comprises, relative to the amino acid sequence shown in SEQ ID NO:1, at least one of the following amino acid substitutions (a1) to (h1): (a1)I87T (b1) I87T and A144G; (c1)I87T and Y78R; (d1)I87T and A144T; (e1)I87T and H146C; (f1)I87T and H146Q; (g1)I87T and F118M; (h1)I87T and M132L.

11. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of A144S; and optionally, includes any one or more amino acid substitutions selected from V94G, Y78R, H146C, I87T, and Q92C. Preferably, the amino acid sequence of the recombinase comprises, relative to the amino acid sequence shown in SEQ ID NO:1, at least one of the following amino acid substitutions (a2) to (f2): (a2)A144S; (b2) A144S and V94G; (c2)A144S and Y78R; (d2)A144S and H146C; (e2)A144S and I87T; (f2)A144S and Q92C.

12. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of L93A; and optionally, includes any one or more amino acid substitutions selected from H146C and K101E. Preferably, the amino acid sequence of the recombinase comprises, relative to the amino acid sequence shown in SEQ ID NO:1, at least one of the following amino acid substitutions (a3) ​​to (c3): (a3)L93A; (b3) L93A and H146C; (c3)L93A and K101E.

13. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of Q92C; and optionally, includes any one or more amino acid substitutions selected from V94G and H146C. Preferably, the amino acid sequence of the recombinase comprises, relative to the amino acid sequence shown in SEQ ID NO:1, at least one of the following amino acid substitutions (a4) to (c4): (a4)Q92C; (b4)Q92C and V94G; (c4)Q92C and H146C.

14. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of M132L; and optionally, includes any one or more amino acid substitutions selected from V94G, Y78R, and H146C. Preferably, the amino acid sequence of the recombinase comprises, relative to the amino acid sequence shown in SEQ ID NO:1, at least one of the following amino acid substitutions (a5) to (d5): (a5)M132L; (b5) M132L and V94G; (c5)M132L and Y78R; (d5)M132L and H146C.

15. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of N148A; and optionally, includes any one or more amino acid substitutions selected from P59D, Y78R, and V94G. Preferably, the amino acid sequence of the recombinase comprises, relative to the amino acid sequence shown in SEQ ID NO:1, at least one of the following amino acid substitutions (a6) to (d6): (a6)N148A; (b6) N148A and P59D; (c6)N148A and Y78R; (d6)N148A and V94G.

16. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes an amino acid substitution of E42W; and optionally, includes any one or more amino acid substitutions selected from T84N, P59D, V94G, K102H, and Y78R. Preferably, the amino acid sequence of the recombinase comprises, relative to the amino acid sequence shown in SEQ ID NO:1, at least one of the following amino acid substitutions: (a7) to (f7) (a7)E42W; (b7)E42W and T84N; (c7)E42W and P59D; (d7)E42W and V94G; (e7)E42W and K102H; (f7)E42W and Y78R.

17. The recombinase according to any one of claims 1-4, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes any one or more amino acid substitutions selected from A3G, V129A, V76C, L4M, or I123T.

18. The recombinase according to claim 17, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes amino acid substitutions of R2Q, A3G, or L4M.

19. The recombinase according to claim 17, wherein, The amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes the substitution of amino acid V76C.

20. A genome editing system, wherein, The genome editing system includes: 1) Recombinase recognition site sequence (RS) containing attP or attB sites; 2) A recombinase and / or an expression construct containing a nucleotide sequence encoding the recombinase, wherein the recombinase includes the recombinase according to any one of claims 1-19.

21. The genome editing system according to claim 20, wherein, The recombinase recognition site sequence (RS) is either naturally occurring or engineered or optimized.

22. The genome editing system according to claim 21, wherein, One or more of the recombinase recognition site sequences (RS) are inserted into the desired location in the genome via a recombinase recognition site integration unit.

23. The genome editing system according to claim 22, wherein, The recombinase recognition site integration unit includes: a CRISPR effector protein or a functional variant thereof and / or an expression construct containing a nucleotide sequence encoding the CRISPR effector protein or a functional variant thereof, and at least one guide RNA and / or at least one expression construct containing a nucleotide sequence encoding the at least one guide RNA.

24. The genome editing system according to claim 23, wherein, The functional variant of the CRISPR effector protein is a CRISPR nuclease that has completely or partially lost its cleavage activity. Preferably, the functional variant of the CRISPR effector protein is a CRISPR cleavage enzyme, such as Cas9-D10A, Cas9-H840A, Cas12a cleavage enzyme, Cas12b cleavage enzyme, or TraC cleavage enzyme.

25. The genome editing system according to claim 24, wherein, The guide RNA comprises a scaffold sequence and a primer binding sequence, as well as a recombinase recognition site sequence (RS) integrating template sequence.

26. The genome editing system according to any one of claims 22-25, wherein, The recombinase recognition site integration unit further includes a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the reverse transcriptase.

27. The genome editing system according to claim 26, wherein, The CRISPR nicking enzyme of the recombinase recognition site integration unit is linked to a reverse transcriptase. The guide RNA interacts with the CRISPR nicking enzyme and targets a desired location in the genome. The CRISPR nicking enzyme nicks the strand of the genome, and the reverse transcriptase incorporates the RS integration template sequence from the guide RNA into the nicking site, thereby inserting at least one recombinase recognition site sequence (RS) that can be recognized by the recombinase at the desired location in the genome.

28. The genome editing system according to claim 27, wherein, The reverse transcriptase is selected from Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcription heteropolymerase (RTX), avian myeloblastoma virus reverse transcriptase (AMV-RT), and rectal eubacterium maturation enzyme RT (MarathonRT).

29. The genome editing system according to any one of claims 22-28, wherein, The recombinase recognition site integration unit is not covalently linked to the recombinase; or, the recombinase recognition site integration unit is covalently linked to the recombinase.

30. The genome editing system according to any one of claims 22-29, wherein, The genome editing system also includes: 3) A donor of the foreign nucleotide sequence to be inserted into the genome, optionally the foreign nucleotide sequence may be about 1 bp to about 50 kb or longer.

31. A fusion protein comprising the genome editing system of any one of claims 20-30, wherein the CRISPR nickase is linked to a reverse transcriptase at its C-terminus, and the reverse transcriptase is linked to a recombinase via a linker.

32. A polynucleotide comprising a nucleotide sequence encoding a recombinase according to any one of claims 1-19, a genome editing system according to any one of claims 20-30, or a fusion protein according to claim 31, wherein optionally, the polynucleotide is RNA, such as mRNA.

33. An expression construct comprising the polynucleotide of claim 32.

34. A cell comprising the recombinase of any one of claims 1-19, the genome editing system of any one of claims 20-30, the fusion protein of claim 31, the polynucleotide of claim 32, or the expression construct of claim 33.

35. A kit comprising the recombinase of any one of claims 1-19, the genome editing system of any one of claims 20-30, the fusion protein of claim 31, the polynucleotide of claim 32, the expression construct of claim 33, or the cell of claim 34.

36. A method for gene editing in an organism or somatic cells, wherein, The method introduces the recombinase of any one of claims 1-19, the polynucleotide of claim 32, the expression construct of claim 33, the genome editing system of any one of claims 20-30, or the fusion protein of claim 31 into an organism or somatic cells.

37. The method of claim 36, wherein, Components 1), 2), and optional 3) of the genome editing system are simultaneously introduced into an organism or somatic cells.

38. The method according to claim 36, wherein, The components 1), 2), or optional 3) of the genome editing system are introduced into an organism or its cells in steps.

39. The method according to any one of claims 36-38, wherein, Component 1) is inserted into the donor construct of the genome or exogenous nucleotide sequence in the same or opposite direction.

40. The method according to any one of claims 36-39, wherein, The method includes recombining the DNA of the genome of an organism or a somatic cell; optionally, the recombining of the DNA of the genome of an organism or a somatic cell includes deleting DNA, flipping DNA in the genome, and / or integrating foreign DNA into the genome.

41. The method according to any one of claims 36-40, wherein, The genome editing system is introduced into cells via methods selected from the following: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), gene gun method, N-acetylgalactosamine (GalNAc) mediated, PEG mediated protoplast transformation, and Agrobacterium tumefaciens mediated transformation.

42. The method according to any one of claims 36-41, wherein, The organism or its cells are derived from mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; or plants, preferably crop plants such as wheat, rice, corn, soybeans, sunflowers, leafy greens, lettuce, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, tobacco, cassava, and potatoes.

Citation Information

Patent Citations

  • Novel CRISPR gene editing system

    CN117187213A

  • Base editing system based on TraC effect protein and application thereof

    CN120026006A

  • Methods and compositions for editing nucleotide sequences

    WO2020191234A1

  • Novel crispr gene editing system

    WO2023232109A1

Cited By

  • Nuclease having improved salt tolerance and / or temperature performance

    US20250092372A1