Bxb1 recombinase with improved activity and use thereof
By optimizing the amino acid sequence of Bxb1 recombinase and constructing a genome editing system, the problem of low recombinase efficiency was solved, achieving efficient genome editing, expanding the application scenarios, and making it suitable for the treatment of animal and plant diseases and trait improvement.
Patent Information
- Application Number
- PCT/CN2025/112734
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-28
- Filing Date
- 2025-08-05
- Publication Date
- 2026-02-12
AI Technical Summary
Existing recombinases have low recombination efficiency and poor specificity in genome editing, and their application scenarios are limited, making it difficult to meet the needs of animal and plant disease treatment and trait improvement.
By optimizing the amino acid sequence of Bxb1 recombinase and introducing specific amino acid substitutions, its recombination activity and efficiency were improved, and a highly active Bxb1 mutant recombinase was developed. Combined with CRISPR effector proteins and reverse transcriptase, a genome editing system was constructed to achieve the insertion, deletion, and flipping of large DNA fragments.
It improves the editing activity and efficiency of recombinases, expands the application scenarios of gene editing tools, and is suitable for genome editing in a variety of organisms, including large-fragment DNA manipulation in animal and plant cells.
Smart Images

Figure CN2025112734_12022026_PF_FP_ABST
Abstract
Description
Bxb1 recombinase with improved activity and applications thereof
[0001] Priority and Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202411068225.0, filed August 5, 2024, entitled “Bxb1 Recombinase with Improved Activity and Applications Thereof,” and Chinese Patent Application No. 202510237975.4, filed February 28, 2025, entitled “Bxb1 Recombinase with Improved Activity and Applications Thereof.” The entire contents of the above-referenced patent applications are incorporated herein by reference in their entirety. TECHNICAL FIELD
[0003] The present application belongs to the field of genetic engineering. Specifically, the present application relates to a Bxb1 recombinase with improved activity and applications thereof. More specifically, the present application provides a Bxb1 mutant recombinase that can act on a genome, as well as a genome editing system and a fusion protein based thereon, a method for genome editing of an organism or cells of an organism using the recombinase, and a genetically modified organism and its offspring produced by the method. BACKGROUND
[0004] Site-specific recombinases catalyze the specific recombination of fragments between two specific DNA sequences, mediate DNA fragment integration, excision or inversion, and play a key role in the life cycle of many microorganisms, including bacteria and bacteriophages.
[0005] Due to the characteristics that the recombinase does not introduce a DNA double-strand break in the process of editing the genome and can edit large fragments of chromosomes, the recombinase can be used as an ideal gene editing tool for large fragment insertion, deletion and inversion of DNA, for the treatment of diseases and the improvement of traits in plants and animals.
[0006] Currently, the recombinases in the prior art have problems such as low recombination efficiency and poor specificity. For example, the recombination activity of tyrosine recombinase is reversible, resulting in low editing efficiency, and the tyrosine family recombinase has few choices, can effectively edit small cargo fragments, and in addition, tyrosine recombinase Cre has cytotoxicity after overexpression. Large serine recombinase (LSR) has the characteristics of irreversible recombination, making it a potential genome editing tool. However, LSR still has problems of low editing efficiency and single variety. So far, only a few LSRs have been discovered, including Bxb1 and PhiC31 recombinases, but their recombination efficiency in plant and animal cells is very limited. Therefore, finding a recombinase with high activity, high recombination efficiency and wide application scenarios is of great significance for expanding the existing DNA large fragment editing system and developing a library of gene editing tools for precise manipulation of target DNA sequences. SUMMARY
[0007] Problems to be solved by the invention
[0008] Site-specific recombinases can directly modify the genome at the target sequence, or work together with other genome editing tools to achieve precise modification of the genome. Finding recombinases with high catalytic activity, high recombination efficiency and wide application scenarios can expand existing gene editing techniques, and have great research potential for the treatment of animal and plant diseases, trait improvement, etc.
[0009] Solution to the problem
[0010] The first aspect of the present application provides a recombinase, wherein the amino acid sequence of the recombinase comprises one or more amino acid substitutions relative to the amino acid sequence shown in SEQ ID NO: 1, wherein the one or more amino acid substitutions include substitutions at positions 2, 3, 4, 5, 6, 7, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 23, 24, 25, 26, 27, 28, 29, 31, 32, 34, 35, 36, 37, 38, 39, 40, 41, 42, 44, 45, 46, 49, 50, 51, 52, 53, 54, 55, 58, 59, 60, 62, 63, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 77, 78, 79, 83, 84, 86, 87, 88, 89, 90, 92, 93, 94, 95, 96, 97, 98, 100, 101, 102, 103, 104, 105, 106, 107, 108, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 122, 123, 124, 125, 126, 128, 129, 132, 134, 135, 137, 138, 139, 142, 143, 144, 145, 146, 147, 148, 149, 150, or 151 of the amino acid sequence shown in SEQ ID NO: 1, and any combination thereof.
[0011] In some embodiments, the recombinase has improved recombination activity relative to the recombinase as shown in SEQ ID NO: 1.
[0012] In some embodiments, the amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions selected from the following: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C, S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V 37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, E42L, E42R, D 45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, D51C, D51S, P52L, P5 2M, P52V, F53H, F53L, F53M, F53R, F53W, D54C, R55A, R55Q, R55W, R55Y, P5 9D, P59G, N60M, A62H, A62N, R63C, R63Q, A66W, F67S, E68D, E68P, E68Q, E6 8R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I7 5F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84 N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89 S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S , A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V , K101Y, K102C, K102H, L103M, L103Q, V104C, V104L, V105E, V105Q, S106N,A107V, T108A, T108G, T108Y, A110C, A110L, A110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, or F118S.
[0013] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, R2Q, A3G, L4M, V129A, I123T, or V76C, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0014] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution T84N, relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more of the amino acid substitutions selected from H146C, A144S, P71M, P71S, Q92C, D51S, E42L, E42R, D51C, E69M, D54C, A66W, V104L.
[0015] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, P71M, Q92C, E42L, E69M, D54C, A66W.
[0016] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of T84N and P71M relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, Q92C.
[0017] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of T84N and P71S relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, Q92C.
[0018] In some embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution selected from the group consisting of (a) to (s) below relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0019] (a) T84N
[0020] (b) T84N and E69M;
[0021] (c) T84N and Q92C;
[0022] (d) T84N and E42L;
[0023] (e) T84N and A66W;
[0024] (f) T84N and D54C;
[0025] (g) T84N and H146C;
[0026] (h) T84N and P71M;
[0027] (i) T84N and P71S;
[0028] (j) T84N and E42R;
[0029] (k) T84N and V104L;
[0030] (l) T84N and D51C;
[0031] (m) T84N and A144S;
[0032] (n) T84N and D51S;
[0033] (o) T84N, P71M and H146C;
[0034] (p) T84N, P71M and Q92C;
[0035] (q) T84N, P71S and H146C;
[0036] (r) T84N, P71S and Q92C.
[0037] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of I87T relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from Y78R, A144T, H146C, H146Q, F118M, M132L, A144G.
[0038] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a1) to (h1) relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0039] (a1) I87T;
[0040] (b1) I87T and A144G;
[0041] (c1) I87T and Y78R;
[0042] (d1) I87T and A144T;
[0043] (e1) I87T and H146C;
[0044] (f1) I87T and H146Q;
[0045] (g1) I87T and F118M;
[0046] (h1) I87T and M132L.
[0047] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, Y78R, H146C, I87T, Q92C.
[0048] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a2) to (f2) relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0049] (a2) A144S;
[0050] (b2) A144S and V94G;
[0051] (c2) A144S and Y78R;
[0052] (d2) A144S and H146C;
[0053] (e2) A144S and I87T;
[0054] (f2) A144S and Q92C.
[0055] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from H146C, K101E.
[0056] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in any one of (a3) to (c3) relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0057] (a3) L93A;
[0058] (b3) L93A and H146C;
[0059] (c3) L93A and K101E.
[0060] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, H146C.
[0061] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in any one of (a4) to (c4) relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0062] (a4) Q92C;
[0063] (b4) Q92C and V94G;
[0064] (c4) Q92C and H146C.
[0065] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, Y78R, H146C.
[0066] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in (a5) to (d5) below relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0067] (a5) M132L;
[0068] (b5) M132L and V94G;
[0069] (c5) M132L and Y78R;
[0070] (d5) M132L and H146C.
[0071] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from P59D, Y78R, V94G.
[0072] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in (a6) to (d6) below relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0073] (a6) N148A;
[0074] (b6) N148A and P59D;
[0075] (c6) N148A and Y78R;
[0076] (d6) N148A and V94G.
[0077] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of E42W relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from T84N, P59D, V94G, K102H, Y78R.
[0078] In some preferred embodiments, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution set forth in (a7) to (f7) below relative to the amino acid sequence set forth in SEQ ID NO: 1:
[0079] (a7) E42W;
[0080] (b7) E42W and T84N;
[0081] (c7) E42W and P59D;
[0082] (d7) E42W and V94G;
[0083] (e7) E42W and K102H;
[0084] (f7) E42W and Y78R.
[0085] In some embodiments, the amino acid sequence of the recombinant enzyme comprises any one or more of the amino acid substitutions selected from A3G, V129A, V76C, L4M, or I123T relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0086] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution of R2Q, A3G, or L4M relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0087] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution of V76C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0088] The second aspect of the application provides a genome editing system, wherein the genome editing system comprises:
[0089] 1) a recombinase recognition site sequence (RS) comprising an attP or attB site;
[0090] 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase, the recombinase comprising the recombinase of the first aspect of the application.
[0091] In some embodiments, the recombinase recognition site sequence (RS) is naturally occurring or engineered or optimized.
[0092] In some embodiments, one or more of the recombinase recognition site sequence (RS) is inserted into a desired location in the genome by a recombinase recognition site integration unit.
[0093] In some embodiments, the recombinase recognition site integration unit comprises a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or the functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.
[0094] In some embodiments, the CRISPR effector protein functional variant is a CRISPR nuclease with full / partial loss of cleavage activity, preferably, the CRISPR effector protein functional variant is a CRISPR nickase, for example, Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase, or TraC nickase.
[0095] In some embodiments, the guide RNA comprises a scaffold sequence and a primer binding sequence, and a recombination enzyme recognition site sequence (RS) integration template sequence.
[0096] In some embodiments, the recombination enzyme recognition site integration unit further comprises a reverse transcriptase and / or an expression construct comprising a nucleotide sequence encoding the reverse transcriptase.
[0097] In some embodiments, the CRISPR nickase of the recombination enzyme recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with the CRISPR nickase and targets a desired location in the genome, wherein the CRISPR nickase makes a nick on a strand of the genome, and the reverse transcriptase incorporates the RS integration template sequence in the guide RNA into the nicked site, thereby inserting at least one recombination enzyme recognition site sequence (RS) recognizable by the recombination enzyme at the desired location of the genome.
[0098] In some embodiments, the reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii Marathonase RT (Marathon RT).
[0099] In some embodiments, the recombination enzyme recognition site integration unit is not covalently linked to a recombination enzyme; or, the recombination enzyme recognition site integration unit is covalently linked to a recombination enzyme.
[0100] In some embodiments, the genome editing system further comprises:
[0101] 3) a donor of an exogenous nucleotide sequence to be inserted into the genome, optionally, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer.
[0102] The third aspect of the present application provides a fusion protein comprising the genome editing system of the second aspect of the present application, wherein the CRISPR nickase is linked to a reverse transcriptase at the C-terminus, and the reverse transcriptase is linked to a recombination enzyme via a linker.
[0103] The fourth aspect of the application provides a polynucleotide comprising a nucleotide sequence encoding the recombinase of the first aspect of the application, the genome editing system of the second aspect of the application, or the fusion protein of the third aspect of the application. Optionally, the polynucleotide is RNA, e.g., mRNA.
[0104] The fifth aspect of the application provides an expression construct comprising the polynucleotide of the fourth aspect of the application.
[0105] The sixth aspect of the application provides a cell comprising the recombinase of the first aspect of the application, the genome editing system of the second aspect of the application, the fusion protein of the third aspect of the application, the polynucleotide of the fourth aspect of the application, or the expression construct of the fifth aspect of the application.
[0106] The seventh aspect of the application provides a kit comprising the recombinase of the first aspect of the application, the genome editing system of the second aspect of the application, the fusion protein of the third aspect of the application, the polynucleotide of the fourth aspect of the application, the expression construct of the fifth aspect of the application, or the cell of the sixth aspect of the application.
[0107] The eighth aspect of the application provides a method of performing gene editing in an organism or a cell of an organism, wherein the method introduces into the organism or the cell of the organism the recombinase of the first aspect of the application, the polynucleotide of the fourth aspect of the application, the expression construct of the fifth aspect of the application, the genome editing system of the second aspect of the application, or the fusion protein of the third aspect of the application.
[0108] In some embodiments, components 1), 2), and optionally 3) of the genome editing system are introduced into the organism or the cell of the organism simultaneously.
[0109] In some embodiments, components 1), 2), or optionally 3) of the genome editing system are introduced into the organism or the cell of the organism stepwise, respectively.
[0110] In some embodiments, the component 1) is inserted into the donor construct of the genomic or exogenous nucleotide sequence in the same direction or in the opposite direction.
[0111] In some embodiments, the method comprises recombining DNA of the genome of the organism or the cell of the organism; optionally, the recombining DNA of the genome of the organism or the cell of the organism comprises deleting DNA, inverting DNA in the genome, and / or integrating exogenous DNA into the genome.
[0112] In some embodiments, the genome editing system is introduced into the cell by a method selected from the group consisting of calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, viral infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, or other viruses), biolistics, N-acetylgalactosamine (GalNAc)-mediated, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.
[0113] In some embodiments, the organism or organism cell is from a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; a plant, preferably a crop plant, for example, wheat, rice, corn, soybean, sunflower, leafy vegetable, lettuce, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.
[0114] Effects of the Invention
[0115] The present inventors, through a large number of experiments and explorations, screened and optimized Bxbl recombinase, and unexpectedly obtained a mutant Bxbl recombinase with high recombination activity, the integration efficiency and inversion efficiency of the mutant Bxbl were higher than those of wild-type Bxbl, and the mutant Bxbl could be used for large-fragment DNA insertion and translocation in the genome. This discovery helps to improve the editing activity of the genome editing system based on the recombinase, and the multiple mutant Bxbl enriches the application scenarios and selection of the genome editing based on the recombinase. BRIEF DESCRIPTION OF DRAWINGS
[0116] Figure 1, Bxbl mutant screening system 1 (evolutionary screening system 1) expression construct and screening method schematic diagram.
[0117] Figure 2, Bxbl mutant screening system 2 (evolutionary screening system 2) expression construct schematic diagram.
[0118] Figure 3, recombinase activity verification system 1 expression construct schematic diagram.
[0119] Figure 4, recombinase activity verification system 2 double pegRNA insertion RS method schematic diagram in the genome.
[0120] Figure 5, recombinase activity verification system 2 recombinase expression construct schematic diagram.
[0121] Figure 6, recombinase activity verification system 2 detection method schematic diagram.
[0122] Figure 7, recombinase activity verification system 2 recombinase activity verification results of recombinant vector 1 in rice cells.
[0123] Figure 8, recombinase activity verification system 2 recombinase activity verification results of recombinant vector 2 in rice cells.
[0124] Figure 9, Recombinase activity verification results of Recombinase Activity Verification System 2 in human cells.
[0125] Figure 10, Recombinase activity verification results of Recombinase Activity Verification System 1 in tobacco for Bxbl multi-mutants.
[0126] Figure 11, Recombinase activity verification results of Recombinase Activity Verification System 1 in rice for Bxbl multi-mutants. DETAILED DESCRIPTION
[0127] I. DEFINITIONS
[0128] In the present application, the scientific and technical terms used herein have the meanings commonly understood by one of ordinary skill in the art, unless indicated otherwise. Also, the terms and experimental procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology used herein are terms and procedures commonly employed by those skilled in the relevant art. For example, the standard recombinant DNA and molecular cloning techniques used in the present application are well known in the art and are described more fully in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter "Sambrook"). Also, for better understanding of the present application, the definitions and explanations of the related terms are provided below.
[0129] As used herein, the term "and / or" encompasses all combinations of the items linked by the term, should be read as if each combination was individually listed. For example, "A and / or B" covers A, B, and "A and B." For example, "A, B, and / or C" covers A, B, C, "A and B," "A and C," "B and C," and "A and B and C."
[0130] The term "comprising" as used herein to describe a protein or nucleic acid sequence means that the protein or nucleic acid can consist of the sequence, or can have additional amino acids or nucleotides at either or both ends of the protein or nucleic acid, but still have the activity described herein. Furthermore, it is clear to one skilled in the art that the methionine encoded by the start codon at the N-terminus of a polypeptide can in some instances (e.g., when expressed in certain expression systems) be retained, but does not materially affect the function of the polypeptide. Thus, the specification and claims herein, in describing specific polypeptide amino acid sequences, encompass sequences that include the methionine encoded by the start codon at the N-terminus, even though they can not include it, and correspondingly, the nucleotide sequences that encode them can include the start codon, and vice versa.
[0131] "Genome" as used herein encompasses not only chromosomal DNA present in the nucleus of a cell, but also organelle DNA present in subcellular components of a cell, such as mitochondria, plastids.
[0132] "Organism" includes any organism suitable for genome editing, preferably a eukaryote. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry such as chickens, ducks, geese; plants including monocots and dicots, for example rice, maize, wheat, sorghum, barley, soybean, peanut, Arabidopsis, etc.
[0133] "Foreign" means a sequence from a foreign species, or, if from the same species, which has been substantially modified by deliberate human intervention from its natural form.
[0134] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and are single- or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single letter designation: "A" for adenosine or deoxyadenosine (RNA or DNA, respectively), "C" for cytosine or deoxycytosine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "I" for inosine, and "X" for any nucleotide.
[0135] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" can also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.
[0136] As used herein, the term "amino acid" can include natural amino acids, non-natural amino acids, amino acid analogs, and all their D and L stereoisomers. Amino acids and abbreviations and English names in the present application are shown below:
[0137] Histidine (His, H); Serine (Ser, S); Glutamic acid (Glu, E); Glutamine (Gln, Q); Glycine (Gly, G); Threonine (Thr, T); Phenylalanine (Phe, F); Aspartic acid (Asp, D); Tyrosine (Tyr, Y); Leucine (Leu, L); Isoleucine (lie, I); Arginine (Arg, R); Alanine (Ala, A); Valine (Val, V); Tryptophan (Trp, W); Methionine (Met, M); Asparagine (Asn, N); Cysteine (Cys, C); Lysine (Lys, K); Proline (Pro, P).
[0138] As used herein, the term "wild type" refers to an object that can be found in nature. For example, a polypeptide or polynucleotide sequence that exists in an organism, can be isolated from a source in nature, and has not been intentionally modified by humans in a laboratory is naturally occurring. As used herein, "naturally occurring" and "wild type" are synonymous.
[0139] As used herein, the term "mutant" refers to a polynucleotide or polypeptide that comprises an alteration (i.e., a substitution, an insertion, and / or a deletion) at one or more (e.g., several) positions relative to a "wild type," or "compared" polynucleotide or polypeptide, wherein a substitution refers to replacing a nucleotide or amino acid occupying a position with a different nucleotide or amino acid. A deletion refers to removing a nucleotide or amino acid occupying a position. An insertion refers to adding a nucleotide or amino acid after nucleotides or amino acids that abut and immediately follow the nucleotide or amino acid occupying a position.
[0140] As used herein, an "amino acid" that is "mutated" includes a substitution, a duplication, a deletion, or an addition of one or more amino acids. In the present application, the term "mutation" refers to an alteration of an amino acid sequence. In a particular embodiment, the term "mutation" refers to a "substitution."
[0141] The term "corresponding to" or "relative to" as used herein has the meaning generally understood by one of ordinary skill in the art. In particular, "corresponding to" means that after homology or sequence identity alignment of two sequences, a position in one sequence corresponds to a specified position in the other sequence. Thus, for example, in the context of "the amino acid residue corresponding to position 150 of the amino acid sequence shown in SEQ ID NO: 1," if a 6xHis tag is added to one end of the amino acid sequence shown in SEQ ID NO: 1, then the corresponding position to position 150 of the amino acid sequence shown in SEQ ID NO: 1 in the resulting mutant can be position 156.
[0142] Sequence "identity" has the meaning commonly understood in the art and can be calculated using published techniques to determine the percent sequence identity between two nucleic acid or polypeptide molecules or regions. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule. (See, e.g., Computational Molecular Biology, Lesk, A.M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A.M., and Griffin, H.G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). Although there are a number of methods for measuring sequence identity between two polynucleotides or polypeptides, the term "identity" is well known to one of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).
[0143] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art and can generally be made without altering the biological activity of the resulting molecule. In general, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a polypeptide does not substantially alter biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0144] As used herein, "expression construct" or "construct" refers to a vector, such as a recombinant vector, suitable for expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to transcription of the nucleotide sequence (e.g., to produce mRNA or functional RNA) and / or translation of the RNA into a precursor or mature protein.
[0145] An "expression construct" of the application can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (e.g., mRNA).
[0146] An "expression construct" of the application can comprise regulatory sequences and a nucleotide sequence of interest of different origin, or regulatory sequences and a nucleotide sequence of interest of the same origin arranged in a manner different from that normally existing in nature.
[0147] As used herein, the term "recombinase" and the like refers to a site-specific enzyme that mediates recombination of DNA between recombinase recognition sequences. During the reaction, an attack on the DNA phosphate backbone is launched by a tyrosine or serine located in the catalytic active center of the recombinase, causing a DNA strand break. Recombinases can be divided into two different families: serine recombinases (e.g., resolvases and invertases) and tyrosine recombinases (e.g., integrases). Some examples of serine recombinases include, but are not limited to, Bxbl, Ilin, Gin, Tn3, beta-six, CinH, ParA, gamma delta, Bxbl, TP901, TGl, Rl, R2, R3, R4, R5, MRll, Al 18, U153, and gp29.
[0148] The term "recombination" refers to excision, integration, inversion, or exchange (e.g., translocation) of a DNA fragment between recombinase recognition sequences. In some embodiments, the recombinase recombination activity boost includes an increase in integration activity or inversion activity of the recombinase.
[0149] As used herein, the term "large serine recombinase," "LSR" is a classification of serine recombinases, which are integrases carried by bacteriophages that integrate large fragments of DNA sequences into the genome of a bacterium by recognizing specific sequences on the bacterial genome and the phage DNA fragment, without the need for any cellular cofactors and without the generation of harmful DNA double-strand breaks (DSBs). Many serine recombinases function to resolve transposition intermediates or to regulate gene expression by inverting regulatory sequences. During the catalytic reaction of a serine recombinase, a complex is formed from DNA and the recombinase, then a serine in the recombinase domain attacks the DNA phosphate backbone causing the DNA to form a nick with a 3'-OH double-strand break end and a 5'-phosphoserine covalently linked to the DNA, the complex is flipped, the double-stranded DNA is re-ligated, and a recombination product is formed. Most serine recombinases have a 150 amino acid catalytic domain at their amino terminus, followed by a small HTH (Helix-Turn-Helix)-DNA binding domain. Large serine recombinases have a similar amino-terminal catalytic domain, but have a larger carboxy-terminal region that varies in size from 300 residues in the bacteriophage R4 and A118 recombinases to 550 residues in the TnpX transposase, and are referred to as large serine recombinases because of the larger carboxy-terminal region.
[0150] As used herein, the terms "attP," "attB" are the phage attachment site (attP) and the bacterial attachment site (attB), respectively, and are collectively referred to herein as "recombinase recognition site sequences" or "recombination sites." The original function of integrases is to recombine short sequences of DNA between the phage attachment site (attP) and the bacterial attachment site (attB), so the recognition site for the target DNA in the gene editing process is called "attB" and the recognition site for the DNA sequence to be edited is called "attP."
[0151] As used herein, the terms "recombinase recognition site sequence," "recombination site," "recognition site," "RS site," "RS sequence," "recognition sequence," "RS" generally refer to a nucleic acid (e.g., DNA) sequence that is recognized (e.g., can be bound by) a recombinase polypeptide, which can be naturally occurring, or engineered or optimized.
[0152] II. Recombinases
[0153] In one aspect, the present application provides a recombinase that is capable of recombining DNA between a recombinase recognition site sequence (RS). In some embodiments, the recombinase is from a bacteriophage. In some embodiments, the recombinase is from a mycolicibacterium smegmatis bacteriophage. In some embodiments, the recombinase is a Bxbl recombinase.
[0154] In some embodiments, the recombinase comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to the amino acid sequence set forth in SEQ ID NO: 1.
[0155] In some embodiments, the amino acid sequence of the recombinase comprises one or more amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1, wherein the one or more amino acid substitutions comprise a substitution at position 2, 3, 4, 5, 6, 7, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 23, 24, 25, 26, 27, 28, 29, 31, 32, 34, 35, 36, 37, 38, 39, 40, 41, 42, 44, 45, 46, 49, 50, 51, 52, 53, 54, 55, 58, 59, 60, 62, 63, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 77, 78, 79, 83, 84, 86, 87, 88, 89, 90, 92, 93, 94, 95, 96, 97, 98, 100, 101, 102, 103, 104, 105, 106, 107, 108, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 122, 123, 124, 125, 126, 128, 129, 132, 134, 135, 137, 138, 139, 142, 143, 144, 145, 146, 147, 148, 149, 150, or 151 of the amino acid sequence set forth in SEQ ID NO: 1, and any combination thereof.
[0156] In some embodiments, the recombinase has increased recombination activity relative to a recombinase as set forth in SEQ ID NO: 1.
[0157] In some embodiments, the amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions selected from the following: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C. S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V 37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, E42L, E42R, D 45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, D51C, D51S, P52L, P5 2M, P52V, F53H, F53L, F53M, F53R, F53W, D54C, R55A, R55Q, R55W, R55Y, P5 9D, P59G, N60M, A62H, A62N, R63C, R63Q, A66W, F67S, E68D, E68P, E68Q, E6 8R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I7 5F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84 N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89 S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S , A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V , K101Y, K102C, K102H, L103M, L103Q, V104C, V104L, V105E, V105Q, S106N,A107V, T108A, T108G, T108Y, A110C, A110L, A110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, or F118S.
[0158] In some embodiments, the amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, and 59 amino acids. 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 1 16, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, or 319 amino acid substitutions.
[0159] In some embodiments, the amino acid sequence of the recombinase, relative to the amino acid sequence shown in SEQ ID NO:1, includes one or more amino acid substitutions selected from the following: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C. S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V 37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, D45N, D45S, A 49C, A49D, V50A, V50Y, D51G, D51K, D51V, P52L, P52M, P52V, F53H, F53L, F5 3M, F53R, F53W, R55A, R55Q, R55W, R55Y, P59D, P59G, N60M, A62H, A62N, R6 3C, R63Q, F67S, E68D, E68P, E68Q, E68R, E69M, E69T, E69Y, P71C, P71M, P7 1S, D73A, D73C, D73G, V74R, V74S, I75F, I75L, A77L, Y78F, Y78G, Y78I, Y7 8R, Y78T, Y78V, R79G, L83C, T84A, T84N, S86A, S86E, I87A, I87D, I87K, I87 S, I87T, I87V, H89F, H89G, H89Q, H89S, Q92C, Q92E, Q92G, Q92I, L93A, L93 F, L93I, V94G, H95W, W96L, W96Q, W96S, A97I, A97S, A97V, E98F, E98N, E98V , H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, L103M, L103Q , V104C, V105E, V105Q, S106N, A107V, T108A, T108G, T108Y, A110C, A110L,A110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T.
[0160] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0161] In some embodiments, the amino acid sequence of the recombinant enzyme has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, or 131 amino acid substitution relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0162] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0163] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitution selected from H146C, A144S, P71M, P71S, Q92C, D51S, E42L, E42R, D51C, E69M, D54C, A66W, V104L.
[0164] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitution selected from H146C, P71M, Q92C, E42L, E69M, D54C, A66W.
[0165] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and E69M relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0166] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0167] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and E42L relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0168] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and A66W relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0169] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and D54C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0170] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0171] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and P71M relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0172] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N, P71M, and any one or more amino acid substitutions selected from H146C, Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0173] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N, P71M, and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0174] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N, P71M, and Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0175] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions T84N and P71S relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0176] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N, P71S amino acid substitutions, and any one or more amino acid substitutions selected from H146C, Q92C, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0177] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N, P71S, and H146C amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0178] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N, P71S, and Q92C amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0179] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and E42R amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0180] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and V104L amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0181] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and D51C amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0182] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and A144S amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0183] In some embodiments, the amino acid sequence of the recombinant enzyme comprises T84N and D51S amino acid substitutions, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0184] In some embodiments, the amino acid sequence of the recombinant enzyme comprises I87T amino acid substitution, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0185] In some embodiments, the amino acid sequence of the recombinant enzyme comprises I87T amino acid substitution, and any one or more amino acid substitutions selected from Y78R, A144T, H146C, H146Q, F118M, M132L, A144G, relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0186] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and A144G relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0187] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0188] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and A144T relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0189] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0190] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and H146Q relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0191] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and F118M relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0192] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions I87T and M132L relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0193] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution A144S relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0194] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution A144S relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more of the amino acid substitutions selected from V94G, Y78R, H146C, I87T, Q92C.
[0195] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions A144S and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0196] In some embodiments, the amino acid sequence of the recombinant enzyme comprises the amino acid substitutions A144S and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0197] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0198] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S and I87T relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0199] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of A144S and Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0200] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0201] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A relative to the amino acid sequence set forth in SEQ ID NO: 1, and one or more amino acid substitutions selected from H146C, K101E.
[0202] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0203] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A and K101E relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0204] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0205] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1, and one or more amino acid substitutions selected from V94G, H146C.
[0206] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0207] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0208] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0209] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and any one or more amino acid substitutions selected from V94G, Y78R, H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0210] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0211] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0212] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L and H146C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0213] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0214] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and any one or more amino acid substitutions selected from P59D, Y78R, V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0215] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and P59D relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0216] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0217] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0218] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of E42W relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0219] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of E42W relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from T84N, P59D, V94G, K102H, Y78R.
[0220] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and T84N relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0221] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and P59D relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0222] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and V94G relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0223] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and K102H relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0224] In some embodiments, the amino acid sequence of the recombinant enzyme comprises amino acid substitutions of E42W and Y78R relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0225] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from A3G, V5A, I7T, I7V, I87T, I87V, T108A, V129A, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A41V, A62Y, P71S, P71I, H89Y, A119Y, F118I, A120R, F118S, or H146Y relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0226] In some embodiments, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions of R2Q, A3G, V129A, V76C, L4M, or I123T relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0227] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of R2Q, A3G, or L4M relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0228] In some embodiments, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of V76C relative to the amino acid sequence set forth in SEQ ID NO: 1.
[0229] In some embodiments, the recombinase recognition site sequence (RS) of the recombinant enzyme provided herein comprises an attP site (SEQ ID NO: 2), an attB site (SEQ ID NO: 3), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the nucleotide sequence set forth in SEQ ID NO: 2-3.
[0230] III. Genome editing system
[0231] In one aspect, the present disclosure provides a genome editing system, wherein the genome editing system comprises:
[0232] 1) a recombinase recognition site sequence (RS) comprising an attP or attB site (hereinafter also referred to as Component 1);
[0233] 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase (hereinafter also referred to as Component 2), wherein the recombinase comprises any one or more of the recombinases described above.
[0234] In some embodiments, the recombinase recognition site sequence (RS) is naturally occurring, or engineered or optimized.
[0235] In some embodiments, one or more recombinase recognition site sequence (RS) is inserted into a desired location in the genome by a recombinase recognition site integration unit.
[0236] A person skilled in the art can select a suitable gene editing tool to insert the RS into a desired location in the genome.
[0237] In some embodiments, the recombinase recognition site integration unit comprises: a sequence-specific DNA cleavage protein and a RS integration template, and / or an expression construct encoding the sequence-specific DNA cleavage protein and the RS integration template. In some embodiments, the recombinase recognition site integration unit specifically recognizes a desired location in a genome and cleaves DNA, followed by insertion of a recombinase recognition site sequence into the desired location in the genome by an endogenous or exogenous repair mechanism of the cell, mediated by the RS integration template.
[0238] In some embodiments, the RS integration template is a donor DNA template or a reverse transcription RNA template, wherein the donor DNA template comprises a desired recombinase recognition site sequence or a complement thereof, and the reverse transcription RNA template comprises a transcription sequence of a desired recombinase recognition site sequence or a complement thereof. In some embodiments, the RS integration template comprises a DNA synthetic template encoding a RS recognizable by a recombinase.
[0239] In some embodiments, the sequence-specific DNA cleavage protein is selected from one or more of a meganuclease (MGN), a zinc-finger nuclease (ZFN), a transcription-activator-like effector nuclease (TALEN), and a CRISPR effector protein.
[0240] In some embodiments, the recombinase recognition site integration unit comprises: a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or the functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.
[0241] In some embodiments, the CRISPR effector protein or the functional variant thereof comprises a CRISPR nuclease and a functional variant thereof.
[0242] As used herein, the term "CRISPR effector protein" generally refers to a nuclease present in a naturally occurring CRISPR system or a functional variant thereof. The term encompasses any effector protein based on a CRISPR system capable of achieving sequence-specific targeting within a cell.
[0243] As used herein, the term "CRISPR nuclease" can be derived from a Cas9 nuclease, including a Cas9 nuclease or a functional variant thereof. The Cas9 nuclease can be a Cas9 nuclease from different species, such as spCas9 from S. pyogenes or SaCas9 derived from S. aureus. "Cas9 nuclease" and "Cas9" are used interchangeably herein to refer to an RNA-guided nuclease including a Cas9 protein or a fragment thereof (e.g., a protein comprising the active DNA cleavage domain of Cas9 and / or the gRNA binding domain of Cas9). Cas9 is a component of the CRISPR / Cas (clustered regularly interspaced short palindromic repeats and associated systems) genome editing system that can target and cleave a DNA target sequence under the guidance of a guide RNA to form a DNA double-strand break (DSB).
[0244] A "CRISPR nuclease" can also be derived from a Cpf1 nuclease, including a Cpf1 nuclease or a functional variant thereof. The Cpf1 nuclease can be a Cpf1 nuclease from different species, such as Cpf1 nuclease from Francisella novicida U112, Acidaminococcus sp. BV3L6 and Lachnospiraceae bacterium ND2006.
[0245] As used herein, a "functional variant" with respect to a CRISPR nuclease means that it retains at least the ability of sequence-specific targeting mediated by a guide RNA. Preferably, the functional variant is a nuclease-inactive variant, i.e., it lacks the double-stranded nucleic acid cleavage activity. However, a CRISPR nuclease that lacks the double-stranded nucleic acid cleavage activity also encompasses a nickase, which forms a nick in a double-stranded nucleic acid molecule, but does not completely cut the double-stranded nucleic acid.
[0246] In some preferred embodiments of the present application, the CRISPR effector protein of the present application has nickase activity (or is referred to as a CRISPR nickase). In some embodiments, the functional variant recognizes a different PAM (protospacer adjacent motif) sequence relative to the wild-type nuclease.
[0247] As used herein, the term “CRISPR nickase” can be derived from a Cas9 of different species, e.g., from S. pyogenes Cas9 (SpCas9), or from S. aureus Cas9 (SaCas9). Simultaneous mutation of the HNH nuclease subdomain and the RuvC subdomain of Cas9, e.g., comprising mutations D10A and H840A, renders Cas9 nuclease-inactive, a nuclease-inactive Cas9 (dCas9). Mutation of only one of the subdomains can render Cas9 to have nickase activity, i.e., a Cas9 nickase (nCas9), e.g., nCas9 with only mutation D10A (or referred to as Cas9-D10A), and nCas9 with only mutation H840A (or referred to as Cas9-H840A). In some embodiments, the CRISPR nickase can be derived from a Cas12 of different species, e.g., Cas12a, Cas12b, Cas12i. In some embodiments, the CRISPR nickase is based on TraC protein that only retains DNA single-strand cleavage activity among transposon and CRISPR-Cas12 intermediate TraC effector protein (wherein TraC nuclease is disclosed in PCT / CN2023 / 097783 (publication number WO / 2023 / 232109), which is incorporated herein by reference).
[0248] In some embodiments, the CRISPR nuclease functional variant is a CRISPR nickase, e.g., Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase, and / or TraC nickase (see patents CN202310646033.2, CN202411698901.2).
[0249] In some embodiments, the recombinase recognition site integration unit further comprises a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the reverse transcriptase. In these embodiments, the recombinase recognition site integration unit can be based on prime editors and iterations thereof, which have demonstrated insertion, deletion, and base conversion with modest editing efficiency (as disclosed in WO 2020 / 191234 Al, incorporated by reference herein), as well as plant prime editors (PPEs), disclosed in Lin, Qiupeng et al. “Prime genome editing in rice and wheat.” Nature biotechnology vol. 38, 5 (2020): 582-585. doi:10.1038 / s41587-020-0455-x; Lin, Qiupeng et al. “High-efficiency prime editing with optimized, paired pegRNAs in plants.” Nature biotechnology vol. 39, 8 (2021): 923-927. doi:10.1038 / s41587-021-00868-w, and the like; dual-ePPEs, disclosed in Sun, Chao et al. “Precise integration of large DNA sequences in plant genomes using PrimeRoot editors.” Nature biotechnology vol. 42, 2 (2024): 316-327. doi:10.1038 / s41587-023-01769-w; TwinPEs, disclosed in Anzalone, Andrew V et al. “Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing.” Nature biotechnology vol. 40, 5 (2022): 731-740. doi:10.1038 / s41587-021-01133-w, and the like, incorporated by reference herein.
[0250] In some embodiments, the recombinase recognition site integration unit further comprises all the technologies in the prior art that can insert specific DNA sequences into DNA fragments at a specific site, including but not limited to: technologies based on CRISPR and transposon systems, such as homologous directed repair (HDR); INTEGRATE system, which is disclosed in Strecker, Jonathan et al. “RNA-guided DNA insertion with CRISPR-associated transposases.” Science (New York, N.Y.) vol. 365, 6448 (2019): 48-53. doi:10.1126 / science.aax9181, and the like; CRISPR-associated transposase (CAST) technology, which is disclosed in Lampe, George D et al. “Targeted DNA integration in human cells without double-strand breaks using CRISPR RNA-guided transposases.” bioRxiv: the preprint server for biology 2023.03.17.533036. 18 Mar. 2023, doi:10.1101 / 2023.03.17.533036. Preprint. and the like, which are incorporated herein by reference.
[0251] As used herein, the term “reverse transcriptase” is an enzyme that directs the synthesis of deoxyribonucleotide triphosphates into complementary DNA (cDNA) using RNA as a template.
[0252] In some embodiments, the reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii mature enzyme RT (Marathon RT).
[0253] In some embodiments, the guide RNA comprises a scaffold sequence, an RS integration template sequence, and a primer binding sequence, the RS integration template sequence being connected to the primer binding sequence.
[0254] As used herein, the terms "guide RNA" and "gRNA" are used interchangeably to refer to an RNA molecule capable of forming a complex with a CRISPR effector protein and capable of targeting the complex to a target sequence due to a certain identity with the target sequence. The guide RNA targets the target sequence through base pairing between the complementary strands of the target sequence. The guide RNA functions in conjunction with the CRISPR effector protein is collectively referred to as "scaffold sequence", for example, the scaffold sequence adopted by Cas9 nuclease or its functional variants is usually composed of crRNA and tracrRNA that form a complex in part complementarily. "Primer binding sequence" and "PBS" are used interchangeably to refer to a sequence of at least 10, at least 15 or at least 20 consecutive nucleotides of the guide RNA complementary to the target sequence. In some embodiments, the recombinase recognition site sequence is linked to the primer binding sequence. In some embodiments, the "guide RNA" is a single guide RNA (sgRNA), a guide RNA of prime editor (pegRNA). It is within the ability of those skilled in the art to design a suitable guide RNA based on the CRISPR nuclease used and the target sequence to be edited.
[0255] In some embodiments, the CRISPR nickase of the recombinase recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with the CRISPR nickase and targets a desired location in the genome, wherein the CRISPR nickase makes a cut in the strand of the genome, and the reverse transcriptase incorporates the RS integration template sequence in the guide RNA into the cut site, thereby inserting at least one RS recognizable by the recombinase at the desired location of the genome.
[0256] In some embodiments, the recombinase site (RS) is selected from an attP site, an attB site. In some embodiments, the attP site comprises a nucleotide sequence set forth in SEQ ID NO: 2, and the attB site comprises a nucleotide sequence set forth in SEQ ID NO: 3.
[0257] In some embodiments, a RS can be inserted into the genome or a fragment thereof of a cell using a nuclease, a gRNA, and / or an integrase. In some embodiments, the RS is carried on a guide RNA. The guide RNA can target any site known in the art. The complementary RS can be operably linked to a gene or nucleic acid sequence of interest of an exogenous DNA or RNA. In some embodiments, one RS is added to the target genome. In some embodiments, more than one RS is added to the target genome.
[0258] In some embodiments, the recombinase recognition site integration unit is not covalently linked to a recombinase.
[0259] In some embodiments, the recombinase recognition site integration unit is covalently linked to the recombinase.
[0260] In some embodiments, the genome editing system further comprises:
[0261] 3) a donor comprising an exogenous nucleotide sequence to be inserted into the genome (hereinafter also referred to as component 3).
[0262] In some embodiments, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer, such as at least 50 bp, at least 100 bp, at least 300 bp, at least 500 bp, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, at least 20 kb, at least 50 kb.
[0263] In some embodiments, the exogenous nucleotide sequence further comprises one or more RSs that can be recognized by the recombinase.
[0264] In some embodiments, the exogenous nucleotide sequence is flanked by one or two RSs.
[0265] IV. Fusion protein
[0266] In another aspect of the present application, a fusion protein / chimeric protein comprising the genome editing system of any one of the above is provided, wherein the sequence-specific DNA cleavage protein, the reverse transcriptase and the recombinase are directly linked or linked via a linker.
[0267] In some specific embodiments, the sequence-specific DNA cleavage protein, such as a CRISPR nickase, is linked at the C-terminus to the reverse transcriptase, which is linked via a linker to the recombinase.
[0268] As used herein, the term "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids, which has no secondary structure above. For example, the linker can be a flexible linker.
[0269] V. Polynucleotide
[0270] In one aspect, the present application provides a polynucleotide comprising a nucleotide sequence encoding the aforementioned recombinase, the aforementioned genome editing system, or the aforementioned fusion protein.
[0271] The polynucleotide of the present application can be in the form of DNA or RNA. The DNA form includes cDNA, genomic DNA, or artificially synthesized DNA. The DNA can be single-stranded or double-stranded. The DNA can be a coding strand or a non-coding strand. The RNA form includes mRNA.
[0272] The polynucleotide encoding the recombinase of the present application includes: a coding sequence encoding only the recombinase; a coding sequence of the recombinase and various additional coding sequences; a coding sequence of the recombinase (and optional additional coding sequences) and non-coding sequences.
[0273] Six, expression construct
[0274] In one aspect, the present application provides an expression construct comprising the aforementioned polynucleotide.
[0275] In some embodiments, the expression construct is selected from the group consisting of viral, bacterial, yeast, plant, mammalian cell expression constructs.
[0276] Seven, cell
[0277] In one aspect, the present application provides a cell comprising the aforementioned recombinase, the aforementioned genome editing system, the aforementioned fusion protein, the aforementioned polynucleotide, or the aforementioned expression construct.
[0278] Eight, kit
[0279] In one aspect, the present application provides a kit comprising the aforementioned recombinase, the aforementioned polynucleotide, the aforementioned genome editing system, the aforementioned fusion protein, the aforementioned expression construct, or the aforementioned cell.
[0280] Nine, method of genome editing
[0281] In another aspect of the present application, a method of genome editing in an organism or an organism cell is provided, selected from: i) introducing into the organism or the organism cell the aforementioned recombinase and a donor comprising an exogenous nucleotide sequence to be inserted into the genome, wherein the target genome of the organism or the organism cell contains a recombinase recognition site sequence (RS) corresponding to the recombinase; or ii) introducing into the organism or the organism cell the aforementioned genome editing system. In some embodiments, the method of genome editing can result in one or more nucleotide substitutions, or one or more nucleotide insertions. In some embodiments, the length of the substitution of the plurality of nucleotides is at least 2 bp, at least 10 bp, at least 100 bp, at least 1 kbp, at least 10 kbp, at least 20 kbp, at least 50 kbp.
[0282] In some embodiments, components 1), 2) and optional component 3) of the genome editing system are introduced into the organism or organism cell simultaneously.
[0283] In some embodiments, components 1), 2) or optional component 3) of the genome editing system are introduced into the organism or organism cell separately in steps.
[0284] In some embodiments, component 2) of the genome editing system is introduced into the organism or organism cell separately.
[0285] The skilled person can select the RS insertion into the genome and / or the foreign nucleotide sequence, and the direction of insertion into the genome or foreign nucleotide sequence, according to the purpose of DNA recombination, such as deletion, inversion, integration, etc.
[0286] In some embodiments, component 1) is inserted into the donor construct of the genome or foreign nucleotide sequence in the same direction or in the opposite direction.
[0287] In some embodiments, the method comprises recombining DNA of the genome of the organism or organism cell.
[0288] In some specific embodiments, the recombining DNA of the genome of the organism or organism cell comprises deleting DNA, inverting DNA in the genome, and / or integrating foreign DNA into the genome.
[0289] In the method of the present application, the megaserrnase or genome editing system can be introduced into the cell by various methods well known to the skilled person.
[0290] In some embodiments, the method of introducing the recombinase or genome editing system into the organism or organism cell is selected from the group consisting of: calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), biolistics, N-acetylgalactosamine (GalNAc)-mediated, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.
[0291] In some embodiments, the organism or organism cell is from a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; a plant, preferably a crop plant, such as wheat, rice, corn, soybean, sunflower, leafy greens, lettuce, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, kiwi, tobacco, cassava, potato.
[0292] In the present application, the target nucleic acid region to be edited can be located at any position of the genome, for example, at a safe harbor of the genome, within a functional gene such as a protein-coding gene, or for example, can be located at a gene expression regulatory region such as a promoter region or an enhancer region, so as to achieve modification of the function of the gene or modification of the gene expression. In some embodiments, the desired nucleotide sequence substitution results in a desired modification of the gene function or modification of the gene expression.
[0293] EMBODIMENTS
[0294] The embodiments of the present application will be described in detail below with reference to the examples, but those skilled in the art will understand that the following examples are only used to illustrate the present application and should not be regarded as limiting the scope of the present application. If no specific conditions are specified in the examples, the conventional conditions or the conditions recommended by the manufacturer are used. If no manufacturer of the reagent or instrument used is specified, it is a conventional product that can be obtained by purchase on the market.
[0295] Example 1 Screening of Bxbl single mutants with improved recombination activity
[0296] The inventors designed two sets of parallel screening systems to screen the screening library of Bxbl single mutants constructed by single-point saturation mutation, and carried out multi-dimensional coverage based on the screening strategy of functional complementation to avoid missing screening.
[0297] Firstly, the inventors designed an evolutionary screening system 1 in tobacco leaves (Figure 1), which was initiated by CaMV35S promoter to express Bxb1, terminated by NOS terminator to terminate expression, used NNK degenerate primer to introduce single-point saturated mutation into the region of 150 amino acids at the N terminus of Bxb1 to construct a screening library, divided the N terminus of Bxb1 into 15 fragments on average, each fragment contained single-point saturated mutation of 10 amino acids, the directions of the two recognition sites (RS) of attB and attP were opposite, a fragment of about 1 kb was reversely inserted between the two sites, after the construction of 15 libraries, Agrobacterium was transformed respectively to prepare Agrobacterium library. Then Agrobacterium was used to infect tobacco leaves, and tobacco leaves were taken after two days, and genome was extracted. PCR was performed using the genome as the template, and two pairs of primers were used for PCR, the first pair of primers was F1+R1 (SEQ ID NO: 11, SEQ ID NO: 13), and product 1 was obtained as a background control; the second pair of primers was F2+R2 (SEQ ID NO: 12, SEQ ID NO: 14), which was used to amplify the fragment containing the Bxb1 variant with sufficient activity which occurred inversion, as product 2. Among them, the higher the proportion in product 2, the higher the activity of the corresponding Bxb1 variant. Then product 1 and 2 were used as templates to continue PCR reaction, and the same pair of primers was used to enrich the short fragment containing Bxb1 variant in product 1 (using primer F1+R1) and product 2 (using primer F2+R2) respectively, and then barcoding was performed for second-generation sequencing, the proportion of different variants in product 1 and product 2 was calculated, and the enrichment degree of active Bxb1 variant in product 2 relative to product 1 was further calculated, and the higher the enrichment degree, the higher the activity of the Bxb1 variant. Each library was biologically repeated three times, and the Bxb1 variant which was enriched in three biological repeats was screened, and the enrichment degree was the average value of three biological repeats.
[0298] The screening results are shown in Table 1, which lists the mutants with an average enrichment degree of more than 10%, proving that these screened mutants have recombinase activity and can be used as potential Bxb1 active mutants.
[0299] Table 1 Candidate Bxb1 mutants with improved recombination activity
[0300] Secondly, the inventors designed an evolution screening system 2 (a in FIG. 2) in tobacco leaves, which consists of two vectors. The first binary vector is a receptor vector (Vector 1, SEQ ID NO: 24), which expresses a truncated firefly luciferase nF-LUC under the control of the cassava vein mosaic virus CSVMV promoter (P-CSVMV), and is terminated by a tobacco mosaic virus-derived CaMV 35S terminator (T-CaMV) at the 3' end, and is linked to an intron sequence and the recognition sequence attB of Bxb1. The second binary vector is a donor vector (Vector 2), which contains a sea cucumber luciferase (R-LUC) expression frame; another recognition sequence attP of Bxb1, an intron and the C-terminal of firefly luciferase (cF-LUC); and a Bxb1 expression frame.
[0301] Firstly, the evolution screening system was detected in tobacco leaves using wild-type Bxb1 (WT_Bxb1) and inactivated Bxb1 (dead_Bxb1) to determine whether the evolution screening system works in tobacco leaves. Vector 1 and Vector 2 (WT_Bxb1 or dead_Bxb1) were transformed into Agrobacterium, and then co-infected tobacco leaves. The sequence of Vector 1 is shown in SEQ ID NO: 24, the sequence of Vector 2 (WT_Bxb1) is shown in SEQ ID NO: 25, and the sequence of Vector 2 (dead_Bxb1) is shown in a in FIG. 2 and SEQ ID NO: 25, in which the WT_Bxb1 sequence in the sequence is replaced by dead_Bxb1. After 48 hours, the infected leaves were taken to detect the luminescence signal, and the results are shown in b in FIG. 2. The leaves expressing WT_Bxb1 detected obvious firefly luciferase signals, while the leaves expressing dead_Bxb1 had almost no signal, indicating that the evolution screening system 2 can be used to screen Bxb1 variants with higher integration activity in tobacco leaves.
[0302] The single point saturation mutation library was constructed by using NNK degenerate primers in the catalytic region of Bxb1 (containing 150 amino acids), and the N-terminal of Bxb1 was divided into 15 fragments on average, each containing a single point saturation mutation of 10 amino acids. After the construction of the 15 libraries, Agrobacterium was transformed, and an Agrobacterium library was prepared. Then the Agrobacterium was used to infect tobacco leaves, and the tobacco leaves were taken after two days. Genomic DNA was extracted from the tobacco leaves. PCR was performed using genomic DNA as a template. Two pairs of primers were used for PCR. The first pair of primers was P1+P3 (SEQ ID NO: 26, SEQ ID NO: 28), and the product 1 was obtained as a background control. The second pair of primers was P2+P3 (SEQ ID NO: 27, SEQ ID NO: 28), and the product 2 was obtained by amplifying the fragment containing the Bxb1 variant with sufficient activity. The Bxb1 variant with high activity theoretically has a high proportion in product 2. Then, product 1 and product 2 were used as templates for further PCR, and the same pair of primers was used to enrich the short fragment containing the Bxb1 variant in product 1 (using primers F1+R1) and product 2 (using primers F2+R2), and then barcode labeling was performed for next-generation sequencing. The proportion of different variants in product 1 and product 2 was calculated, and the enrichment degree of the active Bxb1 variant in product 2 relative to product 1 was further calculated. The higher the enrichment degree, the higher the activity of the Bxb1 variant. Each library was biologically repeated three times, and the Bxb1 variant that was enriched in three biological repeats was selected, and the enrichment degree was the average of three biological repeats.
[0303] Meanwhile, the evolution screening system plasmid was transformed into rice and corn protoplasts using the PEG (polyethylene glycol) method. The protoplasts were collected 48 hours after transformation, and the genomic DNA was extracted. The samples were prepared using the above primers and methods, and next-generation sequencing analysis was performed. Three biological repeats were performed.
[0304] The screening results are shown in Table 2. The mutants that were enriched in three biological repeats and at least in two species were used to prove that these selected mutants have recombinase activity and can be used as potential Bxb1 high-activity mutants.
[0305] Table 2: Candidate Bxb1 single mutants with improved integration activity
[0306] In summary, the screening results of the two parallel screening systems are summarized, and the final screened candidate recombinant activity improved Bxbl mutant sites are: R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C, S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, D45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, P52L, P52M, P52V, F53H, F53L, F53M, F53R, F53W, R55A, R55Q, R55W, R55Y, P59D, P59G, N60M, A62H, A62N, R63C, R63Q, F67S, E68D, E68P, E68Q, E68R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I75F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S, A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, L103M, L103Q, V104C, V105E, V105Q, S106N, A107V, T108A, T108G, T108Y, A110C, A110L, A110M, A110Q,A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, F118S.
[0307] Example 2: Recombination activity verification of Bxbl single mutants
[0308] In molecular biology, Bxb1 recombination activity can be divided into three types according to the spatial arrangement of its recognition sites (attB / attP) and the reaction results: integration, deletion and inversion. The integration type recombination refers to the specific recombination between the attP site on the foreign DNA and the host genome attB site (same direction arrangement) mediated by Bxb1 recombinase, which can insert the foreign DNA into the host genome at a specific site and generate a new hybrid site attL and attR. The deletion type recombination refers to the recognition of two same direction arranged target sites (attB / attP) on the same DNA molecule by the enzyme, which can loop out the middle fragment and generate hybrid sites attL and attR, respectively. The inversion type recombination refers to the catalysis of the 180° inversion of the middle DNA fragment with the same direction of the target site (attB and reverse attP, etc.) on the same DNA molecule. The above three types of activity promotion can be counted as the promotion of the recombination activity of the recombinase. For this purpose, the inventors constructed two sets of verification systems to verify the recombination editing activity of the candidate Bxb1 single mutant sites in Example 1.
[0309] Firstly, the inventors constructed a recombinase activity verification system 1 to characterize the recombination activity. As shown in FIG. 3, the tobacco mosaic virus CaMV 35S promoter (P-CaMV 35S) initiates the expression of the recombinase (e.g. Bxb1 or its variant), the T-NOS terminator terminates the expression of the recombinase (Bxb1 or its variant), the attB and attP sequences are arranged in opposite directions, the reverse CaMV 35S promoter sequence (P-CaMV 35S) is inserted in the middle to initiate the expression of firefly luciferase, a segment from the Agrobacterium tumefaciens 3' untranslated region sequence (T-ORF1 3'UTR) terminates the expression of firefly luciferase, the cassava vein mosaic virus CSVMV promoter (P-CSVMV) initiates the expression of Renilla luciferase, the CaMV 35S terminator (T-CaMV 35S) terminates the expression of Renilla luciferase, only the active recombinase (Bxb1 or its variant) can initiate the expression of firefly luciferase after the sequence of CaMV 35S promoter in the middle of the att site is inverted, the activity of firefly luciferase is detected by a microplate reader, representing the activity of Bxb1 and its variants, and the constant expression of Renilla luciferase is used as an internal reference to correct the expression of firefly luciferase. The expression construct of the recombinase activity quantitative system 1 is shown in FIG. 3. The luciferase detection kit used is Transdetect Double-Luciferase Reporter Assay Kit (Transgen, FR201).
[0310] The data of three biological repeats of fold activity improvement in tobacco leaf relative to wild type Bxbl (Bxbl-WT; SEQ ID NO: 1) were averaged, the activity results were normalized, Bxbl-WT activity as 1, Bxbl mutant activity as fold of 1 were compared. The results are shown in Table 3, in which are preferred Bxbl single mutants, whose editing activity is 1.01-8.86 fold of Bxbl-WT, indicating that the mutation of these preferred sites has significantly improved the recombination activity of Bxbl recombinase.
[0311] Table 3 Verification results of verification system 1 of recombinase activity
[0312] Therefore, by the recombination enzyme activity verification system 1, the inventors verified a batch of single mutation sites with improved recombination activity, which are: R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q.
[0313] Secondly, the inventors constructed a recombination enzyme activity verification system 2 for supplementary verification of editing activity. A recognition site of the recombination enzyme was inserted into a safe harbor site of the genome by using a guide editing tool (Prime Editor, PE) and double pegRNA, and the method principle of inserting the RS is shown in FIG. 4, and another recognition site of the recombination enzyme was placed near the donor large fragment DNA. When the recombination enzyme is expressed, it recognizes the site inserted on the genome and the site on the donor respectively, forms a tetramer, activates the recombination enzyme activity, and integrates the donor large fragment DNA into the genome. In order to facilitate detection of integration efficiency, a R primer on the genome was pre-set on the donor large fragment DNA vector.
[0314] The endogenous recombination activity of the candidate Bxbl mutants was verified in rice and human cells, respectively. The expression construct of the recombinase is shown in FIG. 5 (a for rice cells, b for human cells). The selected site in rice is a safe harbor site GSH1 (kitaake, Chr1:7660637-7661671) on the genome, the integration vector 1 is an 8.5 kb donor vector (SEQ ID NO: 19), the PE tool used is ePPE (SEQ ID NO: 16), and the double pegRNA is shown in SEQ ID NO: 17 and SEQ ID NO: 18; the integration vector 2 is a 5.7 kb donor vector (SEQ ID NO: 35), the PE tool used is ePPE (SEQ ID NO: 16), and the double pegRNA is shown in SEQ ID NO: 36 and SEQ ID NO: 37; the TRAC site is selected in human cells HEK293T, and the integration activity of the variant on a 5.6 kb donor vector (SEQ ID NO: 23) is verified, the PE tool used is PE2 (SEQ ID NO: 20), and the double pegRNA is shown in SEQ ID NO: 21 and SEQ ID NO: 22.
[0315] The activity detection method is shown in FIG. 6. A pair of F and R primers are used to amplify the genomic DNA. If recombination occurs, three sequences can be amplified, one is the wild-type sequence (genome amplicon), the second is the sequence after PE work (attB installed amplicon), and the third is the sequence after the integration of the donor into the genome (insertion of donor). After second-generation sequencing, the number of three kinds of amplicons is counted by using the specific marker sequences of the three sequences, wherein the reads number of the second and third sequences / total reads number = PE efficiency, and the reads number of the third sequence / (reads number of the second and third sequences) = recombination efficiency of the recombinase. The F and R primers used for amplification in rice cells are SEQ ID NO: 33 and SEQ ID NO: 34, respectively; the F and R primers used for amplification in human cells are SEQ ID NO: 30 and SEQ ID NO: 31, respectively.
[0316] Table 4: Correspondence between variant name and mutation type
[0317] The recombination activity detection results of the recombination enzyme activity verification system 2 are shown in FIGS. 7-9, wherein PE represents the efficiency of inserting the recognition site of Bxb1 at the genomic safe harbor site, and IN represents the recombination efficiency of the Bxb1 recombinase. The correspondence between the variant name and the mutation type is shown in Table 4.
[0318] The recombination activity verification results of integrating vector 1 in rice cells are shown in FIG. 7. It is found through verification that the recombination efficiency of four variants V27, V39, V53 and V54 is obviously improved compared with the wild type Bxb1, and the corresponding mutation types are A3G, L4M, V129A and I123T. The recombination activity of the above-mentioned variants is 1.3-1.7 times that of WT_Bxb1.
[0319] The recombination activity verification results of integrating vector 2 in rice cells are shown in FIG. 8. V27 and V39 still have high recombination efficiency, and the recombination efficiency of V1 variant is 1.69 times that of the wild type Bxb1, even slightly higher than that of V27 and V39, which proves that R2Q is also an ideal mutation type for improving efficiency.
[0320] The recombination activity verification results in human cells are shown in FIG. 9. Among them, hV27 and hV39 have a slight improvement compared with WT Bxb1, and the corresponding mutation types are A3G and L4M; the hV30 variant has a significant improvement compared with WT Bxb1, and the corresponding mutation type is V76C, and the recombination activity reaches 1.9 times that of WT_Bxb1.
[0321] Therefore, through the recombination enzyme activity verification system 2, the inventors verified a batch of single mutation sites with improved recombination activity: R2Q, A3G, L4M, V129A, I123T and V76C.
[0322] Based on the above, the single mutation sites with improved recombination activity verified by the two sets of verification systems were summarized, and the Bxbl mutation sites with significantly improved recombination activity were finally obtained as follows: R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, R2Q, A3G, L4M, V129A, I123T and V76C.
[0323] Experimental Example 3: Verification of the activity of Bxbl multi-mutants with improved recombination activity in tobacco
[0324] A series of Bxb1 mutant libraries with multiple mutation sites were obtained by arranging and combining the Bxb1 single mutation sites with high activity in Experimental Example 1 and Experimental Example 2. The activity of the Bxb1 mutant library with multiple mutation sites was identified by using the recombinase activity verification system 1 in Experimental Example 1. The verification method is the same as that in Experimental Example 2. The average value of three biological repeat data of the activity improvement multiple of WT_Bxb1 obtained in tobacco leaves was taken, and the activity results were normalized, with the activity of WT_Bxb1 as 1, and the activity of the Bxb1 mutant as 1 multiple for comparison. Table 5 lists the combinations of mutation sites and activity verification results, wherein the efficiency of the double mutant is improved by 1.72-5.1, and the efficiency of the triple mutant is improved by 3.35-4.48. In these multiple mutants with improved recombination activity, the multiple recombination activity of the mutant with T84N, L93A, I87T or A144S as the reference mutant is more significant, as shown in Figure 10, especially T84N, which has a significant improvement in the single mutation activity of T84N compared with the double mutation combination of H146C, P71M, E42L, E69M, D54C or A66W, and the activity is more significantly improved compared with WT_Bxb1.
[0325] In summary, the above results all show that the double mutant or triple mutant of Bxb1 has the effect of further improving the recombination activity.
[0326] Table 5: Editing of Bxb1 N mutants with improved activity and activity verification results
[0327] Experimental Example 4: Activity verification of Bxb1 multiple mutants with improved recombination activity in rice protoplasts
[0328] The recombinase activity verification system 1 was introduced into rice protoplast cells, and the results showed that the Bxb1 multiple mutant recombinase with improved activity could also occur in rice cells. As shown in Figure 11, the activity of WT_Bxb1 was taken as 1, and the recombination activity of the Bxb1 mutant was taken as 1 multiple for comparison, and the results showed that in rice cells, the double mutant with T84N as the reference mutant and Q92C, H146C, P71M, E42L, E69M or A66W, and the double mutant with I87T as the reference mutant and A144G had significantly improved recombination activity, wherein the combination activity of E69M+T84N was the highest, and the improvement reached 11.69 times of WT_Bxb1.
[0329] It should be noted that although the technical solutions of the present application are introduced by specific examples, those skilled in the art can understand that the present application should not be limited thereto.
[0330] Embodiments of the application have been described above, with the understanding that they are exemplary only and are not restrictive in nature. Many modifications and variations of the described embodiments are possible in light of the above teachings. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the embodiments disclosed herein, which is defined solely by the appended claims. The use of the terms "including," "containing," or "comprising" and variations thereof herein is meant to encompass the items listed thereafter, and any subsequent conjugate, associate, artificial, or natural, compound, element, component, or agent, ingredient, or additive of like kind, without limitation.
[0331] Sequences referred to in the specification:
[0332] >SEQ ID NO: 1 BxBl
[0333] >SEQ ID NO: 2 attP
[0334] wherein: W is A (typically for integration) or T (typically for inversion).
[0335] >SEQ ID NO: 3 attB
[0336] wherein: W is A (typically for integration) or T (typically for inversion).
[0337] >SEQ ID NO: 4 CaMV 35S
[0338] >SEQ ID NO: 5 T-NOS
[0339] >SEQ ID NO: 6 Firefly luciferase
[0340] >SEQ ID NO: 7 T-At ORF1 3'UTR
[0341] >SEQ ID NO: 8 P-CSVMV
[0342] >SEQ ID NO: 9 Renilla luciferase
[0343] >SEQ ID NO: 10 T-CaMV35S
[0344] >SEQ ID NO: 11 Fl
[0345] >SEQ ID NO: 12 F2
[0346] >SEQ ID NO: 13 R1
[0347] >SEQ ID NO: 14 R2
[0348] >SEQ ID NO: 15 dead_Bxbl
[0349] >SEQ ID NO: 16 ePPE vector, rice
[0350] >SEQ ID NO: 17 epegRNA1, rice
[0351] >SEQ ID NO: 18 epegRNA2, rice
[0352] >SEQ ID NO: 19 8.5 kb donor
[0353] >SEQ ID NO: 20 PE2 vector, human cells
[0354] >SEQ ID NO: 21 pegRNA1, human cells
[0355] >SEQ ID NO: 22 pegRNA2, human cells
[0356] >SEQ ID NO: 23 5.6 kb donor
[0357] >SEQ ID NO: 24 Vector1
[0358] >SEQ ID NO: 25 Vector2 (WT_Bxbl)
[0359] >SEQ ID NO: 26 Figure 2P1
[0360] >SEQ ID NO: 27 Figure 2P2
[0361] >SEQ ID NO: 28 Figure 2P3
[0362] >SEQ ID NO: 29 WT_Bxbl expression vector for use in human cells
[0363] >SEQ ID NO: 30 Specific binding sequence of F primer for use in human cells to detect efficiency
[0364] >SEQ ID NO: 31 R primer sequence inserted on donor vector for use in human cells
[0365] >SEQ ID NO: 32 WT_Bxbl expression vector for use in rice
[0366] >SEQ ID NO: 33 Specific binding sequence of F primer for use in rice to detect efficiency
[0367] >SEQ ID NO: 34 R primer sequence inserted on donor vector for use in rice
[0368] >SEQ ID NO: 35 5.7 kb donor for use in rice
[0369] >SEQ ID NO: 36 epegRNA3 for use in rice
[0370] >SEQ ID NO: 37 epegRNA4 for use in rice
Claims
1. A recombinant enzyme, wherein, the amino acid sequence set forth in SEQ ID NO: 1, wherein the one or more amino acid substitutions comprise a substitution at position 2, 3, 4, 5, 6, 7, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 23, 24, 25, 26, 27, 28, 29, 31, 32, 34, 35, 36, 37, 38, 39, 40, 41, 42, 44, 45, 46, 49, 50, 51, 52, 53, 54, 55, 58, 59, 60, 62, 63, 66, 67, 68, 69, 71, 72, 73, 74, 75, 76, 77, 78, 79, 83, 84, 86, 87, 88, 89, 90, 92, 93, 94, 95, 96, 97, 98, 100, 101, 102, 103, 104, 105, 106, 107, 108, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 122, 123, 124, 125, 126, 128, 129, 132, 134, 135, 137, 138, 139, 142, 143, 144, 145, 146, 147, 148, 149, 150, or 151 of the amino acid sequence set forth in SEQ ID NO: 1, or any combination thereof.
2. The recombinant enzyme of claim 1, wherein, the recombinase has increased recombinase activity relative to the recombinase as set forth in SEQ ID NO:
1.
3. The recombinant enzyme of claim 1 or 2, wherein, The amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1 selected from the group consisting of R2A, R2C, R2D, R2H, R2S, R2T, R2W, A3G, A3I, A3P, L4Q, L4V, V5A, V5F, V5G, V5L, V5M, V5S, V6C, V6H, V6K, I7L, I7T, I7V, R11K, R11Q, V12L, T13H, D14N, D14P, D14Q, A15T, T16S, P19F, P19L, P19M, E20A, E20G, E20P, E20V, R21A, R21N, R21V, R21Y, L23G, L23Y, E24C, S25D, S25G, S25N, C26N, Q28G, Q28L, Q28V, Q28W, L29S, L29T, L29W, A31Q, Q32C, Q32R, G34C, G34E, G34S, G34W, W35C, W35G, W35L, W35Y, D36S, V37L, V37M, V38S, G39D, V40W, A41G, A41S, A41V, E42C, E42S, E42W, E42L, E42R, D45N, D45S, A49C, A49D, V50A, V50Y, D51G, D51K, D51V, D51C, D51S, P52L, P52M, P52V, F53H, F53L, F53M, F53R, F53W, D54C, R55A, R55Q, R55W, R55Y, P59D, P59G, N60M, A62H, A62N, R63C, R63Q, A66W, F67S, E68D, E68P, E68Q, E68R, E69M, E69T, E69Y, P71C, P71M, P71S, D73A, D73C, D73G, V74R, V74S, I75F, I75L, A77L, Y78F, Y78G, Y78I, Y78R, Y78T, Y78V, R79G, L83C, T84A, T84N, S86A, S86E, I87A, I87D, I87K, I87S, I87T, I87V, H89F, H89G, H89Q, H89S, Q92C, Q92E, Q92G, Q92I, L93A, L93F, L93I, V94G, H95W, W96L, W96Q, W96S, A97I, A97S, A97V, E98F, E98N, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, L103M, L103Q, V104C, V104L, V105E, V105Q, S106N, A107V, T108A,T108G, T108Y, A110C, A110L, A110M, A110Q, A110R, A110S, H111K, F112V, D113C, T114C, T114G, T114L, T114M, T114N, T114R, T114S, T115A, T115C, T115G, T115H, T115K, T115Q, T115V, T116K, P117A, F118M, F118W, A119G, A119S, A120E, V122G, V122P, V122S, V122T, I123A, A124S, L125C, L125I, L125S, M126C, M126T, V129A, M132A, M132L, M132Q, M132S, E135A, E135I, E135T, E135V, I137N, K138A, K138Q, E139H, E139N, E139V, E139Y, R142N, S143F, S143M, S143Y, A144G, A144S, A144T, A145Q, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147H, F147Y, N148A, N148H, N148Q, N148T, R150F, R150I, R150L, R150M, R150T, A151G, A151K, A151L, A151S, A151T, A3V, V46A, L4F, V6A, A3T, R58G, V76A, V76C, R88G, V105A, V105C, L134F, I137T, L134I, I137V, L4M, R2L, R21F, L44I, R2Q, L4P, T16A, S18P, Q27R, E42G, L44P, N60K, F72L, I75T, Y78H, F72S, L90R, H111Y, S106P, T128A, I123T, I149T, A62Y, P71I, H89Y, A119Y, F118I, A120R, or F118S.
4. The recombinant enzyme according to any one of claims 1 to 3, wherein, the amino acid sequence of the recombinant enzyme comprises one or more amino acid substitutions selected from R2A, R2D, R2T, A3I, V5L, V5S, V6C, V12L, D14Q, A15T, T16S, S25D, S25G, Q28G, Q28W, L29W, Q32C, Q32R, G34E, G34S, W35C, W35L, W35Y, V37L, A41G, A41S, A49C, A49D, V50A, D51G, D51K, D51V, A62H, A62N, R63C, R63Q, F67S, E68R, E69M, P71C, P71M, P71S, D73A, D73C, D73G, V74S, Y78R, Y78T, L83C, T84N, S86E, I87A, I87D, I87K, I87T, I87V, H89F, Q92C, Q92E, Q92G, L93F, H95W, W96Q, E98V, H100S, H100T, K101E, K101L, K101V, K101Y, K102C, K102H, T108Y, F112V, T114L, T114M, T114S, T115A, T115Q, L125I, L125S, M132A, M132L, M132S, E135A, E135V, K138Q, S143M, S143Y, A144G, A144S, A144T, A145S, H146C, H146E, H146Q, H146R, H146S, H146Y, F147Y, N148H, R150F, R150I, R150T, A151G, A151K, A151T, E20P, E20G, R21V, G34W, D36S, G39D, V40W, E42W, P52L, F53H, F53R, P59D, Q92I, L93A, L93I, E98F, V105Q, A110C, A110R, A110M, A110Q, A110L, E139Y, A145Q, R2Q, A3G, L4M, V129A, I123T, or V76C relative to the amino acid sequence set forth in SEQ ID NO:
1.
5. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises the amino acid substitution T84N relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more of the amino acid substitutions selected from H146C, A144S, P71M, P71S, Q92C, D51S, E42L, E42R, D51C, E69M, D54C, A66W, V104L.
6. The recombinant enzyme of claim 5, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, P71M, Q92C, E42L, E69M, D54C, A66W.
7. The recombinant enzyme of claim 5, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N and P71M relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, Q92C.
8. The recombinant enzyme of claim 5, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N and P71S relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, Q92C.
9. The recombinant enzyme of any one of claims 5-8, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of T84N relative to the amino acid sequence set forth in SEQ ID NO: 1, and any one or more amino acid substitutions selected from the group consisting of H146C, P71M, Q92C, E42L, E69M, D54C, A66W. (a) T84N; (b) T84N and E69M; (c) T84N and Q92C; (d) T84N and E42L; (e) T84N and A66W; (f) T84N and D54C; (g) T84N and H146C; (h) T84N and P71M; (i) T84N and P71S; (j) T84N and E42R; (k) T84N and V104L; (l) T84N and D51C; (m) T84N and A144S; (n) T84N and D51S; (o) T84N, P71M and H146C; (p) T84N, P71M and Q92C; (q) T84N, P71S and H146C; (r) T84N, P71S and Q92C.
10. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of I87T relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, any one or more amino acid substitutions selected from the group consisting of Y78R, A144T, H146C, H146Q, F118M, M132L, A144G; Preferably, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of I87T relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, any one or more amino acid substitutions selected from the group consisting of Y78R, A144T, H146C, H146Q, F118M, M132L, A144G; the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of I87T relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, any one or more amino acid substitutions selected from the group consisting of Y78R, A144T, H146C, H146Q, F118M, M132L, A144G; the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of I87T relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, any one or more amino acid substitutions selected from the group consisting of Y78R, A144T, H146C, H146Q, F118M, M132L, A144G; 11. The recombinant enzyme of any one of claims 1-4, wherein, Preferably, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a2) to (f2) amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1: (a2) A144S; (b2) A144S and V94G; (c2) A144S and Y78R; (d2) A144S and H146C; (e2) A144S and I87T; (f2) A144S and Q92C.
12. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of L93A relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from H146C, K101E; Preferably, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a3) to (c3) amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1: (a3) L93A; (b3) L93A and H146C; (c3) L93A and K101E.
13. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of Q92C relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, H146C; Preferably, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a4) to (c4) amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1: (a4) Q92C; (b4) Q92C and V94G; (c4) Q92C and H146C.
14. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of M132L relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from V94G, Y78R, H146C; Preferably, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a5) to (d5) amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1: (a5) M132L; (b5) M132L and V94G; (c5) M132L and Y78R; (d5) M132L and H146C.
15. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of N148A relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from P59D, Y78R, V94G; Preferably, the amino acid sequence of the recombinant enzyme comprises at least one of the following (a6) to (d6) amino acid substitutions relative to the amino acid sequence set forth in SEQ ID NO: 1: (a6) N148A; (b6) N148A and P59D; (c6) N148A and Y78R; (d6) N148A and V94G.
16. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of E42W relative to the amino acid sequence set forth in SEQ ID NO: 1; and, optionally, comprises any one or more amino acid substitutions selected from T84N, P59D, V94G, K102H, Y78R; Preferably, the amino acid sequence of the recombinant enzyme comprises at least one amino acid substitution selected from (a7)-(f7) relative to the amino acid sequence set forth in SEQ ID NO: 1: (a7) E42W; (b7) E42W and T84N; (c7) E42W and P59D; (d7) E42W and V94G; (e7) E42W and K102H; (f7) E42W and Y78R.
17. The recombinant enzyme of any one of claims 1-4, wherein, the amino acid sequence of the recombinant enzyme comprises any one or more amino acid substitutions selected from A3G, V129A, V76C, L4M, or I123T relative to the amino acid sequence set forth in SEQ ID NO:
1.
18. The recombinant enzyme of claim 17, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of R2Q, A3G, or L4M relative to the amino acid sequence set forth in SEQ ID NO:
1.
19. The recombinant enzyme of claim 17, wherein, the amino acid sequence of the recombinant enzyme comprises an amino acid substitution of V76C relative to the amino acid sequence set forth in SEQ ID NO:
1.
20. A genome editing system, wherein, The genome editing system comprises: 1) a recombinase recognition site sequence (RS) comprising an attP or attB site; 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase, the recombinase comprising any one of claims 1-19.
21. The genome editing system of claim 20, wherein, The recombinase recognition site sequence (RS) is naturally occurring, or engineered or optimized.
22. The genome editing system of claim 21, wherein, One or more of the recombinase recognition site sequences (RS) is inserted into a desired location in the genome by a recombinase recognition site integration unit.
23. The genome editing system of claim 22, wherein, The recombinase recognition site integration unit comprises: a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or the functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.
24. The genome editing system of claim 23, wherein, The CRISPR effector protein functional variant is a CRISPR nuclease with full / partial loss of cleavage activity, preferably the CRISPR effector protein functional variant is a CRISPR nickase, for example, Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase, or TraC nickase.
25. The genome editing system of claim 24, wherein, The guide RNA comprises a scaffold sequence and a primer binding sequence, and a recombinase recognition site sequence (RS) integration template sequence.
26. The genome editing system of any one of claims 22-25, wherein, The recombinase recognition site integration unit further comprises a reverse transcriptase and / or an expression construct comprising a nucleotide sequence encoding the reverse transcriptase.
27. The genome editing system of claim 26, wherein, The CRISPR nuclease of the recombinase recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with the CRISPR nuclease and targets a desired location in the genome, wherein the CRISPR nuclease makes a cut in a strand of the genome, and the reverse transcriptase incorporates a RS integration template sequence in the guide RNA into the cut site, thereby inserting at least one recombinase recognition site sequence (RS) recognizable by the recombinase at the desired location in the genome.
28. The genome editing system of claim 27, wherein, The reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii Marathonase RT (Marathon RT).
29. The genome editing system of any one of claims 22-28, wherein, The recombinase recognition site integration unit is not covalently linked to a recombinase; or, the recombinase recognition site integration unit is covalently linked to a recombinase.
30. The genome editing system of any one of claims 22-29, wherein, The genome editing system further comprises: 3) a donor of an exogenous nucleotide sequence to be inserted into the genome, optionally, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer.
31. A fusion protein comprising the genome editing system of any one of claims 20-30, wherein the CRISPR nuclease is linked at the C-terminus to a reverse transcriptase, which is linked via a linker to a recombinase.
32. A polynucleotide comprising a nucleotide sequence encoding the recombinase of any one of claims 1-19, the genome editing system of any one of claims 20-30, or the fusion protein of claim 31, optionally, the polynucleotide is RNA, e.g., mRNA.
33. An expression construct comprising the polynucleotide of claim 32.
34. A cell comprising the recombinase of any one of claims 1-19, the genome editing system of any one of claims 20-30, the fusion protein of claim 31, the polynucleotide of claim 32, or the expression construct of claim 33.
35. A kit comprising the recombinase of any one of claims 1-19, the genome editing system of any one of claims 20-30, the fusion protein of claim 31, the polynucleotide of claim 32, the expression construct of claim 33, or the cell of claim 34.
36. A method of performing gene editing in an organism or cell of an organism, wherein, The method introduces the recombinase of any one of claims 1-19, the polynucleotide of claim 32, the expression construct of claim 33, the genome editing system of any one of claims 20-30, or the fusion protein of claim 31 into an organism or a cell of an organism.
37. The method of claim 36, wherein, Components 1), 2), and optionally 3) of the genome editing system are introduced into an organism or a cell of an organism simultaneously.
38. The method of claim 36, wherein, Components 1), 2), and optionally 3) of the genome editing system are introduced into an organism or a cell of an organism stepwise, respectively.
39. The method of any one of claims 36-38, wherein, The component 1) is inserted into a donor construct of a genomic or exogenous nucleotide sequence in the same direction or in the opposite direction.
40. The method of any one of claims 36-39, wherein, The method comprises recombining DNA of the genome of the organism or organism cell; optionally, the recombining DNA of the genome of the organism or organism cell comprises deleting DNA, inverting DNA in the genome, and / or integrating exogenous DNA into the genome.
41. The method of any one of claims 36-40, wherein, The genome editing system is introduced into the cell by a method selected from the group consisting of calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), biolistic method, N-acetylgalactosamine (GalNAc) mediation, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.
42. The method of any one of claims 36-41, wherein, The organism or organism cell is from a mammal such as human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; plant, preferably crop plant, for example wheat, rice, corn, soybean, sunflower, leafy vegetable, lettuce, sorghum, rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava and potato. The component 1) is inserted into a donor construct of a genomic or exogenous nucleotide sequence in the same direction or in the opposite direction. The method comprises recombining DNA of the genome of the organism or organism cell; optionally, the recombining DNA of the genome of the organism or organism cell comprises deleting DNA, inverting DNA in the genome, and / or integrating exogenous DNA into the genome. The genome editing system is introduced into the cell by a method selected from the group consisting of calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), biolistic method, N-acetylgalactosamine (GalNAc) mediation, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation. The organism or organism cell is from a mammal such as human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; plant, preferably crop plant, for example wheat, rice, corn, soybean, sunflower, leafy vegetable, lettuce, sorghum, rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava and potato.
Citation Information
Patent Citations
Systems, methods, and compositions for site-specific genetic modification using programmable addition (PASTE) implemented with site-specific targeting elements
CN116419975A
Systems, methods, and compositions for site-specific genetic engineering using programmable addition via site-specific targeting elements (PASTE)
US20220154224A1
Serine recombinase systems for site-specific gene editing
WO2023147507A1
Evolved recombinases for editing a genome in combination with prime editing
WO2024168147A2