Isolated nuclease and use thereof

The IS200/IS605 family RNA-mediated endonucleases address the limitations of CRISPR/Cas9 by providing compact and efficient gene editing tools with specific recognition of transposon-associated motifs, enhancing delivery and editing efficiency.

WO2026031760A1PCT designated stage Publication Date: 2026-02-12ASTRAGENOMICS TECHNOLOGIES PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/099634
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-08
Filing Date
2025-06-06
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

The limitations of CRISPR/Cas9 technology, such as large protein size and high off-target effects, hinder its application in clinical medicine and gene editing efficiency, while existing alternatives like Cas12 offer lower efficiency and flexibility.

Method used

Development of RNA-mediated endonucleases from the IS200/IS605 family with a compact protein structure and specific recognition of transposon-associated motifs, providing diverse and efficient gene editing tools.

Benefits of technology

The new nucleases enable efficient gene editing with reduced protein size, allowing for various delivery methods and higher editing efficiency compared to SpCas9 and Cas12, offering more flexible and precise gene editing options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025099634_12022026_PF_FP_ABST
    Figure CN2025099634_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The field of molecular biology, and specifically an isolated nuclease and the use thereof are related to. A nucleic acid and a nucleic acid construct encoding the nuclease, a guide RNA and a nucleic acid construct thereof, and a composition, a recombinant vector, a recombinant host cell and a kit comprising the nuclease are further related to specifically. A method for introducing a double-strand break into a targeting gene of a host cell, a method for deleting, replacing or inserting a targeting gene of a host cell, and a method for obtaining a host cell in which a targeting gene is deleted, replaced or inserted are further related to specifically. The use of the nuclease, the nucleic acid and the nucleic acid construct encoding the nuclease, the guide RNA and the nucleic acid construct thereof, the composition, the recombinant vector, or the recombinant host cell for introducing a double-strand break into a targeting gene of a host cell, deleting, replacing or inserting a targeting gene of a host cell, and preparing a drug or a preparation are further related to specifically.
Need to check novelty before this filing date? Find Prior Art

Description

ISOLATED NUCLEASE AND USE THEREOF

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to PCT Application No. PCT / CN2024 / 110700 filed on August 8, 2024, the entire contents of which are hereby incorporated by reference in their entirety for all purpose.TECHNICAL FIELD

[0003] The present application relates to the field of molecular biology, and specifically to an isolated nuclease and the use thereof. The present application further specifically relates to: a nucleic acid and a nucleic acid construct encoding the nuclease, a guide RNA and a nucleic acid construct thereof, and a composition, a recombinant vector, a recombinant host cell and a kit comprising the nuclease. The present application further specifically relates to: a method for introducing a double-strand break into a targeting gene of a host cell, a method for deleting, replacing or inserting a targeting gene of a host cell, and a method for obtaining a host cell in which a targeting gene is deleted, replaced or inserted. The present application further specifically relates to the use of the nuclease, the nucleic acid and the nucleic acid construct encoding the nuclease, the guide RNA and the nucleic acid construct thereof, the composition, the recombinant vector, or the recombinant host cell for introducing a double-strand break into a targeting gene of a host cell, deleting, replacing or inserting a targeting gene of a host cell, and preparing a drug or a preparation for gene therapy, cell therapy, genome research, and stem cell induction and post-induction differentiation.BACKGROUND

[0004] With the rapid development of modern biotechnology and the advent of post-genome era, people are entering the stage of rewriting or even redesigning genetic information from the stage of reading biological genetic DNA information. The discovery of CRISPR / Cas9 technology has made a revolutionary breakthrough in gene editing technology. CRISPR / Cas9 is an RNA-mediated targeted gene editing tool, which can specifically recognize and cleave different endogenous DNA sequences through reprogramming of sgRNA. Cas9 has two nuclease domains, RuvC and HNH, which are responsible for the cleavage of either strand of DNA respectively. Mutating either of these sites can convert Cas9 into a single-strand Cas9 nickase. Important new technologies concerning Cas9, such as base editing and prime editing, are all designed based on Cas9 nickase.

[0005] However, some shortcomings of CRISPR / Cas9 limit its application: First, the CDS sequence of SpCas9 has a length exceeding 4.1 Kb, which exceeds the maximum effective packaging capacity of adenovirus (AAV) , and therefore it is difficult for the adenovirus-mediated gene delivery; although lentivirus has a stronger packaging capacity than AAV (with an upper loading limit of about 9 kb) , the proportion of proteins in SpCas9 is still too high, limiting the potential for subsequent engineering. These shortcomings seriously restrict the application of SpCas9 in clinical medicine. Subsequently, CRISPR / Cas12 or 12f system with a smaller molecular weight appears, but the editing efficiency of proteins such as Cas12 is not superior to that of SpCas9. Therefore, SpCas9 is still widely accepted and used at present. Second, the PAM sequence of SpCas9, which is the NGG sequence, is relatively simple and has a higher occurrence rate in the genome. Its advantage lies in the flexibility in reprograming sgRNA to complete the recognition and cleavage of different DNA sequences. However, this flexibility also leads to the off-target effects of suboptimal genome editing outcomes.

[0006] Therefore, gene editing technologies realized using RNA-mediated endonuclease, i.e., insertion sequences IscB and TnpB from IS200 / IS605 family, appear subsequently. They are widely distributed in microorganisms and have a more compact protein structure, with a size of about 400 aa that is less than 1 / 3 of SpCas9, so they have greater potential for engineering in terms of the application of enzymes. TnpB cleaves DNA next to the 5’ TTGAT transposon-associated motif (TAM) through reRNA (right element RNA, derived from RE element in ISDra2 transposon) mediation, thereby breaking and mutating the DNA sequence in the genome. The DNA cleavage function of TnpB needs to meet two conditions at the same time: (1) TAM sequence; (2) a sequence located at the 3’ end of reRNA that matches with a targeting gene. Different nucleases can recognize different TAM, and therefore the excavation of more highly active nuclease tools and the verification and detection of their functions can provide more, better and flexible choices for the development of gene editing strategies.

[0007] It should be noted that methods described in this section are not necessarily methods that have been previously conceived or employed. It should not be assumed that any of the methods described in this section is considered to be the prior art just because they are included in this section, unless otherwise indicated expressly. Similarly, the problem mentioned in this section should not be considered to be universally recognized in any prior art, unless otherwise indicated expressly.SUMMARY

[0008] In order to solve the above problems, the present application is intended to find RNA-mediated endonucleases having a suitable protein molecular weight and good gene editing effects, and provide more diverse and specific tools for gene editing.

[0009] The present application provides an isolated nuclease, wherein the nuclease comprises an amino acid sequence as shown in the following formula:

[0010] (X1) (X2) a (X3) (X4) b (X5) (X6) c (X7) (X8) (X9) d (X10) (X11) e (X12) (X13) f (X14) (X15) g (X16) (X17) h (X18) (X19) i (X20) (X21) (X22) j (X23)

[0011] wherein a, b, c, d, e, f, g, h, i, and j are the numbers of amino acids; (X1) , (X3) , (X5) , (X7) , (X8) , (X10) , (X12) , (X14) , (X16) , (X18) , (X20) , (X21) , and (X23) are independently aliphatic amino acids; (X2) is any amino acid, and a is 2; (X4) is any amino acid, and b is 4; (X6) is any amino acid, and c is 10 or 11; (X9) is any amino acid, and d is 2; (X11) is any amino acid, and e is 2; (X13) is any amino acid, and f is 10, 11 or 12; (X15) is any amino acid, and g is 3; (X17) is any amino acid, and h is 1 or 2; (X19) is any amino acid, and i is 5; and (X22) is any amino acid, and j is 1. According to an embodiment of the present application, an isolated nuclease can be provided, wherein the nuclease has a nuclease sequence selected from the following (i) or a variant sequence of the aforementioned nuclease having a nuclease activity in (ii) - (iv) : (i) at least one amino acid sequence as shown in any one of SEQ ID NOs: 1-86; (ii) at least one of sequences obtained by performing deletion, substitution, insertion, or mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids on the amino acid sequence as shown in any one of SEQ ID NOs: 1-86; (iii) at least one of amino acid sequences having at least 70%, 80%, 90%, 95%or 99%identity to the amino acid sequence as shown in any one of SEQ ID NOs: 1-86; and (iv) at least one of sequences obtained by further fusing the amino acid sequence as shown in any one of SEQ ID NOs: 1-86 with other sequences.

[0012] According to an embodiment of the present application, a guide RNA can be provided, wherein the guide RNA comprises a reRNA, the reRNA comprises a nucleotide sequence as shown in any one of SEQ ID NOs: 87-172 or a variant thereof, and the guide RNA can bind to a specific nuclease.

[0013] According to an embodiment of the present application, a nucleic acid can be provided, wherein, the nucleic acid encodes the nuclease described in the present application and / or the guide RNA described in the present application.

[0014] According to an embodiment of the present application, a nucleic acid construct can be provided, comprising the nucleic acid described in the present application, and further comprising a promoter.

[0015] According to an embodiment of the present application, a composition may be provided, wherein, the composition includes: an IS200 / IS605 family nuclease or a functional fragment thereof, or comprises a nucleic acid encoding the IS200 / IS605 family nuclease or the functional fragment thereof, and the nuclease or the functional fragment thereof has endonuclease activity; and a guide RNA, or comprises a nucleic acid encoding the guide RNA, and the guide RNA can bind to a specific nuclease.

[0016] According to an embodiment of the present application, a recombinant vector can be provided, wherein, the recombinant vector comprises the nucleic acid encoding the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, or the composition described in the present application.

[0017] According to an embodiment of the present application, a recombinant host cell can be provided, wherein, the recombinant host cell comprises the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.

[0018] According to an embodiment of the present application, a method for introducing a double-strand break into a targeting gene of a host cell can be provided, wherein the method comprises: delivering the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0019] According to an embodiment of the present application, a method for deleting, replacing or inserting a targeting gene of a host cell can be provided, wherein the method comprises: delivering the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0020] According to an embodiment of the present application, a method for obtaining a host cell in which a targeting gene is deleted, replaced or inserted can be provided, wherein the method comprises: delivering the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0021] According to an embodiment of the present application, the use of the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for introducing a double-strand break into a targeting gene of a host cell can be provided.

[0022] According to an embodiment of the present application, the use of the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for deleting, replacing or inserting a targeting gene of a host cell can be provided.

[0023] According to an embodiment of the present application, the use of the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for preparing a drug or a preparation for gene therapy, cell therapy, genome research, and stem cell induction and post-induction differentiation can be provided.

[0024] According to an embodiment of the present application, a kit can be provided, wherein, the kit comprises the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.

[0025] The protein molecular weight of the nuclease described in the present application is far less than that of SpCas9, about less than one third of the latter, which provides more possibilities for a variety of in vivo delivery in subsequent gene therapy, and can solve the problem of having difficulty in the delivery of SpCas9 caused by the protein size; and compared with AsCas12 which also has a low protein molecular weight, the nuclease has higher gene editing efficiency, which provides the possibility of same becoming a new gene editing application tool; additionally, since different nucleases can recognize different transposon-associated motifs, the novel nuclease discovered in the present application brings more choices for subsequent application scenarios of different scales.

[0026] It should be understood that the content described in this section is not intended to identify critical or important features of the examples of the present application and is not used to limit the scope of the present application. Other features of the present application will be easily understood through the following description.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings exemplarily show embodiments and form a part of the specification, and are used to explain exemplary implementations of the embodiments together with a written description of the specification. The embodiments shown are merely for illustrative purposes and do not limit the scope of the claims. Throughout the accompanying drawings, the same reference numerals denote similar but not necessarily same elements.

[0028] FIG. 1 shows a schematic diagram of an RGS dual fluorescence surrogate reporter system in example 1.

[0029] FIG. 2 shows GFP expression for all the active nuclease candidates of TP_O_18, TP_N_34, TP_K_36, TP_K_57, TP_N_9, TP_N_10, TP_N_11, TP_N_12, TP_N_14, TP_N_15, TP_N_18, TP_N_22, TP_N_28, TP_N_30, TP_N_31, TP_N_35, TP_N_36, TP_N_37, TP_N_42, TP_N_43, TP_N_45, TP_N_46, TP_N_51, TP_N_52, TP_N_54, TP_N_59, TP_O_2, TP_O_5, TP_O_6, TP_O_9, TP_O_11, TP_O_14, TP_O_16, TP_O_17, TP_O_19, TP_O_20, TP_O_21, TP_O_22, TP_O_23, TP_O_24, TP_O_25, TP_O_26, TP_O_27, TP_O_28, TP_O_30, TP_O_32, TP_O_33, TP_O_34, TP_O_35, TP_O_36, TP_O_37, TP_O_39, TP_O_40, TP_O_41, TP_O_42, TP_O_43, TP_O_44, TP_O_45, TP_O_48, TP_O_49, TP_O_50, TP_O_53, TP_O_55, TP_O_56, TP_O_57, TP_O_58, TP_O_59, TP_O_63, TP_O_66, TP_O_67, TP_O_68, TP_O_70, TP_O_71, TP_O_74, TP_O_75, TP_O_76, TP_O_77, TP_O_78, TP_O_79, TP_O_80, TP_P_4, TP_P_8, TP_P_10, TP_P_19, TP_P_47, and TP_P_60. All the results were quantified by flow cytometry assay in example 2.

[0030] FIG. 3 shows endogenous editing efficiency (quantified by the proportion of reads with insertions or deletions at target site) of TP_N_15, TP_N_34, TP_O_5, TP_O_6, TP_O_11, TP_O_17, TP_O_18, TP_O_55, and TP_O_56 in example 3.

[0031] FIG. 4 shows endogenous editing efficiency (quantified by the proportion of reads with insertions or deletions at target site) of the nucleases with same TAM sequences (CCAT or TTTAA) in example 4.

[0032] FIG. 5 shows a schematic diagram illustrating the order of elements in the report vectors in example 5.

[0033] FIG. 6 shows a schematic diagram demonstrating how the reporter vectors work in example 5.

[0034] FIGs. 7, 8, 9, 10, 11, 12, 13, and 14 show YFP expression in example 5 for all the active nuclease candidates of TP_O_18, TP_N_34, TP_K_36, TP_K_57, TP_N_9, TP_N_10, TP_N_11, TP_N_12, TP_N_14, TP_N_15, TP_N_18, TP_N_22, TP_N_28, TP_N_30, TP_N_31, TP_N_35, TP_N_36, TP_N_37, TP_N_42, TP_N_43, TP_N_45, TP_N_46, TP_N_51, TP_N_52, TP_N_54, TP_N_59, TP_O_2, TP_O_5, TP_O_6, TP_O_9, TP_O_11, TP_O_14, TP_O_16, TP_O_17, TP_O_19, TP_O_20, TP_O_21, TP_O_22, TP_O_23, TP_O_24, TP_O_25, TP_O_26, TP_O_27, TP_O_28, TP_O_30, TP_O_32, TP_O_33, TP_O_34, TP_O_35, TP_O_36, TP_O_37, TP_O_39, TP_O_40, TP_O_41, TP_O_42, TP_O_43, TP_O_44, TP_O_45, TP_O_48, TP_O_49, TP_O_50, TP_O_53, TP_O_55, TP_O_56, TP_O_57, TP_O_58, TP_O_59, TP_O_63, TP_O_66, TP_O_67, TP_O_68, TP_O_70, TP_O_71, TP_O_74, TP_O_75, TP_O_76, TP_O_77, TP_O_78, TP_O_79, TP_O_80, TP_P_4, TP_P_8, TP_P_10, TP_P_19, TP_P_47, TP_P_60.DETAILED DESCRIPTION OF EMBODIMENTS

[0035] Unless otherwise indicated or contradicts the context, the terms or expressions used herein should be read in conjunction with the entire content of the present disclosure and as understood by those of ordinary skill in the art. All technical and scientific terms used herein have the same meanings as commonly understood by those of ordinary skill in the art, unless otherwise defined.

[0036] In the present application, the terms “nucleic acid” and “polynucleotide” are used interchangeably, and refer to polymerization forms of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogs thereof.

[0037] In the present application, the terms “polypeptide” and “peptide” are used interchangeably, and refer to polymers of amino acids of any length. Therefore, polypeptides, oligopeptides, proteins, antibodies and enzymes are all included in the definition of polypeptide.

[0038] As described in the present application, the “fragment” of a sequence refers to a portion of a sequence. For example, the fragment of a nucleic acid sequence refers to a portion of the nucleic acid sequence, and the fragment of an amino acid sequence refers to a portion of the amino acid sequence.

[0039] As described in the present application, a “variant” of a sequence is a polynucleotide or polypeptide that differs from a reference polynucleotide or polypeptide, respectively, but retains essential properties. A typical variant of a polynucleotide differs in nucleic acid sequence from another reference polynucleotide, and the differences in nucleic acid sequence may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. A typical variant of a polypeptide differs in amino acid sequence from another reference polypeptide. Generally, the differences are limited so that the sequences of the reference polypeptide and the variant are generally very similar, and are identical in many regions. A variant polypeptide and a reference polypeptide may differ in amino acid sequence by one or more substitutions, additions, deletions in any combination. The substituted or inserted amino acid residue may or may not be a residue encoded by the genetic code. Variants of polynucleotides or polypeptides may be naturally occurring, such as allelic variations, or they may be unknown naturally occurring variants. Non-naturally occurring polynucleotide and polypeptide variants can be produced by mutagenesis techniques, direct synthesis, and other recombinant methods known to the skilled artisan.

[0040] Amino acids are usually classified by the properties of their side chains. For example, side chains may render amino acids weak acids (e.g., amino acids D and E) or weak bases (e.g., amino acids K, R and H) ; and if the side chains are polar, the amino acids become hydrophilic (e.g., amino acids L and I) , or if the side chains are nonpolar, the amino acids become hydrophobic (e.g., amino acids S and C) .

[0041] As described in the present application, the “aliphatic amino acid” has a side chain that is an aliphatic group. Aliphatic groups cause amino acids to be nonpolar and hydrophobic. The aliphatic group is preferably an unsubstituted branched or linear alkyl group. Non-limiting examples of the aliphatic amino acids are A (alanine) , V (valine) , L (leucine) , I (isoleucine) , M (methionine) , D (aspartic acid) , E (glutamic acid) , K (lysine) , R (arginine) , G (glycine) , S (serine) , T (threonine) , C (cysteine) , N (asparagine) , and Q (glutamine) .

[0042] As described in the present application, the “nonpolar amino acid” has a nonpolar side chain that makes the amino acid hydrophobic. Non-limiting examples of the nonpolar amino acid are A (alanine) , V (valine) , L (leucine) , I (isoleucine) , F (phenylalanine) , W (tryptophan) , M (methionine) , P (proline) , and G (glycine) .

[0043] As described in the present application, the “polar amino acid” has a polar side chain that makes the amino acid hydrophilic. Non-limiting examples of the polar amino acid are T (threonine) , S (serine) , C (cysteine) , N (asparagine) , Q (glutamine) , Y (tyrosine) , K (lysine) , R (arginine) , H (histidine) , D (aspartic acid) , and E (glutamic acid) . Polar amino acids can be divided into polar uncharged amino acids or polar charged amino acids.

[0044] As described in the present application, the “polar uncharged amino acid” has a polar side chain of uncharged residues. Non-limiting examples of the polar uncharged amino acid are T (threonine) , S (serine) , C (cysteine) , N (asparagine) , Q (glutamine) , and Y (tyrosine) .

[0045] As described in the present application, the “polar charged amino acid” has a polar side chain of at least one charged residue. Non-limiting examples of the polar charged amino acid are K (lysine) , R (arginine) , H (histidine) , D (aspartic acid) , and E (glutamic acid) . Polar charged amino acids can be divided into positively charged amino acids or negatively charged amino acids.

[0046] As described in the present application, the “positively charged amino acid” has a polar side chain of at least one positively charged residue. Non-limiting examples of the positively charged amino acid are K (lysine) , R (arginine) , and H (histidine) .

[0047] As described in the present application, the “negatively charged amino acid” has a polar side chain of at least one negatively charged residue. Non-limiting examples of the negatively charged amino acid are D (aspartic acid) , and E (glutamic acid) .

[0048] The term “family” as used in the present application refers to a group of nucleic acids or proteins having high structural similarity produced by the same ancestor by means of replication and variation, which usually have related or even the same functions.

[0049] The term “nuclease” described in the present application refers to an enzyme capable of cleaving phosphodiester bonds. Nucleases hydrolyze the phosphodiester bonds in the backbone of nucleic acids. The term “endonuclease” described in the present application refers to an enzyme capable of cleaving phosphodiester bonds between nucleotides.

[0050] The term “guide RNA” described in the present application refers to any RNA molecule that can form a complex with the nuclease described in the present application. For example, the guide RNA can be a molecule that recognizes a targeting gene. In some embodiments of the present application, the guide RNA comprises a reRNA and a targeted sequence, wherein the reRNA can bind to a particular nuclease, and the targeted sequence can be designed to be complementary to a target strand of a targeting gene.

[0051] The term “transposon-associated motif” (TAM) described in the present application refers to a short nucleotide sequence adjacent to a targeting gene, which sequence can be recognized by a complex formed by nuclease and guide RNA described in the present application. If a targeting gene is not adjacent to a transposon-associated motif, the nuclease cannot successfully recognize the targeting gene. Sequences and lengths of the transposon-associated motif in the present application can vary depending on the nuclease.

[0052] The terms “targeting gene” “targeting sequence” “targeting nucleic acid” “gene of interest” , “sequence of interest” and “nucleic acid of interest” described in the present application are used interchangeably, and refer to nucleotide sequences on chromosomal DNA, chloroplast DNA, mitochondrial DNA, plasmid DNA, or any other DNA molecule in the genome of cells, which sequences can be recognized, bound to, and selectively cleaved by a complex formed by the nuclease and guide RNA described in the present application.

[0053] The term “nucleic acid construct” as used in the present application is defined as a single-stranded or double-stranded nucleic acid molecule herein, and preferably refers to an artificially constructed nucleic acid molecule. Optionally, the nucleic acid construct further includes one or more operably linked regulatory sequences, which can direct the expression of a coding sequence in a suitable host cell under compatible conditions. The term “expression” is understood to include any step involved in the production of a protein or polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification and secretion. The term “regulatory sequence” includes all components necessary or advantageous for expression of the polypeptide / protein of the present application. Each regulatory sequence may be naturally present or exogenous to the nucleic acid sequence encoding the protein or polypeptide. These regulatory sequences include, but are not limited to, leader sequences, polyadenylation sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, the regulatory sequences should include promoters and termination signals for transcription and translation. Regulatory sequences with linkers can be provided for the purpose of introduction into specific restriction sites for linking the regulatory sequences to the coding region of a nucleic acid sequence encoding a protein or polypeptide.

[0054] The term “promoter” as used in the present application refers to a polynucleotide sequence that can control the transcription of a coding sequence. Promoter sequences include specific sequences sufficient to enable RNA polymerase to recognize, bind, and initiate transcription. In addition, promoter sequences may include sequences that optionally modulate the recognition, binding and transcription initiation activities of RNA polymerase in the nucleic acid construct provided in the present application. A promoter can affect the transcription of a gene located on the same nucleic acid molecule as the promoter or a gene located on a different nucleic acid molecule from the promoter.

[0055] The term “host cell” as used in the present application include, but are not limited to, an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell. This term includes a progeny of an original cell into which an exogenous nucleic acid fragment has been introduced. Exemplary host cell includes human embryonic kidney cell HEK293T. It is understood that, due to natural, accidental or intentional mutations, the progeny of a single parent cell may not necessarily be identical to the original parent morphologically or in terms of genome or total DNA complement.

[0056] The term “vector” as used in the present application refers to a nucleic acid molecule capable of transporting another nucleic acid molecule connected to it. Examples of vectors include, but are not limited to, plasmids, viruses, bacteria, phages, and insertable DNA fragments. The term “plasmid” refers to a circular double-stranded DNA capable of accepting an exogenous nucleic acid fragment and replicating in prokaryotic or eukaryotic cells.

[0057] Nuclease

[0058] The present application provides an isolated nuclease, wherein the nuclease comprises an amino acid sequence as shown in the following formula:

[0059] (X1) (X2) a (X3) (X4) b (X5) (X6) c (X7) (X8) (X9) d (X10) (X11) e (X12) (X13) f (X14) (X15) g (X16) (X17) h (X18) (X19) i (X20) (X21) (X22) j (X23)

[0060] wherein a, b, c, d, e, f, g, h, i, and j are the numbers of amino acids; (X1) , (X3) , (X5) , (X7) , (X8) , (X10) , (X12) , (X14) , (X16) , (X18) , (X20) , (X21) , and (X23) are independently aliphatic amino acids; (X2) is any amino acid, and a is 2; (X4) is any amino acid, and b is 4; (X6) is any amino acid, and c is 10 or 11; (X9) is any amino acid, and d is 2; (X11) is any amino acid, and e is 2; (X13) is any amino acid, and f is 10, 11 or 12; (X15) is any amino acid, and g is 3; (X17) is any amino acid, and h is 1 or 2; (X19) is any amino acid, and i is 5; and (X22) is any amino acid, and j is 1.

[0061] In some embodiments, the (X1) and (X5) are independently nonpolar amino acids; (X3) , (X7) , (X8) , (X10) , (X12) , (X14) , (X16) , (X18) , (X20) , (X21) , and (X23) are independently polar amino acids.

[0062] In some embodiments, the (X1) is a nonpolar amino acid; (X3) is a positively charged amino acid; (X5) is a nonpolar amino acid; (X7) is a polar uncharged amino acid; (X8) is a polar uncharged amino acid; (X10) is a polar uncharged amino acid; (X12) is a polar uncharged amino acid; (X14) is a positively charged amino acid; (X16) is a polar uncharged amino acid; (X18) is a polar uncharged amino acid; (X20) is a positively charged amino acid; (X21) is a negatively charged amino acid; and / or (X23) is a polar uncharged amino acid.

[0063] According to an embodiment of the present application, an isolated nuclease can be provided, wherein the nuclease has a nuclease sequence selected from the following (i) or a variant sequence of the aforementioned nuclease having a nuclease activity in (ii) - (iv) : (i) at least one amino acid sequence as shown in any one of SEQ ID NOs: 1-86; (ii) at least one of sequences obtained by performing deletion, substitution, insertion, or mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids on the amino acid sequence as shown in any one of SEQ ID NOs: 1-86; (iii) at least one of amino acid sequences having at least 70%, 80%, 90%, 95%or 99%identity to the amino acid sequence as shown in any one of SEQ ID NOs: 1-86; and (iv) at least one of sequences obtained by further fusing the amino acid sequence as shown in any one of SEQ ID NOs: 1-86 with other sequences.

[0064] In some embodiments, the nuclease has a nuclease sequence selected from at least one of the following groups (1) - (6) : (1) at least one amino acid sequence as shown in any one of SEQ ID NOs: 35-54 and 56-61; (2) at least one amino acid sequence as shown in any one of SEQ ID NOs: 1 and 4-23; (3) at least one amino acid sequence as shown in any one of SEQ ID NOs: 63-80; (4) at least one amino acid sequence as shown in any one of SEQ ID NOs: 26-31; (5) at least one amino acid sequence as shown in any one of SEQ ID NOs: 81-86; and / or (6) at least one amino acid sequence as shown in any one of SEQ ID NOs: 2 and 24-25.

[0065] In some embodiments, wherein the nuclease has a nuclease sequence selected from at least one of the following groups (1) - (5) : (1) at least one amino acid sequence as shown in any one of SEQ ID NOs: 35-62; (2) at least one amino acid sequence as shown in any one of SEQ ID NOs: 1-25; (3) at least one amino acid sequence as shown in any one of SEQ ID NOs: 63-80; (4) at least one amino acid sequence as shown in any one of SEQ ID NOs: 26-34; and / or (5) at least one amino acid sequence as shown in any one of SEQ ID NOs: 81-86.

[0066] In some embodiments, the nuclease belongs to the IS200 / IS605 family. In some embodiments, the nuclease belongs to the IS605, or IS1341 subfamily. In some embodiments, the species sources of the nuclease include Bacteria. In some embodiments, the species sources of the nuclease include Actinobacteria or Firmicutes.

[0067] Guide RNA

[0068] According to an embodiment of the present application, a guide RNA can be provided, wherein the guide RNA comprises a reRNA, the reRNA comprises a nucleotide sequence as shown in any one of SEQ ID NOs: 87-172 or a variant thereof, and the guide RNA can bind to a specific nuclease. In some embodiments, the reRNA comprises at least one of nucleotide sequences having at least 70%, 80%, 90%, 95%or 99%identity to the nucleotide sequence as shown in any one of SEQ ID NOs: 87-172. In some embodiments, the reRNA comprises at least one of the nucleotide sequences as shown in any one of SEQ ID NOs: 87-172. In some embodiments, the reRNA is at least one of the nucleotide sequences as shown in any one of SEQ ID NOs: 87-172.

[0069] In some embodiments, the guide RNA further comprises a targeted sequence that can recognize a targeting gene adjacent to a transposon-associated motif. In some embodiments, the targeted sequence is of at least one of 10-50, 10-40, 10-30, or 15-25 nucleotides in length.

[0070] Sequences and lengths of the transposon-associated motif in the present application can vary depending on the nuclease, and the transposon-associated motif can be recognized by a complex formed by the nuclease and guide RNA described in the present application. In some embodiments, the transposon-associated motif comprises a nucleotide sequence selected from at least one of the following groups (1) - (5) : (1) GGGG; (2) CCAT; (3) TTAT; (4) TTTAA; and / or (5) AGGAG.

[0071] Nucleic acid, nucleic acid construct

[0072] According to an embodiment of the present application, a nucleic acid can be provided, wherein, the nucleic acid encodes the nuclease described in the present application and / or the guide RNA described in the present application.

[0073] According to an embodiment of the present application, a nucleic acid construct can be provided, comprising the nucleic acid described in the present application. In some embodiments, the nucleic acid construct further comprising a promoter. The promoter can be any suitable promoter sequence, that is, a nucleic acid sequence that can be recognized by a host cell expressing the nucleic acid sequence. The promoter sequence contains a transcriptional regulatory sequence that mediates the expression of the protein or polypeptide. The promoter can be any nucleic acid sequence having transcriptional activity in a selected host cell, including mutant, truncated and heterozygous promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, Polyhedrin promotor, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0074] In some embodiments, the nucleic acid construct is modified by 5’-end capping and / or 3’-end polyadenylating, and the nucleic acid construct retains the activity of nuclease and / or guide RNA. In some embodiments, the nucleic acid construct is modified by thiophosphate bond modification, 2’-MOE (2-O- (2-methoxyethyl) ) , PNA (peptide nucleic acid) , GNA (glycerol nucleic acid) , LNA (locked nucleic acid) , GalNAc (N-acetylgalactosamine) LNP (lipid nano particle) PNP (peptide nanoparticles) . The modification methods of nucleic acid are known in the art, the entire contents of which are hereby incorporated by reference.

[0075] In some embodiments, the nucleic acid construct further comprises a polyA sequence. PolyA tailing signal sequences well known in the art, as well as various truncated forms of polyA tailing signals, can be used in the present application.

[0076] In some embodiments, the nucleic acid construct further includes any transcription termination sequence, i.e., a sequence that is recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3’-terminus of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that is functional in the host cell of choice can be used in the present invention.

[0077] Optionally, the nucleic acid construct may further include a suitable leader sequence, that is, an untranslated region in the mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5’-terminus of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in the host cell of choice can be used in the present invention.

[0078] Optionally, the nucleic acid construct may further include a propeptide coding region, which encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or a propolypeptide. The propolypeptide is usually inactive and can be converted into a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.

[0079] Optionally, the nucleic acid construct may further include a regulatory sequence that can regulate the expression of the polypeptide according to the growth conditions of the host cell. Examples of the regulatory sequence are systems that turn gene expression on or off in response to chemical or physical stimuli, including in the presence of regulatory compounds. Other examples of the regulatory sequence are those that enable gene amplification. In these instances, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.

[0080] Composition

[0081] According to an embodiment of the present application, a composition may be provided, wherein, the composition includes: an IS200 / IS605 family nuclease or a functional fragment thereof, or comprises a nucleic acid encoding the IS200 / IS605 family nuclease or the functional fragment thereof, and the nuclease or the functional fragment thereof has endonuclease activity; and a guide RNA, or comprises a nucleic acid encoding the guide RNA, and the guide RNA can bind to a specific nuclease.

[0082] In some embodiments, the composition is selected from at least one of the following groups (1) - (87) , and any one of the following groups (1) - (87) comprises: a nuclease-related sequence and a guide RNA-related sequence,

[0083] (1) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 1 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 87;

[0084] (2) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 2 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 88;

[0085] (3) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 3 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 89;

[0086] (4) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 4 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 90;

[0087] (5) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 5 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 91;

[0088] (6) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 6 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 92;

[0089] (7) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 7 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 93;

[0090] (8) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 8 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 94;

[0091] (9) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 9 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 95;

[0092] (10) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 10 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 96;

[0093] (11) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 11 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 97;

[0094] (12) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 12 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 98;

[0095] (13) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 13 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 99;

[0096] (14) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 14 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 100;

[0097] (15) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 15 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 101;

[0098] (16) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 16 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 102;

[0099] (17) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 17 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 103;

[0100] (18) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 18 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 104;

[0101] (19) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 19 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 105;

[0102] (20) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 20 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 106;

[0103] (21) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 21 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 107;

[0104] (22) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 22 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 108;

[0105] (23) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 23 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 109;

[0106] (24) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 24 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 110;

[0107] (25) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 25 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 111;

[0108] (26) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 26 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 112;

[0109] (27) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 27 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 113;

[0110] (28) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 28 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 114;

[0111] (29) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 29 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 115;

[0112] (30) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 30 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 116;

[0113] (31) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 31 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 117;

[0114] (32) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 32 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 118;

[0115] (33) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 33 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 119;

[0116] (34) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 34 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 120;

[0117] (35) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 35 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 121;

[0118] (36) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 36 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 122;

[0119] (37) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 37 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 123;

[0120] (38) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 38 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 124;

[0121] (39) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 39 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 125;

[0122] (40) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 40 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 126;

[0123] (41) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 41 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 127;

[0124] (42) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 42 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 128;

[0125] (43) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 43 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 129;

[0126] (44) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 44 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 130;

[0127] (45) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 45 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 131;

[0128] (46) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 46 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 132;

[0129] (47) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 47 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 133;

[0130] (48) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 48 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 134;

[0131] (49) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 49 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 135;

[0132] (50) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 50 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 136;

[0133] (51) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 51 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 137;

[0134] (52) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 52 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 138;

[0135] (53) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 53 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 139;

[0136] (54) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 54 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 140;

[0137] (55) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 55 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 141;

[0138] (56) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 56 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 142;

[0139] (57) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 57 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 143;

[0140] (58) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 58 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 144;

[0141] (59) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 59 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 145;

[0142] (60) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 60 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 146;

[0143] (61) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 61 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 147;

[0144] (62) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 62 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 148;

[0145] (63) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 63 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 149;

[0146] (64) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 64 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 150;

[0147] (65) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 65 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 151;

[0148] (66) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 66 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 152;

[0149] (67) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 67 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 153;

[0150] (68) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 68 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 154;

[0151] (69) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 69 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 155;

[0152] (70) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 70 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 156;

[0153] (71) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 71 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 157;

[0154] (72) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 72 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 158;

[0155] (73) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 73 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 159;

[0156] (74) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 74 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 160;

[0157] (75) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 75 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 161;

[0158] (76) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 76 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 162;

[0159] (77) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 77 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 163;

[0160] (78) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 78 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 164;

[0161] (79) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 79 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 165;

[0162] (80) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 80 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 166;

[0163] (81) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 81 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 167;

[0164] (82) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 82 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 168;

[0165] (83) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 83 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 169;

[0166] (84) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 84 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 170;

[0167] (85) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 85 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 171;

[0168] (86) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 86 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 172;

[0169] (87) a variant of any one of the aforementioned groups (1) - (86) ,

[0170] wherein the nuclease-related sequence is the amino acid sequence of the variant of the nuclease in each group or a nucleic acid sequence encoding the variant, and the variant has a variant sequence of the aforementioned nuclease having a nuclease activity selected from the following (i) - (iii) :

[0171] (i) at least one of sequences obtained by performing deletion, substitution, insertion, or mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids on the amino acid sequence of the nuclease in each group;

[0172] (ii) at least one of amino acid sequences having at least 70%, 80%, 90%, 95%or 99%identity to the amino acid sequence as shown in any one of SEQ ID NOs: 1-86; and

[0173] (iii) at least one sequence obtained by further fusing the amino acid sequence as shown in any one of SEQ ID NO: 1-86 with other sequences.

[0174] In some embodiments, the guide RNA-related sequence further comprises a targeted sequence that can recognize a targeting gene adjacent to a transposon-associated motif. In some embodiments, the targeted sequence is of at least one of 10-50, 10-40, 10-30, or 15-25 nucleotides in length.

[0175] Sequences and lengths of the transposon-associated motif in the present application can vary depending on the nuclease, and the transposon-associated motif can be recognized by a complex formed by the nuclease and guide RNA described in the present application. In some embodiments, the transposon-associated motif comprises a nucleotide sequence selected from at least one of the following groups (1) - (5) : (1) GGGG; (2) CCAT; (3) TTAT; (4) TTTAA; and / or (5) AGGAG.

[0176] The targeting gene in the present application includes any gene of interest, e.g., a gene of a natural functional protein, an artificial chimeric gene, or a gene of a non-coding RNA. In some embodiments, the gene of a natural functional protein includes a fluorescein reporter gene, a luciferase gene, and a resistance gene. In some embodiments, the artificial chimeric gene includes a gene of a chimeric antigen receptor. In some embodiments, the fluorescein reporter gene includes a gene encoding a green fluorescent protein, a red fluorescent protein, a blue fluorescent protein, or a yellow fluorescent protein. In some embodiments, the luciferase gene includes a gene encoding firefly luciferase or sea kidney luciferase. In some embodiments, the resistance gene includes a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.

[0177] In some embodiments, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA further comprises a promoter. The promoter can be any suitable promoter sequence, that is, a nucleic acid sequence that can be recognized by a host cell expressing the nucleic acid sequence. The promoter sequence contains a transcriptional regulatory sequence that mediates the expression of the protein or polypeptide. The promoter can be any nucleic acid sequence having transcriptional activity in a selected host cell, including mutant, truncated and heterozygous promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, Polyhedrin promotor, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL. In some embodiments, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA further comprises a polyA sequence. PolyA tailing signal sequences well known in the art, as well as various truncated forms of polyA tailing signals, can be used in the present application.

[0178] In some embodiments, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA further comprises any transcription termination sequence that controls the expression of the exogenous nucleic acid fragment, i.e., a sequence that is recognized by a host cell to terminate transcription. Any terminator that is functional in the host cell of choice can be used in the present invention.

[0179] In some embodiments, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA further comprises any transcription termination sequence, i.e., a sequence that is recognized by a host cell to terminate transcription. The termination sequence is operably linked to the 3’-terminus of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that is functional in the host cell of choice can be used in the present invention.

[0180] Optionally, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA may further comprise a suitable leader sequence, i.e., an untranslated region in the mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5’-terminus of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in the host cell of choice can be used in the present invention.

[0181] Optionally, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA may further comprise a propeptide coding region, which encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or a propolypeptide. The propolypeptide is usually inactive and can be converted into a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.

[0182] Optionally, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA may further comprise a regulatory sequence that can regulate the expression of the polypeptide according to the growth conditions of the host cell. Examples of the regulatory sequence are systems that turn gene expression on or off in response to chemical or physical stimuli, including in the presence of regulatory compounds. Other examples of the regulatory sequence are those that enable gene amplification. In these instances, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.

[0183] Recombinant vector, recombinant host cell and kit

[0184] According to an embodiment of the present application, a recombinant vector can be provided, wherein, the recombinant vector comprises the nucleic acid encoding the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, or the composition described in the present application. The recombinant vector can be any suitable vector. In some embodiments, the recombinant vector includes, but is not limited to, a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector. In some embodiments, the recombinant eukaryotic expression plasmid includes pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV. In some embodiments, the recombinant virus vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector. The recombinant vector of the present invention can be constructed using methods well known in the art. For example, depending on the restriction sites contained in the backbone vector used, appropriate restriction sites can be added to both ends of the nucleic acid construct of the present invention, and then loaded into the backbone vector.

[0185] According to an embodiment of the present application, a recombinant host cell can be provided, wherein, the recombinant host cell comprises the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application. The recombinant host cell can be any host cell in which nucleases can be used. In some embodiments, the recombinant host cell includes, but is not limited to, an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the animal cell includes a mammalian cell. In some embodiments, the mammalian cell includes a primary cell (e.g., a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell) , an immortalized cell line (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines) , a cancer cell line (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca) , an embryonic stem cell line (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW. 4, R1, and D3) and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof. In some embodiments, the plant cell includes a monocot cell or a dicot cell. In some embodiments, the monocot cell or the dicot cell includes rice cell, maize cell, or soybean cell.

[0186] According to an embodiment of the present application, a kit can be provided, wherein, the kit comprises the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.

[0187] Method and use

[0188] The nuclease-based gene editing tools and methods provided in the present application can be applied to many fields such as gene therapy, molecular breeding in animals and plants, industrial microorganism engineering, model animal engineering, and scientific research. Particularly in the field of gene therapy, it can be applied for gene knockout based on DNA double-strand breaks in human genome.

[0189] According to an embodiment of the present application, a method for introducing a double-strand break into a targeting gene of a host cell can be provided, wherein the method comprises: delivering the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0190] According to an embodiment of the present application, a method for deleting, replacing or inserting a targeting gene of a host cell can be provided, wherein the method comprises: delivering the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0191] According to an embodiment of the present application, a method for obtaining a host cell in which a targeting gene is deleted, replaced or inserted can be provided, wherein the method comprises: delivering the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0192] The method of delivery into the host cell can be any suitable method. In some embodiments, the delivery method includes but is not limited to cationic liposome delivery, lipoid nanoparticulate delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticulate delivery, polypeptide and protein delivery, retrovirus delivery, lentivirus delivery, adenovirus delivery, adeno-associated virus delivery, electroporation, agrobacterium infection, or gene gun. The methods of cell transfection and culture are routine methods in the art, and appropriate transfection and culture methods can be selected according to different cell types.

[0193] The host cell can be any host cell in which nucleases can be used. In some embodiments, the host cell includes, but is not limited to, an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the animal cell includes a mammalian cell. In some embodiments, the mammalian cell includes a primary cell (e.g., a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell) , an immortalized cell line (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines) , a cancer cell line (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca) , an embryonic stem cell line (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW. 4, R1, and D3) and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof. In some embodiments, the plant cell includes a monocot cell or a dicot cell. In some embodiments, the monocot cell or the dicot cell includes rice cell, maize cell, or soybean cell.

[0194] According to an embodiment of the present application, the use of the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for introducing a double-strand break into a targeting gene of a host cell can be provided.

[0195] According to an embodiment of the present application, the use of the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for deleting, replacing or inserting a targeting gene of a host cell can be provided.

[0196] The host cell can be any host cell in which nucleases can be used. In some embodiments, the host cell includes, but is not limited to, an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the animal cell includes a mammalian cell. In some embodiments, the mammalian cell includes a primary cell (e.g., a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell) , an immortalized cell line (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines) , a cancer cell line (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca) , an embryonic stem cell line (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW. 4, R1, and D3) and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof. In some embodiments, the plant cell includes a monocot cell or a dicot cell. In some embodiments, the monocot cell or the dicot cell includes rice cell, maize cell, or soybean cell.

[0197] According to an embodiment of the present application, the use of the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for preparing a drug or a preparation for gene therapy, cell therapy, genome research, and stem cell induction and post-induction differentiation can be provided.

[0198] The above various embodiments and preferences for the present application can be combined with each other (as long as they are not inherently contradictory to each other) and are suitable for the use of the present application, and the various embodiments formed by such combinations are considered as a part of the present application.

[0199] EXAMPLES

[0200] Exemplary embodiments of the present application are described below in conjunction with the accompanying drawings, where various details of the examples of the present application are included to facilitate understanding. It should be understood that they are considered to be exemplary only and not intended to limit the protection scope of the present application. The protection scope of the present application is only defined by the claims. Therefore, those of ordinary skill in the art should be aware that various changes and modifications can be made to the examples described herein, without departing from the scope of the present application. Likewise, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0201] Unless otherwise stated, the reagents and instruments used in the following examples are conventional products that are commercially available. Unless otherwise stated, experiments are performed under conventional conditions or conditions recommended by the manufacturer.

[0202] Example 1: Construction of nuclease activity detection system

[0203] A set of an RGS dual fluorescence surrogate reporter system was established to verify the activity of candidate nucleases.

[0204] Plasmid 1 consists of a complete set of elements capable of transcribing and expressing candidate nuclease proteins, comprising a constitutive promoter CMV (sequence as shown in SEQ ID NO: 173) that can initiate transcription in an eukaryotic cell, a candidate nuclease sequence (as shown in Table 1) , a 5’-nuclear localization signal peptide sequence (sequence as shown in SEQ ID NO: 174) , a 3’-nuclear localization signal peptide sequence (sequence as shown in SEQ ID NO: 175) , a polyA sequence (sequence as shown in SEQ ID NO: 176) that terminates transcription, and an ampicillin resistance gene sequence (sequence as shown in SEQ ID NO: 177) .

[0205] Method for constructing plasmid 1: The amino acid sequence (or nucleotide sequence) of the candidate nuclease protein was synthesized through conventional gene synthesis by BGI Tech Solutions (Beijing Liuhe) Co., Ltd., with an ECoRI cleavage site inserted into the upstream 5’ end of the sequence, and a BamH1 cleavage site inserted into the downstream 3’ end. Plasmid construction was also performed by the company responsible for the gene synthesis, and the specific construction method was as follows: 1. Preparation of vector. The plasmid backbone of a pcDNA3.1 plasmid vector was subjected to a double enzymatic cleavage digestion reaction using the single restriction endonuclease cleavage sites ECoRI and BamHI on the plasmid vector, a linearized plasmid vector fragment was obtained by agarose gel electrophoresis, and the enzymatic cleavage band was excised from the gel for recovery to obtain the purified linearized plasmid vector fragment. 2. Ligation. The nucleotide sequence of the candidate nuclease protein obtained through conventional gene synthesis was ligated with the linearized pcDNA3.1 vector fragment using a T4 DNA ligase. 3. Transformation and verification. Monoclonal transformants were obtained through a LB agar plate for screening ampicillin resistance, and the correct clone identified by sequencing was used as a candidate plasmid for later use.

[0206] Plasmid 2 comprises a reRNA sequence (as shown in Table 1) , with a 20 nt targeted sequence GCTCGGAGATCATCATTGCG inserted at the 3’ end of the reRNA sequence, a U6 promoter (sequence as shown in SEQ ID NO: 178) , a PBR322 replication origin (sequence as shown in SEQ ID NO: 179) , and an ampicillin resistance gene sequence (sequence as shown in SEQ ID NO: 177) .

[0207] Method for constructing plasmid 2: Guide reRNA was synthesized through conventional gene synthesis by Beijing Tsingke Biotech Co., Ltd. or General Biosystems (Anhui) Co., Ltd. Plasmid construction was also performed by the company responsible for the gene synthesis, and the specific construction method was as follows: 1. Preparation of vector. A pUC19-U6 vector was subjected to enzymatic cleavage using BbsI, a linearized plasmid vector fragment was obtained by agarose gel electrophoresis, and the enzymatic cleavage band was excised from the gel for recovery to obtain the purified linearized plasmid vector fragment. 2. Ligation. The nucleotide sequence of the guide reRNA obtained through gene synthesis was ligated with the linearized pUC19-U6 vector fragment using a ligation method of seamless cloning. 3. Transformation and verification. Monoclonal transformants were obtained through a LB agar plate for screening ampicillin resistance, and the correct clone identified by sequencing was used as a candidate plasmid for later use.

[0208] Plasmid 3 comprises a TAM sequence (as shown in Table 1) , with a 20 nt targeted sequence GCTCGGAGATCATCATTGCG inserted at the 3’ end of the TAM sequence, a CMV promoter (sequence as shown in SEQ ID NO: 180) , an ampicillin resistance gene sequence (sequence as shown in SEQ ID NO: 177) , and a surrogate reporter gene. The surrogate reporter gene can encode two fluorescent proteins (RFP sequence as shown in SEQ ID NO: 181, and GFP sequence as shown in SEQ ID NO: 182) . By means of the insertion of an endonuclease downstream of RFP and the insertion of an endonuclease upstream of GFP, TAM and a 20 nt targeted sequence at the 3’ end of TAM can be recognized. When there is no endonuclease activity according to the detection system, the reporter gene only expresses RFP to indicate the reference gene expression level of the reporter system, while GFP is designed outside the open reading frame (ORF) and therefore is not expressed. When the candidate has endonuclease activity, it can induce a double-strand break at the targeting site before GFP, which leads to the frameshift mutation of the reading frame when DNA is repaired through non-homology end joining (NHEJ) , resulting in GFP shifting from an out of frame state to an in frame state and beginning to express. The stronger the cleavage activity of a nuclease, the higher the proportion of GFP expressed after frameshift. Therefore, the expression intensity of GFP is positively correlated with the cleavage activity of the nuclease. The working mode of the detection system is as shown in FIG. 1.

[0209] Method for constructing plasmid 3: Through an oligo synthesis method, TAM, and a 20 nt targeted sequence with an ECoRI enzymatic cleavage site inserted at the 5’ end of the upstream sequence and a BamH1 enzymatic cleavage site inserted at the 3’ end of the downstream sequence were subjected to whole synthesis. The specific construction was as follows: 1. Preparation of vector. The plasmid backbone of an RGS-pcDNA3.1 plasmid vector was subjected to a double enzymatic cleavage digestion reaction using the single restriction endonuclease cleavage sites ECoRI and BamHI on the plasmid vector, a linearized plasmid vector fragment was obtained by agarose gel electrophoresis, and the enzymatic cleavage band was excised from the gel for recovery to obtain the purified linearized plasmid vector fragment. 2. Ligation. The nucleotide sequence of the guide reRNA obtained through gene synthesis was ligated with the linearized pUC19-U6 vector fragment using a ligation method of seamless cloning. 3. Transformation and verification. Monoclonal transformants were obtained through a LB agar plate for screening ampicillin resistance, and the correct clone identified by sequencing was used as a candidate plasmid for later use.

[0210] Table 1 Plasmid construction related sequences

[0211] Example 2: Detection of nuclease activity

[0212] 2.1 Cell treatment:

[0213] After HEK293T cells (commercially purchased) were cultured to the logarithmic growth phase, they were trypsinized into single cells with 0.25%Trypsin (Thermo) , and added to a 96-well cell culture plate pre-coated with PDL (Sigma) at a cell concentration of 3 × 104 cells / well, and cultured overnight at 37℃ in 5%CO2.

[0214] 2.2 Cell transfection:

[0215] The three functional plasmids described in example 1 (the nuclease plasmid, the reRNA-targeted sequence plasmid and the RGS dual fluorescence reporter system plasmid) were co-transfected into HEK293T cells, wherein 60 ng of the nuclease plasmid, 40 ng of the reRNA-targeted sequence plasmid and 100 ng of the RGS dual fluorescence reporter system plasmid were added to a 96-well cell culture plate, respectively, and transfection was performed using lipofectamineTM 2000 (Invitrogen, Cat. No. 11668019) at a ratio of transfection reagent volume (μL) : plasmid mass (μg) of 2 : 1.

[0216] 2.3 Obtaining results

[0217] After transfection, the cells were cultured for 48 h, then typsinized and collected, and detected by a flow cytometry. The final screening results were analyzied on the basis of the positive expression of GFP.

[0218] 2.4 Detection results

[0219] The results of nuclease activity were obtained by flow cytometry, and the statistical results of the GFP expressions reflecting the activities of all nucleases were as shown in FIG. 2 and Table 2. The results showed that the 86 nucleases (TP_O_18, TP_N_34, TP_K_36, TP_K_57, TP_N_9, TP_N_10, TP_N_11, TP_N_12, TP_N_14, TP_N_15, TP_N_18, TP_N_22, TP_N_28, TP_N_30, TP_N_31, TP_N_35, TP_N_36, TP_N_37, TP_N_42, TP_N_43, TP_N_45, TP_N_46, TP_N_51, TP_N_52, TP_N_54, TP_N_59, TP_O_2, TP_O_5, TP_O_6, TP_O_9, TP_O_11, TP_O_14, TP_O_16, TP_O_17, TP_O_19, TP_O_20, TP_O_21, TP_O_22, TP_O_23, TP_O_24, TP_O_25, TP_O_26, TP_O_27, TP_O_28, TP_O_30, TP_O_32, TP_O_33, TP_O_34, TP_O_35, TP_O_36, TP_O_37, TP_O_39, TP_O_40, TP_O_41, TP_O_42, TP_O_43, TP_O_44, TP_O_45, TP_O_48, TP_O_49, TP_O_50, TP_O_53, TP_O_55, TP_O_56, TP_O_57, TP_O_58, TP_O_59, TP_O_63, TP_O_66, TP_O_67, TP_O_68, TP_O_70, TP_O_71, TP_O_74, TP_O_75, TP_O_76, TP_O_77, TP_O_78, TP_O_79, TP_O_80, TP_P_4, TP_P_8, TP_P_10, TP_P_19, TP_P_47, TP_P_60) in the present application had good activity.

[0220] Meanwhile, a large number of nucleases with inactive or low cleavage activity were also found during the screening process (e.g. TP_A_24 and TP_D_44 in Table 1 of this application) . Compared with these nucleases with inactive or low activity, the cleavage activity of the 86 nucleases of the present application were markedly higher.

[0221] Table 2 The results of nuclease activity in example 2

[0222] Example 3: Detection of editing efficiency at endogenous loci

[0223] 3.1 Construction of plasmids:

[0224] The nuclease plasmid (plasmid 1) comprised a complete set of elements capable of transcribing and expressing candidate nuclease proteins, including a constitutive promoter CMV (sequence as shown in SEQ ID NO: 173) that can initiate transcription in an eukaryotic cell, a candidate nuclease sequence (as shown in Table 1) , a 5’-nuclear localization signal peptide sequence (sequence as shown in SEQ ID NO: 174) , a 3’-nuclear localization signal peptide sequence (sequence as shown in SEQ ID NO: 175) , a polyA sequence (sequence as shown in SEQ ID NO: 176) that terminates transcription, and an ampicillin resistance gene sequence (sequence as shown in SEQ ID NO: 177) . The method for constructing plasmid 1 is described in example 1.

[0225] The reRNA-targeted sequence plasmid (plasmid 4) comprises a reRNA sequence (as shown in Table 1) , with a 20 nt targeted sequence of endogenous gene inserted at the 3’ end of the reRNA sequence (as shown in Table 3) , a U6 promoter (sequence as shown in SEQ ID NO: 178) , a PBR322 replication origin (sequence as shown in SEQ ID NO: 179) , and an ampicillin resistance gene sequence (sequence as shown in SEQ ID NO: 177) . In addition, different targeted sequences of endogenous genes (as shown in Table 2) can identify different targeting genes adjacent to the TAM sequences.

[0226] Method for constructing plasmid 4:

[0227] 1. Preparation of vector. A reRNA plasmid containing BBSI-BBSI fragment (pUC19-U6 -reRNA-BbsI_BbsI) was subjected to enzymatic cleavage using BbsI, a linearized plasmid vector fragment was obtained by agarose gel electrophoresis, and the enzymatic cleavage band was excised from the gel for recovery to obtain the purified linearized plasmid vector fragment. 2. Preparation of 20nt targeted sequences of endogenous genes. Firstly, the 20 bp DNA sequence adjacent to the 3’ end of the TAM sequence was searched in the endogenous gene sequence, and then oligonucleotides of targeted sequences with BbsI excision end were synthesized through primer synthesis. Finally, a double-stranded oligonucleotide with sticky ends was synthesized by annealing bonding. 3. Ligation. The targeted sequence of endogenous gene was ligated with the linearized pUC19-U6 -reRNA-BbsI_BbsI vector fragment using T4 ligase. 4. Transformation and verification. Monoclonal transformants were obtained through a LB agar plate for screening ampicillin resistance, and the correct clone identified by sequencing was used as a candidate plasmid for later use.

[0228] 3.2 Cell treatment:

[0229] After HEK293T cells (commercially purchased) were cultured to the logarithmic growth phase, they were typsinized into single cells with 0.25%Trypsin (Thermo) , and added to a 48-well cell culture plate pre-coated with PDL (Sigma) at a cell concentration of 1 × 105 cells / well, and cultured overnight at 37℃ in 5%CO2.

[0230] 3.3 Cell transfection:

[0231] The two functional plasmids described in 3.1 (the nuclease plasmid and the reRNA-targeted sequence plasmid) were co-transfected into HEK293T cells, wherein 300 ng of the nuclease plasmid and 200 ng of the reRNA-targeted sequence plasmid were added to a 48-well cell culture plate, respectively, and transfection was performed using lipofectamineTM 2000 (Invitrogen, Cat. No. 11668019) at a ratio of transfection reagent volume (μL) : plasmid mass (μg) of 2 : 1.

[0232] 3.4 PCR amplification and NGS second generation sequencing

[0233] After transfection, the cells were cultured for 48 h, then typsinized and collected, and the genome DNA was extracted. PCR primers were designed near the targeted sequence of endogenous gene to amplify a length of about 200bp PCR product including 20nt targeted sequence. The PCR products were sequenced by the next generation sequencing.

[0234] 3.5 Detection results

[0235] The results of endogenous gene editing efficiency were as shown in FIG. 3 and Table 3.

[0236] By analyzing the sequence data generated by the next generation sequencing technology, the endogenous gene editing activity of nuclease was determined by counting the base insertions and deletions (Indel%) generated on the targeted sequence of endogenous gene. The results showed that the nucleases in this application showed good editing efficiency on different endogenous genes.

[0237] Table 3 The results of endogenous gene editing activity in example 3

[0238] Example 4: Comparison of endogenous editing efficiency among the nucleases with same TAM sequences

[0239] In this example, endogenous gene editing efficiency of the nucleases with same TAM sequences was evaluated by the method described in example 3. The related sequences of 5 control nucleases (TP_C_23, TP_I_15, TP_H_9, TP_H_24 and TP_H_11) were as shown in Table 4.

[0240] Table 4 Plasmid construction related sequences

[0241] The nuclease plasmid (plasmid 1) and the reRNA-targeted sequence plasmid (plasmid 4) described in example 3 were used in this example. The different targeted sequences of endogenous genes (as shown in Table 5 and Table 6) can identify different targeting genes adjacent to the same TAM sequences.

[0242] The endogenous gene editing efficiency of the nucleases with same TAM sequences (CCAT) were as shown in FIG. 4 and Table 5. Compared with the nuclease TP_C_23, the efficiency of the nucleases with same TAM sequences in the present application (including TP_O_5, TP_O_6, TP_O_11, TP_O_17, and TP_O_18) were markedly higher.

[0243] Meanwhile, the endogenous gene editing efficiency of the nucleases with same TAM sequences (TTTAA) were as shown in FIG. 4 and Table 6. Compared with the nucleases TP_H_9 and TP_H_24, the efficiency of the nucleases with same TAM sequences in the present application (including TP_O_55, and TP_O_56) were markedly higher.

[0244] Table 5 Endogenous gene editing activity of nucleases with same TAM sequence (CCAT) in example 4

[0245] Table 6 Endogenous gene editing activity of nucleases with same TAM sequence (TTTAA) in example 4

[0246] Example 5: Detection of the nuclease activity in rice protoplasts

[0247] In this example, nuclease activity in rice protoplasts was evaluated using a pair of synthetic YFP gene report vectors (plasmids 5 and 6) , which were constructed using the method described in example 1 (as shown in FIG. 5) .

[0248] Plasmid 5 comprising a promoter ZmUBI (SEQ ID NO: 324) , a candidate nuclease sequence (as shown in Table 1) , a NOS terminator (SEQ ID NO: 325) , a promoter OsU6 (SEQ ID NO: 326) , a reRNA sequence corresponding to a specific nuclease (as shown in Table 1) , a spacer sequence (SEQ ID NO: 327) , and a terminator (SEQ ID NO: 328) .

[0249] In plasmid 6, the YFP sequence (SEQ ID NO: 329) was segmented by spacer sequence (SEQ ID NO: 327) and the TAM sequence (as shown in Table 1) corresponding to a specific nuclease in plasmid 5. And The YFP sequence in the first half overlapped with the YFP sequence in the second half. Plasmid 6 also comprising a promoter 35S (SEQ ID NO: 330) and a terminator (SEQ ID NO: 331) .

[0250] After co-transforming plasmid 5 and plasmid 6 into rice protoplasts, once the spacer sequence in plasmid 6 is cut by nuclease, the partially overlapping fragment (derived from the middle segment of YFP) promotes DSB repair through homologous dependent DNA repair pathway, thus restoring normal YFP gene (as shown in FIG. 6) . Therefore, the cleavage activity of nuclease can be evaluated by observing the number of YFP-positive cells.

[0251] The results of YFP fluorescence were as shown in FIGs. 7-14, which showed that the 86 nucleases in the present application (TP_O_18, TP_N_34, TP_K_36, TP_K_57, TP_N_9, TP_N_10, TP_N_11, TP_N_12, TP_N_14, TP_N_15, TP_N_18, TP_N_22, TP_N_28, TP_N_30, TP_N_31, TP_N_35, TP_N_36, TP_N_37, TP_N_42, TP_N_43, TP_N_45, TP_N_46, TP_N_51, TP_N_52, TP_N_54, TP_N_59, TP_O_2, TP_O_5, TP_O_6, TP_O_9, TP_O_11, TP_O_14, TP_O_16, TP_O_17, TP_O_19, TP_O_20, TP_O_21, TP_O_22, TP_O_23, TP_O_24, TP_O_25, TP_O_26, TP_O_27, TP_O_28, TP_O_30, TP_O_32, TP_O_33, TP_O_34, TP_O_35, TP_O_36, TP_O_37, TP_O_39, TP_O_40, TP_O_41, TP_O_42, TP_O_43, TP_O_44, TP_O_45, TP_O_48, TP_O_49, TP_O_50, TP_O_53, TP_O_55, TP_O_56, TP_O_57, TP_O_58, TP_O_59, TP_O_63, TP_O_66, TP_O_67, TP_O_68, TP_O_70, TP_O_71, TP_O_74, TP_O_75, TP_O_76, TP_O_77, TP_O_78, TP_O_79, TP_O_80, TP_P_4, TP_P_8, TP_P_10, TP_P_19, TP_P_47, TP_P_60) had good cleavage activity in rice protoplasts as well.

[0252] It should be stated that the above are only the preferred examples of the present application and are not intended to limit the present application. For those of ordinary skill in the art, various modifications and changes can be made to the present application. Although the specific embodiments have been described, for the applicant or a person skilled in the art, the substitutions, modifications, changes, improvements, and substantial equivalents of the above embodiments may exist or cannot be foreseen currently. Therefore, the submitted appended claims and claims that may be modified are intended to cover all such substitutions, modifications, changes, improvements, and substantial equivalents. It is important that, as the technology evolves, many elements described herein may be replaced with equivalent elements that appear after the present application.

Claims

1.An isolated nuclease, wherein the nuclease comprises an amino acid sequence as shown in the following formula:(X1) (X2) a (X3) (X4) b (X5) (X6) c (X7) (X8) (X9) d (X10) (X11) e (X12) (X13) f (X14) (X15) g (X16) (X17) h (X18) (X19) i (X20) (X21) (X22) j (X23)wherein,a, b, c, d, e, f, g, h, i, and j are the numbers of amino acids;(X1) , (X3) , (X5) , (X7) , (X8) , (X10) , (X12) , (X14) , (X16) , (X18) , (X20) , (X21) , and (X23) are independently aliphatic amino acids;(X2) is any amino acid, and a is 2;(X4) is any amino acid, and b is 4;(X6) is any amino acid, and c is 10 or 11;(X9) is any amino acid, and d is 2;(X11) is any amino acid, and e is 2;(X13) is any amino acid, and f is 10, 11 or 12;(X15) is any amino acid, and g is 3;(X17) is any amino acid, and h is 1 or 2;(X19) is any amino acid, and i is 5; and(X22) is any amino acid, and j is 1.2.The nuclease according to claim 1, wherein the(X1) and (X5) are independently nonpolar amino acids;(X3) , (X7) , (X8) , (X10) , (X12) , (X14) , (X16) , (X18) , (X20) , (X21) , and (X23) are independently polar amino acids.3.The nuclease according to claim 1 or 2, wherein the(X1) is a nonpolar amino acid;(X3) is a positively charged amino acid;(X5) is a nonpolar amino acid;(X7) is a polar uncharged amino acid;(X8) is a polar uncharged amino acid;(X10) is a polar uncharged amino acid;(X12) is a polar uncharged amino acid;(X14) is a positively charged amino acid;(X16) is a polar uncharged amino acid;(X18) is a polar uncharged amino acid;(X20) is a positively charged amino acid;(X21) is a negatively charged amino acid; and / or(X23) is a polar uncharged amino acid.4.The nuclease according to any one of claims 1-3, wherein the(X1) is L;(X3) is K;(X5) is G;(X7) is S or T;(X8) is S or T;(X10) is C;(X12) is C;(X14) is R;(X16) is C;(X18) is C;(X20) is R;(X21) is D; and / or(X23) is N.5.An isolated nuclease, wherein the nuclease has a nuclease sequence selected from the following (i) or a variant sequence of the aforementioned nuclease having a nuclease activity in (ii) - (iv) :(i) at least one amino acid sequence as shown in any one of SEQ ID NOs: 1-86;(ii) at least one of sequences obtained by performing deletion, substitution, insertion, or mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids on the amino acid sequence as shown in any one of SEQ ID NOs: 1-86;(iii) at least one of amino acid sequences having at least 70%, 80%, 90%, 95%or 99%identity to the amino acid sequence as shown in any one of SEQ ID NOs: 1-86; and(iv) at least one of sequences obtained by further fusing the amino acid sequence as shown in any one of SEQ ID NOs: 1-86 with other sequences.6.The nuclease according to any one of claims 1-5, wherein the nuclease has a nuclease sequence selected from at least one of the following groups (1) - (6) :(1) at least one amino acid sequence as shown in any one of SEQ ID NOs: 35-54 and 56-61;(2) at least one amino acid sequence as shown in any one of SEQ ID NOs: 1 and 4-23;(3) at least one amino acid sequence as shown in any one of SEQ ID NOs: 63-80;(4) at least one amino acid sequence as shown in any one of SEQ ID NOs: 26-31;(5) at least one amino acid sequence as shown in any one of SEQ ID NOs: 81-86; and / or(6) at least one amino acid sequence as shown in any one of SEQ ID NOs: 2 and 24-25.7.The nuclease according to any one of claims 1-5, wherein the nuclease has a nuclease sequence selected from at least one of the following groups (1) - (5) :(1) at least one amino acid sequence as shown in any one of SEQ ID NOs: 35-62;(2) at least one amino acid sequence as shown in any one of SEQ ID NOs: 1-25;(3) at least one amino acid sequence as shown in any one of SEQ ID NOs: 63-80;(4) at least one amino acid sequence as shown in any one of SEQ ID NOs: 26-34; and / or(5) at least one amino acid sequence as shown in any one of SEQ ID NOs: 81-86.8.The nuclease according to any one of claims 1-7 wherein the nuclease belongs to the IS200 / IS605 family.9.The nuclease according to claim 8, wherein the nuclease belongs to the IS605, or IS1341 subfamily.10.The nuclease according to any one of claims 1-9, wherein the species sources of the nuclease include Bacteria, wherein the Bacteria include Actinobacteria or Firmicutes.11.A guide RNA, wherein the guide RNA comprises a reRNA, the reRNA comprises a nucleotide sequence as shown in any one of SEQ ID NOs: 87-172 or a variant thereof, and the guide RNA can bind to a specific nuclease.12.The guide RNA according to claim 11, wherein the reRNA comprises at least one of nucleotide sequences having at least 70%, 80%, 90%, 95%or 99%identity to the nucleotide sequence as shown in any one of SEQ ID NOs: 87-172.13.The guide RNA according to claim 11, wherein the reRNA comprises at least one of the nucleotide sequences as shown in any one of SEQ ID NOs: 87-172.14.The guide RNA according to any one of claims 11-13, wherein the guide RNA further comprises a targeted sequence that can recognize a targeting gene adjacent to a transposon-associated motif.15.The guide RNA according to claim 14, wherein the targeted sequence is of at least one of 10-50, 10-40, 10-30, or 15-25 nucleotides in length.16.The guide RNA according to claim 15, wherein the transposon-associated motif comprises a nucleotide sequence selected from at least one of the following groups (1) - (5) :(1) GGGG;(2) CCAT;(3) TTAT;(4) TTTAA; and / or(5) AGGAG.17.A nucleic acid, wherein, the nucleic acid encodes the nuclease according to any one of claims 1-10 and / or the guide RNA according to any one of claims 11-16.18.A nucleic acid construct, comprising the nucleic acid according to claim 17.19.The nucleic acid according to claim 18, wherein the nucleic acid construct comprises a promoter, wherein the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, Polyhedrin promotor, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.20.The nucleic acid according to claim 18, wherein the nucleic acid construct is modified by 5’-end capping and / or 3’-end polyadenylating, and the nucleic acid construct retains the activity of nuclease and / or guide RNA.21.The nucleic acid according to claim 18, wherein the nucleic acid construct is modified by thiophosphate bond modification, 2’-MOE (2-O- (2-methoxyethyl) ) , PNA (peptide nucleic acid) , GNA (glycerol nucleic acid) , LNA (locked nucleic acid) , GalNAc (N-acetylgalactosamine) , LNP (lipid nano particle) PNP (peptide nanoparticles) .22.A composition, wherein, the composition includes:an IS200 / IS605 family nuclease or a functional fragment thereof, or comprises a nucleic acid encoding the IS200 / IS605 family nuclease or the functional fragment thereof, and the nuclease or the functional fragment thereof has endonuclease activity; anda guide RNA, or comprises a nucleic acid encoding the guide RNA, and the guide RNA can bind to a specific nuclease.23.The composition according to claim 22, wherein the composition is selected from at least one of the following groups (1) - (87) , and any one of the following groups (1) - (87) comprises: a nuclease-related sequence and a guide RNA-related sequence,(1) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 1 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 87;(2) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 2 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 88;(3) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 3 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 89;(4) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 4 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 90;(5) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 5 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 91;(6) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 6 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 92;(7) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 7 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 93;(8) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 8 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 94;(9) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 9 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 95;(10) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 10 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 96;(11) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 11 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 97;(12) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 12 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 98;(13) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 13 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 99;(14) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 14 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 100;(15) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 15 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 101;(16) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 16 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 102;(17) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 17 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 103;(18) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 18 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 104;(19) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 19 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 105;(20) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 20 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 106;(21) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 21 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 107;(22) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 22 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 108;(23) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 23 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 109;(24) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 24 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 110;(25) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 25 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 111;(26) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 26 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 112;(27) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 27 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 113;(28) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 28 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 114;(29) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 29 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 115;(30) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 30 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 116;(31) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 31 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 117;(32) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 32 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 118;(33) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 33 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 119;(34) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 34 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 120;(35) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 35 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 121;(36) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 36 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 122;(37) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 37 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 123;(38) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 38 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 124;(39) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 39 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 125;(40) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 40 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 126;(41) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 41 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 127;(42) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 42 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 128;(43) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 43 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 129;(44) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 44 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 130;(45) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 45 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 131;(46) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 46 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 132;(47) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 47 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 133;(48) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 48 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 134;(49) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 49 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 135;(50) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 50 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 136;(51) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 51 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 137;(52) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 52 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 138;(53) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 53 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 139;(54) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 54 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 140;(55) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 55 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 141;(56) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 56 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 142;(57) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 57 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 143;(58) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 58 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 144;(59) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 59 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 145;(60) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 60 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 146;(61) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 61 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 147;(62) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 62 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 148;(63) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 63 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 149;(64) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 64 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 150;(65) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 65 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 151;(66) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 66 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 152;(67) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 67 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 153;(68) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 68 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 154;(69) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 69 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 155;(70) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 70 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 156;(71) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 71 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 157;(72) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 72 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 158;(73) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 73 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 159;(74) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 74 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 160;(75) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 75 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 161;(76) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 76 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 162;(77) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 77 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 163;(78) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 78 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 164;(79) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 79 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 165;(80) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 80 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 166;(81) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 81 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 167;(82) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 82 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 168;(83) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 83 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 169;(84) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 84 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 170;(85) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 85 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 171;(86) the nuclease-related sequence is an amino acid sequence comprising the sequence as shown in SEQ ID NO: 86 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence as shown in SEQ ID NO: 172;(87) a variant of any one of the aforementioned groups (1) - (86) ,wherein the nuclease-related sequence is the amino acid sequence of the variant of the nuclease in each group or a nucleic acid sequence encoding the variant, and the variant has a variant sequence of the aforementioned nuclease having a nuclease activity selected from the following (i) - (iii) :(i) at least one of sequences obtained by performing deletion, substitution, insertion, or mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids on the amino acid sequence of the nuclease in each group;(ii) at least one of amino acid sequences having at least 70%, 80%, 90%, 95%or 99%identity to the amino acid sequence as shown in any one of SEQ ID NOs: 1-86; and(iii) at least one sequence obtained by further fusing the amino acid sequence as shown in any one of SEQ ID NO: 1-86 with other sequences.24.The composition according to claim 23, wherein the guide RNA-related sequence further comprises a targeted sequence that can recognize a targeting gene adjacent to a transposon-associated motif.25.The composition according to claim 24, wherein the targeted sequence is of at least one of 10-50, 10-40, 10-30, or 15-25 nucleotides in length.26.The composition according to claim 24, wherein the transposon-associated motif comprises a nucleotide sequence selected from at least one of the following groups (1) - (5) :(1) GGGG;(2) CCAT;(3) TTAT;(4) TTTAA; and / or(5) AGGAG.27.A recombinant vector, wherein, the recombinant vector comprises the nucleic acid encoding the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, or the composition according to any one of claims 22-26.28.The recombinant vector according to claim 27, wherein the recombinant vector includes a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector.29.The recombinant vector according to claim 28, wherein the recombinant eukaryotic expression plasmid includes pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV.30.The recombinant vector according to claim 28, wherein the recombinant virus vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector.31.A recombinant host cell, wherein, the recombinant host cell comprises the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, or the recombinant vector according to any one of claims 27-30.32.The recombinant host cell according to claim 31, wherein the recombinant host cell includes an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell.33.The recombinant host cell according to claim 32, wherein the animal cell includes a mammalian cell; wherein the plant cell includes a monocot cell or a dicot cell.34.The recombinant host cell according to claim 33, wherein the mammalian cell includes a primary cell (e.g., a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell) , an immortalized cell line (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines) , a cancer cell line (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca) , an embryonic stem cell line (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW. 4, R1, and D3) and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof; wherein the monocot cell or the dicot cell includes rice cell, maize cell, or soybean cell.35.A method for introducing a double-strand break into a targeting gene of a host cell, wherein the method comprises: delivering the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, or the recombinant vector according to any one of claims 27-30 into a host cell.36.A method for deleting, replacing or inserting a targeting gene of a host cell, wherein the method comprises: delivering the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, or the recombinant vector according to any one of claims 27-30 into a host cell.37.A method for obtaining a host cell in which a targeting gene is deleted, replaced or inserted, wherein the method comprises: delivering the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, or the recombinant vector according to any one of claims 27-30 into a host cell.38.The method according to any one of claims 35-37, wherein the delivery method includes cationic liposome delivery, lipoid nanoparticulate delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticulate delivery, polypeptide and protein delivery, retrovirus delivery, lentivirus delivery, adenovirus delivery, adeno-associated virus delivery, electroporation, agrobacterium infection, or gene gun.39.The method according to any one of claims 35-37, wherein the host cell includes an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell.40.The method according to claim 39, wherein the animal cell includes a mammalian cell; wherein the plant cell includes a monocot cell or a dicot cell.41.The method according to claim 40, wherein the mammalian cell includes a primary cell (e.g., a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell) , an immortalized cell line (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines) , a cancer cell line (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca) , an embryonic stem cell line (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW. 4, R1, and D3) and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof; wherein the monocot cell or the dicot cell includes rice cell, maize cell, or soybean cell.42.Use of the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, the recombinant vector according to any one of claims 27-30, or the recombinant host cell according to any one of claims 31-34 for introducing a double-strand break into a targeting gene of a host cell.43.Use of the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, the recombinant vector according to any one of claims 27-30, or the recombinant host cell according to any one of claims 31-34 for deleting, replacing or inserting a targeting gene of a host cell.44.The use according to any one of claims 42-43, wherein the host cell includes an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell.45.The use according to claim 44, wherein the animal cell includes a mammalian cell; wherein the plant cell includes a monocot cell or a dicot cell.46.The use according to claim 45, wherein the mammalian cell includes a primary cell (e.g., a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell) , an immortalized cell line (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines) , a cancer cell line (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca) , an embryonic stem cell line (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW. 4, R1, and D3) and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof; wherein the monocot cell or the dicot cell includes rice cell, maize cell, or soybean cell.47.Use of the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, the recombinant vector according to any one of claims 27-30, or the recombinant host cell according to any one of claims 31-34 for preparing a drug or a preparation for gene therapy, cell therapy, genome research, and stem cell induction and post-induction differentiation.48.A kit, wherein, the kit comprises the nuclease according to any one of claims 1-10, the guide RNA according to any one of claims 11-16, the nucleic acid according to claim 17, the nucleic acid construct according to any one of claims 18-21, the composition according to any one of claims 22-26, the recombinant vector according to any one of claims 27-30, or the recombinant host cell according to any one of claims 31-34.

Citation Information

Patent Citations

  • Novel TnpB programming nuclease and application thereof

    CN116355878A

  • IS200 / IS60S transposon ISCB mutant protein and application thereof

    CN116656649A

  • TnpB editing system and application thereof

    CN117737034A

  • Tnpb-based genome editor

    WO2024017189A1