An isolated transposase and its uses

By providing highly active transposases and nucleic acid constructs, the random carcinogenicity and size restriction of existing viral tools are solved, and efficient large fragment gene integration and expression of non-virals is achieved, which improves the safety and flexibility of gene therapy.

CN119053695BActive Publication Date: 2025-08-05BEIJING ASTRAGENOMICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480001971.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-03-27
Filing Date
2024-03-12
Publication Date
2025-08-05
Estimated Expiration
2044-03-12

AI Technical Summary

Technical Problem

Existing viral gene integration tools have problems in gene therapy with random risk of carcinogenicity, limited size of exogenous genes, impact of immunogenicity and high production complexity, and lack effective large fragment gene insertion and integration tools.

Method used

An isolated transposase with high transposal activity, including specific amino acid sequences or variants thereof, is provided for nonviral gene integration, combining nucleic acid constructs and recombinant vectors to achieve efficient insertion and integration of exogenous nucleic acid fragments.

Benefits of technology

It has achieved efficient and stable integration and expression of large fragments of exogenous nucleic acids, avoided the disadvantages of viral integration, and provided a more flexible and efficient gene therapy strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119053695B_ABST
    Figure CN119053695B_ABST
Patent Text Reader

Abstract

Provided is a separated transposase and its uses. Also provided are a nucleic acid encoding the transposase, a nucleic acid construct, a nucleic acid group, a nucleic acid group construct, and a composition, a recombinant vector, a recombinant host cell, and a kit containing the transposase. Also provided are a method for introducing an exogenous nucleic acid fragment into the genome of a host cell, a method for editing the genome of a host cell, and a method for obtaining a host cell containing an exogenous nucleic acid fragment in the genome. Also provided are the uses of the transposase, the nucleic acid, the nucleic acid construct, the nucleic acid group, the nucleic acid group construct, the composition, the recombinant vector or the recombinant host cell for introducing a foreign nucleic acid fragment gene into the genome of a host cell or for preparing a drug or a preparation for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the priority of Chinese Patent Application No. 202310304787X, filed with the China National Intellectual Property Administration on March 27, 2023, and the entire content of this Chinese patent application is hereby incorporated by reference in its entirety for all purposes. Technical field

[0003] This application relates to the field of molecular biology, and specifically relates to an isolated transposase and its uses. This application also specifically relates to: a nucleic acid and nucleic acid construct encoding the transposase, a nucleic acid group and nucleic acid group construct, and a composition, recombinant vector, recombinant host cell, and kit containing the transposase. This application also specifically relates to: a method for introducing an exogenous nucleic acid fragment into the genome of a host cell, a method for editing the genome of a host cell, and a method for obtaining a host cell whose genome contains an exogenous nucleic acid fragment. This application also specifically relates to the use of the transposase, the nucleic acid and nucleic acid construct, the nucleic acid group and nucleic acid group construct, the composition, the recombinant vector, or the recombinant host cell in introducing an exogenous nucleic acid fragment gene into the genome of a host cell, or in the use of preparing drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation. Background art

[0004] A transposon is a DNA sequence that can insert or excise within a genome, thereby transferring its own sequence or a complete copy of its own sequence within or between genomes. Transposons are mainly divided into two major categories. The main type involved in this article is type II transposons (DNA transposons), which consist of terminal inverted repeat sequences (TIRs) at both ends and a gene encoding a transposase. Transposons have a "cut-and-paste" transposition mechanism, cutting DNA from a chromosome and directly inserting it into other parts of the genome.

[0005] A transposase is a sequence-specific DNA-binding protein expressed by a DNA transposon sequence and contains catalytic domains that mediate DNA cleavage and ligation. The transposase will recognize and bind to the TIRs at both ends of the transposon to form a synaptic complex, and then remove the DNA transposon from the original site and integrate it into a new site. The transposition activity of a transposon mainly depends on the expression level and activity of the transposase. Therefore, a DNA transposon with high transposase activity is the main requirement for developing gene editing tools based on transposon functions.

[0006] The insertion and integration of large DNA fragments have important application values in the fields of gene therapy, molecular breeding of animals and plants, and the modification of industrial microorganisms. At present, there is a lack of effective tools and systems for the insertion and integration of large DNA fragments in the industry. In recent years, the scientific community has developed some tools and methods capable of inserting and integrating large DNA fragments, but these methods still have some problems. For example, lentiviruses or retroviruses are most commonly used for integrating gene sequences in cellular immunotherapy and gene therapy for genetic diseases. There are already several therapy products based on this for the treatment of tumors and genetic diseases (Aiuti, A., Roncarolo, M. G. and Naldini, L. (2017) Gene therapy for ADA-SCID, the first marketing approval of an ex vivo gene therapy in Europe: paving the road for the next generation of advanced therapy medicinal products. [ADA-SCID gene therapy, the first ex vivo gene therapy marketing approval in Europe: paving the way for the next generation of advanced therapy medicinal products] EMBO Mol. Med. [EMBO Molecular Medicine] 9, 737–740; Aiuti, A. et al. (2009) Gene therapy for immunodeficiency due to adenosine deaminase deficiency. [Gene therapy for adenosine deaminase deficiency-induced immunodeficiency] N. Engl. J. Med. [New England Journal of Medicine] 360, 447–458). However, there are some potential application limitations in using viruses for the integration of large DNA fragments: First, the randomness of viral integration into the genome can pose a carcinogenic risk; second, the size of the foreign genes that a virus can carry is also limited, which is not conducive to the transfer of large therapeutic genes; third, the immunogenicity of the virus may affect the long-term expression of foreign therapeutic genes and subsequent administration; fourth, the production of the virus requires the use of living cells to complete, making the quality control and downstream processing of such products more complex and costly, and there are certain disadvantages in industrialization. Therefore, non-viral large fragment integration can avoid the various drawbacks brought by viral integration and become a valuable tool in gene therapy.

[0007] As a non-viral gene integration tool, DNA transposons can not only achieve the integration and stable expression of large fragments of foreign genes in the host genome, but also avoid negative effects such as immunogenicity. Therefore, some transposons have been used in gene therapy. Although transposons have been proven to exist widely in various fields from prokaryotes to eukaryotes, during evolution, in order to maintain genome stability, a large number of transposon fragments have become silent and inactivated. Currently, a few transposon tools with relatively high activity and value, such as SleepingBeauty (SB), PiggyBac (PB), and Tol2, are used in gene therapy research. Therefore, the discovery of more highly active transposon tools and the verification and detection of their functions can provide more, better, and more flexible options for the development of gene therapy strategies.

[0008] It should be noted that the methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0009] Based on this, in order to seek more advanced and effective non-viral gene integration tools, the present application provides an isolated transposase, wherein the transposase has a transposase sequence selected from the following (i) or a variant sequence of the foregoing transposase having transposase activity in (ii)-(iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-79; (ii) at least one of the sequences obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NOs: 1-79; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95% or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-79; and (iv) at least one of the sequences obtained by further fusing other sequences to the amino acid sequence shown in any one of SEQ ID NOs: 1-79. The transposase provided by the present application has comparable or even higher transposase activity compared with the currently widely used SleepingBeauty (SB) and PiggyBac (PB), providing more or better options for the development of gene integration tools.

[0010] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence shown in formula (I):

[0011] D E(I)

[0012] where D is aspartic acid; and E is glutamic acid.

[0013] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence shown in formula (II):

[0014] D(X1) a H(II)

[0015] where D is aspartic acid; H is histidine; a is the number of amino acids; and (X1) is any amino acid, and a is 5.

[0016] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence shown in formula (III):

[0017] P(X2)(X3)(III)

[0018] where P is proline; X2 is any amino acid; and X3 is aspartic acid or glutamic acid.

[0019] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises at least two of the amino acid sequences shown in formula (I), formula (II) and formula (III).

[0020] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequences shown in formula (I), formula (II) and formula (III).

[0021] According to an embodiment of the present application, a nucleic acid can be provided, wherein the nucleic acid encodes the transposase described in the present application.

[0022] According to an embodiment of the present application, a nucleic acid construct can be provided, which comprises the nucleic acid according to the present application and further comprises a promoter.

[0023] According to an embodiment of the present application, a nucleic acid set can be provided, wherein the 5' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO: 80-158.

[0024] According to an embodiment of the present application, a nucleic acid set can be provided, wherein the 3' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO: 159-237.

[0025] According to an embodiment of the present application, a nucleic acid set can be provided, which includes a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence includes the nucleotide sequence shown in any one of SEQ ID NO: 80-158 or a variant thereof, the 3'-recognition sequence includes the nucleotide sequence shown in any one of SEQ ID NO: 159-237 or a variant thereof, and the nucleic acid set can be recognized by a specific transposase.

[0026] According to an embodiment of the present application, a nucleic acid set construct can be provided, which includes the nucleic acid set described in the present application and also includes an exogenous nucleic acid fragment.

[0027] According to an embodiment of the present application, a composition can be provided, wherein the composition includes: a Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding the Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into the cell genome; and a nucleic acid set, wherein the nucleic acid set can be recognized by a specific transposase or a functional fragment thereof.

[0028] According to an embodiment of the present application, a recombinant vector can be provided, wherein the recombinant vector includes a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, or the composition described in the present application.

[0029] According to an embodiment of the present application, a recombinant host cell can be provided, wherein the recombinant host cell includes the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.

[0030] According to an embodiment of the present application, a method for introducing an exogenous nucleic acid fragment into the genome of a host cell can be provided, wherein the method includes: delivering the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.

[0031] According to an embodiment of the present application, a method for editing a host cell genome can be provided, wherein the method includes: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0032] According to an embodiment of the present application, a method for obtaining a host cell containing an exogenous nucleic acid fragment in its genome can be provided, wherein the method includes: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into a host cell.

[0033] According to an embodiment of the present application, uses of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in introducing an exogenous nucleic acid fragment into the genome of a host cell can be provided.

[0034] According to an embodiment of the present application, uses of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in preparing drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation can be provided.

[0035] According to an embodiment of the present application, a kit can be provided, wherein the kit contains the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.

[0036] It should be understood that the content described in this part is not intended to identify the key or important features of the examples of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings exemplarily illustrate embodiments and form part of the specification, and are used together with the textual description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0038] Figure 1 The schematic diagram of the dual-plasmid vector in the transposon activity detection system in Example 1 is shown. Plasmid 1 is the plasmid expressing the transposase (Tn), and plasmid 2 is the transposon donor plasmid.

[0039] Figure 2Shows the relative transposition efficiency results of TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, T CM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12, SB100X and PiggyBac in 293T cells compared with the inactive transposase (TCM_C_D8) in Example 2.

[0040] Figure 3Shows the clone screening results of TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F11, TCM_E_G4, TCM_E_G6, TCM_E_G7 and TCM_E_G9 in Example 3 in 293T cells. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid.

[0041] Figure 4Shows the detection results of the transposition activities of TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F11, TCM_E_G4, TCM_E_G6, TCM_E_G7 and TCM_E_G9 in 293T cells in Example 3. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid.

[0042] Figure 5Shows the phylogenetic branching diagrams of TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12 and SB100X based on protein sequences.

[0043] Figure 6Shows the protein sequence similarity results between TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12 and SB100X in Example 2. Detailed implementation manner

[0044] Unless otherwise specified or inconsistent with the context, the terms or expressions used herein shall be read in combination with the entire content of this disclosure and as understood by those of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art.

[0045] In this application, the terms "nucleic acid" and "polynucleotide" are used interchangeably and refer to polymeric forms of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogs thereof.

[0046] In this application, the terms "polypeptide" and "peptide" are used interchangeably and refer to polymers of amino acids of any length. Thus, polypeptides, oligopeptides, proteins, antibodies, and enzymes are all included in the definition of polypeptides.

[0047] As used in this application, a "fragment" of a sequence refers to a part of the sequence. For example, a fragment of a nucleic acid sequence refers to a part of the nucleic acid sequence, and a fragment of an amino acid sequence refers to a part of the amino acid sequence.

[0048] As used in this application, a "variant" of a sequence is a polynucleotide or polypeptide that is respectively different from a reference polynucleotide or polypeptide, but retains the basic characteristics. A typical variant of a polynucleotide differs from another reference polynucleotide in the nucleic acid sequence, and the difference in this nucleic acid sequence may or may not change the amino acid sequence of the polypeptide encoded by the reference polynucleotide. A typical variant of a polypeptide differs from another reference polypeptide in the amino acid sequence. Generally, the differences are limited, so the sequences of the reference polypeptide and the variant are generally very similar and identical in many regions. The variant polypeptide and the reference polypeptide may differ in the amino acid sequence by one or more substitutions, additions, deletions in any combination. The substituted or inserted amino acid residues may or may not be residues encoded by the genetic code. Variants of polynucleotides or polypeptides may be naturally occurring, such as allelic variations, or may be naturally occurring variants that are unknown. Non-naturally occurring polynucleotide and polypeptide variants can be produced by mutagenesis techniques, direct synthesis, and other recombinant methods known to those skilled in the art.

[0049] Amino acids are usually classified by the nature of their side chains. For example, the side chain can make an amino acid a weak acid (such as amino acids D and E) or a weak base (such as amino acids K, R, and H); if the side chain is polar, the amino acid becomes a hydrophilic substance (such as amino acids L and I), or if the side chain is non-polar, the amino acid becomes a hydrophobic substance (such as amino acids S and C).

[0050] As used in this application, the term "family" refers to a group of nucleic acids or proteins with relatively high structural similarity that are produced by replication and variation from the same ancestor and usually have related or even identical functions. The "superfamily" refers to a group of nucleic acids or proteins with generally the same structure that are produced by replication and variation from the same ancestor, belong to different families, and usually have different functions.

[0051] As used herein, the term "transposase" refers to a polypeptide that catalyzes the excision of a transposon (comprising an exogenous nucleic acid and the transposase recognition sequences flanking both sides thereof) from a first nucleic acid (a vector comprising a transposase recognition sequence and an exogenous nucleic acid) and the integration thereof into a second nucleic acid, i.e., a target site (e.g., genomic or extrachromosomal DNA in a cell comprising a target site duplication (TSD) sequence). In some embodiments, the transposase binds to at least one terminal inverted repeat (TIR).

[0052] As used herein, the term "recognition sequence" refers to nucleic acid sequences located at both ends of a transposable element and flanking a first nucleic acid sequence that is transposable, wherein the recognition sequence located at the 5' end of the first nucleic acid sequence is referred to as the 5' recognition sequence, and the recognition sequence located at the 3' end of the first nucleic acid sequence is referred to as the 3' recognition sequence. In some embodiments, the recognition sequence comprises at least one terminal inverted repeat that can bind to a transposase.

[0053] As used herein, the term "nucleic acid construct" is defined herein as a single-stranded or double-stranded nucleic acid molecule, preferably an artificially constructed nucleic acid molecule. Optionally, the nucleic acid construct further comprises one or more regulatory sequences operably linked thereto, and the one or more regulatory sequences can direct the expression of a coding sequence in a suitable host cell under compatible conditions. The term "expression" should be understood to include any steps involved in the production of a protein or polypeptide, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion. The term "regulatory sequence" includes all components necessary or advantageous for the expression of the polypeptides / proteins of the present application. Each regulatory sequence can be native or exogenous to the nucleic acid sequence encoding the protein or polypeptide. These regulatory sequences include but are not limited to leader sequences, polyadenylation sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, the regulatory sequences should include a promoter and the start and stop signals for transcription and translation. To introduce specific restriction sites for ligating the regulatory sequences to the coding region of the nucleic acid sequence encoding the protein or polypeptide, regulatory sequences with linkers can be provided.

[0054] As used herein, the term "promoter" refers to a polynucleotide sequence that can control the transcription of a coding sequence. The promoter sequence includes specific sequences sufficient to enable RNA polymerase to recognize, bind, and initiate transcription. In addition, the promoter sequence can include sequences that optionally regulate the recognition, binding, and transcriptional initiation activity of RNA polymerase in the nucleic acid construct or nucleic acid set construct provided in the present application. A promoter can affect the transcription of a gene located on the same nucleic acid molecule as the promoter or a gene located on a different nucleic acid molecule from the promoter.

[0055] As used herein, the term "exogenous nucleic acid fragment" includes any gene of interest or any gene or fragment thereof that is capable of being transposed. In some non-limiting embodiments, the exogenous nucleic acid fragment is from a different source than the terminal repeat sequences, for example, a nucleic acid sequence isolated from an organism different from the organism from which the terminal inverted repeat sequences are derived, i.e., it is an exogenous nucleic acid fragment relative to the terminal inverted repeat sequences. In some non-limiting embodiments, the exogenous nucleic acid fragment is from a different source than the host cell, for example, a nucleic acid sequence isolated from an organism different from the host cell, i.e., it is an exogenous nucleic acid fragment relative to the host cell.

[0056] As used herein, the term "host cell" includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. This term includes the progeny of the original cell into which an exogenous nucleic acid fragment has been introduced. Exemplary host cells include human embryonic kidney cells HEK293T. It should be understood that due to natural, accidental, or intentional mutations, the progeny of a single parental cell may not be identical to the original parent in terms of morphology or in genomic or total DNA complement.

[0057] As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is linked. Examples of vectors include, but are not limited to, plasmids, viruses, bacteria, bacteriophages, and insertable DNA fragments. The term "plasmid" refers to a circular double-stranded DNA capable of accepting an exogenous nucleic acid fragment and capable of replicating in prokaryotic or eukaryotic cells.

[0058] Transposase

[0059] The present application provides an isolated transposase, wherein the transposase has a transposase sequence selected from the following (i) or a variant sequence of the foregoing transposase having transposase activity in (ii)-(iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-79; (ii) at least one of the sequences obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NOs: 1-79; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-79; and (iv) at least one of the sequences obtained by further fusing other sequences to the amino acid sequence shown in any one of SEQ ID NOs: 1-79.

[0060] In some embodiments, the transposase has a transposase sequence selected from at least one of the following groups (1)-(8): (1) at least one of the amino acid sequences shown in any of SEQ ID NO: 11-36 and 59-79; (2) at least one of the amino acid sequences shown in any of SEQ ID NO: 3-4 and 48-55; (3) at least one of the amino acid sequences shown in any of SEQ ID NO: 5-8 and 40-44; (4) at least one of the amino acid sequences shown in any of SEQ ID NO: 1-2 and 45-47; (5) at least one of the amino acid sequences shown in any of SEQ ID NO: 9-10 and 56-58; (6) at least one of the amino acid sequences shown in any of SEQ ID NO: 38-39; (7) the amino acid sequence shown in SEQ ID NO: 37; and (8) an amino acid sequence having at least 70% identity with any of SEQ ID NO: 1-79 in the foregoing (1)-(7).

[0061] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in formula (I):

[0062] D E(I)

[0063] where D is aspartic acid; and E is glutamic acid.

[0064] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in formula (II):

[0065] D(X1) a H(II)

[0066] where D is aspartic acid; H is histidine; a is the number of amino acids; and (X1) is any amino acid, and a is 5.

[0067] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in formula (III):

[0068] P(X2)(X3)(III)

[0069] where P is proline; X2 is any amino acid; and X3 is aspartic acid or glutamic acid.

[0070] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises at least two of the amino acid sequences shown in formula (I), formula (II) and formula (III).

[0071] According to an embodiment of the present application, a separated transposase can be provided, wherein the transposase comprises the amino acid sequences shown in formula (I), formula (II) and formula (III).

[0072] In some embodiments, formula (I) and formula (II) are separated by 80 to 120 amino acids, and formula (II) and formula (III) are separated by 20 to 40 amino acids.

[0073] In some embodiments, the transposase belongs to the Tc1 / mariner superfamily.

[0074] In some embodiments, the transposase belongs to the Tc1, Tc2, Tc4, Mariner, Tigger, Pogo, Fot1, ISRm11, or m44 family.

[0075] In some embodiments, the species source of the transposase includes Arthropoda, Chordata, Cnidaria, Mollusca, or Platyhelminthes. In some embodiments, the species source of the transposase includes Insecta, Actinopterygii, Amphibia, Malacostraca, Arachnida, Chondrichthyes, Sauropsida, Bivalvia, Ascidiacea, Hydrozoa, or Turbellaria. In some embodiments, the species source of the transposase includes Aelia acuminata, Agrypnus murinus, Albula glossodonta, Amblyraja radiata, Anthonomus grandis, Astatotilapia calliptera, Blattella germanica, Bufo gargarizans, Carassius auratus, Cephalopholis sonnerati, Cerceris rybyensis, Cheilinus undulatus, Chelonia mydas, Clitarchus hookeri, Crassostrea gigas, Cromileptes altivelis, Cyprinus carpio, Danio rerio, Drosophila ananassae, Drosophila mojavensis, Epicauta chinensis, white symbol (Folsomia candida), hybrid of Formica aquilonia and Formica polyctena, three-spined stickleback (Gasterosteus aculeatus), five-spotted leaf beetle (Gonioctena quinquepunctata), common tachinid fly (Gymnosoma rotundatum), jumping ant (Harpegnathos saltator), Hemibagrus wyckioides, glassy-winged sharpshooter (Homalodisca vitripennis), common hydra (Hydra vulgaris), elegant spreadwing (Ischnura elegans), Pacific opah (Lampris incognitus), poplar hawkmoth (Laothoe populi), coelacanth (Latimeria chalumnae), bristly carabid beetle (Leistus spinibarbis), black turban (Lottia gigantea), olive tree (Olea europaea subsp. europaea), oak processionary moth (Orgyia antiqua), Hawaiian semaeostomean (Parhyale hawaiensis), Petrochromis sp.'moshi yellow') AB-2019, Philaenus spumarius, Philereme vetulata, Phytophthora ramorum, Polistes metricus, Rana pipiens, Rhinella marina, Rhodnius prolixus, Salmo salar, Schmidtea mediterranea, Seladonia tumulorum, Sesia apiformis, Sitophilus oryzae, Solenopsis invicta, Thalassophryne amazonica, Thecocarcelia acutangulata, Thymallus thymallus, Tribolium castaneum, Vandiemenella viatica, or Zaprionus bogoriensis.

[0076] According to an embodiment of the present application, a nucleic acid can be provided, wherein the nucleic acid encodes the transposase described in the present application

[0077] According to an embodiment of the present application, a nucleic acid construct can be provided, which comprises a nucleic acid encoding the transposase described in the present application. In some embodiments, the nucleic acid construct further comprises a promoter. The promoter can be any suitable promoter sequence, that is, a nucleic acid sequence recognizable by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence having transcriptional activity in the selected host cell, including mutant, truncated, and chimeric promoters, and can be derived from a gene encoding an extracellular or intracellular protein or polypeptide homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0078] In some embodiments, the nucleic acid construct further comprises a polyadenylation [poly(A)] signal sequence. Poly(A) tailing signal sequences well-known in the art and various truncated forms of poly(A) tailing signals can be used in the present application.

[0079] In some embodiments, the nucleic acid construct further comprises any transcription termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3'-end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.

[0080] Optionally, the nucleic acid construct may further comprise a suitable leader sequence, i.e., an untranslated region of the mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.

[0081] Optionally, the nucleic acid construct may further comprise a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or pro-polypeptide. The pro-polypeptide is usually inactive and can be converted into a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.

[0082] Optionally, the nucleic acid construct may further comprise regulatory sequences that can regulate the expression of the polypeptide according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can amplify genes. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.

[0083] Nucleic acid construct

[0084] According to an embodiment of the present application, a nucleic acid set can be provided, which comprises a 5'-recognition sequence, wherein the 5'-recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO: 80-158.

[0085] According to an embodiment of the present application, a nucleic acid set can be provided, which comprises a 3'-recognition sequence, wherein the 3'-recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO: 159-237.

[0086] According to embodiments of the present application, a nucleic acid set can be provided, which includes a 5'-recognition sequence and a 3'-recognition sequence. The 5'-recognition sequence includes a nucleotide sequence shown in any one of SEQ ID NO: 80-158 or a variant thereof, the 3'-recognition sequence includes a nucleotide sequence shown in any one of SEQ ID NO: 159-237 or a variant thereof, and the nucleic acid set can be recognized by a specific transposase.

[0087] In some embodiments, the 5'-recognition sequence or the 3'-recognition sequence includes terminal inverted repeats, and the length of the terminal inverted repeats is at least one of 1-1200 nt, 1-800 nt, 1-600 nt, 1-400 nt, 1-200 nt, 1-100 nt, 5-80 nt, 10-70 nt, or 20-60 nt.

[0088] In some embodiments, the 5'-recognition sequence or the 3'-recognition sequence includes terminal inverted repeats, and the length of the terminal inverted repeats is at least one of 1-800 nt, 20-700 nt, 50-600 nt, 100-500 nt, 150-400 nt, 150-300 nt, or 200-260 nt.

[0089] According to embodiments of the present application, a nucleic acid set construct can be provided. The nucleic acid set construct includes the nucleic acid set described in the present application and also includes an exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment is operably inserted into the nucleic acid set construct through a multiple cloning insertion site. The exogenous nucleic acid fragment can be one or more, and can be the same or different; a promoter can also be inserted to control the expression of the exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment includes any gene of interest or any gene that can be transposed, such as a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA (ncRNA) gene. In some embodiments, the non-coding RNA gene includes various known-function RNAs and unknown-function RNAs such as rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA). In some embodiments, the natural functional protein gene includes a fluorescence-based reporter gene, a luciferase gene, and a resistance gene. In some embodiments, the artificial chimeric gene includes a chimeric antigen receptor gene. In some embodiments, the fluorescence-based reporter gene is selected from at least one of the genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase gene is selected from at least one of the genes encoding firefly luciferase and Renilla luciferase. In some embodiments, the resistance gene is selected from at least one of the genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, and bleomycin resistance.

[0090] In some embodiments, the nucleic acid construct may also insert a promoter to control the expression of the exogenous nucleic acid fragment. The promoter can be any suitable promoter sequence, that is, a nucleic acid sequence recognizable by the host cell expressing the exogenous nucleic acid fragment. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of proteins or polypeptides. The promoter can be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutant, truncated, and hybrid promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoters include CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0091] In some embodiments, the nucleic acid construct also contains any transcriptional termination sequence (i.e., a sequence that can be recognized by the host cell to terminate transcription) to control the expression of the exogenous nucleic acid fragment. Any terminator that can function in the selected host cell can be used in the present invention.

[0092] Optionally, the nucleic acid construct may also contain a suitable leader sequence (i.e., the non-translated region of mRNA that is important for translation in the host cell) to control the expression of the exogenous nucleic acid fragment. The leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.

[0093] Optionally, the nucleic acid construct may also contain a propeptide coding region to control the expression of the exogenous nucleic acid fragment, and the propeptide coding region encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or pro-polypeptide. Pro-polypeptides are usually inactive and can be converted into mature active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.

[0094] Optionally, the nucleic acid construct may also contain regulatory sequences that can regulate the expression of the exogenous nucleic acid fragment according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimulants (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can amplify genes. In these examples, the exogenous nucleic acid fragment should be operably linked to the regulatory sequence.

[0095] Transposition composition

[0096] According to an embodiment of the present application, a composition can be provided, wherein the composition comprises: a Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding the Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into the cell genome; and a nucleic acid group, wherein the nucleic acid group can be recognized by a specific transposase or a functional fragment thereof.

[0097] In some embodiments, the composition is selected from at least one of the following groups (1)-(80), and any one of the following groups (1)-(79) comprises: a transposase-related sequence and a nucleic acid group.

[0098] (1) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:1 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:80, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:159.

[0099] (2) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:2 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:81, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:160.

[0100] (3) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:3 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:82, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:161.

[0101] (4) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:4 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:83, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:162.

[0102] (5) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:5 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:84, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:163;

[0103] (6) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:6 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:85, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:164;

[0104] (7) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:7 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:86, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:165;

[0105] (8) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:8 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:87, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:166;

[0106] (9) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:9 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:88, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:167;

[0107] (10) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:10 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:89, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:168;

[0108] (11) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 11 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 90, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 169;

[0109] (12) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 12 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 91, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 170;

[0110] (13) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 13 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 92, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 171;

[0111] (14) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 14 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 93, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 172;

[0112] (15) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 15 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 94, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 173;

[0113] (16) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 16 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 95, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 174;

[0114] (17) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 17 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 96, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 175;

[0115] (18) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 18 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 97, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 176;

[0116] (19) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 19 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 98, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 177;

[0117] (20) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 20 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 99, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 178;

[0118] (21) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 21 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 100, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 179;

[0119] (22) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 22 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 101, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 180;

[0120] (23) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 23 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 102, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 181;

[0121] (24) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 24 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 103, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 182;

[0122] (25) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 25 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 104, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 183;

[0123] (26) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 26 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 105, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 184;

[0124] (27) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 27 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 106, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 185;

[0125] (28) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 28 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 107, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 186;

[0126] (29) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 29 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 108, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 187;

[0127] (30) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 30 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 109, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 188;

[0128] (31) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 31 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 110, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 189;

[0129] (32) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 32 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 111, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 190;

[0130] (33) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 33 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 112, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 191;

[0131] (34) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 34 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 113, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 192;

[0132] (35) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 35 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 114, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 193;

[0133] (36) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 36 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 115, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 194;

[0134] (37) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 37 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 116, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 195;

[0135] (38) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 38 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 117, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 196;

[0136] (39) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 39 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 118, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 197;

[0137] (40) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 40 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 119, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 198;

[0138] (41) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 41 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 120, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 199;

[0139] (42) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 42 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 121, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 200;

[0140] (43) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 43 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 122, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 201;

[0141] (44) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 44 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 123, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 202;

[0142] (45) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 45 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 124, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 203;

[0143] (46) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 46 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 125, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 204;

[0144] (47) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 47 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 126, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 205;

[0145] (48) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 48 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 127, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 206;

[0146] (49) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 49 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 128, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 207;

[0147] (50) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 50 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 129, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 208;

[0148] (51) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 51 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 130, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 209;

[0149] (52) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 52 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 131, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 210;

[0150] (53) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 53 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 132, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 211;

[0151] (54) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 54 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 133, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 212;

[0152] (55) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 55 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 134, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 213;

[0153] (56) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 56 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 135, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 214;

[0154] (57) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 57 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 136, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 215;

[0155] (58) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 58 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 137, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 216;

[0156] (59) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 59 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 138, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 217;

[0157] (60) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 60 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 139, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 218;

[0158] (61) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 61 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 140, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 219;

[0159] (62) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 62 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 141, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 220;

[0160] (63) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 63 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 142, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 221;

[0161] (64) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 64 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 143, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 222;

[0162] (65) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 65 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 144, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 223;

[0163] (66) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 66 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 145, and the 3'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 224;

[0164] (67) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 67 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 146, and the 3'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 225;

[0165] (68) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 68 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 147, and the 3'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 226;

[0166] (69) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 69 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 148, and the 3'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 227;

[0167] (70) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 70 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 149, and the 3'-recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 228;

[0168] (71) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 71 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 150, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 229;

[0169] (72) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 72 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 151, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 230;

[0170] (73) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 73 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 152, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 231;

[0171] (74) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 74 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 153, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 232;

[0172] (75) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 75 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 154, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 233;

[0173] (76) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 76 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 155, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 234;

[0174] (77) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 77 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 156, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 235;

[0175] (78) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 78 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 157, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 236;

[0176] (79) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 79 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 158, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 237; or

[0177] (80) A variant of any one of the foregoing groups (1)-(79),

[0178] wherein the transposase-related sequence is an amino acid sequence of a variant of the transposase in each group or a nucleic acid sequence encoding said variant, and the variant has a variant sequence of the foregoing transposase having transposase activity selected from the following (i)-(iii):

[0179] (i) At least one of the sequences obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in the amino acid sequence of the transposase in each group;

[0180] (ii) at least one amino acid sequence having at least 70%, 80%, 90%, 95% or 99% identity to any of the amino acid sequences shown in SEQ ID NO: 1-79; and

[0181] (iii) at least one of the sequences obtained by further fusing other sequences to any of the amino acid sequences shown in SEQ ID NO: 1-79.

[0182] In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a promoter. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence recognizable by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of the protein or polypeptide. The promoter can be any nucleic acid sequence having transcriptional activity in the selected host cell, including mutated, truncated and chimeric promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL. In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a poly(A) sequence. The well-known poly(A) tailing signal sequences and various truncated forms of poly(A) tailing signals in the art can be used in the present application.

[0183] In some embodiments, the nucleic acid set further comprises an exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment is operably inserted into the nucleic acid construct through a multiple cloning site. There can be one or more exogenous nucleic acid fragments, which can be the same or different; a promoter can also be inserted to control the expression of the exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment includes any gene of interest or any gene capable of being transposed, such as a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA (ncRNA) gene. In some embodiments, the non-coding RNA gene includes various known-function RNAs and unknown-function RNAs such as rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA). In some embodiments, the natural functional protein gene includes a fluorescence-based reporter gene, a luciferase gene, and a resistance gene. In some embodiments, the artificial chimeric gene includes a chimeric antigen receptor gene. In some embodiments, the fluorescence-based reporter gene includes a gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase gene includes a gene encoding firefly luciferase or Renilla luciferase. In some embodiments, the resistance gene includes a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance. In some embodiments, the nucleic acid set can also insert a promoter to control the expression of the exogenous nucleic acid fragment. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence recognizable by the host cell expressing the exogenous nucleic acid fragment. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutated, truncated, and chimeric promoters, and can be derived from a gene encoding an extracellular or intracellular protein or polypeptide homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0184] In some embodiments, the nucleic acid encoding an amino acid sequence and / or the nucleic acid set further comprises any transcriptional termination sequence to control the expression of the exogenous nucleic acid fragment, i.e., a sequence recognizable by the host cell to terminate transcription. Any terminator that can function in the selected host cell can be used in the present invention.

[0185] In some embodiments, the nucleic acid and / or nucleic acid group encoding the amino acid sequence further comprises any transcriptional termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3' end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.

[0186] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may further comprise a suitable leader sequence, i.e., an untranslated region of mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.

[0187] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may further comprise a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a proenzyme or pro-polypeptide. The pro-polypeptide is usually inactive and can be converted into a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.

[0188] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may further comprise regulatory sequences that can regulate the expression of the polypeptide according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimulants (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can amplify genes. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.

[0189] Recombinant vector, recombinant host cell and kit

[0190] According to embodiments of the present application, a recombinant vector can be provided, wherein the recombinant vector contains a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, or the composition described in the present application. The recombinant vector can be any suitable vector. In some embodiments, the recombinant vector includes but is not limited to a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector. In some embodiments, the recombinant eukaryotic expression plasmid includes pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV. In some embodiments, the recombinant viral vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector. The recombinant vector of the present invention can be constructed by methods well known in the art. For example, appropriate restriction sites can be added to both ends of the nucleic acid construct of the present invention according to the restriction sites contained in the backbone vector used, and then loaded into the backbone vector.

[0191] According to an embodiment of the present application, a recombinant host cell can be provided, wherein the recombinant host cell contains the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application. The recombinant host cell can be any host cell to which the transposase can be applied. In some embodiments, the recombinant host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the animal cells include mammalian cells. In some embodiments, the mammalian cells include primary cells (such as mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratin / ductal cell lines), cancer cell lines (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

[0192] According to an embodiment of the present application, a kit can be provided, wherein the kit contains the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.

[0193] Methods and uses

[0194] The transposase-based large fragment gene insertion and integration tools and methods provided by this application can be applied to multiple fields such as gene and cell therapy, molecular breeding of animals and plants, and industrial microorganism modification. Especially in the field of cell therapy, the transposon system provided by this application can be used for the integration of CAR sequences in cell immunotherapy (CAR-T, CAR-NK, CAR-M, etc.); in the field of gene therapy, the transposon system provided by this application can be used to insert or integrate healthy genes into the cell genome, thereby facilitating the treatment of diseases caused by gene mutations or gene defects; in molecular breeding, the transposon system provided by this application can be used as a tool for breeding many crops such as rice, corn, and wheat, and can also specifically accelerate the breeding process of animals and plants; in industrial microorganism modification, due to the defects of plasmids in gene expression such as instability and easy loss, the transposon system provided by this application can stably integrate genes into the chromosomes of microorganisms.

[0195] According to an embodiment of this application, a method for introducing an exogenous nucleic acid fragment into a host cell genome can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid group described in this application, the nucleic acid group construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0196] According to an embodiment of this application, a method for editing a host cell genome can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid group described in this application, the nucleic acid group construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0197] According to an embodiment of this application, a method for obtaining a host cell whose genome contains an exogenous nucleic acid fragment can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid group described in this application, the nucleic acid group construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0198] The method of delivering into the host cell can be any suitable method. In some embodiments, the delivery methods include, but are not limited to, cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated viral delivery, electroporation, Agrobacterium infection, or gene gun. The methods of cell transfection and culture are conventional methods in the art, and suitable transfection and culture methods can be selected according to different cell types.

[0199] The host cell can be any host cell to which the transposase can be applied. In some embodiments, the host cells include, but are not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the host cells include mammalian cells. In some embodiments, the host cells include primary cells (such as mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, immune cells), immortalized cell lines (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal cell lines), cancer cell lines (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

[0200] According to embodiments of the present application, there can be provided uses of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in introducing an exogenous nucleic acid fragment into the genome of a host cell. The host cell can be any host cell to which the transposase can be applied. In some embodiments, the host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the host cell includes mammalian cells. In some embodiments, the host cell includes primary cells (such as mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, immune cells), immortalized cell lines (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal cell lines), cancer cell lines (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

[0201] According to embodiments of the present application, there can be provided uses of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in the preparation of drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation.

[0202] The various embodiments and preferred options for the present application above can be combined with each other (as long as they are not inherently contradictory to each other), and are applicable to the uses of the present application. All the various embodiments formed by such combinations are regarded as part of the present application.

[0203] Examples

[0204] The following describes exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. It should be understood that they are considered merely exemplary and are in no way intended to limit the protection scope of the present application. The protection scope of the present application is defined only by the claims. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0205] Unless otherwise specified, the reagents and instruments used in the following examples are traditional products that can be obtained commercially. Unless otherwise specified, the experiments are carried out under traditional conditions or conditions recommended by the manufacturer.

[0206] Example 1: Construction of a transposon activity detection system

[0207] We established a fluorescence-based detection system that combines a reporter gene with an antibiotic selection marker to verify the activity of candidate transposons. This system uses a two-plasmid vector for verification, as Figure 1 shown: Plasmid 1 is a plasmid expressing transposase (Tn), which contains a constitutive promoter CMV (sequence shown in SEQ ID NO: 298) that can initiate transcription in eukaryotic cells, the sequence of the candidate transposase (shown in Table 1), and a poly(A) sequence (PA, sequence shown in SEQ ID NO: 299) that terminates transcription; Plasmid 2 is a transposon donor plasmid, which contains a GFP gene (sequence shown in SEQ ID NO: 300), a puromycin resistance selection gene (PuroR, sequence shown in SEQ ID NO: 301), a promoter PGK (sequence shown in SEQ ID NO: 302), P2A (sequence shown in SEQ ID NO: 303), a poly(A) element (sequence shown in SEQ ID NO: 299), and transposon sequences ([ Figure 1 LTF and RTF in, sequences shown in Table 1) that can be specifically recognized by transposase are inserted at both ends of these sequences.

[0208] Table 1 Sequences related to plasmid construction

[0209]

[0210]

[0211]

[0212]

[0213] When two plasmids are co-transfected into HEK293T cells, the transposase gene of plasmid 1 initiates transcription to express the transposase protein. The transposase protein is then recognized and binds to the transposon recognition sequence on plasmid 2, excising all sequences including the transposon recognition sequence, GFP gene, and puromycin resistance gene from the plasmid vector and integrating them into the cell genome. When cells are continuously cultured in a medium containing a certain concentration of puromycin, only cells that have undergone transposition events survive because they contain the puromycin resistance gene in their genome. The transposition activity level of the candidate transposase is reflected by the number of surviving cells or their ability to form monoclonal cells.

[0214] DNA synthesis and plasmid construction method:

[0215] Construction of plasmid 1: The amino acid sequence of the transposase was entrusted to Beijing Tsingke Biotech Co., Ltd. and GENERAL Biosystems (Anhui) Co., Ltd. to synthesize the corresponding DNA sequence, which was cloned into the plasmid vector pICOZ containing the CMV promoter element through the EcoRI site at the 5' end and the NotI site at the 3' end, enabling the transposase gene to be transcribed and subsequently translated into a functional protein in eukaryotic cells under the control of the CMV promoter.

[0216] Construction of plasmid 2: The transposon sequence (including terminal inverted repeats (TIR)) is located on both sides of the transposase open reading frame. The left transposon fragment (LTF) contains all DNA sequences from the target site duplication (TSD) sequence at the 5' end to the sequence before the transposase start codon, and the right transposon fragment (RTF) contains all DNA sequences from the first base after the transposase stop codon to the TSD sequence at the 3' end. In principle, the terminal inverted repeats recognized by the transposase are respectively included in the transposon sequences on both sides. The LTF and RTF sequences were entrusted to BGI Tech Solutions (Beijing Liuhe) Co., Ltd. for synthesis and cloned into the pMV plasmid vector containing elements such as the PGK promoter, puromycin resistance gene (PuroR), P2A, green fluorescent protein (GFP) gene, and poly(A), with the LTF located upstream of the PGK promoter and the RTF located downstream of poly(A).

[0217] Plasmid 1 corresponds to plasmid 2 one by one.

[0218] Example 2: High-throughput screening of transposition activity

[0219] 2.1 Cell treatment (Day 0):

[0220] HEK293T cells (commercially purchased) stably expressing the firefly luciferase gene were established for high-throughput screening assays. After the cells were cultured to the logarithmic growth phase, they were digested and dissociated into single cells with 0.25% trypsin (Thermo), and added to a 96-well cell culture plate pre-coated with PDL (Sigma) at a cell concentration of 1.0×10 4 cells / well, and cultured overnight at 37°C with 5% CO2.

[0221] 2.2 Cell transfection (Day 1):

[0222] The two plasmids corresponding to each transposon system were mixed at a dose of 20 ng for plasmid 1 and 10 ng for plasmid 2, and then mixed with the transfection reagent Lipofectamine 2000 (Thermo) at a ratio of transfection plasmid mass (μg): transfection reagent volume (μL) = 1:2, and allowed to stand at room temperature for 15 min to form a transfection complex. The transfection complex was transferred to the cell culture plate and incubated with the cells, and two parallel tests were performed for each sample to be screened.

[0223] 2.3 Cell screening (Day 3)

[0224] 48 h after transfection, the medium was replaced with DMEM (Thermo) screening medium (this screening medium contains 2 μg / mL puromycin (Invivogen), 10% fetal bovine serum and 1% penicillin / streptomycin (Thermo)), and cultured at 37°C with 5% CO2 for 4 days. Then, the cells were digested into single cells with 0.25% trypsin, diluted at a ratio of 1:5, transferred to another 96-well culture plate pre-coated with PDL, and cultured in DMEM screening medium containing 2 μg / mL puromycin, 10% fetal bovine serum and 1% penicillin / streptomycin at 37°C for 4 days.

[0225] 2.4 Detection of cell viability (Day 11)

[0226] 2.4.1 Preparation of detection reagent: Mix the luciferase assay system (Promega) with PBS at a volume ratio of 1:5. Prepare the detection reagent at a dose of 50 μL / well, and prepare 5 mL of detection reagent for a 96-well plate.

[0227] 2.4.2 Remove the cells that have been screened with puromycin for 8 days from the incubator. After removing the culture medium, add the detection reagent at a dose of 50 μL / well. After incubating for 5 minutes at room temperature in the dark, perform the detection using a multifunctional microplate reader with luminescence detection function. The more cells survive after puromycin screening, the stronger the detected luminescence signal, indicating a higher transposition activity of the sample.

[0228] 2.5 Statistical results

[0229] During high-throughput screening, positive and negative controls are set on each plate. According to the read value of the luminescence signal detected by the microplate reader in each well, calculate the fold change of the read value of each sample (including the positive control) relative to the average read value of the negative control. The level of the calculated value indicates the level of transposition activity. The statistical results of the relative transposition activities of the transposase in this application compared with SB100X, PiggyBac, and inactive transposase are as Figure 2 shown in Table 2.

[0230] Table 2 Results of relative transposition activities in Example 2

[0231]

[0232]

[0233]

[0234]

[0235] Example 3: Transposition activity assay

[0236] 3.1 Cell treatment (Day 0):

[0237] After culturing HEK293T cells (commercially purchased) to the logarithmic growth phase, digest and dissociate them into single cells with 0.25% trypsin (Thermo Fisher Scientific), and add them to a 24-well cell culture plate pre-coated with PDL (Sigma-Aldrich) at a cell concentration of 1.2×10 5 cells / well, and culture overnight at 5% CO2, 37°C.

[0238] 3.2 Cell transfection (Day 1):

[0239] Mix the two plasmids corresponding to each transposon system at a dose of 200 ng for plasmid 1 and 100 ng for plasmid 2, and then mix them with the transfection reagent Lipofectamine 2000 (Thermo Fisher Scientific) at a ratio of transfection plasmid mass (μg): transfection reagent volume (μL) = 1:2, and let it stand at room temperature for 15 min to form a transfection complex. Transfer the transfection complex to the cell culture plate and incubate it with the cells. Two parallel tests are performed for each sample to be screened.

[0240] 3.3 Cell Screening (Day 3)

[0241] After 48 h of transfection, the cells were digested with 0.25% trypsin to dissociate into single cells. The cells were added to DMEM (Thermo Fisher Scientific) screening medium containing 2 μg / mL puromycin (Invitrogen), 10% fetal bovine serum, and antibiotics (1% penicillin / streptomycin, Thermo Fisher Scientific), diluted at a ratio of 1:2000, and transferred to a 6-well culture plate for continued culture. After continuous screening and culture with puromycin-resistant medium for 10 days, clone counting was performed to calculate the transposition activity of the transposase.

[0242] 3.4 Cell Staining (Day 13)

[0243] The cells screened by puromycin and cultured in a 6-well plate were washed with PBS and then fixed with 4% paraformaldehyde at room temperature for 15 min. The waste liquid was discarded, and 0.2% methylene blue staining solution was added to the cells. The cells were stained at room temperature for 1 h. The stained cell clones were washed with PBS and photographed in an imaging system (Bio-Rad). The number of cell clones in each well was counted. The results of clone screening of the transposase in HEK293T cells in this application are as Figure 3 shown. The figure shows the staining results of the surviving cell clones after puromycin resistance screening, indicating that transposition events occurred successfully. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid, serving as a negative control for each transposase sample.

[0244] 3.5 Statistical Results

[0245] The statistical results of the transposition activity are as Figure 4 shown in and Table 3. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid. The y-axis in the figure shows the percentage of the calculated transposition activity. The specific calculation formula is as follows: Transposition efficiency (%) = number of cell clones per well / (number of cells plated per well × transfection efficiency (GFP-positive cell %)) × 100%.

[0246] Table 3 Statistical Results of Transposition Activity in Example 3

[0247]

[0248]

[0249]

[0250] In the implementation process of all the above embodiments, two transposons, SB100X and PiggyBac, were used as positive controls to evaluate the transposition activity of the transposase of the present application. These two transposons are commercially available DNA transposons that are currently protected by patents. The sequences reported in the references by Lajos Ma′te′s et al. (Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates [Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates], Nature Genetics [Nature Genetics], 2009 41(6):753-761) and Cary, L.C. et al. (Transposon mutagenesis of baculoviruses: analysis of Trichoplusia nitransposon IFP2 insertions within the FP-locus of nuclear polyhedrosis viruses [Transposon mutagenesis of baculoviruses: analysis of Trichoplusia nitransposon IFP2 insertions within the FP-locus of nuclear polyhedrosis viruses], Virology [Virology], 1989, 172(1):156-169) were synthesized and cloned into the corresponding plasmid vectors according to the same method as in Example 1.

[0251] The statistical results of the transposition activity of the transposase of the present application are as Figure 2 and Figure 4As shown above. The above results indicate that the 79 transposases (TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12) of the present application have good transposition activity.

[0252] Meanwhile, a large number of inactive or low-transposition-activity transposases were also found during the screening process (such as TCM_A_B8, TCM_A_D6, TCM_A_F7, TCM_B_A8, TCM_B_B8, TCM_B_G5, TCM_C_A8, TCM_C_A11, TCM_C_B11, TCM_C_D2, TCM_C_D8, TCM_D_G3, TCM_D_G8, TCM_E_D4, TCM_E_D5, TCM_E_E4, TCM_E_F4, TCM_E_G1, TCM_E_G2, TCM_E_G3 in Application Form 1 of this application).Compared with these inactive or low-transposition-activity transposases, the transposition activities of 79 transposases (TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12) in the present application are significantly higher, and most of them have transposition activities comparable to or better than SB100X and PiggyBac.

[0253] In addition, Figure 5 The phylogenetic branching diagram of the transposons of the Tc1 / mariner superfamily in the present application based on protein sequences is shown. Figure 6 The protein sequence similarities (%) among the transposons of the Tc1 / mariner superfamily in the present application are shown. The results show that these transposons cover different branches of the superfamily, and SB100X is also included therein.

[0254] It should be noted that the above are only preferred examples of the present application and are not intended to limit the present application. Various modifications and changes can be made to the present application by those of ordinary skill in the art. Although specific embodiments have been described, for the applicant or other persons skilled in the art, there may exist or currently be unforeseen alternatives, modifications, changes, improvements, and substantial equivalents to the above embodiments. Therefore, the appended claims submitted and the claims that may be modified are intended to cover all such alternatives, modifications, changes, improvements, and substantial equivalents. Importantly, as technology evolves, many of the elements described herein can be replaced by equivalent elements that emerge after the present application.

Claims

1. An isolated transposase, wherein the amino acid sequence of the transposase is shown in SEQ ID NO:

61.

2. The transposase according to claim 1, wherein the transposase belongs to the Tc1 / mariner superfamily. The transposase according to claim 2 , wherein the transposase belongs to the Tc1 family. The transposase according to any one of claims 1 to 3, wherein the transposase is derived from the phylum Chordata. The transposase according to claim 4 , wherein the transposase is derived from a species of ray-finned fish. The transposase according to claim 5 , wherein the transposase is derived from Cromileptes altivelis.

7. A nucleic acid, wherein the nucleic acid encodes the transposase according to any one of claims 1 to 6. A nucleic acid construct comprising the nucleic acid according to claim 7 and further comprising a promoter.

9. The nucleic acid construct of claim 8, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

10. The nucleic acid construct according to claim 9, wherein the nucleic acid construct further comprises a poly(A) sequence.

11. A nucleic acid set comprising a 5' recognition sequence, wherein the 5' recognition sequence is the nucleotide sequence shown in SEQ ID NO:

140.

12. A nucleic acid set comprising a 3' recognition sequence, wherein the 3' recognition sequence is the nucleotide sequence shown in SEQ ID NO:

219.

13. A nucleic acid group comprising a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is the nucleotide sequence shown in SEQ ID NO: 140, the 3' recognition sequence is the nucleotide sequence shown in SEQ ID NO: 219, and the nucleic acid group can be recognized by a specific transposase.

14. The nucleic acid set according to any one of claims 11 to 13, wherein the 5' recognition sequence or the 3' recognition sequence comprises an inverted terminal repeat sequence. The nucleic acid set according to claim 14 , wherein the length of the terminal inverted repeat sequence is 1-1200 nt. 16 . A nucleic acid group construct comprising the nucleic acid group according to claim 11 , and further comprising an exogenous nucleic acid fragment.

17. The nucleic acid group construct according to claim 16, wherein the exogenous nucleic acid fragment is operably inserted into the nucleic acid group construct via a multiple cloning insertion site.

18. The nucleic acid group construct according to claim 17, wherein a promoter is further inserted into the nucleic acid group construct, and the promoter controls the expression of the exogenous nucleic acid fragment.

19. The nucleic acid construct according to claim 17, wherein the exogenous nucleic acid fragment comprises a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.

20. The nucleic acid group construct according to claim 19, wherein the natural functional protein genes include a fluorescence-based reporter gene, a luciferase gene, and a resistance gene.

21. The nucleic acid group construct according to claim 20, wherein the fluorescence-based reporter gene is selected from at least one of genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein or yellow fluorescent protein.

22. The nucleic acid group construct according to claim 20, wherein the luciferase gene is selected from at least one of genes encoding firefly luciferase or Renilla luciferase.

23. The nucleic acid group construct according to claim 20, wherein the resistance gene is selected from at least one of genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance or bleomycin resistance.

24. The nucleic acid recombinant construct of claim 19, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.

25. The nucleic acid group construct of claim 18, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

26. A composition, wherein the composition comprises: A Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding the Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into a cell genome; and A nucleic acid group, wherein the nucleic acid group can be recognized by a specific transposase or a functional fragment thereof; The composition comprises: a transposase-associated sequence and a nucleic acid group, wherein the transposase-associated sequence is the amino acid sequence shown in SEQ ID NO: 61 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is the nucleotide sequence shown in SEQ ID NO: 140, and the 3' recognition sequence is the nucleotide sequence shown in SEQ ID NO:

219.

27. The composition of claim 26, wherein the set of nucleic acids further comprises a promoter.

28. The composition of claim 27, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

29. The composition of claim 26, wherein the set of nucleic acids further comprises a poly(A) sequence.

30. The composition of claim 26, wherein the set of nucleic acids further comprises exogenous nucleic acid fragments.

31. The composition of claim 30, wherein the exogenous nucleic acid fragment is operably inserted into the nucleic acid set via a multiple cloning insertion site.

32. The composition of claim 31, wherein the exogenous nucleic acid fragment comprises a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.

33. The composition of claim 32, wherein the natural functional protein gene comprises a fluorescence-based reporter gene, a luciferase gene, or a resistance gene.

34. The composition of claim 33, wherein the fluorescence-based reporter gene comprises a gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein.

35. The composition of claim 33, wherein the luciferase gene comprises a gene encoding firefly luciferase or Renilla luciferase.

36. The composition of claim 33, wherein the resistance gene comprises a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.

37. The composition of claim 32, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.

38. A recombinant vector, wherein the recombinant vector comprises a nucleic acid encoding a transposase according to any one of claims 1-6, a nucleic acid according to claim 7, a nucleic acid construct according to any one of claims 8-10, a nucleic acid group according to any one of claims 11-15, a nucleic acid group construct according to any one of claims 16-25, or a composition according to any one of claims 26-37.

39. The recombinant vector according to claim 38, wherein the recombinant vector comprises a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector.

40. The recombinant vector of claim 39, wherein the recombinant eukaryotic expression plasmid comprises pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV.

41. The recombinant vector of claim 39, wherein the recombinant viral vector comprises a recombinant adenoviral vector, a recombinant adeno-associated viral vector, a recombinant retroviral vector, a recombinant herpes simplex viral vector, or a recombinant vaccinia viral vector.

42. A recombinant host cell, wherein the recombinant host cell comprises the transposase of any one of claims 1-6, a nucleic acid encoding the transposase of any one of claims 1-6, a nucleic acid of claim 7, a nucleic acid construct of any one of claims 8-10, a nucleic acid group of any one of claims 11-15, a nucleic acid group construct of any one of claims 16-25, a composition of any one of claims 26-37, or a recombinant vector of any one of claims 38-41.

43. The recombinant host cell of claim 42, wherein the recombinant host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, or a bacterial cell.

44. The recombinant host cell of claim 43, wherein the animal cell comprises a mammalian cell.

45. The recombinant host cell of claim 44, wherein the mammalian cell comprises a primary cell, an immortalized cell line, a cancer cell line, an embryonic stem cell line and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof.

46. The recombinant host cell of claim 45, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

47. A method for introducing an exogenous nucleic acid fragment into a host cell genome for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1 to 6, the nucleic acid encoding the transposase according to any one of claims 1 to 6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 15 to 16, the composition according to any one of claims 26 to 37, or the recombinant vector according to any one of claims 38 to 41 is delivered into a host cell.

48. A method for editing the genome of a host cell for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1 to 6, the nucleic acid encoding the transposase according to any one of claims 1 to 6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8 to 10, the composition according to any one of claims 26 to 37, or the recombinant vector according to any one of claims 38 to 41 is delivered into a host cell.

49. A method for obtaining a host cell comprising an exogenous nucleic acid fragment in its genome for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1 to 6, the nucleic acid encoding the transposase according to any one of claims 1 to 6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8 to 10, the composition according to any one of claims 26 to 37, or the recombinant vector according to any one of claims 38 to 41 is delivered into a host cell.

50. The method of any one of claims 47-49, wherein the delivery method comprises cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated viral delivery, electroporation, Agrobacterium infection, or a gene gun.

51. The method of any one of claims 47-49, wherein the host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, or a bacterial cell.

52. The method of claim 51, wherein the host cell comprises a mammalian cell.

53. The method of claim 52, wherein the host cell comprises a primary cell, an immortalized cell line, a cancer cell line, an embryonic stem cell line and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof.

54. The method of claim 53, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

55. Use of the transposase according to any one of claims 1-6, the nucleic acid encoding the transposase according to any one of claims 1-6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8-10, the composition according to any one of claims 26-37, the recombinant vector according to any one of claims 38-41, or the recombinant host cell according to any one of claims 42-46 in the preparation of a medicament or agent for introducing an exogenous nucleic acid fragment into the genome of a host cell.

56. The use of claim 55, wherein the host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, or a bacterial cell.

57. The use of claim 56, wherein the host cell comprises a mammalian cell.

58. The method of claim 57, wherein the host cell comprises a primary cell, an immortalized cell line, a cancer cell line, an embryonic stem cell line and cells differentiated therefrom, or an induced pluripotent stem cell line and cells differentiated therefrom.

59. The use according to claim 58, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

60. Use of the transposase according to any one of claims 1-6, the nucleic acid encoding the transposase according to any one of claims 1-6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8-10, the composition according to any one of claims 26-37, the recombinant vector according to any one of claims 38-41, or the recombinant host cell according to any one of claims 42-46 in the preparation of a medicament or preparation for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation.

61. A kit, wherein the kit comprises the transposase of any one of claims 1-6, a nucleic acid encoding the transposase of any one of claims 1-6, a nucleic acid of claim 7, a nucleic acid construct of any one of claims 8-10, a composition of any one of claims 26-37, a recombinant vector of any one of claims 38-41, or a recombinant host cell of any one of claims 42-46.