An isolated transposase and its uses

By providing isolated transposases with high transposable activity, the shortcomings of non-viral gene integration tools in the prior art are solved, and efficient large-fragment gene insertion and integration are achieved, avoiding the defects of viral integration.

CN119053694BActive Publication Date: 2025-06-24BEIJING ASTRAGENOMICS TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480001966.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-03-27
Filing Date
2024-03-26
Publication Date
2025-06-24
Estimated Expiration
2044-03-26

AI Technical Summary

Technical Problem

There is a lack of effective non-viral gene integration tools in the prior art, especially in the insertion and integration of large fragment genes, which have problems such as random carcinogenic risk, limited load size, immunogenic influence and production complexity.

Method used

An isolated transposase is provided with a transposase activity selected from a particular amino acid sequence or variant thereof for catalyzing the insertion and integration of exogenous nucleic acid fragments into the host cell genome.

Benefits of technology

Efficient gene integration is achieved, avoiding the disadvantages brought about by viral integration, providing higher transposal activity and more flexible gene editing options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119053694B_ABST
    Figure CN119053694B_ABST
Patent Text Reader

Abstract

Provided are nucleic acids and nucleic acid constructs, nucleic acid sets and nucleic acid set constructs, compositions, recombinant vectors, recombinant host cells and kits comprising transposases. Further specifically provided is a method for introducing an exogenous nucleic acid fragment into a host cell genome, a method for editing a host cell genome, and a method for obtaining a host cell containing an exogenous nucleic acid fragment in the genome. Further specifically provided are transposases, nucleic acids and nucleic acid constructs, nucleic acid sets and nucleic acid set constructs, compositions, recombinant vectors or recombinant host cells for introducing an exogenous nucleic acid fragment gene into a host cell genome or for preparing a drug or preparation for gene therapy, cell therapy, genomic research or stem cell induction and post-induction differentiation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of Chinese Patent Application No. 2023103046203, filed with the China National Intellectual Property Administration on March 27, 2023. The entire content of this Chinese patent application is hereby incorporated by reference in its entirety for all purposes. Technical field

[0003] This application relates to the field of molecular biology, and specifically to an isolated transposase and its uses. This application also specifically relates to: a nucleic acid and nucleic acid construct encoding the transposase, a nucleic acid group and nucleic acid group construct, and a composition, recombinant vector, recombinant host cell, and kit containing the transposase. This application also specifically relates to: a method for introducing an exogenous nucleic acid fragment into the genome of a host cell, a method for editing the genome of a host cell, and a method for obtaining a host cell containing an exogenous nucleic acid fragment in its genome. This application also specifically relates to the use of the transposase, the nucleic acid and nucleic acid construct, the nucleic acid group and nucleic acid group construct, the composition, the recombinant vector, or the recombinant host cell in introducing a foreign nucleic acid fragment gene into the genome of a host cell, or in the use of preparing drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post - induction differentiation. Background art

[0004] A transposon is a DNA sequence that can insert or excise within a genome, thereby transferring its own sequence or a complete copy of its own sequence within or between genomes. Transposons are mainly divided into two major categories. The ones mainly involved in this article are type II transposons (DNA transposons), which consist of terminal inverted repeats (TIRs) at both ends and a gene encoding a transposase. Transposons have a "cut - and - paste" transposition mechanism, cutting DNA from a chromosome and directly inserting it into other parts of the genome.

[0005] A transposase is a sequence - specific DNA - binding protein expressed by a DNA transposon sequence and contains catalytic domains that mediate DNA cleavage and ligation. The transposase will recognize and bind to the TIRs at both ends of the transposon to form a synaptic complex, and then remove the DNA transposon from the original site and integrate it into a new site. The transposition activity of a transposon mainly depends on the expression level and activity of the transposase. Therefore, a DNA transposon with high transposase activity is the main requirement for developing gene - editing tools based on transposon function.

[0006] The insertion and integration of large DNA fragments have important application values in the fields of gene therapy, molecular breeding of animals and plants, and industrial microbial transformation. At present, there is a lack of effective tools and systems for the insertion and integration of large DNA fragments in the industry. In recent years, the scientific community has developed some tools and methods that can insert and integrate large DNA fragments, but there are still some problems with these methods. For example, lentiviruses or retroviruses are most commonly used for integrating gene sequences in cell immunotherapy and gene therapy for genetic diseases, and several therapy products based on this have been used for the treatment of tumors and genetic diseases (Aiuti, A., Roncarolo, M.G. and Naldini, L. (2017) Gene therapy for ADA-SCID, the first marketing approval of an ex vivo gene therapy in Europe: paving the road for the next generation of advanced therapy medicinal products. [ADA-SCID gene therapy, the first ex vivo gene therapy marketing approval in Europe: paving the way for the next generation of advanced therapy medicinal products] EMBO Mol. Med. [EMBO Molecular Medicine] 9, 737–740; Aiuti, A. et al. (2009) Gene therapy for immunodeficiency due to adenosine deaminase deficiency. [Gene therapy for adenosine deaminase deficiency-induced immunodeficiency] N. Engl. J. Med. [New England Journal of Medicine] 360, 447–458). However, there are some potential application limitations in using viruses for the integration of large DNA fragments: First, the randomness of virus integration into the genome can pose a carcinogenic risk; second, the size of the foreign genes that viruses can carry is also limited, which is not conducive to the transfer of large therapeutic genes; third, the immunogenicity of viruses may affect the long-term expression of foreign therapeutic genes and subsequent administrations; fourth, the production of viruses requires the use of living cells, making the quality control and downstream processing of such products more complex and costly, and there are certain disadvantages in industrialization. Therefore, non-viral large fragment integration can avoid the various drawbacks brought by viral integration and become a valuable tool in gene therapy.

[0007] As a non-viral gene integration tool, DNA transposons can not only achieve the integration and stable expression of large fragments of foreign genes in the host genome, but also avoid negative impacts such as immunogenicity. Therefore, some transposons have been used in gene therapy. Although transposons have been proven to be widely present in various fields from prokaryotes to eukaryotes, during evolution, a large number of transposon fragments have become silent and inactivated in order to maintain genome stability. Currently, a few transposon tools with relatively high activity and value, such as SleepingBeauty (SB), PiggyBac (PB), and Tol2, are used in gene therapy research. Therefore, the discovery of more highly active transposon tools and the verification and detection of their functions can provide more, better, and more flexible options for the development of gene therapy strategies.

[0008] It should be noted that the methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0009] Based on this, in order to seek more advanced and effective non-viral gene integration tools, the present application provides an isolated transposase, wherein the transposase has a transposase sequence selected from the following (i) or a variant sequence of the foregoing transposase having transposase activity in (ii)-(iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-146; (ii) at least one of the sequences obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NOs: 1-146; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-146; and (iv) at least one of the sequences obtained by further fusing other sequences to the amino acid sequence shown in any one of SEQ ID NOs: 1-146. The transposase provided by the present application has equal or even higher transposase activity compared to the currently widely used SleepingBeauty (SB) and PiggyBac (PB), providing more or better options for the development of gene integration tools.

[0010] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence represented by the following formula:

[0011] D E(X1) a K(X2)b G(X3) c K(X4) d G

[0012] Among them, a, b, c, and d are the number of amino acids; D is aspartic acid; E is glutamic acid; K is lysine; G is glycine; (X1) is any amino acid, and a is 17, 18, or 19; (X2) is any amino acid, and b is 3, 4, or 5; (X3) is any amino acid, and c is 1; and (X4) is any amino acid, and d is 17, 18, or 19.

[0013] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence represented by the following formula:

[0014] P(X5) e Y(X6) f D

[0015] Among them, e, f are the number of amino acids; P is proline; Y is tyrosine; D is aspartic acid; (X5) is any amino acid, and e is 5; and (X6) is any amino acid, and f is 7.

[0016] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence represented by the following formula:

[0017] C(X7) g C(X8) h C(X9) i C

[0018] Among them, g, h, i are the number of amino acids; C is cysteine; (X7) is any amino acid, and g is 2, 3, or 4; (X8) is any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28; and (X9) is any amino acid, and i is 2.

[0019] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises at least two of the following amino acid sequences (1)-(3):

[0020] (1) D E(X1) a K(X2) b G(X3) c K(X4) d G;

[0021] (2) P(X5) e Y(X6)f D; or

[0022] (3)C(X7) g C(X8) h C(X9) i C,

[0023] wherein, a, b, c, d, e, f, g, h, i are the number of amino acids; D is aspartic acid; E is glutamic acid; K is lysine; G is glycine; P is proline; Y is tyrosine; D is aspartic acid; C is cysteine; (X1) is any amino acid, and a is 17, 18 or 19; (X2) is any amino acid, and b is 3, 4 or 5; (X3) is any amino acid, and c is 1; (X4) is any amino acid, and d is 17, 18 or 19; (X5) is any amino acid, and e is 5; (X6) is any amino acid, and f is 7; (X7) is any amino acid, and g is 2, 3 or 4; (X8) is any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28; and (X9) is any amino acid, and i is 2.

[0024] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises the following amino acid sequences (1)-(3):

[0025] (1)D E(X1) a K(X2) b G(X3) c K(X4) d G;

[0026] (2)P(X5) e Y(X6) f D; and

[0027] (3)C(X7) g C(X8) h C(X9) i C,

[0028] Wherein, a, b, c, d, e, f, g, h, and i are the number of amino acids; D is aspartic acid; E is glutamic acid; K is lysine; G is glycine; P is proline; Y is tyrosine; D is aspartic acid; C is cysteine; (X1) is any amino acid, and a is 17, 18, or 19; (X2) is any amino acid, and b is 3, 4, or 5; (X3) is any amino acid, and c is 1; (X4) is any amino acid, and d is 17, 18, or 19; (X5) is any amino acid, and e is 5; (X6) is any amino acid, and f is 7; (X7) is any amino acid, and g is 2, 3, or 4; (X8) is any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28; and (X9) is any amino acid, and i is 2.

[0029] According to an embodiment of the present application, a nucleic acid can be provided, wherein the nucleic acid encodes the transposase described in the present application.

[0030] According to an embodiment of the present application, a nucleic acid construct can be provided, which contains the nucleic acid according to the present application and also contains a promoter.

[0031] According to an embodiment of the present application, a nucleic acid set can be provided, which contains a 5' recognition sequence, wherein the 5' recognition sequence contains at least one of the nucleotide sequences shown in SEQ ID NO: 147 - 292.

[0032] According to an embodiment of the present application, a nucleic acid set can be provided, which contains a 3' recognition sequence, wherein the 3' recognition sequence contains at least one of the nucleotide sequences shown in SEQ ID NO: 293 - 438.

[0033] According to an embodiment of the present application, a nucleic acid set can be provided, which contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence contains any one of the nucleotide sequences shown in SEQ ID NO: 147 - 292 or its variant, the 3' recognition sequence contains any one of the nucleotide sequences shown in SEQ ID NO: 293 - 438 or its variant, and the nucleic acid set can be recognized by a specific transposase.

[0034] According to an embodiment of the present application, a nucleic acid set construct can be provided, which contains the nucleic acid set described in the present application and also contains an exogenous nucleic acid fragment.

[0035] According to an embodiment of the present application, a composition can be provided, wherein the composition comprises: a PiggyBac family transposase or a functional fragment thereof, or a nucleic acid encoding a PiggyBac family transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into the cell genome; and a nucleic acid group, wherein the nucleic acid group can be recognized by a specific transposase or a functional fragment thereof.

[0036] According to an embodiment of the present application, a recombinant vector can be provided, wherein the recombinant vector comprises a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, or the composition described in the present application.

[0037] According to an embodiment of the present application, a recombinant host cell can be provided, wherein the recombinant host cell comprises the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.

[0038] According to an embodiment of the present application, a method for introducing an exogenous nucleic acid fragment into a host cell genome can be provided, wherein the method comprises: delivering the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.

[0039] According to an embodiment of the present application, a method for editing a host cell genome can be provided, wherein the method comprises: delivering the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.

[0040] According to an embodiment of the present application, a method for obtaining a host cell having an exogenous nucleic acid fragment in its genome can be provided, wherein the method comprises: delivering the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.

[0041] According to embodiments of the present application, there can be provided the use of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in introducing an exogenous nucleic acid fragment into the genome of a host cell.

[0042] According to embodiments of the present application, there can be provided the use of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in the preparation of drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation.

[0043] According to embodiments of the present application, there can be provided a kit, wherein the kit contains the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.

[0044] It should be understood that the content described in this part is not intended to identify the key or important features of the examples of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0046] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0047] Figure 1 Shows a schematic diagram of the dual-plasmid vector in the transposon activity detection system in Example 1. Plasmid 1 is a plasmid expressing the transposase (Tn), and plasmid 2 is a transposon donor plasmid.

[0048] Figure 2Shows PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1 in 293T cells in Example 2Relative transposition efficiency results of PB04_D2, PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, PB04_F2, SB100X, and PiggyBac. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid.

[0049] Figure 3 , Figure 4 and Figure 5 is Figure 2 a partial enlarged view of.

[0050] Figure 6Shows the clone screening results of PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A4, PB03_A5, PB03_A10, PB03_B3, PB03_B12, PB03_C1, PB03_C7, PB03_C11, PB03_D10, PB03_E6, PB03_F1, PB03_F5, PB03_F6, PB03_F7, PB03_F12, PB03_G3, PB03_G4, PB03_G8, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2, PB04_D3, PB04_D4, PB04_D7, PB04_D9, PB04_D11, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F2, PB04_F3, PB04_F7, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10 and PB04_G12 in Example 3. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid.

[0051] Figure 7Shows the transposition activity detection results of PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A4, PB03_A5, PB03_A10, PB03_B3, PB03_B12, PB03_C1, PB03_C7, PB03_C11, PB03_D10, PB03_E6, PB03_F1, PB03_F5, PB03_F6, PB03_F7, PB03_F12, PB03_G3, PB03_G4, PB03_G8, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2, PB04_D3, PB04_D4, PB04_D7, PB04_D9, PB04_D11, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F2, PB04_F3, PB04_F7, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10 and PB04_G12 in Example 3.

[0052] Figure 8 and Figure 9 is Figure 7 a partial enlarged view of.

[0053] Figure 10Shows PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2 in Example 2Phylogenetic branching diagrams of PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, PB04_F2, and PiggyBac based on protein sequences.

[0054] Figure 11Shows PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2 in Example 2Protein sequence similarity results between PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, PB04_F2, and PiggyBac.

[0055] Figure 12 , Figure 13 and Figure 14 is Figure 11 a partial enlarged view of. Detailed implementation manners

[0056] Unless otherwise specified or inconsistent with the context, the terms or expressions used herein shall be read in conjunction with the entire content of this disclosure and as understood by those of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art.

[0057] In this application, the terms "nucleic acid" and "polynucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogs thereof.

[0058] In this application, the terms "polypeptide" and "peptide" are used interchangeably and refer to a polymer of amino acids of any length. Thus, polypeptides, oligopeptides, proteins, antibodies, and enzymes are all included in the definition of polypeptides.

[0059] As used in this application, a "fragment" of a sequence refers to a part of the sequence. For example, a fragment of a nucleic acid sequence refers to a part of the nucleic acid sequence, and a fragment of an amino acid sequence refers to a part of the amino acid sequence.

[0060] As used herein, a "variant" of a sequence is a polynucleotide or polypeptide that is respectively different from a reference polynucleotide or polypeptide, but retains the basic characteristics. A typical variant of a polynucleotide is different from another reference polynucleotide in the nucleic acid sequence, and the difference in this nucleic acid sequence may or may not change the amino acid sequence of the polypeptide encoded by the reference polynucleotide. A typical variant of a polypeptide is different from another reference polypeptide in the amino acid sequence. Generally speaking, the differences are limited, so the sequences of the reference polypeptide and the variant are generally very similar and identical in many regions. The variant polypeptide and the reference polypeptide may differ in the amino acid sequence by one or more substitutions, additions, deletions in any combination. The substituted or inserted amino acid residues may or may not be residues encoded by the genetic code. Variants of polynucleotides or polypeptides may be naturally occurring, such as allelic variations, or may be naturally occurring variants that are unknown. Non-naturally occurring polynucleotide and polypeptide variants can be produced by mutagenesis techniques, direct synthesis, and other recombinant methods known to those skilled in the art.

[0061] Amino acids are usually classified according to the properties of their side chains. For example, the side chain can make an amino acid a weak acid (such as amino acids D and E) or a weak base (such as amino acids K, R, and H); if the side chain is polar, the amino acid becomes a hydrophilic substance (such as amino acids L and I), or if the side chain is non-polar, the amino acid becomes a hydrophobic substance (such as amino acids S and C).

[0062] As used herein, the term "family" refers to a group of nucleic acids or proteins with relatively high structural similarity that are produced by replication and variation from the same ancestor, and usually have related or even the same functions. The "superfamily" refers to a group of nucleic acids or proteins with generally the same structure that are produced by replication and variation from the same ancestor, which belong to different families and usually have different functions.

[0063] As used herein, the term "transposase" refers to a polypeptide that catalyzes the excision of a transposon (including an exogenous nucleic acid and the transposase recognition sequences on both sides thereof) from a first nucleic acid (a vector containing a transposase recognition sequence and an exogenous nucleic acid) and integrates it into a second nucleic acid, that is, a target site (for example, genomic or extrachromosomal DNA containing a target site duplication (TSD) sequence in a cell). In some embodiments, the transposase binds to at least one terminal inverted repeat (TIR).

[0064] As used herein, the term "recognition sequence" refers to nucleic acid sequences located at both ends of a transposable element and flanking a first nucleic acid sequence that can be transposed. Among them, the recognition sequence located at the 5' end of the first nucleic acid sequence is called the 5' recognition sequence, and the recognition sequence located at the 3' end of the first nucleic acid sequence is called the 3' recognition sequence. In some embodiments, the recognition sequence contains at least one terminal inverted repeat that can bind to the transposase.

[0065] As used herein, the term "nucleic acid construct" is defined herein as a single-stranded or double-stranded nucleic acid molecule, preferably an artificially constructed nucleic acid molecule. Optionally, the nucleic acid construct further comprises one or more regulatory sequences operably linked thereto, which can direct the expression of the coding sequence in a suitable host cell under its compatible conditions. The term "expression" should be understood to include any steps involved in the production of a protein or polypeptide, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion. The term "regulatory sequence" includes all components necessary or advantageous for expressing the polypeptides / proteins of the present application. Each regulatory sequence may be native or foreign to the nucleic acid sequence encoding the protein or polypeptide. These regulatory sequences include but are not limited to leader sequences, polyadenylation [poly(A)] signal sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, the regulatory sequences should include a promoter and the start and stop signals for transcription and translation. To introduce specific restriction sites for ligating the regulatory sequences to the coding region of the nucleic acid sequence encoding the protein or polypeptide, regulatory sequences with linkers may be provided.

[0066] As used herein, the term "promoter" refers to a polynucleotide sequence that can control the transcription of a coding sequence. The promoter sequence includes specific sequences sufficient to enable RNA polymerase to recognize, bind, and initiate transcription. In addition, the promoter sequence may include sequences that optionally regulate the recognition, binding, and transcriptional initiation activities of RNA polymerase in the nucleic acid construct or nucleic acid group construct provided herein. The promoter can affect the transcription of a gene located on the same nucleic acid molecule as the promoter or a gene located on a different nucleic acid molecule from the promoter.

[0067] As used herein, the term "exogenous nucleic acid fragment" includes any gene of interest or any gene or fragment thereof that can be transposed. In some non-limiting embodiments, the exogenous nucleic acid fragment is from a different source than the terminal repeat sequences, for example, a nucleic acid sequence isolated from an organism different from the organism from which the terminal inverted repeat sequences are derived, i.e., an exogenous nucleic acid fragment relative to the terminal inverted repeat sequences. In some non-limiting embodiments, the exogenous nucleic acid fragment is from a different source than the host cell, for example, a nucleic acid sequence isolated from an organism different from the host cell, i.e., an exogenous nucleic acid fragment relative to the host cell.

[0068] As used herein, the term "host cell" includes but is not limited to animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. This term includes the progeny of the original cell into which the exogenous nucleic acid fragment has been introduced. Exemplary host cells include human embryonic kidney cells HEK293T. It should be understood that due to natural, accidental, or intentional mutations, the progeny of a single parental cell may not be identical to the original parent in terms of morphology or genomic or total DNA complement.

[0069] As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is linked. Examples of vectors include, but are not limited to, plasmids, viruses, bacteria, bacteriophages, and insertable DNA fragments. The term "plasmid" refers to a circular double-stranded DNA capable of accepting foreign nucleic acid fragments and replicating in prokaryotic or eukaryotic cells.

[0070] Transposase

[0071] The present application provides an isolated transposase, wherein the transposase has a transposase sequence selected from the following (i) or a variant sequence of the foregoing transposase having transposase activity in (ii)-(iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-146; (ii) at least one of the sequences obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NOs: 1-146; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95% or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-146; and (iv) at least one of the sequences obtained by further fusing other sequences to the amino acid sequence shown in any one of SEQ ID NOs: 1-146.

[0072] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence represented by the following formula:

[0073] D E(X1) a K(X2) b G(X3) c K(X4) d G

[0074] Wherein, a, b, c, and d are the number of amino acids; D is aspartic acid; E is glutamic acid; K is lysine; G is glycine; (X1) is any amino acid, and a is 17, 18 or 19; (X2) is any amino acid, and b is 3, 4 or 5; (X3) is any amino acid, and c is 1; and (X4) is any amino acid, and d is 17, 18 or 19.

[0075] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence represented by the following formula:

[0076] P(X5) e Y(X6) f D

[0077] Among them, e and f are the number of amino acids; P is proline; Y is tyrosine; D is aspartic acid; (X5) is any amino acid, and e is 5; and (X6) is any amino acid, and f is 7.

[0078] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises an amino acid sequence represented by the following formula:

[0079] C(X7) g C(X8) h C(X9) i C

[0080] Among them, g, h, and i are the number of amino acids; C is cysteine; (X7) is any amino acid, and g is 2, 3, or 4; (X8) is any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28; and (X9) is any amino acid, and i is 2.

[0081] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises at least two of the following amino acid sequences (1)-(3):

[0082] (1) D E(X1) a K(X2) b G(X3) c K(X4) d G;

[0083] (2) P(X5) e Y(X6) f D; or

[0084] (3) C(X7) g C(X8) h C(X9) i C,

[0085] Among them, a, b, c, d, e, f, g, h, and i are the number of amino acids; D is aspartic acid; E is glutamic acid; K is lysine; G is glycine; P is proline; Y is tyrosine; D is aspartic acid; C is cysteine; (X1) is any amino acid, and a is 17, 18, or 19; (X2) is any amino acid, and b is 3, 4, or 5; (X3) is any amino acid, and c is 1; (X4) is any amino acid, and d is 17, 18, or 19; (X5) is any amino acid, and e is 5; (X6) is any amino acid, and f is 7; (X7) is any amino acid, and g is 2, 3, or 4; (X8) is any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28; and (X9) is any amino acid, and i is 2.

[0086] According to an embodiment of the present application, an isolated transposase can be provided, wherein the transposase comprises the following amino acid sequences (1)-(3):

[0087] (1) D E(X1) a K(X2) b G(X3) c K(X4) d G;

[0088] (2) P(X5) e Y(X6) f D; and

[0089] (3) C(X7) g C(X8) h C(X9) i C,

[0090] Among them, a, b, c, d, e, f, g, h, and i are the number of amino acids; D is aspartic acid; E is glutamic acid; K is lysine; G is glycine; P is proline; Y is tyrosine; D is aspartic acid; C is cysteine; (X1) is any amino acid, and a is 17, 18, or 19; (X2) is any amino acid, and b is 3, 4, or 5; (X3) is any amino acid, and c is 1; (X4) is any amino acid, and d is 17, 18, or 19; (X5) is any amino acid, and e is 5; (X6) is any amino acid, and f is 7; (X7) is any amino acid, and g is 2, 3, or 4; (X8) is any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28; and (X9) is any amino acid, and i is 2.

[0091] In some embodiments, the transposase belongs to the PiggyBac family.

[0092] In some embodiments, the species sources of the transposase include Arthropoda, Platyhelminthes, Cnidaria, Mollusca, Annelida, or Chordata. In some embodiments, the species sources of the transposase include Insecta, Actinopteri, Amphibia, Rhabditophora, Bivalvia, Hydrozoa, Ascidiacea, Anthozoa, or Clitellata. In some embodiments, the species sources of the transposase include Aedes aegypti, Aelia acuminata, Agrypnus murinus, Anthonomus grandis, Apoderus coryli, Aporophyla lueneburgensis, Atethmia centrago, Blastobasis adustella, Bombyx mori, Calamotropha paludella, Catocala fraxini, Chrysoteuchia culmella, Ciona savignyi, Coptotermes formosanus, Coremacera marginata, Crassostrea gigas, Crassostrea virginica, Cryptotermes secundus, Diabrotica virgifera virgifera, Drosophila bipectinata, Drosophila elegans, Eubasilissa regina, Euschistus heros, Gonioctena quinquepunctata, Gymnosoma rotundatum, Heliconius melpomene, HermetiaHermetia illucens, Hesperophylax magnus, Homalodisca vitripennis, Hydra vulgaris, Hyles vespertilio, Ips nitidus, Ips typographus, Ischnura elegans, Lamprigera yunnana, Lasiommata megera, Limonius californicus, Locusta migratoria, Macaria notata, Malachius bipustulatus, Mamestra brassicae, Marasmarcha lunaedactyla, Marronus borbonicus, Melanotaenia boesemani, Mythimna impura, Nematostella vectensis, Ochropleura plecta, Ocypus olens, Orius insidiosus, Oryzias sinensis, Pachyrhynchus sulphureomaculatus, Parnassius apollo, Periplaneta americana, Philaenus spumarius, Philonthus cognatus, Pieris napi, Pissodes strobi, Platycnemis pennipes, Schistocerca americana, Schistocerca piceifrons, Schmidtea mediterranea, Sesamia nonagrioides, Sesia apiformis, Sitophilus oryzae, Solenopsis invicta, Teleogryllusoccipitalis), Timemashepardi, Timema tahoe, Vandiemenella viatica, Ypsolopha sequella, or Zophobas atratus.

[0093] According to an embodiment of the present application, a nucleic acid can be provided, wherein the nucleic acid encodes the transposase described in the present application.

[0094] According to an embodiment of the present application, a nucleic acid construct can be provided, which contains a nucleic acid encoding the transposase described in the present application. In some embodiments, the nucleic acid construct further contains a promoter. The promoter can be any suitable promoter sequence, that is, a nucleic acid sequence recognizable by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences mediating the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence having transcriptional activity in the selected host cell, including mutant, truncated, and chimeric promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0095] In some embodiments, the nucleic acid construct further contains a poly(A) sequence. Well-known poly(A) tailing signal sequences in the art and various truncated forms of poly(A) tailing signals can be used in the present application.

[0096] In some embodiments, the nucleic acid construct further contains any transcriptional termination sequence, that is, a sequence recognizable by the host cell to terminate transcription. The termination sequence is operably linked to the 3'-end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.

[0097] Optionally, the nucleic acid construct can further contain a suitable leader sequence, that is, an untranslated region of mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.

[0098] Optionally, the nucleic acid construct may further comprise a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is referred to as a proenzyme or pro-polypeptide. Pro-polypeptides are generally inactive and can be converted into mature active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.

[0099] Optionally, the nucleic acid construct may further comprise regulatory sequences that can regulate the expression of the polypeptide according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn gene expression on or off. Other examples of regulatory sequences are those that can cause gene amplification. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.

[0100] Nucleic acid construct

[0101] According to an embodiment of the present application, a nucleic acid set can be provided, which comprises a 5' recognition sequence, wherein the 5' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 147-292.

[0102] According to an embodiment of the present application, a nucleic acid set can be provided, which comprises a 3' recognition sequence, wherein the 3' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 293-438.

[0103] According to an embodiment of the present application, a nucleic acid set can be provided, which comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises any one of the nucleotide sequences shown in SEQ ID NOs: 147-292 or a variant thereof, the 3' recognition sequence comprises any one of the nucleotide sequences shown in SEQ ID NOs: 293-438 or a variant thereof, and the nucleic acid set can be recognized by a specific transposase.

[0104] In some embodiments, the 5' recognition sequence or the 3' recognition sequence comprises terminal inverted repeats, and the length of the terminal inverted repeats is at least one of 1-800 nt, 1-600 nt, 1-400 nt, 1-200 nt, 1-100 nt, 5-50 nt, 5-25 nt, or 10-20 nt.

[0105] According to an embodiment of the present application, a nucleic acid group construct can be provided. The nucleic acid group construct contains the nucleic acid group described in the present application and also contains an exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment is operably inserted into the nucleic acid group construct through a multiple cloning insertion site. There can be one or more exogenous nucleic acid fragments, which can be the same or different; a promoter can also be inserted to control the expression of the exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment includes any gene of interest or any gene that can be transposed, such as a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene. In some embodiments, the non-coding RNA gene includes various known-function RNAs and unknown-function RNAs such as rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA). In some embodiments, the natural functional protein gene includes a fluorescence-based reporter gene, a luciferase gene, and a resistance gene. In some embodiments, the artificial chimeric gene includes a chimeric antigen receptor gene. In some embodiments, the fluorescence-based reporter gene is selected from at least one of the genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase gene is selected from at least one of the genes encoding firefly luciferase and Renilla luciferase. In some embodiments, the resistance gene is selected from at least one of the genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, and bleomycin resistance.

[0106] In some embodiments, a promoter can also be inserted into the nucleic acid group construct to control the expression of the exogenous nucleic acid fragment. The promoter can be any suitable promoter sequence, that is, a nucleic acid sequence recognizable by the host cell expressing the exogenous nucleic acid fragment. The promoter sequence contains a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence with transcriptional activity in the selected host cell, including mutant, truncated, and chimeric promoters, and can be derived from a gene encoding an extracellular or intracellular protein or polypeptide homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0107] In some embodiments, the nucleic acid construct further comprises any transcriptional termination sequence (i.e., a sequence that can be recognized by a host cell to terminate transcription) to control the expression of the foreign nucleic acid fragment. Any terminator that can function in the selected host cell can be used in the present invention.

[0108] Optionally, the nucleic acid construct may further comprise a suitable leader sequence (i.e., the untranslated region of mRNA that is important for translation in the host cell) to control the expression of the foreign nucleic acid fragment. The leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.

[0109] Optionally, the nucleic acid construct may further comprise a propeptide coding region to control the expression of the foreign nucleic acid fragment, and the propeptide coding region encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or pro-polypeptide. The pro-polypeptide is usually inactive and can be converted into a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.

[0110] Optionally, the nucleic acid construct may further comprise a regulatory sequence that can regulate the expression of the foreign nucleic acid fragment according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can cause gene amplification. In these examples, the foreign nucleic acid fragment should be operably linked to the regulatory sequence.

[0111] Transposition composition

[0112] According to embodiments of the present application, a composition can be provided, wherein the composition comprises: a PiggyBac family transposase or a functional fragment thereof, or a nucleic acid encoding a PiggyBac family transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of a foreign nucleic acid fragment into the cell genome; and a nucleic acid construct, wherein the nucleic acid construct can be recognized by a specific transposase or a functional fragment thereof.

[0113] In some embodiments, the composition is selected from at least one of the following groups (1)-(147), and any one of the following groups (1)-(146) comprises: a transposase-related sequence and a nucleic acid construct.

[0114] (1) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:1 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:147, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:293;

[0115] (2) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:2 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:148, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:294;

[0116] (3) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:3 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:149, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:295;

[0117] (4) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:4 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:150, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:296;

[0118] (5) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:5 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:151, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:297;

[0119] (6) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:6 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:152, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:298;

[0120] (7) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:7 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:153, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:299;

[0121] (8) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:8 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:154, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:300;

[0122] (9) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:9 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:155, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:301;

[0123] (10) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:10 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:156, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:302;

[0124] (11) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:11 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:157, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:303;

[0125] (12) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 12 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 158, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 304;

[0126] (13) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 13 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 159, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 305;

[0127] (14) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 14 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 160, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 306;

[0128] (15) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 15 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 161, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 307;

[0129] (16) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 16 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 162, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 308;

[0130] (17) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 17 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 163, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 309;

[0131] (18) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 18 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 164, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 310;

[0132] (19) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 19 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 165, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 311;

[0133] (20) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 20 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 166, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 312;

[0134] (21) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 21 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 167, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 313;

[0135] (22) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 22 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 168, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 314;

[0136] (23) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 23 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 169, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 315;

[0137] (24) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 24 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 170, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 316;

[0138] (25) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 25 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 171, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 317;

[0139] (26) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 26 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 172, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 318;

[0140] (27) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 27 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 173, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 319;

[0141] (28) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 28 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 174, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 320;

[0142] (29) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 29 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 175, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 321;

[0143] (30) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 30 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 176, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 322;

[0144] (31) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 31 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 177, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 323;

[0145] (32) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 32 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 178, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 324;

[0146] (33) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 33 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 179, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 325;

[0147] (34) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 34 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 180, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 326;

[0148] (35) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 35 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 181, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 327;

[0149] (36) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 36 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 182, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 328;

[0150] (37) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 37 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 183, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 329;

[0151] (38) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 38 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 184, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 330;

[0152] (39) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 39 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 185, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 331;

[0153] (40) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 40 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 186, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 332;

[0154] (41) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 41 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 187, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 333;

[0155] (42) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 42 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 188, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 334;

[0156] (43) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 43 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 189, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 335;

[0157] (44) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 44 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 190, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 336;

[0158] (45) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 45 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 191, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 337;

[0159] (46) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 46 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 192, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 338;

[0160] (47) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 47 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 193, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 339;

[0161] (48) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 48 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 194, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 340;

[0162] (49) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 49 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 195, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 341;

[0163] (50) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 50 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 196, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 342;

[0164] (51) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 51 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 197, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 343;

[0165] (52) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 52 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 198, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 344;

[0166] (53) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 53 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 199, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 345;

[0167] (54) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 54 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 200, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 346;

[0168] (55) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 55 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 201, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 347;

[0169] (56) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 56 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 202, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 348;

[0170] (57) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 57 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 203, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 349;

[0171] (58) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 58 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 204, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 350;

[0172] (59) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 59 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 205, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 351;

[0173] (60) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 60 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 206, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 352;

[0174] (61) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 61 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 207, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 353;

[0175] (62) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 62 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 208, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 354;

[0176] (63) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 63 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 209, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 355;

[0177] (64) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 64 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 210, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 356;

[0178] (65) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 65 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 211, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 357;

[0179] (66) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 66 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 212, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 358;

[0180] (67) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 67 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 213, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 359;

[0181] (68) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 68 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 214, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 360;

[0182] (69) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 69 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 215, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 361;

[0183] (70) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 70 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 216, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 362;

[0184] (71) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 71 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 217, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 363;

[0185] (72) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 72 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 218, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 364;

[0186] (73) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 73 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 219, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 365;

[0187] (74) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 74 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 220, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 366;

[0188] (75) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 75 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 221, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 367;

[0189] (76) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 76 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 222, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 368;

[0190] (77) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 77 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 223, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 369;

[0191] (78) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 78 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 224, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 370;

[0192] (79) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 79 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 225, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 371;

[0193] (80) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 80 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 226, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 372;

[0194] (81) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 81 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 227, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 373;

[0195] (82) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 82 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 228, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 374;

[0196] (83) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 83 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 229, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 375;

[0197] (84) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 84 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 230, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 376;

[0198] (85) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 85 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 231, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 377;

[0199] (86) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 86 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 232, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 378;

[0200] (87) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 87 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 233, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 379;

[0201] (88) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 88 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 234, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 380;

[0202] (89) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 89 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 235, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 381;

[0203] (90) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 90 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 236, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 382;

[0204] (91) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 91 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 237, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 383;

[0205] (92) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 92 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 238, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 384;

[0206] (93) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 93 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 239, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 385;

[0207] (94) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 94 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 240, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 386;

[0208] (95) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 95 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 241, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 387;

[0209] (96) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 96 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 242, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 388;

[0210] (97) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 97 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 243, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 389;

[0211] (98) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 98 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 244, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 390;

[0212] (99) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 99 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 245, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 391;

[0213] (100) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 100 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 246, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 392;

[0214] (101) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 101 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 247, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 393;

[0215] (102) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 102 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 248, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 394;

[0216] (103) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 103 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 249, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 395;

[0217] (104) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 104 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 250, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 396;

[0218] (105) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 105 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 251, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 397;

[0219] (106) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 106 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 252, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 398;

[0220] (107) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 107 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 253, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 399;

[0221] (108) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 108 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 254, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 400;

[0222] (109) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 109 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 255, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 401;

[0223] (110) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 110 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 256, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 402;

[0224] (111) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 111 or a nucleic acid encoding said amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 257, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 403;

[0225] (112) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 112 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 258, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 404;

[0226] (113) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 113 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 259, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 405;

[0227] (114) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 114 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 260, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 406;

[0228] (115) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 115 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 261, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 407;

[0229] (116) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 116 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 262, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 408;

[0230] (117) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 117 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 263, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 409;

[0231] (118) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 118 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 264, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 410;

[0232] (119) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 119 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 265, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 411;

[0233] (120) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 120 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 266, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 412;

[0234] (121) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 121 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 267, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 413;

[0235] (122) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 122 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 268, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 414;

[0236] (123) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 123 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 269, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 415;

[0237] (124) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 124 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 270, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 416;

[0238] (125) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 125 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 271, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 417;

[0239] (126) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 126 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 272, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 418;

[0240] (127) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 127 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 273, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 419;

[0241] (128) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 128 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 274, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 420;

[0242] (129) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 129 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 275, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 421;

[0243] (130) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 130 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 276, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 422;

[0244] (131) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 131 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 277, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 423;

[0245] (132) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 132 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 278, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 424;

[0246] (133) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 133 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 279, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 425;

[0247] (134) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 134 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 280, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 426;

[0248] (135) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 135 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 281, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 427;

[0249] (136) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 136 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 282, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 428;

[0250] (137) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 137 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 283, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 429;

[0251] (138) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 138 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 284, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 430;

[0252] (139) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 139 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 285, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 431;

[0253] (140) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 140 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 286, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 432;

[0254] (141) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 141 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 287, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 433;

[0255] (142) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 142 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 288, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 434;

[0256] (143) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 143 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 289, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 435;

[0257] (144) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 144 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 290, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 436;

[0258] (145) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 145 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 291, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 437;

[0259] (146) The transposase-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 146 or a nucleic acid encoding said amino acid sequence; and the nucleic acid group comprises a 5'-recognition sequence and a 3'-recognition sequence, wherein the 5'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 292, and the 3'-recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 438; or

[0260] (147) A variant of any one of the foregoing groups (1)-(146),

[0261] Among them, the transposase-related sequence is the amino acid sequence of a variant of each group of transposases or the nucleic acid sequence encoding the variant, and the variant has a variant sequence of the foregoing transposase with transposase activity selected from the following (i)-(iii):

[0262] (i) At least one of the sequences obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in the amino acid sequence of each group of transposases;

[0263] (ii) At least one of the amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with any of the amino acid sequences shown in SEQ ID No: 1-146; and

[0264] (iii) At least one of the sequences obtained by further fusing other sequences to any of the amino acid sequences shown in SEQ ID No: 1-146.

[0265] In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a promoter. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence recognizable by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of the protein or polypeptide. The promoter can be any nucleic acid sequence having transcriptional activity in the selected host cell, including mutant, truncated, and chimeric promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL. In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a poly(A) sequence. Well-known poly(A) tailing signal sequences in the art and various truncated forms of poly(A) tailing signals can be used in the present application.

[0266] In some embodiments, the nucleic acid set further comprises exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments are operably inserted into the nucleic acid set through a multiple cloning site. There can be one or more exogenous nucleic acid fragments, which can be the same or different; a promoter can also be inserted to control the expression of the exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments include any gene of interest or any gene capable of being transposed, such as a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene. In some embodiments, the non-coding RNA genes include various known-function RNAs and unknown-function RNAs such as rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA). In some embodiments, the natural functional protein genes include fluorescence-based reporter genes, luciferase genes, and resistance genes. In some embodiments, the artificial chimeric genes include chimeric antigen receptor genes. In some embodiments, the fluorescence-based reporter genes include genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase genes include genes encoding firefly luciferase or Renilla luciferase. In some embodiments, the resistance genes include genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance. In some embodiments, the nucleic acid set can also insert a promoter to control the expression of the exogenous nucleic acid fragments. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence recognizable by the host cell for expressing the exogenous nucleic acid fragments. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of proteins or polypeptides. The promoter can be any nucleic acid sequence with transcriptional activity in the selected host cell, including mutated, truncated, and chimeric promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoters include CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0267] In some embodiments, the nucleic acid encoding the amino acid sequence and / or the nucleic acid set further comprises any transcription termination sequence to control the expression of the exogenous nucleic acid fragments, i.e., a sequence that can be recognized by the host cell to terminate transcription. Any terminator that can function in the selected host cell can be used in the present invention.

[0268] In some embodiments, the nucleic acid and / or nucleic acid set encoding the amino acid sequence further comprises any transcriptional termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3' end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.

[0269] Optionally, the nucleic acid and / or nucleic acid set encoding the amino acid sequence may further comprise a suitable leader sequence, i.e., an untranslated region of the mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.

[0270] Optionally, the nucleic acid and / or nucleic acid set encoding the amino acid sequence may further comprise a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a proenzyme or pro-polypeptide. The pro-polypeptide is usually inactive and can be converted into a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.

[0271] Optionally, the nucleic acid and / or nucleic acid set encoding the amino acid sequence may further comprise regulatory sequences that can regulate the expression of the polypeptide according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can cause gene amplification. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.

[0272] Recombinant vector, recombinant host cell and kit

[0273] According to an embodiment of the present application, a recombinant vector can be provided, wherein the recombinant vector contains a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, or the composition described in the present application. The recombinant vector can be any suitable vector. In some embodiments, the recombinant vector includes, but is not limited to, a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector. In some embodiments, the recombinant eukaryotic expression plasmid includes pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV. In some embodiments, the recombinant viral vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector. The recombinant vector of the present invention can be constructed by methods well known in the art. For example, appropriate restriction sites can be added to both ends of the nucleic acid construct of the present invention according to the restriction sites contained in the backbone vector used, and then loaded into the backbone vector.

[0274] According to embodiments of the present application, a recombinant host cell can be provided, wherein the recombinant host cell comprises the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleicidome described in the present application, the nucleicidome construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application. The recombinant host cell can be any host cell to which the transposase can be applied. In some embodiments, the recombinant host cell includes, but is not limited to, an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the animal cell includes a mammalian cell. In some embodiments, the mammalian cell includes a primary cell (such as a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell), an immortalized cell line (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell line), a cancer cell line (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), an embryonic stem cell line (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and its differentiated cells, or an induced pluripotent stem cell line and its differentiated cells.

[0275] According to embodiments of the present application, a kit can be provided, wherein the kit comprises the transposase described in the present application, a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleicidome described in the present application, the nucleicidome construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.

[0276] Methods and uses

[0277] The transposase-based large fragment gene insertion and integration tools and methods provided by this application can be applied to multiple fields such as gene and cell therapy, molecular breeding of animals and plants, and industrial microorganism transformation. Especially in the field of cell therapy, the transposon system provided by this application can be applied to cell immunotherapy (CAR-T, CAR-NK, CAR-M, etc.) for the integration of CAR sequences; in the field of gene therapy, the transposon system provided by this application can be used to insert or integrate healthy genes into the cell genome, thereby facilitating the treatment of diseases caused by gene mutations or gene defects; in molecular breeding, the transposon system provided by this application can be used as a tool for breeding many crops such as rice, corn, and wheat, and can also specifically accelerate the breeding process of animals and plants; in the transformation of industrial microorganisms, due to the defects of plasmids in gene expression such as instability and easy loss, the transposon system provided by this application can stably integrate genes into the chromosomes of microorganisms.

[0278] According to an embodiment of this application, a method for introducing an exogenous nucleic acid fragment into the genome of a host cell can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid group described in this application, the nucleic acid group construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0279] According to an embodiment of this application, a method for editing the genome of a host cell can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid group described in this application, the nucleic acid group construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0280] According to an embodiment of this application, a method for obtaining a host cell whose genome contains an exogenous nucleic acid fragment can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid group described in this application, the nucleic acid group construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0281] The method of delivering into the host cell can be any suitable method. In some embodiments, the delivery methods include, but are not limited to, cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated viral delivery, electroporation, Agrobacterium infection, or gene gun. The methods of cell transfection and culture are conventional methods in the art, and suitable transfection and culture methods can be selected according to different cell types.

[0282] The host cell can be any host cell to which the transposase can be applied. In some embodiments, the host cells include, but are not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the animal cells include mammalian cells. In some embodiments, the mammalian cells include primary cells (such as mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines), cancer cell lines (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

[0283] According to embodiments of the present application, there can be provided the use of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in introducing an exogenous nucleic acid fragment into the genome of a host cell. The host cell can be any host cell to which the transposase can be applied. In some embodiments, the host cell includes, but is not limited to, an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the animal cell includes a mammalian cell. In some embodiments, the mammalian cell includes a primary cell (such as a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, and an immune cell), an immortalized cell line (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratin / duct / cell line), a cancer cell line (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), an embryonic stem cell line (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and its differentiated cells, or an induced pluripotent stem cell line and its differentiated cells.

[0284] According to embodiments of the present application, there can be provided the use of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application in the preparation of drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation.

[0285] The various embodiments and preferred options for the present application above can be combined with each other (as long as they are not inherently contradictory to each other), and are applicable to the uses of the present application. All the various embodiments formed by such combinations are regarded as part of the present application.

[0286] Examples

[0287] The following describes exemplary embodiments of the present application with reference to the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. It should be understood that they are merely exemplary and are in no way intended to limit the scope of protection of the present application. The scope of protection of the present application is defined only by the claims. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0288] Unless otherwise specified, the reagents and instruments used in the following examples are traditional products that can be obtained commercially. Unless otherwise specified, the experiments are carried out under traditional conditions or conditions recommended by the manufacturer.

[0289] Example 1: Construction of a transposon activity detection system

[0290] We established a fluorescence-based reporter gene combined with an antibiotic selection marker detection system to verify the activity of candidate transposons. This system uses a two-plasmid vector for verification, as Figure 1 shown: Plasmid 1 is a plasmid expressing a transposase (Tn), which contains a constitutive promoter CMV (the sequence is as shown in SEQ ID NO: 499) that can initiate transcription in eukaryotic cells, the sequence of the candidate transposase (as shown in Table 1), and a poly(A) sequence (PA, the sequence is as shown in SEQ ID NO: 500) that terminates transcription; Plasmid 2 is a transposon donor plasmid, which contains a GFP gene (the sequence is as shown in SEQ ID NO: 501), a puromycin resistance selection gene (PuroR, the sequence is as shown in SEQ ID NO: 502), a promoter PGK (the sequence is as shown in SEQ ID NO: 503), P2A (the sequence is as shown in SEQ ID NO: 504), a Poly(A) element (the sequence is as shown in SEQ ID NO: 500), and transposon sequences ( Figure 1 LTF and RTF in, the sequences are shown in Table 1) that can be specifically recognized by the transposase are inserted at both ends of these sequences.

[0291] Table 1 Sequences related to plasmid construction

[0292]

[0293]

[0294]

[0295]

[0296]

[0297]

[0298]

[0299] When two plasmids are co-transfected into HEK293T cells, the transposase gene of plasmid 1 starts to transcribe and express the transposase protein. The transposase protein is then recognized and binds to the transposon recognition sequence on plasmid 2, excises all sequences including the transposon recognition sequence, GFP gene, and puromycin resistance gene from the plasmid vector, and integrates them into the cell genome. When the cells are continuously cultured in a medium containing a certain concentration of puromycin, only the cells that have undergone the transposition event survive because they contain the puromycin resistance gene in their genome. The transposition activity level of the candidate transposase is reflected by the number of surviving cells or their ability to form monoclonal cells.

[0300] DNA synthesis and plasmid construction methods:

[0301] Construction of plasmid 1: The amino acid sequence of the transposase was commissioned to Beijing Tsingke Biotech Co., Ltd. and GENERAL Biosystems (Anhui) Co., Ltd. to synthesize the corresponding DNA sequence, which was cloned into the plasmid vector pICOZ containing the CMV promoter element through the EcoRI site at the 5' end and the NotI site at the 3' end, enabling the transposase gene to be transcribed and subsequently translated into a functional protein under the control of the CMV promoter in eukaryotic cells.

[0302] Construction of plasmid 2: The transposon sequence (including terminal inverted repeats) is located on both sides of the transposase open reading frame. The left transposon fragment (LTF) contains all DNA sequences from the target site duplication (TSD) sequence at the 5' end to the sequence before the start codon of the transposase. The right transposon fragment (RTF) contains all DNA sequences from the first base after the stop codon of the transposase to the TSD sequence at the 3' end. In principle, the terminal inverted repeats recognized by the transposase are respectively included in the transposon sequences on both sides. BGI Tech Solutions (Beijing Liuhe) Co., Ltd. was commissioned to synthesize the LTF and RTF fragments, which were respectively cloned into the pMV plasmid vector containing elements such as the PGK promoter, puromycin resistance gene (PuroR), P2A, green fluorescent protein gene (GFP), and poly(A), with the LTF located upstream of the PGK promoter and the RTF located downstream of poly(A).

[0303] Plasmid 1 corresponds to plasmid 2 one by one.

[0304] Example 2: High-throughput screening of transposition activity

[0305] 2.1 Cell treatment (Day 0):

[0306] HEK293T cells (commercially purchased) stably expressing the firefly luciferase gene were established for high-throughput screening assays. After the cells were cultured to the logarithmic growth phase, they were digested into single cells with 0.25% trypsin (Thermo), and added to a 96-well cell culture plate pre-coated with PDL (Sigma) at a cell concentration of 1.0×10 4 cells / well, and cultured overnight at 5% CO2 and 37°C.

[0307] 2.2 Cell transfection (Day 1):

[0308] The two plasmids corresponding to each transposon system were mixed at a dose of 20 ng for plasmid 1 and 10 ng for plasmid 2, and then mixed with the transfection reagent Lipofectamine 2000 (Thermo) at a ratio of transfection plasmid mass (μg): transfection reagent volume (μL) = 1:2, and allowed to stand at room temperature for 15 min to form a transfection complex. The transfection complex was transferred to the cell culture plate and incubated with the cells, and two parallel tests were performed for each sample to be screened.

[0309] 2.3 Cell screening (Day 3)

[0310] 48 h after transfection, the medium was replaced with DMEM (Thermo) screening medium (this screening medium contains 2 μg / mL puromycin (Invivogen), 10% fetal bovine serum and 1% penicillin / streptomycin (Thermo)), and cultured for 4 days at 37°C and 5% CO2. Then, the cells were digested into single cells with 0.25% trypsin, diluted at a ratio of 1:5, transferred to another 96-well culture plate pre-coated with PDL, and cultured for 4 days in DMEM screening medium containing 2 μg / mL puromycin, 10% fetal bovine serum and 1% penicillin / streptomycin at 37°C.

[0311] 2.4 Detection of cell viability (Day 11)

[0312] 2.4.1 Preparation of detection reagent: Mix the luciferase assay system (Promega) and PBS at a volume ratio of 1:5. Prepare the detection reagent at a dose of 50 μL / well, and prepare 5 mL of the detection reagent for a 96-well plate.

[0313] 2.4.2 Take out the cells that have been screened with puromycin for 8 days from the incubator. After removing the medium, add the detection reagent at a dose of 50 μL / well. After incubating for 5 minutes at room temperature in the dark, perform the detection using a multifunctional microplate reader with luminescence detection function. The more cells survive after puromycin screening, the stronger the detected luminescence signal, indicating higher transposition activity of the sample.

[0314] 2.5 Statistical results

[0315] During high-throughput screening, set positive and negative controls on each plate. According to the read value of the luminescence signal detected by the microplate reader in each well, calculate the fold change of the read value of each sample relative to the average read value of the positive control (SB100X). Then divide the calculated fold change of each sample by the calculated fold change of the inactive transposase (PB03_D1) to obtain the relative transposition activity of all transposases, as Figure 2 shown in Table 2.

[0316] Table 2 Results of relative transposition activity in Example 2

[0317]

[0318]

[0319]

[0320]

[0321]

[0322] Example 3: Transposition activity assay

[0323] 3.1 Cell treatment (Day 0):

[0324] After culturing HEK293T cells (commercially purchased) to the logarithmic growth phase, digest and dissociate them into single cells with 0.25% trypsin (Thermo Fisher Scientific), and add them to a 24-well cell culture plate pre-coated with PDL (Sigma-Aldrich) at a cell concentration of 1.2×10 5 cells / well, and culture overnight at 5% CO2 and 37°C.

[0325] 3.2 Cell transfection (Day 1):

[0326] Mix the two plasmids corresponding to each transposon system at a dose of 200 ng for plasmid 1 and 100 ng for plasmid 2, and then mix them with the transfection reagent Lipofectamine 2000 (Thermo Fisher Scientific) at a ratio of transfection plasmid mass (μg): transfection reagent volume (μL) = 1:2. Let it stand at room temperature for 15 min to form a transfection complex. Transfer the transfection complex to a cell culture plate and incubate it with the cells. Two parallel tests are performed for each sample to be screened.

[0327] 3.3 Cell screening (Day 3)

[0328] After 48 h of transfection, digest the cells with 0.25% trypsin to dissociate them into single cells. Add the cells to the DMEM (Thermo Fisher Scientific) screening medium containing 2 μg / mL puromycin (Invitrogen), 10% fetal bovine serum, and 1% penicillin / streptomycin (Thermo Fisher Scientific), dilute them at a ratio of 1:2000, and transfer them to a 6-well culture plate for further culture. Continuously screen and culture with puromycin-resistant medium for 10 days, then perform clone counting to calculate the transposition activity of the transposase.

[0329] 3.4 Cell staining (Day 13)

[0330] Wash the cells screened by puromycin and cultured in a 6-well plate with PBS, and then fix them with 4% paraformaldehyde at room temperature for 15 min. Discard the waste liquid, add 0.2% methylene blue staining solution to the cells. Stain the cells at room temperature for 1 h. Wash the stained cell clones with PBS and take pictures in an imaging system (BioRad). Count the number of cell clones in each well. The results of clone screening of the transposase in HEK293T cells in this application are as Figure 6 shown. The figure shows the staining results of the surviving cell clones after puromycin resistance screening, indicating that the transposition event has occurred successfully. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid, serving as a negative control for each transposase sample.

[0331] 3.5 Statistical results

[0332] The statistical results of the transposition activity are as Figure 7 shown in and Table 3. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, and Tn- indicates transfection of only the donor plasmid. The y-axis in the figure shows the percentage of the calculated transposition activity. The specific calculation formula is as follows: Transposition efficiency (%) = number of cell clones per well / (number of cells plated per well × transfection efficiency (GFP-positive cell %)) × 100%.

[0333] Table 3 Statistical results of transposition activity in Example 3

[0334]

[0335]

[0336]

[0337]

[0338] During the implementation of all the above embodiments, two transposons, SB100X and PiggyBac, were used as positive controls to evaluate the transposition activity of the transposase of the present application. These two transposons are commercially available DNA transposons that are currently protected by patents. The sequences reported in the references by Lajos Ma′te′s et al. (Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates [Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates], Nature Genetics [Nature Genetics], 2009 41(6):753-761) and Cary, L.C. et al. (Transposon mutagenesis of baculoviruses: analysis of Trichoplusia nitransposon IFP2 insertions within the FP-locus of nuclear polyhedrosis viruses [Transposon mutagenesis of baculoviruses: analysis of Trichoplusia nitransposon IFP2 insertions within the FP-locus of nuclear polyhedrosis viruses], Virology [Virology], 1989, 172(1):156-169) were synthesized and cloned into the corresponding plasmid vectors according to the same method as in Example 1.

[0339] The statistical results of the transposition activity of the transposase of the present application are as Figure 2 and Figure 7As shown. The above results indicate that the 146 transposases of the present application (PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11,PB04_D1, PB04_D2, PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, and PB04_F2) have good transposition activity.、

[0340] Meanwhile, a large number of inactive or low transposase activities were also found during the screening process (such as PB01_A5, PB01_A7, PB01_B1, PB01_B3, PB01_B6, PB01_B8, PB01_E3, PB01_E9, PB01_F3, PB01_F12, PB02_B1, PB02_B4, PB02_B11, PB02_B12, PB02_C12, PB02_D8, PB02_E2, PB02_F2, PB02_F4, PB03_D1 in Application Form 1 of this application). Compared with these inactive or low transposase activities, 146 transposases of this application (PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3The transposition activities of PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2, PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, and PB04_F2) are significantly higher, and most of them have comparable or better transposition activities compared to SB100X and PiggyBac.

[0341] In addition, Figure 10 The evolutionary cladogram of the transposons of the PiggyBac superfamily in the present application based on protein sequences is shown. Figure 11 The protein sequence similarities (%) among the transposons of the PiggyBac superfamily in the present application are shown. The results show that these transposons cover different branches of the superfamily, and PiggyBac is also included therein.

[0342] It should be noted that the above are only preferred examples of the present application and are not intended to limit the present application. Various modifications and changes can be made to the present application by those of ordinary skill in the art. Although specific embodiments have been described, for the applicant or other persons skilled in the art, there may exist or currently be unforeseen alternatives, modifications, changes, improvements, and substantial equivalents of the above embodiments. Therefore, the appended claims submitted and the claims that may be modified are intended to cover all such alternatives, modifications, changes, improvements, and substantial equivalents. Importantly, with the evolution of technology, many elements described herein can be replaced by equivalent elements that emerge after the present application.

Claims

1. An isolated transposase, wherein the amino acid sequence of the transposase is shown in SEQ ID NO:

86.

2. The transposase according to claim 1, wherein the transposase belongs to the PiggyBac family.

3. The transposase according to claim 1, wherein the species origin of the transposase comprises Arthropoda. The transposase according to claim 3 , wherein the species origin of the transposase comprises Insecta. 5 . The transposase according to claim 4 , wherein the species source of the transposase comprises Philaenus spumarius .

6. A nucleic acid, wherein the nucleic acid encodes the transposase according to any one of claims 1 to 5. A nucleic acid construct comprising the nucleic acid according to claim 6 .

8. nucleic acid construct according to claim 7, wherein the nucleic acid construct further comprises a promoter.

9. The nucleic acid construct of claim 8, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

10. The nucleic acid construct according to claim 7, wherein the nucleic acid construct further comprises a poly(A) sequence.

11. A nucleic acid set comprising a 5' recognition sequence, wherein the 5' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO:

232.

12. A nucleic acid set comprising a 3' recognition sequence, wherein the 3' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO:

378.

13. A nucleic acid group comprising a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO: 232, the 3' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO: 378, and the nucleic acid group can be recognized by a specific transposase.

14. The nucleic acid set according to any one of claims 11-13, wherein the 5' recognition sequence or the 3' recognition sequence comprises a terminal inverted repeat sequence, and the length of the terminal inverted repeat sequence is at least one of 1-800nt, 1-600nt, 1-400nt, 1-200nt, 1-100nt, 5-50nt, 5-25nt, or 10-20nt. 15 . A nucleic acid group construct, comprising the nucleic acid group according to any one of claims 11 to 14, and further comprising an exogenous nucleic acid fragment.

16. The nucleic acid group construct according to claim 15, wherein the exogenous nucleic acid fragment is operably inserted into the nucleic acid group construct via a multiple cloning insertion site.

17. The nucleic acid group construct according to claim 16, wherein a promoter is further inserted into the nucleic acid group construct, and the promoter controls the expression of the exogenous nucleic acid fragment.

18. The nucleic acid group construct according to claim 16, wherein the exogenous nucleic acid fragment comprises a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.

19. The nucleic acid group construct according to claim 18, wherein the natural functional protein genes include a fluorescence-based reporter gene, a luciferase gene, and a resistance gene.

20. The nucleic acid construct of claim 19, wherein the fluorescence-based reporter gene is selected from at least one of genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein.

21. The nucleic acid construct according to claim 19, wherein the luciferase gene is selected from at least one of genes encoding firefly luciferase or Renilla luciferase.

22. The nucleic acid group construct according to claim 19, wherein the resistance gene is selected from at least one of genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance or bleomycin resistance.

23. The nucleic acid recombinant construct of claim 18, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.

24. The nucleic acid construct of claim 17, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

25. A composition, wherein the composition comprises: A PiggyBac family transposase or a functional fragment thereof, or a nucleic acid encoding a PiggyBac family transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into a cell genome; and A nucleic acid group, wherein the nucleic acid group can be recognized by a specific transposase or a functional fragment thereof, The composition comprises: a transposase-associated sequence and a nucleic acid group, wherein the transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 86 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 232, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:

378.

26. The composition of claim 25, wherein the set of nucleic acids further comprises a promoter.

27. The composition of claim 26, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

28. The composition of claim 25, wherein the set of nucleic acids further comprises a poly(A) sequence.

29. The composition of any one of claims 25-28, wherein the set of nucleic acids further comprises exogenous nucleic acid fragments.

30. The composition of claim 29, wherein the exogenous nucleic acid fragment is operably inserted into the nucleic acid set via a multiple cloning insertion site.

31. The composition of claim 30, wherein the exogenous nucleic acid fragment comprises a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.

32. The composition of claim 31, wherein the native functional protein gene comprises a fluorescence-based reporter gene, a luciferase gene, or a resistance gene.

33. The composition of claim 32, wherein the fluorescence-based reporter gene comprises a gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein.

34. The composition of claim 32, wherein the luciferase gene comprises a gene encoding firefly luciferase or Renilla luciferase.

35. The composition of claim 32, wherein the resistance gene comprises a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.

36. The composition of claim 31, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.

37. A recombinant vector, wherein the recombinant vector comprises a nucleic acid encoding a transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid group according to any one of claims 11-14, a nucleic acid group construct according to any one of claims 15-24, or a composition according to any one of claims 25-36.

38. The recombinant vector of claim 37, wherein the recombinant vector comprises a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector.

39. The recombinant vector of claim 38, wherein the recombinant eukaryotic expression plasmid comprises pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV.

40. The recombinant vector of claim 38, wherein the recombinant viral vector comprises a recombinant adenoviral vector, a recombinant adeno-associated viral vector, a recombinant retroviral vector, a recombinant herpes simplex viral vector, or a recombinant vaccinia viral vector.

41. A recombinant host cell, wherein the recombinant host cell comprises a transposase according to any one of claims 1-5, a nucleic acid encoding a transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid group according to any one of claims 11-14, a nucleic acid group construct according to any one of claims 15-24, a composition according to any one of claims 25-36, or a recombinant vector according to any one of claims 37-40.

42. The recombinant host cell of claim 41, wherein the recombinant host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell.

43. The recombinant host cell of claim 42, wherein the animal cell comprises a mammalian cell.

44. The recombinant host cell of claim 43, wherein the mammalian cell comprises a primary cell, an immortalized cell line, a cancer cell line, an embryonic stem cell line and cells differentiated therefrom, or an induced pluripotent stem cell line and cells differentiated therefrom.

45. The recombinant host cell of claim 44, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

46. ​​A method for introducing an exogenous nucleic acid fragment into a host cell genome for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1-5, the nucleic acid encoding the transposase according to any one of claims 1-5, the nucleic acid according to claim 6, the nucleic acid construct according to any one of claims 7-10, the nucleic acid group according to any one of claims 11-14, the nucleic acid group construct according to any one of claims 15-24, the composition according to any one of claims 25-36, or the recombinant vector according to any one of claims 37-40 is delivered into a host cell.

47. A method for editing a host cell genome for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1-5, the nucleic acid encoding the transposase according to any one of claims 1-5, the nucleic acid according to claim 6, the nucleic acid construct according to any one of claims 7-10, the nucleic acid group according to any one of claims 11-14, the nucleic acid group construct according to any one of claims 15-24, the composition according to any one of claims 25-36, or the recombinant vector according to any one of claims 37-40 is delivered into a host cell.

48. A method for obtaining a host cell containing an exogenous nucleic acid fragment in its genome for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1-5, the nucleic acid encoding the transposase according to any one of claims 1-5, the nucleic acid according to claim 6, the nucleic acid construct according to any one of claims 7-10, the nucleic acid group according to any one of claims 11-14, the nucleic acid group construct according to any one of claims 15-24, the composition according to any one of claims 25-36, or the recombinant vector according to any one of claims 37-40 is delivered into a host cell.

49. The method of any one of claims 46-48, wherein the delivery method comprises cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated virus delivery, electroporation, Agrobacterium infection, or gene gun.

50. The method of any one of claims 46-48, wherein the host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell.

51. The method of claim 50, wherein the animal cells comprise mammalian cells.

52. The method of claim 51, wherein the mammalian cells comprise primary cells, immortalized cell lines, cancer cell lines, embryonic stem cell lines and differentiated cells thereof, or induced pluripotent stem cell lines and differentiated cells thereof.

53. The method of claim 52, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

54. Use of a transposase according to any one of claims 1-5, a nucleic acid encoding a transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid group according to any one of claims 11-14, a nucleic acid group construct according to any one of claims 15-24, a composition according to any one of claims 25-36, a recombinant vector according to any one of claims 37-40, or a recombinant host cell according to any one of claims 41-45 in the preparation of a drug or agent for introducing an exogenous nucleic acid fragment into the host cell genome.

55. The use of claim 54, wherein the host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell.

56. The use according to claim 55, wherein the animal cell comprises a mammalian cell.

57. The use according to claim 56, wherein the mammalian cells comprise primary cells, immortalized cell lines, cancer cell lines, embryonic stem cell lines and differentiated cells thereof, or induced pluripotent stem cell lines and differentiated cells thereof.

58. The use according to claim 57, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

59. Use of a transposase according to any one of claims 1-5, a nucleic acid encoding a transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid group according to any one of claims 11-14, a nucleic acid group construct according to any one of claims 15-24, a composition according to any one of claims 25-36, a recombinant vector according to any one of claims 37-40, or a recombinant host cell according to any one of claims 41-45 in the preparation of a drug or preparation for gene therapy, cell therapy, or genomic research.

60. A kit, wherein the kit comprises a transposase according to any one of claims 1-5, a nucleic acid encoding a transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid group according to any one of claims 11-14, a nucleic acid group construct according to any one of claims 15-24, a composition according to any one of claims 25-36, a recombinant vector according to any one of claims 37-40, or a recombinant host cell according to any one of claims 41-45.

Citation Information

Patent Citations

  • Transposition of nucleic acid constructs into eukaryotic genomes with a transposase from amyelois

    CN112513277A