An isolated transposase AG-P3G4 and uses thereof

CN120574801BActive Publication Date: 2026-08-28BEIJING ASTRAGENOMICS TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510731095.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-03-27
Filing Date
2024-03-26
Publication Date
2026-08-28
Estimated Expiration
2044-03-26

AI Technical Summary

Technical Problem

尽管转座子已被证实广泛存在于从原核到真核的各个领域中,但是在进化过程中,为了保持基因组的稳定性,有大量的转座子片段变为沉默失活

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120574801B_ABST
    Figure CN120574801B_ABST
Patent Text Reader

Abstract

Provided are nucleic acids and nucleic acid constructs encoding transposases, nucleic acid panels and nucleic acid panel constructs, compositions, recombinant vectors, recombinant host cells, and kits comprising transposases. Further specifically provided are a method for introducing an exogenous nucleic acid fragment into a host cell genome, a method for editing a host cell genome, and a method for obtaining a host cell containing an exogenous nucleic acid fragment in a genome. Further specifically provided are transposases, nucleic acids and nucleic acid constructs, nucleic acid panels and nucleic acid panel constructs, compositions, recombinant vectors, or recombinant host cells for use in introducing an exogenous nucleic acid fragment gene into a host cell genome or in the preparation of a medicament or a preparation for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 2024800019669, entitled "An isolated transposase and its use", filed on March 26, 2024.

[0002] Cross-references to related applications

[0003] This application claims priority to Chinese Patent Application No. 2023103046203, filed with the China National Intellectual Property Administration on March 27, 2023, the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0004] This application relates to the field of molecular biology, specifically to an isolated transposase and its uses. This application also specifically relates to: a nucleic acid encoding the transposase and a nucleic acid construct, a nucleome and a nucleome construct, and a composition comprising the transposase, a recombinant vector, a recombinant host cell, and a kit. This application further specifically relates to: a method for introducing a foreign nucleic acid fragment into the genome of a host cell, a method for editing the genome of a host cell, and a method for obtaining a host cell whose genome contains a foreign nucleic acid fragment. This application also specifically relates to the use of the transposase, the nucleic acid and its construct, the nucleome and its construct, the composition, the recombinant vector, or the recombinant host cell in introducing a foreign nucleic acid fragment gene into the genome of a host cell, or in the preparation of drugs or formulations for gene therapy, cell therapy, genome research, or stem cell induction and post-differentiation. Background Technology

[0005] A transposon is a DNA sequence that can be inserted into or removed from the genome, thereby transferring its own sequence or a complete copy of its own sequence within or between genomes. Transposons are mainly divided into two categories; this article primarily focuses on type II transposons (DNA transposons), which consist of terminal inverted repeats (TIRs) at both ends and a gene encoding a transposase. Transposons possess a "cut-and-paste" transposition mechanism, cutting DNA from the chromosome and inserting it directly into other parts of the genome.

[0006] Transposases are sequence-specific DNA-binding proteins expressed from DNA transposon sequences, containing catalytic domains that mediate DNA breakage and ligation. Transposases recognize and bind to the TIRs at both ends of the transposon, forming a protrusion complex, which then removes the DNA transposon from its original site and integrates it into a new site. The transposon's transposition activity depends primarily on the expression level and activity of the transposase. Therefore, DNA transposons with high transposase activity are a key requirement for developing transposon-based gene editing tools.

[0007] The insertion and integration of large gene fragments has significant applications in gene therapy, molecular breeding of plants and animals, and industrial microbial modification. Currently, effective tools and systems for large gene fragment insertion and integration are lacking in industry. In recent years, the scientific community has developed some tools and methods capable of inserting and integrating large gene fragments, but these methods still have some limitations. For example, lentiviruses or retroviruses are most commonly used to integrate gene sequences in cell immunotherapy and gene therapy for hereditary diseases. Several therapeutic products based on these have been used to treat tumors and hereditary diseases (Aiuti, A., Roncarolo, MG and Naldini, L. (2017) Gene therapy for ADA-SCID, the first marketing approval of an ex vivo gene therapy in Europe: paving the road for the next generation of advanced therapy medicinal products. [ADA-SCID gene therapy, the first marketing approval of an ex vivo gene therapy in Europe: paving the road for the next generation of advanced therapy medicinal products] EMBO Mol. Med. [EMBO Molecular Medicine] 9, 737–740; Aiuti, A. et al. (2009) Gene therapy for immunodeficiency due to adenosine deaminase deficiency. [Gene therapy for immunodeficiency due to adenosine deaminase deficiency] N. Engl. J. Med. [New England Journal of Medicine] 360, 447–458). However, using viruses for large-fragment gene integration has some potential limitations: First, the randomness of viral genome integration poses a risk of oncology; second, the size of the foreign gene that a virus can carry is also limited, which is not conducive to the transfer of large therapeutic gene fragments; third, the immunogenicity of the virus may affect the long-term expression of the foreign therapeutic gene and re-administration; fourth, virus production requires the use of living cells, making quality control and downstream processing of such products more complex and costly, thus posing a certain disadvantage in industrialization. Therefore, non-viral large-fragment integration can avoid the various drawbacks of viral integration and become a valuable tool in gene therapy.

[0008] As a non-viral gene integration tool, DNA transposons can not only achieve the integration and stable expression of large fragments of exogenous genes into the host genome, but also avoid negative impacts such as immunogenicity. Therefore, some transposons have already been used in gene therapy. Although transposons have been shown to be widely present in various regions from prokaryotes to eukaryotes, during evolution, in order to maintain genome stability, a large number of transposon fragments have become silent and inactive. Currently, a few highly active and valuable transposon tools, such as SleepingBeauty (SB), PiggyBac (PB), and Tol2, are used in gene therapy research. Therefore, the discovery of more highly active transposon tools and the validation and detection of their functions can provide more, better, and more flexible options for the development of gene therapy strategies.

[0009] It should be noted that the methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0010] Based on this, in order to seek more advanced and more efficient non-viral gene integration tools, this application provides an isolated transposase, wherein the transposase has a transposase sequence selected from (i) or (ii)-(iv) of the aforementioned transposase having transposase activity: (i) at least one of the amino acid sequences shown in any one of SEQ ID NO:1-146; (ii) at least one of the sequences obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NO:1-146; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95% or 99% identity with the amino acid sequence shown in any one of SEQ ID NO:1-146; and (iv) at least one of the sequences obtained by further fusing other sequences with the amino acid sequence shown in any one of SEQ ID NO:1-146. The transposase provided in this application has equal or even higher transposable activity compared to the widely used Sleeping Beauty (SB) and PiggyBac (PB), providing more or better options for the development of gene integration tools.

[0011] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in the following formula:

[0012] DE(X1) a K(X2)b G(X3) c K(X4) d G

[0013] Where a, b, c, and d represent the number of amino acids; D represents aspartic acid; E represents glutamic acid; K represents lysine; G represents glycine; (X1) represents any amino acid, and a is 17, 18, or 19; (X2) represents any amino acid, and b is 3, 4, or 5; (X3) represents any amino acid, and c is 1; and (X4) represents any amino acid, and d is 17, 18, or 19.

[0014] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in the following formula:

[0015] P(X5) e Y(X6) f D

[0016] Where e and f are the number of amino acids; P is proline; Y is tyrosine; D is aspartic acid; (X5) is any amino acid, and e is 5; and (X6) is any amino acid, and f is 7.

[0017] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in the following formula:

[0018] C(X7) g C(X8) h C(X9) i C

[0019] Where g, h, and i represent the number of amino acids; C represents cysteine; (X7) represents any amino acid, and g is 2, 3, or 4; (X8) represents any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28; and (X9) represents any amino acid, and i is 2.

[0020] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises at least two of the following amino acid sequences (1)-(3):

[0021] (1)DE(X1) a K(X2) b G(X3) c K(X4) d G;

[0022] (2)P(X5) e Y(X6)f D; or

[0023] (3)C(X7) g C(X8) h C(X9) i C,

[0024] Where a, b, c, d, e, f, g, h, and i represent the number of amino acids; D represents aspartic acid; E represents glutamic acid; K represents lysine; G represents glycine; P represents proline; Y represents tyrosine; D represents aspartic acid; C represents cysteine; (X1) represents any amino acid, and a is 17, 18, or 19; (X2) represents any amino acid, and b is 3, 4, or 5; (X3) represents any amino acid, and c is 1; (X4) represents any amino acid, and d is 17, 18 or 19; (X5) is any amino acid and e is 5; (X6) is any amino acid and f is 7; (X7) is any amino acid and g is 2, 3 or 4; (X8) is any amino acid and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28; and (X9) is any amino acid and i is 2.

[0025] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the following amino acid sequences (1)-(3):

[0026] (1)DE(X1) a K(X2) b G(X3) c K(X4) d G;

[0027] (2)P(X5) e Y(X6) f D; and

[0028] (3)C(X7) g C(X8) h C(X9) i C,

[0029] Where a, b, c, d, e, f, g, h, and i represent the number of amino acids; D represents aspartic acid; E represents glutamic acid; K represents lysine; G represents glycine; P represents proline; Y represents tyrosine; D represents aspartic acid; C represents cysteine; (X1) represents any amino acid, and a is 17, 18, or 19; (X2) represents any amino acid, and b is 3, 4, or 5; (X3) represents any amino acid, and c is 1; (X4) represents any amino acid, and d is 17, 18 or 19; (X5) is any amino acid and e is 5; (X6) is any amino acid and f is 7; (X7) is any amino acid and g is 2, 3 or 4; (X8) is any amino acid and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28; and (X9) is any amino acid and i is 2.

[0030] According to embodiments of this application, a nucleic acid may be provided, wherein the nucleic acid encodes the transposase described in this application.

[0031] According to embodiments of this application, a nucleic acid construct may be provided, which comprises the nucleic acid according to this application and also includes a promoter.

[0032] According to embodiments of this application, a nucleic acid genome comprising a 5' recognition sequence, wherein the 5' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO:147-292.

[0033] According to embodiments of this application, a nucleic acid genome comprising a 3' recognition sequence, wherein the 3' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO:293-438.

[0034] According to embodiments of this application, a nucleic acid set can be provided, comprising a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NO:147-292, and the 3' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NO:293-438, and the nucleic acid set can be recognized by a specific transposase.

[0035] According to embodiments of this application, a nucleic acid genome construct can be provided, which includes the nucleic acid genome described in this application and also includes exogenous nucleic acid fragments.

[0036] According to embodiments of this application, a composition may be provided, wherein the composition comprises: a PiggyBac family transposase or a functional fragment thereof, or a nucleic acid encoding a PiggyBac family transposase or a functional fragment thereof, wherein the transposase or functional fragment thereof has the function of catalyzing the insertion of a foreign nucleic acid fragment into the cell genome; and a nucleic acid set, wherein the nucleic acid set is recognizable by a specific transposase or functional fragment thereof.

[0037] According to embodiments of this application, a recombinant vector may be provided, wherein the recombinant vector comprises a nucleic acid encoding a transposase described in this application, a nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid genome described in this application, a nucleic acid genome construct described in this application, or a composition described in this application.

[0038] According to embodiments of this application, a recombinant host cell may be provided, wherein the recombinant host cell comprises the transposase described in this application, a nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid genome described in this application, a nucleic acid genome construct described in this application, a composition described in this application, or a recombinant vector described in this application.

[0039] According to embodiments of this application, a method for introducing exogenous nucleic acid fragments into the genome of a host cell can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0040] According to embodiments of this application, a method for editing the genome of a host cell can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, or the recombinant vector described in this application into a host cell.

[0041] According to embodiments of this application, a method for obtaining a host cell whose genome contains exogenous nucleic acid fragments can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0042] According to embodiments of this application, the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, the recombinant vector described in this application, or the recombinant host cell described in this application may be provided for the use of introducing exogenous nucleic acid fragments into the host cell genome.

[0043] According to embodiments of this application, the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, the recombinant vector described in this application, or the recombinant host cell described in this application may be used in the preparation of drugs or formulations for gene therapy, cell therapy, genome research, or stem cell induction and post-induction differentiation.

[0044] According to embodiments of this application, a kit may be provided, wherein the kit comprises the transposase described in this application, a nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid genome described in this application, a nucleic acid genome construct described in this application, a composition described in this application, a recombinant vector described in this application, or a recombinant host cell described in this application.

[0045] It should be understood that the descriptions in this section are not intended to identify key or essential features of the present application, nor are they intended to limit the scope of the application. Other features of the present application will become readily apparent from the following description. Attached Figure Description

[0046] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0047] Figure 1 A schematic diagram of the dual-plasmid vector in the transposon activity detection system of Example 1 is shown. Plasmid 1 is a plasmid expressing transposase (Tn), and plasmid 2 is a transposon donor plasmid.

[0048] Figure 2shows PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1 in 293T cells in Example 2The relative transposon efficiency results for PB04_D2, PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, PB04_F2, SB100X, and PiggyBac. Tn+ indicates co-transfection of the transposase plasmid and donor plasmid, and Tn- indicates transfection of only the donor plasmid. ,

[0049] Figure 3 , Figure 4 and Figure 5 yes Figure 2 A magnified view of a portion of the image.

[0050] Figure 6The following are examples of PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, and PB02_D1 in Example 3. 2. PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, P B03_A4, PB03_A5, PB03_A10, PB03_B3, PB03_B12, PB03_C1, PB03_C7, PB03_C11, PB03_D10, PB03_E6, PB03_F1, PB03_F5, PB03 _F6, PB03_F7, PB03_F12, PB03_G3, PB03_G4, PB03_G8, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A1 2. PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2, PB Cloning screening results for PB04_D3, PB04_D4, PB04_D7, PB04_D9, PB04_D11, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F2, PB04_F3, PB04_F7, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, and PB04_G12. Tn+ indicates co-transfection of the transposase plasmid and donor plasmid, while Tn- indicates transfection of only the donor plasmid.

[0051] Figure 7The following are examples of PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, and PB02_D1 in Example 3. 2. PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB 03_A4, PB03_A5, PB03_A10, PB03_B3, PB03_B12, PB03_C1, PB03_C7, PB03_C11, PB03_D10, PB03_E6, PB03_F1, PB03_F5, PB03_ F6, PB03_F7, PB03_F12, PB03_G3, PB03_G4, PB03_G8, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12 , PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2, PB0 Transposable activity detection results for 4_D3, PB04_D4, PB04_D7, PB04_D9, PB04_D11, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F2, PB04_F3, PB04_F7, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10 and PB04_G12.

[0052] Figure 8 and Figure 9 yes Figure 7 A magnified view of a portion of the image.

[0053] Figure 10shows PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2 in Example 2Evolutionary branching diagrams based on protein sequences for PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, PB04_F2, and PiggyBac.

[0054] Figure 11Shows PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2 in Example 2Protein sequence similarity results between PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, PB04_F2, and PiggyBac.

[0055] Figure 12 , Figure 13 and Figure 14 yes Figure 11 A magnified view of a portion of the image. Detailed Implementation

[0056] Unless otherwise stated or contradicted by the context, the terms or expressions used herein should be read in conjunction with the entire disclosure and as understood by one of ordinary skill in the art. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0057] In this application, the terms "nucleic acid" and "polynucleotide" are used interchangeably to refer to a polymeric form of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogues.

[0058] In this application, the terms "polypeptide" and "peptide" are used interchangeably to refer to a polymer of amino acids of any length. Therefore, polypeptides, oligopeptides, proteins, antibodies, and enzymes are all included in the definition of polypeptide.

[0059] As described in this application, a “fragment” of a sequence refers to a portion of a sequence. For example, a fragment of a nucleic acid sequence refers to a portion of a nucleic acid sequence, and a fragment of an amino acid sequence refers to a portion of an amino acid sequence.

[0060] As described in this application, a sequence “variant” is a polynucleotide or polypeptide that differs from a reference polynucleotide or polypeptide, but retains its essential characteristics. A typical variant of a polynucleotide differs from another reference polynucleotide in its nucleic acid sequence, and this difference may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. A typical variant of a polypeptide differs from another reference polypeptide in its amino acid sequence. Generally, the differences are limited, and thus the sequences of the reference polypeptide and the variant are very similar overall, sharing many regions. Variant polypeptides and reference polypeptides may differ in their amino acid sequences through one or more substitutions, additions, or deletions in any combination. The substituted or inserted amino acid residues may or may not be genetically encoded residues. Variants of polynucleotides or polypeptides may be naturally occurring, such as allelic variations, or may be unknown naturally occurring variants. Non-naturally occurring polynucleotide and polypeptide variants can be produced using mutagenesis techniques, direct synthesis, and other recombinant methods known to those skilled in the art.

[0061] Amino acids are typically classified by the nature of their side chains. For example, side chains can make amino acids weak acids (e.g., amino acids D and E) or weak bases (e.g., amino acids K, R, and H); if the side chain is polar, the amino acid is hydrophilic (e.g., amino acids L and I), or if the side chain is nonpolar, the amino acid is hydrophobic (e.g., amino acids S and C).

[0062] As used in this application, the term "family" refers to a group of nucleic acids or proteins that are structurally similar and have been generated from the same ancestor through replication and mutation. They typically have related or even identical functions. The term "superfamily" refers to a group of nucleic acids or proteins that are structurally similar and have been generated from the same ancestor through replication and mutation. They belong to different families and typically have different functions.

[0063] As used in this application, the term "transposase" refers to a polypeptide that catalyzes the excision of a transposon (containing the foreign nucleic acid and transposase recognition sequences flanking it) from a first nucleic acid (a vector containing a transposase recognition sequence and the foreign nucleic acid) and its integration into a second nucleic acid, i.e., a target site (e.g., genomic or extrachromosomal DNA containing a target site duplication (TSD) sequence in a cell). In some embodiments, the transposase binds to at least one terminal inverted repeat (TIR) ​​sequence.

[0064] As used in this application, the term "recognition sequence" refers to the nucleic acid sequences located at both ends of the transposable element and flanking the first nucleic acid sequence of the transposable element. The recognition sequence located at the 5' end of the first nucleic acid sequence is called the 5' recognition sequence, and the recognition sequence located at the 3' end of the first nucleic acid sequence is called the 3' recognition sequence. In some embodiments, the recognition sequence comprises at least one terminal inverted repeat sequence that can bind to a transposase.

[0065] The term "nucleic acid construct" as used herein is defined as a single-stranded or double-stranded nucleic acid molecule, preferably an artificially constructed nucleic acid molecule. Optionally, the nucleic acid construct may also include one or more regulatory sequences operatively linked to direct the expression of the coding sequence in a suitable host cell under compatible conditions. The term "expression" should be understood to include any step involved in protein or peptide production, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion. The term "regulatory sequence" includes all components necessary or advantageous for the expression of the peptide / protein of this application. Each regulatory sequence may be naturally present or exogenous to the nucleic acid sequence encoding the protein or peptide. These regulatory sequences include, but are not limited to, leader sequences, poly(A) signal sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, the regulatory sequence shall include a promoter and initiation and termination signals for transcription and translation. To introduce specific restriction sites for linking the regulatory sequence to the coding region of the nucleic acid sequence encoding the protein or peptide, a regulator-linked regulatory sequence may be provided.

[0066] As used in this application, the term "promoter" refers to a polynucleotide sequence that can control the transcription of a coding sequence. A promoter sequence includes a specific sequence sufficient to enable RNA polymerase to recognize, bind to, and initiate transcription. Furthermore, the promoter sequence may include sequences that optionally regulate the recognition, binding, and transcription initiation activity of RNA polymerase in the nucleic acid construct or nucleome construct provided in this application. A promoter can affect the transcription of genes located on the same nucleic acid molecule as the promoter or genes located on different nucleic acid molecules than the promoter.

[0067] The term "exogenous nucleic acid fragment" as used in this application includes any gene of interest or any gene capable of transposition, or a fragment thereof. In some non-limiting embodiments, the exogenous nucleic acid fragment originates from a different source than the terminal repeat sequence; for example, it is a nucleic acid sequence isolated from an organism different from the terminal inverted repeat sequence, i.e., the exogenous nucleic acid fragment is foreign relative to the terminal inverted repeat sequence. In some non-limiting embodiments, the exogenous nucleic acid fragment originates from a different source than the host cell; for example, it is a nucleic acid sequence isolated from an organism different from the host cell, i.e., the exogenous nucleic acid fragment is foreign relative to the host cell.

[0068] The term "host cell" as used in this application includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. This term includes the progeny of the original cell into which a foreign nucleic acid fragment has been introduced. An exemplary host cell includes the human embryonic kidney cell HEK293T. It should be understood that, due to natural, accidental, or intentional mutations, the progeny of a single-parent cell may not necessarily be identical to the original parent in terms of morphology or in terms of genome or total DNA complementarity.

[0069] As used in this application, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule linked to it. Examples of vectors include, but are not limited to, plasmids, viruses, bacteria, bacteriophages, and insertable DNA fragments. The term "plasmid" refers to a circular double-stranded DNA capable of accepting exogenous nucleic acid fragments and replicating in prokaryotic or eukaryotic cells.

[0070] transposase

[0071] This application provides an isolated transposase having a transposase sequence selected from (i) or (ii)-(iv) of the aforementioned transposase having transposase activity: (i) at least one of the amino acid sequences shown in any one of SEQ ID NO: 1-146; (ii) at least one of the sequences obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in any one of the amino acid sequences shown in SEQ ID NO: 1-146; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with any one of the amino acid sequences shown in SEQ ID NO: 1-146; and (iv) at least one of the sequences obtained by further fusing other sequences with any one of the amino acid sequences shown in SEQ ID NO: 1-146.

[0072] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in the following formula:

[0073] DE(X1) a K(X2) b G(X3) c K(X4) d G

[0074] Where a, b, c, and d represent the number of amino acids; D represents aspartic acid; E represents glutamic acid; K represents lysine; G represents glycine; (X1) represents any amino acid, and a is 17, 18, or 19; (X2) represents any amino acid, and b is 3, 4, or 5; (X3) represents any amino acid, and c is 1; and (X4) represents any amino acid, and d is 17, 18, or 19.

[0075] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in the following formula:

[0076] P(X5) e Y(X6) f D

[0077] Where e and f are the number of amino acids; P is proline; Y is tyrosine; D is aspartic acid; (X5) is any amino acid, and e is 5; and (X6) is any amino acid, and f is 7.

[0078] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the amino acid sequence shown in the following formula:

[0079] C(X7) g C(X8) h C(X9) i C

[0080] Where g, h, and i represent the number of amino acids; C represents cysteine; (X7) represents any amino acid, and g is 2, 3, or 4; (X8) represents any amino acid, and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28; and (X9) represents any amino acid, and i is 2.

[0081] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises at least two of the following amino acid sequences (1)-(3):

[0082] (1)DE(X1) a K(X2) b G(X3) c K(X4) d G;

[0083] (2)P(X5) e Y(X6) f D; or

[0084] (3)C(X7) g C(X8) h C(X9) i C,

[0085] Where a, b, c, d, e, f, g, h, and i represent the number of amino acids; D represents aspartic acid; E represents glutamic acid; K represents lysine; G represents glycine; P represents proline; Y represents tyrosine; D represents aspartic acid; C represents cysteine; (X1) represents any amino acid, and a is 17, 18, or 19; (X2) represents any amino acid, and b is 3, 4, or 5; (X3) represents any amino acid, and c is 1; (X4) represents any amino acid, and d is 17, 18 or 19; (X5) is any amino acid and e is 5; (X6) is any amino acid and f is 7; (X7) is any amino acid and g is 2, 3 or 4; (X8) is any amino acid and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28; and (X9) is any amino acid and i is 2.

[0086] According to embodiments of this application, an isolated transposase can be provided, wherein the transposase comprises the following amino acid sequences (1)-(3):

[0087] (1)DE(X1) a K(X2) b G(X3) c K(X4) d G;

[0088] (2)P(X5) e Y(X6) f D; and

[0089] (3)C(X7) g C(X8) h C(X9) i C,

[0090] Where a, b, c, d, e, f, g, h, and i represent the number of amino acids; D represents aspartic acid; E represents glutamic acid; K represents lysine; G represents glycine; P represents proline; Y represents tyrosine; D represents aspartic acid; C represents cysteine; (X1) represents any amino acid, and a is 17, 18, or 19; (X2) represents any amino acid, and b is 3, 4, or 5; (X3) represents any amino acid, and c is 1; (X4) represents any amino acid, and d is 17, 18 or 19; (X5) is any amino acid and e is 5; (X6) is any amino acid and f is 7; (X7) is any amino acid and g is 2, 3 or 4; (X8) is any amino acid and h is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28; and (X9) is any amino acid and i is 2.

[0091] In some embodiments, transposases belong to the PiggyBac family.

[0092] In some embodiments, the species sources of the transposase include Arthropoda, Platyhelminthes, Cnidaria, Mollusca, Annelida, or Chordata. In some embodiments, the species sources of the transposase include Insecta, Actinopteris, Amphibia, Rhabditophora, Bivalvia, Hydrozoa, Ascidiacea, Anthozoa, or Clitellata. In some embodiments, the species sources of the transposase include Aedes aegypti, Aelia acuminata, Agrypnus murinus, Anthonomus grandis, Apoderus coryli, Aporophyla lueneburgensis, Atethmia centrago, Blastobasis adustella, Bombyx mori, Calamotropha paludella, Catocala fraxini, Chrysoteuchia culmella, Ciona savignyi, Coptotermes formosanus, Coremacera marginata, Crassostrea gigas, and Crassostrea gibbon. *Cryptotermes secundus*, *Diabrotica virgifera virgifera*, *Drosophila bipectinata*, *Drosophila elegans*, *Eubasilissa regina*, *Euschistusheros*, *Gonioctena quinquepunctata*, *Gymnosoma rotundatum*, *Heliconius melpomene*, and *Hermetia*illucens), Hesperophylaxmagnus, Homalodisca vitripennis, Hydra vulgaris, Hyles vespertilio, Ips nitidus, Ips typographus, Ischnura elegans, Lamprigera yunnana, Lasiommatamegera, Limonius californicus, Locusta migratoria, Macaria notata, Malachius bipustulatus, Mamestrabrassicae, Marasmarcha lunaedactyla, Marronus borbonicus, Melanotaenia boesemani, Mythimna impura, Nematostellavectensis, Ochropleura plecta, Ocypus olens, Oriusinsidiosus, Oryzias sinensis, Pachyrhynchus sulphureomaculatus, Parnassius apollo, Periplaneta americana, Philaenus spumarius, Philonthus cognatus, Pieris napi, Pissodes strobi, Platycnemis pennipes, Schistocerca americana, Schistocercapiceifrons, Schmidtea The following insect species are listed: mediterranea, stalk borer (Sesamianonagrioides), bee-like clearwing moth (Sesia apiformis), rice weevil (Sitophilus oryzae), invasive red imported fire ant (Solenopsis invicta), and black-faced oil cricket (Teleogryllus).(Occipitalis), Timemashepardi, Timema tahoe, Vandiemenella viatica, Ypsolopha sequella, or Zopobas atratus.

[0093] According to embodiments of this application, a nucleic acid may be provided, wherein the nucleic acid encodes the transposase described in this application.

[0094] According to embodiments of this application, a nucleic acid construct comprising a nucleic acid encoding the transposase described in this application can be provided. In some embodiments, the nucleic acid construct further comprises a promoter. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by a host cell expressing a nucleic acid sequence. The promoter sequence contains a transcriptional regulatory sequence mediating protein or polypeptide expression. The promoter can be any nucleic acid sequence that is transcriptionally active in a selected host cell, including mutated, truncated, and heterozygous promoters, and can be derived from a gene encoding an extracellular or intracellular protein or polypeptide that is homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrome protein promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0095] In some embodiments, the nucleic acid construct further comprises a poly(A) sequence. Poly(A) tailing signal sequences and various truncated forms of poly(A) tailing signals known in the art can be used in this application.

[0096] In some embodiments, the nucleic acid construct further comprises any transcription termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operatively attached to the 3' end of a nucleic acid sequence encoding a protein or polypeptide. Any terminator that can function in a selected host cell can be used in this invention.

[0097] Optionally, the nucleic acid construct may also include a suitable leader sequence, i.e., an untranslated region of mRNA that is crucial for translation in the host cell. The leader sequence is operatively linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in this invention.

[0098] Optionally, the nucleic acid construct may also include a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a prozymogen or polypeptide precursor. Polypeptides precursors are typically inactive and can be converted into mature, active polypeptides by catalytic or autocatalytic cleavage of the propeptide.

[0099] Optionally, the nucleic acid construct may also include a regulatory sequence that can adjust peptide expression according to the growth status of the host cell. Examples of regulatory sequences are systems that respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn gene expression on or off. Other examples of regulatory sequences are those that enable gene amplification. In these examples, the nucleic acid sequence encoding the protein or peptide should be operatively linked to the regulatory sequence.

[0100] Nucleic acid constructs

[0101] According to embodiments of this application, a nucleic acid genome comprising a 5' recognition sequence, wherein the 5' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO:147-292.

[0102] According to embodiments of this application, a nucleic acid genome comprising a 3' recognition sequence, wherein the 3' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NO:293-438.

[0103] According to embodiments of this application, a nucleic acid set can be provided, comprising a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NO:147-292, and the 3' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NO:293-438, and the nucleic acid set can be recognized by a specific transposase.

[0104] In some embodiments, the 5' identification sequence or the 3' identification sequence includes a terminal inverted repeat sequence, the length of which is at least one of 1-800nt, 1-600nt, 1-400nt, 1-200nt, 1-100nt, 5-50nt, 5-25nt, or 10-20nt.

[0105] According to embodiments of this application, a nucleic acid genome construct can be provided, comprising the nucleic acid genome described in this application and further comprising exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments are operatively inserted into the nucleic acid genome construct via multiple cloning insertion sites. There may be one or more exogenous nucleic acid fragments, which may be identical or different; a promoter may also be inserted to control the expression of the exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments include any gene of interest or any gene capable of transposition, such as a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene. In some embodiments, non-coding RNA genes include various RNAs with known functions, such as rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA), as well as RNAs with unknown functions. In some embodiments, natural functional protein genes include fluorescence-based reporter genes, luciferase genes, and resistance genes. In some embodiments, artificial chimeric genes include chimeric antigen receptor genes. In some embodiments, the fluorescence-based reporter gene is selected from at least one gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase gene is selected from at least one of the genes encoding firefly luciferase and kidney luciferase. In some embodiments, the resistance gene is selected from at least one of the genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, and bleomycin resistance.

[0106] In some embodiments, the nucleotide genome construct may also insert a promoter to control the expression of a foreign nucleic acid fragment. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by the host cell expressing the foreign nucleic acid fragment. The promoter sequence contains a transcriptional regulatory sequence that mediates protein or peptide expression. The promoter can be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutated, truncated, and heterozygous promoters, and can be derived from genes encoding extracellular or intracellular proteins or peptides that are homologous or heterologous to those of the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedral protein promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0107] In some embodiments, the nucleic acid genome construct also includes any transcription termination sequence (i.e., a sequence that can be recognized by the host cell to terminate transcription) to control the expression of the exogenous nucleic acid fragment. Any terminator that can function in the selected host cell can be used in this invention.

[0108] Optionally, the nucleic acid genome construct may also include a suitable leader sequence (i.e., an untranslated region of mRNA that is crucial for translation in the host cell) to control the expression of the exogenous nucleic acid fragment. The leader sequence is operatively linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in this invention.

[0109] Optionally, the nucleic acid construct may also include a propeptide coding region that controls the expression of the exogenous nucleic acid fragment. The propeptide coding region encodes an amino acid sequence located at the N-terminus of the polypeptide. The resulting polypeptide is called a prozymogen or polypeptide precursor. Polypeptides precursors are usually inactive and can be converted into mature, active polypeptides by cleavage of the propeptide through catalysis or autocatalysis.

[0110] Optionally, the nucleic acid genome construct may also include regulatory sequences that can regulate the expression of exogenous nucleic acid fragments according to the growth status of the host cell. Examples of regulatory sequences are systems that respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn gene expression on or off. Other examples of regulatory sequences are those that enable gene amplification. In these examples, the exogenous nucleic acid fragment should be operatively linked to the regulatory sequence.

[0111] Transposition Composition

[0112] According to embodiments of this application, a composition may be provided, wherein the composition comprises: a PiggyBac family transposase or a functional fragment thereof, or a nucleic acid encoding a PiggyBac family transposase or a functional fragment thereof, wherein the transposase or functional fragment thereof has the function of catalyzing the insertion of a foreign nucleic acid fragment into the cell genome; and a nucleic acid set, wherein the nucleic acid set is recognizable by a specific transposase or functional fragment thereof.

[0113] In some embodiments, the composition is selected from at least one group of groups (1)-(147) below, and any one group of groups (1)-(146) below comprises: transposase-associated sequences and nucleomes.

[0114] (1) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:1 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:147 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:293;

[0115] (2) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:2 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:148 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:294;

[0116] (3) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:3 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:149 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:295;

[0117] (4) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:4 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:150 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:296;

[0118] (5) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:5 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:151 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:297;

[0119] (6) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:6 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:152 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:298;

[0120] (7) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:7 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:153 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:299;

[0121] (8) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:8 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:154 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:300;

[0122] (9) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:9 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:155 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:301;

[0123] (10) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:10 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:156 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:302;

[0124] (11) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:11 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:157 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:303;

[0125] (12) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:12 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:158 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:304;

[0126] (13) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:13 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:159 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:305;

[0127] (14) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:14 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:160 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:306;

[0128] (15) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:15 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:161 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:307;

[0129] (16) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:16 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:162, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:308;

[0130] (17) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:17 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:163 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:309;

[0131] (18) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:18 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:164 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:310;

[0132] (19) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:19 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:165 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:311;

[0133] (20) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:20 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:166 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:312;

[0134] (21) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:21 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:167 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:313;

[0135] (22) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:22 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:168 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:314;

[0136] (23) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:23 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:169 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:315;

[0137] (24) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:24 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:170 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:316;

[0138] (25) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:25 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:171 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:317;

[0139] (26) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:26 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:172 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:318;

[0140] (27) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:27 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:173 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:319;

[0141] (28) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:28 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:174 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:320;

[0142] (29) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:29 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:175 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:321;

[0143] (30) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:30 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:176 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:322;

[0144] (31) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:31 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:177 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:323;

[0145] (32) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:32 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:178 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:324;

[0146] (33) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:33 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:179 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:325;

[0147] (34) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:34 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:180 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:326;

[0148] (35) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:35 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:181 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:327;

[0149] (36) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:36 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:182 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:328;

[0150] (37) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:37 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:183 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:329;

[0151] (38) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:38 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:184 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:330;

[0152] (39) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:39 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:185 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:331;

[0153] (40) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:40 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:186 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:332;

[0154] (41) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:41 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:187 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:333;

[0155] (42) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:42 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:188 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:334;

[0156] (43) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:43 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:189 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:335;

[0157] (44) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:44 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:190 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:336;

[0158] (45) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:45 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:191 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:337;

[0159] (46) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:46 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:192 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:338;

[0160] (47) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:47 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:193 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:339;

[0161] (48) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:48 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:194 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:340;

[0162] (49) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:49 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:195 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:341;

[0163] (50) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:50 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:196 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:342;

[0164] (51) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:51 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:197 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:343;

[0165] (52) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:52 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:198 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:344;

[0166] (53) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:53 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:199 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:345;

[0167] (54) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:54 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:200 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:346;

[0168] (55) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:55 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:201 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:347;

[0169] (56) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:56 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:202 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:348;

[0170] (57) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:57 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:203 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:349;

[0171] (58) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:58 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:204 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:350;

[0172] (59) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:59 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:205 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:351;

[0173] (60) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:60 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:206 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:352;

[0174] (61) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:61 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:207 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:353;

[0175] (62) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:62 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:208 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:354;

[0176] (63) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:63 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:209 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:355;

[0177] (64) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:64 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:210 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:356;

[0178] (65) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:65 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:211 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:357;

[0179] (66) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:66 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:212 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:358;

[0180] (67) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:67 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:213 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:359;

[0181] (68) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:68 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:214 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:360;

[0182] (69) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:69 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:215 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:361;

[0183] (70) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:70 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:216 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:362;

[0184] (71) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:71 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:217 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:363;

[0185] (72) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:72 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:218 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:364;

[0186] (73) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:73 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:219 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:365;

[0187] (74) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:74 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:220 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:366;

[0188] (75) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:75 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:221 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:367;

[0189] (76) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:76 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:222 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:368;

[0190] (77) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:77 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:223 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:369;

[0191] (78) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:78 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:224 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:370;

[0192] (79) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:79 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:225 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:371;

[0193] (80) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:80 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:226 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:372;

[0194] (81) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:81 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:227 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:373;

[0195] (82) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:82 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:228 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:374;

[0196] (83) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:83 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:229 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:375;

[0197] (84) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:84 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:230 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:376;

[0198] (85) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:85 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:231 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:377;

[0199] (86) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:86 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:232 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:378;

[0200] (87) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:87 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:233 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:379;

[0201] (88) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:88 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:234 ​​and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:380;

[0202] (89) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:89 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:235 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:381;

[0203] (90) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:90 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:236 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:382;

[0204] (91) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:91 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:237 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:383;

[0205] (92) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:92 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:238 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:384;

[0206] (93) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:93 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:239 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:385;

[0207] (94) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:94 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:240 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:386;

[0208] (95) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:95 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:241 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:387;

[0209] (96) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:96 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:242 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:388;

[0210] (97) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:97 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:243 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:389;

[0211] (98) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:98 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:244 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:390;

[0212] (99) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:99 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:245 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:391;

[0213] (100) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:100 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:246 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:392;

[0214] (101) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:101 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:247 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:393;

[0215] (102) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:102 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:248 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:394;

[0216] (103) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:103 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:249 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:395;

[0217] (104) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:104 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:250 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:396;

[0218] (105) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:105 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:251 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:397;

[0219] (106) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:106 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:252 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:398;

[0220] (107) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:107 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:253 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:399;

[0221] (108) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:108 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:254 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:400;

[0222] (109) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:109 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:255 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:401;

[0223] (110) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:110 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:256 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:402;

[0224] (111) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:111 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:257 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:403;

[0225] (112) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:112 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:258 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:404;

[0226] (113) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:113 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:259 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:405;

[0227] (114) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:114 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:260 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:406;

[0228] (115) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:115 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:261 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:407;

[0229] (116) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:116 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:262 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:408;

[0230] (117) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:117 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:263 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:409;

[0231] (118) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:118 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:264 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:410;

[0232] (119) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:119 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:265 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:411;

[0233] (120) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:120 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:266 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:412;

[0234] (121) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:121 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:267 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:413;

[0235] (122) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:122 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:268 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:414;

[0236] (123) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:123 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:269 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:415;

[0237] (124) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:124 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:270 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:416;

[0238] (125) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:125 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:271 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:417;

[0239] (126) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:126 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:272 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:418;

[0240] (127) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:127 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:273 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:419;

[0241] (128) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:128 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:274 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:420;

[0242] (129) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:129 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:275 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:421;

[0243] (130) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:130 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:276 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:422;

[0244] (131) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:131 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:277 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:423;

[0245] (132) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:132 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:278 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:424;

[0246] (133) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:133 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:279 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:425;

[0247] (134) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:134 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:280 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:426;

[0248] (135) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:135 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:281 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:427;

[0249] (136) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:136 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:282 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:428;

[0250] (137) The transposase-associated sequence is an amino acid sequence containing the sequence shown in SEQ ID NO:137 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set contains a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:283 and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:429;

[0251] (138) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:138 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:284 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:430;

[0252] (139) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:139 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:285 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:431;

[0253] (140) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:140 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:286 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:432;

[0254] (141) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:141 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:287 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:433;

[0255] (142) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:142 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:288 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:434;

[0256] (143) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:143 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:289 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:435;

[0257] (144) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:144 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:290 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:436;

[0258] (145) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:145 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:291 and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:437;

[0259] (146) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:146 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:292, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:438; or

[0260] (147) A variant of any of the aforementioned groups (1)-(146),

[0261] Wherein, the transposase-related sequence is an amino acid sequence of a variant of each group of transposases or a nucleic acid sequence encoding the variant, wherein the variant has a variant sequence of the aforementioned transposase having transposase activity selected from (i)-(iii) below:

[0262] (i) At least one of the amino acid sequences obtained by deletion, substitution, insertion, or mutation of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in each group of transposases;

[0263] (ii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequences shown in any one of SEQ ID No: 1-146; and

[0264] (iii) At least one of the sequences obtained by further fusing other sequences with the amino acid sequence shown in any one of SEQ ID No:1-146.

[0265] In some embodiments, the nucleic acid encoding an amino acid sequence further comprises a promoter. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by the host cell expressing the nucleic acid sequence. The promoter sequence contains a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutated, truncated, and heterozygous promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides that are homologous or heterologous to those of the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedral protein promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL. In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a poly(A) sequence. Poly(A) tailing signal sequences and various truncated forms of poly(A) tailing signals known in the art can be used in this application.

[0266] In some embodiments, the nucleotide genome further comprises exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments are operatively inserted into the nucleotide genome via multiple cloning insertion sites; there may be one or more exogenous nucleic acid fragments, which may be identical or different; a promoter may also be inserted to control the expression of the exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments include any gene of interest or any gene capable of transposition, such as a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene. In some embodiments, non-coding RNA genes include a variety of RNAs with known functions, such as rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA), as well as RNAs with unknown functions. In some embodiments, natural functional protein genes include fluorescence-based reporter genes, luciferase genes, and resistance genes. In some embodiments, artificial chimeric genes include chimeric antigen receptor genes. In some embodiments, fluorescence-based reporter genes include genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, luciferase genes include genes encoding firefly luciferase or kidney luciferase. In some embodiments, resistance genes include genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance. In some embodiments, the nucleome may also include a promoter inserted to control the expression of a foreign nucleic acid fragment. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by the host cell expressing the foreign nucleic acid fragment. The promoter sequence contains a transcriptional regulatory sequence that mediates protein or polypeptide expression. The promoter can be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutated, truncated, and heterozygous promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides that are homologous or heterologous to those of the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrome protein promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

[0267] In some embodiments, the nucleic acid and / or nucleic acid group encoding the amino acid sequence further includes any transcription termination sequence that controls the expression of the exogenous nucleic acid fragment, i.e., a sequence that can be recognized by the host cell to terminate transcription. Any terminator that can function in the selected host cell can be used in this invention.

[0268] In some embodiments, the nucleic acid and / or nucleic acid group encoding the amino acid sequence further includes any transcription termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operatively attached to the 3' end of the nucleic acid sequence encoding a protein or polypeptide. Any terminator that can function in the selected host cell can be used in this invention.

[0269] Optionally, the nucleic acid and / or nucleic acid set encoding the amino acid sequence may also include a suitable leader sequence, i.e., an untranslated region of mRNA that is crucial for translation in the host cell. The leader sequence is operatively linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in this invention.

[0270] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may also include a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a prozymogen or polypeptide precursor. Polypeptides precursors are typically inactive and can be converted into mature, active polypeptides by catalytic or autocatalytic cleavage of the propeptide.

[0271] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may also include a regulatory sequence that can regulate peptide expression according to the growth status of the host cell. Examples of regulatory sequences are those that respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn gene expression on or off. Other examples of regulatory sequences are those that can amplify a gene. In these examples, the nucleic acid sequence encoding the protein or peptide should be operatively linked to the regulatory sequence.

[0272] Recombinant vectors, recombinant host cells, and reagent kits

[0273] According to embodiments of this application, a recombinant vector can be provided, wherein the recombinant vector comprises a nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleotide genome described in this application, the nucleotide genome construct described in this application, or a composition described in this application. The recombinant vector can be any suitable vector. In some embodiments, the recombinant vector includes, but is not limited to, a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector. In some embodiments, the recombinant eukaryotic expression plasmid includes pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV. In some embodiments, the recombinant viral vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector. The recombinant vector of the present invention can be constructed using methods known in the art. For example, suitable restriction sites can be added to both ends of the nucleic acid construct of the present invention according to the restriction sites contained in the backbone vector used, and then loaded into the backbone vector.

[0274] According to embodiments of this application, a recombinant host cell can be provided, wherein the recombinant host cell comprises the transposase described in this application, a nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid genome described in this application, a nucleic acid genome construct described in this application, a composition described in this application, or a recombinant vector described in this application. The recombinant host cell can be any host cell to which the transposase can be applied. In some embodiments, the recombinant host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, animal cells include mammalian cells. In some embodiments, mammalian cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratin / duct / cell lines), and cancer cell lines (e.g., HeLa, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW). 620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

[0275] According to embodiments of this application, a kit may be provided, wherein the kit comprises the transposase described in this application, a nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid genome described in this application, a nucleic acid genome construct described in this application, a composition described in this application, a recombinant vector described in this application, or a recombinant host cell described in this application.

[0276] Methods and uses

[0277] The transposase-based large-fragment gene insertion and integration tools and methods provided in this application can be applied to multiple fields such as gene and cell therapy, molecular breeding of plants and animals, and industrial microbial modification. Particularly in the field of cell therapy, the transposon system provided in this application can be used for the integration of CAR sequences in cell immunotherapy (CAR-T, CAR-NK, CAR-M, etc.); in the field of gene therapy, the transposon system provided in this application can be used to insert or integrate healthy genes into the cell genome, thereby facilitating the treatment of diseases caused by gene mutations or defects; in molecular breeding, the transposon system provided in this application can serve as a tool for breeding many crops such as rice, corn, and wheat, and can also be used to specifically accelerate the breeding process of plants and animals; in industrial microbial modification, since plasmids have defects such as instability and easy loss in gene expression, the transposon system provided in this application can stably integrate genes into the chromosomes of microorganisms.

[0278] According to embodiments of this application, a method for introducing exogenous nucleic acid fragments into the genome of a host cell can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0279] According to embodiments of this application, a method for editing the genome of a host cell can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, or the recombinant vector described in this application into a host cell.

[0280] According to embodiments of this application, a method for obtaining a host cell whose genome contains exogenous nucleic acid fragments can be provided, wherein the method includes: delivering the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, or the recombinant vector described in this application into the host cell.

[0281] The method of delivery into host cells can be any suitable method. In some embodiments, the delivery method includes, but is not limited to, cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, peptide and protein delivery, retroviral delivery, lentiviral delivery, adenovirus delivery, adeno-associated virus delivery, electroporation, Agrobacterium infection, or gene gun. The cell transfection and culture methods are conventional in the art, and appropriate transfection and culture methods can be selected according to different cell types.

[0282] The host cell can be any host cell to which the transposase can be applied. In some embodiments, the host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, animal cells include mammalian cells. In some embodiments, mammalian cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratin / duct / cell lines), and cancer cell lines (e.g., HeLa, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW). 620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

[0283] According to embodiments of this application, the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleome described in this application, the nucleome construct described in this application, the composition described in this application, the recombinant vector described in this application, or the recombinant host cell described in this application can be used to introduce exogenous nucleic acid fragments into the host cell genome. The host cell can be any host cell to which the transposase can be applied. In some embodiments, the host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, animal cells include mammalian cells. In some embodiments, mammalian cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratin / duct / cell lines), and cancer cell lines (e.g., HeLa, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW). 620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

[0284] According to embodiments of this application, the transposase described in this application, the nucleic acid encoding the transposase described in this application, the nucleic acid described in this application, the nucleic acid construct described in this application, the nucleic acid genome described in this application, the nucleic acid genome construct described in this application, the composition described in this application, the recombinant vector described in this application, or the recombinant host cell described in this application may be used in the preparation of drugs or formulations for gene therapy, cell therapy, genome research, or stem cell induction and post-induction differentiation.

[0285] The various embodiments and preferences described above can be combined with each other (as long as they are not inherently contradictory) and are suitable for the purposes of this application. All embodiments formed by such combinations are considered part of this application.

[0286] Example

[0287] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present application, including various details of the embodiments to aid understanding. It should be understood that these are considered merely exemplary and are in no way intended to limit the scope of protection of this application. The scope of protection of this application is defined only by the claims. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0288] Unless otherwise stated, all reagents and instruments used in the following examples are conventional products that are commercially available. Unless otherwise stated, experiments were conducted under conventional conditions or conditions recommended by the manufacturer.

[0289] Example 1: Construction of a transposon activity detection system

[0290] We developed a fluorescence-based reporter gene-antibiotic selection marker-based detection system to validate the activity of candidate transposons. This system was validated using a dual-plasmid vector, such as... Figure 1 As shown: Plasmid 1 is a plasmid expressing transposase (Tn), containing a constitutive promoter CMV (sequence shown in SEQ ID NO:499) that can initiate transcription in eukaryotic cells, the sequence of the candidate transposase (as shown in Table 1), and a poly(A) sequence (PA, sequence shown in SEQ ID NO:500) that terminates transcription; Plasmid 2 is a transposon donor plasmid containing the GFP gene (sequence shown in SEQ ID NO:501), the puromycin resistance selection gene (PuroR, sequence shown in SEQ ID NO:502), the promoter PGK (sequence shown in SEQ ID NO:503), P2A (sequence shown in SEQ ID NO:504), and a poly(A) element (sequence shown in SEQ ID NO:500), wherein transposon sequences specifically recognized by transposases are inserted at both ends of these sequences. Figure 1 The LTF and RTF sequences are shown in Table 1.

[0291] Table 1. Sequences related to plasmid construction

[0292]

[0293]

[0294]

[0295]

[0296]

[0297]

[0298]

[0299] When both plasmids are co-transfected into HEK293T cells, the transposase gene in plasmid 1 initiates transcription to express the transposase protein. The transposase protein is then recognized and binds to the transposon recognition sequence on plasmid 2. This transposon recognition sequence, along with the GFP gene and the puromycin resistance gene, is excised from the plasmid vector and integrated into the cell genome. When cells are continuously cultured in a medium containing a certain concentration of puromycin, only cells that have undergone transposition survive because they contain the puromycin resistance gene in their genome. The transposition activity level of candidate transposases is reflected by the number of surviving cells or their ability to form monoclonal cells.

[0300] DNA synthesis and plasmid construction methods:

[0301] Construction of plasmid 1: The amino acid sequence of the transposase was synthesized by Beijing Tsingke Biotech Co., Ltd. and GENERAL Biosystems (Anhui) Co., Ltd. The corresponding DNA sequence was cloned into the plasmid vector pICOZ containing CMV promoter elements through the 5' EcoRI site and the 3' NotI site, so that the transposase gene could be transcribed in eukaryotic cells under the control of the CMV promoter and subsequently translated into a functional protein.

[0302] Construction of plasmid 2: The transposon sequences (containing terminal inverted repeat sequences) are located on both sides of the transposase open reading frame. The left transposon fragment (LTF) contains the entire DNA sequence from the target site repeat (TSD) sequence at the 5' end to the sequence before the transposase start codon. The right transposon fragment (RTF) contains the entire DNA sequence from the first base after the transposase stop codon to the TSD sequence at the 3' end. In principle, the terminal repeat sequences recognized by the transposase are contained in the transposon sequences on both sides. The LTF and RTF fragments were synthesized by BGI Tech Solutions (Beijing Liuhe) Co., Ltd., and cloned into pMV plasmid vectors containing elements such as the PGK promoter, puromycin resistance gene (PuroR), P2A, green fluorescent protein gene (GFP), and poly(A), so that the LTF is located upstream of the PGK promoter and the RTF is located downstream of poly(A).

[0303] Plasmid 1 and plasmid 2 are in one-to-one correspondence.

[0304] Example 2: High-throughput screening for transposable activity

[0305] 2.1 Cell treatment (Day 0):

[0306] HEK293T cells (commercially purchased) stably expressing the firefly luciferase gene were established for high-throughput screening assays. After cells reached the logarithmic growth phase, they were digested and discretized into single cells using 0.25% trypsin (Thermo). Cells were then cultured at a concentration of 1.0 × 10⁻⁶ cells / cells. 4 Cells / well were added to 96-well cell culture plates pre-coated with PDL (Sigma) and incubated overnight at 37°C with 5% CO2.

[0307] 2.2 Cell transfection (Day 1):

[0308] For each transposon system, two plasmids were mixed at a dosage of 20 ng for plasmid 1 and 10 ng for plasmid 2. These were then mixed with Lipofectamine 2000 (Thermo Fisher Scientific) transfection reagent at a ratio of plasmid mass (μg): reagent volume (μL) of 1:2. The mixture was incubated at room temperature for 15 min to form a transfection complex. The transfection complex was then transferred to cell culture plates and incubated with cells. Two parallel tests were performed on each sample to be screened.

[0309] 2.3 Cell screening (Day 3)

[0310] Forty-eight hours after transfection, the culture medium was replaced with DMEM (Thermo Fisher Scientific) selection medium (containing 2 μg / mL puromycin (Invivogen), 10% fetal bovine serum, and 1% penicillin / streptomycin (Thermo Fisher Scientific)) and cultured at 37°C and 5% CO2 for 4 days. Then, the cells were digested into single cells with 0.25% trypsin, diluted 1:5, and transferred to another 96-well plate pre-coated with PDL. These were then cultured at 37°C for 4 days in DMEM selection medium containing 2 μg / mL puromycin, 10% fetal bovine serum, and 1% penicillin / streptomycin.

[0311] 2.4 Cell viability assay (Day 11)

[0312] 2.4.1 Preparation of detection reagents: ... The luciferase assay system (Promega) was mixed with PBS at a volume ratio of 1:5. Assay reagents were prepared at a dose of 50 μL / well, and 5 mL of reagents were prepared for use in 96-well plates.

[0313] 2.4.2 Remove cells screened with puromycin for 8 days from the incubator. After removing the culture medium, add the detection reagent at a dose of 50 μL / well. After incubating in the dark at room temperature for 5 minutes, perform detection using a multi-functional microplate reader with luminescent detection function. The more cells that survive after puromycin screening and the stronger the detected luminescent signal, the higher the transposition activity of the sample.

[0314] 2.5 Statistical Results

[0315] During high-throughput screening, positive and negative controls were set up on each plate. Based on the readings of the luminescent signals detected by a microplate reader in each well, the fold change of each sample's reading relative to the average reading of the positive control (SB100X) was calculated. This calculated fold change for each sample was then divided by the calculated fold change for the inactive transposase (PB03_D1) to obtain the relative transposase activities of all transposases, such as... Figure 2 As shown in Table 2.

[0316] Table 2 Results of relative transposition activity in Example 2

[0317]

[0318]

[0319]

[0320]

[0321]

[0322]

[0323] Example 3: Transposable Activity Assay

[0324] 3.1 Cell treatment (Day 0):

[0325] HEK293T cells (commercially purchased) were cultured to the logarithmic growth phase, then digested and dispersed into single cells using 0.25% trypsin (Thermo Fisher Scientific), and cultured at a cell concentration of 1.2 × 10⁻⁶ cells / year. 5 Cells / well were added to 24-well cell culture plates pre-coated with PDL (Sigma) and incubated overnight at 37°C with 5% CO2.

[0326] 3.2 Cell transfection (Day 1):

[0327] For each transposon system, two plasmids were mixed at a dosage of 200 ng for plasmid 1 and 100 ng for plasmid 2. These were then mixed with Lipofectamine 2000 (Thermo Fisher Scientific) transfection reagent at a ratio of plasmid mass (μg):transfection reagent volume (μL) of 1:2. The mixture was incubated at room temperature for 15 min to form a transfection complex. The transfection complex was then transferred to a cell culture plate and incubated with cells. Two parallel tests were performed on each sample to be screened.

[0328] 3.3 Cell screening (Day 3)

[0329] Forty-eight hours after transfection, cells were digested with 0.25% trypsin to separate them into single cells. Cells were then added to DMEM (Thermo Fisher Scientific) selection medium containing 2 μg / mL puromycin (Inveggie), 10% fetal bovine serum, and 1% penicillin / streptomycin (Thermo Fisher Scientific), diluted 1:2000, and transferred to 6-well plates for further culture. After 10 days of continuous selection culture in puromycin-resistant medium, clones were counted, and transposase activity was calculated.

[0330] 3.4 Cell staining (day 13)

[0331] Cells selected via puromycin and cultured in 6-well plates were washed with PBS and then fixed with 4% paraformaldehyde for 15 min at room temperature. The waste solution was discarded, and 0.2% methylene blue staining solution was added to the cells. The cells were stained at room temperature for 1 h. The stained cell clones were washed with PBS and photographed using an imaging system (BioRad). The number of cell clones in each well was counted. The results of transposase clone selection in HEK293T cells in this application are as follows: Figure 6 As shown in the figure, the staining results of surviving cell clones after puromycin resistance selection indicate that the transposition event occurred successfully. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, while Tn- indicates transfection of only the donor plasmid, serving as a negative control for each transposase sample.

[0332] 3.5 Statistical Results

[0333] Statistical results of transposable activity are as follows: Figure 7 As shown in Table 3. Tn+ indicates co-transfection with both the transposase plasmid and the donor plasmid, while Tn- indicates transfection with only the donor plasmid. The y-axis in the figure shows the calculated percentage of transposition activity, calculated using the following formula: Transposition efficiency (%) = Number of cell clones per well / (Number of cells per well × Transfection efficiency (GFP-positive cells%)) × 100%.

[0334] Table 3 Statistical results of transposable activity in Example 3

[0335]

[0336]

[0337]

[0338]

[0339] In the implementation of all the above embodiments, two transposons, SB100X and PiggyBac, were used as positive controls to evaluate the transposase activity of the present application. These two transposons are commercially available DNA transposons that are currently protected by patents. The sequences reported in the references Lajos Ma'te's et al. (Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates), Nature Genetics, 2009 41(6):753-761 and Cary, LC et al. (Transposon mutagenesis of baculoviruses: analysis of Trichoplusia nitransposon IFP2 insertions within the FP-locus of nuclear polyhedrosisviruses), Virology, 1989, 172(1):156-169) were synthesized and cloned into the corresponding plasmid vectors using the same method as in Example 1.

[0340] The statistical results of the transposase activity of this application are as follows: Figure 2 and Figure 7as shown in the figure. The above results show that the 146 transposases of the present application (PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB02_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02_E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB03_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F2, PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3, PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11,PB04_D1, PB04_D2, PB04_D3, PB04_D4, PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, and PB04_F2) have good transposition activity.

[0341] Meanwhile, a large number of inactive or low transposase activity transposases were also found during the screening process (e.g., PB01_A5, PB01_A7, PB01_B1, PB01_B3, PB01_B6, PB01_B8, PB01_E3, PB01_E9, PB01_F3, PB01_F12, PB02_B1, PB02_B4, PB02_B11, PB02_B12, PB02_C12, PB02_D8, PB02_E2, PB02_F2, PB02_F4, PB03_D1 in Table 1 of this application). Compared to these inactive or low-transpositional-activity transposases, the 146 transposases in this application (PB01_B9, PB01_B10, PB01_B11, PB01_B12, PB01_C3, PB01_C4, PB01_C7, PB01_C8, PB01_C9, PB01_C12, PB01_D1, PB01_D3, PB01_D4, PB02_A2, PB02_A4, PB02_A9, PB02_A12, PB02_B2, PB02_B8, PB02_B9, PB0 2_C6, PB02_C11, PB02_D3, PB02_D12, PB02_E1, PB02_E4, PB02_E5, PB02_E6, PB02_E7, PB02_E8, PB02_E9, PB02_E10, PB02 _E12, PB02_F5, PB02_F11, PB03_A1, PB03_A2, PB03_A3, PB03_A4, PB03_A5, PB03_A6, PB03_A8, PB03_A9, PB03_A10, PB03_ A11, PB03_A12, PB03_B2, PB03_B3, PB03_B4, PB03_B5, PB03_B6, PB03_B8, PB03_B10, PB03_B11, PB03_B12, PB03_C1, PB0 3_C2, PB03_C3, PB03_C4, PB03_C5, PB03_C7, PB03_C8, PB03_C9, PB03_C10, PB03_C11, PB03_D3, PB03_D4, PB03_D5, PB03_ D7, PB03_D8, PB03_D9, PB03_D10, PB03_E1, PB03_E6, PB03_E7, PB03_E8, PB03_E9, PB03_E10, PB03_E11, PB03_F1, PB03_F 2. PB03_F3, PB03_F4, PB03_F5, PB03_F6, PB03_F7, PB03_F8, PB03_F9, PB03_F11, PB03_F12, PB03_G1, PB03_G2, PB03_G3,PB03_G4, PB03_G6, PB03_G7, PB03_G8, PB03_G9, PB03_G10, PB03_G11, PB03_G12, PB04_A1, PB04_A3, PB04_A5, PB04_A7, PB04_A10, PB04_A11, PB04_ A12, PB04_B5, PB04_B7, PB04_B9, PB04_B10, PB04_B12, PB04_C3, PB04_C5, PB04_C8, PB04_C9, PB04_C11, PB04_D1, PB04_D2, PB04_D3, PB04_D4, PB0 The transposable activities of PB04_D6, PB04_D9, PB04_E3, PB04_E7, PB04_E8, PB04_E9, PB04_E10, PB04_E12, PB04_F1, PB04_F3, PB04_F7, PB04_F8, PB04_F10, PB04_F11, PB04_F12, PB04_G1, PB04_G4, PB04_G7, PB04_G8, PB04_G10, PB04_G12, PB04_D7, PB04_D11, and PB04_F2 were significantly higher, and most of them had comparable or better transposable activities compared to SB100X and PiggyBac.

[0342] also, Figure 10 This paper illustrates the evolutionary branching of transposons in the PiggyBac superfamily based on protein sequences. Figure 11 The protein sequence similarity (%) among transposons of the PiggyBac superfamily in this application is shown. The results show that these transposons cover different branches of the superfamily, including PiggyBac.

[0343] It should be noted that the above are merely preferred examples of this application and are not intended to limit this application. Various modifications and changes can be made to this application by those skilled in the art. Although specific embodiments have been described, alternatives, modifications, alterations, improvements, and substantial equivalents of the above embodiments may exist or be unforeseeable to the applicant or other persons skilled in the art. Therefore, the appended claims and any possible modified claims are intended to cover all such alternatives, modifications, alterations, improvements, and substantial equivalents. Importantly, as technology evolves, many elements described herein can be replaced by equivalent elements that appear after this application.

Claims

1. An isolated transposase, wherein the amino acid sequence of the transposase is shown in SEQ ID NO:

94.

2. The transposase according to claim 1, wherein the transposase belongs to the PiggyBac family.

3. The transposase according to claim 1, wherein the species source of the transposase includes arthropods.

4. The transposase according to claim 3, wherein the species source of the transposase includes insects.

5. The transposase according to claim 4, wherein the species source of the transposase includes the cotton boll weevil. (Anthonomus grandis) .

6. A nucleic acid, wherein the nucleic acid encodes a transposase according to any one of claims 1-5.

7. A nucleic acid construct comprising the nucleic acid according to claim 6.

8. The nucleic acid construct according to claim 7, wherein the nucleic acid construct further comprises a promoter.

9. The nucleic acid construct according to claim 8, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedral protein promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

10. The nucleic acid construct according to claim 7, wherein the nucleic acid construct further comprises a poly(A) sequence.

11. A nucleic acid genome comprising a 5' recognition sequence, wherein the 5' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO:

240.

12. A nucleic acid genome comprising a 3' recognition sequence, wherein the 3' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO:

386.

13. A nucleic acid genome comprising a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO: 240, the 3' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO: 386, and the nucleic acid genome is recognizable by a specific transposase.

14. The nucleotide genome according to any one of claims 11-13, wherein the 5' recognition sequence or the 3' recognition sequence comprises a terminal inverted repeat sequence, the length of which is at least one selected from 1-800 nt, 1-600 nt, 1-400 nt, 1-200 nt, 1-100 nt, 5-50 nt, 5-25 nt, or 10-20 nt.

15. A nucleic acid genome construct comprising a nucleic acid genome according to any one of claims 11-14, and further comprising exogenous nucleic acid fragments.

16. The nucleotide genome construct of claim 15, wherein the exogenous nucleic acid fragment is operatively inserted into the nucleotide genome construct via a multiple cloning insertion site, the nucleotide genome construct comprising one or more exogenous nucleic acid fragments, the plurality of exogenous nucleic acid fragments being identical or different; the nucleotide genome construct further inserting a promoter that controls the expression of the exogenous nucleic acid fragment.

17. The nucleic acid genome construct according to claim 16, wherein the exogenous nucleic acid fragment includes a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.

18. The nucleic acid genome construct according to claim 17, wherein the natural functional protein gene includes a fluorescence-based reporter gene, a luciferase gene, and an resistance gene.

19. The nucleic acid genome construct according to claim 18, wherein the fluorescence-based reporter gene is selected from at least one gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein.

20. The nucleic acid genome construct according to claim 18, wherein the luciferase gene is selected from at least one of the genes encoding firefly luciferase or kidney luciferase.

21. The nucleic acid genome construct according to claim 18, wherein the resistance gene is selected from at least one of the genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.

22. The nucleic acid genome construct according to claim 17, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.

23. The nucleic acid genome construct according to claim 16, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrome protein promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

24. A composition comprising: PiggyBac family transposases or functional fragments thereof, or nucleic acids encoding PiggyBac family transposases or functional fragments thereof, wherein the transposases or functional fragments thereof have the function of catalyzing the insertion of exogenous nucleic acid fragments into the cellular genome; and Nucleome, wherein the nucleome can be recognized by specific transposases or functional fragments thereof; The composition comprises: a transposase-associated sequence and a nucleic acid set, wherein the transposase-associated sequence is the amino acid sequence shown in SEQ ID NO:94 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:240, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:

386.

25. The composition of claim 24, wherein the nucleic acid group further comprises a promoter.

26. The composition of claim 25, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrome promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

27. The composition of claim 24, wherein the nucleic acid group further comprises a poly(A) sequence.

28. The composition according to any one of claims 24-27, wherein the nucleic acid group further comprises exogenous nucleic acid fragments.

29. The composition of claim 28, wherein the exogenous nucleic acid fragment is operatively inserted into the nucleotide genome via a multiple cloning insertion site, the nucleotide genome construct comprising one or more exogenous nucleic acid fragments, the plurality of exogenous nucleic acid fragments being identical or different; the nucleotide genome construct further inserting a promoter that controls the expression of the exogenous nucleic acid fragment.

30. The composition of claim 29, wherein the exogenous nucleic acid fragment comprises a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.

31. The composition of claim 30, wherein the natural functional protein gene comprises a fluorescence-based reporter gene, a luciferase gene, or a resistance gene.

32. The composition of claim 31, wherein the fluorescence-based reporter gene comprises a gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein.

33. The composition of claim 31, wherein the luciferase gene comprises a gene encoding firefly luciferase or kidney luciferase.

34. The composition of claim 31, wherein the resistance gene comprises a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.

35. The composition of claim 30, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.

36. The composition of claim 29, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedral protein promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.

37. A recombinant vector, wherein the recombinant vector comprises a nucleic acid encoding a transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid genome according to any one of claims 11-14, a nucleic acid genome construct according to any one of claims 15-23, or a composition according to any one of claims 24-36.

38. The recombinant vector according to claim 37, wherein the recombinant vector comprises a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector.

39. The recombinant vector according to claim 38, wherein the recombinant eukaryotic expression plasmid comprises pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV.

40. The recombinant vector according to claim 38, wherein the recombinant viral vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector.

41. A recombinant host cell, wherein the recombinant host cell comprises a transposase according to any one of claims 1-5, a nucleic acid encoding the transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid genome according to any one of claims 11-14, a nucleic acid genome construct according to any one of claims 15-23, a composition according to any one of claims 24-36, or a recombinant vector according to any one of claims 37-40.

42. The recombinant host cell according to claim 41, wherein the recombinant host cell includes animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells.

43. The recombinant host cell of claim 42, wherein the animal cell comprises mammalian cells.

44. The recombinant host cell according to claim 43, wherein the mammalian cell includes primary cells, immortalized cell lines, cancer cell lines, embryonic stem cell lines and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.

45. The recombinant host cell according to claim 44, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial cell line, hTERT immortalized epithelial cell line, hTERT immortalized fibroblast cell line, hTERT immortalized keratinocyte cell line, and hTERT immortalized duct cell line. The cancer cell lines include HeLa, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

46. ​​A method for introducing a foreign nucleic acid fragment into the genome of a host cell for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1-5, the nucleic acid encoding the transposase according to any one of claims 1-5, the nucleic acid according to claim 6, the nucleic acid construct according to any one of claims 7-10, the nucleic acid genome according to any one of claims 11-14, the nucleic acid genome construct according to any one of claims 15-23, the composition according to any one of claims 24-36, or the recombinant vector according to any one of claims 37-40 is delivered into a host cell.

47. A method for editing the genome of a host cell for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1-5, the nucleic acid encoding the transposase according to any one of claims 1-5, the nucleic acid according to claim 6, the nucleic acid construct according to any one of claims 7-10, the nucleic acid genome according to any one of claims 11-14, the nucleic acid genome construct according to any one of claims 15-23, the composition according to any one of claims 24-36, or the recombinant vector according to any one of claims 37-40 is delivered into a host cell.

48. A method for obtaining host cells containing exogenous nucleic acid fragments in their genome for non-diagnostic or therapeutic purposes, wherein the method comprises: The transposase according to any one of claims 1-5, the nucleic acid encoding the transposase according to any one of claims 1-5, the nucleic acid according to claim 6, the nucleic acid construct according to any one of claims 7-10, the nucleic acid genome according to any one of claims 11-14, the nucleic acid genome construct according to any one of claims 15-23, the composition according to any one of claims 24-36, or the recombinant vector according to any one of claims 37-40 is delivered into a host cell.

49. The method according to any one of claims 46-48, wherein the delivery method comprises cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenovirus delivery, adeno-associated virus delivery, electroporation, Agrobacterium infection, or gene gun.

50. The method according to any one of claims 46-48, wherein the host cell comprises animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells.

51. The method of claim 50, wherein the animal cell comprises a mammalian cell.

52. The method according to claim 51, wherein the mammalian cells include primary cells, immortalized cell lines, cancer cell lines, embryonic stem cell lines and cells differentiated therefrom, or induced pluripotent stem cell lines and cells differentiated therefrom.

53. The method according to claim 52, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial cell line, hTERT immortalized epithelial cell line, hTERT immortalized fibroblast cell line, hTERT immortalized keratinocyte cell line, and hTERT immortalized duct cell line. The cancer cell lines include HeLa, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

54. The use of the transposase according to any one of claims 1-5, the nucleic acid encoding the transposase according to any one of claims 1-5, the nucleic acid according to claim 6, the nucleic acid construct according to any one of claims 7-10, the nucleotide genome according to any one of claims 11-14, the nucleotide genome construct according to any one of claims 15-23, the composition according to any one of claims 24-36, the recombinant vector according to any one of claims 37-40, or the recombinant host cell according to any one of claims 41-45 in the preparation of a medicament or reagent for introducing exogenous nucleic acid fragments into the host cell genome.

55. The use according to claim 54, wherein the host cell comprises animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells.

56. The use according to claim 55, wherein the animal cell comprises a mammalian cell.

57. The use according to claim 56, wherein the mammalian cells include primary cells, immortalized cell lines, cancer cell lines, embryonic stem cell lines and cells differentiated therefrom, or induced pluripotent stem cell lines and cells differentiated therefrom.

58. The use according to claim 57, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial cell line, hTERT immortalized epithelial cell line, hTERT immortalized fibroblast cell line, hTERT immortalized keratinocyte cell line, and hTERT immortalized duct cell line. The cancer cell lines include HeLa, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.

59. The use of a transposase according to any one of claims 1-5, a nucleic acid encoding a transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid genome according to any one of claims 11-14, a nucleic acid genome construct according to any one of claims 15-23, a composition according to any one of claims 24-36, a recombinant vector according to any one of claims 37-40, or a recombinant host cell according to any one of claims 41-45 in the preparation of a drug or formulation for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation.

60. A kit comprising a transposase according to any one of claims 1-5, a nucleic acid encoding the transposase according to any one of claims 1-5, a nucleic acid according to claim 6, a nucleic acid construct according to any one of claims 7-10, a nucleic acid genome according to any one of claims 11-14, a nucleic acid genome construct according to any one of claims 15-23, a composition according to any one of claims 24-36, a recombinant vector according to any one of claims 37-40, or a recombinant host cell according to any one of claims 41-45.

Citation Information

Patent Citations

  • Separated transposase and application thereof

    CN119053694A