Isolated transposases and their use
Isolated transposases with specific sequences address the limitations of viral vectors by enabling stable integration of large gene fragments, enhancing gene therapy efficacy and simplifying production.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ASTRAGENOMICS TECHNOLOGY CO LTD
- Filing Date
- 2024-03-12
- Publication Date
- 2026-04-10
AI Technical Summary
Current gene therapy methods using viral vectors face limitations such as random integration in the genome, limited size of exogenous genes, immunogenicity, and complex production processes, making them unsuitable for large fragment gene incorporation.
Development of isolated transposases with specific amino acid sequences or variants, compatible with existing tools like Sleeping Beauty and PiggyBac, for targeted integration of large exogenous nucleic acid fragments into host genomes, avoiding viral drawbacks.
The transposases provide stable and efficient integration of large gene fragments, reducing cancer risks and immunogenicity, and simplifying production processes, offering advanced gene therapy tools.
Smart Images

Figure 2026511241000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the priority of Chinese Patent Application No. 202310304787X, filed with the China National Intellectual Property Administration on March 27, 2023, the entire content of which is incorporated herein by reference in its entirety for all purposes.
[0002] Technical Field This application relates to the field of molecular biology, specifically to isolated transposases and their use. More specifically, this application relates to nucleic acids and nucleic acid constructs encoding transposases, nucleic acid sets and nucleic acid set constructs, as well as compositions, recombinant vectors, recombinant host cells and kits containing transposases. More specifically, this application relates to methods for introducing exogenous nucleic acid fragments into the genome of host cells, methods for editing the genome of host cells, and methods for obtaining host cells containing exogenous nucleic acid fragments in their genomes. More specifically, this application relates to the use of transposases, nucleic acids and nucleic acid constructs, nucleic acid sets and nucleic acid set constructs, compositions, recombinant vectors, or recombinant host cells for introducing exogenous nucleic acid fragment genes into the genome of host cells or for preparing drugs or formulations for gene therapy, cell therapy, genomic research, or stem cell induction and post - induction differentiation.
Background Art
[0003] A transposon is a DNA sequence that can insert into or excise from the genome to transfer its own sequence or a complete copy of its own sequence within or between genomes. Transposons are classified into two main categories and are mainly referred to as type II transposons (DNA transposons) in this specification, which consist of terminal inverted repeat sequences (TIRs) at both ends and a gene encoding a transposase. Transposons have a "cut - and - paste" transfer mechanism in which DNA is cut from a chromosome and directly inserted into another part of the genome.
[0004] Transposases are sequence-specific DNA-binding proteins expressed by DNA transposon sequences, containing catalytic domains that mediate DNA cleavage and ligation. Transposases recognize and bind to TIRs at both ends of a transposon, forming a bulge complex, which then removes the DNA transposon from its original site and integrates it into a new site. Transposation activity of a transposon depends primarily on the expression level and activity of the transposase. Therefore, DNA transposons with high transposase activity are a key requirement for the development of transposon function-based gene editing tools.
[0005] Gene insertion and large fragment integration have significant application value in fields such as gene therapy, molecular breeding of plants and animals, and the manipulation of industrial microorganisms. Currently, the industry lacks effective tools and systems for the insertion and integration of large fragment genes. In recent years, the scientific community has developed several tools and methods that can insert and integrate large fragment genes, but these methods still have several problems. For example, in cellular immunotherapy and gene therapy for genetic diseases, lentiviruses or retroviruses are most commonly used to incorporate gene sequences, and based on this, there are several therapeutic products for the treatment of tumors and genetic diseases (Aiuti, A., Roncarolo, MG and Naldini, L. (2017) Gene therapy for ADA-SCID, the first marketing approval of an ex vivo gene therapy in Europe: paving the road for the next generation of advanced therapy medicinal products. EMBO Mol. Med. 9, 737-740; Aiuti, A. et al. (2009) Gene therapy for immunodeficiency due to adenosine deaminase deficiency. N.Engl. J. Med. 360, 447-458). However, there are some potential limitations in the application of using viruses to incorporate large gene fragments. Firstly, the randomness of viral integration in the genome poses a risk of cancer; secondly, the size of exogenous genes that the virus can possess is also limited, which does not facilitate the introduction of large therapeutic fragment genes; thirdly, the immunogenicity of the virus may affect the long-term expression and re-administration of exogenous therapeutic genes; and fourthly, viral production must be completed with the help of living cells, which makes quality control and downstream processing of such products more complex and expensive, resulting in several disadvantages in terms of industrialization.Therefore, incorporating large non-viral fragments can avoid the various drawbacks caused by viral incorporation and could be a valuable tool in gene therapy. [Overview of the project] [Problems that the invention aims to solve]
[0006] As non-viral gene integration tools, DNA transposons can achieve integration into the host genome and stable expression of large fragments of exogenous genes, while also avoiding adverse effects such as immunogenicity. Therefore, several transposons are used in gene therapy. Transposons have been proven to be widely present in various fields from prokaryotes to eukaryotes; however, during evolution, many transposon fragments silently become inactive to maintain genomic stability. Currently, several highly active and valuable transposon tools, such as Sleeping Beauty (SB), PiggyBac (PB), and Tol2, are used in gene therapy research. Therefore, the discovery of more highly active transposon tools, along with the validation and detection of their functions, could provide better and more flexible choices for the development of gene therapy strategies.
[0007] It should be noted that the methods described in this section are not necessarily previously conceived or adopted. Unless otherwise specified, none of the methods described in this section should be considered prior art simply because they are included in this section. Similarly, the problems mentioned in this section should not be considered universally recognized in any prior art unless otherwise explicitly stated. [Means for solving the problem]
[0008] Based on this, in order to seek more advanced and effective non-viral gene integration tools, this application provides an isolated transposase having a transposase sequence selected from (i) below or a variant sequence of the above transposase having transposase activity in (ii) to (iv): (i) at least one amino acid sequence shown in any one of SEQ ID NOs: 1 to 79; (ii) at least one sequence obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids to the amino acid sequence shown in any one of SEQ ID NOs: 1 to 79; (iii) at least one amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1 to 79; and (iv) at least one sequence obtained by further fusing the amino acid sequence shown in any one of SEQ ID NOs: 1 to 79 with another sequence. The transposases provided in this application are compatible with or have even higher transposase activity compared to the currently widely used Sleeping Beauty (SB) and PiggyBac (PB), providing more or better options for the development of gene integration tools.
[0009] According to one embodiment of this application, an isolated transposase may be provided, the transposase having formula (I): DE(I) The formula contains the amino acid sequence shown, where D is aspartic acid and E is glutamic acid.
[0010] According to one embodiment of this application, an isolated transposase of formula (II): D(X1) a H(II) The amino acid sequence is as shown, where D is aspartic acid, H is histidine, a is the number of amino acids, and (X1) is any amino acid, and a is 5.
[0011] According to one embodiment of this application, an isolated transposase can be provided, formula (III): P(X2)(X3)(III) It contains the amino acid sequence shown, In the formula, P is proline, X2 is any amino acid, and X3 is aspartic acid or glutamic acid.
[0012] According to one embodiment of this application, an isolated transposase comprising at least two of the amino acid sequences shown in formulas (I), (II), and (III) can be provided.
[0013] According to one embodiment of this application, an isolated transposase comprising the amino acid sequences shown in formulas (I), (II), and (III) can be provided.
[0014] According to one embodiment of this application, a nucleic acid encoding the transposase described in this application may be provided.
[0015] According to one embodiment of the present application, a nucleic acid construct comprising the nucleic acid according to the present application and further comprising a promoter may be provided.
[0016] According to one embodiment of the present application, a nucleic acid set comprising a 5' recognition sequence is provided, the 5' recognition sequence comprising at least one of the nucleotide sequences shown in SEQ ID NOs: 80-158.
[0017] According to one embodiment of the present application, a nucleic acid set comprising a 3' recognition sequence is provided, the 3' recognition sequence comprising at least one of the nucleotide sequences shown in SEQ ID NOs: 159-237.
[0018] According to one embodiment of the present application, a nucleic acid set including a 5' recognition sequence and a 3' recognition sequence can be provided. The 5' recognition sequence includes a nucleotide sequence shown in any one of SEQ ID NOs: 80 to 158 or a variant thereof. The 3' recognition sequence includes a nucleotide sequence shown in any one of SEQ ID NOs: 159 to 237 or a variant thereof. The nucleic acid set can be recognized by a specific transposase.
[0019] According to one embodiment of the present application, a nucleic acid set construct including the nucleic acid set described in the present application and further including an exogenous nucleic acid fragment can be provided.
[0020] According to one embodiment of the present application, a composition can be provided. The composition is a nucleic acid encoding a Tc1 / mariner superfamily transposase or a functional fragment thereof, or a PiggyBac family transposase or a functional fragment thereof that encodes a Tc1 / mariner superfamily transposase or a functional fragment thereof, and the transposase or the functional fragment thereof has a function of catalyzing the insertion of an exogenous nucleic acid fragment into the genome of a cell, or a nucleic acid encoding a PiggyBac family transposase or a functional fragment thereof; and a nucleic acid set that can be recognized by a specific transposase or a functional fragment thereof.
[0021] According to an embodiment of the present application, a recombinant vector including the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, or the composition described in the present application can be provided.
[0022] According to an embodiment of the present application, a recombinant host cell including the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application can be provided.
[0023] According to one embodiment of the present application, there is provided a method for introducing an exogenous nucleic acid fragment into the genome of a host cell, the method comprising delivering to the host cell the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.
[0024] According to one embodiment of the present application, there is provided a method for editing the genome of a host cell, the method comprising delivering to the host cell the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.
[0025] According to one embodiment of the present application, there is provided a method for obtaining a host cell containing an exogenous nucleic acid fragment in its genome, the method comprising delivering to the host cell the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.
[0026] According to one embodiment of the present application, there is provided the use of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid set described in the present application, the nucleic acid set construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for introducing an exogenous nucleic acid fragment into the genome of a host cell.
[0027] According to one embodiment of this application, the use of a transposase described in this application, a nucleic acid encoding a transposase described in this application, a nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid set described in this application, a nucleic acid set construct described in this application, a composition described in this application, a recombinant vector described in this application, or a recombinant host cell described in this application may be provided for preparing a drug or formulation for gene therapy, cell therapy, genome research, or stem cell induction and post-induction differentiation.
[0028] According to embodiments of this application, a kit comprising the transposase described in this application, a nucleic acid encoding the transposase described in this application, a nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid set described in this application, a nucleic acid set construct described in this application, a composition described in this application, a recombinant vector described in this application, or a recombinant host cell described in this application may be provided.
[0029] It should be understood that the information in this section is not intended to identify any significant or important features of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will be readily apparent from the following description.
[0030] The accompanying drawings illustrate embodiments and form part of this specification and are used together with the description herein to illustrate exemplary implementations of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. Throughout the accompanying drawings, the same reference numerals indicate similar elements, but are not necessarily identical. [Brief explanation of the drawing]
[0031] [Figure 1] A schematic diagram of the two plasmid vectors in the transposon activity detection system of Example 1 is shown. Plasmid 1 is a plasmid expressing transposase (Tn), and Plasmid 2 is a transposon donor plasmid. [Figure 2]TCM_C_D8 (TCM_C_D8) (TCM_C_D8) (TCM_A2) TCM_A_B2、TCM_A_B5、TCM_A_C6、TCM_A_C9、TCM_A_D1、TCM_A _D3、TCM_A_D4、TCM_A_F8、TCM_A_G6、TCM_B_C4、TCM_B_C10、TCM_B_C11、TCM_B_C12、TCM_B_D1、TCM_B_D4、TCM_B_D1、TCM_B_D4、TCM_D_M_D1 B_D6、TCM_B_D7、TCM_B_D8、TCM_B_D10、TCM_B_E1、TCM_B_E4、TCM_B_E6、TCM_B_F2、TCM_B_F3、TCM_B_F5、TCM_B_F3、TCM_B_F5、TCM_F_10 _B_F11、TCM_B_G2、TCM_B_G10、TCM_B_G11、TCM_C_A12、TCM_C_B2、TCM_C_B8、TCM_C_C3、TCM_D_B11、TCM_D_9、TCM_A_1 TCM_E_A4、TCM_E_A5、TCM_E_A6、TCM_E_A7、TCM_E_A8、TCM_E_B6、TCM_E_B8、TCM_E_B10、TCM_E_B11、TCM_E_2_B11 1、TCM_E_C2、TCM_E_C3、TCM_E_C5、TCM_E_C6、TCM_E_C8、TCM_E_C10、TCM_E_C11、TCM_E_C12、TCM_E_D6、TCM_E_C12、TCM_E_D6、TCM_D_E_7 D12、TCM_E_E6、TCM_E_E7、TCM_E_E8、TCM_E_E10、TCM_E_E11、TCM_E_E12、TCM_E_F2、TCM_E_F3、TCM_E_F5、TCM_E_E_F3 _E_F8、TCM_E_F11、TCM_E_G4、TCM_E_G5、TCM_E_G6、TCM_E_G 7、TCM_E_G9、TCM_E_G12、SB100X 、PiggyBac wireless connection. [Figure 3]TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4 in 293T cells of Example 3, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_ B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, T CM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM _D_B11, TCM_D_G9, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B1 0, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TC The cloning screening results for M_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F11, TCM_E_G4, TCM_E_G6, TCM_E_G7, and TCM_E_G9 are shown. Tn+ indicates co-transfection with both the transposase plasmid and the donor plasmid, and Tn- indicates transfection with the donor plasmid only. [Figure 4]TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4 in 293T cells of Example 3 , TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TC M_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F 5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3 , TCM_D_B11, TCM_D_G9, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_ E_B10, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C The results of the detection of transposase activity for 12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F11, TCM_E_G4, TCM_E_G6, TCM_E_G7, and TCM_E_G9 are shown. Tn+ represents co-transfection with both the transposase plasmid and the donor plasmid, and Tn- represents transfection with the donor plasmid only. [Figure 5]TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_ A_C6、TCM_A_C9、TCM_A_D1、TCM_A_D3、TCM_A_D4、TCM_A_F 8、TCM_A_G6、TCM_B_C4、TCM_B_C10、TCM_B_C11、TCM_B_C12、TCM_B_D1、TCM_B_D4、TCM_B_D5、TCM_B_D6、TCM_B_D5、TCM_B_D6、TCM_D_7 CM_B_D8、TCM_B_D10、TCM_B_E1、TCM_B_E4、TCM_B_E6、TCM_B_F2、TCM_B_F3、TCM_B_F5、TCM_B_F10、TCM_B_F_F10、TCM_B_F_F10 B_G2、TCM_B_G10、TCM_B_G11、TCM_C_A12、TCM_C_B2、TCM_C_B8、TCM_C_C3、TCM_D_B11、TCM_D_G9、TCM_E_A1_M_E A4、TCM_E_A5、TCM_E_A6、TCM_E_A7、TCM_E_A8、TCM_E_B6、TCM_E_B8、TCM_E_B10、TCM_E_B11、TCM_E_B12、TCM_E_1、TCM_E_1 TCM_E_C2、TCM_E_C3、TCM_E_C5、TCM_E_C6、TCM_E_C8、TCM_E_C10、TCM_E_C11、TCM_E_C12、TCM_E_D6、TCM_E_D7、TCM_E_C5 _E_D12、TCM_E_E6、TCM_E_E7、TCM_E_E8、TCM_E_E10、TCM_E_E11、TCM_E_E12、TCM_E_F2、TCM_E_F3、TCM_E_E_F2、TCM_E_F3 _F7、TCM_E_F8、TCM_E_F11、TCM_E_G4、TCM_E_G5、TCM_E_G 6、TCM_E_G7、TCM_E_G9、TCM_E_G12、SB100X. [Figure 6]TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A _C6、TCM_A_C9、TCM_A_D1、TCM_A_D3、TCM_A_D4、TCM_A_F8 TCM_A_G6、TCM_B_C4、TCM_B_C10、TCM_B_C11、TCM_B_C12、TCM_B_D1、TCM_B_D4、TCM_B_D5、TCM_B_D6、TCM_B_D5、TCM_B_D6、TCM_D_7 _B_D8、TCM_B_D10、TCM_B_E1、TCM_B_E4、TCM_B_E6、TCM_B_F2、TCM_B_F3、TCM_B_F5、TCM_B_F10、TCM_B_F_F10、TCM_B_F_G_GTC 2、TCM_B_G10、TCM_B_G11、TCM_C_A12、TCM_C_B2、TCM_C_B8、TCM_C_C3、TCM_D_B11、TCM_D_G9、TCM_E_A1_ATC_AT_4 CM_E_A5、TCM_E_A6、TCM_E_A7、TCM_E_A8、TCM_E_B6、TCM_E_B8、TCM_E_B10、TCM_E_B11、TCM_E_B12、TCM_E_1、TCM_E_M_C_M _C2、TCM_E_C3、TCM_E_C5、TCM_E_C6、TCM_E_C8、TCM_E_C10、TCM_E_C11、TCM_E_C12、TCM_E_D6、TCM_E_D7、TCM_E_D6 、TCM_E_E6、TCM_E_E7、TCM_E_E8、TCM_E_E10、TCM_E_E11、TCM_E_E12、TCM_E_F2、TCM_E_F3、TCM_E_F5、TCM_E_7、TCM_E_7 _E_F8、TCM_E_F11、TCM_E_G4、TCM_E_G5、TCM_E_G6、TCM_E _G7、TCM_E_G9、TCM_E_G12 SB100X 2019-2014
Breakfast Story
[0032] Unless otherwise indicated or inconsistent with the context, terms or expressions used herein should be read in conjunction with the entirety of this disclosure and to be understood by those skilled in the art. All technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art unless otherwise defined.
[0033] In this application, the terms “nucleic acid” and “polynucleotide” are used interchangeably and refer to polymeric forms of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogues thereof.
[0034] In this application, the terms “polypeptide” and “peptide” are used interchangeably and refer to polymers of amino acids of any length. Therefore, polypeptides, oligopeptides, proteins, antibodies, and enzymes are all included in the definition of polypeptide.
[0035] As described in this application, a “fragment” of a sequence refers to a portion of a sequence. For example, a fragment of a nucleic acid sequence refers to a portion of a nucleic acid sequence, and a fragment of an amino acid sequence refers to a portion of an amino acid sequence.
[0036] As described in this application, a “variant” of a sequence is a polynucleotide or polypeptide that is different from a reference polynucleotide or polypeptide but retains essential properties. A typical variant of a polynucleotide has a different nucleic acid sequence from another reference polynucleotide, and this difference in nucleic acid sequence may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. A typical variant of a polypeptide has a different amino acid sequence from another reference polypeptide. Generally, the differences are limited such that the sequences of the reference polypeptide and the variant are generally very similar and identical in many regions. The amino acid sequences of the variant polypeptide and the reference polypeptide may differ by one or more substitutions, additions, or deletions in any combination. The substituted or inserted amino acid residues may or may not be residues encoded by the genetic code. Variants of polynucleotides or polypeptides may be naturally occurring, such as allele variants, or unknown naturally occurring variants. Polynucleotide and polypeptide variants that do not exist naturally can be produced by mutagenetic techniques, direct synthesis, and other recombination methods known to those skilled in the art.
[0037] Amino acids are typically classified according to the properties of their side chains. For example, the side chains can make an amino acid a weak acid (e.g., amino acids D and E) or a weak base (e.g., amino acids K, R, and H). If the side chain is polar, the amino acid becomes hydrophilic (e.g., amino acids L and I), or if the side chain is nonpolar, the amino acid becomes hydrophobic (e.g., amino acids S and C).
[0038] As used in this application, the term “family” refers to a group of nucleic acids or proteins that have high structural similarity and are produced by the same ancestor through replication and mutation, and which usually have related or the same function. A “superfamily” refers to a group of nucleic acids or proteins that have nearly identical structures and are produced by the same ancestor through replication and mutation, but which belong to different families and usually have different functions.
[0039] As used in this application, the term “transposase” refers to a polypeptide that catalyzes the excision of a transposon (containing an exogenous nucleic acid and transposase recognition sequences on both sides thereof) from a first nucleic acid (a vector containing a transposase recognition sequence and an exogenous nucleic acid) and its incorporation into a second nucleic acid, i.e., genomic DNA or extrachromosomal DNA containing an intracellular target site duplication (TSD) sequence. In some embodiments, the transposase binds to at least one terminal inverted repeat (TIR).
[0040] As used in this application, the term “recognition sequence” refers to nucleic acid sequences located at both ends of a transposable element and adjacent to the transposable first nucleic acid sequence, where a recognition sequence located at the 5' end of the first nucleic acid sequence is called a 5' recognition sequence, and a recognition sequence located at the 3' end of the first nucleic acid sequence is called a 3' recognition sequence. In some embodiments, the recognition sequence includes at least one terminally reversed repeat sequence that can bind to a transposase.
[0041] As used in this application, the term “nucleic acid construct” is defined herein as a single-stranded or double-stranded nucleic acid molecule, preferably an artificially constructed nucleic acid molecule. Optionally, a nucleic acid construct further comprises one or more operablely linked regulatory sequences that can direct the expression of a coding sequence in a suitable host cell under compatibility conditions. The term “expression” is understood to include, but is not limited to, any steps involved in the production of a protein or polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion. The term “regulatory sequence” includes all components necessary or advantageous for the expression of the polypeptide / protein of this application. Each regulatory sequence may be naturally present in the nucleic acid sequence encoding a protein or polypeptide, or may be exogenous. These regulatory sequences include, but are not limited to, leader sequences, polyadenylated sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, a regulatory sequence should include promoters for transcription and translation, as well as start and termination signals. Regulatory sequences having linkers may be provided for the purpose of introducing the regulatory sequence into a specific restriction site for linking the regulatory sequence to the coding region of the nucleic acid sequence encoding a protein or polypeptide.
[0042] As used in this application, the term “promoter” refers to a polynucleotide sequence capable of controlling the transcription of a coding sequence. A promoter sequence includes a sufficiently specific sequence to enable RNA polymerase to recognize, bind to, and initiate transcription. Furthermore, a promoter sequence may include sequences that optionally modulate the recognition, binding, and transcription initiation activity of RNA polymerase in the nucleic acid construct or nucleic acid set construct provided in this application. A promoter may affect the transcription of genes located on the same nucleic acid molecule as the promoter or genes located on a different nucleic acid molecule than the promoter.
[0043] As used in this application, the term “exogenous nucleic acid fragment” includes any gene of interest or any transferable gene or fragment thereof. In some non-limiting embodiments, the exogenous nucleic acid fragment is of a different origin from terminal repeats, for example, a nucleic acid sequence isolated from an organism different from the terminal inverted repeat, i.e., the exogenous nucleic acid fragment is exogenous with respect to terminal inverted repeats. In some non-limiting embodiments, the exogenous nucleic acid fragment is of a different origin from host cells, for example, a nucleic acid sequence isolated from an organism different from the host cell, i.e., the exogenous nucleic acid fragment is exogenous with respect to host cells.
[0044] As used in this application, the term “host cell” includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. The term includes the offspring of the original cell into which an exogenous nucleic acid fragment has been introduced. An example host cell is the human embryonic kidney cell HEK293T. It is understood that, due to natural, accidental, or intentional mutations, the offspring of a single parental cell may not necessarily be morphologically or with respect to the original parent in terms of genomic or total DNA complement.
[0045] As used in this application, the term “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is ligated. Examples of vectors include, but are not limited to, plasmids, viruses, bacteria, phages, and insertable DNA fragments. The term “plasmid” refers to a circular double-stranded DNA that can receive an exogenous nucleic acid fragment and replicate in prokaryotic or eukaryotic cells.
[0046] Transposase This application provides an isolated transposase having a transposase sequence selected from (i) below or a variant sequence of the above transposase having transposase activity in (ii) to (iv): (i) at least one amino acid sequence shown in any one of SEQ ID NOs: 1 to 79; (ii) at least one sequence obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the amino acid sequence shown in any one of SEQ ID NOs: 1 to 79; (iii) at least one amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1 to 79; (iv) at least one sequence obtained by further fusing the amino acid sequence shown in any one of SEQ ID NOs: 1 to 79 with another sequence.
[0047] In some embodiments, the transposase has a transposase sequence selected from the following groups (1) to (8): (1) at least one amino acid sequence shown in any one of SEQ ID NOs: 11 to 36 and 59 to 79; (2) at least one amino acid sequence shown in any one of SEQ ID NOs: 3 to 4 and 48 to 55; (3) at least one amino acid sequence shown in any one of SEQ ID NOs: 5 to 8 and 40 to 44; (4) at least one amino acid sequence shown in any one of SEQ ID NOs: 1 to 2 and 45 to 47; (5) at least one amino acid sequence shown in any one of SEQ ID NOs: 9 to 10 and 56 to 58; (6) at least one amino acid sequence shown in any one of SEQ ID NOs: 38 to 39; (7) at least one amino acid sequence shown in any one of SEQ ID NOs: 37; and (8) a transposase sequence selected from at least one amino acid sequence having at least 70% identity with any one of SEQ ID NOs: 1 to 79 in (1) to (7) above.
[0048] According to one embodiment of this application, an isolated transposase may be provided, the transposase having formula (I): DE(I) It contains the amino acid sequence shown, In the formula, D is aspartic acid and E is glutamic acid.
[0049] According to one embodiment of this application, an isolated transposase of formula (II): D(X1) a H(II) It contains the amino acid sequence shown, In the formula, D is aspartic acid, H is histidine, a is the number of amino acids, and (X1) is any amino acid, and a is 5.
[0050] According to one embodiment of this application, an isolated transposase can be provided, formula (III): P(X2)(X3)(III) It contains the amino acid sequence shown, In the formula, P is proline, X2 is any amino acid, and X3 is aspartic acid or glutamic acid.
[0051] According to one embodiment of this application, an isolated transposase comprising at least two of the amino acid sequences shown in formulas (I), (II), and (III) can be provided.
[0052] According to one embodiment of this application, an isolated transposase comprising the amino acid sequences shown in formulas (I), (II), and (III) can be provided.
[0053] In some embodiments, formulas (I) and (II) are spaced 80 to 120 amino acids apart, and formulas (II) and (III) are spaced 20 to 40 amino acids apart.
[0054] In some embodiments, the transposase belongs to the Tc1 / mariner superfamily.
[0055] In some embodiments, the transposase belongs to the Tc1, Tc2, Tc4, Mariner, Tigger, Pogo, Fot1, ISRm11, or m44 family.
[0056] In some embodiments, the species source of the transposase includes arthropods, chordates, cnidarians, mollusks, or flatworms. In some embodiments, the species source of the transposase includes insects, ray-finned fishes, amphibians, malacostraca, arachnida, chondrichthyes, sauropsida, bivalvia, ascidiacea, hydrozoans, or rhabditophoras. In some embodiments, the transposase species source is Aelia acuminata, Agrypnus murinus, Albula glossodonta, Amblyraja radiata, Anthonomus grandis, Astatotilapia calliptera, Blatella germanica, Bufo gargarizans, Carassius auratus, Cephalopholis sonnerati, Cerceris rybyensis, Cheilinus undulatus, and Chelonia midas. mydas), Clitarchus hookeri, Crassostrea gigas, Cromileptes altivelis, Cyprinus carpio, Danio rerio, Drosophila ananassae, Drosophila mojavensismojavensis), Epicauta chinensis, Folsomia candida, Formica aquilonia x Formica polyctena, Gasterosteus aculeatus, Gonioctena quinquepunctata, Gymnosoma rotundatum, Harpegnathos saltator, Hemibagrus wyckioides, Homalodisca vitripennis, Hydra vulgaris, Ischnura elegans elegans), Lampris incognitus, Laothoe populi, Latimeria chalumnae, Leistus spinibarbis, Lottia gigantea, Olea europaea subsp. europaea, Orgyia antiqua, Parhyale hawaiensis, Petrochromis sp. "Moshi Yellow" AB-2019), Philaenus spumarius, Philereme vetulata, Phytophthora ramorum ramorum), Polistes metricus, Rana pipiens, Rhinella marina, Rhodnius prolixus, Salmo salar, Schmidtea mediterraneaThis includes mediterranea, Seladonia tumulorum, Sesia apiformis, Sitophilus oryzae, Solenopsis invicta, Thalassophryne amazonica, Thecocarcelia acutangulata, Thymallus thymallus, Tribolium castaneum, Vandiemenella viatica, or Zaprionus bogoriensis.
[0057] According to one embodiment of this application, a nucleic acid encoding the transposase described in this application may be provided.
[0058] According to one embodiment of this application, a nucleic acid construct comprising a nucleic acid encoding the transposase described herein may be provided. In some embodiments, the nucleic acid construct further comprises a promoter. The promoter may be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by a host cell expressing the nucleic acid sequence. The promoter sequence comprises a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter may be any nucleic acid sequence having transcriptional activity in a selected host cell (including mutant, cleaved, and heterozygous promoters) and may originate from a gene encoding a homologous or heterologous extracellular or intracellular protein or polypeptide in the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
[0059] In some embodiments, the nucleic acid construct further comprises a polyadenylated [poly(A)] signal sequence. Poly(A) tailing signal sequences known in the art, as well as various cleavage forms of the poly(A) tailing signal, can be used in this application.
[0060] In some embodiments, the nucleic acid construct further includes an optional transcription termination sequence, i.e., a sequence recognized by a host cell to terminate transcription. The termination sequence is operably ligated to the 3' end of a nucleic acid sequence encoding a protein or polypeptide. Any terminator that is functional in a selected host cell can be used in the present invention.
[0061] Optionally, the nucleic acid construct may further include a suitable leader sequence, i.e., an untranslated region in mRNA that is important for translation in the host cell. The leader sequence is operably ligated to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in selected host cells can be used in the present invention.
[0062] Optionally, the nucleic acid construct may further include a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or propolypeptide. Propolypeptides are typically inactive and can be converted to mature, active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.
[0063] Optionally, nucleic acid constructs may further include regulatory sequences that can adjust polypeptide expression according to host cell growth conditions. Examples of regulatory sequences include systems that switch gene expression on or off in response to chemical or physical stimuli, including the presence of a regulatory compound. Other examples of regulatory sequences are those that enable gene amplification. In these examples, the nucleic acid sequence encoding a protein or polypeptide should be operably ligated to the regulatory sequence.
[0064] nucleic acid construct According to one embodiment of the present application, a nucleic acid set comprising a 5' recognition sequence is provided, the 5' recognition sequence comprising at least one of the nucleotide sequences shown in SEQ ID NOs: 80-158.
[0065] According to one embodiment of the present application, a nucleic acid set comprising a 3' recognition sequence is provided, the 3' recognition sequence comprising at least one of the nucleotide sequences shown in SEQ ID NOs: 159-237.
[0066] According to one embodiment of the present application, a nucleic acid set comprising a 5' recognition sequence and a 3' recognition sequence may be provided, wherein the 5' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NOs: 80 to 158, and the 3' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NOs: 159 to 237, and the nucleic acid set may be recognized by a specific transposase.
[0067] In some embodiments, the 5' recognition sequence or 3' recognition sequence includes at least one terminal inverted repeat having lengths of 1-1200 nt, 1-800 nt, 1-600 nt, 1-400 nt, 1-200 nt, 1-100 nt, 5-80 nt, 10-70 nt, or 20-60 nt.
[0068] In some embodiments, the 5' recognition sequence or 3' recognition sequence includes at least one terminal inverted repeat with lengths of 1-800 nt, 20-700 nt, 50-600 nt, 100-500 nt, 150-400 nt, 150-300 nt, or 200-260 nt.
[0069] According to one embodiment of this application, a nucleic acid set construct comprising the nucleic acid set described herein and further comprising an exogenous nucleic acid fragment may be provided. In some embodiments, the exogenous nucleic acid fragment is operably inserted into the nucleic acid set construct via a polyclonal insertion site, and there may be one or more exogenous nucleic acid fragments, which may be the same or different, and a promoter may also be inserted to control the expression of the exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment includes any gene of interest or any transferable gene, such as a gene for a native functional protein, an artificial chimeric gene, or a non-coding RNA gene. In some embodiments, the non-coding RNA (ncRNA) gene includes various RNAs with known functions and RNAs with unknown functions, such as rRNA, tRNA, small interfering RNA (siRNA), nuclear small RNA (snRNA), nucleolar small RNA (snoRNA), and microRNA (miRNA). In some embodiments, the gene for a native functional protein includes a fluorescent reporter gene, a luciferase gene, and a resistance gene. In some embodiments, the artificial chimeric gene includes a gene for a chimeric antigen receptor. In some embodiments, the fluorescent reporter gene is selected from at least one of the genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase gene is selected from at least one of the genes encoding firefly luciferase and sea kidney luciferase. In some embodiments, the resistance gene is selected from at least one of the genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, and bleomycin resistance.
[0070] In some embodiments, promoters can also be inserted into nucleic acid set constructs to control the expression of exogenous nucleic acid fragments. The promoter may be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by a host cell expressing an exogenous nucleic acid fragment. The promoter sequence includes a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter may be any nucleic acid sequence (including mutant, cleaved, and heterozygous promoters) that has transcriptional activity in selected host cells and may originate from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
[0071] In some embodiments, the nucleic acid set construct further includes an optional transcription termination sequence (i.e., a sequence recognized by a host cell to terminate transcription) for controlling the expression of an exogenous nucleic acid fragment. Any terminator that is functional in a selected host cell can be used in the present invention.
[0072] Optionally, the nucleic acid set construct may further include a suitable leader sequence (i.e., an untranslated region in mRNA important for translation in the host cell) for controlling the expression of an exogenous nucleic acid fragment. The leader sequence is operably ligated to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in selected host cells can be used in the present invention.
[0073] Optionally, the nucleic acid set construct may further include a propeptide coding region for controlling the expression of an exogenous nucleic acid fragment, the propeptide coding region encoding an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or propolypeptide. Propolypeptides are typically inactive and can be converted to mature, active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.
[0074] Optionally, a nucleic acid set construct may further include a regulatory sequence capable of regulating the expression of an exogenous nucleic acid fragment according to the growth conditions of the host cell. Examples of regulatory sequences include systems that turn gene expression on or off in response to chemical or physical stimuli, including the presence of a regulatory compound. Other examples of regulatory sequences are those that enable gene amplification. In these cases, the exogenous nucleic acid fragment should be operably ligated to the regulatory sequence.
[0075] Transition composition According to one embodiment of the present application, a composition may be provided, comprising a Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding a Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into the cellular genome; and a set of nucleic acids that can be recognized by a particular transposase or functional fragment thereof.
[0076] In some embodiments, the composition is selected from at least one of the following groups (1) to (80), and any one of the following groups (1) to (79) comprises a transposase-associated sequence and nucleic acid set. (1) The transposase-related sequence is an amino acid sequence or nucleic acid encoding an amino acid sequence, the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 80, and the 3' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 159. (2) The transposase-related sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 81, and the 3' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 160. (3) The transposase-related sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 82, and the 3' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 161. (4) The transposase-related sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 83, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 162. (5) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 84, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 163. (6) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 85, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 164. (7) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 86, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 165. (8) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 87, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 166. (9) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 88, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 167. (10) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 89, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 168. (11) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 90, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 169. (12) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 91, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 170. (13) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 92, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 171. (14) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 93, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 172. (15) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 94, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 173. (16) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 95, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 174. (17) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 96, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 175. (18) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 97, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 176. (19) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 98, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 177. (20) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 99, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 178. (21) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 100, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 179. (22) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 101, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 180. (23) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 102, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 181. (24) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 103, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 182. (25) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 104, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 183. (26) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 105, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 184. (27) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 106, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 185. (28) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 107, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 186. (29) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 108, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 187. (30) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 109, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 188. (31) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 110, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 189. (32) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 111, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 190. (33) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 112, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 191. (34) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 113, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 192. (35) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 114, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 193. (36) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 115, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 194. (37) The transposase-associated sequence is an amino acid sequence or nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 116, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 195. (38) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 117, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 196. (39) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 118, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 197. (40) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 119, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 198. (41) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 120, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 199. (42) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 121, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 200. (43) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 122, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 201. (44) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 123, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 202. (45) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 124, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 203. (46) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 125, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 204. (47) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 126, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 205. (48) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 127, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 206. (49) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 128, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 207. (50) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 129, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 208. (51) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 130, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 209. (52) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 131, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 210. (53) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 132, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 211. (54) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 133, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 212. (55) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 134, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 213. (56) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 135, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 214. (57) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 136, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 215. (58) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 137, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 216. (59) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 138, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 217. (60) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 139, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 218. (61) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 140, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 219. (62) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 141, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 220. (63) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 142, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 221. (64) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 143, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 222. (65) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 144, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 223. (66) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 145, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 224. (67) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 146, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 225. (68) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 147, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 226. (69) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 148, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 227. (70) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 149, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 228. (71) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 150, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 229. (72) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 151, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 230. (73) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 152, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 231. (74) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 153, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 232. (75) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 154, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 233. (76) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 155, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 234. (77) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 156, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 235. (78) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 157, and the 3' recognition sequence being a nucleotide sequence including the sequence shown in SEQ ID NO: 236. (79) The transposase-associated sequence is an amino acid sequence or a nucleic acid encoding an amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 158, and the 3' recognition sequence is a nucleotide sequence including the sequence shown in SEQ ID NO: 237, or (80) Any one variant from the above group (1) to (79) And, The transposase-related sequence is the amino acid sequence of the transposase variant for each group, or the nucleic acid sequence encoding the variant, and the variants are as follows: (i)~(iii): (i) at least one sequence obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the amino acid sequence of each group of transposases; (ii) at least one amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1 to 79; and (iii) At least one sequence obtained by further fusing the amino acid sequence shown in any one of sequence numbers 1 to 79 with another sequence. It has a variant sequence of the above transposase having transposase activity selected from the above.
[0077] In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a promoter. The promoter may be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by a host cell expressing the nucleic acid sequence. The promoter sequence comprises a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter may be any nucleic acid sequence (including mutant, cleaved, and heterozygous promoters) that has transcriptional activity in selected host cells and may originate from genes encoding extracellular or intracellular proteins or polypeptides homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL. In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a poly(A) sequence. Poly(A) tailing signal sequences known in the art, as well as various cleavage forms of the poly(A) tailing signal, can be used in this application.
[0078] In some embodiments, the nucleic acid set further comprises exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments are operably inserted into the nucleic acid construct via polyclonal insertion sites, and there may be one or more exogenous nucleic acid fragments, which may be the same or different, and promoters may also be inserted to control the expression of the exogenous nucleic acid fragments. In some embodiments, the exogenous nucleic acid fragments include any gene of interest or any transferable gene, e.g., a gene for a native functional protein, an artificial chimeric gene, or a non-coding RNA gene. In some embodiments, the non-coding RNA (ncRNA) gene includes a variety of RNAs with known functions and RNAs with unknown functions, such as rRNA, tRNA, small interfering RNA (siRNA), nuclear small RNA (snRNA), nucleolar small RNA (snoRNA), and microRNA (miRNA). In some embodiments, the gene for a native functional protein includes a fluorescent reporter gene, a luciferase gene, and a resistance gene. In some embodiments, the artificial chimeric gene includes a gene for a chimeric antigen receptor. In some embodiments, the fluorescence-based reporter gene includes a gene encoding a green fluorescent protein, a red fluorescent protein, a blue fluorescent protein, or a yellow fluorescent protein. In some embodiments, the luciferase gene includes a gene encoding a firefly luciferase or a sea kidney luciferase. In some embodiments, the resistance gene includes a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance. In some embodiments, a promoter can also be inserted into the nucleic acid set to control the expression of an exogenous nucleic acid fragment. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by a host cell expressing an exogenous nucleic acid fragment. The promoter sequence includes a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence (including mutant, cleaved, and heterozygous promoters) that has transcriptional activity in selected host cells and can originate from a gene encoding a homologous or heterologous extracellular or intracellular protein or polypeptide in the host cell.In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
[0079] In some embodiments, the nucleic acid encoding an amino acid sequence and / or a set of nucleic acids further includes an optional transcription termination sequence that controls the expression of an exogenous nucleic acid fragment, i.e., a sequence recognized by the host cell to terminate transcription. Any terminator that is functional in a selected host cell can be used in the present invention.
[0080] In some embodiments, the nucleic acid encoding an amino acid sequence and / or a set of nucleic acids further comprises an optional transcription termination sequence, i.e., a sequence recognized by the host cell to terminate transcription. The termination sequence is operably ligated to the 3' end of the nucleic acid sequence encoding a protein or polypeptide. Any terminator that is functional in selected host cells can be used in the present invention.
[0081] Optionally, a nucleic acid encoding an amino acid sequence and / or a set of nucleic acids may further include a suitable leader sequence, i.e., an untranslated region in mRNA that is important for translation in the host cell. The leader sequence is operably ligated to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in selected host cells can be used in the present invention.
[0082] Optionally, a nucleic acid encoding an amino acid sequence and / or a set of nucleic acids may further include a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or propolypeptide. Propolypeptides are typically inactive and can be converted to mature, active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.
[0083] Optionally, nucleic acids encoding amino acid sequences and / or sets of nucleic acids may further include regulatory sequences that can regulate polypeptide expression according to host cell growth conditions. Examples of regulatory sequences include systems that turn gene expression on or off in response to chemical or physical stimuli, including the presence of regulatory compounds. Other examples of regulatory sequences are those that enable gene amplification. In these examples, the nucleic acid sequence encoding a protein or polypeptide should be operably ligated to the regulatory sequence.
[0084] Recombinant vectors, recombinant host cells, and kits According to embodiments of this application, recombinant vectors comprising nucleic acids encoding a transposase described herein, nucleic acids described herein, nucleic acid constructs described herein, nucleic acid sets described herein, nucleic acid set constructs described herein, or compositions described herein may be provided. The recombinant vector may be any suitable vector. In some embodiments, the recombinant vector may include, but is not limited to, a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector. In some embodiments, the recombinant eukaryotic expression plasmid may include pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV. In some embodiments, the recombinant viral vector may include a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector. Recombinant vectors of the present invention can be constructed using methods well known in the art. For example, depending on the restriction sites contained in the backbone vector used, appropriate restriction sites may be added to both ends of the nucleic acid construct of the present invention, and then loaded into the backbone vector.
[0085] Embodiments of this application may provide recombinant host cells comprising a transposase described in this application, a nucleic acid encoding a transposase described in this application, a nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid set described in this application, a nucleic acid set construct described in this application, a composition described in this application, or a recombinant vector described in this application. The recombinant host cell may be any host cell on which the transposase can be used. In some embodiments, the recombinant host cell may include, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, animal cells may include mammalian cells. In some embodiments, mammalian cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / tubule / cell lines), and cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A54). 9, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK and Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1 and D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.
[0086] According to embodiments of this application, a kit comprising the transposase described in this application, a nucleic acid encoding the transposase described in this application, a nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid set described in this application, a nucleic acid set construct described in this application, a composition described in this application, a recombinant vector described in this application, or a recombinant host cell described in this application may be provided.
[0087] Method and Use The transposase-based tools and methods for large fragment gene insertion and integration provided in this application can be applied to many fields, including gene and cell therapy, molecular breeding in animals and plants, and industrial microbial engineering. In particular, in the field of cell therapy, the transposase system provided in this application can be applied to the integration of CAR sequences in cellular immunotherapy (CAR-T, CAR-NK, CAR-M, etc.), in the field of gene therapy, the transposase system provided in this application can be used to insert or integrate healthy genes into the genome of cells, thereby facilitating the treatment of diseases caused by gene mutations or gene defects, from the perspective of molecular breeding, the transposase system provided in this application can be used as a tool for breeding many crops such as rice, maize, and wheat, and can also accelerate the breeding process in animals and plants in a targeted manner, and from the perspective of industrial microbial engineering, the transposase system provided in this application can stably integrate genes into microbial chromosomes, overcoming drawbacks such as plasmid instability and easy loss in gene expression.
[0088] According to one embodiment of the present application, a method for introducing an exogenous nucleic acid fragment into the genome of a host cell may be provided, the method comprising delivering to a host cell a transposase described in the present application, a nucleic acid encoding a transposase described in the present application, a nucleic acid described in the present application, a nucleic acid construct described in the present application, a nucleic acid set described in the present application, a nucleic acid set construct described in the present application, a composition described in the present application, or a recombinant vector described in the present application.
[0089] According to one embodiment of the present application, a method for editing the genome of a host cell may be provided, the method comprising delivering a transposase described in the present application, a nucleic acid encoding a transposase described in the present application, a nucleic acid described in the present application, a nucleic acid construct described in the present application, a nucleic acid set described in the present application, a nucleic acid set construct described in the present application, a composition described in the present application, or a recombinant vector described in the present application to a host cell.
[0090] According to one embodiment of the present application, a method may be provided for obtaining a host cell containing an exogenous nucleic acid fragment in its genome, comprising delivering to a host cell a transposase described in the present application, a nucleic acid encoding a transposase described in the present application, a nucleic acid described in the present application, a nucleic acid construct described in the present application, a nucleic acid set described in the present application, a nucleic acid set construct described in the present application, a composition described in the present application, or a recombinant vector described in the present application.
[0091] The method of delivery to host cells can be any suitable method. In some embodiments, the delivery method includes, but is not limited to, cationic liposome delivery, lipoid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenovirus delivery, adeno-associated virus delivery, electroporation, Agrobacterium infection, or gene gun. Methods of cell transfection and culture are routine in the art, and suitable transfection and culture methods can be selected according to different cell types.
[0092] The host cell can be any host cell on which a transposase can be used. In some embodiments, the host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the host cell includes mammalian cells. In some embodiments, the host cell includes primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / tubule / cell lines), and cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549). This includes SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.
[0093] According to one embodiment of this application, the use of a transposase described in this application, a nucleic acid encoding a transposase described in this application, a nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid set described in this application, a nucleic acid set construct described in this application, a composition described in this application, a recombinant vector described in this application, or a recombinant host cell described in this application may be provided for introducing an exogenous nucleic acid fragment into the genome of a host cell. The host cell may be any host cell on which the transposase can be used. In some embodiments, the host cell may include, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the host cell may include mammalian cells. In some embodiments, the host cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / tubule / cell lines), and cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549). This includes SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.
[0094] According to one embodiment of this application, the use of a transposase described in this application, a nucleic acid encoding a transposase described in this application, a nucleic acid described in this application, a nucleic acid construct described in this application, a nucleic acid set described in this application, a nucleic acid set construct described in this application, a composition described in this application, a recombinant vector described in this application, or a recombinant host cell described in this application may be provided for preparing a drug or formulation for gene therapy, cell therapy, genome research, or stem cell induction and post-induction differentiation.
[0095] The various embodiments and preferences of this application described above can be combined with one another (inso that they are not essentially contradictory to each other) and are suitable for use in this application, and the various embodiments formed by such combinations are considered to be part of this application. [Examples]
[0096] The exemplary embodiments of this application are described below in conjunction with the accompanying drawings and include various details of the examples of this application for the sake of ease of understanding. It should be understood that these are illustrative only and are not intended to limit the scope of protection of this application. The scope of protection of this application is defined solely by the claims. Therefore, those skilled in the art should be aware that various changes and modifications can be made to the examples described herein without departing from the scope of this application. Similarly, for the sake of clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0097] Unless otherwise specified, the reagents and equipment used in the following examples are commercially available conventional products. Unless otherwise specified, experiments are conducted under conventional conditions or conditions recommended by the manufacturer.
[0098] Example 1: Construction of a transposon activity detection system To validate the activity of candidate transposons, a set of fluorescence-based reporter gene detection systems combined with antibiotic screening markers was established. The system was validated using two plasmid vectors, as shown in Figure 1: Plasmid 1 was a plasmid expressing transposase (Tn) containing a constitutive promoter CMV (sequence number 298) that can initiate transcription in eukaryotic cells, a candidate transposase sequence (shown in Table 1), and a poly(A) sequence (PA, sequence number 299) that terminates transcription; Plasmid 2 was a transposon donor plasmid containing a GFP gene (sequence number 300), a puromycin resistance screening gene (PuroR, sequence number 301), a promoter PGK (sequence number 302), a P2A (sequence number 303), and a Poly(A) element (sequence number 299), with transposon sequences (LTF and RTF in Figure 1, sequences shown in Table 1) that can be specifically recognized by the transposase inserted at both ends of these sequences.
[0099] [Table 1-1]
[0100] [Table 1-2]
[0101] [Table 1-3]
[0102] When both plasmids were co-transfected into HEK293T cells, transcription of the transposase gene from plasmid 1 was initiated, leading to the expression of the transposase protein. The transposase protein then recognized and bound to the transposon recognition sequence on plasmid 2. All sequences, including the transposon recognition sequence, the GFP gene, and the puromycin resistance gene, were then cleaved from the plasmid vector and incorporated into the cell genome. When the cells were continuously cultured in a medium containing a constant concentration of puromycin, only cells that experienced transposation events survived because they contained the puromycin resistance gene in their genome. The transposase activity level of the candidate transposases was reflected in the number of surviving cells or their ability to form monoclonal cells.
[0103] Methods for DNA synthesis and plasmid construction: Plasmid 1 Construction: DNA sequences corresponding to the amino acid sequence of the transposase were synthesized by Beijing Tsingke Biotech Co., Ltd. and GENERAL Biosystems (Anhui) Co., Ltd., and cloned into the plasmid vector pICOZ containing the CMV promoter element via the EcoRI site at the 5' end and the NotI site at the 3' end. The transposase gene was then transcribed in eukaryotic cells under the control of the CMV promoter and subsequently translated into a functional protein.
[0104] Plasmid 2 Construction: The transposon sequences (including terminal inversion repeats (TIRs)) are located on both sides of the open reading frame of the transposase. The left transposon fragment (LTF) contains the entire DNA sequence from the 5' target site duplication (TSD) sequence to the sequence before the transposase start codon, while the right transposon fragment (RTF) contains the entire DNA sequence from the first base after the transposase stop codon to the 3' TSD sequence. In principle, the terminal repeats on both sides recognized by the transposase are contained in the respective transposon sequences on both sides. The LTF and RTF sequences were synthesized by BGI Tech Solutions (Beijing Liuhe) Co., Ltd., and cloned into pMV plasmid vectors containing elements such as the PGK promoter, puromycin resistance gene (PuroR), P2A, green fluorescent protein (GFP) gene, and poly(A), with the LTF located upstream of the PGK promoter and the RTF located downstream of poly(A).
[0105] Plasmid 1 corresponds to one plasmid 2.
[0106] Example 2: High-throughput screening of transfer activity 2.1 Cell treatment (Day 0): HEK293T cells (commercially available) that stably express the firefly luciferase gene were established for a high-throughput screening assay. When cultured to the logarithmic growth phase, the cells were digested with 0.25% Trypsin (Thermo), dispersed into single cells, and placed in a 96-well cell culture plate pre-coated with PDL (Sigma) at a cell concentration of 1.0 × 10⁶. 4 The cells were added to each well and incubated overnight at 37°C under 5% CO2 conditions.
[0107] 2.2 Cell transfection (Day 1): Two plasmids corresponding to each transposon system were mixed at doses of 20 ng of plasmid 1 and 10 ng of plasmid 2. The mixture was then combined with the transfection reagent Lipofectamine 2000 (Thermo) in a volume ratio of 1:2 (plasmid mass (μg):transfection reagent (μL)), and allowed to stand at room temperature for 15 minutes to form a transfection complex. The transfection complexes were transferred to cell culture plates and incubated with cells. Two parallel tests were performed for each sample to be screened.
[0108] 2.3 Cell screening (Day 3) Forty-eight hours after transfection, the culture medium was replaced with DMEM (Thermo) screening medium containing 2 μg / mL puromycin (Invivogen), 10% fetal bovine serum, and 1% penicillin / streptomycin (Thermo), and the cells were cultured at 37°C and 5% CO2 for four days. Subsequently, the cells were digested into single cells with 0.25% trypsin, diluted in a 1:5 ratio, and transferred to another 96-well culture plate pre-coated with PDL. These cells were then cultured at 37°C for four days in DMEM screening medium containing 2 μg / mL puromycin, 10% fetal bovine serum, and 1% penicillin / streptomycin.
[0109] 2.4 Detection of cell viability (Day 11) 2.4.1 Preparation of detection reagent: Steady-Glo® Luciferase Assay System (Promega) was mixed with PBS in a volume ratio of 1:5. The detection reagent was prepared in a volume of 50 μL / well, and 5 mL of the detection reagent was prepared for a 96-well plate.
[0110] 2.4.2 Cells screened with puromycin for 8 days were removed from the incubator. After removing the culture medium, the detection reagent was added at a dose of 50 μL / well. After incubation in the dark at room temperature for 5 minutes, a multifunctional microplate reader with luminescence detection capabilities was used for detection. The more cells that survived after puromycin screening, the stronger the detected luminescence signal, indicating higher translocation activity of the sample.
[0111] 2.5 Statistical results During high-throughput screening, positive and negative controls were set in each plate. Based on the readings of the luminescence signals detected by the microplate reader in each well, the multiplier change in the readings of each sample (including the positive control) relative to the average reading of the negative control was calculated. The level of the calculated value indicates the level of transposase activity. Statistical results of the relative transposase activity of the transposase of this application compared to SB100X, PiggyBac, and inactive transposases are shown in Figure 2 and Table 2.
[0112] [Table 2-1]
[0113] [Table 2-2]
[0114] Example 3: Transfer Activity Assay 3.1 Cell treatment (Day 0): After culturing commercially available HEK293T cells until the logarithmic growth phase, they were digested with 0.25% trypsin (Thermo), dispersed into single cells, and placed in 24-well cell culture plates pre-coated with PDL (Sigma) at a cell concentration of 1.2 × 10⁶. 5 The cells were added to each well and incubated overnight at 37°C under 5% CO2 conditions.
[0115] 3.2 Cell transfection (Day 1): Two plasmids corresponding to each transposon system were mixed at doses of 200 ng of plasmid 1 and 100 ng of plasmid 2. The mixture was then combined with the transfection reagent Lipofectamine 2000 (Thermo) in a volume ratio of 1:2 (plasmid mass (μg):transfection reagent (μL)) and allowed to stand at room temperature for 15 minutes to form a transfection complex. The transfection complexes were transferred to cell culture plates and incubated with cells. Two parallel tests were performed for each sample to be screened.
[0116] 3.3 Cell screening (Day 3) 48 hours after transfection, the cells were digested, dispersed into single cells with 0.25% trypsin, and added to DMEM (Thermo) screening medium diluted 1:2000 with 2 μg / mL puromycin (Invivogen), 10% fetal bovine serum, and an antibiotic (1% penicillin / streptomycin, Thermo). The cells were then transferred to 6-well culture plates for further cultivation. After 10 days of continuous screening culture in puromycin-resistant medium, clones were counted and the transposase translocation activity was calculated.
[0117] 3.4. Cell staining (Day 13) Cells were screened with puromycin, cultured in 6-well plates, washed with PBS, and then fixed with 4% paraformaldehyde at room temperature for 15 minutes. The waste solution was discarded, and 0.2% methylene blue stain was added to the cells. The cells were stained at room temperature for 1 hour. The stained cell clones were washed with PBS and imaged using an imaging system (BioRad). The number of cell clones in each well was counted. The cloning and screening results of the transposase of this application in HEK293T cells are shown in Figure 3. The staining results of viable cell clones after puromycin resistance screening are shown in the figure, indicating that the transposition event occurred successfully. For each sample, Tn+ represents simultaneous transfection with the transposase plasmid and donor plasmid, and Tn- represents transfection with the donor plasmid only as a negative control.
[0118] 3.5 Statistical results The statistical results for transfection activity are shown in Figure 4 and Table 3. Tn+ represents co-transfection with both the transposase plasmid and the donor plasmid, while Tn- represents transfection with the donor plasmid only. The y-axis in the figure shows the calculated percentage of transfection activity, and the specific calculation formula was as follows: Transfection efficiency (%) = Number of cell clones per well / (Number of seeded cells per well × Transfection efficiency (GFP-positive cells %)) × 100%.
[0119] [Table 3-1]
[0120] [Table 3-2]
[0121] During the execution of all the above examples, two transposons, SB100X and PiggyBac, were used as positive controls to evaluate the transposase activity of the transposases of this application. These two transposons are recently patented commercial DNA transposons and were synthesized using the same method as in Example 1, referencing sequences reported by Ma'te's et al. (Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates, Nature Genetics, 2009 41(6):753-761) and Cary, L et al. (Transposon mutagenesis of baculoviruses: analysis of Trichoplusia ni transposon IFP2 insertions within the FP-locus of nuclear polyhedrosis viruses, Virology, 1989, 172(1):156-169), and cloned into corresponding plasmid vectors.
[0122] The statistical results of the transposase activity of the transposases of this application are shown in Figures 2 and 4. From these results, the 79 transposases of this application (TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_ B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11 , TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM _E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C 1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, and TCM_E_G12) were shown to have good transfer activity.
[0123] On the other hand, numerous transposases with inactive or low transposase activity were also found during the screening process (e.g., TCM_A_B8, TCM_A_D6, TCM_A_F7, TCM_B_A8, TCM_B_B8, TCM_B_G5, TCM_C_A8, TCM_C_A11, TCM_C_B11, TCM_C_D2, TCM_C_D8, TCM_D_G3, TCM_D_G8, TCM_E_D4, TCM_E_D5, TCM_E_E4, TCM_E_F4, TCM_E_G1, TCM_E_G2, TCM_E_G3 in Table 1 of this application).Compared to these transposases which are inactive or have low transposase activity, the transposase activity of the 79 transposases of this application is remarkably high (TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM _B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F1 0, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1 , TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TC M_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_ E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12), most of them were comparable to or better than those of the SB100X and PiggyBac.
[0124] Furthermore, Figure 5 shows the evolutionary branching diagram of the PiggyBac superfamily transposons in this application based on protein sequences. Figure 6 shows the protein sequence similarity (%) between transposons of the Tc1 / mariner superfamily in this application. The results show that these transposons cover different branches of the superfamily, including SB100X.
[0125] It should be noted that the above are merely preferred examples of this application and are not intended to limit it. Those skilled in the art can make various modifications and changes to this application. While specific embodiments have been described, it is not possible, or even currently foreseeable, that the applicant or those skilled in the art may have substitutions, modifications, changes, improvements, and substantial equivalents of the above embodiments. Therefore, the submitted and potentially modified claims are intended to cover all such substitutions, modifications, changes, improvements, and substantial equivalents. It is important to note that as the art advances, many of the elements described herein may be replaced by equivalent elements appearing after this application.
Claims
1. An isolated transposase sequence selected from isolated transposases, wherein the transposase has a variant sequence of the transposase having transposase activity in (i) or (ii) to (iv) below: (i) At least one amino acid sequence shown in any one of Sequence IDs 1 to 79; (ii) At least one sequence obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the amino acid sequence shown in any one of sequence numbers 1 to 79; (iii) at least one amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequence shown in any one of Sequence IDs 1 to 79; and (iv) At least one sequence obtained by further fusing the amino acid sequence shown in any one of sequence numbers 1 to 79 with another sequence. 。
2. The transposase according to claim 1, wherein the transposase has a transposase sequence selected from at least one of the following groups (1) to (8): (1) At least one amino acid sequence shown in any one of SEQ ID NOs: 11-36 and 59-79; (2) At least one amino acid sequence shown in any one of SEQ ID NOs: 3-4 and 48-55; (3) At least one amino acid sequence shown in any one of SEQ ID NOs. 5-8 and 40-44; (4) At least one amino acid sequence shown in any one of SEQ ID NOs: 1-2 and 45-47; (5) At least one amino acid sequence shown in any one of SEQ ID NOs: 9-10 and 56-58; (6) At least one amino acid sequence shown in any one of sequence numbers 38 to 39; (7) At least one amino acid sequence shown in any one of Sequence ID No. 37; and (8) An amino acid sequence having at least 70% identity with any one of the sequence numbers 1 to 79 of (1) to (7) above.
3. An isolated transposase wherein the transposase is of formula (I): DE(I) It contains the amino acid sequence shown, During the ceremony, D is aspartic acid, E is glutamic acid. Isolated transposase.
4. An isolated transposase, wherein the transposase is of formula (II): ︤(︸ 1 ) a ((=).、 It contains the amino acid sequence shown, During the ceremony, D is aspartic acid, H is histidine, a is the number of amino acids, and (X 1 ) is any amino acid, and a is 5. Isolated transposase.
5. An isolated transposase, wherein the transposase is of formula (III): P(X 2 )(X 3 )(III): It contains the amino acid sequence shown, During the ceremony, P is proline, X 2 is any amino acid, and X 3 It is aspartic acid or glutamic acid. Isolated transposase.
6. An isolated transposase comprising at least two of the amino acid sequences shown in formulas (I), (II), and (III).
7. An isolated transposase comprising the amino acid sequences shown in formulas (I), (II), and (III).
8. The transposase according to claim 6 or 7, wherein formula (I) and formula (II) are separated by 80 to 120 amino acids, and formula (II) and formula (III) are separated by 20 to 40 amino acids.
9. The transposase according to any one of claims 1 to 8, wherein the transposase belongs to the Tc1 / mariner superfamily.
10. The transposase according to claim 9, wherein the transposase belongs to the Tc1, Tc2, Tc4, Mariner, Tigger, Pogo, Fot1, ISRm11, or m44 family.
11. The transposase according to any one of claims 1 to 10, wherein the species source of the transposase includes arthropods, chordates, cnidarians, mollusks, or flatworms.
12. The transposase according to claim 11, wherein the species source of the transposase includes insects, ray-finned fishes, amphibians, malacostraca, arachnida, chondrichthyes, sauropsida, bivalvia, ascidiacea, hydrozoa, or rhabditophora.
13. The species source of the aforementioned transposase is Aelia acuminata, Agrypnus murinus, Albula grossodonta, Amblyraja radiata, Anthonymus grandis, Astatotilapia calliptera, Blatella germanica, Bufo gargarizans, and Carassius auratus. auratus), Cephalophoris sonnerati, Cerceris rybyensis, Cheilinus undulatus, Chelonia mydas, Clitarchus hookeri, Crassostrea gigas, Cromileptes altivelis, Cyprinus carpio, Danio rerio, Drosophylla ananasae ananassae), Drosophylla mojavensis, Epicauta chinensis, Folsomia candida, Formica aquilonia x Formica polyctena, Gasterosteus aculeatus, Gonioctena quinquepunctata, Gymnosoma rotundatum rotundatum, Harpegnathos saltator, Hemibagras wickchioideswyckioides), Homalodisca vitripennis, Hydra vulgaris, Ischnura elegans, Lampris incognitus, Laothoe populi, Latimeria chalumnae, Leistus spinibarbis, Lottia gigantea, Olea europaea subspecies europaea subsp. europaea), Orgyia antiqua, Parhyale hawaiensis, Petrochromis sp. "Moshi Yellow" AB-2019, Philaenus spumarius, Philereme vetulata, Phytophora ramorum, Polistes metricus, Rana pipiens, Rhinella marina marina), Rhodnius prolixus, Salmo salar, Schmidtea mediterranea, Seladonia tumulorum, Sesia apiformis, Sitophilus oryzae, Solenopsis invicta, Thalassophryne amazonica, Thecocalceria actangrata acutangulata, Thymallus thymallus, Tribolium castaneum, Vandiemenella biaticaThe transposase according to claim 12, comprising Zaprionus vaatica or Zaprionus bogoriensis.
14. A nucleic acid encoding a transposase according to any one of claims 1 to 13.
15. A nucleic acid construct comprising the nucleic acid described in claim 14, further comprising a promoter.
16. The nucleic acid construct according to claim 15, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
17. The nucleic acid construct according to claim 15, wherein the nucleic acid construct further comprises a poly(A) sequence.
18. A nucleic acid set comprising a 5' recognition sequence, wherein the 5' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 80 to 158.
19. A nucleic acid set comprising a 3' recognition sequence, wherein the 3' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 159 to 237.
20. A nucleic acid set comprising a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NOs: 80 to 158, and the 3' recognition sequence comprises a nucleotide sequence or a variant thereof shown in any one of SEQ ID NOs: 159 to 237, and the nucleic acid set is recognizable by a specific transposase.
21. The nucleic acid set according to any one of claims 18 to 20, wherein the 5' recognition sequence or the 3' recognition sequence includes a terminal inverted repeat.
22. The nucleic acid set according to claim 21, wherein the terminal inversion repeat has a length of at least one of 1 to 1200 nt, 1 to 800 nt, 1 to 600 nt, 1 to 400 nt, 1 to 200 nt, 1 to 100 nt, 5 to 80 nt, 10 to 70 nt, 20 to 60 nt, 20 to 700 nt, 50 to 600 nt, 100 to 500 nt, 150 to 400 nt, 150 to 300 nt, or 200 to 260 nt.
23. A nucleic acid set construct comprising the nucleic acid set according to any one of claims 18 to 22, further comprising an exogenous nucleic acid fragment.
24. The nucleic acid set construct according to claim 23, wherein the exogenous nucleic acid fragments are operably inserted into the nucleic acid set construct via polyclonal insertion sites, and there may be one or more exogenous nucleic acid fragments that are the same or different, and a promoter can be inserted to control the expression of the exogenous nucleic acid fragments.
25. The nucleic acid set construct according to claim 24, wherein the exogenous nucleic acid fragment comprises any gene of interest or any transferable gene, such as a gene for a naturally occurring functional protein, an artificial chimeric gene, or a gene for a non-coding RNA.
26. The nucleic acid set construct according to claim 25, wherein the genes for the natural functional protein include a fluorescent reporter gene, a luciferase gene, and a resistance gene.
27. The nucleic acid set construct according to claim 26, wherein the fluorescent reporter gene is selected from at least one of genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein.
28. The nucleic acid set construct according to claim 26, wherein the luciferase gene is selected from at least one of the genes encoding firefly luciferase or sea kidney luciferase.
29. The nucleic acid set construct according to claim 26, wherein the resistance gene is selected from at least one of the genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.
30. The nucleic acid set construct according to claim 25, wherein the artificial chimeric gene includes a gene for a chimeric antigen receptor.
31. The nucleic acid set construct according to claim 24, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
32. A composition, wherein the composition is A Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding the Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into the cellular genome; and A set of nucleic acids that can be recognized by a specific transposase or its functional fragment. A composition containing the following:
33. The composition is selected from at least one of the following groups (1) to (80), and any one of the following groups (1) to (79) comprises a transposase-related sequence and a nucleic acid set. (1) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 1 or a nucleic acid encoding the amino acid sequence, the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 80, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
159. (2) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 2 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 81, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
160. (3) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 3 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 82, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
161. (4) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 4 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 83, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
162. (5) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 5 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 84, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
163. (6) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 6 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 85, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
164. (7) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 7 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 86, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
165. (8) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 8 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 87, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
166. (9) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 9 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 88, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
167. (10) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 10 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 89, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
168. (11) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 11 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 90, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
169. (12) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 12 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 91, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
170. (13) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 13 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 92, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
171. (14) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 14 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 93, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
172. (15) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 15 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 94, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
173. (16) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 16 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 95, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
174. (17) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 17 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 96, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
175. (18) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 18 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 97, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
176. (19) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 19 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 98, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
177. (20) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 20 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 99, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
178. (21) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 21 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 100, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
179. (22) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 22 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 101, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
180. (23) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 23 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 102, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
181. (24) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 24 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 103, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
182. (25) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 25 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 104, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
183. (26) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 26 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 105, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
184. (27) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 27 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 106, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
185. (28) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 28 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 107, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
186. (29) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 29 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 108, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
187. (30) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 30 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 109, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
188. (31) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 31 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 110, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
189. (32) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 32 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 111, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
190. (33) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 33 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 112, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
191. (34) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 34 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 113, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
192. (35) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 35 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 114, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
193. (36) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 36 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 115, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
194. (37) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 37 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 116, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
195. (38) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 38 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 117, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
196. (39) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 39 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 118, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
197. (40) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 40 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 119, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
198. (41) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 41 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 120, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
199. (42) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 42 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 121, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
200. (43) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 43 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 122, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
201. (44) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 44 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 123, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
202. (45) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 45 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 124, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
203. (46) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 46 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 125, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
204. (47) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 47 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 126, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
205. (48) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 48 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 127, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
206. (49) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 49 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 128, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
207. (50) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 50 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 129, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
208. (51) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 51 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 130, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
209. (52) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 52 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 131, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
210. (53) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 53 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 132, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
211. (54) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 54 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 133, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
212. (55) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 55 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 134, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
213. (56) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 56 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 135, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
214. (57) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 57 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 136, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
215. (58) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 58 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 137, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
216. (59) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 59 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 138, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
217. (60) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 60 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 139, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
218. (61) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 61 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 140, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
219. (62) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 62 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 141, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
220. (63) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 63 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 142, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
221. (64) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 64 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 143, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
222. (65) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 65 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 144, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
223. (66) The transposase-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 66 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 145, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO:
224. (67) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 67 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 146, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
225. (68) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 68 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 147, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
226. (69) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 69 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 148, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
227. (70) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 70 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 149, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
228. (71) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 71 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 150, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
229. (72) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 72 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 151, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
230. (73) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 73 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 152, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
231. (74) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 74 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 153, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
232. (75) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 75 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 154, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
233. (76) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 76 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 155, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
234. (77) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 77 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 156, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
235. (78) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 78 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 157, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No.
236. (79) The transposase-related sequence is an amino acid sequence containing the sequence shown in Sequence ID No. 79 or a nucleic acid encoding the amino acid sequence, wherein the nucleic acid set includes a 5' recognition sequence and a 3' recognition sequence, the 5' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 158, and the 3' recognition sequence is a nucleotide sequence containing the sequence shown in Sequence ID No. 237, or (80) Any one variant from the aforementioned groups (1) to (79) And, The transposase-related sequence is the amino acid sequence of the variant of the transposase in each group or the nucleic acid sequence encoding the variant, and the variant is as follows (i) to (iii): (i) at least one sequence obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the amino acid sequence of the transposase of each group; (ii) at least one amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity with the amino acid sequence shown in any one of Sequence IDs 1 to 79; and (iii) At least one sequence obtained by fusing the amino acid sequence shown in any one of Sequence IDs 1 to 79 with another sequence. The composition according to claim 32, comprising a variant sequence of the transposase having transposase activity selected from the above.
34. The composition according to claim 32 or 33, wherein the nucleic acid set further comprises a promoter.
35. The composition according to claim 34, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
36. The composition according to claim 32 or 33, wherein the nucleic acid set further comprises a poly(A) sequence.
37. The composition according to any one of claims 32 to 36, wherein the nucleic acid set further comprises an exogenous nucleic acid fragment.
38. The composition according to claim 37, wherein the exogenous nucleic acid fragment is operably inserted into the nucleic acid set via a polyclonal insertion site, and there may be one or more exogenous nucleic acid fragments that are the same or different, and a promoter can be inserted to control the expression of the exogenous nucleic acid fragment.
39. The composition according to claim 38, wherein the exogenous nucleic acid fragment comprises any gene of interest or any transferable gene, such as a gene for a natural functional protein, an artificial chimeric gene, or a gene for a non-coding RNA.
40. The composition according to claim 39, wherein the gene for the natural functional protein comprises a fluorescent reporter gene, a luciferase gene, or a resistance gene.
41. The composition according to claim 40, wherein the fluorescence-based reporter gene comprises a gene encoding a green fluorescent protein, a red fluorescent protein, a blue fluorescent protein, or a yellow fluorescent protein.
42. The composition according to claim 40, wherein the luciferase gene comprises a gene encoding firefly luciferase or sea kidney luciferase.
43. The composition according to claim 41, wherein the resistance gene comprises a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.
44. The composition according to claim 39, wherein the artificial chimeric gene comprises a gene for a chimeric antigen receptor.
45. The composition according to claim 38, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
46. A recombinant vector comprising a nucleic acid encoding a transposase according to any one of claims 1 to 12, a nucleic acid according to claim 13, a nucleic acid construct according to any one of claims 14 to 16, a nucleic acid set according to any one of claims 17 to 21, a nucleic acid set construct according to any one of claims 22 to 30, or a composition according to any one of claims 32 to 45.
47. The recombinant vector according to claim 46, wherein the recombinant vector comprises a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector.
48. The recombinant vector according to claim 47, wherein the recombinant eukaryotic expression plasmid comprises pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV.
49. The recombinant vector according to claim 47, wherein the recombinant viral vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector.
50. A recombinant host cell comprising a transposase according to any one of claims 1 to 12, a nucleic acid encoding the transposase according to any one of claims 1 to 12, a nucleic acid according to claim 13, a nucleic acid construct according to any one of claims 14 to 16, a nucleic acid set according to any one of claims 17 to 21, a nucleic acid set construct according to any one of claims 22 to 30, a composition according to any one of claims 32 to 45, or a recombinant vector according to any one of claims 46 to 49.
51. The recombinant host cell according to claim 50, wherein the recombinant host cell includes an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell.
52. The recombinant host cell according to claim 51, wherein the animal cell includes a mammalian cell.
53. The mammalian cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / tubule / cell lines), and cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-1).
5. Recombinant host cells according to claim 52, comprising HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.
54. A method for introducing an exogenous nucleic acid fragment into the genome of a host cell, comprising delivering to the host cell a transposase according to any one of claims 1 to 12, a nucleic acid encoding the transposase according to any one of claims 1 to 12, a nucleic acid according to claim 13, a nucleic acid construct according to any one of claims 14 to 16, a nucleic acid set according to any one of claims 17 to 21, a nucleic acid set construct according to any one of claims 22 to 30, a composition according to any one of claims 32 to 45, or a recombinant vector according to any one of claims 46 to 49.
55. A method for editing the genome of a host cell, comprising delivering to the host cell a transposase according to any one of claims 1 to 12, a nucleic acid encoding the transposase according to any one of claims 1 to 12, a nucleic acid according to claim 13, a nucleic acid construct according to any one of claims 14 to 16, a nucleic acid set according to any one of claims 17 to 21, a nucleic acid set construct according to any one of claims 22 to 30, a composition according to any one of claims 32 to 45, or a recombinant vector according to any one of claims 46 to 49.
56. A method for obtaining a host cell containing an exogenous nucleic acid fragment in its genome, comprising delivering to a host cell a transposase according to any one of claims 1 to 12, a nucleic acid encoding the transposase according to any one of claims 1 to 12, a nucleic acid according to claim 13, a nucleic acid construct according to any one of claims 14 to 16, a nucleic acid set according to any one of claims 17 to 21, a nucleic acid set construct according to any one of claims 22 to 30, a composition according to any one of claims 32 to 45, or a recombinant vector according to any one of claims 46 to 49.
57. The method according to any one of claims 54 to 56, wherein the delivery method includes cationic liposome delivery, lipoid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retrovirus delivery, lentivirus delivery, adenovirus delivery, adeno-associated virus delivery, electroporation, Agrobacterium infection, or gene gun.
58. The method according to any one of claims 54 to 56, wherein the host cell includes an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell.
59. The method according to claim 58, wherein the host cell includes a mammalian cell.
60. The host cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / tubule / cell lines), and cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT- The method according to claim 59, comprising 15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK and Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1 and D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.
61. Use of a transposase according to any one of claims 1 to 12, a nucleic acid encoding the transposase according to any one of claims 1 to 12, the nucleic acid according to claim 13, the nucleic acid construct according to any one of claims 14 to 16, the nucleic acid set according to any one of claims 17 to 21, the nucleic acid set construct according to any one of claims 22 to 30, the composition according to any one of claims 32 to 45, the recombinant vector according to any one of claims 46 to 49, or the recombinant host cell according to any one of claims 50 to 53, for introducing an exogenous nucleic acid fragment into the genome of a host cell.
62. The use according to claim 61, wherein the host cell includes an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell, or a bacterial cell.
63. The use according to claim 62, wherein the host cell includes a mammalian cell.
64. The host cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / tubule / cell lines), and cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT- The use according to claim 63, comprising 15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK and Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1 and D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells.
65. Use of a transposase according to any one of claims 1 to 12, a nucleic acid encoding the transposase according to any one of claims 1 to 12, a nucleic acid according to claim 13, a nucleic acid construct according to any one of claims 14 to 16, a nucleic acid set according to any one of claims 17 to 21, a nucleic acid set construct according to any one of claims 22 to 30, a composition according to any one of claims 32 to 45, a recombinant vector according to any one of claims 46 to 49, or a recombinant host cell according to any one of claims 50 to 53, for preparing drugs or preparations for gene therapy, cell therapy, genome research, or stem cell induction and post-induction differentiation.
66. A kit comprising a transposase according to any one of claims 1 to 12, a nucleic acid encoding the transposase according to any one of claims 1 to 12, a nucleic acid according to claim 13, a nucleic acid construct according to any one of claims 14 to 16, a nucleic acid set according to any one of claims 17 to 21, a nucleic acid set construct according to any one of claims 22 to 30, a composition according to any one of claims 32 to 45, a recombinant vector according to any one of claims 46 to 49, or a recombinant host cell according to any one of claims 50 to 53.