Separated transposase AG-TEF11 and application thereof
By providing a highly active transposase tool, the problem of non-viral gene integration in existing technologies is solved, the stable integration and expression of exogenous nucleic acid fragments in host cells is achieved, the defects mediated by viruses are avoided, and a better gene therapy strategy is provided.
Patent Information
- Application Number
- CN202510735313.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-27
- Filing Date
- 2024-03-12
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-03-12
AI Technical Summary
The existing technology lacks efficient non-viral gene integration tools. Virus-mediated gene therapy has problems such as random carcinogenic risks, limited exogenous gene size, immunogenicity and high production complexity, which limit the stable integration and expression of large gene fragments.
Provided is an isolated transposase with high transposition activity, comprising a specific amino acid sequence or a variant thereof, for use in non-viral gene integration, which avoids the drawbacks of viral integration by recognizing and integrating exogenous nucleic acid fragments into the host cell genome.
It achieves efficient and stable integration and expression of exogenous nucleic acid fragments in the host cell genome, avoids virus-mediated immunogenicity and randomness risks, and provides more flexible gene therapy options.
Smart Images

Figure CN120591233A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202480001971X, filed on March 12, 2024, with the invention name “A separated transposase and its use”.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to Chinese Patent Application No. 202310304787X filed with the State Intellectual Property Office of China on March 27, 2023, and the entire contents of this Chinese patent application are hereby incorporated by reference in their entirety for all purposes. Technical Field
[0004] The present application relates to the field of molecular biology, and specifically to an isolated transposase and its use. The present application also specifically relates to: a nucleic acid and a nucleic acid construct encoding the transposase, a nucleic acid group and a nucleic acid group construct, and a composition, a recombinant vector, a recombinant host cell and a kit comprising the transposase. The present application also specifically relates to: a method for introducing an exogenous nucleic acid fragment into a host cell genome, a method for editing a host cell genome, and a method for obtaining a host cell whose genome contains an exogenous nucleic acid fragment. The present application also specifically relates to the use of the transposase, the nucleic acid and the nucleic acid construct, the nucleic acid group and the nucleic acid group construct, the composition, the recombinant vector, or the recombinant host cell in introducing an exogenous nucleic acid fragment gene into the host cell genome, or in the preparation of a drug or preparation for gene therapy, cell therapy, genomic research, or stem cell induction and post-induction differentiation. Background Art
[0005] Transposons are DNA sequences that can insert or excise within a genome, thereby transferring their own sequence or intact copies within or between genomes. Transposons are divided into two main categories. This article focuses on type II transposons (DNA transposons), which consist of terminal inverted repeats (TIRs) at both ends and a gene encoding a transposase. Transposons employ a "cut-and-paste" transposition mechanism, excising DNA from chromosomes and inserting it directly into other parts of the genome.
[0006] Transposases are sequence-specific DNA-binding proteins expressed from DNA transposon sequences. They contain catalytic domains that mediate DNA breakage and ligation. Transposases recognize and bind to the TIRs at both ends of the transposon, forming a protrusion complex. The transposition activity of a transposon is primarily determined by the expression level and activity of the transposase. Therefore, DNA transposons with high transposase activity are a key requirement for developing transposon-based gene editing tools.
[0007] The insertion and integration of large gene fragments has important applications in gene therapy, plant and animal molecular breeding, and industrial microbial engineering. Currently, industry lacks effective tools and systems for large gene insertion and integration. In recent years, the scientific community has developed several tools and methods capable of inserting and integrating large genes, but these methods still present several challenges. For example, lentivirus or retrovirus are most commonly used to integrate gene sequences in cellular immunotherapy and gene therapy for genetic diseases, and several therapeutic products based on them have been used to treat tumors and genetic diseases (Aiuti, A., Roncarolo, MG and Naldini, L. (2017) Gene therapy for ADA-SCID, the first marketing approval of an ex vivo gene therapy in Europe: paving the road for the next generation of advanced therapy medicinal products. [ADA-SCID gene therapy, the first marketing approval of an ex vivo gene therapy in Europe: paving the road for the next generation of advanced therapy medicinal products.] EMBO Mol. Med. [EMBO Molecular Medicine] 9, 737–740; Aiuti, A. et al. (2009) Gene therapy for immunodeficiency due to adenosine deaminase deficiency. [Gene therapy for immunodeficiency due to adenosine deaminase deficiency] N. Engl. J. Med. [New England Journal of Medicine] 360, 447–458). However, the use of viruses for large-scale gene integration has some potential application limitations: first, the random nature of viral integration into the genome poses a risk of carcinogenesis; second, the size of the exogenous gene that a virus can carry is also limited, making it unfavorable for the transfer of large therapeutic genes; third, the immunogenicity of viruses may affect the long-term expression of exogenous therapeutic genes and the ability to administer them again; and fourth, viral production requires the use of living cells, making quality control and downstream processing of such products more complex and costly, presenting certain disadvantages in terms of industrialization. Therefore, non-viral large-scale gene integration can avoid the various drawbacks of viral integration and become a valuable tool in gene therapy.
[0008] As a non-viral gene integration tool, DNA transposons not only enable the integration and stable expression of large exogenous gene fragments in the host genome, but also avoid negative effects such as immunogenicity. Therefore, some transposons have been used in gene therapy. Although transposons have been shown to be widespread across the prokaryotic and eukaryotic spectrum, a large number of transposon fragments have become silent and inactive during evolution to maintain genomic stability. Currently, a few highly active and valuable transposon tools, such as SleepingBeauty (SB), PiggyBac (PB), and Tol2, are being used in gene therapy research. Therefore, the discovery of more highly active transposon tools and the verification and testing of their functions will provide more, better, and more flexible options for the development of gene therapy strategies.
[0009] It should be noted that the approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply because it is included in this section. Similarly, unless otherwise indicated, the issues mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0010] Based on this, in order to seek more advanced and more efficient non-viral gene integration tools, the present application provides an isolated transposase, wherein the transposase has a transposase sequence selected from the following (i) or a variant sequence of the aforementioned transposase with transposase activity among (ii) to (iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-79; (ii) at least one of the sequences obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NOs: 1-79; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95% or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-79; and (iv) at least one of the sequences obtained by further fusion of the amino acid sequence shown in any one of SEQ ID NOs: 1-79 with other sequences. The transposase provided in this application has comparable or even higher transposition activity than the currently widely used Sleeping Beauty (SB) and PiggyBac (PB), providing more or better options for the development of gene integration tools.
[0011] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises an amino acid sequence represented by formula (I):
[0012] DE (I)
[0013] wherein D is aspartic acid; and E is glutamic acid.
[0014] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises an amino acid sequence represented by formula (II):
[0015] D(X1) a H (II)
[0016] wherein D is aspartic acid; H is histidine; a is the number of amino acids; and (X1) is any amino acid, and a is 5.
[0017] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises an amino acid sequence represented by formula (III):
[0018] P(X2)(X3)(III)
[0019] wherein P is proline; X2 is any amino acid; and X3 is aspartic acid or glutamic acid.
[0020] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises at least two of the amino acid sequences represented by formula (I), formula (II) and formula (III).
[0021] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises the amino acid sequences shown in Formula (I), Formula (II) and Formula (III).
[0022] According to an embodiment of the present application, a nucleic acid may be provided, wherein the nucleic acid encodes the transposase described in the present application.
[0023] According to an embodiment of the present application, a nucleic acid construct may be provided, which comprises the nucleic acid according to the present application and a promoter.
[0024] According to an embodiment of the present application, a nucleic acid set may be provided, wherein the 5' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 80-158.
[0025] According to an embodiment of the present application, a nucleic acid set may be provided, wherein the 3' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 159-237.
[0026] According to an embodiment of the present application, a nucleic acid group can be provided, which comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises a nucleotide sequence shown in any one of SEQ ID NOs: 80-158 or a variant thereof, and the 3' recognition sequence comprises a nucleotide sequence shown in any one of SEQ ID NOs: 159-237 or a variant thereof, and the nucleic acid group can be recognized by a specific transposase.
[0027] According to an embodiment of the present application, a nucleic acid group construct may be provided, wherein the nucleic acid group construct comprises the nucleic acid group described in the present application and further comprises an exogenous nucleic acid fragment.
[0028] According to an embodiment of the present application, a composition may be provided, wherein the composition comprises: a Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding the Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into a cell genome; and a nucleic acid group, wherein the nucleic acid group can be recognized by a specific transposase or a functional fragment thereof.
[0029] According to an embodiment of the present application, a recombinant vector may be provided, wherein the recombinant vector comprises a nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, or the composition described in the present application.
[0030] According to an embodiment of the present application, a recombinant host cell can be provided, wherein the recombinant host cell comprises the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.
[0031] According to an embodiment of the present application, a method for introducing an exogenous nucleic acid fragment into a host cell genome can be provided, wherein the method comprises: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.
[0032] According to an embodiment of the present application, a method for editing the genome of a host cell can be provided, wherein the method comprises: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.
[0033] According to an embodiment of the present application, a method for obtaining a host cell containing an exogenous nucleic acid fragment in its genome can be provided, wherein the method comprises: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.
[0034] According to the embodiments of the present application, the use of the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application for introducing exogenous nucleic acid fragments into the host cell genome can be provided.
[0035] According to the embodiments of the present application, the use of the transposase described herein, the nucleic acid encoding the transposase described herein, the nucleic acid described herein, the nucleic acid construct described herein, the nucleic acid group described herein, the nucleic acid group construct described herein, the composition described herein, the recombinant vector described herein, or the recombinant host cell described herein in the preparation of drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post-induced differentiation can be provided.
[0036] According to an embodiment of the present application, a kit can be provided, wherein the kit comprises the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.
[0037] It should be understood that the content described in this section is not intended to identify the key or important features of the examples of this application, nor is it intended to limit the scope of this application. Other features of this application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.
[0039] Figure 1 The figure shows a schematic diagram of the dual-plasmid vector in the transposon activity detection system in Example 1. Plasmid 1 is a plasmid expressing transposase (Tn), and plasmid 2 is a transposon donor plasmid.
[0040] Figure 2The results of Example 2 show that TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D 6. TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B _F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_ C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM Relative transposition efficiency results of TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12, SB100X, and PiggyBac.
[0041] Figure 3In Example 3, TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_ B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C _C3, TCM_D_B11, TCM_D_G9, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TC Clone screening results for M_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F11, TCM_E_G4, TCM_E_G6, TCM_E_G7, and TCM_E_G9. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, while Tn- indicates transfection of the donor plasmid alone.
[0042] Figure 4The results of Example 3 show that TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_C6, TCM_A_D7, TCM_A_D8, TCM_A_G9, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_C6 CM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B _F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_ C3, TCM_D_B11, TCM_D_G9, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, T CM_E_B10, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_ Transposition activity assay results for E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F11, TCM_E_G4, TCM_E_G6, TCM_E_G7, and TCM_E_G9. Tn+ indicates co-transfection of the transposase plasmid and the donor plasmid, while Tn- indicates transfection of the donor plasmid alone.
[0043] Figure 5Shows the phylogenetic tree based on protein sequences of TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12 and SB100X in Example 2.
[0044] Figure 6The protein sequences of Example 2 are shown as TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_C8, TCM_B_C9, TCM_A_D1 M_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B _G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A 4. TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, T CM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_ Protein sequence similarity results between E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12, and SB100X. DETAILED DESCRIPTION
[0045] Unless otherwise specified or contradicted by the context, the terms or expressions used herein should be read in conjunction with the entire disclosure and as understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0046] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein to refer to a polymeric form of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogs thereof.
[0047] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to amino acid polymers of any length. Thus, polypeptides, oligopeptides, proteins, antibodies, and enzymes are all included within the definition of polypeptide.
[0048] As described in this application, a "fragment" of a sequence refers to a portion of the sequence. For example, a fragment of a nucleic acid sequence refers to a portion of the nucleic acid sequence, and a fragment of an amino acid sequence refers to a portion of the amino acid sequence.
[0049] As used herein, a "variant" of a sequence is a polynucleotide or polypeptide that differs from a reference polynucleotide or polypeptide, respectively, but retains essential properties. A typical variant of a polynucleotide differs from another reference polynucleotide in nucleic acid sequence, and this difference in nucleic acid sequence may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. A typical variant of a polypeptide differs from another reference polypeptide in amino acid sequence. Generally, the differences are limited so that the sequences of the reference polypeptide and the variant are generally very similar and identical in many regions. The variant polypeptide and the reference polypeptide may differ in amino acid sequence by one or more substitutions, additions, or deletions in any combination. The substituted or inserted amino acid residues may or may not be residues encoded by the genetic code. Polynucleotide or polypeptide variants may be naturally occurring, such as allelic variations, or they may be unknown naturally occurring variants. Non-naturally occurring polynucleotide and polypeptide variants can be produced by mutagenesis techniques, direct synthesis, and other recombinant methods known to the skilled artisan.
[0050] Amino acids are often classified by the properties of their side chains. For example, a side chain can make an amino acid a weak acid (e.g., amino acids D and E) or a weak base (e.g., amino acids K, R, and H); if the side chain is polar, the amino acid becomes a hydrophile (e.g., amino acids L and I); or if the side chain is nonpolar, the amino acid becomes a hydrophobe (e.g., amino acids S and C).
[0051] As used herein, the term "family" refers to a group of nucleic acids or proteins with high structural similarity that are derived from a common ancestor through replication and mutation, and which typically have related or even identical functions. The term "superfamily" refers to a group of nucleic acids or proteins with roughly identical structures that are derived from a common ancestor through replication and mutation, but which belong to different families and typically have different functions.
[0052] As used herein, the term "transposase" refers to a polypeptide that catalyzes the excision of a transposon (comprising an exogenous nucleic acid and flanked by transposase recognition sequences) from a first nucleic acid (a vector comprising a transposase recognition sequence and an exogenous nucleic acid) and its integration into a second nucleic acid, i.e., a target site (e.g., genomic or extrachromosomal DNA in a cell comprising a target site duplication (TSD) sequence). In some embodiments, the transposase binds to at least one inverted terminal repeat (TIR).
[0053] As used herein, the term "recognition sequence" refers to nucleic acid sequences located at both ends of a transposable element and flanking a first transposable nucleic acid sequence. The recognition sequence located 5' to the first nucleic acid sequence is referred to as a 5' recognition sequence, and the recognition sequence located 3' to the first nucleic acid sequence is referred to as a 3' recognition sequence. In some embodiments, the recognition sequence comprises at least one inverted terminal repeat sequence that can bind to a transposase.
[0054] As used herein, the term "nucleic acid construct" is defined herein as a single-stranded or double-stranded nucleic acid molecule, preferably an artificially constructed nucleic acid molecule. Optionally, the nucleic acid construct further comprises one or more operably linked regulatory sequences capable of directing expression of the coding sequence in a suitable host cell under compatible conditions. The term "expression" should be understood to encompass any step involved in protein or polypeptide production, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion. The term "regulatory sequence" encompasses all components necessary or advantageous for expression of the polypeptide / protein of the present application. Each regulatory sequence may be naturally occurring or exogenous to the nucleic acid sequence encoding the protein or polypeptide. These regulatory sequences include, but are not limited to, a leader sequence, a polyadenylation sequence, a propeptide sequence, a promoter, a signal sequence, and a transcription terminator. At a minimum, a regulatory sequence includes a promoter and initiation and termination signals for transcription and translation. The regulatory sequences may be provided with linkers to introduce specific restriction sites for ligating the regulatory sequence to the coding region of the nucleic acid sequence encoding the protein or polypeptide.
[0055] The term "promoter" as used herein refers to a polynucleotide sequence that can control the transcription of a coding sequence. The promoter sequence includes a specific sequence that is sufficient to enable RNA polymerase recognition, binding, and transcription initiation. In addition, the promoter sequence can include a sequence that optionally regulates the recognition, binding, and transcription initiation activity of RNA polymerase in the nucleic acid construct or nucleic acid group construct provided herein. The promoter can affect the transcription of a gene that is located on the same nucleic acid molecule as the promoter or a gene that is located on a different nucleic acid molecule than the promoter.
[0056] As used herein, the term "exogenous nucleic acid fragment" includes any gene of interest or any gene capable of being transposed, or a fragment thereof. In some non-limiting embodiments, the exogenous nucleic acid fragment is of a different origin than the terminal repeat sequence, for example, a nucleic acid sequence isolated from an organism other than the inverted terminal repeat sequence, i.e., it is an exogenous nucleic acid fragment relative to the inverted terminal repeat sequence. In some non-limiting embodiments, the exogenous nucleic acid fragment is of a different origin than the host cell, for example, a nucleic acid sequence isolated from an organism other than the host cell, i.e., it is an exogenous nucleic acid fragment relative to the host cell.
[0057] As used herein, the term "host cell" includes, but is not limited to, animal cells, plant cells, algae cells, fungal cells, yeast cells, or bacterial cells. This term includes the progeny of an original cell into which an exogenous nucleic acid fragment has been introduced. Exemplary host cells include human embryonic kidney cells, HEK293T cells. It should be understood that the progeny of a single parental cell are not necessarily identical to the original parent in morphology or in terms of genome or total DNA complement, due to natural, accidental, or intentional mutations.
[0058] As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is attached. Examples of vectors include, but are not limited to, plasmids, viruses, bacteria, bacteriophages, and insertable DNA segments. The term "plasmid" refers to a circular double-stranded DNA that is capable of accepting exogenous nucleic acid segments and replicating in prokaryotic or eukaryotic cells.
[0059] transposase
[0060] The present application provides an isolated transposase, wherein the transposase has a transposase sequence selected from the following (i) or a variant sequence of the aforementioned transposase having transposase activity among (ii) to (iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-79; (ii) at least one of the sequences obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids from the amino acid sequence shown in any one of SEQ ID NOs: 1-79; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95% or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-79; and (iv) at least one of the sequences obtained by further fusion of the amino acid sequence shown in any one of SEQ ID NOs: 1-79 with other sequences.
[0061] In some embodiments, the transposase has a transposase sequence selected from at least one of the following groups (1)-(8): (1) at least one of the amino acid sequences set forth in any one of SEQ ID NOs: 11-36 and 59-79; (2) at least one of the amino acid sequences set forth in any one of SEQ ID NOs: 3-4 and 48-55; (3) at least one of the amino acid sequences set forth in any one of SEQ ID NOs: 5-8 and 40-44; (4) at least one of the amino acid sequences set forth in any one of SEQ ID NOs: 1-2 and 45-47; (5) at least one of the amino acid sequences set forth in any one of SEQ ID NOs: 9-10 and 56-58; (6) at least one of the amino acid sequences set forth in any one of SEQ ID NOs: 38-39; (7) the amino acid sequence set forth in SEQ ID NO: 37; and (8) an amino acid sequence having at least 70% identity with any one of SEQ ID NOs: 1-79 in the aforementioned (1)-(7).
[0062] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises an amino acid sequence represented by formula (I):
[0063] DE(I)
[0064] wherein D is aspartic acid; and E is glutamic acid.
[0065] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises an amino acid sequence represented by formula (II):
[0066] D(X1) a H(II)
[0067] wherein D is aspartic acid; H is histidine; a is the number of amino acids; and (X1) is any amino acid, and a is 5.
[0068] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises an amino acid sequence represented by formula (III):
[0069] P(X2)(X3)(III)
[0070] wherein P is proline; X2 is any amino acid; and X3 is aspartic acid or glutamic acid.
[0071] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises at least two of the amino acid sequences represented by formula (I), formula (II) and formula (III).
[0072] According to an embodiment of the present application, an isolated transposase may be provided, wherein the transposase comprises the amino acid sequences shown in Formula (I), Formula (II) and Formula (III).
[0073] In some embodiments, Formula (I) and Formula (II) are separated by 80 to 120 amino acids, and Formula (II) and Formula (III) are separated by 20 to 40 amino acids.
[0074] In some embodiments, the transposase belongs to the Tc1 / mariner superfamily.
[0075] In some embodiments, the transposase belongs to the Tc1, Tc2, Tc4, Mariner, Tigger, Pogo, Fot1, ISRm11, or m44 family.
[0076] In some embodiments, the species of the transposase include Arthropoda, Chordata, Cnidaria, Mollusca, or Platyhelminthes. In some embodiments, the species of the transposase include Insecta, Actinopterygii, Amphibia, Malacora, Arachnida, Chondrichthyes, Sauropoda, Bivalvia, Ascidian, Hydrozoa, or Turbellaria. In some embodiments, the species of origin of the transposase include Aelia acuminata, Agrypnus murinus, Albula glossodonta, Amblyraja radiata, Anthonomus grandis, Astatotilapia calliptera, Blattella germanica, Bufo gargarizans, Carassius auratus, Cephalophoris sonnerati, Cerceris rybyensis, Cheilinus undulatus, Chelonia mydas, Clitarchus hookeri, Crassostrea gigas, Cromileptes altivelis, Cyprinus carpio, Danio rerio, and others. rerio), Drosophila ananassae, Drosophila mojavensis, Epicauta chinensis, and White Folsomia candida, Northern Ant Hybrid Bald-backed Forest Ant (Formicaaquilonia x Formica polyctena), Three-spined Stickleback (Gasterosteus aculeatus), Five-spotted Horned Leaf Beetle (Gonioctena quinquepunctata), Common Membrane-bellied Fly (Gymnosoma rotundatum), Sickle Ant (Harpegnathos saltator), Hemibagrus wyckioides, Sharpshooter Leafhopper (Homalodisca vitripennis), Hydra (Hydra vulgaris), Ischnura elegans, Pacific Small-eyed Moonfish (Lampris incognitus), Poplar Moth (Laothoe populi), Latimeria chalumnae, Leistus spinibarbis, Lotus Snail (Lottia gigantea), Olive (Olea europaea) subsp. europaea), Orgyia antiqua, Parhyale hawaiensis, Petrochromis sp.'moshi yellow') AB-2019, forage cicada (Philaenus spumarius), gray-winged inchworm (Philereme vetulata), sudden oak death fungus (Phytophthora ramorum), long-legged wasp (Polistes metricus), leopard frog (Rana pipiens), giant cane toad (Rhinella marina), red cone bug (Rhodnius prolixus), Atlantic salmon (Salmo salar), Mediterranean planarian (Schmidtea mediterranea), mountain tunnel wasp (Seladonia tumulorum), bee-shaped clearwing moth (Sesiaapiformis), rice weevil (Sitophilus oryzae), invasive red imported fire ant (Solenopsis invicta), Amazon canine toadfish (Thalassophryne amazonica), Thecocarcelia acutangulata, grayling (Thymallus thymallus), red flour beetle (Tribolium castaneum), Vandiemenellaviatica, or Zaprionus bogoriensis.
[0077] According to an embodiment of the present application, a nucleic acid may be provided, wherein the nucleic acid encodes the transposase described in the present application.
[0078] According to embodiments of the present application, a nucleic acid construct may be provided, comprising a nucleic acid encoding the transposase described herein. In some embodiments, the nucleic acid construct further comprises a promoter. The promoter may be any suitable promoter sequence, i.e., a nucleic acid sequence that is recognized by a host cell expressing the nucleic acid sequence. The promoter sequence contains a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter may be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutant, truncated, and hybrid promoters, and may be derived from a gene encoding an extracellular or intracellular protein or polypeptide that is homologous or heterologous to the host cell. In some embodiments, the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
[0079] In some embodiments, the nucleic acid construct further comprises a polyadenylation [poly(A)] signal sequence. Poly(A) tailing signal sequences well known in the art and various truncated forms of poly(A) tailing signals can be used in the present application.
[0080] In some embodiments, the nucleic acid construct further comprises any transcription termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3' end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.
[0081] Optionally, the nucleic acid construct may also include a suitable leader sequence, i.e., an untranslated region of mRNA that is very important for translation of the host cell. The leader sequence may be operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell may be used in the present invention.
[0082] Optionally, the nucleic acid construct may further comprise a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is referred to as a proenzyme or propolypeptide. A propolypeptide is generally inactive and can be converted to a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.
[0083] Optionally, the nucleic acid construct may also include regulatory sequences that regulate polypeptide expression according to the growth of the host cell. Examples of regulatory sequences are systems that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to open or close gene expression. Other examples of regulatory sequences are those that enable gene amplification. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.
[0084] Nucleic acid constructs
[0085] According to an embodiment of the present application, a nucleic acid set may be provided, comprising a 5' recognition sequence, wherein the 5' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 80-158.
[0086] According to an embodiment of the present application, a nucleic acid set may be provided, comprising a 3' recognition sequence, wherein the 3' recognition sequence comprises at least one of the nucleotide sequences shown in SEQ ID NOs: 159-237.
[0087] According to an embodiment of the present application, a nucleic acid group can be provided, which comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises a nucleotide sequence shown in any one of SEQ ID NOs: 80-158 or a variant thereof, and the 3' recognition sequence comprises a nucleotide sequence shown in any one of SEQ ID NOs: 159-237 or a variant thereof, and the nucleic acid group can be recognized by a specific transposase.
[0088] In some embodiments, the 5' recognition sequence or the 3' recognition sequence comprises an inverted terminal repeat sequence having a length of at least one of 1-1200 nt, 1-800 nt, 1-600 nt, 1-400 nt, 1-200 nt, 1-100 nt, 5-80 nt, 10-70 nt, or 20-60 nt.
[0089] In some embodiments, the 5' recognition sequence or the 3' recognition sequence comprises an inverted terminal repeat sequence having a length of at least one of 1-800 nt, 20-700 nt, 50-600 nt, 100-500 nt, 150-400 nt, 150-300 nt, or 200-260 nt.
[0090] According to an embodiment of the present application, a nucleic acid group construct can be provided, comprising the nucleic acid group described herein and an exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment is operably inserted into the nucleic acid group construct via a multiple cloning insertion site. The exogenous nucleic acid fragment may be one or more and may be the same or different; a promoter may also be inserted to control the expression of the exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment includes any gene of interest or any gene capable of transposition, such as a naturally occurring functional protein gene, an artificial chimeric gene, or a non-coding RNA (ncRNA) gene. In some embodiments, non-coding RNA genes include rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA), among other RNAs with known functions and RNAs of unknown functions. In some embodiments, naturally occurring functional protein genes include fluorescence-based reporter genes, luciferase genes, and resistance genes. In some embodiments, artificial chimeric genes include chimeric antigen receptor genes. In some embodiments, the fluorescence-based reporter gene is selected from at least one of genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase gene is selected from at least one of genes encoding firefly luciferase and Renilla luciferase. In some embodiments, the resistance gene is selected from at least one of genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, and bleomycin resistance.
[0091] In certain embodiments, the nucleic acid group construct can also insert a promoter to control the expression of the exogenous nucleic acid fragment. The promoter can be any suitable promoter sequence, i.e., a nucleotide sequence that can be recognized by the host cell expressing the exogenous nucleic acid fragment. The promoter sequence contains a transcriptional regulatory sequence that mediates protein or polypeptide expression. The promoter can be any nucleotide sequence with transcriptional activity in the selected host cell, including mutant, truncated and hybrid promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides that are homologous or heterologous to the host cell. In certain embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
[0092] In some embodiments, the nucleic acid construct further comprises any transcription termination sequence (i.e., a sequence that can be recognized by the host cell to terminate transcription) to control the expression of the exogenous nucleic acid fragment. Any terminator that can function in the selected host cell can be used in the present invention.
[0093] Optionally, the nucleic acid construct may also include a suitable leader sequence (i.e., an untranslated region of an mRNA that is important for translation in the host cell) to control the expression of the exogenous nucleic acid fragment. The leader sequence may be operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell may be used in the present invention.
[0094] Optionally, the nucleic acid construct may further comprise a propeptide coding region to control expression of the exogenous nucleic acid fragment. The propeptide coding region encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is referred to as a zymogen or propolypeptide. A propolypeptide is generally inactive and can be converted to a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.
[0095] Optionally, the nucleic acid construct may further comprise regulatory sequences that regulate the expression of the exogenous nucleic acid fragment according to the growth of the host cell. Examples of regulatory sequences are systems that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to open or close gene expression. Other examples of regulatory sequences are those that enable gene amplification. In these examples, the exogenous nucleic acid fragment should be operably linked to the regulatory sequence.
[0096] Transposition composition
[0097] According to an embodiment of the present application, a composition may be provided, wherein the composition comprises: a Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding the Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into a cell genome; and a nucleic acid group, wherein the nucleic acid group can be recognized by a specific transposase or a functional fragment thereof.
[0098] In some embodiments, the composition is selected from at least one of the following groups (1)-(80), any one of the following groups (1)-(79) comprising: a transposase-associated sequence and a nucleic acid group,
[0099] (1) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 1 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 80, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 159;
[0100] (2) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 2 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 81, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 160;
[0101] (3) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 3 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 82, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 161;
[0102] (4) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 4 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 83, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 162;
[0103] (5) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 5 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 84, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 163;
[0104] (6) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 6 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 85, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 164;
[0105] (7) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 7 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 86, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 165;
[0106] (8) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 8 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 87, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 166;
[0107] (9) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 9 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 88, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 167;
[0108] (10) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 10 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 89, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 168;
[0109] (11) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 11 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 90, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 169;
[0110] (12) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 12 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 91, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 170;
[0111] (13) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 13 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 92, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 171;
[0112] (14) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 14 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 93, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 172;
[0113] (15) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 15 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 94, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 173;
[0114] (16) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 16 or a nucleic acid encoding the amino acid sequence; and the nucleic acid set comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 95, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 174;
[0115] (17) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 17 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 96, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 175;
[0116] (18) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 18 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 97, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 176;
[0117] (19) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 19 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 98, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 177;
[0118] (20) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 20 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 99, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 178;
[0119] (21) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 21 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 100, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 179;
[0120] (22) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 22 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 101, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 180;
[0121] (23) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 23 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 102, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 181;
[0122] (24) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 24 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 103, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 182;
[0123] (25) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 25 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 104, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 183;
[0124] (26) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 26 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 105, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 184;
[0125] (27) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 27 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 106, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 185;
[0126] (28) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 28 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 107, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 186;
[0127] (29) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 29 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 108, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 187;
[0128] (30) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 30 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 109, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 188;
[0129] (31) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 31 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 110, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 189;
[0130] (32) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 32 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 111, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 190;
[0131] (33) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 33 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 112, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 191;
[0132] (34) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 34 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 113, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 192;
[0133] (35) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 35 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 114, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 193;
[0134] (36) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 36 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 115, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 194;
[0135] (37) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 37 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 116, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 195;
[0136] (38) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 38 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 117, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 196;
[0137] (39) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 39 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 118, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 197;
[0138] (40) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 40 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 119, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 198;
[0139] (41) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 41 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 120, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 199;
[0140] (42) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 42 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 121, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 200;
[0141] (43) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 43 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 122, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 201;
[0142] (44) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 44 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 123, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 202;
[0143] (45) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 45 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 124, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 203;
[0144] (46) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 46 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 125, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 204;
[0145] (47) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 47 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 126, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 205;
[0146] (48) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 48 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 127, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 206;
[0147] (49) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 49 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 128, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 207;
[0148] (50) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 50 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 129, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 208;
[0149] (51) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 51 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 130, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 209;
[0150] (52) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 52 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 131, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 210;
[0151] (53) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 53 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 132, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 211;
[0152] (54) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 54 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 133, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 212;
[0153] (55) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 55 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 134, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 213;
[0154] (56) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 56 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 135, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 214;
[0155] (57) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 57 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 136, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 215;
[0156] (58) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 58 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 137, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 216;
[0157] (59) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 59 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 138, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 217;
[0158] (60) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 60 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 139, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 218;
[0159] (61) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 61 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 140, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 219;
[0160] (62) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 62 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 141, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 220;
[0161] (63) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 63 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 142, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 221;
[0162] (64) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 64 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 143, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 222;
[0163] (65) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 65 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 144, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 223;
[0164] (66) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 66 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 145, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 224;
[0165] (67) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 67 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 146, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 225;
[0166] (68) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 68 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 147, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 226;
[0167] (69) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 69 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 148, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 227;
[0168] (70) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 70 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 149, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 228;
[0169] (71) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 71 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 150, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 229;
[0170] (72) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 72 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 151, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 230;
[0171] (73) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 73 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 152, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 231;
[0172] (74) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 74 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 153, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 232;
[0173] (75) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 75 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 154, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 233;
[0174] (76) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 76 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 155, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 234;
[0175] (77) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 77 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 156, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 235;
[0176] (78) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 78 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 157, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 236;
[0177] (79) The transposase-associated sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 79 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 158, and the 3' recognition sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 237; or
[0178] (80) A variant of any one of the aforementioned groups (1) to (79),
[0179] The transposase-related sequence is an amino acid sequence of a variant of each group of transposases or a nucleic acid sequence encoding the variant, wherein the variant has a variant sequence of the aforementioned transposase having transposase activity selected from the following (i)-(iii):
[0180] (i) at least one of the sequences obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the amino acid sequence of each group of transposases;
[0181] (ii) at least one of the amino acid sequences that is at least 70%, 80%, 90%, 95% or 99% identical to the amino acid sequence shown in any one of SEQ ID NOs: 1-79; and
[0182] (iii) at least one of the sequences obtained by further fusion of the amino acid sequence represented by any one of SEQ ID NOs: 1-79 with another sequence.
[0183] In certain embodiments, the nucleic acid encoding the amino acid sequence further comprises a promoter. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by the host cell expressing the nucleic acid sequence. The promoter sequence contains a transcriptional regulatory sequence that mediates protein or polypeptide expression. The promoter can be any nucleic acid sequence with transcriptional activity in the selected host cell, including mutant, truncated and hybrid promoters, and can be derived from genes encoding extracellular or intracellular proteins or polypeptides that are homologous or heterologous to the host cell. In certain embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL. In some embodiments, the nucleic acid encoding the amino acid sequence further comprises a poly(A) sequence. Poly(A) tailing signal sequences well known in the art and various truncated forms of poly(A) tailing signals can be used in the present application.
[0184] In some embodiments, the nucleic acid set further comprises an exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment is operably inserted into the nucleic acid construct via a multiple cloning insertion site. The exogenous nucleic acid fragment may be one or more and may be the same or different. A promoter may also be inserted to control the expression of the exogenous nucleic acid fragment. In some embodiments, the exogenous nucleic acid fragment includes any gene of interest or any gene capable of transposition, such as a naturally occurring functional protein gene, an artificial chimeric gene, or a non-coding RNA (ncRNA) gene. In some embodiments, non-coding RNA genes include a variety of RNAs with known functions and RNAs of unknown functions, such as rRNA, tRNA, small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), and microRNA (miRNA). In some embodiments, naturally occurring functional protein genes include fluorescence-based reporter genes, luciferase genes, and resistance genes. In some embodiments, artificial chimeric genes include chimeric antigen receptor genes. In some embodiments, fluorescence-based reporter genes include genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, luciferase genes include genes encoding firefly luciferase or Renilla luciferase. In some embodiments, the resistance gene includes a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance. In some embodiments, the nucleic acid group can also be inserted into a promoter to control the expression of the exogenous nucleic acid fragment. The promoter can be any suitable promoter sequence, that is, a nucleic acid sequence that can be recognized by the host cell expressing the exogenous nucleic acid fragment. The promoter sequence contains a transcriptional regulatory sequence that mediates the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence that has transcriptional activity in the selected host cell, including mutant, truncated, and hybrid promoters, and can be derived from a gene encoding an extracellular or intracellular protein or polypeptide that is homologous or heterologous to the host cell. In some embodiments, the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
[0185] In some embodiments, the nucleic acid and / or nucleic acid group encoding the amino acid sequence further comprises any transcription termination sequence to control the expression of the exogenous nucleic acid fragment, i.e., a sequence that can be recognized by the host cell to terminate transcription. Any terminator that can function in the selected host cell can be used in the present invention.
[0186] In some embodiments, the nucleic acid and / or nucleic acid group encoding the amino acid sequence further comprises any transcription termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3' end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.
[0187] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may also include a suitable leader sequence, i.e., an untranslated region of an mRNA that is important for translation in the host cell. The leader sequence may be operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in the selected host cell may be used in the present invention.
[0188] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may further comprise a propeptide coding region that encodes the amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is referred to as a proenzyme or propolypeptide. A propolypeptide is generally inactive and can be converted to a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide.
[0189] Optionally, the nucleic acid and / or nucleic acid group encoding the amino acid sequence may also include regulatory sequences that regulate the expression of the polypeptide according to the growth of the host cell. Examples of regulatory sequences are those that respond to chemical or physical stimuli (including in the presence of regulatory compounds) to open or close gene expression systems. Other examples of regulatory sequences are those that enable gene amplification. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.
[0190] Recombinant vectors, recombinant host cells and kits
[0191] According to an embodiment of the present application, a recombinant vector may be provided, wherein the recombinant vector comprises a nucleic acid encoding a transposase described herein, a nucleic acid described herein, a nucleic acid construct described herein, a nucleic acid group described herein, a nucleic acid group construct described herein, or a composition described herein. The recombinant vector may be any suitable vector. In some embodiments, the recombinant vector includes, but is not limited to, a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector. In some embodiments, the recombinant eukaryotic expression plasmid includes pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV. In some embodiments, the recombinant viral vector includes a recombinant adenoviral vector, a recombinant adeno-associated viral vector, a recombinant retroviral vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector. The recombinant vector of the present invention may be constructed using methods well known in the art. For example, appropriate restriction sites may be added to both ends of the nucleic acid construct of the present invention, based on the restriction sites contained in the backbone vector used, and then loaded into the backbone vector.
[0192] According to an embodiment of the present application, a recombinant host cell can be provided, wherein the recombinant host cell comprises a transposase as described herein, a nucleic acid encoding a transposase as described herein, a nucleic acid as described herein, a nucleic acid construct as described herein, a nucleic acid group as described herein, a nucleic acid group construct as described herein, a composition as described herein, or a recombinant vector as described herein. The recombinant host cell can be any host cell to which the transposase can be applied. In some embodiments, the recombinant host cell includes but is not limited to an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the animal cell comprises a mammalian cell. In some embodiments, mammalian cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines), cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW 620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and differentiated cells thereof, or induced pluripotent stem cell lines and differentiated cells thereof.
[0193] According to an embodiment of the present application, a kit can be provided, wherein the kit comprises the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, the recombinant vector described in the present application, or the recombinant host cell described in the present application.
[0194] Methods and uses
[0195] The large-fragment gene insertion and integration tools and methods based on transposase provided in this application can be applied to multiple fields such as gene and cell therapy, animal and plant molecular breeding, and industrial microbial transformation. In particular, in the field of cell therapy, the transposition system provided in this application can be applied to the integration of CAR sequences in cell immunotherapy (CAR-T, CAR-NK, CAR-M, etc.); in the field of gene therapy, the transposition system provided in this application can be used to insert or integrate healthy genes into the cell genome, thereby facilitating the treatment of diseases caused by gene mutations or gene defects; in molecular breeding, the transposition system provided in this application can be used as a tool for breeding many crops such as rice, corn, and wheat, and can also be used to accelerate the breeding process of animals and plants in a targeted manner; in industrial microbial transformation, due to the defects of plasmids in gene expression such as instability and easy loss, the transposition system provided in this application can stably integrate genes into the chromosomes of microorganisms.
[0196] According to an embodiment of the present application, a method for introducing an exogenous nucleic acid fragment into a host cell genome can be provided, wherein the method comprises: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.
[0197] According to an embodiment of the present application, a method for editing the genome of a host cell can be provided, wherein the method comprises: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.
[0198] According to an embodiment of the present application, a method for obtaining a host cell containing an exogenous nucleic acid fragment in its genome can be provided, wherein the method comprises: delivering the transposase described in the present application, the nucleic acid encoding the transposase described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the nucleic acid group described in the present application, the nucleic acid group construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application into the host cell.
[0199] The method of delivery into the host cell can be any suitable method. In some embodiments, delivery methods include, but are not limited to, cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated viral delivery, electroporation, Agrobacterium infection, or gene gun delivery. Cell transfection and culture methods are conventional methods in the art, and appropriate transfection and culture methods can be selected based on the cell type.
[0200] The host cell can be any host cell to which the transposase can be applied. In some embodiments, the host cell includes but is not limited to an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the host cell includes a mammalian cell. In some embodiments, the host cell includes a primary cell (e.g., a mesenchymal stem cell, an endothelial cell, an epithelial cell, a fibroblast, a keratinocyte, a melanocyte, a smooth muscle cell, an immune cell), an immortalized cell line (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell line), a cancer cell line (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW62, etc.), or a combination thereof. 0, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and differentiated cells thereof, or induced pluripotent stem cell lines and differentiated cells thereof.
[0201] According to embodiments of the present application, there may be provided a transposase described herein, a nucleic acid encoding a transposase described herein, a nucleic acid described herein, a nucleic acid construct described herein, a nucleic acid group described herein, a nucleic acid group construct described herein, a composition described herein, a recombinant vector described herein, or a recombinant host cell described herein for the introduction of an exogenous nucleic acid fragment into a host cell genome. The host cell may be any host cell to which the transposase can be applied. In some embodiments, the host cell includes but is not limited to an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell. In some embodiments, the host cell comprises a mammalian cell. In some embodiments, host cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines), cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW62, 0, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and differentiated cells thereof, or induced pluripotent stem cell lines and differentiated cells thereof.
[0202] According to the embodiments of the present application, the use of the transposase described herein, the nucleic acid encoding the transposase described herein, the nucleic acid described herein, the nucleic acid construct described herein, the nucleic acid group described herein, the nucleic acid group construct described herein, the composition described herein, the recombinant vector described herein, or the recombinant host cell described herein in the preparation of drugs or preparations for gene therapy, cell therapy, genomic research, or stem cell induction and post-induced differentiation can be provided.
[0203] The various embodiments and preferences for the present application described above can be combined with each other (as long as they are not inherently contradictory to each other) and are applicable to the purposes of the present application. The various embodiments formed by such combinations are considered to be part of the present application.
[0204] Example
[0205] The following describes exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. It should be understood that they are considered to be merely exemplary and are in no way intended to limit the scope of protection of the present application. The scope of protection of the present application is defined solely by the claims. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0206] Unless otherwise stated, all reagents and instruments used in the following examples are commercially available conventional products. Unless otherwise stated, experiments were performed according to conventional conditions or conditions recommended by the manufacturer.
[0207] Example 1: Construction of a transposon activity detection system
[0208] We have established a detection system based on fluorescence reporter gene combined with antibiotic selection marker to verify the activity of candidate transposons. This system uses a double plasmid vector for verification, such as Figure 1 As shown: Plasmid 1 is a plasmid for expressing transposase (Tn), comprising a constitutive promoter CMV (sequence shown in SEQ ID NO: 298) capable of initiating transcription in eukaryotic cells, a sequence of a candidate transposase (as shown in Table 1), and a poly(A) sequence (PA, sequence shown in SEQ ID NO: 299) for terminating transcription; Plasmid 2 is a transposon donor plasmid, comprising a GFP gene (sequence shown in SEQ ID NO: 300), a puromycin resistance selection gene (PuroR, sequence shown in SEQ ID NO: 301), a promoter PGK (sequence shown in SEQ ID NO: 302), a P2A (sequence shown in SEQ ID NO: 303), and a poly(A) element (sequence shown in SEQ ID NO: 299), wherein transposon sequences specifically recognized by the transposase are inserted at both ends of these sequences ( Figure 1 The LTF and RTF in the ,sequence are shown in Table 1).
[0209] Table 1 Sequences related to plasmid construction
[0210]
[0211]
[0212]
[0213]
[0214] When the two plasmids are co-transfected into HEK293T cells, the transposase gene on plasmid 1 initiates transcription and produces the transposase protein. This protein then recognizes and binds to the transposon recognition sequence on plasmid 2, excising the transposon recognition sequence, along with all other sequences, including the GFP gene and the puromycin resistance gene, from the plasmid vector and integrating them into the cell genome. When cells are continuously cultured in medium containing a certain concentration of puromycin, only cells that have undergone transposition survive because they contain the puromycin resistance gene in their genomes. The level of transposition activity of a candidate transposase is reflected by the number of surviving cells or their ability to form monoclonal clones.
[0215] DNA synthesis and plasmid construction methods:
[0216] Construction of Plasmid 1: The amino acid sequence of the transposase was commissioned to Beijing Tsingke Biotech Co., Ltd. and GENERAL Biosystems (Anhui) Co., Ltd. to synthesize the corresponding DNA sequence. The sequence was then cloned into the plasmid vector pICOZ, which already contained a CMV promoter element, via the 5' EcoRI site and the 3' NotI site. This allowed the transposase gene to be transcribed in eukaryotic cells under the control of the CMV promoter and subsequently translated into a functional protein.
[0217] Construction of Plasmid 2: The transposon sequence (including the inverted terminal repeats (TIRs)) flanks the transposase open reading frame. The left transposon fragment (LTF) encompasses the entire DNA sequence from the target site repeat (TSD) sequence at the 5' end to the sequence preceding the transposase start codon, while the right transposon fragment (RTF) encompasses the entire DNA sequence from the first base after the transposase stop codon to the TSD sequence at the 3' end. In principle, the terminal repeats recognized by the transposase are contained within the flanking transposon sequences. BGI Tech Solutions (Beijing Liuhe) Co., Ltd. was commissioned to synthesize the LTF and RTF sequences. These sequences were cloned into the pMV plasmid vector, which already contains the PGK promoter, the puromycin resistance gene (PuroR), P2A, the green fluorescent protein (GFP) gene, and poly(A) residues. The LTF is positioned upstream of the PGK promoter and the RTF is downstream of the poly(A) residue.
[0218] Plasmid 1 and plasmid 2 correspond one to one.
[0219] Example 2: High-throughput screening of transposition activity
[0220] 2.1 Cell treatment (Day 0):
[0221] HEK293T cells (commercially purchased) stably expressing the firefly luciferase gene were established for high-throughput screening assays. After the cells were cultured to the logarithmic growth phase, they were digested with 0.25% trypsin (Thermo) and dissociated into single cells. The cells were plated at a cell concentration of 1.0 × 10 4 Cells were added per well into a 96-well cell culture plate pre-coated with PDL (Sigma), and cultured overnight at 37°C with 5% CO2.
[0222] 2.2 Cell transfection (Day 1):
[0223] The two plasmids corresponding to each transposon system were mixed at a dose of 20 ng for plasmid 1 and 10 ng for plasmid 2. The mixture was then mixed with the transfection reagent Lipofectamine 2000 (Thermo Fisher Scientific) at a ratio of 1:2:1 (plasmid mass (μg) to transfection reagent volume (μL). The mixture was allowed to stand at room temperature for 15 minutes to form a transfection complex. The transfection complex was transferred to a cell culture plate and incubated with cells. Two replicates were performed for each sample to be screened.
[0224] 2.3 Cell Screening (Day 3)
[0225] 48 h after transfection, the culture medium was replaced with DMEM (Thermo Fisher Scientific) selection medium (the selection medium contains 2 μg / mL puromycin (Invivogen), 10% fetal bovine serum and 1% penicillin / streptomycin (Thermo Fisher Scientific)) and cultured for 4 days at 37°C, 5% CO2. Then, the cells were digested into single cells with 0.25% trypsin, diluted at a ratio of 1:5, transferred to another 96-well culture plate pre-coated with PDL, and cultured at 37°C in DMEM selection medium containing 2 μg / mL puromycin, 10% fetal bovine serum and 1% penicillin / streptomycin for 4 days.
[0226] 2.4 Cell viability detection (Day 11)
[0227] 2.4.1 Preparation of detection reagents: Steady- Luciferase assay system (Promega) was mixed with PBS at a volume ratio of 1:5. The detection reagent was prepared at a dose of 50 μL / well, and 5 mL of the detection reagent was prepared for a 96-well plate.
[0228] 2.4.2 Remove the cells from the incubator after 8 days of puromycin selection. After removing the culture medium, add 50 μL / well of the detection reagent. After incubation in the dark at room temperature for 5 minutes, assay using a multi-function microplate reader with luminescence detection. The more cells that survive puromycin selection, the stronger the luminescence signal detected, indicating a higher transposition activity in the sample.
[0229] 2.5 Statistical Results
[0230] During high-throughput screening, positive and negative controls were set on each plate. Based on the reading values of the luminescent signal detected by the microplate reader in each well, the fold change of the reading value of each sample (including the positive control) relative to the average reading value of the negative control was calculated. The level of the calculated value indicates the level of transposition activity. The statistical results of the relative transposition activity of the transposase of the present application compared to SB100X, PiggyBac and inactive transposase are shown in Figure 2. Figure 2 and as shown in Table 2.
[0231] Table 2 Results of relative transposition activity in Example 2
[0232]
[0233]
[0234]
[0235]
[0236] Example 3: Transposition activity assay
[0237] 3.1 Cell treatment (Day 0):
[0238] HEK293T cells (commercially purchased) were cultured to the logarithmic growth phase and then digested with 0.25% trypsin (Thermo Fisher Scientific) to separate into single cells. The cells were plated at a concentration of 1.2 × 10 5 Cells were added per well into a 24-well cell culture plate pre-coated with PDL (Sigma) and cultured overnight at 37°C with 5% CO2.
[0239] 3.2 Cell transfection (Day 1):
[0240] The two plasmids corresponding to each transposon system were mixed at a dose of 200 ng for plasmid 1 and 100 ng for plasmid 2. The mixture was then mixed with the transfection reagent Lipofectamine 2000 (Thermo Fisher Scientific) at a ratio of 1:2 (plasmid mass (μg): transfection reagent volume (μL). The mixture was allowed to stand at room temperature for 15 minutes to form a transfection complex. The transfection complex was transferred to a cell culture plate and incubated with cells. Each sample to be screened was tested in duplicate.
[0241] 3.3 Cell screening (Day 3)
[0242] 48 hours after transfection, cells were dissociated into single cells using 0.25% trypsin. Cells were diluted 1:2000 in DMEM (Thermo Fisher Scientific) selection medium supplemented with 2 μg / mL puromycin (Invitrogen), 10% fetal bovine serum, and antibiotics (1% penicillin / streptomycin, Thermo Fisher Scientific) and transferred to 6-well culture plates for further culture. After 10 days of selection and culture using puromycin-resistant medium, colonies were counted and the transposition activity of the transposase was calculated.
[0243] 3.4 Cell staining (Day 13)
[0244] The cells selected by puromycin and cultured in 6-well plates were washed with PBS and then fixed with 4% paraformaldehyde for 15 minutes at room temperature. The waste liquid was discarded and 0.2% methylene blue staining solution was added to the cells. The cells were stained for 1 hour at room temperature. The stained cell clones were washed with PBS and photographed in an imaging system (BioRad). The number of cell clones in each well was counted. The results of the clone screening of transposase in HEK293T cells in this application are shown in Figure 2. Figure 3 The figure shows staining of surviving cell clones after puromycin resistance selection, indicating successful transposition. Tn+ indicates co-transfection of the transposase plasmid and donor plasmid, while Tn- indicates transfection of the donor plasmid alone, serving as a negative control for each transposase sample.
[0245] 3.5 Statistical Results
[0246] The statistical results of transposition activity are as follows Figure 4 As shown in Table 3. Tn+ indicates co-transfection of the transposase plasmid and donor plasmid, and Tn- indicates transfection of the donor plasmid alone. The y-axis of the figure shows the calculated percentage of transposition activity, calculated as follows: Transposition efficiency (%) = number of cell colonies per well / (number of cells plated per well × transfection efficiency (% GFP-positive cells)) × 100%.
[0247] Table 3 Statistical results of transposition activity in Example 3
[0248]
[0249]
[0250]
[0251] During the implementation of all the above examples, two transposons, SB100X and PiggyBac, were used as positive controls to evaluate the transposition activity of the transposase of the present application. These two transposons are commercially used DNA transposons that are currently protected by patents. The sequences reported in Lajos Ma'te's et al. (Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates, Nature Genetics, 2009 41(6):753-761) and Cary, LC et al. (Transposon mutagenesis of baculoviruses: analysis of Trichoplusia nitransposon IFP2 insertions within the FP-locus of nuclear polyhedrosisviruses, Virology, 1989, 172(1):156-169) were synthesized and cloned into the corresponding plasmid vectors according to the same method as in Example 1.
[0252] The statistical results of the transposition activity of the transposase in this application are as follows Figure 2 and Figure 4As shown above. The above results indicate that the 79 transposases (TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D7, TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E_A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TCM_E_E7, TCM_E_E8, TCM_E_E10, TCM_E_E11, TCM_E_E12, TCM_E_F2, TCM_E_F3, TCM_E_F5, TCM_E_F7, TCM_E_F8, TCM_E_F11, TCM_E_G4, TCM_E_G5, TCM_E_G6, TCM_E_G7, TCM_E_G9, TCM_E_G12) of the present application have good transposition activity.
[0253] At the same time, a large number of transposases with inactive or low transposition activity were also found during the screening process (for example, TCM_A_B8, TCM_A_D6, TCM_A_F7, TCM_B_A8, TCM_B_B8, TCM_B_G5, TCM_C_A8, TCM_C_A11, TCM_C_B11, TCM_C_D2, TCM_C_D8, TCM_D_G3, TCM_D_G8, TCM_E_D4, TCM_E_D5, TCM_E_E4, TCM_E_F4, TCM_E_G1, TCM_E_G2, and TCM_E_G3 in Table 1 of this application).Compared with these transposases with no activity or low transposition activity, the 79 transposases of the present application (TCM_A_A2, TCM_A_B2, TCM_A_B5, TCM_A_C6, TCM_A_C9, TCM_A_D1, TCM_A_D3, TCM_A_D4, TCM_A_F8, TCM_A_G6, TCM_B_C4, TCM_B_C10, TCM_B_C11, TCM_B_C12, TCM_B_D1, TCM_B_D4, TCM_B_D5, TCM_B_D6, TCM_B_D 7. TCM_B_D8, TCM_B_D10, TCM_B_E1, TCM_B_E4, TCM_B_E6, TCM_B_F2, TCM_B_F3, TCM_B_F5, TCM_B_F10, TCM_B_F11, TCM_B_ G2, TCM_B_G10, TCM_B_G11, TCM_C_A12, TCM_C_B2, TCM_C_B8, TCM_C_C3, TCM_D_B11, TCM_D_G9, TCM_E_A1, TCM_E_A4, TCM_E _A5, TCM_E_A6, TCM_E_A7, TCM_E_A8, TCM_E_B6, TCM_E_B8, TCM_E_B10, TCM_E_B11, TCM_E_B12, TCM_E_C1, TCM_E_C2, TCM_ E_C3, TCM_E_C5, TCM_E_C6, TCM_E_C8, TCM_E_C10, TCM_E_C11, TCM_E_C12, TCM_E_D6, TCM_E_D7, TCM_E_D12, TCM_E_E6, TC TCM_E_G7, TCM_E_G9, TCM_E_G12) had significantly higher transposition activities, and most of them had comparable or better transposition activities than SB100X and PiggyBac.
[0254] also, Figure 5 The figure shows the evolutionary branching diagram of the transposons of the Tc1 / mariner superfamily based on protein sequences in this application. Figure 6 The protein sequence similarity (%) between the transposons of the Tc1 / mariner superfamily in this application is shown. The results show that these transposons cover different branches of the superfamily, including SB100X.
[0255] It should be noted that the above are only preferred examples of the present application and are not intended to limit the present application. For those of ordinary skill in the art, various modifications and changes can be made to the present application. Although specific embodiments have been described, it is possible that there are or it is currently impossible for the applicant or other persons skilled in the art to foresee substitutions, modifications, changes, improvements and substantial equivalents of the above-described embodiments. Therefore, the appended claims submitted and the claims that may be amended are intended to cover all such substitutions, modifications, changes, improvements and substantial equivalents. It is important that as technology evolves, many of the elements described herein can be replaced by equivalent elements that appear after the present application.
Claims
1. An isolated transposase, wherein the transposase comprises the amino acid sequence shown in SEQ ID NO:
73.
2. The transposase according to claim 1, wherein the transposase belongs to the Tc1 / mariner superfamily. The transposase according to claim 2 , wherein the transposase belongs to the Tc1 family. The transposase according to any one of claims 1 to 3 , wherein the species origin of the transposase comprises Arthropoda. The transposase according to claim 4 , wherein the species origin of the transposase comprises Insecta. The transposase according to claim 5 , wherein the species of origin of the transposase comprises Gymnosoma rotundatum .
7. A nucleic acid, wherein the nucleic acid encodes the transposase according to any one of claims 1 to 6. A nucleic acid construct comprising the nucleic acid according to claim 7 and further comprising a promoter.
9. The nucleic acid construct of claim 8, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
10. The nucleic acid construct according to claim 9, wherein the nucleic acid construct further comprises a poly(A) sequence.
11. A nucleic acid set comprising a 5' recognition sequence, wherein the 5' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO:
152.
12. A nucleic acid set comprising a 3' recognition sequence, wherein the 3' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO:
231.
13. A nucleic acid group comprising a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO: 152, the 3' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO: 231, and the nucleic acid group can be recognized by a specific transposase.
14. The nucleic acid set according to any one of claims 11 to 13, wherein the 5' recognition sequence or the 3' recognition sequence comprises an inverted terminal repeat sequence.
15. The nucleic acid set according to claim 14, wherein the length of the terminal inverted repeat sequence is at least one of 1-1200nt, 1-800nt, 1-600nt, 1-400nt, 1-200nt, 1-100nt, 5-80nt, 10-70nt, 20-60nt, 20-700nt, 50-600nt, 100-500nt, 150-400nt, 150-300nt, or 200-260nt. 16 . A nucleic acid group construct comprising the nucleic acid group according to claim 11 , and further comprising an exogenous nucleic acid fragment.
17. The nucleic acid group construct according to claim 16, wherein the exogenous nucleic acid fragment is operably inserted into the nucleic acid group construct through a multiple cloning insertion site, and the exogenous nucleic acid fragment may be one or more and may be the same or different; a promoter may also be inserted to control the expression of the exogenous nucleic acid fragment.
18. The nucleic acid construct according to claim 17, wherein the exogenous nucleic acid fragment comprises a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.
19. The nucleic acid group construct according to claim 18, wherein the natural functional protein genes include a fluorescence-based reporter gene, a luciferase gene, and a resistance gene.
20. The nucleic acid construct according to claim 19, wherein the fluorescence-based reporter gene is selected from at least one of genes encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein or yellow fluorescent protein.
21. The nucleic acid construct according to claim 19, wherein the luciferase gene is selected from at least one of genes encoding firefly luciferase or Renilla luciferase.
22. The nucleic acid group construct according to claim 19, wherein the resistance gene is selected from at least one of genes encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance or bleomycin resistance.
23. The nucleic acid recombinant construct of claim 18, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.
24. The nucleic acid group construct of claim 17, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
25. A composition, wherein the composition comprises: A Tc1 / mariner superfamily transposase or a functional fragment thereof, or a nucleic acid encoding the Tc1 / mariner superfamily transposase or a functional fragment thereof, wherein the transposase or the functional fragment thereof has the function of catalyzing the insertion of an exogenous nucleic acid fragment into a cell genome; and A nucleic acid group, wherein the nucleic acid group can be recognized by a specific transposase or a functional fragment thereof; The composition comprises: a transposase-associated sequence and a nucleic acid group, wherein the transposase-associated sequence comprises the amino acid sequence shown in SEQ ID NO: 73 or a nucleic acid encoding the amino acid sequence; and the nucleic acid group comprises a 5' recognition sequence and a 3' recognition sequence, wherein the 5' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO: 152, and the 3' recognition sequence comprises the nucleotide sequence shown in SEQ ID NO:
231.
26. The composition of claim 25, wherein the set of nucleic acids further comprises a promoter.
27. The composition of claim 26, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
28. The composition of claim 25, wherein the set of nucleic acids further comprises a poly(A) sequence.
29. The composition of claim 25, wherein the set of nucleic acids further comprises exogenous nucleic acid fragments.
30. The composition according to claim 29, wherein the exogenous nucleic acid fragment is operably inserted into the nucleic acid group through a multiple cloning insertion site, and the exogenous nucleic acid fragment may be one or more and may be the same or different; a promoter may also be inserted to control the expression of the exogenous nucleic acid fragment.
31. The composition of claim 30, wherein the exogenous nucleic acid fragment comprises a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene.
32. The composition of claim 31, wherein the natural functional protein gene comprises a fluorescence-based reporter gene, a luciferase gene, or a resistance gene.
33. The composition of claim 32, wherein the fluorescence-based reporter gene comprises a gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein.
34. The composition of claim 32, wherein the luciferase gene comprises a gene encoding firefly luciferase or Renilla luciferase.
35. The composition of claim 32, wherein the resistance gene comprises a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.
36. The composition of claim 31, wherein the artificial chimeric gene comprises a chimeric antigen receptor gene.
37. The composition of claim 30, wherein the promoter comprises CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
38. A recombinant vector, wherein the recombinant vector comprises a nucleic acid encoding a transposase according to any one of claims 1-6, a nucleic acid according to claim 7, a nucleic acid construct according to any one of claims 8-10, a nucleic acid group according to any one of claims 11-15, a nucleic acid group construct according to any one of claims 16-24, or a composition according to any one of claims 25-37.
39. The recombinant vector according to claim 38, wherein the recombinant vector comprises a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector.
40. The recombinant vector of claim 39, wherein the recombinant eukaryotic expression plasmid comprises pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV.
41. The recombinant vector of claim 39, wherein the recombinant viral vector comprises a recombinant adenoviral vector, a recombinant adeno-associated viral vector, a recombinant retroviral vector, a recombinant herpes simplex viral vector, or a recombinant vaccinia viral vector.
42. A recombinant host cell, wherein the recombinant host cell comprises the transposase of any one of claims 1-6, a nucleic acid encoding the transposase of any one of claims 1-6, a nucleic acid of claim 7, a nucleic acid construct of any one of claims 8-10, a nucleic acid group of any one of claims 11-15, a nucleic acid group construct of any one of claims 16-24, a composition of any one of claims 25-37, or a recombinant vector of any one of claims 38-41.
43. The recombinant host cell of claim 42, wherein the recombinant host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell.
44. The recombinant host cell of claim 43, wherein the animal cell comprises a mammalian cell.
45. The recombinant host cell of claim 44, wherein the mammalian cell comprises a primary cell, an immortalized cell line, a cancer cell line, an embryonic stem cell line and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof.
46. The recombinant host cell of claim 45, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.
47. A method for introducing an exogenous nucleic acid fragment into a host cell genome, wherein the method comprises: The transposase according to any one of claims 1-6, the nucleic acid encoding the transposase according to any one of claims 1-6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 15-16, the nucleic acid group according to any one of claims 11-15, the nucleic acid group construct according to any one of claims 16-24, the composition according to any one of claims 25-37, or the recombinant vector according to any one of claims 38-41 is delivered into a host cell.
48. A method for editing the genome of a host cell, wherein the method comprises: The transposase according to any one of claims 1-6, the nucleic acid encoding the transposase according to any one of claims 1-6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8-10, the nucleic acid group according to any one of claims 11-15, the nucleic acid group construct according to any one of claims 16-24, the composition according to any one of claims 25-37, or the recombinant vector according to any one of claims 38-41 is delivered into a host cell.
49. A method for obtaining a host cell comprising an exogenous nucleic acid fragment in its genome, wherein the method comprises: The transposase according to any one of claims 1-6, the nucleic acid encoding the transposase according to any one of claims 1-6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8-10, the nucleic acid group according to any one of claims 11-15, the nucleic acid group construct according to any one of claims 16-24, the composition according to any one of claims 25-37, or the recombinant vector according to any one of claims 38-41 is delivered into a host cell.
50. The method of any one of claims 47-49, wherein the delivery method comprises cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated viral delivery, electroporation, Agrobacterium infection, or a gene gun.
51. The method of any one of claims 47-49, wherein the host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell.
52. The method of claim 51, wherein the host cell comprises a mammalian cell.
53. The method of claim 52, wherein the host cell comprises a primary cell, an immortalized cell line, a cancer cell line, an embryonic stem cell line and differentiated cells thereof, or an induced pluripotent stem cell line and differentiated cells thereof.
54. The method of claim 53, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.
55. Use of the transposase according to any one of claims 1-6, the nucleic acid encoding the transposase according to any one of claims 1-6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8-10, the nucleic acid group according to any one of claims 11-15, the nucleic acid group construct according to any one of claims 16-24, the composition according to any one of claims 25-37, the recombinant vector according to any one of claims 38-41, or the recombinant host cell according to any one of claims 42-46 for introducing an exogenous nucleic acid fragment into the host cell genome.
56. The method of claim 55, wherein the host cell comprises an animal cell, a plant cell, an algae cell, a fungal cell, a yeast cell, or a bacterial cell.
57. The use of claim 56, wherein the host cell comprises a mammalian cell.
58. The method of claim 57, wherein the host cell comprises a primary cell, an immortalized cell line, a cancer cell line, an embryonic stem cell line and cells differentiated therefrom, or an induced pluripotent stem cell line and cells differentiated therefrom.
59. The use according to claim 58, wherein: The primary cells include mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell lines include HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The cancer cell lines include Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca; The embryonic stem cell lines include H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.
60. Use of the transposase according to any one of claims 1-6, the nucleic acid encoding the transposase according to any one of claims 1-6, the nucleic acid according to claim 7, the nucleic acid construct according to any one of claims 8-10, the nucleic acid group according to any one of claims 11-15, the nucleic acid group construct according to any one of claims 16-24, the composition according to any one of claims 25-37, the recombinant vector according to any one of claims 38-41, or the recombinant host cell according to any one of claims 42-46 in the preparation of a drug or preparation for gene therapy, cell therapy, genomic research, or stem cell induction and post-induced differentiation.
61. A kit, wherein the kit comprises the transposase of any one of claims 1-6, a nucleic acid encoding the transposase of any one of claims 1-6, a nucleic acid of claim 7, a nucleic acid construct of any one of claims 8-10, a nucleic acid set of any one of claims 11-15, a nucleic acid set construct of any one of claims 16-24, a composition of any one of claims 25-37, a recombinant vector of any one of claims 38-41, or a recombinant host cell of any one of claims 42-46.
Citation Information
Patent Citations
Transposase polypeptides and uses thereof
CN107532174A
Replicative transposon system
CN109072258A
Method for detecting T cell activation degree of antibody drug taking CD3 as target
CN112458149A
Cell penetrating transposase
CN113661247A
Swivel-based therapy
CN115698268A