Activated DNA transposon system and method of use thereof
Engineered transposable elements with specific terminal repeat sequences and transposases overcome limitations of existing DNA delivery methods by enabling efficient insertion into cellular genomes, enhancing gene therapy and research applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- INST OF ZOOLOGY CHINESE ACAD OF SCI
- Filing Date
- 2024-07-05
- Publication Date
- 2026-04-28
AI Technical Summary
Current methods for introducing DNA into cells, such as calcium phosphate, liposomes, and virus-mediated strategies, face limitations including size constraints, efficiency, and immunological issues, making them unsuitable for diverse applications in gene expression and therapy.
Development of engineered transposable elements with specific 5'- and 3'-terminal repeat sequences, along with a transposase, enabling efficient insertion of heterologous nucleic acids into cellular DNA, suitable for various cell types, including mammalian cells.
The engineered transposable elements provide enhanced transposability, allowing for efficient insertion of heterologous nucleic acids into cellular genomes, facilitating applications like gene therapy and gene discovery research.
Smart Images

Figure 0007852942000036 
Figure 0007852942000037 
Figure 0007852942000038
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to international patent application number PCT / CN2020 / 082087, filed on 30 March 2020, which is incorporated herein by reference in its entirety.
[0002] Reference to sequence listings compliant with WIPO standard ST.25. The contents of the following ASCII text file submitted in the original application are incorporated herein by reference: Computer-Readable Format (CRF) of the Sequence Listing (filename: 182452000641SEQLIST.txt, filed date: March 29, 2021, size: 244kb).
[0003] Technical field This application relates to the field of genetics in general. More specifically, it relates to transposable elements and their use. [Background technology]
[0004] Typical methods for introducing DNA into cells include DNA coagulants such as calcium phosphate and polyethylene glycol, lipid-containing reagents such as liposomes and multilayer vesicles, and virus-mediated strategies. However, such methods have limitations. For example, there are size limitations associated with DNA coagulants and virus-mediated strategies. Furthermore, the amount of nucleic acid that can be transfected into cells is limited in viral strategies. Not all methods are advantageous for inserting delivered nucleic acid into the cellular nucleic acid; while DNA coagulation methods and lipid-containing reagents are relatively easy to prepare, inserting nucleic acid into viral vectors is cumbersome. Virus-mediated strategies may be cell type or tissue type specific, and immunological problems arise when using virus-mediated strategies in vivo.
[0005] One suitable tool for overcoming these problems is through transposable elements. Transposable elements (TEs, transposons, or jumping genes) are DNA sequences that can alter the sequence in a genome by changing their position in nucleic acids, thereby generating or reversing mutations. Transposable elements make up a large portion of many eukaryotic genomes. For example, about 50% of the human genome is derived from transposable elements, while other genomes, such as those of plants, have a higher proportion of DNA derived from transposable elements. Transposable elements are usually classified into two classes: Class 1 and Class 2. Representatives of Class 1 are retrotransposons, which include 1) long-chain terminal repeat (LTR) retrotransposons such as endogenous retroviruses (ERVs), and 2) non-LTR retrotransposons such as long-chain scattered repeat (LINE) and short-chain scattered repeat (SINE) sequences. Class 2 transposons (TEs) include 1) “cut-and-paste” DNA transposons characterized by terminal repeat sequences (TRs, also called terminal reverse repeat sequences TIRs) that are transposable by transposases, and 2) non-“cut-and-paste” DNA transposons, such as helicron and Pollinton. Class 2 TEs are widely present and active in many eukaryotes, but not all of these TEs are transposable. Recent examples of active transposables include members of the hAT and piggyBa superfamilies that have shown signs of translocation over the past millions of years. However, currently, only a limited number of active transposables are available for gene expression research and gene therapy.
[0006] Therefore, we need a novel transposable element suitable for introducing DNA into cells, and via this transposable element, efficiently introducing heterologous sequences of different sizes into the cellular nucleic acid, or DNA into the cellular genome. A method and system for insertion are required. [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] This application provides engineered transposable elements, gene transfer systems containing said engineered transposable elements, and methods and kits using these. It also provides methods for inserting heterologous nucleic acids into target nucleic acids in vitro or intracellularly. The compositions, systems, and methods described herein are useful for a variety of applications, including tagging, genome engineering, and gene expression studies. [Means for solving the problem]
[0008] In one embodiment, the present invention provides a modified transposable element comprising a 5'-terminal repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3'-terminal repeat sequence (3'TR) in the order of 5' to 3', wherein the 5'TR comprises a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90, a variant thereof, or a fragment thereof, and the 3'TR comprises a nucleic acid sequence selected from SEQ ID NO: 27-52 and 91-102, a variant thereof, or a fragment thereof, and the transposable element provides a modified transposable element exhibiting transposable activity that enables insertion of heterogeneous nucleic acid into cellular DNA. In some embodiments, the 5'TR comprises a nucleic acid sequence having at least about 90% sequence identity with the nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90. In some embodiments, the 5'TR comprises a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90. In some embodiments, the 3'TR includes a nucleic acid sequence having at least about 90% sequence identity with a nucleic acid sequence selected from SEQ ID NOs: 27-52 and 91-102. In some further embodiments, the 3'TR includes a nucleic acid sequence selected from SEQ ID NOs: 27-52 and 91-102.
[0009] In any embodiment relating to the above-described manipulated transposable element, the transposable element further comprises a 5' target site repeat sequence (TSD) ligated to the 5' side of the 5' TR or a 3' TSD ligated to the 3' side of the 3' TR. In some embodiments, the nucleic acid sequences of the 5' TSD and the 3' TSD are identical. In some embodiments, the 5' TSD comprises a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NO: 191-206, and the 3' TSD comprises a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NO: 191-206.
[0010] In any embodiment relating to the above-described manipulated transposable element, the 5'TR includes a nucleic acid sequence having at least about 90% sequence identity with a nucleic acid sequence selected from SEQ ID NO: 3, 8, 11, 12, 13, 16, 22, 23, 29, and 82, and the 3'TR includes a nucleic acid sequence having at least about 90% sequence identity with a nucleic acid sequence selected from SEQ ID NO: 29, 34, 37, 38, 39, 42, 48, 49, 91, and 94.
[0011] In any embodiment relating to the above-described manipulated transposable element, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, Tc1-1_AG, Tc1-1_PM, Tc1-4_Xt, Tc1-15_Xt, Mariner-6_AMi, or Mariner-3_Crp. In some embodiments, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM. In some embodiments, the 5'TR contains the nucleic acid sequence of SEQ ID NO: 3, and the 3'TR contains the nucleic acid sequence of SEQ ID NO: 29. In some embodiments, the 5'TR contains the nucleic acid sequence of SEQ ID NO: 8, and the 3'TR contains the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the 5'TR contains the nucleic acid sequence of SEQ ID NO: 11, and the 3'TR contains the nucleic acid sequence of SEQ ID NO: 37. In some embodiments, the 5'TR contains the nucleic acid sequence of SEQ ID NO: 12, and the 3 'TR contains the nucleic acid sequence of SEQ ID NO: 38. In some embodiments, the 5'TR contains the nucleic acid sequence of SEQ ID NO: 16, and the 3'TR contains SEQ ID NO: Contains 42 nucleic acid sequences.
[0012] In any embodiment relating to the above-described manipulated transposable element, the heterogeneous nucleic acid includes a coding sequence. In some further embodiments, the heterogeneous nucleic acid further includes a promoter operably linked to the coding sequence.
[0013] In any embodiment relating to the above-described manipulated transposable element, the transposable activity of the transposable element is higher than that of the piggyBac(PB) transposon, Sleeping Beauty(SB) transposon, and / or TcBuster transposon.
[0014] In any embodiment relating to the above-described manipulated transposable element, the cells are animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the cells are mammalian cells. In some embodiments, the mammalian cells are selected from immune cells (e.g., T cells), liver cells, tumor cells, stem cells, fertilized eggs, muscle cells, and skin cells. In some embodiments, the cells are human cells. In some embodiments, the transposable element exhibits higher transposable activity in human embryonic kidney 293T (293T) cells than in HeLa cells.
[0015] In any embodiment relating to the manipulated transposable element described above, the transposable element is present in a vector. In some further embodiments, the vector is a plasmid vector or a viral vector.
[0016] Another aspect of the present invention provides a gene transfer system comprising 1) an engineered transposer relating to any of the above transposers, and 2) a transposase, or a nucleic acid encoding a transposase. In some embodiments, the transposase comprises an amino acid sequence or a variant thereof selected from SEQ ID NO: 53-78 and 103-114.
[0017] In a further embodiment, the present invention provides a gene transfer system comprising 1) an engineered transposer and 2) a transposase or a nucleic acid encoding a transposase, wherein the transposer comprises a 5'-terminal repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3'-terminal repeat sequence (3'TR) in the order of 5' to 3', the transposer exhibits transposability activity that enables insertion of the heterogeneous nucleic acid into cellular DNA, and the transposase comprises an amino acid sequence or variant thereof selected from SEQ ID NO: 53-78 and 103-114. In some embodiments, the 5'TR comprises a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 1-26 and 79-90, and the 3'TR comprises a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 27-52 and 91-102.
[0018] In some embodiments relating to any of the gene transfer systems described above, the transposable element includes a 5'TR having a nucleic acid sequence having at least about 90% sequence identity with a nucleic acid sequence selected from SEQ ID NO: 3, 8, 11, 12, 13, 16, 22, 23, 29, and 82, and the 3'TR includes a nucleic acid sequence having at least about 90% sequence identity with a nucleic acid sequence selected from SEQ ID NO: 29, 34, 37, 38, 39, 42, 48, 49, 91, and 94. In some embodiments, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, Tc1-1_AG, Tc1-1_PM, Tc1-4_Xt, Tc1-15_Xt, Mariner-6_AMi, or Mariner-3_Crp. In some embodiments, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM. In some embodiments, The 5'TR contains the nucleic acid sequence with SEQ ID NO: 3, the 3'TR contains the nucleic acid sequence with SEQ ID NO: 29, and the transposase contains the amino acid sequence with SEQ ID NO: 55. In some embodiments, the 5'TR contains the nucleic acid sequence with SEQ ID NO: 8, the 3'TR contains the nucleic acid sequence with SEQ ID NO: 34, and the transposase contains the amino acid sequence with SEQ ID NO: 60. In some embodiments, the 5'TR contains the nucleic acid sequence with SEQ ID NO: 11, the 3'TR contains the nucleic acid sequence with SEQ ID NO: 37, and the transposase contains the amino acid sequence with SEQ ID NO: 63. In some embodiments, the 5'TR contains the nucleic acid sequence with SEQ ID NO: 12, the 3'TR contains the nucleic acid sequence with SEQ ID NO: 38, and the transposase contains the amino acid sequence with SEQ ID NO: 64. In some embodiments, the 5'TR contains the nucleic acid sequence of SEQ ID NO: 16, the 3'TR contains the nucleic acid sequence of SEQ ID NO: 42, and the transposase contains the amino acid sequence of SEQ ID NO: 68.
[0019] In some embodiments according to any of the above gene transfer systems, the gene transfer system includes a nucleic acid encoding a transposase. In some further embodiments, the nucleic acid encoding the transposase and the transposable element are in different vectors. In some further specific embodiments, the nucleic acid encoding the transposase and the transposable element are in the same vector.
[0020] In a further aspect, the present application provides a method for inserting a heterologous nucleic acid into a target nucleic acid, comprising contacting the target nucleic acid with a transposable element according to any of the above engineered transposable elements or a gene transfer system according to any of the above gene transfer systems. In some embodiments, the method is performed in vitro. In some embodiments, the target nucleic acid is inside a cell. In some embodiments, the target nucleic acid is genomic DNA.
[0021] In some embodiments according to any of the above methods, the target nucleic acid is inside a cell, and the cell is an animal cell, a plant cell, an algal cell, a fungal cell, a yeast cell or a bacterial cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the mammalian cell is selected from immune cells (e.g., T cells), liver cells, tumor cells, stem cells, fertilized eggs, muscle cells and skin cells. In some embodiments, the insertion of the heterologous nucleic acid inactivates the gene of the cell.
[0022] In some embodiments according to any of the above methods, the heterologous nucleic acid encodes a protein. In some embodiments, the protein is selected from a reporter protein, an engineered receptor, a cytokine, an antibiotic resistance protein, an antigen and a therapeutic protein.
[0023] In some embodiments according to any of the above methods, the heterologous nucleic acid encodes RNA. In some embodiments, the RNA is selected from therapeutic RNA, small interfering RNA (siRNA), microRNA, short hairpin RNA (shRNA), long non-coding RNA (lincRNA), and guide RNA (gRNA). In some embodiments, the heterologous nucleic acid encodes one or more molecules.
[0024] In some embodiments according to any of the above methods, the length of the heterologous nucleic acid is about 300,0 00 bases (kb) or less, for example, about 10 kb to about 300 kb, or about 100 base pairs (bp) to about 10 kb, or about 100 bp to about 5 kb, or about 100 bp to about 2 kb, or about 2 kb to about 300 kb.
[0025] In some examples according to any of the above methods, the insertion is random.
[0026] In a further aspect, the present application provides a kit comprising a transposon according to any of the above engineered transposons, or a gene delivery system according to any of the above gene delivery systems, and a protocol for inserting a heterologous nucleic acid into a target nucleic acid.
Brief Description of the Drawings
[0027] [Figure 1] A diagram explaining the identification of active transposons (TE) is shown. [Figure 2] A summary of the 131 identified transposons (TE) is shown, and these transposons (TE) have an open reading frame (ORF) of 300 amino acids (aa) or more in length, a transposase (Tn) domain, and an average difference of 25% or less. The pie chart explains the distribution of the five superfamilies of the 131 identified TEs. [Figure 3]Exemplary binary constructs for screening active transposases are shown. The upper construct is an auxiliary construct expressing a transposase and contains, in 5' to 3' order, a cytomegalovirus (CMV) promoter, a transposase (Tn) gene, and a polyA(pA) signal. The lower construct is a donor construct and contains, in 5' to 3' order, a phosphoglycerate kinase (PGK) promoter, a 5' target site repeat (5'TSD) sequence, a 5' terminal repeat (5'TR) sequence, a sequence encoding puromycin-resistance-enhancing GFP fusion protein (Puro-eGFP), a 3' terminal repeat (3'TR) sequence, a 3' target site repeat (3'TSD) sequence, and a polyA(pA) signal. [Figure 4]Figures 4A-4B show the transposability activity of identified transposable elements with a difference of 5% or less compared to control TEs including piggyBac, Hyper piggyBac, and SB100X in HEK293T (293T) and HeLa cells. Figure 4A shows the transposability activity of the following identified transposable elements in HEK293T (293T) and HeLa cells compared to control TEs (including piggyBac, Hyper piggyBac, SB100X, and TcBuster (TB)) and negative controls (NCs): 1_hAT-2_AG, 2_AgaP12, 3_P3_AG, 4_HAT2_CI, 5_HAT5_CI, 6_hAT-6_DR, 7_Chaplin1_DR, 8_Harbinger-4_XT, 9_HAT1_AG, 10_AgaP15, 11_IS4EU-1_DR, 12_POGO, 14_Tc1-8B_DR, 15_HOBO, 17_hAT -6_PM, 18_MARISP1, 19_hAT-7_XT, 20_BARI_DM, 21_Mariner-3_PM, 22_hAT-3_XT, 23_P1_AG, 24_IS4EU-2_DR, 25_Tc1-3_Xt , 26_Mariner-1_PM, 27_P2_AG, 28_hAT-5_DR, 29_Tc1-3_FR, 30_Mariner-4_XT, 32_Tc1-5_Xt, 33_Tc1-10_Xt and 34_hAT-1_D.Figure 4B shows the transposability activity of the following identified transposable elements in HEK293T (293T) and HeLa cells compared to control TEs (including piggyBac, Hyper piggyBac, SB100X, and TcBuster (TB)) and negative controls (NC): 35_Mariner2_AG, 36_Tc1-1_Xt, 37_Tc1-1_AG, 38_Mariner-1_XT, 39_Tc1DR3_Xt, 40_hAT-1B_PM, 41_hAT-3_PM, 42_hAT-6B_PM, 43_Tc1-1_PM, 44_TC1_XL, 45_Mariner-4_AMi, 46_Myotis_hAT1 , 47_hAT-7_PM, 48_TC1_FR4, 49_hAT-8_PM, 50_piggyBac1_Mm, 51_Tc1-11_Xt, 52_Tc1~16_Xt, 53_TC1-2_DM, 54_Tc1-4_Xt, 55_TC1_DM, 56_Tc1- 15_Xt, 59_TC1_FR2, 60_Mariner-6B_AMi, 61_Tc1-9_Xt, 63_hAT-9_XT, 64_Mariner-5_XT, 65_S2_DM, 67_PROTOP, 68_Tc1-8_Xt and 69_Tc1-12_Xt. [Figure 5]Figures 5A-5B show the transposability activity of identified transposable elements that show a difference greater than 5% in HEK293T(293T). Figure 5A shows the transposability activity of the following identified transposable elements in HEK293T(293T): 70_Mariner-6_AMi, 71_hAT-3_Gav, 72_piggyBac2_Mm, 73_Mariner-2_XT, 74_hAT-4_Crp, 75_OposCharlie2, 76_hAT-1_AMi, 77_Mariner-2_AMi, 78_hAT-13_AMi, 80_hAT-12_AMi, 81_Mariner-7_Croc, 82_piggyBac-2_XT, 83_Mariner-3_AMi, 84_piggyBac1_CI, 86_piggyBac-1_AMi, 87_ Mariner-5_AMi, 89_Mariner-6_Crp, 90_Mariner-3_Crp, 91_hAT-3_AMi, 92_TC1_FR1, 93_hAT-5_Croc, 94_hAT-8_AMi, 95_PARIS, 96_Mariner-2_PM, 97_piggyBac-1_XT , 98_Tigger1, 99_Mariner-1_Crp, 101_S_DM, 102_hAT-12_Crp, 103_hAT-1_PM, 104_Senkusha1, 105_Tc1-2_PM, 106_Harbinger-2_AMi, 108_hAT-2_Gav and 110_Tigger2T.Figure 5B shows the transposability activity of the following identified transposable elements in HEK293T (293T): 111_TC1-4_DR, 116_TIGGER17, 117_TIGGE3, 118_HAT-14_CRP, 119_MARINER-9_CRP, 124_HAT-111_AMI, 126_MARINER-1_AMI, 126_MARINER-10_CRP, 139_hAT-17_Croc, 140_hAT-4_AMi, 142_Tigger4, 144_Tigger7, 156_Tigger2, 180_hAT-19_Crp, 183_hAT-17B_Croc, 188_Tc1-13_Xt, 19 7_hAT-19B_Croc, 199_MarsTigger8, 212_Tigger5, 222_hAT-10_XT, 237_hAT-6_AMi, 245_MarsTigger1c, 246_Tigger17c, 258_Harbinger-1_AMi, 260_Tc1-14_Xt, 271_ Kanga1, 275_Arthur1, 295_Harbinger-1_Crp, 314_Zaphod3, 320_Joey1, 331_Zaphod, 342_Harbinger-3_AMi, 348_Harbinger-1B_Crp, 349_MARWOLEN1 and 351_Zaphod2. [Figure 6] Figures 6A and 6B show the translocation (i.e., introduction) efficiency of various identified TEs compared with control TEs piggyBac, hyperpiggyBac, and SB100X. Figure 6A shows the translocation (introduction) efficiency of various identified TEs compared with control TEs piggyBac, hyperpiggyBac, and SB100X in HEK293T (293T) cells. Figure 6B shows the translocation (introduction) efficiency of various identified TEs compared with control TEs piggyBac, hyperpiggyBac, and SB100X in HeLa cells. [Figure 7]Figures 7A-7C show phylogenetic trees illustrating the phylogenetic relationships between identified transposases belonging to different superfamilies and their control TEs. Figure 7A shows a phylogenetic tree illustrating the phylogenetic relationships between the following 14 identified hAT superfamily transposases and the control TE TcBuster: 103_hAT-1_PM, 17_hAT-6_PM, 46_Myotis_hAT1, 28_hAT-5_DR, 22_hAT-3_XT, 41_hAT-3_PM, 47_hAT-7_PM, 180_hAT-19_Crp, 1_hAT-2_AG, 9_HAT1_AG, 139_hAT-17_Croc, 183_hAT-17B_Croc, 63_hAT-9_XT, and 222_hAT-10_XT. Figure 7B shows a phylogenetic tree illustrating the phylogenetic relationships between two identified piggyBac superfamily transposases, 82_piggyBac-2_XT and 86_piggyBac-1_AMi, and the control TE TcBuster. Figure 7C shows a phylogenetic tree illustrating the phylogenetic relationships between 22 identified TcMariner superfamily transposases and the control TE SB100X. [Figure 8] Figures 8A and 8B show the transpossession activity of the identified transpossession factors and the five most active TEs out of 131 candidates, compared to a control TE including piggyBac, hyperpiggyBac, and SB100X, and a negative control (NC), in HEK293T (293T) cells and HeLa cells. Figure 8A shows the transpossession (introduction) efficiency of various identified TEs in 293T cells, compared to a control TE including piggyBac, hyperpiggyBac, and SB100X, and a negative control (NC). Figure 8B shows the transpossession (introduction) efficiency of various identified TEs in HeLa cells, compared to a control TE including piggyBac, hyperpiggyBac, and SB100X, and a negative control (NC). [Figure 9]Figures 9A and 9B show the transpossession activity of the identified transposable elements and the top 5 most active TEs out of 131 candidates, compared to a control TE including piggyBac, hyperpiggyBac, and SB100X, and a negative control (NC), in Hct116 and K562 cells. Figure 9A shows the transpossession (introduction) efficiency of various identified TEs in Hct116 cells, compared to a control TE including piggyBac, hyperpiggyBac, and SB100X, and a negative control (NC). Figure 9B shows the transpossession (introduction) efficiency of various identified TEs in K562 cells, compared to a control TE including piggyBac, hyperpiggyBac, and SB100X, and a negative control (NC). [Figure 10] Figures 10A-10B show the transposability activity of identified transposable elements 14-Tc1-8B_DR, 29-Tc1-3_FR, 35-Mariner2_AG, 36-Tc1-1_Xt, 37-Tc1-1_AG, 43-Tc1-1_PM, 52-Tc1-16_Xt, 54-Tc1-4_Xt, and 56-Tc1-15_Xt in primary T cells compared to the control TE SB100X. Figure 10A shows the results of transposability measurements of primary T cells transfected with a plasmid containing the EF1α promoter and the CopGFP gene, where the presence of GFP-positive T cells indicates successful intracellular transposability. Figure 10B shows the results of transposability measurements of primary T cells transfected with a plasmid containing the EF1α promoter and the 019CAR-P2A-eGFP gene, where the presence of GFP-positive T cells indicates successful intracellular transposability. [Figure 11] Figures 11A and 11B show the proportions based on helper plasmids and transposable element plasmids in 293T cells and HeLa cells, compared to the control TE piggyBac and negative control (NC), and the transposable activity of the identified transposable element SB100X and the five most active TEs out of 131 candidates. Figure 11A shows the transposable (introduced) efficiency of various identified TEs in 293T cells compared to the control TE SB100X and negative control (NC). Figure 11B shows the transposable (introduced) efficiency of various identified TEs in HeLa cells compared to the control TE SB100X and negative control (NC). [Figure 12]Figure 12 shows the loading capacity of the identified transposable element SB100X and the five most active TEs out of 131 candidates in 293T cells, compared to the control TEs piggyBac, piggyBac, hyperpiggy Bac, and SB100X. Figure 12 also shows the transpossession (introduction) efficiency of each identified TE compared to the control TEs piggyBac, piggyBac, hyperpiggyBac, and SB100X. [Figure 13] This shows the common sequences of genomic insertion loci within a 40 bp window around target sites in genomes derived from TE primitive species and stable transmutation K562 cell lines. [Figure 14] Figures 14A and 14B show the insertion frequencies of identified transposable elements in the closest genes, including piggyBac, hyperpiggyBac, SB100X, and the five most active TEs out of 131 candidates. Figure 14A shows the insertion frequencies within the 0 bp to 1 Mb window, and Figure 14B shows the enrichment factor for random insertions within the 50 kb window. [Figure 15] Figure 15 shows the enrichment factor for random insertions of identified transposable elements, including piggyBac, hyperpiggy Bac, SB100X, and the five most active TEs out of 131 candidates, at the relative position from the 5' end to the 3' end of the gene body. [Figure 16] Figure 16 shows the enrichment factor for random insertions of identified transposable elements, including piggyBac, hyperpiggyBac, SB100X, and the five most active TEs out of 131 candidates, for genes at various expression levels, with level 7 representing maximum expression and level 0 representing minimum expression. [Figure 17] Figure 17 shows the enrichment factor for random insertions of identified transposable elements, including piggyBac, hyperpiggy Bac, SB100X, and the five most active TEs out of 131 candidates, around the transcription start site (TSS) within a 5kb window. [Figure 18]Figure 18 shows the enrichment factor for random insertion of identified transposable elements, including piggyBac, hyperpiggyBac, SB100X, and the five most active TEs out of 131 candidates, under various chromatin states. [Modes for carrying out the invention]
[0028] This application provides engineered transposable elements, gene transfer systems containing engineered transposable elements, and methods and kits using them. This application is at least in part based on the identification of novel transposable elements from a wide range of species analyzed by pan-genomic bioinformatics (e.g., Table 2), and the surprising result that many of the identified transposable elements can be efficiently transferred in human cells. The disclosed compositions, systems, and methods are useful for inserting heterologous nucleic acids into target nucleic acids, including introducing heterologous DNA into the genome of a cell. The transposable elements and gene transfer systems described herein are useful for a variety of applications, including gene therapy and gene discovery research.
[0029] Therefore, in one embodiment, the present invention comprises a 5'-terminal repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3'-terminal repeat sequence (3'TR) in the order from 5' to 3', wherein the 5'TR comprises a nucleic acid sequence selected from SEQ ID NO: 1~26 and 79~90, its variant, or its fragment, and the 3'TR comprises a nucleic acid sequence selected from SEQ ID NO: 27~52 and 91~102, its variant, or its fragment, and the above transposable element is a heterogeneous nucleic acid cell D This invention provides a modified transposable factor that exhibits transposable activity, enabling insertion into NA.
[0030] In another embodiment, the present invention provides a gene transfer system comprising 1) an engineered transposer and 2) a transposase or a nucleic acid encoding a transposase, wherein the transposer comprises a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3' repeat sequence (3'TR), and the transposer exhibits transposability activity that enables insertion of the heterogeneous nucleic acid into cellular DNA. In some embodiments, the transposer is an engineered transposer. In some embodiments, the transposase comprises an amino acid sequence or a variant thereof selected from SEQ ID NO: 53-78 and 103-114.
[0031] I. Definition As used herein, the terms “transposon,” “transposable element,” or “TE” mean a polynucleotide that can be cleaved from a primary nucleic acid (i.e., a donor nucleic acid such as a vector) and incorporated into a target site (e.g., a secondary nucleic acid in a cell, genome, or extrachromosomal DNA). A transposon includes a nucleic acid sequence ligated to the cis-active nucleic acid sequence side of the transposable element's terminal, such that at least one cis-active nucleic acid sequence is located at 5' of the nucleic acid sequence and at least one cis-active nucleic acid sequence is located at 3' of the nucleic acid sequence. The cis-active nucleic acid sequence includes at least one terminal repeat sequence (TR, also called an inverse terminal repeat sequence (ITR) or terminal reverse repeat sequence (TIR)) at each terminal of the transposable element to which the transposase binds. The transposable elements described herein may or may not include an open reading frame (ORF) encoding the transposase.
[0032] As used herein, the term “transfer” means that a transposable element is altered from its position in the primary nucleic acid (e.g., the vector) and incorporated into a target site (e.g., a secondary nucleic acid within a cell, or genome or extrachromosomal DNA).
[0033] As used herein, the term “terminal repeat” or “TR” refers to nucleic acid sequences located at both ends of a transposable element and ligated to the second nucleic acid sequence side of the transposable element. A TR located at the 5' (upstream) end of the second nucleic acid sequence is called a 5'TR, and a TR located at the 3' (downstream) end of the second nucleic acid sequence is called a 3'TR. In class 2 transposable elements, TRs are complementary to each other.
[0034] As used herein, the terms “target site repeat” or “TSD” mean a nucleic acid sequence appearing at the insertion site of a transposator. TSDs may be expressed by viscous end DNA repair resulting from alternating cleavage of the target DNA double strand by a transposase. TSDs are ligated to the TR side within the transposator. A TSD located at 5' of a 5'TR is a 5'TSD. A TSD located at 3' of a 3'TR is a 3'TSD.
[0035] As used herein, the term “transposase” means a polypeptide that catalyzes the cleavage of a transposon from a primary nucleic acid (e.g., a vector) and its integration into a target site (e.g., a secondary nucleic acid within a cell, or genomic or extrachromosomal DNA). In some embodiments, the transposase binds to one or two terminal repeat sequences.
[0036] As used herein, “left transposon fragment” or “LTF” refers to the fragment in a naturally occurring transposable element from the 5'TSD to the start codon of the transposase ORF sequence. As used herein, “right transposon fragment” or “RTF” refers to the fragment in a naturally occurring transposable element from the stop codon to the 3'TSD of the transposase ORF sequence.
[0037] The terms “nucleic acid,” “polynucleotide,” and “nucleic acid sequence” include deoxyribonucleotides, ribonucleotides, combinations thereof, and analogues thereof, of any length. This refers to a polymerized form of nucleotides, and they can be used interchangeably. "Oligonilocyte" and "oligo" can be used interchangeably and refer to short polynucleotides having approximately 50 or fewer nucleotides.
[0038] As used herein, "heteronucleotide" means a DNA or RNA sequence of a different origin than the reference nucleic acid sequence. For example, in the case of transposable elements, the heteronucleotide originates from a different source than the terminal repeat sequence. For example, a nucleic acid sequence isolated from a different organism than the terminal repeat sequence is considered a heteronucleotide relative to the terminal repeat sequence.
[0039] As used herein, the term “operably ligated” means a nucleic acid sequence that has a functional relationship with another nucleic acid sequence. For example, if a coding sequence is operably ligated to a promoter sequence, it usually means that the promoter can facilitate the transcription of the coding sequence. Being operably ligated means that the DNA sequences to be ligated are usually contiguous, and if two protein coding regions must be ligated, they are contiguous and within the reading frame. Because enhancers function thousands of base pairs away from promoters and intron sequences are of variable length, some nucleic acid sequences may be operably ligated but discontinuous.
[0040] For nucleic acid sequences, the "percentage of sequence identity (%)" is defined as the percentage of nucleotides in a candidate sequence that are identical to the nucleotides in a specific nucleic acid sequence after sequence alignment, and a cap may be applied as needed to maximize the percentage of sequence identity. For peptide, polypeptide, or protein sequences, the "percentage of sequence homology (%)" is the percentage of amino acid residues in a candidate sequence that have the same substitutions as amino acid residues in a specific peptide or amino acid sequence after sequence alignment, and a cap may be applied as needed to maximize the percentage of sequence homology. Alignment for determining the percentage of amino acid sequence identity can be, for example, BLAST, BLAST-2, ALIGN, or MEGALIGN. TMThis can be carried out in various ways within the scope of the art, such as using commonly available computer software like (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measurement alignment, including any algorithm necessary to achieve maximum alignment over the entire length of the sequences being compared.
[0041] As used herein, the term “vector” means a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is ligated. Examples of vectors include, but are not limited to, bacteria, plasmids, phages, cosmids, episomes, viruses, and insertable DNA fragments, i.e., fragments that can be inserted into the host cell genome by homologous recombination.
[0042] As used herein, the term "plasmid" means a circular double-stranded DNA capable of receiving foreign DNA fragments and replicating in prokaryotic or eukaryotic cells.
[0043] The terms “polypeptide” and “peptide” are interchangeable and refer to amino acid polymers of any length. Therefore, for example, the terms peptide, oligopeptide, protein, antibody, and enzyme are included in the definition of polypeptide. Polymers may be linear or branched, may contain modified amino acids, or may be cleaved by non-amino acids. Proteins may have one or more polypeptides. This term also includes modified amino acid polymers. Any other operations are possible, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or binding with labeling components.
[0044] As used herein, “variant” is interpreted as a polynucleotide or polypeptide that is different from a reference polynucleotide or polypeptide but retains its basic properties. Typical variants of a polynucleotide differ from the nucleic acid sequence of other reference polynucleotides. Nucleic acid sequence changes in a variant may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. As described below, nucleotide changes can result in substitution, addition, deletion, fusion, and cleavage of amino acids in the polypeptide encoded by the reference sequence. Typical variants of a polypeptide differ from the amino acid sequence of other reference polypeptides. Usually, the differences are limited, and therefore the sequences of the reference polypeptide and the variant are very similar overall and identical in many regions. The amino acid sequences of the variant and the reference polypeptide may differ in the form of one or more substitutions, additions, or deletions in any combination. The substituted or inserted amino acid residues may or may not be amino acid residues encoded by the genetic code. Variants of polynucleotides or polypeptides may or may not occur naturally, such as allele variants. Non-natural variants of polynucleotides and polypeptides can be prepared by mutagenesis techniques, direct synthesis, and other recombinant methods known to those skilled in the art.
[0045] As used herein, a “fragment” of a sequence means a portion of a sequence. For example, a fragment of a nucleic acid sequence means a portion of a nucleic acid sequence, and a fragment of an amino acid sequence means a portion of an amino acid sequence.
[0046] As used herein, the terms “genetic circuit,” “biological circuit,” or “synthetic circuit” refer to a set of biological components designed to perform a logical function. Generally, a genetic circuit requires an input to be activated, and the genetic circuit produces an output in response to the input.
[0047] As used herein, the term “manipulated” means any operation that causes a detectable change in a polynucleotide or polypeptide, including, but not limited to, the insertion, deletion, and substitution of any part of a polynucleotide or amino acid sequence.
[0048] As used herein, the term “transfer efficiency” means the efficiency with which a transposable element inserts a heterologous nucleic acid into a target cell population. For example, transfer efficiency may be determined by transfecting a target cell population with a plasmid containing a transposable element that includes a gene encoding a reporter gene or a select marker, such as an antibiotic resistance gene (e.g., puromycin), and measuring the number of cells that express the gene product encoded by the reporter gene or select marker, for example, the number of antibiotic-resistant cells.
[0049] As used herein, the terms “transfected,” “transformed,” or “transduced” mean the process of transferring or introducing a foreign nucleic acid into a host cell. A “transfected,” “transformed,” or “transduced” cell is a cell that has been transfected, transformed, or transduced with a foreign nucleic acid. As used herein, the terms “transduction” and “transfection” include all methods known in the art for introducing DNA into a cell to express a desired protein or molecule using an infectious agent (e.g., a virus) or other means. Chemical transfection methods such as those using calcium phosphate, dendrimers, liposomes, or cationic polymers (e.g., DEAE-glucan or polyethyleneimine) in addition to viruses or virus-like preparations; non-chemical methods such as electroporation, cell squeeze, acoustic perforation, phototransfection, puncture transfection, protoplast fusion, plasmid or transposon delivery; gene gun, magnetic transfection or Particle-based methods such as magnetic-assisted transfection and the use of particle shocks; mixed methods such as nuclear transfection.
[0050] The term "in vivo" refers to the process of obtaining cells within the body of the organism from which they are acquired. "Ex vivo" or "in vitro" refers to the process of obtaining cells outside the body of the organism from which they are acquired.
[0051] Furthermore, the embodiments of the present invention described herein include embodiments consisting of and / or substantially consisting of.
[0052] In this specification, when a number or parameter is mentioned with "approximately," it includes (describes) a change in the number or parameter itself. For example, a description that mentions "approximately X" includes a description of "X."
[0053] Where used herein, the term "otherwise" refers to a number or parameter that is "different" from the number or parameter in general. For example, the statement that the method is not used to treat type X cancer means that it is used to treat cancers other than type X.
[0054] As used in this specification, the term "approximately X to Y" has the same meaning as "approximately X to approximately Y".
[0055] As used herein and in the appended claims, the singular forms "one kind," "one," and "above" include the plural form unless otherwise specified in the context.
[0056] II. Manipulated Transposable Factors One aspect of the present invention provides an engineered transposer comprising a 5'-terminal repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3'-terminal repeat sequence (3'TR) from 5' to 3'. In some embodiments, the transposer exhibits transposability activity that enables in vitro insertion of the heterogeneous nucleic acid into a target nucleic acid (e.g., DNA). In some embodiments, the engineered transposer exhibits transposability activity that enables intracellular insertion of the heterogeneous nucleic acid into a target nucleic acid (e.g., mammalian or plant DNA). In some embodiments, the transposer includes a nucleic acid sequence encoding a transposase. In some embodiments, the transposer does not include a nucleic acid sequence encoding a transposase.
[0057] In some embodiments, the manipulated transposer includes a 5'-to-3' target site repeat sequence (5'TSD), a 5'TR, a heterogeneous nucleic acid, a 3'TR, and a 3'TSD. In some embodiments, the transposer exhibits transposability activity that enables in vitro insertion of the heterogeneous nucleic acid into a target nucleic acid (e.g., DNA). In some embodiments, the manipulated transposer exhibits transposability activity that enables intracellular insertion of the heterogeneous nucleic acid into a target nucleic acid (e.g., mammalian or plant DNA). In some embodiments, the transposer includes a nucleic acid sequence encoding a transposase. In some embodiments, the transposer includes a nucleic acid sequence encoding a transposase.
[0058] In some embodiments, the present application provides novel transposable elements of a broad species origin that can be efficiently transposed within human cells. The transposable elements described herein provide direct experimental evidence for mammalian cleavage DNA transposons with innate activity.
[0059] A list of exemplary transposable elements and their corresponding left-terminal fragments (LTFs), right-terminal fragments (RTFs), transposases, 5'-terminal repeats (TRs), 3'TRs, 5'-target site repeats (TSDs), and 3'TSD sequences are listed in Tables 1-2 and the sequence listing. Table 1 shows 131 TEs identified by bioinformatics analysis disclosed herein, of which 11 TEs (TE IDs: 4, 5, 7, 12, 18, 20, 22) It includes (25, 26, 28, and 30), has a length of 3000 bp or less, a MITE copy number greater than 10, and a mean difference less than 1%, making it applicable to high-efficiency genome engineering. Table 2 shows the active TEs verified using transfer analysis experiments in the human cell line HEK293T and HeLa.
[0060] In some embodiments, the transposase is a class 2 transposase. Class 2 transposases may be classified as a superfamily based on the correlation and shared structural features of the transposases, including the length of the terminal repeat (TR) and the target site repeat (TSD) ligated to the TR side during the integration process. It is assumed that the transposases of this application may originate from various appropriate TE superfamilies and / or families. In some embodiments, the transposase originates from the hAT superfamily. In some embodiments, the transposase originates from the P superfamily. In some embodiments, the transposase originates from the PIF-Harbinger superfamily. In some embodiments, the transposase originates from the piggyBac superfamily. In some embodiments, the transposase originates from the TcMariner superfamily. In some embodiments, the 5'TR is the inverse complement of the 3'TR. In some embodiments, the 5'TR is not the inverse complement of the 3'TR.
[0061] In some embodiments, the manipulated transposable element includes a 5'TR containing a nucleic acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% sequence identity with a 5'TR of a transposable element from a superfamily selected from hAT, P, PIF-Harbinger, piggyBac, and TcMariner. In some embodiments, the manipulated transposable element includes a 3'TR containing a nucleic acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% sequence identity with a 3'TR of a transposable element from a superfamily selected from hAT, P, PIF-Harbinger, piggyBac, and TcMariner.
[0062] In some embodiments, the manipulated transposable element comprises a 5'TR and a 3'TR, wherein the 5'TR comprises a nucleic acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% sequence identity with the 5'TR of a transposable element from a superfamily selected from hAT, P, PIF-Harbinger, piggyBac, and TcMariner, and the 3'TR comprises a nucleic acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 95%, 99%, or 100% sequence identity with the 3'TR of a transposable element from a superfamily selected from hAT, P, PIF-Harbinger, piggyBac, and TcMariner.
[0063] In some embodiments, the manipulated transposable element comprises a 5'TR, 3'TR, LTF, RTF, transposase, 5'TSD, and / or 3'TSD derived from any one TE in Table 1. In some embodiments, the manipulated transposable element comprises a 5'TR, 3'TR, LTF, RTF, transposase, 5'TSD, and / or 3'TSD derived from any one TE in Table 2. In some embodiments, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, Tc1-1_AG, Tc1-1_PM, Tc1-4_Xt, Tc1-15_Xt, Mariner-6_AMi, or Mariner-3_Crp. In some embodiments, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0064] In some embodiments, the manipulated transposable element comprises an LTF containing a nucleic acid sequence, a variant thereof, or a fragment thereof, selected from SEQ ID NO: 115-140 and 167-178.
[0065] In some embodiments, the manipulated transposable element comprises an RTF containing a nucleic acid sequence, a variant thereof, or a fragment thereof, selected from SEQ ID NO: 141-190.
[0066] In some embodiments, the manipulated transposable element comprises an LTF and an RTF, wherein the LTF has sequence identity of at least approximately 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more with a nucleic acid sequence selected from SEQ ID NO: 115-140 and 167-178, and the RTF is SEQ The nucleic acid sequence selected from ID NO: 141-190 has at least approximately 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity. In some embodiments, the manipulated transposable element comprises an LTF containing a nucleic acid sequence selected from SEQ ID NO: 115-140 and 167-178, and an RTF containing a nucleic acid sequence selected from SEQ ID NO: 141-190.
[0067] In some embodiments, the manipulated transposable element includes the 5' TR of the LTF, and the LTF includes a nucleic acid sequence selected from SEQ ID NOs: 115-140 and 167-178. In some embodiments, the manipulated transposable element includes the 3' TR of the RTF, and the RTF includes a nucleic acid sequence selected from SEQ ID NOs: 141-190.
[0068] In some embodiments, the manipulated transposable element comprises a 5'TR, a variant of a 5'TR, or a fragment of a 5'TR of an LTF, wherein the LTF comprises a nucleic acid sequence selected from SEQ ID NOs: 115-140 and 167-178, and a 3'TR, a variant of a 3'TR, or a fragment of a 3'TR of an RTF, wherein the RTF comprises a nucleic acid sequence selected from SEQ ID NOs: 141-190.
[0069] In some embodiments, the manipulated transposable element includes a 5' TR having a nucleic acid sequence having at least about 90% sequence identity with a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90. In some embodiments, the manipulated transposable element of the present invention includes a 5' TR having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90. In some embodiments, the 5' TR includes a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90. In some embodiments, the manipulated transposable element includes a 3' TR having a complementary sequence as the 5' TR.
[0070] In some embodiments, the manipulated transposable element includes a 3' TR having a nucleic acid sequence with at least about 90% sequence identity to a nucleic acid sequence selected from SEQ ID NO: 27-52 and 91-102. In some embodiments, the manipulated transposable element includes a 3' TR having at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to a nucleic acid sequence selected from SEQ ID NO: 27-52 and 91-102. In some embodiments, the 3' TR includes a nucleic acid sequence selected from SEQ ID NO: 27-52 and 91-102. In some embodiments, the manipulated transposable element includes a 5' TR having a complementary sequence as the 3' TR.
[0071] In some embodiments, the manipulated transposable element has 1) a 5'TR having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90, and 2) SEQ ID NO: 27-52 and It includes a nucleic acid sequence selected from 91 to 102 and a 3'TR having at least approximately 90% sequence identity (for example, at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more).
[0072] In some embodiments, the manipulated transposable element includes: 1) a 5' TR having at least about 90% sequence identity (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) with a nucleic acid sequence selected from SEQ ID NO: 1 to 26; and 2) a 3' TR having at least about 90% sequence identity (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) with a nucleic acid sequence selected from SEQ ID NO: 27 to 52.
[0073] In some embodiments, the manipulated transposable element includes: 1) a 5' TR having at least about 90% sequence identity (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) with a nucleic acid sequence selected from SEQ ID NO: 79-90; and 2) a 3' TR having at least about 90% sequence identity (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) with a nucleic acid sequence selected from SEQ ID NO: 91-102.
[0074] This specification also considers engineered transposable elements comprising any one variant or fragment of the 5'TR and / or 3'TR, or LTF and / or RTF, as described in Tables 1 and 2 of this specification. In some embodiments, the variants include substitutions of approximately 50, 40, 35, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or one or fewer nucleotides. In some embodiments, the fragments include at least approximately 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides.
[0075] In some embodiments, the manipulated transposable element includes a 5'TSD. In some embodiments, the manipulated transposable element includes a 3'TSD. In some embodiments, the manipulated transposable element includes both a 5'TSD and a 3'TSD. In some embodiments, the manipulated transposable element does not include a 5'TSD. In some embodiments, the manipulated transposable element does not include a 3'TSD. In some embodiments, the manipulated transposable element does not include either a 5'TSD or a 3'TSD. In some embodiments, the 5'TSD is the same as the 3'TSD. In some embodiments, the 5'TSD is different from the 3'TSD.
[0076] In some embodiments, the manipulated transposator includes a nucleic acid sequence encoding a transposase. In some embodiments, the manipulated transposator does not include a nucleic acid sequence encoding a transposase. In some embodiments, the transposase is derived from the same species as the 5'TR and 3'TR sequences. In some embodiments, the transposase is a natural transposase for the 5'TR and 3'TR sequences. In some embodiments, the transposase is a manipulated transposase based on a natural transposase for the 5'TR and 3'TR sequences.
[0077] A transposase catalyzes the cleavage of a transposon from a donor polynucleotide (e.g., a vector) and then the integration of the transposon into a target nucleic acid, such as the genome or extrachromosomal DNA of a target cell. In some embodiments, the transposase binds to the terminal repeats of the transposase factor.
[0078] In some embodiments, the transposase comprises an amino acid sequence, a variant thereof, or a fragment thereof, selected from SEQ ID NO: 53-78 and 103-114.
[0079] In some embodiments of this application, transposase variants described in Tables 1-2 and the sequence listing are considered. The term transposase variant refers to a transposase polypeptide that differs from a reference transposase polypeptide (e.g., a naturally occurring transposase polypeptide) in the addition, deletion, cleavage, and / or substitution of at least one amino acid residue, but retains transposase activity. In some embodiments, the transposase polypeptide variant differs from the reference transposase polypeptide in one or more substitutions, which may be conserved or non-conservative as known in the art. In some embodiments, the variant transposase contains an amino acid sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity or similarity with the sequence corresponding to the reference transposase. In some embodiments, the mutant transposases contain approximately 50, 40, 35, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or one or fewer amino acid substitutions.
[0080] Functional fragments having amino acid deletions or mutants having amino acid additions of the transposases described herein are also considered. In some embodiments, the length of the transposase fragment is at least about 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, or 700 or more amino acid residues. In some embodiments, the amino acid addition or deletion occurs at the C-terminus and / or N-terminus of the reference transposase. In some embodiments, the amino acid addition or deletion occurs at an internal location, for example, in the flexible loop of the reference transposase. In some embodiments, an amino acid deletion (e.g., N-terminal and / or C-terminal truncation) may contain approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175 or more amino acids, encompassing all and ranges between these values. In some embodiments, the mutant transposase includes an N-terminal or C-terminal purification tag, a selection marker (e.g., an antibiotic resistance gene), or a reporter gene (e.g., a fluorescent reporter gene).
[0081] As described above, the transposase polypeptide of the present invention may be modified by many means, such as amino acid substitution, deletion, cleavage, and insertion. Methods of such operations are generally known in the art. For example, amino acid sequence variants of the reference polypeptide may be produced by DNA mutation. Methods of mutagenesis and nucleic acid sequence modification are known in the art. For example, see Kunkel (1985, Proc. Natl. Acad. Sci. USA. 82: 488-492), Kunkel et al., (1987, Methods in Enzymol, 154: 367-382), US Pat. No. 4,873,192, Watson, JD et al., (Molecular Biology of the Gene, Fourth Edition, Benjamin / Cummings, Menlo Park, Calif., 1987) and the references cited therein. Guidance on appropriate amino acid substitutions that do not affect the biological activity of the target protein is described in the model in Dayhoff et al., (1978) Atlas of Protein Sequence and Structure (Natl. Biomed. Res. Found., Washington, DC).
[0082] In some embodiments, the transposase is codon-optimized compared to a reference transposase. In some embodiments, the transposase is codon-optimized for expression in mammalian cells, such as human cells. In some embodiments, the transposase is codon-optimized for expression in plant cells.
[0083] In some embodiments, the transposase contains an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with an amino acid sequence selected from SEQ ID NO: 53-78 and 103-114.
[0084] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, where the LTF contains the nucleic acid sequence of SEQ ID NO: 115; and 2) a 3'TR of the RTF, where the RTF contains the nucleic acid sequence of SEQ ID NO: 141. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence of TAGGC (SEQ ID NO: 1), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence of GCCTA (SEQ ID NO: 27), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 1, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 27. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 1, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 27. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence GTATGGAC (SEQ ID NO: 191) and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 191. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element is SEQ ID 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 115; and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 141. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 53 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 53. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 53 (transposase of hAT-2_AG) MASNVATSSSIWDHFTRNGTKAKCRYCYNEIAFTKGSTSNLKRHMKAKHPSISLERQHLASTTSPSTQNSEPGPSCSSRPPNLISNFFKQPMNSETKRKLDMMLLKLICKDCLPLSIVESEAFKTFVGCLNQNYDLPTRKNVSNALLP SIYNEILVKVQGEVRNATSIALTTDGWTNVNNTSFLGLTAHFIDNDYKLRSCLLECSEISLSHSGQNIAAWIKEVIIKYEIQDKIVGIVTDNAANMKSAARELEFNHVTCFAHSLHLIVKDAIKKSIISTVDEVKRIVMYFKKSPKATN ELADTQSKFNLPNLKLKQDVPTRWNSTYDMLNRFYKNKIAIVACADKLNTLDPIKIDWAILEHSLNALKIFDVATNMVSAEKNITVSHVGLLSKMIIRKLNETDYTIPELGNLVTHLKEGVKKRLEIYSNNQIIAKSMLLDPRIKKQGF HEEPLKYRETYELIIQELIPFQTPSMNAQEPHNIDNEANLLLGEFITWVNNAECEIESPIELAKNELNSFLKIRNIDIKNDPLEWWRIHSSKYPSIYALAKTIICIPGTSVPCERLFSKAGQIYSDKRSRLHPKKFKEIIFIQQNVDKF
[0085] In some embodiments, the transposable element includes: 1) the 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 116; 2) the 3' TR of the RTF. The above RTF includes the nucleic acid sequence of SEQ ID NO: 142. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of TAG (SEQ ID NO: 2), a variant thereof, or a fragment thereof, and 2) a 3'TR containing the nucleic acid sequence of CTA (SEQ ID NO: 28), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 2, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 28. In some embodiments, the transposable element comprises 1) a 5'TRt containing the nucleic acid sequence of SEQ ID NO: 2, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 28. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of ATGTGAAC (SEQ ID NO: 192) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 192. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 116, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 142. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 54 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 54. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 54 (Transposase of HAT1_AG) MMAPTNATTSPVWDHFSPVETGAKCLYCLKVFKYTKGTTSNLKRHLNLVHKTVPYLKQKQPIPQTINIDDEAGPSAVNFQPSNQYFNSNMSIQGYLKKPINSETKKVLDRMLLDLICKECLPFNLVESEIFKKFVYTLNPNYIMPTRKSL SNALLPSVYNQEFEKAKEKLSTAKAIAITSDGWTNLNQISFFALTGHYIDENCKLSSILIECSEFENPHSGRNIANWIQGTLNKFDIEDKIVAMVTDNASNMKAASTELNFCHIPCFAHTLNLIVRDAIKKSVLPVVEEVKRVVMLFKKSP KASQMLADTQKKLNLDQLKMIQEVSTRWNSGYDMLNRFYKNKIALSCADSLKMKISLESHDWEAIEQIVRVLKYFYSATNIVSAQKYITISHVGLLCNVLLTKTSQFRNDEDIAENIQNLVALLIEGLQNKLKIYRSNEQILKSMILDPR IKQLGFQDDVEKFKNICESIISELLPLQKPAVEVEKVVKKVSKDVDMLFGDLLKNKGAQNYKTPRQIAENELHQYLSVENIDLENDPLLWWKEHQVLYPSLYTLAMSTLCIPGTSVPCERLFSKAGQIYSEKRSRLAPKKLQEILFIQQNA
[0086] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 117; and 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 143. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 3, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 29, a variant thereof, or a fragment thereof. In some embodiments, the transposable element comprises: 1) a 5'TR having a nucleic acid sequence having at least approximately 90% (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 3; and 2) a distribution of at least approximately 90% (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 29. The transposterior element includes a 3'TR having a nucleic acid sequence having sequence identity. In some embodiments, the transposterior element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 3, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 29. In some embodiments, the transposterior element further includes a 5'TSD containing the nucleic acid sequence of TA (SEQ ID NO: 193) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposterior element does not include a 5'TSD and / or a 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 117, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 143. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 55 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% (for example, at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 55. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 3 (Tc1-8B_DR's 5'TR) CAGCGGGGAAAATAAGTATTGACACATCAGCATTTTTATCAGTAAGGGGATTTCTAAGTGGGCTACTGACACAAAATTCCTACCAGATGTAGCCATCAAGCCAAATATTGAATTCATACAAAGAAATCAGAACATTTAAGTATACAAGTTGAGTCATAATAAATAAAGTGAAATGACACAGGGAATAAGTATTGAACAC SEQ ID NO: 29(Tc1-8B_DRの3'TR) TGTGTTCAATACTTATTCCCTGTGTCATTTCACTTTATTATTATGACTCAACTTGTATAACTTAAATGTTCTGATTTCTTTGATGAATTCAATATTTGGCTTGATGGCTACATCTGGTAGGAATTTTGTGTCAGTAGCCCACTTAGAAATCCCCTTACTGATAAAAATGCTGATGTGTCAAATACTTATTTTCCCCGCTG SEQ ID NO: 55(Tc1-8B_DRのTRANSポザーゼ) MMGKNKELSQDLRSLIVEKHFDGNGYRRRISRMNLVPVSTVGAIIRKWKKHKFTINRPRSGAPRKIPVRGVQRIIRRVLQEPRTTRAELQEDLASAGTIVSKKTISNALNHHGIHARSPRKTPLLNKKHVEARLKFAKQHLEKPVDYWETIVWSDESKIELFGSHSTHHVW RRNGTAHHPKNTIPTVKFGGGSIMVWGCFSARGTGRLHIIEGRMNGEMYRDILDKNLLPSTRKLKMKRGWTFQQDNDPKHKAKETMKWFQRKKIKLLEWPSQSPDLNPIENLWRELKIKVHKRGPRNLQDLKTVCVEEWARITPEQCRRLVSPYKRRLEAVITNKGFSTKY
[0087] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 118; and 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 144. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 4, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 30, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 4, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 30. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 4, and 2) SEQ The transposterior factor includes a 3'TR containing the nucleic acid sequence ID NO: 30. In some embodiments, the transposterior factor further includes a 5'TSD containing the nucleic acid sequence TTAGAG (SEQ ID NO: 195) and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 195. In some embodiments, the transposterior factor does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 118, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 144. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 56 or a variant thereof. In some embodiments, the transposase is at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%) of the amino acid sequence of SEQ ID NO: 56. It contains an amino acid sequence having sequence identity of 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more. In some embodiments, the transposable element includes a nucleic acid sequence encoding a transposase. In some embodiments, the transposable element does not include a nucleic acid sequence encoding a transposase. SEQ ID NO: 56 (transposase of hAT-6_PM) MAKRKKDDEYRTFQDEWTEEFAFVERAGSAVCLICNDKITSLKRSNVKRHFDTRHATFASKYPAGDSRKKACQELLSRVQASQQQLRVWTRQGDYNSASFAGSLAIVRNGKPFTDGEYAKTFMLDVANELFDDLPNKDKIIKRIQDMPLSARTVHDRTIVMANKVEETQVKDINAAPFFSLALDESTDVSHLSQFSVIARYAVGDTLREESLAVLPLKGSTRGEDLFKSFMEFAQEKSLPMDKLLSVCTDGAPCMVGKNKGFVALLREHENRPILSFHCIIHQEALCAQMCDRQFGEVMSLVIRVINFIVARALNDRQFKTLLDEVVNNYPGLLLHSNVRWLSRGKVLSRFAACLNEIRTFLEMKGVGHPELADTEWLLKFYYLVDLTEHLNQLNVKMQGIGNTVLSLQQAVFAFENKLELFITDLETGRLLHFEKLSQFKDACTASEPIQNLDLHQLAGCTSSLLQSFKARFGEFREHTRLFKFITHPNECSLNTADLNYIPGVSVRDFEAEVADLKASDMWVNKFKSLNEDLERITRQKAELASKHMWTEMKKLQPEDQLILKTWNALPVTYHTLQRVSIAVLTMFGSTYACEQSFSHLKNIKSNLRSRLTDESLNACMKLNLTKYQPDYKDISKSMQHQKSH
[0088] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 119; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 145. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CAGGGG (SEQ ID NO: 5), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 31, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 5, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 5, and 2) a 3'TR containing the nucleic acid sequence CCCCTG (SEQ ID NO: 31). In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence CTGTATAG (SEQ ID NO: 196) and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 196. The transposable element does not contain the 5'TSD and / or 3'TSD. In some embodiments, the transposable element comprises an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 119, and 2) a nucleic acid sequence having at least about 90% ( For example, the transposer includes an RTF containing a nucleic acid sequence having sequence identity of at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the transposer relates to a transposase containing the amino acid sequence of SEQ ID NO: 57 or a variant thereof. In some embodiments, the transposase contains an amino acid sequence having sequence identity of at least about 80% (for example, at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) of the amino acid sequence of SEQ ID NO: 57. In some embodiments, the transposer includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposer does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 57 (transposase of hAT-3_XT) MTSSQPAVKRKIDDEHRQFQEKWEMQYFFVEHRGIPTCLICAEKVAVHKEYNLKRHYSTKHAEECAKYQGDERAKRVASLKACLMRQQDFFKKATKENVASVQASYMVSEMIAKAGKPFTEGEFVKCCMLQVASIICPE KKGQFSKISLSANTVAERISDMSSDIYHQLCEKAKCFDAYSVALDESTDITGTAQLTIYVRGVDCNFELTEELLTIIPMHGQTTANEIFHHLCDAIENAGLPWKRFVGIITDGAPSMTGRKNGLVALVKKKLEEEGIEEE AIALHCIIHQQALCSKCLPCDNVMSVVVKCVNQIRSRGLTHRRFRAFLEEMGSEYGDVLYFTEVRWLSRGNVLKRFFELREEVKAFMEKNGKAVSELSDHKWLMDLAFLVDITQRLNVLNKMLQGQGQLVSAAYDNVRAF STKLVLWKSQLSQTNLCHFPACKELVDAGIPFSGEKYVDAIFKLEKEFDHRFADFKTHRATFQIFVDPFSFDVQDAPPVLQMELIDLQCNSDIKAKFREMSGKADTHVQFLRELPPSFPELSRMFKRTMCLFGSTYLCEK
[0089] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence with SEQ ID NO: 120; and 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence with SEQ ID NO: 146. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence with SEQ ID NO: 6, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence with SEQ ID NO: 32, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 6, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 6, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 120, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 146. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 58 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 58. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator encodes the transposase. It does not contain nucleic acid sequences. SEQ ID NO: 6 (Tc1-3_Xt's 5'TR) CAGTGGAGGAAATAATTATTTGACCCCTCACTGATTTTGTAAGTTTGTCCAATGACAAAGAAATGAAAAGTCTCAGAACAGTATCATTTCAATGGTAGGTTTATTTTAACAGTGGCAGATAGCACATCAAAAGGAAAATCGAAAAAATAACTTTAAATAAAAGATAGCAACTGATTTGCATTTCATTGAGTGAAATAAGTTTTTGAACCC SEQ ID NO: 32 (3'TR of Tc1-3_Xt) GGGTTCAAAAACTTATTTCACTCAATGAAATGCAAATCAGTTGCTATCTTTTATTTAAAGTTATTTTTTCGATTTTCCTTTTGATGTGCTATCTGCCACTGTTAAAATAAACCTACCATTGAAATGATACTGTTCTGAGACTTTTCATTTCTTTGTCATTGGACAAACTTACAAAATCAGTGAGGGGTCAAATAATTATTTCCTCCACTG SEQ ID NO: 58 (Transposase of Tc1-3_Xt) MGKTKELSKDVRDKIVDLHKAGMGYKTISKKLGEKVTTVGAIVRKWKEHKMTINRPRSGAPRKISPRGVSMILRKVKKHPRTTREELVNDLKLAGTTVTKKTIGNTLHRNGLKSCRARKVPLLKKAHVQARLKFANEHLNDSVSDWEKVLWSDETKIELFGINSTRCVWRKKNAAYDPQNTVPTVKHGGGNILLWGCFSAKGTGQLIRINGKMDGAMYREILNDNLPSARKLKMGRGWVFQHDNDPKHTAKATKEWLKKKHIKVMEWPSQSPDLNPIENLWRELKLRVAQRQPRNLRDLEMICKEEWTNIPKPKMCANLVINYKKRLTSVLANKGFSTKY
[0090] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 121; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 147. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CAGGGGTGGCGAACC (SEQ ID NO: 7), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence GGTTCGCCACCCCTG (SEQ ID NO: 33), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 7, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 33. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 7, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 33. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of GCTATAC (SEQ ID NO: 194) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 194. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 121, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 147.In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 59 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 59. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not comprise a nucleic acid sequence encoding the transposase. SEQ ID NO: 59 (transposase of hAT-5_DR) MAESGKRTAKRKYEDERRTFLSEWEDLYFFVERNGKPFCLICQTSLSHFKASNLERHFTSLHSAVAREFPKGSELRKHKVKTLKGQAEKQTQLFRKFTKHSETVTLASYQLAWNIARAKKPYLEGEFVKKCLSDAVAILCPENENLKRSVKDLQ LSRHTVEQRISDIDNSVETHLLSDLQKCQYFSIALDESCDVQDKPQLAIFVRFVSEDCTIREELLDIVPLKDRTRGIDLKETLMTVVEKANLQLSKLTAIVTDGAPAMLGSERGLVGLCKADDRFPAFWTFHCIIHQEHLVSKKLNLDHIMKPV LEIVNFVRTHALNHRQFKNLIDELDEDLPSDLLFHCAVRWLSRGHVLSRFFELLNPVKLFLAEKHKEYPELHDPQWISDLAFLVDVLHYLNGLNVDLQGKLKMLPDLVQSVFAFVNKLKLFKTHLQKRDYTHFPTLLKASGQEAEVVKGSTARY ATLLENLEQSFEERFSNLRQKRQQITFLINPFTAESGCLKAPLVEDEAASQLEMIELSEDDRLKSVLREGTVEFWKIVPVERYPNVKQAALKLLSMFGSTYVCESLFSTLKLVKSKHRSVLTDTHVKELLRVATTEYEPDLKKIVETKECQVSH
[0091] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 122; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 148. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 8, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 34, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 8, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 8, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TR containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 122, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 148. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 60 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 60. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 8 (Tc1-3_FR 5'TR) CAGTGAGAGTAAAAAAGTATTTGATCCTTGCTGATTTTGTTGGTTTGTCCACTAATAAAGACATGATCATTCTATACTTTTAATGGTAGATGTATTCTAACATGGAGAGACAGAATATCAAAAAGAAAATC SEQ ID NO: 34 (Tc1-3_FR 3'TR) GATTTTCTTTTTGATATTCTGTCTCTCCATGTTAGAATACATCTACCATTAAAAGTATAGAATGATCATGTCTTTATTAGTGGACAAACCAACAAAATCAGCAAGGGATCAAATACTTTTTACTCTCACTG SEQ ID NO: 60 (Transposase of Tc1-3_FR) MGKTKELSQDLRDRIVDLHKSGMGYKNISKLLGIKVTTIGAIVRKFKKYNMTINRPRSGAPQKISPRGVAMIMRTVRNRPATTQQELVNDLQAAGTKVTRKTIGNTLRRNGLKSCSARKVPLLKKAHVEARLKYANDHLKDAQSDWEKVLWSDETKIE LFGLNSTRHVWRKKNAAYDPRNTVPTVKHGGGNIMFWGCFSAKGTGLLHRITGKMDGAMYRAVLRDNLLPSARKLKMGRGWVFQHDNDPKHTAKATKEWLKKNHIKVMEWPSQSPDLNPIENLWRELKVRVAKRQPTNLNDLERICKEEWAKIPPDMCANLVVNYNKRLTAVLANNGFATKY
[0092] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 123; 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 149. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 9, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 35, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes: 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence SEQ ID NO: 9; and 2) SEQ The transposable element includes a 3' TR having a nucleic acid sequence with at least approximately 90% sequence identity (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) of the nucleic acid sequence with ID NO: 35. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence with SEQ ID NO: 9, and 2) SEQ It includes a 3'TR containing the nucleic acid sequence ID NO: 35. In some embodiments, the transposable element includes a 5'TSD containing the nucleic acid sequence SEQ ID NO: 193, and SEQ The transposase comprises a 3'TSD containing the nucleic acid sequence ID NO: 193. In some embodiments, the transposase does not include a 5'TSD and / or a 3'TSD. In some embodiments, the transposase comprises an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 123, and an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 149. In some embodiments, the transposase relates to a transposase containing the amino acid sequence SEQ ID NO: 61 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 61. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 9 (Tc1-5_Xt's 5'TR) CAAACCGGATTCCAAAAAAGTTGGGACACTAAACAAATTGTGAATAAAAACTGAACGCAATGATGTGGAGGTGCCAACTTCTAATATTTTATTCAGAATAGAACATAAATCACGGAACAAAAGTTTAAACTGAGAAAATGTACCATTTTAAGGGAAAAATATGTTGATTCAGAATTTCATGGTGTCAACAAATCCCAAAAAAGTTGGGACAAG SEQ ID NO: 35 (3'TR of Tc1-5_Xt) CTTGTCCCAACTTTTTGGGATTTGTTGACACCATGAAATTCTGAATCAACATATTTTTCCCTTAAATGGTACATTTTCTCAGTTTAAACTTTTGTTCCGTGATTTATGTTCTATTCTGAATAAAATATTAGAGTTGGCACCTCCACATCATTGCGTTCAGTTTTTATTCACAATTTGTTTAGTGTCCCCAACTTTTTTGGAATCCGGGTTG SEQ ID NO: 61 (Tc1-5_XtのTRANSポザーゼ) MIGYKKFSEWQCLSEAKMGRGSPIPTMLRRKIVEQYQKGVTQRKIAKILHLSSSTVHNIIRRFRESGTISVRKGQGRKTILDARDLRALKRHCTTNRNATVKEITEWAQEYFQKPLSVNTIHRAIRRCQLKLYSAKKKPFLSKIHKLRRFHWARDHLKWSVAKWKTVLWSDESRFEVL FGNLGRHVIRTKEDKDNPSCYQRSVQKPASLMVWGCMSACGMGSLHVWKGSINAEKYIQVLEQHMLPSRRHLFQGRPCIFQQDNARPHSASITTSWLRRRRIRVLKWPVCSPDLSPIENIWRIIKRKVRQRRPKTIEQLEACIRQEWESIPIPKLEKLVSSVPRRLLSVVRRRGDATQW
[0093] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 124; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 150. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 10, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 36, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 10, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 36. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 10, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 36. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 124, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 150. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 62 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the amino acid sequence of SEQ ID NO: 62. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 10 (c1-10_Xt's 5'TRT) CAGTGGCTTGCAAAAGTATTCGGCCCCCTTGAACTTTTCCACATTTTGTCACATTACAGCCACAAACATGAATCAATTTTATTGGAATTCCACGTGAAAGACCAATACAAAGTGGTGTACACGTGAGAAGTGGAACGAAAATCATACATGATTCCAAACATTTTTTACAAATAAATAACTGCAAAGTGGGGTGTGCGTAATTATTCAGCCCCCT SEQ ID NO: 36 (c1-10_Xt's 3'TR T) AGGGGGCTGAATAATTACGCACACCCCACTTTGCAGTTATTTATTTGTAAAAAATGTTTGGAATCATGTATGATTTTCGTTCCACTTCTCACGTGTACACCACTTTGTATTGGTCTTTCACGTGGAATTCCAATAAAATTGATTCATGTTTGTGGCTGTAATGTGACAAAATGTGGAAAAGTTCAAGGGGCCGAATACTTTTGCAAGCCACTG SEQ ID NO: 62 (Transposase t of Tc1-10_X) MKSKEHTRQVRDKVIEKFKAGLYKKISKALNIPRSTVQAIIQKWKEYGTTVNLPRQGRPPKLTGRTRRALIRNAAKRPMVTLDELQRSTAQVGESVHRTTISRALHKVGLYGRVARRKPLLTENHKKSRLQFATSHVGDTANMWKKVLWSDETKMELFGQNAKRYVWR KTNTAHHSEHTIPTVKYGGGSIMLWGCFSSAGTGKLVRVDGKMDGAKYRAILEENLLESAKDLRLGRRFTFQQDNDPKHKARATMEWFKTKHIHVLEWPSQSPDLNPIENLWQDLKTAVHKRCPSNLTELELFCKEEWARISVSRCAKLVETYPKRLAAVIAAKGGSTKY
[0094] In some embodiments, the transposable element includes: 1) the 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 125; 2) the 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 151. In some embodiments, the transposable element is 1) TACAGTGTCGGACAAATC(SEQ I The transposable element comprises: 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 11), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence of GATTTGTCCGACACTGTA (SEQ ID NO: 37), a variant thereof, or a fragment thereof. In some embodiments, the transposable element comprises: 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 11; and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 37. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 11, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 37. In some embodiments, the transposable element further includes a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193, and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 125, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 151. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 63 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 63. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 63 (Transposase for Mariner2_AG) MGRSAHSTQQQRLDIKRLSAAGYSQRKIAEILGRSKTFVYNALHSTGTKIPTGRPRKTSARDDARMTRLCKADPFKSARAIRDELQLSVSDRTVQRRLFENNLVGRNPRKVPLLRRCHVQARLQFAREHYDWAGENLNKWRNVLWSDESKVNLVGSDGKRFVRRPK NTAYRPQYTLKTVKHGGGNIMVWACFSWYGVGPIFWIKDIMDQHRYLNIIQTVMLPHAEWEMPLKWQFMHDNDPKHTAKAVKKWFVDQKIDVMNWPAQSPDLNPIENLWKIVKAKLPPAGSRTKEKLWKHIENAWYSIPPSTCKSLVESMPKRMRAVIRNNGHATKY
[0095] In some embodiments, the transposable element includes: 1) a 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 126; and 2) a 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 152. In some embodiments, the transposable element includes: 1) a 5' TR containing the nucleic acid sequence SEQ ID NO: 12, a variant thereof, or a fragment thereof; and 2) a 3' TR containing the nucleic acid sequence SEQ ID NO: 38, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 12, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 38. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 12, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 38. In some embodiments, the transposable element includes a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193, and S The transposase comprises a 3'TSD containing the nucleic acid sequence of EQ ID NO: 193. In some embodiments, the transposase does not contain a 5'TSD and / or a 3'TSD. In some embodiments, the transposase comprises an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 126, and an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 152. In some embodiments, the transposase relates to a transposase containing the amino acid sequence of SEQ ID NO: 64 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 64. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 12 (Tc1-1_Xt's 5'TR) CAGTGGTGTGAAAAACTATTTGCCCCCTTCCTGATTTCTTATTCTTTTGCATGTTTGTCACAC SEQ ID NO: 38 (3'TR of Tc1-1_Xt) GTGTGACAAACATGCAAAAGAATAAGAAATCAGGAAGGGGGCAAATAGTTTTTCACACCACTG SEQ ID NO: 64 (Transposase of Tc1-1_Xt) MPRSKEIQEQMRTKVIEIYQSGKGYKAISKALGLQRTTVRAIIHKWQKHGTVVNLPRSGRPTKITPRAQRQLIREATKDPRTTSKELQASLASIKVSVHDSTIRKRLGKNGLHGRFPRRKPLLSKKNIKARLNFAKHLNDCQDFWENTLWTDETKVELFGRCVSRYIWR KSNTAFQKKNIIPTVKYGGGSVMVWGCFAASGPGRLAVIDGTMNSTVYQKILKENVRPSVRQLKLKRSWVLQQDNDPKHTSKSTSEWLKKNKMKTLEWPSQSPDLNPIEMLWHDLKKAVHARKPSNKAELQQFCKDEWAKIPPERCKRLVASYRKRLIAVIAAKGGPTSY
[0096] In some embodiments, the transposable element includes: 1) the 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 127; 2) the 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 153. In some embodiments, the transposable element includes: 1) CACTGGTGGACAT(SEQ ID NO: 13) A 5'TR comprising a nucleic acid sequence, a variant thereof, or a fragment thereof; and 2) a 3'TR comprising the nucleic acid sequence ATGTCCACCAGTG (SEQ ID NO: 39), a variant thereof, or a fragment thereof. In some embodiments, the transposable element comprises 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 13; and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 39. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 13, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 39. In some embodiments, the transposable element further includes a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193, and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises an LTF having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 127, and 2) a sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 153. The transposase comprises an RTF containing a nucleic acid sequence having sequence identity. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 65 or a variant thereof. In some embodiments, the transposase contains an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 65. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not contain a nucleic acid sequence encoding the transposase. SEQ ID NO: 65 (Transposase of Tc1-1_AG) MGRGKHCTPEERKHIQGLYRENVPIKTICKAFGRSRTFVDNAIRSEATGKSTGRPRKTTADVDAQIVEMIRADPFKTCTRIKQELGLQVSAKTVSRRLHAAGFCARRPRKVRKLLPHHVEARIRFAEEHLAASIFWWSKIIFSDESRINLDGSDGIKYVWRFPNQA YHPKNTIKTLSHGGGHVMVWGCFSWHGTGPLFRINGTLNSEGYRKILSRKMLPYARQQFGDEEHYIFQHDNDSKHTSRTVKCYLANQDVQVLPWPALSPDLNPIENLWSTLKRQLKNQPARSADDLWTRCKVMWERIPRSECRNLIGDMAKRCQEVIANNGHQIDR
[0097] In some embodiments, the transposable element includes: 1) a 5' TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 128; and 2) a 3' TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 154. In some embodiments, the transposable element includes: 1) a 5' TR containing the ATATACAC (SEQ ID NO: 14) nucleic acid sequence, a variant thereof, or a fragment thereof; and 2) a 3' TR containing the GTGTATAT (SEQ ID NO: 40) nucleic acid sequence, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 14, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 40. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 14, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 40. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of AT (SEQ ID NO: 198) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 198. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 128, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 154.In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 66 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 66. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not comprise a nucleic acid sequence encoding the transposase. SEQ ID NO: 66 (Transposase of Tc1DR3_Xt) MGKKGDLSNFERGMVVGARRAGLSISQSAQLLGFSRTTISRVYKEWCEKGKTSSMRQSCGRKCLVDARGQRRMGRLIQADRRATLTEITTRYNRGMQQSICEATTRTTLRRMGYNSRRPHRVPLISTTNRKKRLQFAQAHQNWTVEDWKNVAWSDESR FLLRHSNGRVRIWRKQNENMDPSCLVTTVQAGGGGVMVWGMFSWHTLGPLVPIGHRLNATAYLSIVSDHVHPFMTTMYPSSDGYFQQDNAPCHKARIISNWFLEHDNEFTVLKWPPQSPDLNPIEHLWDVVERELRALDVHPTNLHQLQDAILSIWANISKECFQHLVESMPRRIKAVLKAKGGQTPY
[0098] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 129; 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 155. In some embodiments, the transposable element includes: 1) a 5'TR containing the CAGGGGTCACCAAACT (SEQ ID NO: 15) nucleic acid sequence, its variant or fragment, and 2) SEQ The transposable element comprises: 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 15; and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 41. In some embodiments, the transposable element comprises: 1) a 5'TR having the nucleic acid sequence of SEQ ID NO: 15; and 2) AGTTTGGTGACCCCTG(SEQ The transposterior factor includes a 3'TR containing the nucleic acid sequence of ID NO: 41). In some embodiments, the transposterior factor further includes a 5'TSD containing the nucleic acid sequence of CTCTAGAC (SEQ ID NO: 199) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 199. In some embodiments, the transposterior factor does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 129, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 155. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 67 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 67. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 67 (transposase of hAT-3_PM) MATKKRKVDAECRVFNEEWGVKYFFVETKDQKASCVICSESVAVLKEYNLRRHYETKHLSTYSKFSGKLRSEKYESMKRGLETQRNLFTRKFAENESVTRTSYKIVHKMAERGKPFTDGNFIKECMMEAANDLCPERADLFGSISLSASS VVRRTEELGENIVLQIREKARNLLWYSLALDESTDLSSTSQLLVFIRGVNLDFQITEELASVCSMHGTTTGKDIFMEVQKTLQDYNLQWNQLRGVTVDGGKNMAGVRKGLVGQIRTQLEDLQIPGALFIHCIIHQQALCGKDLDISCVLKP VVSAVNFIRGALNHRQFQAFLEEMDSDFCDLPYHTAVRWLSCGKVLFRFYKLRNEIDVFLTEKDRADPQLSDPTWLSKLSFLVDITSHMNELNLKLQGKDNLVCDLYRIIKGFRRKLSLFEAQLEGENFSHFHCFKEFCATIAEDVNLD FPKKIIRDLKKHFLKRFSDLDRIESDILLFQNPFDCNLDDVPVELQLELIDLQANDLLKEKHREGKLVEFYRCLPDVEFPKLKKFAAGMASVFGTTYVCEQTFSKMKYVKSTHRTRLTDEHLKAILLIGCSNSKPNIDDILKAKRQFHKSH
[0099] In some embodiments, the transposable element includes: 1) the 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 130; 2) the 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 156. In some embodiments, the transposable element includes: 1) CAGGGGTCACCAAACT (SEQ ID NO: 16) A 5' TR containing a nucleic acid sequence, a variant thereof, or a fragment thereof; and 2) A 3' TR containing the nucleic acid sequence AGTTTGGTGACCCCTG (SEQ ID NO: 42), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 16; and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence of SEQ ID NO: 42. In some embodiments, the transposable element comprises 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 16, and 2) SEQ ID The transposterior element comprises a 3'TR containing the nucleic acid sequence NO: 42. In some embodiments, the transposterior element further comprises a 5'TSD containing the nucleic acid sequence SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 193. In some embodiments, the transposterior element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 130, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 156. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 68 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 68. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 68 (Transposase of Tc1-1_PM) MATKKRKVDAECRVFNEEWGVKYFFVETKDQKASCVICSESVAVLKEYNLRRHYETKHLSTYSKFSGKLRSEKYESMKRGLETQRNLFTRKFAENESVTRTSYKIVHKMAERGKPFTDGNFIKECMMEAANDLCPERADLFGSISLSASS VVRRTEELGENIVLQIREKARNLLWYSLALDESTDLSSTSQLLVFIRGVNLDFQITEELASVCSMHGTTTGKDIFMEVQKTLQDYNLQWNQLRGVTVDGGKNMAGVRKGLVGQIRTQLEDLQIPGALFIHCIIHQQALCGKDLDISCVLKP VVSAVNFIRGALNHRQFQAFLEEMDSDFCDLPYHTAVRWLSCGKVLFRFYKLRNEIDVFLTEKDRADPQLSDPTWLSKLSFLVDITSHMNELNLKLQGKDNLVCDLYRIIKGFRRKLSLFEAQLEGENFSHFHCFKEFCATIAEDVNLD FPKKIIRDLKKHFLKRFSDLDRIESDILLFQNPFDCNLDDVPVELQLELIDLQANDLLKEKHREGKLVEFYRCLPDVEFPKLKKFAAGMASVFGTTYVCEQTFSKMKYVKSTHRTRLTDEHLKAILLIGCSNSKPNIDDILKAKRQFHKSH
[0100] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 131; 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 157. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CAGGC (SEQ ID NO: 17), a variant thereof, or a fragment thereof; and 2) GCCTG (SEQ ID The transposable element comprises: 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 17; and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 43. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 17, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 43. In some embodiments, the transposable element further includes a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193, and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 131, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 157. In some embodiments, the transposable element includes SEQ ID NO: The present invention relates to a transposase comprising the amino acid sequence 69 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 69. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. SEQ ID NO: 69 (Transposase of Mariner-4_AMi) MSAKRKSMTSPGGSAKKQRKAIDLEMKMKIINDYEAGKKVKAIARDLELAHSTISTILKDKDRVKEAVQASTGFKAIITRQRKGLIHEMEKLLAIWFDDQIQKRMPMSLLIIQAKARSIFETLKAREGEESTETFTA SRGWFQRFRRRRFNVHNRSISGEAASADVEAAEKFVDQFDEIIEKGGYRPEQIFNVDETGLFWKKMPERSYIHKEAKAMPGFKAFKDRVTLLLGGNVAGFKLKPFLIHRSENPRALKQVSKHTLPVYYRANSKAWMTQA LFEDWFINCFIPSVKHYCLEKGVPFKIILLLDNAPGHPQHLDDLHPDVKVVYLPKNTTAILQPMDQGAIATFKVYYLRATFSKAVAATESDEVTLRDFWKSYNILHCIKNIESAWEGVTEKCMQGIWKKCLKVFVNN FEGFDKDEHVDVINKKIVELANVLNLDVEVKDIEELVEYVEGELTNEDLIELEAQQHLEEEEEEEEERMEEVQKKFTVNGLAGVFSKVNAAILELEGMDPNVERFTKVERQMNELLRCYCEIYEEKKKKKRKLQSRPH
[0101] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 132; 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 158. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CAGTGATGGCGAACCT (SEQ ID NO: 18), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence AGGTTCGCCATCACTG (SEQ ID NO: 44), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes: 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence SEQ ID NO: 18; and 2) SEQ ID The transposterior factor comprises a nucleic acid sequence NO: 44 and a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity. In some embodiments, the transposterior factor comprises 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 18 and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 44. In some embodiments, the transposterior factor further comprises a 5'TSD containing the nucleic acid sequence GTCTAGAG (SEQ ID NO: 197) and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 197. In some embodiments, the transposterior factor does not include a 5'TSD and / or a 3'TSD. In some embodiments, the transposable element comprises an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 132, and 2) a nucleic acid sequence having at least about 90% sequence identity with the nucleic acid sequence of SEQ ID NO: 158 The transposase includes an RTF containing a nucleic acid sequence having sequence identity of at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 70 or a variant thereof. In some embodiments, the transposase contains an amino acid sequence having sequence identity of at least about 80% (for example, at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more). In some embodiments, the transposase includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 70 (Transposase of Myotis_hAT1) MENQKNKKARLNDKGSGSSSSRPFQDIWTEMYGIVDRNNRSLCILCNETVVSRTWNVKRHFETNHCQLLKKSIDERKEYISRQLHLYKSQSNSIIKFVRCSTNLTSASYCIAHSIAQHGKALSDGEFIKETFLRCAPALFHDMKNKDDIIKRISELPLSR NTIKDRIIDLNKNVQHQLKKDLNSCKYFSISLDETTDVTSNARLAIIARYSDGLTMREELIKLESVPISTSGNEICKVVIKTFSDLNIDISKIVSVTTDGAPNMVGKNVGFLKLFMEAIRHPLVPFHCIIHQEVLCAKSGFSELNDLMSVVTKIVNFIASR PLHKREFSALLQEVDSTYSGLLMFNNVRWLSRGKVLERFVECFEEITVFIENKDLANFPQINDHKWVSNLMFFTDLSVHMNELNLKLQGFGKSIDVMFGYIKSFESKLKIFKRDIETKTYKYFPRIKKYFEKATTTSVMESILMYQNIVNSLFEQFSERFD QFRSLEQTIKIIKYPDVVVYSTLELKDFQWMQIDDLEMQLADFQGSIWTHVFVDLRSKLENLERNRLNNEEECNYEQEILSAWNRVPDTFSDIKNLAMALLTIFSSTYFCETLFSALNNIKTNKRNRLTDEVSGACLALKCTKYEPAIEELADKLQHQKSH
[0102] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 133; 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 159. In some embodiments, the transposable element includes: 1) a 5'TR containing the CAG (SEQ ID NO: 19) nucleic acid sequence, a variant thereof, or a fragment thereof; and 2) a CTG (SEQ ID NO: The transposterior element comprises 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 19, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 45. In some embodiments, the transposterior element comprises 1) a 5'TR having the nucleic acid sequence of SEQ ID NO: 19, and 2) a 3'TR having the nucleic acid sequence of SEQ ID NO: 45. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence GTCTAGAC (SEQ ID NO: 200) and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 200. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element is SEQ ID 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 133; and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 159. In some embodiments, the transposase is associated with the amino acid sequence of SEQ ID NO: 71 or a variant thereof. In some embodiments, the transposase contains an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 71. In some embodiments, The above transposable element includes a nucleic acid sequence encoding a transposase. In some embodiments, the above transposable element does not include a nucleic acid sequence encoding a transposase. SEQ ID NO: 71 (transposase of hAT-7_PM) MSGSRKRKVDNECRVFNTEWTTKYFFTEVQSKAVCLICRETVAVFKEYNISRHFATKHANYASKQSTQERAATAQRLAANLQTQQHFFHRQTAIQESTTKASFLVAFEIAKASKPFSEGEFVKECMVQTADILCPEIKSKFEKVSLSRRT VTRRVELIDENIASQLNKKSDSFELYSLALDESTDVKDTAQLLIFIRGIDDSFAITEEFLTMESLKGTTRGEDLYNQVSAVIERMKLPWSKLVNVTTDGSPNLTGKNVGLLKRIQNKVKEENPDQDLIFLHCIIHQESLCKSVLQLNHVVN PAVKLVNFIRARGLQHRQFITFLEETDADHQDLLYHSRVRWLSLGKVLQRVWELKEDIIAFLELMGKSDEFPELSDKNWLSDFAFAVDIFSHMNELNVKLQGKDQFVHDMYKHVKAFKSKLTLFSRQIANKSFAHFPTLAMQEEAPRNAK KYSKSLEDLHGEFCRRFSDFENIEQSLQLVSCPLSQDSETAPQELQLELIDLQSDSVLKEKFNSVKLNDFYASLNRATFPNLRRTAQKMLTLFGSTYVCEQTFSVMNANKARHRSKLTDQHLRSILRIATTKITPDLDALAKMGDQQHCSH
[0103] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 134; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 160. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 20, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 46, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 20, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 46. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 20, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 46. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 134, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 160. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 72 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 72. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 20 (Tc1-11_Xt's 5'TR) CAGTGGATATAAAAAGTCTACACCCCCTGGTAAAATGTCAGGTTCCTGTGCTGTACAAAAATGAGACAAAGATAAATCATTTCAGAACTTTTTCCACCTTTAATGTGACCTATAAACTGTACCACTCAATTGAAAAACAAACTGAAATCTTTAGG SEQ ID NO: 46 (3'TR of Tc1-11_Xt) CCTAAAAGATTTCAGTTTGTTTTCAATTGAGTGGTACAGTTTATAGGTCACATTAAAGGTGGAAAAAGTTCTGAAATGATTTATCTTTGTCTCATTTTTGTACAGCACAGGAACCTGACATTTTACCAGGGGTGTGTAGACTTTTTATATCCACTG SEQ ID NO: 72 (Transposase of Tc1-11_Xt) MVHRELPKHQRDLIVKRYQSGEGYKRISKALDIPWNTVKTVIIKWRKYGTTVTLPRTGRPSKIDEKTRRKLVREATKRPTATLKELQEYLASTGCVVHVTTISRILHMSGLWGRVARRKPFLTKKNIQARLHFAKTHLKSPKSMWEKVLWSDETKVELFGHNSKKYVWRKNNTAHHQKNTIPTVKHGGGSIMLWGCFSSAGTGALVKIEGIMNSSKYQSILAQNLQASARKLNMRNFIFQHDNDPKHTSKSTKEWLHRKKIKVLEWPSQSPDLNPIENLWGDLKRAVHRRCPRNLTDLECFCKEEWANLAKSKCAMLIDSYPKRLSAVIKSKGASTKY
[0104] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 135; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 161. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 21, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 47, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 21, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 47. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 21, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 47. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 135, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 161. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 73 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 73. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 21 (Tc1-16_Xt 5'TR) CAGTCATGGCCAAAATTGTTGGCACCCCAGAAATTTTTCCAGAAAATCAAGTATTTCTCACAGAAAAGTATTGCAGTAACACATGTTTTGCTATACACATGTTTATTCCCTTTGTGTGTATTGGAACAGAACAAAAAAGGGAGGAAAAAAAGCAAATTGGACATAATGTCACACAAAACTCCAAAAATGGGCTGGACAAAATTATTGGCACCCTT SEQ ID NO: 47 (Tc1-16_Xt 3'TR) AAGGGTGCCAATAATTTTGTCCAGCCCATTTTTGGAGTTTTGTGTGACATTATGTCCAATTTGCTTTTTTTCCTCCCTTTTTTGTTCTGTTCCAATACACACAAAGGGAATAAACATGTGTATAGCAAAACATGTGTTACTGCAATACTTTTCTGTGAGAAATACTTGATTTTCTGGAAAAATTTCTGGGGTGCCAACAATTTTGGCCATGACTG SEQ ID NO: 73 (Transposase of Tc1-16_Xt) MDNRKRRRELSEDLRTKIVEKYQQSQGYKSISRDLDLPLSTVRNIIKKFATHGTVANLPGRGRKRKIDERLQRRIVRMVDKQPQTSSKEIQAVLQAQGASVSARTIRRHLNEMKRYGRRPRRTPLLTQRHKKARLQFAKMYLSKPQSFWENVLWTDETKIELFGKAHHSTVYRKRNEAYKEKNTVPTVKYGGGSMMFWGCFAASGTGCLECVQGIMKSEDYQRILGRTVEPSVRKLGLRPRSWVFQQDNDPKHTSKSTQKWMATKRVRVLKWPAMSPDLNPIEHLWRDLKIAVGKRRPSNKRDLEQFAKEEWSKIPGERCKKLIDGYRKRLISVIFSKGCATKY
[0105] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 136; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 162. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 22, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 48, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 22, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 48. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 22, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 48. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 136, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 162. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 74 or a variant thereof.In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 74. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 22 (Tc1-4_Xt 5'TR) CAGTGAGGAACATAAGTATTTGAACACCCTGCGATTTTGCAAGTTCTCCCACTTAGAAATCATGGAGGGGTCTGAAATTCACATTGTAGGTGCATTCCCACTGTGAGAGACAGCATTTAAAAAAAAAATTCAGGAAATCACATTGTATGATTTTTAAAGAATGTATTTGTATTGCACTGCTGCACATAAGTATTTGAACACCTG SEQ ID NO: 48 (3'TR of Tc1-4_Xt) CAGGTGTTCAAATACTTATGTGCAGCAGTGCAATACAAATACATTCTTTAAAAATCATACAATGTGATTTCCTGAATTTTTTTTTTAAATGCTGTCTCTCACAGTGGGAATGCACCTACAATGTGAATTTCAGACCCCTCCATGATTTCTAAGTGGGAGAACTTGCAAAATCGCAGGGTGTTCAAATACTTATGTTCCTCACTG SEQ ID NO: 74 (Transposase of Tc1-4_Xt) MGKTKELSKDTRDKIVDLHKAGKGYGAIAKQLGENRSTVGAIVRKWKRLKTTVSLPRTGAPCKISPRGVSLMIRKVRNQPRTTREELVNDMKRAGTTVSKVTVGRTLRRHGFKSCIARKVPLLKSSHVQARLKFANDHLDDPEEAWEKVMWSDETKVE LFGLNFTRRVWRKNKDELHPKNTIPTVKHGGGNIMLWGCFSAKGTGRLHCIKERMNGAMYCEILSNNLLPSVRALKMGRGWVFQHDNDPKHTARITKEWLRKRHIKVLEWPSQSPDLNPIENLWRELKLRVAQRQPRNLTDLEEICVEEWAKIPVAVCANLVKNYRKRLTSVIANKGFCTKY
[0106] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 137; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 163. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CACTGCTCAAAAAAATAAAGGGAACAC (SEQ ID NO: 23), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence GTGTTCCCTTTATTTTTTTTGAGCAGTG (SEQ ID NO: 49), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 23, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 49. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 23, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 49. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 49. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 137, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 163.In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 75 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 75. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not comprise a nucleic acid sequence encoding the transposase. SEQ ID NO: 75 (Transposase of Tc1-15_Xt) MRRSLQPTQMAQVVQLIQDGTSMRAVARRFAVSVSVVSRAWRRYQETGQYIRRRGGGRRRATTQQQDRYLRLCARRNRRSTARALQNDLQRATNVHVSAQTVRNRLHEGGMRARRPQVGVVLTAQHRAGRLAFAREHQDWQIRHWRPVLFTDESRFTLSTCDRRDRVWRR RGERSAACNILQHDRFGSGSVMVWGGISLGGRTALHVLARGSLTAIRYRDEILRPLVRPYAGAVGPGFLLMQDNARPHVAGVCQQFLQDEGIDAMDWPARSPDLNPIEHIWDIMSRSIHQRHVAPQTVQELVDALVQVWEEIPQETIRHLIRSMPRRCREVIQARGGHTHY
[0107] In some embodiments, the transposable element includes: 1) a 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 138; 2) a 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 164. In some embodiments, the transposable element includes: 1) a 5' TR containing the nucleic acid sequence CACTGCTCAAAAAAATTAGAGGAACACTT (SEQ ID NO: 24), a variant thereof, or a fragment thereof; and 2) the nucleic acid sequence AAGTGTTTCCTCTAATTTTTTTTGAGCAGTG (SEQ ID NO: 50), a variant thereof, or a fragment thereof. Includes 3'TR and, in some embodiments, the above transposable element is 1) SEQ ID NO: 24 comprises a 5'TR having a nucleic acid sequence with at least approximately 90% sequence identity (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%), and a 3'TR having a nucleic acid sequence with at least approximately 90% sequence identity (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) of SEQ ID NO: 50. In some embodiments, the transposable element comprises 1) SEQ ID 1) A 5'TR containing the nucleic acid sequence NO: 24, and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 50. In some embodiments, the above transposable element is SEQ ID The transposable element further comprises a 5'TSD containing the nucleic acid sequence NO: 193 and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 138, and an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 164. In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 76 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 76. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not comprise a nucleic acid sequence encoding the transposase. SEQ ID NO: 76 (Transposase of TC1_FR2) MVRYLDATEVAQVVQLLQDGTSIRAVARRFAVSPSTVSRAWRRFQETGSYSRRAGQGRRRSLNPQQDRYLLLCARRNRMSTARALQNDLQQATGVNVSDQTIRNRLHEGGLRARRPVVGPVLTARHRRARLAFAIDHQNWQLRHWRPVLFTDESRFNLSTCDRRERVWRC RGERYAACNIIQHDRFGGGSVMVWGGISLEGRTDLYRLDNGTLTAIRYRDEILGPIVRTYAGAVGPGFLLVHDNARPHVARVCRQFLEDEGIDTIDWPPRSPDLNPIEHLWDIMFRSIRRRQVDPQTVQELSDALVQIWEEIPQDTIRRLIRSMPRRCQACVQARGGHTNY
[0108] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 139; 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 165. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence of TAG (SEQ ID NO: 25), a variant thereof, or a fragment thereof; and 2) CTA (SEQ ID NO: The transposable element comprises 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 25, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 51. In some embodiments, the transposable element comprises 1) a 5'TR having a nucleic acid sequence having SEQ ID NO: 25, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 51. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of ATCATCAT (SEQ ID NO: 201) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 201. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element is SEQ I 1) The transposase comprises an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 139, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 165. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 77 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 77. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 77 (transposase of hAT-9_XT)
[0109] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 140; 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 168. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CCGTATTTTCCGCACTATAAGGCGCACC (SEQ ID NO: 26), a variant thereof, or a fragment thereof; 2) a 3'TR containing the nucleic acid sequence GGTGCGCCTTATAGTGCGGAAAATACGG (SEQ ID NO: 52), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes: 1) SEQ ID NO: The transpossession factor comprises: 1) a 5' TR having nucleic acid sequences with at least approximately 90% sequence identity (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) of nucleic acid sequences with SEQ ID NO: 52 and a 3' TR having nucleic acid sequences with at least approximately 90% sequence identity (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) of nucleic acid sequences. In some embodiments, the transpossession factor comprises: 1) SEQ ID NO: It includes a 5'TR containing 26 nucleic acid sequences and a 3'TR containing the nucleic acid sequence SEQ ID NO: 52. In some embodiments, the transposable element is SEQ ID NO: The transposterior factor further comprises a 5'TSD containing the nucleic acid sequence 193 and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 193. In some embodiments, the transposterior factor does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposterior factor contains the nucleic acid sequence SEQ ID NO: 140 and at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%). 1) an LTF containing a nucleic acid sequence having sequence identity with 1) the nucleic acid sequence of SEQ ID NO: 168 and an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with SEQ ID NO: 168. In some embodiments, the transposator is associated with a transposase containing the amino acid sequence of SEQ ID NO: 78 or a variant thereof. In some embodiments, the transposase contains an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with SEQ ID NO: 78. In some embodiments, the transposator contains a nucleic acid sequence encoding the transposase. In some embodiments, the transposable element does not contain a nucleic acid sequence encoding a transposase. SEQ ID NO: 78 (Transposase for Mariner-5_XT) MAPKKRHAYEAQFKLQAASYAVVNGNRAAAKEFNINESMVRKWRKQENELRQVKKTKQSFRGNKARWPQLEDQLEQWVIEQRTAGRSVSTVTIRLKATTIAQDLEIE HFQGGPSWCFRFMKRRHLSIRARTTVAQQLPADYKEKMAIFRTYCSNKITDKKIQPNHITNMDEVPLTFDIPVNHTVEIKGTSTVSIRTTGHEKSAFTVVLSCHGNG QKLPPMVIFKRKTLPKEKFPAGVIVKANQKGWMDEEKMREWLREVYVKRPDGFFHTSPSLLICDSMRAHLTATVKKQVKQMNSELAIIPGGLTKELQPLDVGVNRAF KVKLRTAWERWMTDGEHTFTKTGKQRRASYATICEWIVDAWAKVSAITVVRAFAKTGIIAEQPPGNDTGNETDSDNDEREPGMFDGEIAQLFNSDTEDEDFDGFVGED
[0110] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 167; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 179. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CCGTATTTTCTC (SEQ ID NO: 79), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence GAGAAAATACGG (SEQ ID NO: 91), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 79, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 91. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 79, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 91. In some embodiments, the transposable element includes SEQ The transposterior factor further comprises a 5'TSD containing the nucleic acid sequence ID NO: 193 and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 193. In some embodiments, the transposterior factor does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposterior factor comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 167, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 179. In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 103 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 103. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. stomach. SEQ ID NO: 103 (Transposase of Mariner-6_AMi) MPKRKSYTFDFKLSVIDYAKKHGNRAAGREFSVDEKSVREWRKEEDILKTLNPRKRARRGKHCKWPTLEANLKEWVVAQRESNKVVSTIAIRQKARVMANEMKITDFGGGCNWVHKFMHRNNLAVRSRTTVGQKLPDDWEEKLTKFRDFVRKECRVHNLCPADIINMDEVPISFDVPATRTVDQRGKKTIAISTTGHERTNFTVVLACTASGVKLKPMVIFKRVTMPREKIPNGVAVICNKKGWMNEDIMKEWTERCFRTRQGGFFAPKSLLIFDAMAAHKHSSVQKQINKAGAHIAVIPGGLTCKLQPLDISVNHSFKCFMRKEWEKWMSQGVHSFTPSGKQRRATYTEVCNWIVAAWRAVKPSSIINGFFKAGILDGASGEESDYTDSEDYADELDDAAFVDAYFGESSYDTDFEGFPASEKEEAPNWV
[0111] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 168; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 180. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence CCCTTT (SEQ ID NO: 80), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence AAAGGG (SEQ ID NO: 92), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 80, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 92. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 80, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 92. In some embodiments, the transposable element is TTAA (SEQ ID NO: 1) The transposterior factor further includes a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 202) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 202. In some embodiments, the transposterior factor does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposterior factor includes an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 168, and an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 180. In some embodiments, the transposterior factor is SEQ The present invention relates to a transposase comprising the amino acid sequence of ID NO: 104 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 104. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. SEQ ID NO: 104 (transposase for piggyBac-2_XT) MAKRFYSAEEAAAHCMASSSEEFSGSDSEYVPPASESDSSTEESWCSSSTVSALEEPMEVDEDVDDLEDQEAGDRADAAAGGEPAWGPPCNFPPEIPPFTTVPGVKVDTSNFEPINFFQLFMTEAILQDMVLYTNVYAEQYLTQNPL PRYARAHAWHPTDIAEMKRFVGLTLAMGLIKANSLESYWDTTTVLSIPVFSATMSRNRYQLLLRFLHFNNNATAVPPDQPGHDRLHKLRPLIDSLSERFAAVYTPCQNICIDESLLLFKGRLQFRQYIPSKRARYGIKFYKLCESSS GYTSYFLIYEGKDSKLDPPGCPPDLTVSGKIVWELISPLLGQGFHLYVDNFYSSIPLFTALYCLDTPACGTINRNRKGLPRALLDKKLNRGETYALRKNELLAIKFFDKKNVFMLTSIHDESVIREQRVGRPPKNKPLCSKEYSKYM GGVDRTDQLQHYYNATRKTRAWYKKVGIYLIQMALRNSYIVYKAAVPGPKLSYYKYQLQILPALLFGGVEEQTVPEMPPSDNVARLIGKHFIDTLPPTPGKQRPQKGCKVCRKRGIRRDTRYYCPKCPRNPGLCFKPCFEIYHTQLHY
[0112] In some embodiments, the transposable element includes: 1) a 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 169; and 2) a 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 181. In some embodiments, the transposable element includes: 1) a 5' TR containing the nucleic acid sequence CCTTCATACG TTCCCATG (SEQ ID NO: 81), a variant thereof, or a fragment thereof; and 2) a 3' TR containing the nucleic acid sequence CATGAGAACG GATGAGGG (SEQ ID NO: 93), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 81, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 93. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 81, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 93. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 202 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 202. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 169, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 181.In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 105 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 105. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not comprise a nucleic acid sequence encoding the transposase. SEQ ID NO: 105 (transposase of piggyBac-1_AMi) MRRQRIFSAEACLAQVTGDDSDTSTDEILSQGSEDADFVPDMGDCSEQEDHASTEASSSSDEEELVPQAEVMPTLPVGKGRGPARGRRGRPRRATAGRGAEPLPPPPENIADDDTFRGRNGRVWGKSPQLIHHRRTAKDIVHESAGPTRESNKDTV VETFHLFLSQDILTVIIRETNREAERRFREWNDSHPENRKEWKKLDETELKAYIGLLLLSGLYHGGHEALEELWEQRTGRPCFRATMSLKRFRQITNYLRFDNKLTREERRAVDKLAAFRDIWTQFVEQLRKWYIPGIHLTVDEQLIAFRGRCPFRQ YMPSKPAKYGLKIWWCCDAATSFPLSGDIYLGKQPEAPRQGNLGTQVIEKLVHPWYKTGRNIVGDNFFTSTDLAERLLQQDLTYVGTIRKNKADIPNEMQANRRRETFSSIFAFDGELTLVSYVPKKSKVVLLLSSMHHDLSVGGEQRKPDIIHDY NKNKGGVDNLDHLVRMTTCRRKINRWPMTLFFNMLDVAAIASLIIWLGHNANWKQRSHLRRRRFFLQELGEQLIDGQIKRRMMDTRNLSKSVKSGMTALNYKLPQRTPVEPSASMSGRCAFCPRTDDRKTRKRCEDCKRFACVHHSTSLLLCEECNI
[0113] In some embodiments, the transposable element includes: 1) a 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 170; 2) a 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 182. In some embodiments, the transposable element includes: 1) a 5' TR containing the nucleic acid sequence CAGTAGAACC CCG (SEQ ID NO: 82), a variant thereof, or a fragment thereof; and 2) a 3' TR containing the nucleic acid sequence CGGGGTTCTACTG (SEQ ID NO: 94), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes: 1) the nucleic acid sequence SEQ ID NO: 82 and at least about 90% (for example, at least about 1) A 5'TR having a nucleic acid sequence having sequence identity of 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more; and 2) a 3'TR having a nucleic acid sequence having sequence identity of at least approximately 90% (for example, at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the nucleic acid sequence of SEQ ID NO: 94. In some embodiments, the transposable element comprises 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 82; and 2) SEQ ID NO: The transposterior element comprises a 3'TR containing the nucleic acid sequence 94. In some embodiments, the transposterior element further comprises a 5'TSD containing the nucleic acid sequence SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 193. In some embodiments, the transposterior element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 170, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 182. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 106 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 106. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 106 (Transposase of Mariner-3_Crp) MRRPSQMELRAAWVPKQDSGFSLSVAIVFKMSASNTDSTSASASTRKRKAITLEEKLEVVKRYERGEKTKEIRCATGLSESTLRTIRDNAEKIKESCKSATRLSTARVSLTRSAIMEKMERMLTTWIEHQNQEHVPLSTQVIQAKARSIYEVLKRD DPEAKPFNASAGWFDRFKARHGFHNLKLTGEAAAADKSAADKFPSILQAAIEEGSYDPRQVFNLDETGLFWKRMPSRTFISVEERTAPGFKAAKDRCTLLLGANAAGDFKLKPLMVYHAENPRALKGYAKGHLPVYWRANRKGWVTGSLFEDYFSSL LHLELRAYCDRQNLPFKILLLLDNAPGHPPSLGDLSENIKVLFLPPNTTSLIQPMDQGVIRAFKSYYLRRTFKKLINETESKEKANVKEFWKAFNIKMAIDIIGDAWNDVTQQCLNGVWKNIWPAVVNDFRGFAADEFISDARHEIVEMAKSVGFEQ VDEENVAELLDSHREELSNEDLLELDRERHEEEEPMEENERPSPRVLTMKGMADAFKLLDGYLAFFDENDPNRERSAKVARLMKEANACYTELYREKRRCSVQSSLDKFFRKRDSTISTQSTDSQPSTSGGEPATSGGEISPPSFDVFDLDEDSAPH
[0114] In some embodiments, the transposable element includes: 1) the 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 171; 2) the 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 183. In some embodiments, the transposable element includes: 1) CAGTGGTTCT TAACCT (SEQ ID The transposable element comprises: 1) a 5'TR containing the nucleic acid sequence NO: 83), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence AGGTTAAGAA CCACTG (SEQ ID NO: 95), a variant thereof, or a fragment thereof. In some embodiments, the transposable element comprises: 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence SEQ ID NO: 83; and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the nucleic acid sequence SEQ ID NO: 95. In some embodiments, the transposable element comprises 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 83, and 2) SEQ It includes a 3'TR containing the nucleic acid sequence ID NO: 95. In some embodiments, the transposable element contains 5 nucleic acid sequences including CTCTAGAG (SEQ ID NO: 203). 1) The transposable element further comprises a 'TSD' and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 203. In some embodiments, the transposable element does not include a 5'TSD and / or a 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 171, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 183. In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 107 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 107. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposase does not comprise a nucleic acid sequence encoding the transposase. SEQ ID NO: 107 (transposase of hAT-1_PM) MSGKRRKWSDEYVQHGFSCITERDGSERPICMICNAKLSNSSLAPAKLKEHFLKLHGDGEYKNTTLAEFRVKRARFDERATLPALGFVPVNKPILTASYEVAYLIAKQGKPHTIGETLVKPAALKMANLILGTAAEGKLSQIPLSNDTVSDRIEDMSKDILAQVVADLISSPAKFSLQLDETTDVSDLSQLAVFVRYVKDDMIKEEFLFCKPLTTTTKAADVKKLVDDFFRDNNLSWDMVSAVCTDGAPVMLGRKSGFGALVKADAPHIIVTHCILHRHALATKTLPPKMAEVLKTVVECVNYVRKSALRHRIFSELCKEMGSEFEVLLLHSNIRWLSRGQVLNRVFAMRVELALFLQEHQHCHADCFKNPEFILILAYMADIFAALNHLNQQMQGGGVNIMEAEEHLKAFQKKLPLWKRRTENNNFANFPLLDDCASKIEDVSGIGDISVSTELKQTIATHLDELAKSLDGYFPTRESYPAWVRQPFTFAVETADVNDEFLDEIIEIQQSQVQQQLFRTTTLSTFWCRQMNKYPVIAKKALDFFIPFVTTYLCEQSFSRMLDIKTKKRNRLCCENDMRVALAKVKPRISELISQRQQQKSH
[0115] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 172; and 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 184. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence TAGGGCTGTG CGAAA (SEQ ID NO: 84), a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence TTTCGCACAG CCCTA (SEQ ID NO: 96), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 84, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 96. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 84, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 96. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 196 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 96. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 172, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 184. In this configuration, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 108 or a variant thereof. In some embodiments, the transposase is SEQ ID NO: The transposable element includes an amino acid sequence having sequence identity of at least approximately 80% (for example, at least approximately 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of ID NO: 108. In some embodiments, the transposable element includes a nucleic acid sequence encoding a transposase. In some embodiments, the transposable element does not include a nucleic acid sequence encoding a transposase. SEQ ID NO: 108 (transposase of hAT-17_Croc) MTVSSTPISLPMLSLPVPEEELELDLDLSAIAREILGKEGESSEFVLHTPQVSPRARSPSPEGTSEEAEVEVAEASGSAESAPPVAVPLPAEEGEKAASAHAPAPQKRAGSVVWDHFELADDPTYAICLHCRAKTSRGKGTKHMSTTGMLMHLRRHHPLALAPPQPGTSGSVPKGKSPSRSKAPVPPKQRQATLEQWGKGGQKGGRVARASEITQSIGEMLAVDGQPFSLVERPGFRRLMALVAPSYQVPARTTFSRTVVPSLYEACREYLREELRKAGPQVALHFTSDIWSSRGGDHAYLSLTGHWCDQSGRRWALLQAEVMDESHTATEIMGAMNRLVQGWLVGQGELTRGFMVTDNGANMVKAVRDANFVGIRCVAHKFHLIVRDALEGDRAAGDGASTTSKLISKCRKVAGYFHRSIKGGKMLRDKQAELNIPQHKIMQDVETRWNSTYLMLERLVEQQKAIHEMALLGEIGISGPLNRAEWDTISQILVVLKPFLEATETLSAGDALLSQVIPVVRELQSQMESFQGINVPGWGKPLSPDVQALVRRLKEGIRKRLDPLRSSTVHVLAATCDPRVKGSVCGSQSLNQWTEVLVKRVREAERRRRGDVEEGDPLSRASTPSTSQSPPPPQALPMWAKGMASMVGSRGTRPHHQAGSAQALVAAYLAEDVEPLLCDPLAYWASRSQMWPDLATVAREHLSCPPTSVPSERVFSIAGDVVTPHRTSLDPGLVEQLVFLKVNLPLLGFPKLQVQTE
[0116] In some embodiments, the transposable element includes: 1) a 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 173; and 2) a 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 185. In some embodiments, the transposable element includes: 1) a 5' TR containing the nucleic acid sequence CAGGTTGAG (SEQ ID NO: 85), a variant thereof, or a fragment thereof; and 2) a 3' TR containing the nucleic acid sequence CTCAACCTG (SEQ ID NO: 97), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 85, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 97. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 85, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 97. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 173, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 185. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence of SEQ ID NO: 109 or a variant thereof.In some embodiments, the transposase contains an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 109. In some embodiments, the transposase contains a nucleic acid sequence encoding the transposase. Several embodiments. In this state, the above-mentioned transposable element does not contain a nucleic acid sequence that encodes a transposase. SEQ ID NO: 109 (Transposase of Tigger4) MSKRPADTPMGNSDKKKRKHLCLSIAQKVKLLEKLDSGVSVKRLTEEYGVGMTTIYDLKKQKDKLLKFYAESDEQKLMKNRKTLHKAKNEDLDRVLKEWIRQRRSEHMPLNGMLIMKQAKIYHDELKIEGNCEYSTGW LQKFKKRHGIKFLKICGDKASADHEAAEKFIDEFAKVIADENLTPEQVYNADETSLFWRYCPRKTLTTADETAPTGIKDAKDRITVLGCANAAGTHKCKLAVIGKSLRPRCFQGVNFLPVHYYANKKAWITRDIFSDWF HKHFVPAARAHCREAGLDDDCKILLFLDNCSAHPPAEILIKNNVYAMYFPPNVTSLIQPCDQGILRSMKSKYKNTFLNSMLAAVNRGVGVEGFQKEFSMKDAIYAVANAWNTVTKDTVVHAWHNLWPATMFSDDDEQG GDFEGFRMSSEKKMMSDLLTYAKNIPSESVSKLEEVDIEEVFNIDNEAPVVHSLTDGEIAEMVLNQGDRDNSDDEDDVVNTAEKVPIDDMVKMCDGLIEGLEQRAFITEQEIMSVYKIKERLLRQKPLLMRQMTLEETF
[0117] In some embodiments, the transposable element includes: 1) a 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 173; and 2) a 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 185. In some embodiments, the transposable element includes: 1) a 5' TR containing the nucleic acid sequence CAGTAGTCCC CCCTTATCCG CGG (SEQ ID NO: 86), a variant thereof, or a fragment thereof; and 2) a 3' TR containing the nucleic acid sequence CCGCGGATAA GGGGGGACTA CTG (SEQ ID NO: 98), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 86, and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 98. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 86, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 98. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193 and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element is SEQ ID NO: The present invention comprises: 1) an LTF containing a nucleic acid sequence having at least approximately 90% sequence identity (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) of nucleic acid sequence 173; and 2) an RTF containing a nucleic acid sequence having at least approximately 90% sequence identity (e.g., at least approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) of nucleic acid sequence SEQ ID NO: 185. In some embodiments, the transposase is associated with a transposase containing the amino acid sequence SEQ ID NO: 110 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 110. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 110 (Transposase of Tigger7) MAPKRKSSDAGNSDMPKRSRKVLPLSEKVKVLDLIRKEKKSYAEVAKIYGKNESSIREIVKKEKEIRASFAVAPQTAKVTATVRDKCLVKMEKALNLWVEDMNRKRVPIDGNVLRQKALSLYEDFSKGSPETSDTKPFT ASKGWLHRFRNRFGLKNIKITGEAASANEEAAATFPAELKKLIKEKGYHPKQVFNCDETGLFWKKMPNRTYIHKSAKEAPGHKTWKDRLTLVLCGNAAGHMIKPGVVYRAKNPRALKNKNKNYLPVFWQHNQKAWVTAIL FMEWFHQCFIPEVKKYLEEEGLEFKVLLIIDNAPGHPESVCYENENVEVVFLPPNTTSLLQPLDQGIIRFVKATYTRLVFDRIRSAIDADPNLDIMQCWKSFTIADAITFIKAAMDELKPETVNACWKNLWSEVVNDFK GFPGIDGEVRKIIHAARQVGGEGFADMLDEEVEEHIEGHREVLTNEELEELVESSTEEEEDEEETEAEPAMWTLPKFAEVFRIAQTLKDKIMEYDPRMERSIKVTRMITEGLQPLQQHFDELKRKRQQLPITMFFQKISA KNLQLSRIPNHRHRLLLTSNHRHRHGSMIQDHPKQMILLLTYRQKVNSSLTLRHNAYVIHLTSSHHVGIVSSHIITRRRVSTVQ
[0118] In some embodiments, the transposable element includes: 1) a 5'TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 175; 2) a 3'TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 187. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence TAGGG (SEQ ID NO: 87), a variant thereof, or a fragment thereof; and 2) CCCTA (SEQ ID The transposable element comprises: 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 87; and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 99. In some embodiments, the transposable element comprises: 1) a 5'TR having a nucleic acid sequence having SEQ ID NO: 87; and 2) a 3'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 99. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence TTTATAAT (SEQ ID NO: 204) and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 204. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 175, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 187. In some embodiments, the transposase is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 111 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 111. In some embodiments, the transposase comprises a nucleic acid sequence encoding the transposase.In some embodiments, the transposable element does not contain a nucleic acid sequence encoding a transposase. SEQ ID NO: 111 (transposase of hAT-19_CrP) MNAKKSSKTAKYVSALDRARQYPAGTLHADGGRLFCTSCNVTLDVSRKTSIDRHLESEAHMKRKAAAEAEGQSKKQATVSSLFKRTTESSLARREAAFSLVEAFAAANIPLEKLDHPKLRDYLNQNVPNAGSFPRANKLRQDYLP AVIASHVQSTKAALAEYHSFSIVADEATDPQDNYVLHILFVPGSWTDRDLPVFQTEPVHLTSVNFSTVSQAIVKVIVSYGVDFNKISAFLSDNAMYMCKAFSDVLQGMLPNAVHVTCNAHILSLVSDIWRTQFPEVDSLVSSFKKA FKHCPGRKVRYRESVASQSGDSHGTVPLPPEPVLTRWNSWFNAVRYHAKHMQHYQQFVSEELEITPSTQCMLKLRDLLGRGDIQDQVTFVADNATRFMELLTWFEGRQAVIHHAHNKLMDLISWVEDRASQELSSSCPRENQRRRD LFQAIAVKLNQYYRYDPSLPPRFRQPAAPFLSAVRILDPNQVPLLSCDPQLLKAIPGWGKEHDNERAAYVNFVKECQMPVPVPVFWDSVCDRFPHLYALARRCLTIPTNSVDAERAISVYGQVFTPQRQRMSPATAAGCSMLAFNT
[0119] In some embodiments, the transposable element includes: 1) a 5' TR of the LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 176; 2) a 3' TR of the RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 188. In some embodiments, the transposable element includes: 1) a 5' TR containing the nucleic acid sequence TAGG (SEQ ID NO: 88), a variant thereof, or a fragment thereof; and 2) a 3' TR containing the nucleic acid sequence CCTA (SEQ ID NO: 100), a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 88, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 100. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 88, and 2) a 3' TR containing the nucleic acid sequence of SEQ ID NO: 100. In some embodiments, the transposable element further comprises a 5'TSD containing the nucleic acid sequence CTATATAG (SEQ ID NO: 205) and a 3'TSD containing the nucleic acid sequence SEQ ID NO: 205. In some embodiments, the transposable element does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposable element comprises 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 176, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence SEQ ID NO: 188. In some embodiments, the transposer is associated with a transposase comprising the amino acid sequence of SEQ ID NO: 112 or a variant thereof. In some embodiments, the transposase comprises an amino acid sequence having at least about 80% (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) sequence identity with the amino acid sequence of SEQ ID NO: 112. In some embodiments, the transposer comprises a nucleic acid sequence encoding the transposase. In some embodiments, the transposer does not comprise a nucleic acid sequence encoding the transposase. SEQ ID NO: 112(hAT-17B_CrocのTRANSポザーゼ) MTVSSTPISSPSVSLPVPEEELDLEAGDLSALAKEILGTEGESEFVLHTPQSSPRARSPSPEGTSEEVEVVEASGSAESAPPVVVPLPAEEGDKAGSPPSTQKRGGSVVWDHFEVADDPKYAICQHCRRQISRGKEVKHFTTTAMLLLRQHPLALAPPQPGTSGSVPKGKSPSRSKAPVPTKQRQA TLERWGKAADKGGRVARASEVTQSIGMLALDGQPFSLVERPGFRRLMALVAPSYQVPARTTFSRTVVPSLYEACREYLREERLKAGPQVALHFTSDIWSSRGGDHAYLSLTGHWCDQSGRRWALLQAEVMDESHTAREIMVAMNRMVQGWLVGQGELTRGFMVTDNGANMVKAVRDANFVGIRCVAHK LHLIVRDALEGDRAAGDGAVTTTQLISKCRKVAGYFHRSIKGGKMLRDKQAELSVPQHKIMQDVETRWNSTYLMLERLVEQQKAIHEMALLGEIGISGPNLKAEWDTISQILVVLKPFLEATETLSAGDALLSQVIPVVRELENQMEKFQGIDVPGWGKPLSPEVQALVKRLKEGIKKWLDPLRSSTV HVLAGMCDPRVKGSMCTGSTKTLDHWTEVLVNRVREAEGQRRGDVEEGDPLSRASTPSTSQSPPRQPLPMWAKGMASMLGSRRRTRPQPRAGSAEASVATYLAEDVEPLDCDPLAYWATRSQMWQDLATVAREHLSCPPTSVPSERVFSIAGDVVTPHRTSLDPGLVEQLVFLKVNLPLLGFPELHLQTE
[0120] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 177; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 189. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 89, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 101, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 89, and 2) a 3' TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 101. In some embodiments, the transposable element includes 1) a 5' TR containing the nucleic acid sequence of SEQ ID NO: 89, and 2 The transposterior factor includes a 3'TR containing the nucleic acid sequence of SEQ ID NO: 101. In some embodiments, the transposterior factor further includes a 5'TSD containing the nucleic acid sequence of ATTAATAG (SEQ ID NO: 206) and a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 206. In some embodiments, the transposterior factor does not include the 5'TSD and / or the 3'TSD. In some embodiments, the transposator includes 1) an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 177, and 2) an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 189. In some embodiments, the transposator relates to a transposase containing the amino acid sequence of SEQ ID NO: 113 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 113. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 113 (transposase of hAT-10_XT)
[0121] In some embodiments, the transposable element includes: 1) a 5'TR of an LTF, wherein the LTF contains the nucleic acid sequence SEQ ID NO: 178; and 2) a 3'TR of an RTF, wherein the RTF contains the nucleic acid sequence SEQ ID NO: 190. In some embodiments, the transposable element includes: 1) a 5'TR containing the nucleic acid sequence SEQ ID NO: 90, a variant thereof, or a fragment thereof; and 2) a 3'TR containing the nucleic acid sequence SEQ ID NO: 102, a variant thereof, or a fragment thereof. In some embodiments, the transposable element includes 1) a 5'TR having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 90, and 2) a 3'TRv having a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some embodiments, the transposable element includes 1) a 5'TR containing the nucleic acid sequence of SEQ ID NO: 90, and 2) a 3'TR containing the nucleic acid sequence of SEQ ID NO: 102. In some embodiments, the transposable element includes a 5'TSD containing the nucleic acid sequence of SEQ ID NO: 193. The transposase further comprises a 3'TSD containing the nucleic acid sequence of SEQ ID NO: 193. In some embodiments, the transposase does not include a 5'TSD and / or a 3'TSD. In some embodiments, the transposase comprises an LTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 178, and an RTF containing a nucleic acid sequence having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the nucleic acid sequence of SEQ ID NO: 190. In some embodiments, the transposase relates to a transposase containing the amino acid sequence of SEQ ID NO: 114 or a variant thereof. In some embodiments, the transposase includes an amino acid sequence having at least about 80% sequence identity (e.g., at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more) with the amino acid sequence of SEQ ID NO: 114. In some embodiments, the transposator includes a nucleic acid sequence encoding the transposase. In some embodiments, the transposator does not include a nucleic acid sequence encoding the transposase. SEQ ID NO: 90 (MARWOLEN1 5'TR) TACAGTGGCGAGCAAAACTGAGTGCATGTTCACAACTCTCACTTTGGCCATCAATAAAAACACAACCGCTTATGCAAATTCAATAATTTTTTTTATATTATTAATCTTAATTAATTTATAACTTATTTTATGAGAAAAAACAAATTTAGCATATATATTCAAAGACTTAGAATAAAAAAACTAAATAAACTCACATAAACTCAATTTTGCACGTTT SEQ ID NO: 102 (MARWOLEN1 3'TR) AAACGTCAAAATTGAGTTTATGTGAGTTTATTTAGTTTTTTTATTCTAAGTCTTTGAATATATATGCTAAATTTGTTTTTTCTCATAAAATAAGTTATAAATTAATTAAGATTAATAATATAAAAAAAATTATTGAATTTGCATAAGCGGTTGTGTTTTTATTGATGCCAAAGTGAGAGTTGTGAACATGCACTCAGTTTTGCTCGCCACTGTA SEQ ID NO: 114 (Transposase of MARWOLEN1) MRKLSGEIQNNIVSLTESGNSTRQIAERLGISQSAVVSIQKRRKLTPKPQVSGRKKLLKDSDARLMMSEMMRKNKNLTPKGACLAINKNVSEWTARRALQDIGYMSIVKKNKPALSDKNVKARLKFAKDHKNWTIDDWKRVVWSDESKFNRFQSDGKQYCWIRPGDRV QRHHVKQTVKHGGGNIMVWGCFTWWHIGPLQLVEGIMKKEDYLRILQTNLPNYFDKCAYPEKDIIFQQDGDPKHTAKIVKEWIGKQHFQLMEWPAQSPDLNPIENLWSIVKRRLGQYDSAPKNMGDLWERVAVEWSRIPQDILRNLVESMPKRVTEVIVNKGLWTKY
[0122] heterologous nucleic acid The manipulated transposable elements described herein are suitable for the transposition of various heterologous nucleic acids. In some embodiments, the heterologous nucleic acid is DNA. In some embodiments, the heterologous nucleic acid is double-stranded. In some embodiments, the heterologous nucleic acid contains one or more modified nucleotides. In some embodiments, the heterologous nucleic acid is unmodified.
[0123] The heterologous nucleic acid in the transposable element may have a variety of suitable lengths. In some embodiments, the length of the heterologous nucleic acid is at least about 1kb, 2kb, 10kb, 20kb, 30kb, 40kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 150kb, 200kb, 250kb, 300kb, 350kb, 400kb, 450kb, 500kb, 600kb, 700kb, 800kb, 900kb, 1000kb or more. In some embodiments, the length of the heterologous nucleic acid does not exceed any of the following: 1000kb, 900kb, 800kb, 700kb, 600kb, 500kb, 450kb, 400kb, 350kb, 300kb, 250kb, 200kb, 150kb, 100kb, 90kb, 80kb, 70kb, 60kb, 50kb, 40kb, 30kb, 20kb, 10kb, 5kb, 2kb, or 1kb. In some embodiments, the length of the heterologous nucleic acid is approximately 100bp to approximately 1kb, approximately 1kb to approximately 2kb, approximately 2kb to approximately 5kb. b, approximately 5kb to approximately 10kb, approximately 100bp to approximately 5kb, approximately 100bp to approximately 2kb, approximately 2kb to approximately 10kb, approximately 1kb to approximately 10kb, approximately 10kb to approximately 20kb, approximately 20kb to approximately 50kb, approximately 50kb to approximately 100kb, approximately 1kb to approximately 100kb, approximately 150kb to approximately 200kb, approximately 200kb to approximately 300kb, approximately 300kb to approximately 4 The lengths are approximately 00kb, about 400kb to about 500kb, about 500kb to about 600kb, about 600kb to about 700kb, about 700kb to about 80kb, about 800kb to about 900kb, about 900kb to about 1000kb, about 10kb to about 100kb, about 100kb to about 500kb, about 500kb to about 1000kb, or about 10kb to about 500kb. In some embodiments, the length of the heterologous nucleic acid is about 10kb to about 300kb nucleotides. In some embodiments, the length of the heterologous nucleic acid is about 100bp to about 300kb nucleotides.
[0124] The heterogeneous nucleic acid may contain one or more coding sequences, each containing one, two, three, four, five, six, ten, or more coding sequences. In this application, any suitable coding sequence may be used, and the coding sequence can code for any suitable target biological product. In some embodiments, the coding sequence codes for an RNA molecule. In some embodiments, the coding sequence codes for a polypeptide, such as a protein. In some embodiments, the heterogeneous nucleic acid includes a first coding sequence that codes for a first protein and a second coding sequence that codes for a second protein. In some embodiments, the heterogeneous nucleic acid includes a first coding sequence that codes for a first RNA and a second coding sequence that codes for a second RNA. In some embodiments, the heterogeneous nucleic acid includes a first coding sequence that codes for a protein and a second coding sequence that codes for RNA.
[0125] In some embodiments, the coding sequence encodes a therapeutic protein. In some embodiments, the coding sequence encodes a therapeutic antibody, including a monoclonal antibody, a multispecific antibody, and an antibody fragment. In some embodiments, the coding sequence encodes a cytokine. In some embodiments, the coding sequence encodes an antigen. In some embodiments, the coding sequence encodes a therapeutic agent useful for gene therapy.Examples of therapeutic proteins useful in gene therapy include adenosine deaminase, an enzyme affected by lysosomal storage disorders, apolipoprotein E, brain-derived neurotrophic factor (BDNF), bone morphogenetic protein 2 (BMP-2), bone morphogenetic protein 6 (BMP-6), bone morphogenetic protein 7 (BMP-7), cardiomyocyte trophic factor 1 (CT-1), CD22, CD40, ciliary body trophic factor (CNTF), CCL1-CCL28, CXCL1-CXCL17, CXCL1, CXCL2, CX3CL1, vascular endothelial growth factor (VEGF), dopamine, erythropoietin, factor IX, factor VIII, epidermal growth factor (EGF), estrogen, FAS-ligand, fibroblast growth factor 1 (FGF-1), fibroblast growth factor 2 (FGF-2), and fibroblasts. Cell growth factor 4 (FGF-4), fibroblast growth factor 5 (FGF-5), fibroblast growth factor 6 (FGF-6), fibroblast growth factor 7 (FGF-7), fibroblast growth factor 10 (FGF-10), Flt-3, granulocyte colony-stimulating factor (G-CSF), granulocyte-macrophage-stimulating factor (GM-CSF), growth hormone, hepatocyte growth factor (HGF), interferon α (IFN-a), interferon β (IFN-b), interferon γ (IFNg), insulin, glucagon, insulin-like growth factor 1 (IGF-1), insulin-like growth factor 2 (IGF-2), interleukin 1 (IL-1), interleukin 2 (IL-2), interleukin 3 (IL-3), interleukin 4 (IL-4), interleukin 5 (IL-5), Interleukin 6 (IL-6), Interleukin 7 (IL-7), Interleukin 8 (IL-8), Interleukin 9 (IL-9), Interleukin 10 (IL-10), Interleukin 11 (IL-11), Interleukin 12 (IL-12), Interleukin 13 (IL-13), Interleukin 15 (IL-15), Interleukin 17 (IL-17), Interleukin 19 (IL-19), Macrophage Colony Stimulating Factor (M-CS). F) Monocyte chemoattractant protein 1 (MCP-1), macrophage inflammatory protein 3a (MIP-3a), macrophage inflammatory protein 3b (MIP-3b), nerve growth factor (NGF), neurotrophic factor 3 (NT-3), neurotrophic factor 4 (NT-4), parathyroid hormone, platelet-derived growth factor AA (PDGF-AA), platelet-derived growth factor AB (PDGF-AB), platelet-derived growth factor BB (PDGF-BB), platelet-derived growth factor CC (PDGF-CC), platelet-derived growth factor DD (PDGF-DD), RANTES, stem cell factor (SCF), stromal cell This includes, but is not limited to, cell-derived growth factor 1 (SDF-1), transforming growth factor α (TGF-a), transforming growth factor β (TGF-b), tumor necrosis factor α (TNF-a), Wnt1, Wnt2, Wnt2b / 13, Wnt3, Wnt3a, Wnt4, Wnt5a, Wnt5b, Wnt6, Wnt7a, Wnt7b, Wnt7c, Wnt8, Wnt8a, Wnt8b, Wnt8c, Wnt10a, Wnt10b, Wnt11, Wnt14, Wnt15 or Wnt16, Sonic Hedgehog, Desert Hedgehog, and Indian Hedgehog.
[0126] In some embodiments, the above coding sequence encodes an engineered receptor such as a chimeric antigen receptor (CAR) or an engineered T cell receptor (TCR).
[0127] As used herein, “chimeric antigen receptor” or “CAR” means an engineered receptor that specifically implants one or more antigens into a cell such as a T cell. CARs are also called “artificial T cell receptors,” “chimeric T cell receptors,” or “chimeric immune receptors.” In some embodiments, a CAR comprises an extracellular variable domain of an antibody specific to a tumor antigen and an intracellular signaling domain of a T cell or other receptor, such as one or more costimulatory domains. “CAR-T” refers to a T cell that expresses a CAR.
[0128] In some specific embodiments, the above coding sequence encodes a chimeric antigen receptor (CAR). Many chimeric antigen receptors are known to those skilled in the art and can be applied to the present invention. Furthermore, CARs specific to any cell surface marker can be constructed, for example, by utilizing antigen-binding fragments or antibody variable domains of antibody molecules. Any method for generating CARs can be used herein. See, for example, US6,410,319, US7,446,191, US7,514,537, US9765342B2, WO 2002 / 077029, WO2015 / 142675, US2010 / 065818, US 2010 / 025177, US 2007 / 059298 and Berger C. et al., J. Clinical Investigation 118: 1 294-308 (2008), which are incorporated herein by reference.
[0129] As used herein, “T cell receptor” or “TCR” means an endogenous or recombinant T cell receptor comprising an extracellular antigen-binding domain that binds to a specific antigen peptide bound to an MHC molecule. In some embodiments, the TCR comprises a TCRα polypeptide chain and a TCRβ polypeptide chain. In some embodiments, the TCR specifically binds to a tumor antigen. “TCR-T” means a T cell expressing a recombinant TCR. The term “recombinant” means a biomolecule such as a gene or protein that is (1) removed from its naturally occurring environment, (2) unrelated to all or part of the polynucleotide whose gene was found in nature, (3) operably ligated to a polynucleotide unrelated to nature, or (4) not naturally occurring. The term “recombinant” may be used to mean a cloned DNA isolate, a chemically synthesized polynucleotide analog, or a polynucleotide analog biosynthesized by a heterologous lineage, and proteins and / or mRNA encoded by such nucleic acids.
[0130] In some specific embodiments, the above coding sequence encodes an engineered T cell receptor (TCR). In some embodiments, the engineered and modified TCR is specific to a tumor antigen. In some embodiments, the tumor antigen is an intracellular protein of a tumor cell. It is derived from a substance. Many TCRs specific to tumor antigens (including tumor-associated antigens) have been described, including, for example, the NY-ESO-1 cancer-testis antigen, the p53 tumor inhibitor antigen, TCRs of tumor antigens in melanoma (e.g., MARTI, gp100), leukemia (e.g., WT1, minor histocompatibility antigens), and breast cancer (e.g., HER2, NY-BR1). Any TCR known in the art may be used in this application. In some embodiments, the TCR has enhanced affinity for the tumor antigen. Exemplary TCRs and methods for producing TCRs are described, for example, in US5830755 and Kessels et al. Immunotherapy through TCR gene transfer. Nat. Immunol. 2, 957-961 (2001).
[0131] In some embodiments, the above coding sequence encodes a selection marker. A “selection marker” is a gene that is expressed to produce a detectable phenotype and facilitates the detection of host cells having a heterologous nucleic acid encoding the selection marker inserted into a target nucleic acid (e.g., genomic DNA). In some embodiments, the selection marker confers resistance to antibiotics such as puromycin. Other non-limiting examples of selection markers include drug resistance genes and nutritional markers. For example, a selection marker may be a gene that confers resistance to antibiotics selected from the group consisting of ampicillin, kanamycin, erythromycin, chloramphenicol, gentamicin, kasugamycin, rifampicin, spectinomycin, D-cycloserine, nalidixic acid, streptomycin, and tetracycline. Other non-limiting examples of selection markers include adenosyl deaminase, aminoglycoside phosphotransferase, dihydrofolate reductase, hygromycin-B-phosphotransferase, thymidine kinase, and xanthine-guanine phosphoribosyltransferase. Examples of selection markers suitable for mammalian cells include DHFR, thymidine kinase, metallothionein-I and -II, preferably primate metallothionein genes, adenosine deaminase, and ornithine decarboxylase. In some specific embodiments, the heterologous nucleic acids include the coding sequence of a puromycin resistance gene.
[0132] In some embodiments, the above coding sequence is a reporter gene. Since a "reporter gene" is a gene that codes for a detectable product, detection of the reporter gene product can be used to evaluate the function of a target nucleic acid. A reporter gene can be fused with any suitable target nucleic acid (e.g., a promoter, target gene, selection marker, and / or terminal repeat of a transposator) to enable detection of whether the target nucleic acid has been expressed or altered (e.g., cleaved by a transposase) under given conditions. Non-limiting examples of reporter genes include 3-galactosidase, 3-glucuronidase, glutathione-S-transferase (GST), wasabi peroxidase (HRP), luciferase, chloramphenicol acetyltransferase (CAT), secreted alkaline phosphatase (SEAP), green fluorescent protein (GFP, e.g., eGFP), red fluorescent protein (RFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), catechol 2,3-oxygenase (xylE), The invention also includes autofluorescent proteins, including blue fluorescent protein (BFP). In some embodiments, the heterogeneous nucleic acid includes a coding sequence encoding enhanced green fluorescent protein (eGFP). In some embodiments, the coding sequence may encode multiple biological products or a fusion protein. In some embodiments, the heterogeneous nucleic acid of the invention includes a coding sequence encoding a puromycin resistance-enhancing green fluorescent protein (eGFP) fusion protein.
[0133] In some embodiments, the above code sequence encodes a transposase.
[0134] In some embodiments, the above coding sequence encodes a polypeptide useful for genome editing. Genome editing generates specific double-strand breaks (DSBs) at desired locations in the genome. This can be achieved by using nucleases that repair cleavage induced by homologous-directed repair (HDR) (e.g., homologous recombination) or non-homologous end joining (NHEJ) by utilizing the cell's endogenous mechanisms. Any suitable nuclease, including but not limited to CRISPR-related protein (Cas, e.g., Cas9) nucleases, zinc finger nucleases (ZFN, e.g., FokI), transcription activator-like effector nucleases (TALEN, e.g., TALE), homing end nucleases, and their variants, can be introduced into cells to induce genome editing of the target DNA sequence (Shukla et al. (2009) Nature 459: 437-441; Townsend et al (2009) Nature 459: 442-445). In some embodiments, the above coding sequence encodes the Cas9 polypeptide.
[0135] In some embodiments, the above coding sequence encodes an RNA molecule. The RNA molecule may be protein-coding RNA such as messenger RNA (mRNA), or non-protein-coding RNA including but not limited to transporter RNA (tRNA) and ribosomal RNA (rRNA), microRNA (miRNA), small RNA such as small interfering RNA (siRNA), short hairpin RNA (shRNA) or piwi-interacting RNA (piRNA), and long non-coding RNA (lincRNA). Small RNAs such as microRNA and siRNA are important in the process of RNA interference (RNAi). RNAi is a gene regulatory process that interferes with small RNAs by post-transcriptional degradation or inhibition of translation, thereby inhibiting the expression of target genes that are normally expressed. For a detailed description of RNAi techniques, see, for example, US Pat. No. 6,326,527;6,452,067;6,573,099;6,753,139; and 6,777,588. In some embodiments, the above coding sequence encodes a regulatory RN It encodes A. In some embodiments, the above coding sequence encodes an RNAi molecule. In some embodiments, the above coding sequence encodes shRNA. In some embodiments, the above coding sequence encodes miRNA.
[0136] In some embodiments, the above coding sequence can encode RNA molecules useful for genome editing. Examples of such RNA molecules include, but are not limited to, CRISPR RNA (crRNA), transactivating crRNA (tracrRNA), guide RNA (gRNA), and single guide RNA (sgRNA).
[0137] In some embodiments, the heterologous nucleic acid further comprises one or more regulatory elements that modulate the expression of the coding sequence. The regulatory elements are intended to be used in conjunction with the methods and constructs described herein. The term “regulatory element” is intended to include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression regulatory elements (e.g., transcription termination signals such as polyadenylation (poly-A) signals and poly-U sequences). Such regulatory elements are, for example, those described in Goeddel, Gene Expression Technology: This method is described in Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990).
[0138] Promoters are important regulatory elements that guide the expression pattern of coding sequences. In some embodiments, the heterologous nucleic acid includes a promoter operably ligated to the coding sequence. Any suitable promoter can be used in this application. In some embodiments, the promoter is an endogenous promoter. In some embodiments, the promoter is a heterologous promoter. Multiple promoters for gene expression in mammalian cells have been explored, and any promoter known in the art can be used in this application. Promoters can be broadly classified into constitutive promoters or regulatory promoters such as inductive promoters. In some embodiments, the heterologous nucleic acid includes a coding sequence (e.g., a transposase coding sequence) operably ligated to a constitutive promoter. In some embodiments, the heterologous nucleic acid includes a coding sequence (e.g., a transposase coding sequence) operably ligated to an inductive promoter.
[0139] Constitutive promoters enable the constitutive expression of heterologous nucleic acids in host cells. Exemplary constitutive promoters considered herein include, but are not limited to, the cytomegalovirus (CMV) promoter, human elongation factor-1α (hEF1α), ubiquitin C promoter (UbiC), phosphoglycerin kinase promoter (PGK), monkey virus 40 early promoter (SV40), and chicken β-actin promoter conjugated to CMV early enhancer (CAGG). Such constitutive promoters have been extensively compared in numerous studies for their efficiency in driving transgene expression. For example, Michael C. Milone et al. compared the efficiency of CMV, hEF1α, UbiC, and PGK in driving chimeric antigen receptor expression in human primary T cells and concluded that the hEF1α promoter not only induces the highest level of transgenic expression but is also ideally maintained in CD4 and CD8 human T cells (Molecular Therapy, 17(8): 1453-1464 (2009)). In some embodiments, the promoter in the heterologous nucleic acid is the CAG promoter. Figure 3 shows an exemplary engineered transposer containing a heterologous nucleic acid sequence encoding a transposase or selection marker / reporter gene driven by a constitutive promoter. Here, the promoter is either the CMV promoter or the PGK promoter.
[0140] In some applications, and in some gene therapies, it is anticipated that using promoters with moderate or weak expression patterns of coding sequences, instead of strongly expressing promoters (e.g., CMV promoters), may be necessary to avoid or reduce transposase-based autoregulatory events collectively known as overproduction inhibition (OPI).
[0141] Regulatory promoters, such as inductive promoters, enable the expression of heterologous nucleic acids in certain conditions, such as specific developmental stages, tissue types, or subcellular locations. Various types of regulatory promoters are known in this field, including inductive, tissue-specific, cell-type specific, and cell cycle-specific promoters (see, for example, Sambrook and Russell, 2001). Inductive promoters fall into the category of regulatory promoters. Inductive promoters can be induced by one or more conditions, such as physical conditions, the microenvironment of the engineered mammalian cell, the physiological state of the engineered mammalian cell, inducers (i.e., inducing agents), or a combination thereof.
[0142] In some embodiments, expressing a coding sequence using a promoter is limited to a subset of cell types, cell lineages, or tissues, or to specific areas. Desirable only during developmental stages. Examples include the B29 promoter (B cell expression), low transcription factor (CBFa2) promoter (stem cell expression), CD14 promoter (monocyte expression), CD43 promoter (leukocyte and platelet expression), CD45 promoter (hematopoietic cell expression), CD68 promoter (macrophage expression), endothelial glycoprotein promoter (endothelial cell expression), fms-related tyrosine kinase 1 (FLT1) promoter (endothelial cell expression), integrins, α2b (ITGA2B) promoter (megakaryocyte expression), and intracellular adhesion. This includes, but is not limited to, molecule 2 (ICAM-2) promoter (expressed in endothelial cells), interferon-β (IFN-β) promoter (expressed in hematopoietic cells), β-globin LCR (expressed in erythrocytes), globin promoter (expressed in erythrocytes), β-globin promoter (expressed in erythrocytes), α-globin HS40 enhancer (expressed in erythrocytes), ankyrin-1 promoter (expressed in erythrocytes), and Biscot-Aldrich syndrome protein (WASP) promoter (expressed in hemoglobin topology cells).
[0143] Embodiments of the methods described herein may utilize terminator sequences. A terminator sequence comprises a segment of nucleic acid sequence that labels the end of a gene or operon during transcription. This sequence mediates transcription termination by providing a signal to newly synthesized mRNA that triggers the process of releasing mRNA from the transcription complex. These processes This includes direct interaction between the mRNA secondary structure and the complex and / or indirect activity of the recruited termination factor. The release of the transcription complex releases RNA polymerase and associated transcriptional mechanisms, initiating transcription of new mRNA. Terminator sequences include those known in the art. In some embodiments of the present application, the terminator sequence is a polyadenylation (poly-A) signal.
[0144] In some embodiments, the heterologous nucleic acid includes at least one restriction endonuclease recognition site, e.g., a restriction site, which is used as an insertion site for the foreign nucleic acid. Various restriction sites are known in the art, including but not limited to HindIII, PstI, SalI, AccI, HincII, XbaI, BamHI, SmaI, XmaI, KpnI, SacI, EcoRI, and others. In some embodiments, the restriction site is a multicloning site (MCS, also called a multilinker), i.e., a densely arranged series or array of sites recognized by multiple different restriction endonucleases (e.g., those described above). In other embodiments, the heterologous nucleic acid of the present application includes recombinant enzyme recognition sites such as LoxP, FRT, or AttB / AttP sites, recognized by Cre, Flp, and PhiC31 recombinant enzymes, respectively.
[0145] In some embodiments, the heterogeneous nucleic acid includes a tag sequence. The tag sequence may be used to identify a molecule or to provide a site for capturing a molecule, for example, by hybridization.
[0146] In some embodiments, the heterogeneous nucleic acid includes a barcode sequence. “Barcode sequence” means a nucleic acid having a sequence useful for identifying and / or distinguishing one or more first molecules and one or more second molecules bound to this nucleic acid barcode. Nucleic acid barcode sequences are typically very short, for example, about 5 to 20 base pairs in length, and can be bound to one or more target molecules of interest or their amplification products. Nucleic acid barcode sequences may be single-stranded or double-stranded.
[0147] In some embodiments, the heterogeneous nucleic acid includes a unique molecular identifier (UMI). As used herein, the terms “unique molecular identifier” or “UMI” mean a nucleic acid sequence that can be used to identify and / or distinguish one or more first molecules and one or more second molecules to which the UMI is bound. UMIs are generally short, for example, about 5 to 20 bases in length, and can be bound to one or more target molecules of interest or their amplification products. UMIs may be single-stranded or double-stranded. In some embodiments, both nucleic acid barcode sequences and UMIs are incorporated into the nucleic acid target molecule or its amplification product. Typically, UMIs are used to distinguish similar types of molecules within a population or group, while nucleic acid barcode sequences are used to distinguish populations or groups of molecules. In some embodiments where UMIs and nucleic acid barcode sequences are used together, the UMI sequence length is shorter than that of the nucleic acid barcode sequence. In some embodiments where UMIs and nucleic acid barcode sequences are used together, the UMI is incorporated into the target nucleic acid or its amplification product before being incorporated into the nucleic acid barcode sequence. In some embodiments, when UMI and nucleic acid barcode sequences are used together, the UMI is incorporated into the target nucleic acid or its amplified product, and then the nucleic acid barcode sequence is incorporated into the UMI or its amplified product.
[0148] Transfer activity In some embodiments, the transposable elements of the present invention exhibit transposable activity in vitro or intracellularly.
[0149] The transposable activity may be detected by many techniques known to those skilled in the art, such as the excision of the transposable element from the vector, the incorporation of the transposable element into the cell's genome or extrachromosomal DNA, and the reverse An example of measuring the ability of transposases to bind to repeat sequences is, for example, Ivies et al. Cell, 91, 501-510 (1997), WO 98 / 40510 (Hackett et al.), WO 99 / 25817 (Hackett et al. See al., and WO00 / 68399 (Mclvor et al.).
[0150] In some embodiments, transposase measurement is based on the transcomplementarity of two components in a transposator system, where one component comprises a lateral linkage terminal repeat selection marker / reporter gene (donor), and the other component expresses a transposase that identifies and ligates the terminal repeat to perform transposase (helper). For example, Figure 3 shows an exemplary binary construct for screening active transposators. The upper construct is a transposase expression support construct containing, from 5' to 3', a cytomegalovirus (CMV) promoter, a transposase (Tn) gene, and a poly(A) signal. The lower construct is a donor construct containing, from 5' to 3', a phosphoglycerate kinase (PGK) promoter, a 5' target site repeat (TSD) sequence, a 5' terminal repeat (5'TR) sequence, a sequence encoding puromycin and enhanced GFP fusion (Puro-eGFP), a 3' terminal repeat (3'TR) sequence, a 3' target site repeat (TSD) sequence, and a poly(A) signal. In transfection testing, co-transfection of cultured mammalian cells (e.g., human 293T, HeLa, or Hct116 cells) with a donor plasmid and an auxiliary plasmid or control plasmid is performed. The number of cell clones expressing puromycin resistance due to chromosomal integration and puromycin resistance gene expression is used as an indicator of gene transfer efficiency and is represented by methylene blue staining. For suspension cells such as K562 and primary T cells, transfection activity can be evaluated from GFP reporter gene-positive cells after electroporation. Figures 4A-4B show the colony counts from methylene blue staining results based on transfection testing, and indicate the transfection efficiency of TE evaluated compared to the control. Transfection efficiency in such measurements is also called "transfer efficiency."
[0151] In some embodiments, the transduction efficiency of the manipulated transposable element is at least about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or higher. In some embodiments, the transduction efficiency of the manipulated transposable element is measured in human cells such as human 293T, HeLa, Hct116, K562, or primary T cells.
[0152] In some embodiments, the transposability activity of the manipulated transposers is higher than that of the piggyBac(PB) transposon, Sleeping Beauty(SB) transposon, and / or TcBuster(TB) transposon. In some embodiments, the transposability activity of the manipulated transposers is higher than that of the piggyBac(PB) transposon, for example, by about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 2x, 3x, 5x, 10x or more, as determined in mammalian cells by reporter gene-based transposability measurements (as described in Example 2, for example). In some embodiments, the transposability activity of the manipulated transpossessor is, for example, about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 2x, 3x, 5x, 10x, or more higher than the transposability activity of the Sleeping Beauty (SB) transposon, as determined in mammalian cells by a reporter gene-based transposability measurement (e.g., as described in Example 2). In some embodiments, the transposability activity of the manipulated transpossessor is, for example, at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 2x, 3x, 5x, 10x, or more higher than the transposability activity of the TcBuster (TB) transposon, as determined in mammalian cells by a reporter gene-based transposability measurement (e.g., as described in Example 2). In some embodiments, the transposability activity of the manipulated transposers is higher than that of PB transposons and SB transposons. In some embodiments, the transposability activity of the manipulated transposers is higher than that of PB transposons and TB transposons. In some embodiments, the transposability activity of the manipulated transposers is higher than that of TB transposons and SB transposons. In some embodiments, the transposability activity of the manipulated transposers is higher than that of PB transposons, SB transposons, and TB transposons.
[0153] In some embodiments, the transposability activity of the manipulated transposable element is evaluated in mammalian cells. In some embodiments, the mammalian cells are HeLa cells. In some embodiments, the mammalian cells are human embryonic kidney 293T (293T). In some embodiments, the mammalian cells are K562 cells. In some embodiments, the mammalian cells are Hct116 cells. In some embodiments, the mammalian cells are human T cells, e.g., donor-derived primary T cells. In some embodiments, the manipulated transposable element has higher transposability activity in 293T cells than in HeLa cells, for example, the transposability activity in 293T cells is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 2x, 3x, 5x, 10x or more higher than the transposability activity in HeLa cells.
[0154] cell The manipulated transposable elements described herein have transposable activity in various cells and can be used to insert heterologous nucleic acids into target nucleic acids of any suitable cell.
[0155] In some embodiments, the cells are isolated cells. In some embodiments, the cells are in a cell culture. In some embodiments, the cells are ex vivo. In some embodiments, the cells are obtained from a living organism and maintained in a cell culture. In some embodiments, the cells are single-celled organisms. Cells can be classified into different types based on their origin, tissue of origin, morphology, function, histological markers, expression profile, etc.
[0156] In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are bacterial cells or derived from bacterial cells. In some embodiments, the cells are archaeal cells or derived from archaeal cells. In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are plant cells or derived from plant cells. In some embodiments, the cells are fungal cells or derived from fungal cells. In some embodiments, the cells are animal cells or derived from animal cells. In some embodiments, the cells are invertebrate cells or derived from invertebrate cells. In some embodiments, the cells are vertebrate cells or derived from vertebrate cells. In some embodiments, the cells are mammalian cells or derived from mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the cells are zebrafish (D. rerio) cells. In some embodiments, the cells are rodent cells. In some embodiments, the cells are produced by synthetic means and may be called artificial cells. In some embodiments, the cells are related to animal species derived from 5'TR, 3'TR, and / or transposases. In some embodiments, the cells are independent of the animal species from which the 5'TR, 3'TR, and / or transposases originate.
[0157] In some embodiments, the cells are bacterial, yeast, fungal, algal, plant, or animal cells. In some embodiments, the cells are isolated from a natural source, such as a tissue biopsy. In some embodiments, the cells are isolated from a cell line cultured in vitro. In some embodiments, the cells are genetically modified cells. In some embodiments, the cells are proliferating, differentiating, or otherwise in the nucleus. These are seed cells that experience both.
[0158] In some embodiments, the cells are animal cells derived from organisms selected from cattle, sheep, goats, horses, pigs, deer, chickens, ducks, geese, rabbits, and fish.
[0159] In some embodiments, the cells are plant cells of biological origin selected from corn, wheat, barley, oats, rice, soybeans, oil palm, safflower, sesame, tobacco, flax, cotton, sunflower, pearl millet, foxtail millet, sorghum, rapeseed, cannabis, vegetable crops, fodder crops, commercial crops, woody crops, and biomass crops.
[0160] In some embodiments, the cells are mammalian cells, including cells derived from humans, domesticated and farm animals, zoos, exercise areas, or pets such as dogs (C. familiaris), horses, cats, and cows. In some embodiments, the cells are human cells. In some embodiments, the human cells are human embryonic kidney 293T (HEK293T or 293T) cells or HeLa cells.
[0161] In some embodiments, the cells are derived from primary cells. For example, cultures of primary cells can be passaged 0, 1, 2, 4, 5, 10, 15 or more times. In some embodiments, primary cells are obtained from an organism by any known method. For example, leukocytes can be collected by blood component replacement, leukocyte separation, density gradient separation, etc. Cells derived from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, and stomach can be collected by biopsy. A suitable solution can be used to disperse or suspend the collected cells. Such solutions are generally equilibrium salt solutions (e.g., physiological saline, phosphate-buffered salt (PBS), Hanks equilibrium salt solution, etc.) and can be easily supplemented with calf serum or other naturally occurring factors along with an acceptable low concentration buffer. The buffer may include HEPES, phosphate buffer, lactate buffer, etc. The cells can be used immediately or stored (e.g., by freezing). Frozen cells are thawable and reusable. Cells can be frozen in DMSO, serum, culture medium buffers (e.g., 10% DMSO, 50% serum, 40% buffered medium), and / or other common solutions used to preserve cells at freezing temperatures.
[0162] In some embodiments, the cells described above are derived from cell lines. Various cell lines are known in the art. Examples of cell lines include, but are not limited to, 293T, MF7, K562, HeLa, and their transgenic varieties. Cell lines can be obtained from various sources known to those skilled in the art (see, for example, the American Center for Typical Cell Cultures Depository (ATCC) (Manassus, Va.)).
[0163] In some embodiments, the cells include adherent cells. In some embodiments, the cells include differentiated adherent cells. In some embodiments, the cells include undifferentiated adherent cells. In some embodiments, the cells include pluripotent stem cells. In some embodiments, the cells include non-adherent cells.
[0164] In some embodiments, the cells are derived from epithelial tissue, muscle tissue, nerve tissue, connective tissue, or any combination thereof. In some embodiments, the cells are derived from tissues selected from the liver, gastrointestinal tract, pancreas, kidney, lung, trachea, blood vessels, skeletal muscle, heart, skin, smooth muscle, connective tissue, cornea, urogenital system, mammary gland, reproductive organs, endothelium, epithelium, fibroblasts, nerves, Schwann cells, fat, bone, bone marrow, cartilage, pericytes, mesothelium, endocrine, matrix, lymph, blood, endoderm, ectoderm, mesoderm, and combinations thereof. In some embodiments, the cells are derived from connective tissue (e.g., loose connective tissue, dense connective tissue, elastic tissue, reticular connective tissue, and adipose tissue), muscle tissue (e.g., skeletal muscle, smooth muscle, and The cells are derived from tissues selected from myocardium, urogenital tissue, gastrointestinal tissue, lung tissue, bone tissue, nerve tissue, and epithelial tissue (e.g., simple epithelium and multilayer epithelium), endoderm-derived tissue, mesoderm-derived tissue, and ectoderm-derived tissue, or any combination thereof. In some embodiments, the cells are tumor-derived.
[0165] In some embodiments, the cells are selected from liver cells, gastrointestinal cells, pancreatic cells, kidney cells, lung cells, tracheal cells, vascular cells, skeletal muscle cells, cardiomyocytes, skin cells, smooth muscle cells, connective tissue cells, corneal cells, urinary germ cells, mammary gland cells, germ cells, endothelial cells, epithelial cells, fibroblasts, nerve cells, Schwann cells, adipocytes, osteocytes, bone marrow cells, chondrocytes, pericytes, mesothelial cells, cells derived from endocrine tissue, stromal cells, stem cells, progenitor cells, lymphocytes, blood cells, cells derived from endoderm, cells derived from ectoderm, cells derived from mesoderm, undifferentiated cells (e.g., stem cells and progenitor cells), tumor cells, iPS cells, and combinations thereof.
[0166] In some embodiments, the cells are immune cells such as T cells, B cells, natural killer (NK) cells, dendritic cells (DCs), and macrophages. In some embodiments, the cells are human T cells obtained from a patient or donor. In some embodiments, the cells are immune cells selected from cytotoxic T cells, helper T cells, natural killer (NK) T cells, iNK-T cells, NK-T-like cells, αβ-T cells, αγδ-T cells, tumor-infiltrating T cells, and dendritic cell (DC)-activated T cells. In some embodiments, the cells are immune cells modified using the manipulated transposable elements or gene transfer systems of the present invention. In some embodiments, the modified immune cells are CAR-T cells. In some embodiments, the modified immune cells are TCR-T cells.
[0167] In some embodiments, the cells of the present application are mammalian cells. In some embodiments, the mammalian cells are human HKT293 cells or HeLa cells. In some further embodiments, the transposability activity of the transposable element in 293T cells is higher than that in HeLa cells. In some embodiments, the mammalian cells are selected from immune cells, liver cells, tumor cells, stem cells, fertilized eggs, muscle cells, and skin cells.
[0168] In some embodiments, the cells are stem cells or progenitor cells. The cells may include stem cells (e.g., adult stem cells, embryonic stem cells, iPS cells) and progenitor cells (e.g., cardiac progenitor cells, neural progenitor cells, etc.). The cells may also include mammalian stem cells and progenitor cells, such as rodent stem cells, rodent progenitor cells, human stem cells, and human progenitor cells.
[0169] In some embodiments, the cells are diseased cells. Diseased cells may have altered metabolism, gene expression, and / or morphological characteristics. Diseased cells may be cancer cells, diabetic cells, and apoptotic cells. Diseased cells may be cells derived from the diseased organism.
[0170] In some embodiments, the cells of the present invention belong to a target cell type useful for gene therapy. Examples of target cell types include hematopoietic stem cells, hematopoietic progenitor cells, myeloid progenitor cells, lymphocyte progenitor cells, platelet-producing progenitor cells, erythrocyte progenitor cells, granulocyte-producing progenitor cells, monocyte progenitor cells, megakaryocytes, megakaryoblasts, megakaryocytes, blood coagulation cells / platelets, proerythroblasts, basophilic erythroblasts, polychromatic erythroblasts, orthochromatic erythroblasts, polychromatic erythrocytes, erythrocytes (erythrocytes or RBCs), basophilic promyelocytes, basophilic myeloid cells, basophilic metamyelocytes, basophils, neutral promyelocytes, neutral metamyelocytes, neutral metamyelocytes, neutrophils, eosinophilic promyelocytes, eosinophilic myeloid cells, macrophages, dendritic cells, lymphoblasts, immature lymphocytes, natural killer (NK) cells, small lymphocytes, T lymphocytes, B lymphocytes, plasma cells, and lymphoidendritic cells. In preferred embodiments, the target cell type is one or more erythrocytes, such as proerythroblasts, basophilic erythrocytes, polystained erythrocytes, normostained erythrocytes, polystained erythrocytes, and red blood cells (RBCs). These include chromatic erythroblasts, normochromatic erythroblasts, polychromatic erythrocytes, and red blood cells (RBCs).
[0171] vector In some embodiments, the nucleic acid sequences encoding the manipulated transposase and / or transposase are present in one or more vectors.
[0172] In this application, a variety of suitable vectors can be used. In some embodiments, the vectors are plasmid vectors, cosmid vectors, artificial chromosomes (e.g., bacterial artificial chromosomes, yeast artificial chromosomes, or mammalian artificial chromosomes), phages, baculoviruses, retroviruses, lentiviruses, adenoviruses, cowpox viruses, semliki forest virus, adeno-associated virus (AAV) vectors, and other viral vectors, all of which are publicly known and commercially available. Typically, a suitable vector includes a replication origin that functions in at least one organism, a promoter sequence, a convenient restriction endonuclease site, and one or more selection markers.
[0173] In some embodiments, the vector is a plasmid. In some embodiments, the plasmid may be transformed into bacteria for storage or amplification, or transfected into mammalian cells.
[0174] Methods for introducing this vector into mammalian cells are known in the field. This vector can be introduced into host cells by physical, chemical, and / or biological methods. In this application, it is anticipated that various vector types and vector delivery methods can be used individually or in combination.
[0175] Physical methods for introducing vectors into host cells include calcium phosphate precipitation, lipid transfection, particle impact, microinjection, and electroporation. Methods for producing cells containing vectors and / or exogenous nucleic acids are known in the art. See, for example, Sambrook et al. (2001) Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York. In some embodiments, vectors are introduced into cells by electroporation.
[0176] Biological methods for introducing heterologous nucleic acids into host cells include the use of DNA and RNA vectors. Viral vectors have become the most widely used method for inserting genes into mammalian cells (e.g., human cells).
[0177] Chemical methods for introducing vectors into host cells include polymer complexes, nanocapsules, microspheres, beads, and colloidal dispersion systems such as lipid systems including oil-in-water emulsions, micelles, mixed micelles, and liposomes. An example of a colloidal system used as an in vitro delivery vector is liposomes (e.g., artificial membrane vesicles).
[0178] In some embodiments, the vector is a viral vector. Examples of viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus (AAV) vectors, lentiviral vectors, retroviral vectors, cowpox vectors, herpes simplex virus vectors, and their derivatives. Viral vector technology is publicly known in the art, for example, Sambrook et al. (2001, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York), and other viruses. It is described in the manuals of science and molecular biology. In some embodiments, viral vector delivery transposable elements can be used in gene therapy that combines the efficient gene delivery of viral vectors with the stability of gene expression achieved by transposable elements. For example, Yant, Stephen R., et al. “Transposition from a gutless adeno-transpos See "on vector stabilizes transgene expression in vivo." Nature biotechnology 20.10 (2002): 999-1005.
[0179] III. Gene Transfer Systems Another aspect of the present application provides a gene transfer system comprising 1) an engineered transposable element (e.g., any of the transposable elements described herein) and 2) a transposase, or a nucleic acid encoding a transposase.
[0180] In some embodiments, a gene transfer system is provided comprising: 1) an engineered transposator comprising a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3' repeat sequence (3'TR), wherein the 5'TR comprises a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90, a variant thereof, or a fragment thereof, and the 3'TR comprises a nucleic acid sequence selected from SEQ ID NO: 27-52 and 91-102, a variant thereof, or a fragment thereof; and 2) a transposase. In some embodiments, the transposator exhibits transposability activity that enables insertion of heterogeneous nucleic acids into the DNA of cells (e.g., mammalian cells or plant cells). In some embodiments, the engineered transposator is derived from one of the TEs in Table 2. In some embodiments, the engineered transposator is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0181] In some embodiments, a gene transfer system is provided comprising: 1) an engineered transposer comprising a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3' repeat sequence (3'TR), wherein the 5'TR comprises a nucleic acid sequence selected from SEQ ID NO: 1-26 and 79-90, a variant thereof, or a fragment thereof, and the 3'TR comprises a nucleic acid sequence selected from SEQ ID NO: 27-52 and 91-102, a variant thereof, or a fragment thereof; and 2) a nucleic acid (e.g., DNA or RNA) encoding a transposase. In some embodiments, the transposer exhibits transposability activity that enables insertion of heterogeneous nucleic acid into the DNA of cells (e.g., mammalian cells or plant cells). In some embodiments, the engineered transposer is derived from one of the TEs in Table 2. In some embodiments, the engineered transposer is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0182] In some embodiments, 1) a manipulated transposable element comprising a 5'-terminal repeat sequence (5'TR), a heteronucleotide, and a 3'-terminal repeat sequence (3'TR) from 5' to 3', and 2) SEQ ID The present invention provides a gene transfer system comprising a transposase containing an amino acid sequence or a variant thereof selected from NO: 53-78 and 103-114. In some embodiments, the transposase exhibits transposase activity that enables insertion of heterologous nucleic acids into the DNA of cells (e.g., mammalian cells or plant cells). In some embodiments, the manipulated transposase is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposase is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0183] In some embodiments, 1) a manipulated transposable element comprising a 5'-terminal repeat sequence (5'TR), a heteronucleotide, and a 3'-terminal repeat sequence (3'TR) from 5' to 3', and 2) SEQ ID The present invention provides a gene transfer system comprising a nucleic acid (e.g., DNA or RNA) encoding a transposase containing an amino acid sequence or variant thereof selected from NO: 53-78 and 103-114. In some embodiments, the transposase exhibits transposase activity that enables insertion of heterologous nucleic acids into the DNA of cells (e.g., mammalian cells or plant cells). In some embodiments, the manipulated transposase is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposase is Tc1-8B_DR, Tc1-3_FR, Derived from Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0184] In some embodiments, the gene transfer system comprises 1) an engineered transposator and 2) a transposase, wherein the transposator comprises a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3' repeat sequence (3'TR), the 5'TR comprising a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NOs: 1-26 and 79-90, the 3'TR comprising a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NOs: 27-52 and 91-102, and the transposase comprising an amino acid sequence or a variant thereof selected from SEQ ID NOs: 53-78 and 103-114. In some embodiments, the transposator exhibits transposability activity that enables insertion of heterogeneous nucleic acids into the DNA of cells (e.g., mammalian cells or plant cells). In some embodiments, the engineered transposator is derived from any one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0185] In some embodiments, the gene transfer system comprises 1) an engineered transposator and 2) a nucleic acid (e.g., DNA or RNA) encoding a transposase, wherein the transposator comprises a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3' repeat sequence (3'TR), the 5'TR comprising a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NOs: 1-26 and 79-90, the 3'TR comprising a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NOs: 27-52 and 91-102, and the transposase comprising an amino acid sequence or a variant thereof selected from SEQ ID NOs: 53-78 and 103-114. In some embodiments, the transposator exhibits transposability activity that enables insertion of heterogeneous nucleic acid into the DNA of cells (e.g., mammalian cells or plant cells). In some embodiments, the engineered transposator is derived from any one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0186] In some embodiments, a gene transfer system is provided comprising: 1) an engineered transposable element comprising a 5'-terminal repeat sequence (5'TR), a heterologous nucleic acid, and a 3'-terminal repeat sequence (3'TR) from 5' to 3', wherein the 5'TR comprises a nucleic acid sequence with SEQ ID NO: 3, a variant thereof, or a fragment thereof, and the 3'TR comprises a nucleic acid sequence with SEQ ID NO: 29, a variant thereof, or a fragment thereof; and 2) a transposase comprising an amino acid sequence with SEQ ID NO: 55, or a nucleic acid encoding a transposase.
[0187] In some embodiments, a gene transfer system is provided comprising: 1) an engineered transposable element comprising a 5'-terminal repeat sequence (5'TR), a heterologous nucleic acid, and a 3'-terminal repeat sequence (3'TR) from 5' to 3', wherein the 5'TR comprises a nucleic acid sequence with SEQ ID NO: 8, a variant thereof, or a fragment thereof, and the 3'TR comprises a nucleic acid sequence with SEQ ID NO: 34, a variant thereof, or a fragment thereof; and 2) a transposase comprising an amino acid sequence with SEQ ID NO: 60, or a nucleic acid encoding a transposase.
[0188] In some embodiments, 1) a 5'-terminal repeat sequence (5'TR), a heteronucleotide, and a 3'-terminal repeat sequence (3'TR) are included, where the 5'TR comprises the nucleic acid sequence of SEQ ID NO: 11, its variant or fragment, and the 3'TR comprises SEQ ID NO: The present invention provides a gene transfer system comprising: 1) a manipulated transposer containing 37 nucleic acid sequences, their variants, or fragments thereof; and 2) a transposase containing the amino acid sequence SEQ ID NO: 63, or a nucleic acid encoding the transposase.
[0189] In some embodiments, 1) a 5'-terminal repeat sequence (5'TR), a heteronucleotide, and a 3'-terminal repeat sequence (3'TR) are included, where the 5'TR comprises a nucleic acid sequence with SEQ ID NO: 12, a variant thereof, or a fragment thereof, and the 3'TR comprises a SEQ ID NO: The present invention provides a gene transfer system comprising: 1) an engineered transposer containing 38 nucleic acid sequences, their variants, or fragments thereof; and 2) a transposase containing an amino acid sequence with SEQ ID NO: 64, or a nucleic acid encoding a transposase.
[0190] In some embodiments, 1) a 5'-to-3' repeat sequence (5'TR), a heteronucleotide, and a 3'-to-3' repeat sequence (3'TR) are included, where the 5'TR comprises the nucleic acid sequence of SEQ ID NO: 13, its variant or fragment, and the 3'TR comprises SEQ ID NO: The present invention provides a gene transfer system comprising: 1) a manipulated transposer containing 39 nucleic acid sequences, their variants, or fragments thereof; and 2) a transposase containing the amino acid sequence SEQ ID NO: 65, or a nucleic acid encoding the transposase.
[0191] In some embodiments, 1) from 5' to 3', the 5' terminal repeat sequence (5'TR), a heteronucleotide, and the 3' terminal repeat sequence (3'TR) are included, where the 5'TR comprises the nucleic acid sequence of SEQ ID NO: 16, its variant or fragment, and the 3'TR comprises SEQ ID NO: The present invention provides a gene transfer system comprising: 1) a manipulated transposer containing 42 nucleic acid sequences, their variants, or fragments thereof; and 2) a transposase containing the amino acid sequence SEQ ID NO: 68, or a nucleic acid encoding the transposase.
[0192] In some embodiments, 1) a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3'-to-3' repeat sequence (3'TR) are included, where the 5'TR comprises the nucleic acid sequence of SEQ ID NO: 22, its variant or fragment, and the 3'TR comprises SEQ ID NO: The present invention provides a gene transfer system comprising 1) a manipulated transposer containing 48 nucleic acid sequences, their variants, or fragments thereof, and 2) a transposase containing an amino acid sequence with SEQ ID NO: 74, or a nucleic acid encoding a transposase.
[0193] In some embodiments, 1) from 5' to 3', the 5' terminal repeat sequence (5'TR), heteronucleotide and 3' terminal repeat sequence (3'TR) are included, where the 5'TR comprises the nucleic acid sequence of SEQ ID NO: 23, its variant or fragment, and the 3'TR comprises SEQ ID NO: The present invention provides a gene transfer system comprising: 1) a manipulated transposer containing 49 nucleic acid sequences, their variants, or fragments thereof; and 2) a transposase containing the amino acid sequence of SEQ ID NO: 75, or a nucleic acid encoding the transposase.
[0194] In some embodiments, 1) from 5' to 3', the 5' terminal repeat sequence (5'TR), heterogeneous nucleic acid and 3' terminal repeat sequence (3'TR) are included, where the 5'TR comprises the nucleic acid sequence of SEQ ID NO: 79, its variant or fragment, and the 3'TR comprises SEQ ID NO: The present invention provides a gene transfer system comprising: 1) a manipulated transposer containing 91 nucleic acid sequences, their variants, or fragments; and 2) a transposase containing the amino acid sequence SEQ ID NO: 103, or a nucleic acid encoding the transposase.
[0195] In some embodiments, 1) from 5' to 3', the 5' terminal repeat sequence (5'TR), heterogeneous nucleic acid and 3' terminal repeat sequence (3'TR) are included, where the 5'TR comprises the nucleic acid sequence of SEQ ID NO: 82, its variant or fragment, and the 3'TR comprises SEQ ID NO: The present invention provides a gene transfer system comprising: 1) a manipulated transposer containing 94 nucleic acid sequences, their variants, or fragments thereof; and 2) a transposase containing an amino acid sequence with SEQ ID NO: 106, or a nucleic acid encoding a transposase.
[0196] The gene transfer system described herein includes a transposase. The transposase is a transposase. The transposase may exist as a lipeptide. Alternatively, the transposase may exist in the form of a polynucleotide containing a coding sequence that encodes the transposase. The polynucleotide may be RNA, for example, mRNA encoding the transposase, or DNA, for example, a coding sequence that encodes the transposase. When the transposase exists as a coding sequence that encodes the transposase, in some embodiments, the coding sequence may be present in the same vector containing the transposase factor, i.e., cis. In some embodiments, the gene transfer system comprises a first vector containing the manipulated transposase factor and a second vector containing the transposase coding sequence, i.e., trans.
[0197] In some embodiments, the gene transfer system comprises 1) a vector containing an engineered transposable element (any of the engineered transposable elements described herein) and 2) a transposase.
[0198] In some embodiments, the gene transfer system comprises 1) an engineered transposase (any of the engineered transposases described herein) and 2) a nucleic acid encoding a transposase, wherein the nucleic acid is DNA. In some embodiments, the engineered transposase and the nucleic acid reside in the same vector. In some embodiments, the engineered transposase and the nucleic acid reside in separate vectors.
[0199] In some embodiments, the gene transfer system comprises 1) an engineered transposase (any of the engineered transposases described herein) and 2) a nucleic acid encoding a transposase, wherein the nucleic acid is RNA. In some embodiments, the engineered transposase and the nucleic acid reside in the same vector. In some embodiments, the engineered transposase and the nucleic acid reside in separate vectors.
[0200] In the gene transfer system of the present specification, there are many potential and appropriate combinations of intracellular delivery methods for the engineered transposon and the transposase or nucleic acid encoding the transposase. The transposase can be delivered as DNA, RNA, or protein. The engineered transposon and the transposase can be delivered together or separately. For example, the transposon gene and the transposase gene can be included together in the same recombinant viral genome, and the two parts of a single infection delivery system are such that the expression of the transposase instructs the cleavage of the transposon from the recombinant viral genome for subsequent integration into the cell chromosome. In another example, the transposase and the transposon can be delivered separately by a non-viral combination such as a viral and / or lipid-containing reagent. In these cases, the transposon and / or transposase gene can be delivered via a recombinant virus. In some embodiments, the transposase is instructed to release the transposon from its donor DNA for integration at the target location.
[0201] The transposase can be provided to the cell as a protein or a nucleic acid encoding the transposase protein. The nucleic acid encoding the transposase protein can be in the form of DNA or RNA. The protein can be introduced into the cell alone or introduced into a vector such as a plasmid or viral vector. Further, the nucleic acid encoding the transposase protein can be stably or transiently integrated into the genome of the cell to promote transient or extended expression of the transposase protein intracellularly. Further, a promoter or other expression control region can be operably linked to the nucleic acid encoding the transposase protein to quantitatively or tissue-specifically regulate the expression of this protein. In some embodiments, the transposase protein includes a DNA binding domain, a catalytic domain (having transposase activity), and / or a nuclear localization signal (NLS).
[0202] Therefore, various methods and materials can be used to deliver the gene transfer system of the present application into cells.
[0203] For example, one method is to deliver the gene transfer system using a plasmid vector. For example, this system can be composed of two types of plasmids, a helper plasmid carrying a transposase expression cassette and a donor plasmid carrying a transposable element. After transfection, both plasmids are directed to the cell nucleus, enabling the production of transposase-encoding RNA from the helper plasmid, and then excision of the transposable element from the donor plasmid facilitated by the transposase subunit introduced into the nucleus. Such a method may be further improved by arranging the transposase gene and the transposable element on a single plasmid, which was originally called a helper-independent transposable element-transposase vector. Alternatively, transfected in vitro transcribed mRNA can serve as a rich source of transposase, thereby eliminating the risk of producing cells with extended transposase expression.
[0204] Therefore, in some embodiments, the above transposable element is present in a first vector (donor vector), and the nucleic acid encoding the transposase is present in a second vector (helper vector). In some embodiments, the first vector and the second vector are used to co-transfect cells for transfer.
[0205] Another method involves delivering gene transfer systems using viral vectors. While viral vectors may have immunogenic or carcinogenic issues, high delivery efficiency is sometimes required in certain applications. The components of the gene transfer system may be supported and transmitted by a viral capsid, thereby conferring the ability to incorporate genes into other additional vectors, such as adenovirus or herpes simplex virus vectors, and establish long-term transgenic expression. The viral capsid can provide vector stability, tissue-specific transposable element delivery, and intermembrane transport, and the transposable elements facilitate the incorporation of the viral vector according to the characteristic incorporation spectrum of the transposable elements. Viral vector-based delivery methods and techniques are known in the field. For example, viral vector-based transposition of the Sleeping Beauty transposon system was first confirmed in the liver of mice carrying adenovirus vectors, and recent studies have demonstrated the applicability of this method to larger animals. Adeno-associated virus vectors are also suitable as vectors for the Sleeping Beauty system. For example, Yant, Stephen R., et al. “Transposition from a gutless adeno-transposon stabilize vectors transgene expression in vivo.” Nature biotechnology 20.10 (2002): 999-1005,Hausl, Martin A., et al. “Hyperactive sleeping beauty transposase enables persistent phenotypic correction in mice and a canine model for hemophilia B.” Molecular Therapy 18.11 (2010): 1896-1906, and Zhang, Wenli, et al. “Hybrid adeno-associated See "Viral vectors utilizing transposase-mediated somatic integration for stable transgene expression in human cells." PloS one 8.10 (2013).
[0206] As described above, the gene delivery system of this application can be delivered to host cells by physical, chemical, biological, or a combination thereof. Various delivery vectors and methods are envisioned to be used in this application alone or in combination to achieve desired results for several applications. Therefore, those skilled in the art can most suitably utilize various delivery methods and techniques with various modifications to suit specific envisioned applications such as gene therapy. In order to ensure the successful implementation of gene therapy, stable integration of therapeutic transgenic genes into the genome of affected tissue should be achieved to provide long-term and cost-effective treatment. For example, efficient and low immunogenic gene delivery into the patient's body To achieve this, synthetic compounds and plasmids can be used in combination to deliver DNA to cells. Liposomes and other nanoparticles are sufficient to accomplish this task. For example, two plasmids can be delivered to a patient: one providing the expression of a transposase (helper plasmid) and the other providing a transposable element containing a therapeutic transgenic gene (donor plasmid). These DNAs may be complexed with liposomes and administered by extra-gastrointestinal injection. Upon entering the cell, the transposase can bind to the transposable element of the donor plasmid, excise it, and integrate it into the genome. This insertion becomes stable and permanent. While the helper and donor plasmids may eventually be lost by cellular and host defense mechanisms, the genome-integrated transposable element containing the therapeutic transgenic gene is a stable and permanent modification. Furthermore, the transient nature of these plasmids reduces excessive transposition and minimizes the risk of cancer.
[0207] More detailed and exemplary methods and techniques for delivering transposable elements are described, for example, in Skipper, Kristian Alsbjerg, et al. “DNA transposon-based gene vehicles-scenes from an evolutionary drive.” Journal of biomedical science 20.1 (2013): 92.
[0208] IV. Method The present invention further provides a method for inserting heterologous nucleic acid into a target nucleic acid, comprising contacting the target nucleic acid with an engineered transposable element or a gene transfer system, wherein the engineered transposable element comprises heterologous nucleic acid relating to any of the engineered transposable elements described herein, and the gene transfer system comprises heterologous nucleic acid relating to any of the gene transfer systems described herein. This method is performed in vitro or intracellularly.
[0209] In vitro method In some embodiments, a method is provided for in vitro insertion of a heteronucleotide into a target nucleic acid, comprising contacting a target nucleic acid with a manipulated transposase, wherein the manipulated transposase comprises 1) a manipulated transposase and 2) a transposase, wherein the transposase comprises a 5'-terminal repeat sequence (5'TR), a heteronucleotide, and a 3'-terminal repeat sequence (3'TR) from 5' to 3', the 5'TR comprising a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 1-26 and 79-90, the 3'TR comprising a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 27-52 and 91-102, and the transposase comprising an amino acid sequence or variant thereof selected from SEQ ID NO: 53-78 and 103-114. In some embodiments, the target nucleic acid is circular DNA. In some embodiments, the target nucleic acid is linear DNA. In some embodiments, the manipulated transposase is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0210] The methods described herein, including cell-free systems, may be used for in vitro transposition. For example, Goryshin, Igor Yu, and William S. Reznikoff. "Tn5 in vitro transposition." See Journal of Biological Chemistry 273.13 (1998): 7367-7374. Transposable elements exhibiting transposable activity are useful in the construction of second-generation sequencing (NGS) libraries, for example, in tagging methods.
[0211] In some embodiments, the process involves contacting a target nucleic acid with a manipulated transposase, the manipulated transposase comprising 1) a manipulated transposase and 2) a transposase, the transposase comprising a heteronucleotide including a 5'-to-3' repeat sequence (5'TR), a barcode sequence, and a 3' repeat sequence (3'TR), the 5'TR being a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NO: 1-26 and 79-90. The method provides a method for producing multiple barcoded nucleic acids from a target nucleic acid, wherein the 3'TR comprises a nucleic acid sequence selected from SEQ ID NO: 27-52 and 91-102, a variant thereof, or a fragment thereof, and the transposase comprises an amino acid sequence selected from SEQ ID NO: 53-78 and 103-114 or a variant thereof, thereby providing multiple barcoded nucleic acids. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is cDNA. In some embodiments, the target nucleic acid is amplified DNA. In some embodiments, the heterologous nucleic acid further comprises a primer sequence. In some embodiments, the method further comprises amplifying multiple barcoded nucleic acids to provide a nucleic acid sequencing library. In some embodiments, the method further comprises sequencing the nucleic acid sequencing library. In some embodiments, the method comprises contacting the target nucleic acid with a plurality of manipulated transposers, each manipulated transposer comprising a separate barcode sequence. In some embodiments, the nucleic acid sequencing library produced using the in vitro method described herein retains adjacency information in the target nucleic acid sequence. In some embodiments, the manipulated transposable element is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0212] In some embodiments, a tagging method is provided using either the transposase and / or TR sequence described herein. A tagging method using the Tn5 transposon is known in the art, for example, as described in US9080211B2, which is incorporated herein by reference in its entirety. The tagging method described herein uses a transposome complex composition.
[0213] In some embodiments, a transpososome complex composition is provided comprising a transposase containing an amino acid sequence or a variant thereof selected from SEQ ID NO: 53-78 and 103-114, and a heterogeneous nucleic acid containing one or two TR sequences and a tag sequence. In some embodiments, the transpososome complex contains a single heterogeneous nucleic acid that forms a hairpin. In some embodiments, the hairpin contains a cleavable site.
[0214] In some embodiments, the transposome complex comprises two transposases that bind to two different heteronucleotides. In some embodiments, a transposome complex composition is provided comprising a first transposase that binds to a first heteronucleotide comprising a 5'TR sequence and a first tag sequence, and a second transposase that binds to a second heteronucleotide comprising a 3'TR sequence and a second tag sequence. In some embodiments, the 5'TR comprises a nucleic acid sequence, variant thereof, or a fragment thereof selected from SEQ ID NO: 1-26 and 79-90, and the 3'TR comprises a nucleic acid sequence, variant thereof, or a fragment thereof selected from SEQ ID NO: 27-52 and 91-102. In some embodiments, the first tag sequence is different from the second tag sequence. In some embodiments, the manipulated transposer is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposer is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0215] In some embodiments, the method involves contacting a target nucleic acid with a plurality of transposome complexes, the transposome complex comprising (1) a first transposome complex comprising a first transposase and a first heterologous nucleic acid comprising a TR sequence and a first tag sequence, and (2) a second transposome complex comprising a second transposase and a second heterologous nucleic acid comprising a TR sequence and a second tag sequence, wherein the first tag sequence differs from the second tag sequence, the transposase comprises an amino acid sequence or a variant thereof selected from SEQ ID NO: 53-78 and 103-114, the first heterologous nucleic acid and the second heterologous nucleic acid are inserted into the target nucleic acid, and the target nucleic acid is The method provides a method for producing a nucleic acid (e.g., DNA) fragment library having a first tag sequence and a second tag sequence for a target nucleic acid, wherein the nucleic acid fragment is fragmented into multiple nucleic acid fragments, each nucleic acid fragment containing one of a first or second nucleic acid ligated to the 5' end of each nucleic acid fragment, thereby providing a library of nucleic acid fragments. In some embodiments, the transposome complex in (1) contains two first heterologous nucleic acids, and the transposome complex in (2) contains two second heterologous nucleic acids. In some embodiments, the 5'TR contains a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 1-26 and 79-90, and the 3'TR contains a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 27-52 and 91-102. In some embodiments, the method further includes amplifying the nucleic acid fragment. In some embodiments, the method further includes sequencing the nucleic acid fragment or its amplicon. In some embodiments, the manipulated transposable element is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0216] In some embodiments, the manipulated transposable element or transposomal complex is inserted into the target nucleic acid randomly, i.e., without bias toward a specific sequence basis. In some embodiments, the manipulated transposable element or transposomal complex is inserted into the target nucleic acid more randomly than PB, SB, or TB transposons.
[0217] In some embodiments, the manipulated transposers or transposomal complexes are preferentially inserted into spatially free regions of the target nucleic acid, such as open chromatin regions, nucleolus, or regions that do not bind to other DNA-binding proteins. Therefore, the in vitro methods described herein can be used to produce nucleic acid sequence libraries for measurement to study epigenomics (e.g., chromatin remodeling or DNA methylation). This includes, but is not limited to, chromatin transposase accessibility sequencing (ATAC-seq), target and sub-tagged cleavage (CUT&TAG), transposase accessibility chromatin and DNA methylation measurement (ATAC-Me), and transposase-mediated chromatin cyclization measurement (Trac-looping). ATAC-seq, CUT&TAG, ATAC-Me, and Trac-looping measurements using Tn5 transposons are, for example, seen in Buenrostro, Jason D., et al. "Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position.” Nature methods (2013) 10(2): 1213;Kaya-Okur HS, et al., “CUT&Tag for efficient epigenomic profiling of small samples and single Nature Communications (2019), 10(1):1-10;Barnett KR et al. “ATAC-Me Captures Prolonged DNA Methylation of Dynamic Chromatin Accessibility Loci during Cell Fate Transitions.” Molecular Cell, 2020; and Lai B. et al. “Trac-looping "Measures genome structure and chromatin accessibility," described in Nature Methods, 2018, 15(9): 741, is incorporated herein by reference in its entirety. In some embodiments, the sequencing library product is used for chromatin transposase accessibility sequencing (ATAC-seq).
[0218] In some embodiments, the manipulated transposable element or transposomal complex is useful for tagging spatially adjacent regions of target nucleic acids (e.g., genomic DNA). Two spatially adjacent regions may contain the same tag sequence pair at their ends. Therefore, the in vitro method described herein is used to produce a fluorescently labeled probe, which is used to perform in situ hybridization with chromatin interaction boundaries in genomic DNA, such as fluorescent in situ hybridization (FISH) based on transposases. For example, see Zhang X. et al. "Im aging chromatin interactions at sub-kilobase resolution via Tn5-FISH,” bioRxiv, This is described in 2019:601690, and its entirety is incorporated herein by reference.
[0219] Intracellular methods In some embodiments, the method provides for inserting heterologous nucleic acid into an intracellular target nucleic acid, comprising contacting a target nucleic acid with a manipulated transposase, wherein the manipulated transposase comprises 1) a manipulated transposase and 2) a transposase, wherein the transposase comprises a 5'-to-3' repeat sequence (5'TR), a heterologous nucleic acid, and a 3' repeat sequence (3'TR), wherein the 5'TR comprises a nucleic acid sequence, variant thereof, or a fragment thereof selected from SEQ ID NO: 1-26 and 79-90, the 3'TR comprises a nucleic acid sequence, variant thereof, or a fragment thereof selected from SEQ ID NO: 27-52 and 91-102, and the transposase comprises an amino acid sequence or variant thereof selected from SEQ ID NO: 53-78 and 103-114. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is intrachromosomal DNA. In some embodiments, the manipulated transposase is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0220] In some embodiments, it includes contacting a target nucleic acid with an engineered transposable element, and the engineered transposable element includes 1) the engineered transposable element and 2) a transposase. The transposable element includes a 5'-terminal repeat sequence (5'TR), a heterologous nucleic acid, and a 3'-terminal repeat sequence (3'TR) from 5' to 3'. The 5'TR includes a nucleic acid sequence selected from SEQ ID NOs: 1 to 26 and 79 to 90, a variant thereof, or a fragment thereof. The 3'TR includes a nucleic acid sequence selected from SEQ ID NOs: 27 to 52 and 91 to 102, a variant thereof, or a fragment thereof. The transposase includes an amino acid sequence selected from SEQ ID NOs: 53 to 78 and 103 to 114 or a variant thereof, and provides a method for inserting a heterologous nucleic acid into a target nucleic acid in a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the mammalian cell is an animal cell, such as a rodent cell. In some embodiments, the mammalian cell is an immune cell, such as a T cell. In some embodiments, the method is performed ex vivo. In some embodiments, the engineered transposable element is derived from any of the TEs in Table 2. In some embodiments, the engineered transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt or Tc1-1_PM.
[0221] In some embodiments, the method provides for inserting a heteronucleotide into a target nucleic acid in a plant cell, comprising contacting a target nucleic acid with a manipulated transposator, the manipulated transposator comprising 1) a manipulated transposator and 2) a transposase, the transposator comprising a 5'-to-3' repeat sequence (5'TR), a heteronucleotide, and a 3' repeat sequence (3'TR), the 5'TR comprising a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NO: 1-26 and 79-90, the 3'TR comprising a nucleic acid sequence, a variant thereof, or a fragment thereof selected from SEQ ID NO: 27-52 and 91-102, and the transposase comprising an amino acid sequence or a variant thereof selected from SEQ ID NO: 53-78 and 103-114. In some embodiments, the plant cell is a crop cell. In some embodiments, the manipulated transposator is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0222] The transposable elements or gene transfer systems described herein can be introduced into one or more cells using any of the various techniques known in the art, for example, micro This includes, but is not limited to, injection, binding of nucleic acid fragments to lipid vesicles (e.g., cationic lipid vesicles), particle impact, electroporation, DNA agglutination reagents (e.g., calcium phosphate, polylysine, polyethyleneimine), or incorporating nucleic acid fragments into a viral vector and bringing the viral vector into contact with cells. When a viral vector is used, the viral vector may include any of several viral vectors known in the art, including viral vectors selected from retroviral vectors, adenovirus vectors, or adeno-associated virus (AAV) vectors.
[0223] It is assumed that heterologous nucleic acids may contain multiple operons, or that heterologous nucleic acids may encode multiple biological products. In some embodiments, the heterologous nucleic acids encode gene circuits. An exemplary gene circuit is a set of elements (the “outputs” of each element) that are each transcribed and / or translated to produce mRNA or protein. Element outputs can interact with other parts (e.g., to regulate transcription or translation) or with other molecules within the cell (e.g., small molecules, DNA, RNA, or proteins present in the cellular environment). For example, the circuit may be a metabolic pathway or a gene cascade, and it may be naturally occurring or unnatural, or artificially engineered. Each element in the circuit may include components or gene modules such as promoters, ribosome-binding sites (RBS), coding sequences (CDS), and / or terminators. These components may be interconnected or assembled in different ways to realize different elements, and the final elements may be combined in different ways to create different circuits or pathways. In addition to these elements, the circuit may include other types of molecules present in the cell or in the cellular environment that interact with the components.
[0224] Therefore, in some embodiments, the method provides a way to insert heterologous nucleic acid into a target nucleic acid within a cell, comprising: 1) an engineered transposer; and 2) a transposase, wherein the transposer comprises a 5'-to-3' repeat sequence (5'TR), a heterologous nucleic acid, and a 3' repeat sequence (3'TR), the 5'TR comprising a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 1-26 and 79-90; the 3'TR comprising a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 27-52 and 91-102; and the transposase comprising an amino acid sequence or variant thereof selected from SEQ ID NO: 53-78 and 103-114, or a heterologous nucleic acid encoding a gene circuit. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is intrachromosomal DNA. In some embodiments, the engineered transposer is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0225] Genetic circuits are used in gene therapy. The design, use, and techniques of genetic circuits are well-known in this field. Furthermore, one can refer to, for example, Brophy, Jennifer AN, and Christopher A. Voigt. "Principles of genetic circuit design." Nature methods 11.5 (2014): 508.
[0226] The method can be applied to any suitable cell type. In some embodiments, the cells are bacterial, yeast, fungal, algal, plant, or animal cells. In some embodiments, the cells are isolated from a natural source, such as a tissue biopsy. In some embodiments, the cells are isolated from a cell line cultured in vitro. In some embodiments, the cells are genetically modified cells. In some embodiments, the cells are seed cells that undergo proliferation, differentiation, or both within the nucleus. be.
[0227] In some embodiments, the cells are animal cells derived from organisms selected from cattle, sheep, goats, horses, pigs, deer, chickens, ducks, geese, rabbits, and fish.
[0228] In some embodiments, the cells are plant cells of biological origin selected from corn, wheat, barley, oats, rice, soybeans, oil palm, safflower, sesame, tobacco, flax, cotton, sunflower, pearl millet, foxtail millet, sorghum, rapeseed, cannabis, vegetable crops, fodder crops, commercial crops, woody crops, and biomass crops.
[0229] In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the human cells are human embryonic kidney 293T (HEK293T or 293T) cells or HeLa cells. In some embodiments, the mammalian cells are selected from immune cells, liver cells, tumor cells, stem cells, fertilized eggs, muscle cells, and skin cells.
[0230] In some embodiments, the cells are immune cells selected from the group consisting of cytotoxic T cells, helper T cells, natural killer (NK) T cells, iNK-T cells, NK-T-like cells, αγδ T cells, tumor-infiltrating T cells, and dendritic cell (DC)-activated T cells. In some embodiments, this method generates modified immune cells such as CAR-T cells or TCR-T cells.
[0231] In some embodiments, the heterologous nucleic acid is inserted into the cell's genome. In some further embodiments, the insertion of the heterologous nucleic acid inactivates the cell's genes. In some embodiments, the heterologous nucleic acid encodes a protein or RNAi molecule.
[0232] Furthermore, the aforementioned heterogeneous nucleic acids can encode RNA molecules useful for genome editing. Examples of such RNA molecules include, but are not limited to, CRISPR RNA (crRNA), transactivating crRNA (tracrRNA), guide RNA (gRNA), and single guide RNA (sgRNA).
[0233] In some embodiments, the heterogeneous nucleic acid encodes a biological product selected from the group consisting of reporter proteins, antigen-specific receptors, therapeutic proteins, antibiotic resistance proteins, RNAi molecules, cytokines, kinases, antigens, antigen-specific receptors, cytokine receptors, and suicide polypeptides. For example, the heterogeneous nucleic acid can encode a receptor specific to a tumor-associated antigen. T cells manipulated in this way can recognize and specifically kill tumor cells expressing the tumor-associated antigen. In another example, the heterogeneous nucleic acid encodes a hygromycin resistance protein, allowing for the establishment of a hygromycin-resistant cell line. Alternatively, the heterogeneous nucleic acid may have no biological function and can be used to block the function of other genes by inserting itself into essential genes.
[0234] In some specific embodiments, the heterogeneous nucleic acid encodes a therapeutic protein useful for gene therapy. In some embodiments, the heterogeneous nucleic acid encodes a therapeutic antibody. In some embodiments, the heterogeneous nucleic acid encodes an engineered receptor such as a chimeric antigen receptor (CAR) or an engineered TCR.
[0235] In some embodiments, the heterogeneous nucleic acid includes one or more multicleaning sites (MCS) to facilitate the insertion of a target polynucleotide ("cargo gene").
[0236] In some embodiments, this method is performed ex vivo. In some embodiments, transduced or transfected cells (e.g., mammalian cells) are grown ex vivo after introducing heterologous nucleic acids into the cells. In some embodiments, transduced or transfected cells are cultured for at least about 1, 2, 3, 4, 5, 6, 7, 10, 12, or 14 days for growth. In some embodiments, transduced or transfected cells are cultured for no more than about 1, 2, 3, 4, 5, 6, 7, 10, 12, or 14 days. In some embodiments, transduced or transfected cells are further evaluated or screened to select manipulated cells.
[0237] Reporter genes or selection markers can be used to identify potentially transfected cells and to evaluate the function of regulatory sequences. Generally, a reporter gene is a gene that is not present in or expressed by the receptor organism or tissue, and the expression of the polypeptide it encodes is demonstrated by several readily detectable properties, such as enzymatic activity. Reporter gene expression is measured at appropriate time points after the DNA has been introduced into receptor cells. Suitable reporter genes may include genes encoding luciferase, β-galactosidase, chloramphenicol acetyltransferase, secreted alkaline phosphatase, or the green fluorescent protein gene (e.g., Ui-Tei et al. FEBS Letters 479: 79-82 (2000)). Suitable expression systems are publicly known and can be created using known techniques or commercially available.
[0238] Other methods for confirming the presence of heterologous nucleic acids within cells include molecular biological assays, such as Southern blotting, Northern blotting, RT-PCR, and PCR, which are well known to those skilled in the art; and biochemical assays, such as the detection of the presence or absence of specific peptides by immunological methods (e.g., ELISA and protein blotting).
[0239] This application envisions a method for generating near-isogenic lineages of mammalian cells to study genetic variation. This application also considers genome modification of microorganisms, cells, plants, animals, or synthetic organisms to produce biomedical, agricultural, and industrially useful products. These methods can be used as biological research tools for understanding genomes, such as gene knockout or knock-in studies.
[0240] Cells modified with heterologous nucleic acids using any of the methods described herein, and organisms (e.g., animals, plants, or fungi) comprising or produced from such cells are also provided.
[0241] target nucleic acid The methods described herein are suitable for inserting heterologous nucleic acids into various target nucleic acids. In some embodiments, the target nucleic acid is DNA. In some embodiments, the target nucleic acid is single-stranded. In some embodiments, the target nucleic acid is double-stranded. In some embodiments, the target nucleic acid includes single-stranded and double-stranded regions. In some embodiments, the target nucleic acid is linear. In some embodiments, the target nucleic acid is circular. In some embodiments, the target nucleic acid includes one or more modified nucleotides, such as methylated nucleotides, damaged nucleotides, or nucleotide analogs. In some embodiments, the target nucleic acid is unmodified. In some embodiments, the target nucleic acid is bound to one or more proteins, such as nucleoli.
[0242] The target nucleic acid may be of any length, for example, at least one of approximately 100 bp, 200 bp, 500 bp, 1000 bp, 2000 bp, 5000 bp, 10 kb, 20 kb, 50 kb, 100 kb, 200 kb, 500 kb, or 1 Mb or more. Several implementations Morphologically, the target nucleic acid does not exceed any of the following: approximately 500kb, 200kb, 100kb, 50kb, 40kb, 30kb, 20kb, 10kb, 5kb, 2kb, 1kb, 500bp, or 200bp. In some embodiments, the target nucleic acid is any of the following: approximately 100bp to 500bp, 500bp to 1kb, 100bp to 1kb, 1kb to 2kb, 100bp to 5kb, 100bp to 10kb, 100bp to 20kb, 1kb to 5kb, 1kb to 10kb, 1kb to 20kb, 20kb to 100kb, or 100kb to 1Mb. The target nucleic acid may also contain any sequence. In some embodiments, the target nucleic acid enriches a specific sequence that is a hotspot for transfer in the manipulated transposable element or gene delivery system described herein. In some embodiments, the target nucleic acid is AT enriched and has, for example, an AT content of at least about 40%, 45%, 50%, 55%, 60%, 65%, or more. In some embodiments, the target nucleic acid is not AT enriched. In some embodiments, the target nucleic acid is not enriched for specific hotspot sequences because the manipulated transposable elements or gene delivery systems described herein do not prefer to insert heterologous nucleic acids into specific sequences or sequence motifs. In some embodiments, the target nucleic acid has one or more secondary or higher-order structures. In some embodiments, the target nucleic acid is not in an aggregated state, such as in chromatin.
[0243] In some embodiments, the target nucleic acid is located inside a cell. In some embodiments, the target nucleic acid is located in the cell nucleus. In some embodiments, the target nucleic acid is endogenous to the cell. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is chromosomal DNA. In some embodiments, the target nucleic acid is a protein-coding gene or its functional region, such as a coding region, or a regulatory element such as a promoter, enhancer, or 5' or 3' untranslated region. In some embodiments, the target nucleic acid is a non-coding gene such as a transposon, miRNA, tRNA, ribosomal RNA, ribozyme, or lincRNA. In some embodiments, the target nucleic acid is a plasmid.
[0244] In some embodiments, the target nucleic acid is exogenous to the cell. In some embodiments, the target nucleic acid is a viral nucleic acid, such as viral DNA. In some embodiments, the target nucleic acid is a horizontally transferred plasmid. In some embodiments, the target nucleic acid is integrated into the cell's genome. In some embodiments, the target nucleic acid is not integrated into the cell's genome. In some embodiments, the target nucleic acid is an intracellular plasmid. In some embodiments, the target nucleic acid is present in an extrachromosomal array.
[0245] In some embodiments, the target nucleic acid is an isolated nucleic acid, such as isolated DNA. In some embodiments, the target nucleic acid is present in a cell-free environment. In some embodiments, the target nucleic acid is an isolated vector, such as a plasmid. In some embodiments, the target nucleic acid is an isolated linear DNA fragment.
[0246] V. Kits and Articles This application also provides kits and articles comprising any transposable element or gene transfer system described herein. In some embodiments, the kit includes a protocol for inserting a heterologous nucleic acid into a target nucleic acid (e.g., using any of the methods described herein). In some embodiments, the kit is used for in vitro insertion of a heterologous nucleic acid into a target nucleic acid. In some embodiments, the kit is used for inserting a heterologous nucleic acid into a target nucleic acid within a cell, such as a mammalian cell or a plant cell. The above-described kits and articles herein are useful for modifying target nucleic acids in in vitro or exo, genetic research, and gene therapy.
[0247] In some embodiments, a kit is provided comprising an engineered transposable element, the engineered transposable element comprising a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3' repeat sequence (3'TR), wherein the 5'TR comprises a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 1-26 and 79-90, and the 3'TR comprises a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 27-52 and 91-102, the transposable element exhibits transposable activity that enables insertion of the heterogeneous nucleic acid into the DNA of cells (e.g., mammalian or plant cells). In some embodiments, the heterogeneous nucleic acid comprises one or more multicloning sites (MCS) to facilitate insertion of the target polynucleotide ("cargo gene"). In some embodiments, the engineered transposable element is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable elements are derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0248] In some embodiments, a kit is provided comprising a gene transfer system comprising 1) an engineered transposer and 2) a transposase or nucleic acid encoding a transposase, wherein the transposer comprises a 5'-to-3' repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3' repeat sequence (3'TR), the 5'TR comprising a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 1-26 and 79-90, the 3'TR comprising a nucleic acid sequence, variant thereof, or fragment thereof selected from SEQ ID NO: 27-52 and 91-102, the transposase comprising an amino acid sequence, or variant thereof selected from SEQ ID NO: 53-78 and 103-114, and the transposer exhibits transposability activity that enables insertion of the heterogeneous nucleic acid into the DNA of a cell (e.g., a mammalian cell or a plant cell). In some embodiments, the heterogeneous nucleic acid comprises one or more multiclonings (MCS) to facilitate insertion of the target polynucleotide ("cargo gene"). In some embodiments, the manipulated transposable element is derived from one of the TEs in Table 2. In some embodiments, the manipulated transposable element is derived from Tc1-8B_DR, Tc1-3_FR, Mariner2_AG, Tc1-1_Xt, or Tc1-1_PM.
[0249] In some embodiments, the kit includes one or more reagents used in any of the methods described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reactive or storage buffers. The reagents may be provided in a form that can be used in a particular assay, or in a form that requires the addition of one or more other components before use (e.g., in a concentrated or lyophilized form). The buffers may include, but are not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof, and any other buffer. In some embodiments, the kit includes media, buffers, reagents, etc., for enabling the proliferation or induction of cells modified using the above-described manipulated transposable elements or gene transduction systems. In some embodiments, the kit includes buffers, reagents, etc., for isolating and / or producing target nucleic acids modified using the above-described manipulated transposable elements or gene transduction systems. In some embodiments, the kit includes primers and reagents for producing sequencing libraries using the above-described manipulated transposable elements or gene transduction systems.
[0250] The kit employs appropriate packaging. Appropriate packaging includes, but is not limited to, vials, bottles, wide-mouth bottles, and flexible packaging (e.g., sealed polyester film bags or plastic bags). The kit may selectively provide additional components such as buffers and explanatory information. Accordingly, the present application also provides articles including vials (e.g., sealed vials), bottles, wide-mouth bottles, flexible packaging, etc.
[0251] Examples The following examples are provided merely to illustrate the present invention and are not intended to limit the invention in any way. The following examples and detailed descriptions are provided explanatoryly, not limitingly.
[0252] Example 1: Identification of candidate active transfer factors This example describes the computer identification of candidate transposable elements (TEs, transposons) across species. Transposons are abundant in various species. However, only a small number of transposons exhibit transposability activity in mammalian cells. Therefore, there is a need for a systematic method to identify candidate active transposons useful as reagents for genome engineering and gene therapy. This example focuses on identifying transposons with terminal reverse repeat sequences (TIRs, also called terminal repeats TRs), but this method may be used to identify any other type of transposon.
[0253] Materials and methods Repbase is the most common transposon database, containing 38,000 transposon sequences from various eukaryotes (Bao, W., KK Kojima, and O. Kohany, Repbase Update, a database of repetitive elements in eukaryotic genomes. Mob DNA, 2015. 6:p.11). Many of these prototype sequences are common sequences that reconstruct each transposon family and approximate the activation state of their ancestors. Therefore, common sequences include Sleeping Beauty (Ivics, Z., et al., Molecular reconstruction of Sleeping Beauty, a Tc1-like transposon from fish, and its transposition in Transgenic and genetic, as in human cells. Cell, 1997. 91(4): p. 501-10) This is useful for experimental reconstruction of active transposons for therapeutic purposes. As the first step of this project, all common sequences will be downloaded from Repbase (version Rebase24.02) for candidate transposon screening. This is because common sequences do not necessarily contain the complete transposase gene (especially in the case of highly degenerated, ancient transposons). Since they do not function as active transposons in other systems (e.g., mammalian cells), their identified TE transposition activity in primitive species was evaluated using several parameters, including key domains encoded by the transposase, mean sequence differences between common and family members, copy number, conserved terminal reverse repeat (TR) sequences, and conserved target site repeats (TSDs) ligated to the TE side. To obtain the above information, genome sequences from 100 animals were obtained from the UCSC Genome Browser (Haeussler, M., et al., The UCSC Genome Browser database: 2019). Download update. Nucleic Acids Res, 2019. 47(D1): pp. D853-D858). The mu sequences are repeatedly masked using a common sequence from Repbase. Active transposons are defined as follows: 1) The candidate transposon matches the common sequence from beginning to end; and 2) The length of the candidate transposon reaches 90% of the length of the common transposon to ensure there are no obvious deletions within the transposon.
[0254] result Figure 1 shows a flowchart of the bioinformatics pipeline and the number of candidate active transfer factors at each stage of the pipeline. From a total of 100 animal genomes, 26,853,019 copies of DNA transposons (mainly TIR transposons) were identified. These numerous copies are fragments degenerated from active transposons. The total number of full-length DNA transposon copies was 1,895,466, which were mapped to 1,577 common DNA transposons within the Repbase. It is. Next, these transposons were examined to determine whether or not they contained transposase genes. ORFfinder detects open reading frames (ORFs). The length cutoff value is set to 300 amino acids. This protein sequence is then used for searching for important domains in the PfamHMM library using PfamScan (Madeira, F., et al., The EMBL-EBI search and sequence analysis tools APIs in 2019. Nucleic Acids Res, 2019. 47(W1): p. W636-W641). After this filtering step, 131 transposons with Tn domains remain in the pipeline. Next, RepeatMasker was used to determine the copy number for each transposon (Smit, AFA, Hubley, R & Green, P. RepeatMasker Open-4.0.2013-2015). The mean sequence difference between various common transposons and their family members was calculated, and this difference ranged from 0 to 24.4% across different species. As shown in Table 1, the mean difference for 62 identified TEs was less than 5%, while the mean difference for 69 identified TEs was greater than 5%. These 131 active transposons were found to be distributed across five superfamilies (Figure 2). Of the 62 transposons identified in Table 1, we found that a total of 11 transposons met the following more stringent criteria and are therefore highly suitable for genome engineering: 1) The transposon length is less than 3000 bp. 2) The number of miniature reverse repeat transposers (MITEs) within the transposon exceeds 10. 3) The mean difference value of the transposons is less than 1% (TE IDs: 4, 5, 7, 12, 18, 20, 22, 25, 26, 28, and 30, see Table 1). In summary, these data represent 131 candidate active transposons identified by pan-genomic bioinformatics analysis of 100 genomes. Therefore, this example demonstrates the success of developing a robust bioinformatics pipeline for identifying candidate active transposers.
[0255] Example 2: Verification of candidate active transfer factors This example describes the experimental verification of the candidate active transfer factors identified in Example 1.
[0256] Materials and methods DNA synthesis and plasmid construction Mammalian codon-optimized transposase ORFs with EcoRI and NotI laterally linked were synthesized and cloned into CMV-hyPBase vectors (K. Yusa, et al., A hyperactive piggyBac transposase for mammalian applications. Proc Natl Acad Sci USA, 2011. 108(4): p. 1531-6). These vectors were then used under the human CMV promoter. This is a helper plasmid for assisting transposase expression. The transposon donor plasmid contains left and right transposon fragments located on either side of the antibiotic resistance gene. As used herein, the left transposon fragment (LTF) refers to the fragment of the transposon ORF sequence from the 5'TSD to the TE sequence, specifically the start codon. As used herein, the right transposer fragment (RTF) refers to the fragment from the stop codon of the transposase ORF sequence to the 3'TSD of the TE sequence. Typically, the TR sequence is located within the left and right transposon fragments. For example, the 5'TR sequence is within the LTF sequence, and the 3'TR sequence is within the RTF sequence. The LTF and RTF sequences used in this experiment were synthesized by Qinglan Biotechnology Inc. and cloned into pMV vectors. The 5' and 3' multi-cleaning sites (MCS) were also synthesized in the transposon fragments for cargo gene cloning. The TSD sequence is located on the outermost side. The donor plasmid for metastasis screening in mammalian cells carries a P2A-binding puromycin resistance gene and an enhanced GFP gene expressed by a PGK promoter. Figure 3 shows an exemplary set of auxiliary and donor constructs for validating active transposable elements.
[0257] Metastasis in mammalian cells Four mammalian cell lines, HEK293T (also known as 293T), HeLa, Hct116, and K562, were used for screening active transposons. The 293T and HeLa cell lines were maintained in DMEM medium and supplemented with 10% fetal bovine serum and 1% penicillin / streptomycin. The Hct116 and K562 cell lines were maintained in RPMI 1640 medium and supplemented with 10% fetal bovine serum and 1% penicillin / streptomycin. In addition to these cell lines, CD3 was analyzed using the EasySep Human T Cell Enrichment kit. + T cells were isolated, and then monocytes were collected by histopaque-1077 (Sigma-Aldrich) gradient separation. +T cells were cultured in X-Vivo 15 medium (Lonza) supplemented with 5% (v / v) heat-inactivated fetal bovine serum, 2 mM L-glutamine, and 1 mM sodium pyruvate. For metastasis measurement, 1.2×10 5 HEK293T cells, 0.7×10 5 HeLa cells or 1.0×10 5 Hct116 cells were seeded into each well of a 24-well plate 18 hours before transfection. Using Lipo3000, 200 ng of helper plasmid and 100 ng of donor plasmid, or only 100 ng of donor plasmid were delivered to individual cell lines. Two days after transfection, the cell number was counted and the transfection efficiency was measured by FACS. Next, 1 / 100 th HEK293T cells after transfection, 1 / 100 th HeLa cells after transfection or 1 / 100 th Hct116 cells after transfection were transferred to a 100-mm plate and selected with puromycin (0.5 μg / ml) for 10 days (HEK-293T cell line) or 14 days (HeLa cell line and Hct116 cell line). After puromycin screening, the cells were washed once with 5 ml of cold PBS, fixed with 4% PFA for 15 minutes, and then stained with 0.2% methylene blue (in PBS) for 1 hour (Wu et al., piggyBac is a flexible and highly active transposon as compared to Sleeping Beauty, Tol2, and Mos1 in mammalian cells. PNAS, 2006. 103: p. 15008-15013). Finally, the remaining non-specifically stained ones were washed away with PBS. Single-stained colonies were counted using Image J software. The metastasis efficiency was calculated based on the total number of transfected cells and the transfection efficiency calculated previously. For suspension cells such as the K562 cell line and CD3+ T cells, translocation activity was evaluated based on the percentage of GFP-positive cells 14 days after electroporation with plasmids diluted to a very small proportion.
[0258] Construction of TE insertion libraries and bioinformatics analysis Stable conversion using DNeasy blood and tissue kit (Qiagen, Germany) Genomic DNA was isolated from trans-K562 cells and cut to an average length of 600 bp using a Covaris M220 sonicator (Covaris, USA). End repair was performed, linker ligation and amplification were carried out by nested PCR, and purification and sequencing were performed on an Illumina HiSeq sequencer. Next, the distance to the nearest gene and TSS, as well as the different chromatin states, were compared for random insertion sites in the upstream and downstream regions of the primary sequence linked to the TE integration site and the vector integration site.
[0259] result The verification items include three experiments. In the first iteration, we examined 11 transfectas with a difference of less than 1% and a MITE copy number greater than 10 (i.e., TE IDs: 4, 5, 7, 12, 18, 20, 22, 25, 26, 28, and 30). While we do not wish to be bound by any theory, we hypothesize that low difference values and high MITE copy numbers indicate a high probability that the transposable element is recently active in a certain species (see R. Mitra, et al., Functional characterization of piggyBat from the bat Myotis lucifugus unveils an active mammalian DNA transposon. Proc Natl Acad Sci USA, 2013. 110(1): p. 234-9). During plasmid construction, we discovered that two transposon TR sequences from different sources were available. One was from a common sequence provided by a database, and the other was obtained through alignment of an autonomous sequence with a MITE sequence. To test whether the different sources of TR sequences would affect the measurement results, we designed two sets of donor plasmids for transcriptional measurements, using TR sequences from different sources, respectively. As a result, of the 11 candidate TEs tested in the initial validation, three were found to be active. Tc1-3_Xt (TE ID 25) was found to be active in both HEK293T cells and HeLa cells, with a transduction efficiency of approximately 9.6% in HEK293T cells and approximately 3.6% in HeLa cells, which is about half that of piggyBac (18.4% in HEK293T cells and 8.3% in HeLa cells). Two TEs were found to be active in only one of the two cell lines tested: hAT-3_XT (TE ID 22) had an activity of approximately 1% in HeLa cells, and hAT-5_DR (TE ID 28) had an activity of approximately 2% in HEK-293T cells. Since there was no significant difference in transduction efficiency between the two donor plasmid groups with TR sequences from different sources, subsequent experiments relied solely on TR sequences from common sequences in the database. In the following experiment, the translocation activity of 48 candidate TEs selected from Table 1 was evaluated in 293T cell lines and HeLa cell lines (of which 16 TEs were tested only in 293T cells). With a surprisingly high success rate, 22 of the 48 candidate TEs tested were found to be active. Of these, eight TEs—Tc1-8B_DR (TE ID 14), Tc1-3_FR (TE ID 29), Mariner2_AG (TE ID 35), Tc1-1_Xt (TE ID 36), Tc1-1_AG (TE ID 37), Tc1-1_PM (TE ID 43), Tc1-4_Xt (TE ID 54), and Tc1-15_Xt (TE ID 56)—had translocation efficiencies equal to or better than piggyBac. For reference, piggyBac had an induction efficiency of 17.44% in 293T cells and 10.25% in HeLa cells. For some active TEs, the translocation efficiency was far higher than expected, making the standard dilutions of the cells insufficient to keep viable colonies separated on the measurement plate for accurate counting of individual colonies. As a result, the number of stained colonies counted by the imaging software was underestimated. For these TEs, even though the staining of the detection plate clearly indicated higher translocation efficiency, the translocation efficiency calculated based on the underestimated colony count simply underestimated the actual translocation efficiency. In other words, the translocation activity of 59 candidate TEs was evaluated from two experiments. Figures 4A-4B show the translocation efficiency of the evaluated TEs compared to control TEs. A total of 25 TEs were found to be active in human cell lines. Of the 25 active transposable elements (TEs), 9 belonged to the hAT superfamily and 16 belonged to the TcMariner superfamily. The transposability activity of 8 active TEs was comparable to or greater than that of piggyBac. All 8 of these highly efficient active TEs belonged to the TcMariner superfamily, indicating that a larger number of active transposons are distributed within this superfamily. Furthermore, 9 of the 25 active TEs were derived from the tropical clawed frog (Xenopus tropicalis). No significant relationship was found between the transposability activity of TEs and the difference value or MITE copy number. In the third experiment, the remaining 72 candidate TEs selected from Table 1 were evaluated for their translocation activity in 293T cell lines and HeLa cell lines. Of the candidate TEs, only 69 with a difference greater than 5% in 293T cells were tested. Of the 72 candidate TEs tested, 13 were found to be active. The results are shown in Figures 6A-6B. Table 2 summarizes the active transposases (TEs) based on the above experiments, including a total of 38 TEs whose activity was verified in both 293T cell lines and HeLa cell lines, or in only one of both cell lines. Figures 7A-7C show phylogenetic trees of TEs belonging to different superfamilies, identified based on transposase sequences. Between various transposases within each superfamily... Tables 3-5 show the percentage of sequence similarity. Furthermore, five DNA TEs (Tc1-8B_DR (TE ID 14), Tc1-3_FR (TE ID 29), Mariner2_AG (TE ID 35), Tc1-1_Xt (TE ID 36), and Tc1-1_PM (TE ID 43) showed higher translocation activity than the multiple-optimized SB100X. The five TEs with the highest activity also showed higher translocation activity in HEK293 T cells, HeLa cells, and HCT116 cells (Figures 8A-9B), and Figures 10A-10B show the translocation activity of TEs in primary T cells. Three TEs (Tc1-8B_DR (TE ID 14), Tc1-3_FR (TE ID 29), Tc1-1_PM (TE ID 43) In 43), the phenomenon known as overproduction inhibition (OPI) did not occur. Figures 11A-11B show that transposase activity increases as the ratio of helper plasmids encoding transposases to transposon plasmids is optimized. Figure 12 shows the cargo capacities of the five most active TEs compared to the identified control TEs, piggyBAC, hyperPiggyBac, and SB100X, including gene lengths of 2 kb, 5 kb, 10 kb, and 20 kb. Comparing active and inactive TEs among 131 candidate TEs, active TEs exhibited lower average diversity, slightly longer predicted TIR sequences, and a significantly increased number of autonomous TEs. These differences help identify features that can be used in bioinformatics pipelines to identify active TEs, including in preliminary and large-scale screening settings. The integration sites searched for in transposon integration sites reveal the highly preferred TA target site dinucleotides of the five most active transposons (TEs) (Figure 13). The target site dinucleotides of TEs belong to the Tc / Mariner superfamily and are highly conserved. Figures 14A-18 show the frequency of integration into genomic features, including distance to the gene and transcription start site (TSS), and different chromatin states, by comparing computer-generated random data with control TEs and piggyBAC. The data showed that all five most active TEs had very low preference for integration near the gene sequence and no preference for upstream or downstream sequences of the gene and TSS. Regarding gene expression levels and chromatin states, the five most active TEs also showed a nearly random integration pattern and were often not inserted into regions of excessive activity.
[0260] In summary, these data demonstrate the successful establishment of a robust bioinformatics pipeline for identifying candidate active transposable elements (TEs) and the successful experimental validation of the identified TEs. The transposability efficiencies and integration patterns of various TEs suggest their potential usefulness in genomic engineering applications and gene therapy.
[0261] [Table 1-1]
[0262] [Table 1-2]
[0263] [Table 1-3]
[0264] [Table 1-4]
[0265] Table 1-5
[0266] Table 1-6
[0267] Table 1-7
[0268] Table 1-8
[0269] Table 1-9
[0270] Table 1-10
[0271] Table 1-11
[0272] Table 1-12
[0273] Table 1-13
[0274] Table 1-14
[0275] Table 1-15
[0276] Table 1-16
[0277] Table 1-17
[0278] Table 1-18
[0279] Table 1-19
[0280] Table 1-20
[0281] Table 1-21
[0282] Table 1-22
[0283] Table 1-23
[0284] Table 1-24
[0285] Table 1-25
[0286] Table 1-26
[0287] Table 2-1
[0288] Table 2-2
[0289] Table 2-3
[0290] Table 2-4
[0291] Table 3
[0292] Table 4
[0293] Table 5-1
[0294] Table 5-2
[0295] Table 5-3
[0296] [Sequence Listing] SEQUENCE LISTING <110> Institute of Zoology, Chinese Academy of Sciences <120> ACTIVE DNA TRANSPOSON SYSTEMS AND METHODS FOR USE THEREOF <130> 18245-20006.41 <150> PCT / CN2020 / 082087 <151> 2020-03-30 <160> 206 <170> FastSEQ for Windows Version 4.0 <210> 1 <211> 5 <212> DNA <213> Anopheles gambiae <400> 1 taggc 5 <210> 2 <211> 3 <212> DNA <213> Anopheles gambiae <400> 2 tag 3 <210> 3 <211> 201 <212> DNA <213> Danio rerio <400> 3 cagcggggaa ataagtatt tgacacatca gcattttat cagtaagggg atttctaagt 60 gggctactga cacaaatttc ctaccagatg tagccatca gccaaatatt gattcatac 120 aaagaatca gaacttaa gtatacaagt tgagtcataa taaataagt gaatgacac 180 scream scream a 201 <210> 4 <400> 4 000 <210> 5 <211> 6 <212> DNA <213> Xenopus tropicalis <400> 5 step 6 <210> 6 <211> 210 <212> DNA <213> Xenopus tropicalis <400> 6 cagtggagga aaattatt tgacccctca ctgattttgt aagtttgtcc atgacaaag 60 aaatgaaag tctcagaaca gtatcattc aatggtaggt ttatttac agtggcagat 120 agcacatca aaggaaatc gaaaaaataa ctttaaataa aagatagcaa ctgatttgca 180 slap slap slap slapcc 210 <210> 7 <211> 15 <212> DNA <213> Danio rerio <400> 7 caggggtggc gaacc 15 <210> 8 <211> 131 <212> DNA <213> Takifugu rubripes <400> 8 cagtgagagt aaaaagtatt tgatcccttg ctgattttgt tggtttgtcc actaataaag 60 acatgatcat tctatacttt taatggtaga tgtattctaa catggagaga cagaatatca 120 aaaaaaat c 131 <210> 9 <211> 213 <212> DNA <213> Xenopus tropicalis <400> 9 caaaccggat tccaaaaaag ttgggacact aaacaaattg tgaataaaaa ctgaacgcaa 60 tgatgtggag gtgccaactt ctaatatttt attcagaata gaacataaat cacggaacaa 120 aagtttaaac tgagaaaatg taccatttta agggaaaaat atgttgattc agaatttcat 180 ggtgtcaaca aatcccaaaa aagttgggac aag 213 <210> 10 <211> 214 <212> DNA <213> Xenopus tropicalis <400> 10 cagtggcttg caaaagtatt cggccccctt gaacttttcc acattttgtc acattacagc 60 cacaaacatg aatcaatttt attggaattc cacgtgaaag accaatacaa agtggtgtac 120 acgtgagaag tggaacgaaa atcatacatg attccaaaca ttttttacaa ataaataact 180 gcaaagtggg gtgtgcgtaa ttattcagcc ccct 214 <210> 11 <211> 18 <212> DNA <213> Anopheles gambiae <400> 11 tacagtgtcg gacaaatc 18 <210> 12 <211> 63 <212> DNA <213> Xenopus la...
Claims
1. A manipulated transposable element, in the order from 5' to 3', It comprises a 5'-terminal repeat sequence (5'TR), a heterogeneous nucleic acid, and a 3'-terminal repeat sequence (3'TR), The 5'TR contains a nucleic acid sequence having at least 90% sequence identity with the nucleic acid sequence of SEQ ID NO:
11. The 3'TR contains a nucleic acid sequence having at least 90% sequence identity with the nucleic acid sequence of SEQ ID NO:
37. Transposable element.
2. Further comprising left transposon fragments (LTFs) and right transposon fragments (RTFs), The 5'TR contains the nucleic acid sequence of SEQ ID NO: 11, and the 3'TR contains the nucleic acid sequence of SEQ ID NO:
37. The LTF contains the nucleic acid sequence of SEQ ID NO: 125, and the RTF contains the nucleic acid sequence of SEQ ID NO:
151. The transposable element according to claim 1.
3. The transposable factor according to claim 1 or 2, wherein the transposable activity of the transposable factor is higher than that of the piggyBac (PB) transposon, the Sleeping Beauty (SB) transposon, and / or the TcBuster (TB) transposon.
4. The transposable element according to claim 3, wherein the cell is an animal cell, plant cell, algal cell, fungal cell, yeast cell, or bacterial cell.
5. The transposable element according to any one of claims 1 to 4, wherein the transposable activity of the transposable element in human embryonic kidney 293T (293T) cells is higher than the transposable activity in HeLa cells.
6. The transposable element according to claim 5, wherein the transposable element is present in the vector, and the vector is a plasmid vector or a viral vector.
7. A gene transfer system comprising: 1) an engineered transposer according to any one of claims 1 to 6; and 2) a transposase, or a nucleic acid encoding a transposase, wherein the transposase comprises the amino acid sequence of SEQ ID NO: 63 or a variant thereof, wherein the amino acid sequence of the variant has at least 90% sequence identity with the amino acid sequence of SEQ ID NO:
63.
8. The gene transfer system according to claim 7, wherein the 5'TR contains the nucleic acid sequence of SEQ ID NO: 11, the 3'TR contains the nucleic acid sequence of SEQ ID NO: 37, and comprises an LTF containing the nucleic acid sequence of SEQ ID NO: 125 and an RTF containing the nucleic acid sequence of SEQ ID NO: 151, and the transposase contains the amino acid sequence of SEQ ID NO:
63.
9. The gene transfer system according to claim 7 or 8, wherein the gene transfer system includes a nucleic acid encoding the transposase, and the transposable factor and the nucleic acid encoding the transposase are in different vectors or the same vector.
10. A method for inserting a heterologous nucleic acid into a target nucleic acid in an organism other than a human, or a method for inserting a heterologous nucleic acid into a target nucleic acid ex vivo or in vitro, A method comprising inserting the heterologous nucleic acid into the target nucleic acid by contacting the target nucleic acid with a transposable element according to any one of claims 1 to 6 or a gene transfer system according to any one of claims 7 to 9.
11. The method according to claim 10, wherein the target nucleic acid is located inside a cell, the target nucleic acid is genomic DNA, and / or the cell is an animal cell, plant cell, algal cell, fungal cell, yeast cell or bacterial cell.
12. A method according to claim 11, wherein the genes of the cells are inactivated by insertion of the heterologous nucleic acid, The aforementioned heterogeneous nucleic acid encodes a protein, or The method wherein the aforementioned heterogeneous nucleic acid encodes RNA.
13. The protein is selected from the group consisting of reporter proteins, engineered receptors, cytokines, antibiotic resistance proteins, antigens, and therapeutic proteins, and / or The RNA is selected from the group consisting of therapeutic RNA, small interfering RNA (siRNA), microRNA, short hairpin RNA (shRNA), long non-coding RNA (LINCRNA), and guide RNA (gRNA). The method according to claim 12.
14. The method according to any one of claims 10 to 13, wherein the length of the heterologous nucleic acid is 2 kb to 300 kb, and / or the insertion is random.
15. A kit comprising an engineered transposable element according to any one of claims 1 to 6, or a gene transfer system according to any one of claims 7 to 9, and a protocol for inserting a heterologous nucleic acid into a target nucleic acid.