Transposases and their use

Fusion proteins with N-terminal deletions and double CRDs, along with specific mutations, enhance the thermal stability and integration/excision efficiency of transposases for site-specific genetic manipulation.

JP2026517869APending Publication Date: 2026-06-02POSEIDA THERAPEUTICS INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
POSEIDA THERAPEUTICS INC
Filing Date
2024-05-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The need for site-specific transposases for gene editing remains unmet, with existing PiggyBac transposases lacking optimal thermal stability and integration/excision efficiency.

Method used

Development of fusion proteins comprising transposase domains with N-terminal deletions and double cysteine-rich domains (CRDs), along with specific mutations for enhanced thermal stability and activity, and transposons with symmetrical inverted-end repeat sequences for site-specific integration.

Benefits of technology

The fusion proteins and transposons provide improved thermal stability and enhanced integration/excision efficiency, enabling precise genetic manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026517869000001_ABST
    Figure 2026517869000001_ABST
Patent Text Reader

Abstract

This disclosure generally relates to transposase domains, and more particularly to transposase domains including amino-terminal deletions, as well as transposase domains and DNA-targeting domains that form essential heterodimers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 501,233, filed May 10, 2023, and U.S. Provisional Patent Application No. 63 / 604,996, filed December 1, 2023, each of which is incorporated herein by reference in its entirety.

[0002] Reference to Electronically Submitted Sequence Listing This application includes a sequence listing submitted in XML format via the Patent Center, which is incorporated herein by reference in its entirety. The XML copy created on May 8, 2024, is named "POTH - 084_001WO_SeqList_ST26.xml" and is 271,738 bytes in size.

[0003] Field The present disclosure generally relates to transposase domains, particularly transposase domains containing a double - cysteine - rich domain (CRD), as well as transposase domains containing an N - terminal deletion, transposase domains forming an essential heterodimer, and fusion proteins comprising a transposase domain and a DNA - targeting domain. Methods of using the fusion proteins for site - specific translocation are also provided.

Background Art

[0004] Background Transposases can be used to introduce non - endogenous DNA sequences into genomic DNA and are advantageous over other gene - editing methods in many respects.

[0005] PiggyBac transposase consists of several protein domains. Binding to the ITR of the transposon is mediated by the DNA dimerization and binding domain (DDBD) together with the cysteine-rich C-terminal domain (CRD). Since the DDBD and CRD are also involved in protein dimerization, they play a dual role. Binding of the transposase dimer to the ITR positions the catalytic domain of the transposase at the TTAA cleavage site adjacent to the transposon. A second transposase dimer is thought to bind further along the ITR, distal to the TTAA. This dimer is not involved in catalyzing the transposase reaction, but its presence can stabilize the hairpin structure of the transposon in the transpososome.

[0006] PiggyBac transposase was originally isolated from the genome of the cabbage looper Trichoplusia ni. Since then, several point mutations have been identified that enhance the transposition activity of PiggyBac when used as a genome editing tool. For example, Super PiggyBac (SPB) transposase contains four point mutations, I30V, G165S, M282V and N538K. Another version of PiggyBac transposase, called "hyperactive PiggyBac transposase" or hyPBase, has been reported (Yusa et al. (2011), PNAS 108(4):1531-1536). HyPBase contains three additional mutations (S103P, S509G, N571S) in addition to the four point mutations (I30V, G165S, M282V, N538K) found in SPB.

[0007] Furthermore, a PiggyBac integration-deficient (or excision-only) mutant called "PBx" has been reported to contain two mutations (R372A, K375A). Converting a positively charged amino acid to an uncharged residue at this position results in a transposase that cannot interact with the DNA target. PBx can be constructed on top of SPB or HyPBase high-activity mutants and is integrated into ssSPB. In addition, in the case of PBx (but not SPB), the D450N mutation provides a further boost to excision activity. Alternative versions of PBx can be constructed by converting positions 372 and 375 to alternative amino acids, theoretically allowing for adjustment of the strength of interaction with target DNA. Using site-saturated mutagenesis, the R372H mutation was found to improve the integration efficiency of ssSPB.

[0008] However, the need for site-specific transposases for use in, for example, gene editing, remains unmet. This specification provides an improved version of the piggyBac transposase, including a dual CRD domain, point mutations that may improve the thermal stability of the protein, and additional high-activity mutations. [Overview of the Initiative]

[0009] summary In one embodiment, this specification provides a fusion protein comprising, in N-terminus to C-terminus, a nuclear localization signal (NLS), a DNA targeting domain, and a first transposase domain comprising any of the amino acid sequences of SEQ ID NOs. 28-69. In some embodiments, the first transposase domain comprises any of the amino acid sequences of SEQ ID NOs. 28-48. In some embodiments, the first transposase domain comprises any of the amino acid sequences of SEQ ID NOs. 49-69. In some embodiments, the first transposase domain comprises the amino acid sequence of SEQ ID NOs. 30 or 38. In some embodiments, the first transposase domain comprises the amino acid sequence of SEQ ID NOs. 51 or 59.

[0010] In some embodiments, the DNA targeting domain includes one or more zinc finger motifs. In some embodiments, the DNA targeting domain includes the sequence of Sequence ID No. 74. In some embodiments, the DNA targeting domain includes one or more TAL domains. In some embodiments, the DNA targeting domain binds to a nucleic acid sequence encoding GFP, a LINE1 repeat element, or a zinc finger 268 (ZFM268) binding site.

[0011] In another embodiment, this specification provides a fusion protein comprising (a) a TAL array and (b) a modified Super piggyBac transposase ("SPB") comprising an N-terminal deletion and a second cysteine-rich domain (CRD) fused to the C-terminus of the SPB transposase, wherein the C-terminus of the TAL array is fused to the N-terminal amino acid of the N-terminal deletion SPB to produce a TAL array-N-terminal deletion SPB fusion protein. In some embodiments, the fusion protein further comprises a GS or GGGS linker positioned between the TAL array and the N-terminal deletion SPB. In some embodiments, the SPB includes an N-terminal deletion containing the deletion of amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103.

[0012] In another embodiment, this specification provides a fusion protein comprising (a) a TAL array and (b) a modified Super piggyBac transposase ("SPB") comprising an N-terminal deletion, one or more integrated deletion PBx mutations, and a second cysteine-rich domain (CRD) fused to the C-terminus of the SPB transposase, wherein the C-terminus of the TAL array is fused to the N-terminal amino acid of the N-terminal deletion SPB to produce a TAL array-N-terminal deletion SPB fusion protein. In some embodiments, the fusion protein further comprises a GS or GGGS linker positioned between the TAL array and the N-terminal deletion SPB. In some embodiments, the modified Super piggyBac transposase comprises any of the amino acid sequences of SEQ ID NOs. 28-69.

[0013] In another embodiment, a fusion protein is provided herein, comprising, in N-terminus to C-terminus, a DNA targeting domain and a first transposase domain comprising the sequence shown in SEQ ID NO: 1 or 3, wherein the first transposase domain comprises the deletion of the 83-103 most N-terminal amino acids of SEQ ID NO: 1 or 73. In some embodiments, the DNA targeting domain comprises one or more zinc finger motifs. In some embodiments, the DNA targeting domain comprises one or more TAL domains. In some embodiments, the DNA targeting domain binds to nucleic acid sequences encoding GFP, zinc finger 268 (ZFM268), phenylalanine hydroxylase (PAH), beta-2-microglobulin (B2M), or a LINE1 repeat element.

[0014] In some embodiments, the first transposase domain and the DNA targeting domain are linked by a linker. In some embodiments, the linker includes the sequence GGGGS.

[0015] In some embodiments, the first transposase domain includes an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103. In some embodiments, the first transposase domain includes (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R, or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D for SEQ ID NO: 1 or 73, with the numbering beginning at the 12th residue of SEQ ID NO: 1 and the 1st residue of SEQ ID NO: 73.

[0016] In some embodiments, the fusion protein further comprises a second transposase domain at the C-terminus of the first transposase domain, the second transposase domain comprising the sequence shown in SEQ ID NO: 1 or 73. In some embodiments, the second transposase domain comprises deletions of N-terminal amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103 of SEQ ID NO: 1 or 73. In some embodiments, the second transposase domain includes (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R, or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D for SEQ ID NO: 1 or 73, with the numbering beginning at the 12th residue of SEQ ID NO: 1 and the 1st residue of SEQ ID NO: 73.

[0017] In another embodiment, polynucleotides comprising nucleic acid sequences encoding the fusion protein described herein are provided herein.

[0018] In another embodiment, a transposon comprising a symmetric left-end (LE) inverted-end repeat sequence (ITR) and a right-end (RE) inverted-end repeat sequence (ITR), wherein the nucleotide sequences of the LE ITR and RE ITR are each sequence number 6, is provided herein.

[0019] In another embodiment, a transposon is provided herein that comprises a symmetric left-end (LE) inverted-end repeat (ITR) and a right-end (RE) inverted-end repeat (ITR), wherein the nucleotide sequence of the LE ITR comprises SEQ ID NO: 6, and the nucleotide sequence of the RE ITR comprises SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 21, or SEQ ID NO: 98. In some embodiments, the transposon comprises a nucleotide sequence encoding a therapeutic protein. In some embodiments, the transposon comprises a promoter sequence that controls the expression of the therapeutic protein.

[0020] In another embodiment, a method for site-specific integration of a therapeutic gene into one or more genomic loci of a cell is provided herein, comprising co-introducing a transposon and a polynucleotide described herein into the cell.

[0021] In another embodiment, a method for modifying the genome of a cell is provided herein, comprising providing the cell with a fusion protein described herein, wherein the cell includes a modified binding site comprising, in the order 5' to 3', a reversed sequence of the target site relative to the DNA targeting domain, a first spacer, a TTAA target integration site for the SPB, a second spacer, and a complement of the sequence of the target site relative to the DNA targeting domain. [Brief explanation of the drawing]

[0022] [Figure 1]Figure 1 shows a schematic diagram of the dual reporter plasmid design used to confirm the excision and integration rates using each mutant transposon. When using the H-2kk GFP transposon reporter (Reporter 1), an increase in H2kk expression is observed when transposon excision increases. When using Reporter 2, an increase in GFP expression is observed when transposon integration increases. In an alternative design for Reporter 2, an increase in firefly luciferase expression is observed when transposon excision increases, and an increase in NanoLuc is observed when transposon integration increases.

[0023] [Figure 2] Figure 2 is a schematic diagram of a dual reporter plasmid design used to confirm the excision and integration rates of transposases containing dual CRD domains using the DDBD binding domain binding site (1 bp difference) of the RE ITR, using a transposon containing a wild-type RE ITR or a modified symmetric RE ITR (0 bp), or a modified symmetric RE ITR containing an additional 1 bp, 2 bp, 3 bp, or 3 bp separating the DDBD binding site and the CRD binding site, or a modified symmetric RE ITR containing an additional 3 bp separating the DDBD binding site and the CRD binding site.

[0024] [Figure 3] Figure 3 shows the excision and integration activity of TAL-ssSPB mutations, including thermal stability mutations.

[0025] [Figure 4] Figure 4 shows the TAL target sequence, TAL target length, and distance from the TTAA integration site of the structure described in Example 8. [Modes for carrying out the invention]

[0026] Detailed explanation This specification provides transposase domains and fusion proteins containing them, particularly transposase domains with N-terminal deletions and transposase domains containing a double C-terminal cysteine-rich domain (CRD). The fusion proteins containing the transposase domains may be further mutated to form essential heterodimers. Methods for producing transposase domains and fusion proteins, modified cells using the fusion proteins provided herein, and therapeutic methods using such cells are also provided.

[0027] The transposase domains provided herein may be, for example, wild-type transposase domains or embedded-deficient (excision only) transposase domains.

[0028] Fusion proteins comprising one or more transposase domains and DNA targeting domains are also provided herein. In some embodiments, the fusion protein further comprises a protein stabilization domain.

[0029] Transposons comprising symmetrical left-end (LE) inverted-end repeat sequences (ITRs) and right-end (RE) inverted-end repeat sequences (ITRs) are also provided herein.

[0030] Transposase domain containing a double cysteine-rich domain (CRD) The dimerization and DNA-binding domain (DDBD) allows the transposase to bind to the inverted end repeat (ITR) at the end of the transposon and is also involved in protein dimerization. The catalytic domain catalyzes the rearrangement reaction, while the insertion domain allows the transposase to interact with the target DNA into which the transposon is incorporated. At the C-terminus, the cysteine-rich domain (CRD) is attached to the rest of the protein by a linker approximately 20 amino acids long. Like the DDBD, the CRD is also involved in ITR binding and protein dimerization. When the PiggyBac transposase binds to the left-end (LE) and right-end (RE) ITRs, protein dimerization links the two ITRs to form a synaptic hairpin complex. This arrangement is necessary for the rearrangement to occur when the transposase trans-cleaves the adjacent TTAA sequence.

[0031] DDBD and CRD bind to ITR in a sequence-specific manner. DDBD interacts with approximately 10 bp of DNA located 6 bp inward from the TTAA sequence adjacent to the transposon. DDBD binding is symmetric, with one transposase monomer binding to the LE ITR and the second monomer binding to the RE ITR. The CRD domain of the first dimer binds to a 19 bp sequence on the LE ITR located immediately distal to the DDBD binding site. CRD binding is asymmetric, with both CRD domains of the first dimer interacting only with the 19 bp sequence on the LE ITR. The RE ITR contains the second DDBD binding sequence followed by a 19 bp CRD binding sequence starting 34 bp inward from the TTAA. The first dimer, which binds proximal to the TTAA, catalyzes the rearrangement reaction.

[0032] In one embodiment, a transposase domain comprising a second cysteine-rich domain (CRD) is provided herein. In some embodiments, the second CRD domain is ligated to the C-terminus to generate a transposase domain comprising a double CRD. In some embodiments, the transposase domain comprising a double CRD comprises an N-terminal deletion. In some embodiments, the transposase domain comprising a double CRD is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a highly active piggyBac transposase domain. In preferred embodiments, the transposase domain comprising a double CRD is a Super piggyBac® transposase domain (SPB). Non-limiting examples of SPB transposases are described in detail in U.S. Patents 6,218,182, 6,962,810, 8,399,643 and International Publication No. 2010 / 099296, each of which is incorporated herein by reference in whole with respect to examples of transposase domains that may be used in connection with the fusion proteins described herein.

[0033] In some embodiments, the transposase domain containing a double CRD is the Super PiggyBac transposase (SPB) domain. An exemplary wild-type SPB sequence with an NLS is shown in Sequence ID No. 1, with the NLS in italics, highly active mutations in bold, and the cysteine-rich domain (CRD) underlined. Sequence numbering of the SPB transposase domain to represent deletions and mutations begins at residue 12 in Sequence ID No. 1. TIFF2026517869000002.tif77170(Sequence ID 1)

[0034] An exemplary wild-type SPB sequence containing a second C-terminal CRD domain linked via the AGGG peptide linker sequence (SEQ ID NO: 27) is shown in SEQ ID NO: 3, with NLS in italics, highly active mutations in bold, linker sequences in lowercase, and each of the two cysteine-rich domains (CRDs) underlined. Sequence numbering of the SPB transposase domain to represent deletions and mutations begins at residue 12 of SEQ ID NO: 3. TIFF2026517869000003.tif85170 (Sequence ID 3)

[0035] The amino acid sequence of the SPB transposase domain, which includes a dual CRD domain without NLS, is shown in SEQ ID NO: 73. (Sequence ID 73) The transposase domains used in the fusion proteins described herein can be isolated or derived from insects, vertebrates, crustaceans, or urochordates, as further described in International Publication No. 2019 / 173636 and International Patent Application No. 2019 / 049816. In a preferred embodiment, the SPB transposase domain is isolated or derived from the insect Trichoplusia ni (GenBank accession number AAA87375) or the silkworm (GenBank accession number BAD11135).

[0036] In some embodiments, the transposase domain is an integration-deficient transposase domain. An integration-deficient transposase domain is a transposase that can excise the corresponding transposon but incorporates the excised transposon at a lower frequency than the corresponding wild-type transposase. Examples of integration-deficient transposases are disclosed in U.S. Patents 6,218,185, 6,962,810, 8,399,643 and International Publication No. 2019 / 173636, each of which is incorporated herein by reference in whole with respect to examples of transposase domains that may be used in connection with the fusion proteins described herein. A list of integration-deficient amino acid substitutions is disclosed in U.S. Patent 10,041,077, which is incorporated herein by reference in whole with respect to examples of mutations that may be introduced into the transposase domains described herein. Wild-type SPB can be de-integrated by introducing mutations, such as K93A, R372A, K375A, R376A, and / or D450N (numbering starts at residue 12 for SEQ ID NO: 1). The introduction of mutations R372A, K375A, R376A, and D450N is thought to de-integrate transposase but retain cleavage function. The amino acid sequence of the NLS-free de-integration PBx transposase domain is shown in SEQ ID NO: 70. (Sequence ID 70)

[0037] The amino acid sequence of the embedded / deleted PBx transposase domain, which does not contain NLS but includes a second CRD domain sequence (SEQ ID NO: 2) linked to the C-terminus of the PBx sequence via the AGGG linker sequence (SEQ ID NO: 27), is shown in SEQ ID NO: 71. (Sequence ID 71)

[0038] The amino acid sequence of the embedded / deleted PBx transposase domain, which does not contain NLS but includes a second CRD domain sequence (SEQ ID NO: 14) ligated to the C-terminus of the PBx sequence, is shown in SEQ ID NO: 72.

[0039] (Sequence ID 72)

[0040] Transposase domain including N-terminal deletion In some embodiments, double CRD transposase domains (e.g., SPB transposase domains or PBx transposase domains) are provided herein that include a partial deletion of the amino terminus (also called the "N terminus," "N-terminal domain," or "NTD") of the transposase domain. While we do not wish to be bound by theory, in the context of tandem dimer transposases (or dimers comprising the two fusion proteins described herein), it is conceivable that the N-terminal domain of the transposase (e.g., SPB) may introduce steric hindrance between the two dimers of a tandem dimer, or between two pairs of dimers, or between a dimer and DNA.

[0041] In some embodiments, the N-terminal deletion is approximately 20 amino acids, approximately 40 amino acids, approximately 60 amino acids, approximately 80 amino acids, approximately 100 amino acids, or approximately 115 amino acids. In some embodiments, the N-terminal deletion is approximately 15-25 amino acids, approximately 25-35 amino acids, approximately 35-45 amino acids, approximately 45-55 amino acids, approximately 55-65 amino acids, approximately 65-75 amino acids, approximately 75-85 amino acids, approximately 85-95 amino acids, approximately 95-105 amino acids, or approximately 105-120 amino acids.

[0042] In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-83 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-84 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-85 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-86 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-87 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-88 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-89 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-90 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-91 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-92 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-93 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-94 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-95 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-96 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-97 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-98 compared to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of amino acids 1-99 at the N-terminus compared to SEQ ID NOs. 71-73.In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-100 relative to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-101 relative to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-102 relative to SEQ ID NOs. 71-73. In some embodiments, the transposase domain includes deletions of N-terminal amino acids 1-103 relative to SEQ ID NOs. 71-73.

[0043] Exemplary sequences of the PBx transposase domain, including a double CRD domain with deletions of amino acids 1-93 at the N-terminus, are shown in SEQ ID NOs. 38 and 59. (SEQ ID NO: 38) (Sequence ID 59)

[0044] Exemplary sequences of the PBx transposase domain, including the N-terminal deletion and a second piggyBac CRD domain added via the AGGG linker, are shown in SEQ ID NOs. 28-48 of Table 1. [Table 1] TIFF2026517869000005.tif255170TIFF2026517869000006.tif255170TIFF2026517869000007.tif255170TIFF2026517869000008.tif137170

[0045] Further exemplary sequences of the PBx transposase domain, including the N-terminal deletion and the second piggyBac CRD domain (SEQ ID NO: 14), are shown in SEQ ID NOs: 49-69 in Table 2. [Table 2] TIFF2026517869000010.tif255170TIFF2026517869000011.tif255170TIFF2026517869000012.tif254170TIFF2026517869000013.tif118170

[0046] Transposase domains containing thermal stability mutations and / or high activity mutations Highly active mutations The transposase domains provided herein may contain one or more highly active mutants.

[0047] In some embodiments, the SPB transposase domain provided herein includes one or more highly active mutations (in addition to those present in SPB). In some embodiments, the SPB transposase domain provided herein includes the S103P mutation. In some embodiments, the SPB transposase domain provided herein includes the R372H mutation. In some embodiments, the SPB transposase domain provided herein includes the S509G mutation. In some embodiments, the SPB transposase domain provided herein includes the N571S mutation. In some embodiments, the SPB transposase domain provided herein includes the S103P mutation, the R372H mutation, the S509G mutation, and the N571S mutation.

[0048] In some embodiments, the PBx transposase domain provided herein includes one or more highly active mutations (in addition to those present in PBx). In some embodiments, the PBx transposase domain provided herein includes the S103P mutation. In some embodiments, the PBx transposase domain provided herein includes the R372H mutation. In some embodiments, the PBx transposase domain provided herein includes the S509G mutation. In some embodiments, the PBx transposase domain provided herein includes the N571S mutation. In some embodiments, the PBx transposase domain provided herein includes the S103P mutation, the R372H mutation, the S509G mutation, and the N571S mutation. In some embodiments, the PBx transposase domain including the R372H mutation includes the amino acid sequence shown in SEQ ID NO: 133. In some embodiments, the PBx transposase domain including the S103P, S509G, and N571S mutations includes the amino acid sequence shown in SEQ ID NO: 134. In some embodiments, the PBx transposase domain, including the S103P, S509G, N571S, and R372H mutations, contains the amino acid sequence shown in SEQ ID NO: 135.

[0049] The S103P, S509G, R372H, and / or N571S mutations can also be introduced into any of the cleavage-type SPBs or PBx (e.g., transposase domains containing any one of SEQ ID NOs. 28-69) provided herein. Those skilled in the art will understand that the residue numbering depends on the size of the cleavage.

[0050] The S103P, S509G, R372H, and / or N571S mutations can also be introduced into any of the transposase domains containing dual CRDs provided herein. For example, an SPB transposase domain containing the sequence shown in SEQ ID NOs: 1, 3, or 73 may further contain the S103P, S509G, R372H, and / or N571S mutations.

[0051] Similarly, a PBx transposase domain containing the sequence shown in any one of SEQ ID NOs: 70-72 may further contain the S103P, R372H, S509G, and / or N571S mutations. In some embodiments, a PBx domain containing a double CRD and the R372H mutation contains the amino acid sequence shown in SEQ ID NO: 138. In some embodiments, a PBx domain containing a double CRD and the S103P, S509G, and N571S mutations contains the amino acid sequence shown in SEQ ID NO: 139. In some embodiments, a PBx domain containing a double CRD and the S103P, S509G, N571S, and R372H mutations contains the amino acid sequence shown in SEQ ID NO: 140.

[0052] Similarly, the S103P, S509G, R372H, and / or N571S mutations may be introduced into any of the fusion proteins described below. For example, a fusion protein containing a transposase domain and a DNA-binding domain may further contain the S103P, S509G, R372H, and / or N571S mutations.

[0053] Thermal stability variation The transposase domains provided herein may contain one or more thermally stable mutations.

[0054] In some embodiments, the SPB transposase domains provided herein include one or more thermostable mutations. In some embodiments, the SPB transposase domains provided herein include one or more of the following mutations: I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D, M298L, or any subgroup thereof. In some embodiments, the SPB transposase domain provided herein includes the I182L mutation. In some embodiments, the SPB transposase domain provided herein includes the S301A mutation. In some embodiments, the SPB transposase domain provided herein includes the C420M mutation. In some embodiments, the SPB transposase domain provided herein includes the M185L mutation. In some embodiments, the SPB transposase domain provided herein includes the R315K mutation. In some embodiments, the SPB transposase domain provided herein includes the D421H mutation. In some embodiments, the SPB transposase domain provided herein includes the F200W mutation, and in some embodiments, the SPB transposase domain provided herein includes the Q318G mutation. In some embodiments, the SPB transposase domain provided herein includes the N427D mutation. In some embodiments, the SPB transposase domain provided herein includes the V207I mutation. In some embodiments, the SPB transposase domain provided herein contains the E331R mutation. In some embodiments, the SPB transposase domain provided herein contains the Q434E mutation. In some embodiments, the SPB transposase domain provided herein contains the M226F mutation.In some embodiments, the SPB transposase domain provided herein includes the V336I mutation. In some embodiments, the SPB transposase domain provided herein includes the V436I mutation. In some embodiments, the SPB transposase domain provided herein includes the I231T mutation. In some embodiments, the SPB transposase domain provided herein includes the S373K mutation. In some embodiments, the SPB transposase domain provided herein includes the I474L mutation. In some embodiments, the SPB transposase domain provided herein includes the V240K mutation. In some embodiments, the SPB transposase domain provided herein includes the V381E mutation. In some embodiments, the SPB transposase domain provided herein includes the K500R mutation. In some embodiments, the SPB transposase domain provided herein includes the Q254N mutation. In some embodiments, the SPB transposase domain provided herein includes the T392S mutation. In some embodiments, the SPB transposase domain provided herein includes the S513P mutation. In some embodiments, the SPB transposase domain provided herein includes the A263E mutation. In some embodiments, the SPB transposase domain provided herein includes the A411N mutation. In some embodiments, the SPB transposase domain provided herein includes the K525P mutation. In some embodiments, the SPB transposase domain provided herein includes the S289A mutation. In some embodiments, the SPB transposase domain provided herein includes the S419T mutation. In some embodiments, the SPB transposase domain provided herein includes the R567D mutation. In some embodiments, the SPB transposase domain provided herein includes the M298L mutation.

[0055] In some embodiments, the PBx transposase domains provided herein include one or more thermally stable mutations. In some embodiments, the PBx transposase domains provided herein include one or more of the following mutations: I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, 1231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D, M298L, or any subgroup thereof. In some embodiments, the PBx transposase domain provided herein includes the I182L mutation. In some embodiments, the PBx transposase domain provided herein includes the S301A mutation. In some embodiments, the PBx transposase domain provided herein includes the C420M mutation. In some embodiments, the PBx transposase domain provided herein includes the M185L mutation. In some embodiments, the PBx transposase domain provided herein includes the R315K mutation. In some embodiments, the PBx transposase domain provided herein includes the D421H mutation. In some embodiments, the PBx transposase domain provided herein includes the F200W mutation, and in some embodiments, the PBx transposase domain provided herein includes the Q318G mutation. In some embodiments, the PBx transposase domain provided herein includes the N427D mutation. In some embodiments, the PBx transposase domain provided herein includes the V207I mutation. In some embodiments, the PBx transposase domain provided herein contains the E331R mutation. In some embodiments, the PBx transposase domain provided herein contains the Q434E mutation. In some embodiments, the PBx transposase domain provided herein contains the M226F mutation.In some embodiments, the PBx transposase domain provided herein includes the V336I mutation. In some embodiments, the PBx transposase domain provided herein includes the V436I mutation. In some embodiments, the PBx transposase domain provided herein includes the I231T mutation. In some embodiments, the PBx transposase domain provided herein includes the S373K mutation. In some embodiments, the PBx transposase domain provided herein includes the I474L mutation. In some embodiments, the PBx transposase domain provided herein includes the V240K mutation. In some embodiments, the PBx transposase domain provided herein includes the V381E mutation. In some embodiments, the PBx transposase domain provided herein includes the K500R mutation. In some embodiments, the PBx transposase domain provided herein includes the Q254N mutation. In some embodiments, the PBx transposase domain provided herein includes the T392S mutation. In some embodiments, the PBx transposase domain provided herein includes the S513P mutation. In some embodiments, the PBx transposase domain provided herein includes the A263E mutation. In some embodiments, the PBx transposase domain provided herein includes the A411N mutation. In some embodiments, the PBx transposase domain provided herein includes the K525P mutation. In some embodiments, the PBx transposase domain provided herein includes the S289A mutation. In some embodiments, the PBx transposase domain provided herein includes the S419T mutation. In some embodiments, the PBx transposase domain provided herein includes the R567D mutation. In some embodiments, the PBx transposase domain provided herein includes the M298L mutation.

[0056] In some embodiments, the PBx transposase domain containing the M298L mutation includes the amino acid sequence shown in SEQ ID NO: 132.

[0057] The I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L mutations can also be introduced into any of the cleavage-type SPBs or PBxs provided herein (e.g., SEQ ID NOs. 28-69). Those skilled in the art will understand that the residue numbering depends on the size of the cleavage.

[0058] The I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L mutations can also be introduced into any of the transposase domains containing the dual CRDs provided herein. For example, the SPB transposase domain containing the sequence shown in SEQ ID NO: 1 or 3 may further include the I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L mutations (numbering starting from residue 12 of SEQ ID NO: 1 or 3).

[0059] Similarly, PBx transposase domains containing the sequence shown in any one of sequence numbers 70-72 may further contain the I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V4361, 1231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L mutations. In some embodiments, the PBx domain containing the dual CRD and the M298L mutation includes the amino acid sequence shown in SEQ ID NO: 137.

[0060] Similarly, the I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L mutations can be introduced into any of the fusion proteins listed below. For example, fusion proteins containing a transposase domain and a DNA-binding domain may further include the I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V4361, I23IT, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L mutations. LINE Targeted TAL Arrays and I182L, M185L, F200W, V207I, M226F, 1231T, V240K, Q254N, A263E, S289A, M298L, S301A, R315K, Q318G, E331R, V336I, S373K, V381E, T392S, A411N, S419T, C420M, D421H, N427D, Q434E, V436I, I474L, K500R, S513P The sequences of the fusion proteins containing the SPB transposase domain with the K525P and R567D mutations are described in SEQ ID NOs. 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, or 129, respectively.

[0061] A combination of highly active mutations and thermally stable mutations The highly active and thermally stable mutations described herein can be freely combined. Therefore, the transposase domain may contain one, two, or all of the highly active mutants S103P, R372H, S509G, and / or N571S, and any or all of the thermally stable mutants I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D, and / or M298L.

[0062] In some embodiments, the SPB transposase domain provided herein comprises the M298L heat-stable mutation and one or more of the S103P, R372H, S509G, and / or N571S high-activity mutations. In some embodiments, the SPB transposase domain provided herein comprises the M298L heat-stable mutation and one or more of the S103P, R372H, S509G, and N571S high-activity mutations. In some embodiments, the SPB transposase domain provided herein comprises the M298L heat-stable mutation and one or more of the S103P, R372H, S509G, and N571S high-activity mutations. In some embodiments, the PBx transposase domain provided herein comprises the M298L heat-stable mutation and one or more of the S103P, R372H, S509G, and / or N571S high-activity mutations. In some embodiments, the PBx transposase domains provided herein include the M298L thermally stable mutant and the S103P, S509G, and N571S highly active mutants. In some embodiments, the PBx transposase domains provided herein include the M298L thermally stable mutant and the S103P, R372H, S509G, and N571S highly active mutants.

[0063] In some embodiments, the PBx transposase domain, including the S103P, S509G, N571S, R372H, and M298L mutations, contains the amino acid sequence shown in SEQ ID NO: 136.

[0064] Any combination of highly active mutants and thermally stable mutants can also be introduced into any of the transposase domains containing dual CRDs provided herein. For example, an SPB transposase domain containing the sequence shown in SEQ ID NO: 1 or 3 may contain one or more of the highly active mutants S103P, R372H, S509G and N571S, and I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V33 It may further include one or more of the following thermostable mutations: 6I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L (numbering begins at residue 12 of SEQ ID NO: 1 or 3).

[0065] Similarly, PBx transposase domains containing the sequence shown in any one of sequence numbers 70-72 contain one or more of the highly active mutants S103P, R372H, S509G and N571S, and I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E33 It may further include one or more of the following thermostable mutations: 1R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L. In some embodiments, the PBx domain, which includes a double CRD and the S103P, S509G, N571S, R372H and M298L mutations, contains the amino acid sequence shown in SEQ ID NO: 141.

[0066] Similarly, any combination of highly active and thermally stable mutations can be introduced into any of the fusion proteins described below. For example, a fusion protein containing a transposase domain and a DNA-binding domain may include one or more of the highly active mutations S103P, R372H, S509G, and N571S, and I182L, S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q It may further include one or more of the thermal stability variants 434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L.

[0067] Fusion protein containing a transposase domain Fusion proteins comprising one or more transposase domains described herein are also provided herein.

[0068] In some embodiments, fusion proteins comprising an SPB or PBx domain containing a second C-terminal CRD domain and a DNA targeting domain are provided herein. The DNA targeting domain will be described in more detail below. In some embodiments, fusion proteins comprising an N-terminal deletion SPB or PBx domain containing a second C-terminal CRD domain and a DNA targeting domain and a protein stabilization domain (PSD) are provided herein. The PSD will be described in more detail below.

[0069] In some embodiments, the fusion protein provided herein comprises, in the order of N-terminus to C-terminus, a PSD, a DNA targeting domain, an N-terminal deletion, and a transposase domain including a double CRD domain.

[0070] DNA targeting domain The transposase domains and fusion proteins provided herein may further comprise one or more DNA targeting domains. The DNA targeting domains may be bound to the C-terminus or N-terminus of the transposase domain of the fusion protein. In preferred embodiments, the DNA targeting domain is bound to the N-terminus of the transposase domain, for example, a transposase domain containing an N-terminal deletion. While we do not wish to be bound by theory, it is believed that the addition of a DNA targeting domain to a transposase domain improves site-specific transposase activity by targeting the transposase fused to the DNA targeting domain to a target site. In some embodiments, insertion of a DNA targeting domain improves site-specific transposase activity by at least twofold, at least threefold, at least fourfold, or at least fivefold compared to the same transposase domain without the DNA targeting domain.

[0071] Any DNA targeting domain known in the art may be used in connection with the transposase domains, fusion proteins, and tandem dimer transposases described herein, but is not limited to CRISPR, zinc finger motifs, TALE, and transcription factors. In some embodiments, the DNA targeting domain includes one, two, or three zinc finger motifs. In some embodiments, the three zinc finger motifs are adjacent to the GGGGS linker. In some embodiments, the three zinc finger motifs adjacent to the GGGGS linker are arranged in the following sequence shown in Sequence ID No. 74: The sequence comprises GGGGSERPYACPVESCDRRFSRSDELTRHIRIHTGQKPFQCRICMRNFSRSDHLTTHIRTHTGEKPFACDICGRKFARSDERKRHTKIHLRQKDGGGGS (Sequence ID 74) or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity with it.

[0072] In certain embodiments, fusion proteins comprising an N-terminal deletion, an NLS, and a transposase domain containing three zinc finger motifs are provided herein. In some embodiments, the NLS comprises or comprises the sequence shown in SEQ ID NO: 88.

[0073] In some embodiments, the DNA targeting domain is a TAL array. Xanthomonas-derived TALEs (transcriptional activator-like effectors) typically contain a 288-amino acid N-terminus, followed by an array of approximately 34 amino acid repeats (variable number), and then a 278-amino acid C-terminus, although cleavage types are documented in the literature (see, e.g., Miller et al., Nat Biotechnol 29, 143-148 (2011)). TALs fused to FokI nucleases (called TALENs) most often contain cleavage at both the N-terminus and the C-terminus. For example, the first 152 amino acids of the N-terminus are often removed, leaving 136 amino acids (called delta-152, SEQ ID NO: 75), and the C-terminus is often cleaved, leaving 63 amino acids (called +63, SEQ ID NO: 76).

[0074] TAL contains an array of 34 amino acids that are repeated a variable number of times. By changing two amino acids at positions 12 and 13, we investigate which nucleotides the TAL repeat recognizes. This feature allows us to program the TAL array to bind to specific DNA sequences. Amino acids NG recognize T, NI recognize A, NN recognize G or A, HD recognize C, NK recognize G, and NS recognize A, C, G, or T. When the two amino acids at positions 12 and 13 are replaced with NS, NA, or a single S, the TAL module recognizes any of the four nucleotides without distinction. When the two amino acids at positions 12 and 13 are replaced with a single N, this module can correspond to the recognition of 5-methylcytosine (5mC), for example, as it occurs in CpG sequences in the mammalian genome. Other amino acids within the 34-residue repeat can also be changed. For example, position 11 is often changed to N for repeats that recognize G. Also, positions 4 and 32 are often changed not to investigate binding specificity, but to reduce the repeatability of the array. The number of 34-amino acid repeats in the array determines the length of the recognized DNA sequence (one protein repeat binds to one DNA bp). Furthermore, the last bp is recognized by a "half-array" which has 20 amino acids instead of 34.

[0075] Furthermore, the N-terminal domain of TAL (e.g., SEQ ID NO: 75) recognizes and requires a T directly located at 5' of the target DNA sequence. Mutations of the TAL N-terminal domain have been documented in the literature and no longer require the 5'T (Lamb et al., Nucleic Acids Res. 2013 Nov;41(21):9779-85.doi:10.1093 / nar / gkt754.Epub 2013 Aug 26.PMID:23980031;PMCID:PMC3834825.). For example, the NT-G variant requires 5'G instead of 5'T (SEQ ID NO: 77), while the NT-PN variant does not require any specific 5' nucleotide (SEQ ID NO: 78). These variant N-terminal domain sequences can be used to provide further sequence options that can be targeted using TAL arrays.

[0076] Exemplary TAL modules are shown in Sequence IDs 79-82, where X is any amino acid. TAL Module Version 1: LTPDQVVAIAXXXGGKQALETVQRLLPVLCQDHG (Sequence ID 79) TAL Module Version 2: LTPEQVVAIAXXXGGKQALETVQRLLPVLCQAHG (Sequence ID 80) ·TAL Module · Version 3”LTPDQVVAIAXXXGGKQALETVQRLLPVLCQAHG (Sequence ID 81) TAL Module Version 4: LTPAQVVAIAXXXGGKQALETVQRLLPVLCQDHG (Sequence ID 82)

[0077] An exemplary TAL half-module is shown in Sequence ID No. 83, where X is any amino acid. LTPEQVVAIAXXXGGRPALE

[0078] Pairs of TAL arrays targeting sequences within desired genes can be designed, corresponding modules selected, and pooled together using "Golden Gate Assembly" to assemble each TAL array in-frame. Alternatively, the nucleotide sequences encoding the assembled TAL arrays can be synthesized de novo. The DNA sequences encoding the TAL arrays generated herein can be further codon-optimized using the GeneArt algorithm (Thermo Fisher).

[0079] When designing left and right TAL arrays containing an N-terminal domain that recognizes T and a TAL C-terminal domain that fuses to an N-terminal deletion transposase sequence (i.e., TAL-ssSPB or TAL-PBx, described in detail below), one TAL array recognizes the 5' sequence of TTAA, and the other TAL array recognizes the 3' sequence of TTAA. Since the 5' sequence of TTAA is almost always different from the 3' sequence of TTAA in the genomic DNA target, TAL-ssSPB is most often used as a heterodimer consisting of two different TAL domains that recognize two different DNA sequences. Furthermore, the sequence recognized by the TAL array is not directly adjacent to TTAA. Instead, it is separated from TTAA by a spacer of a given bp length, e.g., a 12bp, 13bp, or 14bp spacer.

[0080] TAL arrays can target any desired DNA sequence (e.g., genomic DNA sequence). It will be apparent to those skilled in the art that any left TAL array for a given target can be combined with any right TAL array for the same target.

[0081] In some embodiments, the TAL array targets ZFN268. Exemplary sequences of TAL arrays functioning as left and right arrays targeting ZFN268 are shown in Sequence ID No. 84. In some embodiments, the TAL array targeting ZFN268 binds to a nucleic acid molecule containing the sequence shown in Sequence ID No. 85.

[0082] In some embodiments, the TAL array targets the LINE1 repeat element. An exemplary sequence of the left TAL array targeting the LINE1 repeat element is shown in Sequence ID 89.

[0083] LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQ DHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGRPALE (SEQ ID NO: 89)

[0084] An exemplary sequence of the right TAL array targeting LINE1 is shown in SEQ ID NO: 90. LTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIANNNGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQ DHGLTPEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGRPALE (Sequence number 90)

[0085] The DNA targeting domain may be fused to or ligated to the N-terminus of a transposase domain, including an N-terminal deletion. For example, the DNA targeting domain may be inserted into the transposase domain at an appropriate location in the N-terminal region of the transposase domain.

[0086] The DNA targeting domain can be inserted at the N-terminus of the transposase domain. In some embodiments, the DNA targeting domain is inserted between the 82nd and 83rd amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 83rd and 84th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 84th and 85th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 85th and 86th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 86th and 87th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 87th and 88th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 88th and 89th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 89th and 90th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 90th and 91st amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 91st and 92nd amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 92nd and 93rd amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 93rd and 94th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 94th and 95th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 95th and 96th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 96th and 97th amino acids of one of the sequence numbers 71-73.In some embodiments, the DNA targeting domain is inserted between the 97th and 98th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 98th and 99th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 99th and 100th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 100th and 101st amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 101st and 102nd amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 102nd and 103rd amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 103rd and 104th amino acids of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain is inserted between the 104th and 105th amino acids of one of sequence numbers 71-73. In some embodiments, the DNA targeting domain includes the sequence of sequence number 74, or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto. In some embodiments, the DNA targeting domain includes TAL. The transposase domain may further include NLS.

[0087] In some embodiments, the DNA targeting domain substitutes the 83rd amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 84th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 85th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 86th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 87th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 88th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 89th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 90th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 91st amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 92nd amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 93rd amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 94th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 95th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 96th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 97th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 98th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 99th amino acid of one of the sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 100th amino acid of one of the sequence numbers 71-73.In some embodiments, the DNA targeting domain substitutes the 101st amino acid of one of sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 102nd amino acid of one of sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 103rd amino acid of one of sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 104th amino acid of one of sequence numbers 71-73. In some embodiments, the DNA targeting domain substitutes the 105th amino acid of one of sequence numbers 71-73. In some embodiments, the DNA targeting domain includes the sequence of sequence number 74, or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto. In some embodiments, the DNA targeting domain includes TAL. The transposase domain may further include NLS.

[0088] Sequence ID 86 shows an exemplary sequence of a fusion protein containing a dual CRD domain, a transposase domain with a 93-amino acid N-terminal deletion, an NLS, and three zinc finger motifs adjacent to the GGGGS linker, with the NLS shown in italics, the sequence containing the three zinc finger motifs and the GGGGS linker underlined, the transposase domain with a 93-amino acid N-terminal deletion in bold, the CRD linker sequence in lowercase, and the second CRD domain (Sequence ID 2) in bold italics. TIFF2026517869000014.tif85170 (Sequence ID 86)

[0089] Sequence ID 87 shows an exemplary sequence of a fusion protein containing a dual CRD domain, an embedded-deficient transposase domain with a 93-amino acid N-terminal deletion, an NLS, and three zinc finger motifs adjacent to the GGGGS linker, with the NLS shown in italics, the sequence containing the three zinc finger motifs and the GGGGS linker underlined, the transposase domain with the 93-amino acid N-terminal deletion in bold, the linker sequence in lowercase, and the second CRD domain in bold italics. TIFF2026517869000015.tif84170 (Sequence ID 87)

[0090] Exemplary sequences of the fusion protein, including a dual CRD domain, an embedded / defective transposase domain with a 93-amino acid N-terminal deletion, an NLS, and a LINE1 left L2 TAL array, are shown in SEQ ID NOs: 17 and 19.

[0091] Exemplary sequences of the fusion protein, including a dual CRD domain and an embedded deficient transposase domain with a 93-amino acid N-terminal deletion, NLS, and LINE1 right TAL array, are shown in SEQ ID NOs: 18 and 20.

[0092] Protein stabilization domain In some embodiments, the fusion proteins provided herein may further include a protein stabilization domain (PSD). The PSD, if present, is preferably bound to the N-terminus of the DNA targeting domain. While we do not wish to be bound by theory, it is thought that the addition of a PSD may improve protein stability or the stability of the transposase tetramer-DNA complex.

[0093] The PSD can be approximately the same size as the N-terminal deletion of the transposase domain. For example, in some embodiments, the N-terminal deletion of the transposase domain contains amino acids 1-93, while the PSD contains 92 amino acids.

[0094] In some embodiments, the PSD contains amino acids 1-83 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-83 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-84 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-84 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-85 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-85 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-86 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-86 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-87 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-87 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-88 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-88 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-89 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-89 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-90 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-90 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-91 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-91 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-92 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-92 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-93 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-93 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-94 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-94 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-95 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-95 of SEQ ID NO: 71 or 72.In some embodiments, the PSD contains amino acids 1-96 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-96 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-97 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-97 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-98 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-98 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-99 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-99 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-100 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-100 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-101 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-101 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-102 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-102 of SEQ ID NO: 71 or 72. In some embodiments, the PSD contains amino acids 1-103 of SEQ ID NO: 73. In some embodiments, the PSD contains amino acids 1-103 of SEQ ID NO: 71 or 72.

[0095] In some embodiments, the PSD is an array Includes GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRG (Sequence ID 91).

[0096] Accordingly, this specification provides fusion proteins comprising, in order from the N-terminus to the C-terminus, a nuclear localization signal (NLS), a PSD, a DNA targeting domain, and a transposase domain including an N-terminal deletion, compared to the sequences shown in Sequence IDs 71-73.

[0097] Exemplary sequences of fusion proteins containing a PSD, NLS, DNA targeting domain, and transposase domains including an N-terminal deletion and a second C-terminal CRD domain are shown in SEQ ID NO: 92 (PBx transposase domain) and SEQ ID NO: 93 (SPB transposase domain), where the NLS (here, PKKKRKV) is shown in italics, the PSD in bold and underlined, the DNA targeting domain (here, three zinc finger motifs adjacent to the GGGGS linker) in underlined, the N-terminal deletion transposase domain (here, PBx) in bold, the linker sequence in lowercase, and the second CRD domain in bold and italicized. TIFF2026517869000016.tif99170 (Sequence ID 92) TIFF2026517869000017.tif98170 (Sequence ID 93)

[0098] Nuclear localization signals In some embodiments, the transposase domains and fusion proteins provided herein may include an in-frame nuclear localization sequence (NLS). Examples of transposases fused to nuclear localization signals are disclosed in U.S. Patents 6,218,185, 6,962,810, 8,399,643 and International Publication No. 2019 / 173636. In some embodiments, the NLS includes the sequence PKKKRKV (SEQ ID NO: 88). In certain embodiments, the in-frame NLS is located upstream (N-terminus) of a transposase domain including an N-terminal deletion.

[0099] Generally, the NLS is preferably located at the N-terminus of the fusion protein. In some embodiments, the NLS is fused to or ligated to the N-terminus of the transposase domain. In some embodiments, the NLS is fused to or ligated to the N-terminus of the DNA targeting domain. In some embodiments, the NLS is fused to or ligated to the N-terminus of the PSD.

[0100] In certain embodiments, the in-frame NLS is directly fused to the amino terminus of the transposase domain containing the N-terminal deletion. In some embodiments, the NLS is bound to the N-terminus of the transposase domain containing the N-terminal deletion via a linker (e.g., a GGGGS linker or a GGS linker).

[0101] In some embodiments, the initiating methionine is introduced before the NLS. In some embodiments, additional alanine residues are introduced before and / or after the NLS to ensure in-frame translation. Thus, the residue numbering in SEQ ID NO: 1 begins at the 12th residue of SEQ ID NO: 1 for the purpose of identifying deleted and mutated residues. In SEQ ID NOs: 71-73, which are sequences of SPB and PBx version 1 and PBx version 2, respectively, containing a second CRD domain that does not contain the NLS, the residue numbering begins at the 1st residue for the purpose of identifying deleted and mutated residues.

[0102] Essential heterodimer In another embodiment, an essential heterodimeric transposase comprising two fusion proteins is provided herein, each fusion protein comprising a DNA targeting domain and a transposase domain. In some embodiments, both fusion proteins comprise a DNA targeting domain, which targets and binds to a DNA sequence adjacent to the DNA sequence that is the insertion site targeted by the transposase. The DNA targeting domain may bind to the C-terminus or N-terminus of the fusion protein.

[0103] Accordingly, in some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a first DNA targeting domain, a linker and a first transposase domain, and (b) a second fusion protein comprising a DNA targeting domain, a linker and a second transposase domain, wherein the first DNA targeting domain and the second DNA targeting domain are different, and the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposing charges that enable the two fusion proteins to form a complex.

[0104] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising, in the order of N-terminus to C-terminus, a first NLS, a first DNA targeting domain, and a first transposase domain including an N-terminal deletion; and (b) a second fusion protein comprising, in the order of N-terminus to C-terminus, a second NLS, a second DNA targeting domain, and a second transposase domain including an N-terminal deletion, wherein the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposing charges that enable the two fusion proteins to form a complex. In some embodiments, the first and / or second transposase domains are SPB domains. In some embodiments, the first and / or second transposase domains are PBx transposase domains. In some embodiments, the first and / or second transposase domains include an N-terminal deletion at amino acids 83, 84, 85, 86, 87, 88, 89, 90, 91, 21, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103. In some embodiments, the first and second transposase domains include the sequence of SEQ ID NO: 38 or 59. In some embodiments, the first and / or second DNA targeting domains include one or more zinc finger motifs. In some embodiments, the first and / or second DNA targeting domains include one zinc finger motif. In some embodiments, the first and / or second DNA targeting domains include two zinc finger motifs. In some embodiments, the first and / or second DNA targeting domains include three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domains include the sequence of SEQ ID NO: 74. In some embodiments, the first and / or second DNA targeting domain includes one or more TAL motifs.

[0105] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising, in the order N-terminus to C-terminus, a first NLS, a first PSD, a first DNA targeting domain, and a first transposase domain including an N-terminal deletion; and (b) a second fusion protein comprising, in the order N-terminus to C-terminus, a second NLS, a second PSD, a second DNA targeting domain, and a second transposase domain including an N-terminal deletion, wherein the transposase domains of the first fusion protein and the transposase domain of the second fusion protein have opposing charges that enable the two fusion proteins to form a complex. In some embodiments, the first and / or second transposase domains are SPB domains. In some embodiments, the first and / or second transposase domains are PBx transposase domains. In some embodiments, the first and second transposase domains comprise the sequence of Sequence ID No. 38 or 59. In some embodiments, the first and / or second PSD includes the sequence of SEQ ID NO: 91. In some embodiments, the first and / or second DNA targeting domain includes three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domain includes the sequence of SEQ ID NO: 74. In some embodiments, the first and / or second DNA targeting domain includes a TAL motif.

[0106] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a first NLS and a first transposase domain comprising the sequences of SEQ ID NOs. 71-73 in N-terminal to C-terminal order, and (b) a second fusion protein comprising a second NLS and a second transposase domain comprising the sequences of SEQ ID NOs. 71-73 in N-terminal to C-terminal order, wherein the first and second transposase domains comprise DNA targeting domains, and the transposase domains of the first and second fusion proteins have opposing charges that enable the two fusion proteins to form a complex. In some embodiments, the first and / or second transposase domains are SPB domains. In some embodiments, the first and / or second transposase domains are PBx transposase domains. In some embodiments, the first and second transposase domains comprise the sequences of SEQ ID NOs. 71-73. In some embodiments, the first and / or second PSD includes the sequence of SEQ ID NO: 91. In some embodiments, the first and / or second DNA targeting domain includes three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domain includes the sequence of SEQ ID NO: 74. In some embodiments, the first and / or second DNA targeting domain includes a TAL motif. In some embodiments, the first DNA targeting domain replaces residues 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103 of the first transposase domain of SEQ ID NOs: 71-73. In some embodiments, the second DNA targeting domain replaces the 83rd, 84th, 85th, 86th, 87th, 88th, 89th, 90th, 91st, 92nd, 93rd, 94th, 95th, 96th, 97th, 98th, 99th, 100th, 101st, 102nd, or 103rd residue of the second transposase domain of SEQ ID NOs.

[0107] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a first transposase domain, a linker, and a first DNA targeting domain, wherein the transposase domain of the first fusion protein comprises the same amino acid sequence as shown in any one of SEQ ID NOs. 28 to 69, and (b) a second fusion protein comprising a first transposase domain, a linker, and a second DNA targeting domain, wherein the transposase domain of the second fusion protein comprises the same amino acid sequence as shown in any one of SEQ ID NOs. 28 to 69.

[0108] In another embodiment, fusion proteins comprising a transposase domain are provided herein that can form an essential heterodimer with another fusion protein comprising a transposase domain. While not wishing to be bound by theory, it is conceivable that two such fusion proteins would assemble to form an essential heterodimer structure linked by a combination of charge interactions, hydrogen bonds, pi-cation pairs, and hydrophobic interactions. Such an essential heterodimer structure is referred herein to as an "essential heterodimer transposase." Thus, each essential heterodimer transposase comprises two transposase domains. In some embodiments, the two fusion proteins provided herein form a complex comprising (a) a first fusion protein comprising a transposase domain and (b) a second fusion protein comprising a transposase domain, wherein the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposing charges that enable the two fusion proteins to form a complex.

[0109] By introducing charged residues to amino acids that contribute to dimerization with a second fusion protein, it is possible to design a pair of fusion proteins that associate only with each other to form an essential heterodimer in a predetermined configuration. By introducing mutations that enable only one configuration of the essential heterodimer, it becomes possible to introduce a DNA targeting domain into the fusion protein, resulting in increased specificity of the transposase domain. Introducing a DNA targeting domain into a fusion protein that can dimerize in any configuration, including homodimerization, yields two copies of the same DNA targeting domain present in the dimeric transposase. However, only one of these DNA targeting domains interacts with DNA, while the other potentially sterically interferes with the transposase-DNA interaction. Any suitable DNA targeting domain described herein or known in the art may be used in the fusion proteins described herein.

[0110] Those skilled in the art will be able to easily identify mutations in transposase domains that confer positive or negative charges. For fusion proteins containing transposase domains, the crystal structure published by Chen et al. (Nat Commun 11, 3446 (2020)) can be used to identify pairs of transposase domain residues adjacent to the dimer formed by two such fusion proteins. Creating positively charged and negatively charged transposase domains by altering the charge of such residue pairs can be achieved using standard techniques such as site-directed mutagenesis.

[0111] For example, one or more of M185, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593, and / or F594 can be mutated with an SPB transposase domain (for example, an SPB shown in SEQ ID NO: 1 or 73, with numbering starting from the 12th residue of SEQ ID NO: 1 and the 1st residue of SEQ ID NO: 73) to generate an SPB-transposase domain or an SPB+transposase domain. Similarly, one or more of M185, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593, and / or F594 can be mutated in the PBx transposase domain (for example, in the PBx transposase domains of SEQ ID NOs. 70-72 to generate PBx- or PBx+ transposase domains).

[0112] To achieve the formation of essential heterodimers, a pair of mutations can be introduced into the fusion protein or transposase domain to generate positively and negatively charged fusion protein or transposase domains that can then interact to form heterodimers. In some embodiments, the mutated residue pairs are shown in Table 3. For example, one or more mutations listed in the column labeled "Protein 1" may be introduced into the first SPB or PBx domain, and one or more corresponding mutations listed in the column labeled "Protein 2" may be introduced into the second SPB or PBs domain. In some embodiments, members of the residue pair are mutated to have opposite charges. [Table 3]

[0113] To introduce a positive charge, amino acids with uncharged side chains, such as methionine, or amino acids with negatively charged side chains, such as aspartic acid, can be replaced with positively charged amino acids, such as lysine or arginine. To introduce a negative charge, amino acids with positively charged side chains, such as arginine or lysine, or amino acids with hydrophobic side chains, such as leucine, can be replaced with negatively charged amino acids, such as aspartic acid or glutamic acid.

[0114] In certain embodiments, to generate an SPB+ fusion protein, one or more of the following mutations are introduced into one or both of the SPB transposase domains of the fusion protein provided herein (e.g., SPBs shown in SEQ ID NO: 1 or 73, numbering beginning at the 12th residue of SEQ ID NO: 1 and the 1st residue of SEQ ID NO: 73): M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, the SPB+ transposase domain includes the M185R mutation and the D198K mutation. In some embodiments, the SPB+ transposase domain includes the M185R mutation and the D201R mutation. In some embodiments, the SPB+ transposase domain includes the D197K mutation and the D201R mutation. In some embodiments, the SPB+ transposase domain includes the D198K mutation and the D201R mutation. In some embodiments, the SPB+ transposase domain includes the M185R mutation, the D198K mutation, and the D201R mutation.

[0115] In certain embodiments, to generate a PBx+ fusion protein, one or more of the following mutations are introduced into one or both of the PBx transposase domains of the fusion protein provided herein (e.g., the PBx transposase domains of SEQ ID NOs. 70-72): M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, the PBx+ transposase domain includes the M185R mutation and the D198K mutation. In some embodiments, the PBx+ transposase domain includes the M185R mutation and the D201R mutation. In some embodiments, the PBx+ transposase domain includes the D197K mutation and the D201R mutation. In some embodiments, the PBx+ transposase domain includes the D198K mutation and the D201R mutation. In some embodiments, the PBx+ transposase domain includes the M185R mutation, the D198K mutation, and the D201R mutation.

[0116] In certain embodiments, to generate an SPB-fusion protein, one or more of the following mutations are introduced into one or both of the SPB transposase domains of the fusion protein provided herein (e.g., SPBs shown in SEQ ID NO: 1 or 73, numbering beginning at the 12th residue of SEQ ID NO: 1 and the 1st residue of SEQ ID NO: 73): L204D, L204E, K500D, K500E, R504E, and R504D. In some embodiments, the SPB-transposase domain includes the L204E mutation and the K500D mutation. In some embodiments, the SPB-transposase domain includes the L204E mutation and the R504D mutation. In some embodiments, the SPB-transposase domain includes the K500 mutation and the R504D mutation. In some embodiments, the SPB-transposase domain includes the L204E mutation, the K500D mutation, and the R504D mutation.

[0117] In certain embodiments, to generate a PBx-fusion protein, one or more of the following mutations are introduced into one or both of the PBx transposases of the fusion proteins provided herein (e.g., the PBx transposase domains of SEQ ID NOs. 70-72): L204D, L204E, K500D, K500E, R504E, and R504D. In some embodiments, the PBx-transposase domain includes the L204E and K500D mutations. In some embodiments, the PBx-transposase domain includes the L204E and R504D mutations. In some embodiments, the PBx-transposase domain includes the K500 and R504D mutations. In some embodiments, the PBx-transposase domain includes the L204E, K500D, and R504D mutations.

[0118] SPB+, SPB-, PBx+, and PBx- fusion proteins and transposase domains may further comprise N-terminal deletions of the transposase domains described herein. Thus, in some embodiments, SPB+ fusion proteins comprising first and second SPB+ transposase domains are provided herein, where the SPB+ transposase domain comprises N-terminal deletions of approximately 20 amino acids, approximately 40 amino acids, approximately 60 amino acids, approximately 80 amino acids, approximately 100 amino acids, or approximately 115 amino acids. In some embodiments, the SPB+ transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the SPB+ transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the SPB+ transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the SPB+ transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the SPB+ transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 88 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 89 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 90 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 91 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 95 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 96 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 97 amino acids. In some embodiments, the SPB+ transposase domain includes a 98-amino acid N-terminal deletion.In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 99 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 100 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 101 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 102 amino acids. In some embodiments, the SPB+ transposase domain includes an N-terminal deletion of 103 amino acids.

[0119] In some embodiments, SPB-fusion proteins comprising an SPB-transposase domain are provided herein, the SPB-transposase domain comprising an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 81 amino acids, about 82 amino acids, about 83 amino acids, about 84 amino acids, about 85 amino acids, about 86 amino acids, about 87 amino acids, about 88 amino acids, about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the SPB-transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the SPB-transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 85 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 86 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 87 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 88 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 89 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 90 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 91 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the SPB-transposase domain includes a 95-amino acid N-terminal deletion. In some embodiments, the SPB-transposase domain includes a 96-amino acid N-terminal deletion. In some embodiments, the SPB-transposase domain includes a 97-amino acid N-terminal deletion.In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 98 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 99 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 100 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 101 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 102 amino acids. In some embodiments, the SPB-transposase domain includes an N-terminal deletion of 103 amino acids.

[0120] In some embodiments, PBx+ fusion proteins comprising a PBx+ transposase domain are provided herein, the PBx+ transposase domain comprising an N-terminal deletion of approximately 20 amino acids, approximately 40 amino acids, approximately 60 amino acids, approximately 80 amino acids, approximately 100 amino acids, or approximately 115 amino acids. In some embodiments, the PBx+ transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the PBx+ transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the PBx+ transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the PBx+ transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the PBx+ transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the PBx+ transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the PBx+ transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 90 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 91 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 95 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 96 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 97 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 98 amino acids. In some embodiments, the PBx+ transposase domain includes an N-terminal deletion of 99 amino acids. In some embodiments, the PBx+ transposase domain includes a 100-amino acid N-terminal deletion.In some embodiments, the PBx+ transposase domain includes a 101-amino acid N-terminal deletion. In some embodiments, the PBx+ transposase domain includes a 102-amino acid N-terminal deletion. In some embodiments, the PBx+ transposase domain includes a 103-amino acid N-terminal deletion.

[0121] In some embodiments, PBx-fusion proteins comprising a PBx-transposase domain are provided herein, the PBx-transposase domain comprising an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 81 amino acids, about 82 amino acids, about 83 amino acids, about 84 amino acids, about 85 amino acids, about 86 amino acids, about 87 amino acids, about 88 amino acids, about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the PBx-transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the PBx-transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 85 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 86 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 87 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 88 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 89 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 90 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 91 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the PBx-transposase domain includes a 95-amino acid N-terminal deletion. In some embodiments, the PBx-transposase domain includes a 96-amino acid N-terminal deletion. In some embodiments, the PBx-transposase domain includes a 97-amino acid N-terminal deletion.In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 98 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 99 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 100 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 101 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 102 amino acids. In some embodiments, the PBx-transposase domain includes an N-terminal deletion of 103 amino acids.

[0122] Transposons containing symmetrically reversed terminal repeat sequences (ITRs) Transposons including left-end (LE) and right-end (RE) symmetric ITRs described herein are also provided herein.

[0123] SPB transposases bind asymmetrically to 35 bp LE ITRs (SEQ ID NO: 4) and 63 bp RE ITRs (SEQ ID NO: 5), while double CRD SPBs bind symmetrically as dimers to transposons that have LE ITR sequences at both ends. Transposons containing a second copy of the 35 bp LE ITR instead of the RE ITR are necessary for transposon recognition, as transposons are not well recognized when using piggyBac transposases containing double CRD domains for wild-type ITRs (see Example 3). An exemplary symmetric RE ITR is referred to as symmetric ITR 0 bp (SEQ ID NO: 6). An exemplary symmetric RE ITR containing a DDBD binding site for the RE ITR is referred to as symmetric ITR SNP (SEQ ID NO: 98).

[0124] In some embodiments, the transposon symmetry RE ITR includes an additional base pair within the RE ITR between a 10 bp DDBD binding sequence (SEQ ID NO: 11) and a CRD binding sequence. In some embodiments, each symmetry ITR includes one additional base pair: CCCTAGAAAGATAGTCATGCGTAAAATTGACGCATG (SEQ ID NO: 7). In some embodiments, each symmetry ITR includes two additional base pairs: CCCTAGAAAGATAGTCATTGCGTAAAATTGACGCATG (SEQ ID NO: 8). In some embodiments, each symmetry ITR includes three additional base pairs: CCCTAGAAAGATAGTCATATGCGTAAAATTGACGCATG (SEQ ID NO: 9).

[0125] In some embodiments, the transposon symmetry RE ITR includes an additional base pair within the RE ITR between the 10 bp DDBD binding sequence (SEQ ID NO: 12) and the CRD binding sequence. In some embodiments, each symmetry ITR has one additional base pair for the RE ITR pair: The symmetry ITR includes CCCTAGAAAGATAATCATGCGTAAAATTGACGCATG (SEQ ID NO: 21). In some embodiments, each symmetry ITR includes two additional base pairs for the RE ITR pair: CCCTAGAAAGATAATCATTGCGTAAAATTGACGCATG (SEQ ID NO: 98). In some embodiments, each symmetry ITR includes three additional base pairs for the RE ITR pair: CCCTAGAAAGATAATCATATGCGTAAAATTGACGCATG (SEQ ID NO: 10).

[0126] Transposons containing one of the LE ITRs and RE ITRs containing additional base pairs can be transposed and incorporated by site-directed transposition using SPB and PBx transposases containing a double CRD domain, as well as DNA-binding domain fusion proteins.

[0127] nucleic acid Polynucleotides comprising nucleic acid sequences encoding the fusion proteins described herein are also provided herein. In some embodiments, these polynucleotides are isolated.

[0128] The isolated polynucleotides of this disclosure can be prepared using (a) recombinant methods, (b) synthetic techniques, (c) purification techniques, and / or (d) a combination thereof, as is well known in the art.

[0129] Methods for constructing nucleic acids encoding transposase domains containing N-terminal deletions as described herein are well known in the art or are described herein, for example, PCR-based mutagenesis.

[0130] The fusions of this disclosure can be produced using methods known in the art or any suitable method described herein.

[0131] The isolated polynucleotides of this disclosure, e.g., RNA, cDNA, genomic DNA, or any combination thereof, can be obtained from a biological source using any number of cloning methodologies known to those skilled in the art. In some embodiments, oligonucleotide probes that selectively hybridize to the polynucleotides of this disclosure under stringent conditions are used to identify desired sequences in a cDNA or genomic DNA library.

[0132] Methods for amplifying RNA or DNA are well known in the art and can be used in accordance with this disclosure without excessive experimentation, based on the teachings and guidelines presented herein. Known methods for amplifying DNA or RNA include polymerase chain reaction (PCR) and related amplification processes (e.g., U.S. Patents 4,683,195, 4,683,202, 4,800,159, 4,965,188 (Mullis et al.), 4,795,699 and 4,921,794 (Tabor et al.), 5,142,033 (Innis), 5,122,464 (Wilson et al.), 5,091,310 (Innis), 5,066,584 (Gyllensten et al.), 4 Examples include, but are not limited to, U.S. Patent Nos. 889,818 (Gelfand et al.), Nos. 4,994,370 (Silver et al.), Nos. 4,766,067 (Biswas), Nos. 4,656,134 (Ringold), and RNA-mediated amplification using antisense RNA against a target sequence as a template for double-stranded DNA synthesis (U.S. Patent No. 5,130,238 (Malek et al.), trade name NASBA). The entire contents of these references are incorporated herein by reference (see, for example, Ausubel or Sambrook mentioned above).

[0133] For example, polymerase chain reaction (PCR) technology can be used to directly amplify the sequences of the polynucleotides and related genes of this disclosure from a genomic DNA or cDNA library. PCR and other in vitro amplification methods are also useful for purposes such as cloning nucleic acid sequences encoding proteins to be expressed, creating nucleic acids to be used as probes to detect the presence of desired mRNA in a sample, nucleic acid sequencing, and other purposes. Examples of techniques sufficient to guide those skilled in the art through in vitro amplification methods can be found in Berger, Sambrook (n.), Ausubel (n.), Mullis (n.), et al., U.S. Patent No. 4,683,202 (1987), and Innis, et al., PCR Protocols: A Guide to Methods and Applications, Eds., Academic Press Inc., San Diego, Calif. (1990). Commercial kits for genomic PCR amplification are known in the art. See, for example, the Advantage-GC Genomic PCR Kit (Clontech). Furthermore, to improve the yield of longer PCR products, for example, the T4 gene 32 protein (Boehringer Mannheim) can be used.

[0134] The polynucleotides of this disclosure can also be prepared by direct chemical synthesis using known methods (see, for example, Ausubel et al.). Chemical synthesis generally yields single-stranded oligonucleotides, which can be converted to double-stranded DNA by hybridization with a complementary sequence or by polymerization using DNA polymerase with the single strand as a template. Those skilled in the art will understand that while the chemical synthesis of DNA is limited to a base sequence of about 100 bases, longer base sequences can be obtained by ligating shorter base sequences.

[0135] Expression vectors and host cells This disclosure also relates to vectors containing the polynucleotides of this disclosure, host cells genetically engineered with recombinant vectors, and the generation of at least one protein backbone by recombinant techniques, as is well known in the art. See, for example, Sambrook et al. and Ausubel et al., cited herein, respectively, which are fully incorporated herein by reference.

[0136] Polynucleotides can be optionally conjugated to vectors containing selectable markers for replication in a host. Generally, plasmid vectors are introduced in precipitates such as calcium phosphate precipitates or in complexes with charged lipids. If the vector is a virus, it can be packaged in vitro using a suitable packaging cell line and transduced into host cells.

[0137] The DNA insert must be operablely ligated to a suitable promoter. In some embodiments, the promoter is the EF-1α promoter. The expression construct further includes transcription start, termination, and ribosome binding sites for translation in the transcribed region. The coding portion of the mature transcript expressed by the construct preferably includes a translation start codon (e.g., ATG) at the beginning of the mRNA to be translated and a appropriately placed stop codon (e.g., UAA, UGA, or UAG) at the end, with UAA and UAG being preferred for expression in mammalian or eukaryotic cells.

[0138] The expression vector preferably contains at least one selectable marker, but is optional. Such markers include, but are not limited to, ampicillin, zeosin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / Geneticin (neo gene), DHFR (encoding Dihydrofolate Reductase and conferring resistance to Methotrexate), mycophenolic acid, or glutamine synthase (GS, U.S. Patent Nos. 5,122,464, 5,770,359, and 5,827,739), blasticidine (bsd gene), resistance genes for eukaryotic cell culture, and ampicillin, zeosin (Sh Examples include the bla gene, puromycin (pac gene), hygromycin B (hygB gene) G418 / Geneticin (neo gene), kanamycin, spectinomycin, streptomycin, carbenicillin, bleomycin, erythromycin, polymyxin B, or tetracycline resistance genes (the above patents are fully incorporated herein by reference). Suitable culture media and conditions for the above host cells are known in the art. Suitable vectors will be readily apparent to those skilled in the art. The introduction of vector constructs into host cells can be carried out by calcium phosphate transfection, DEAE-dextran-mediated transfection, cationic lipid-mediated transfection, electroporation, transduction, infection, or other known methods. Such methods are described in the art, for example, in Sambrook, Chapters 1-4 and 16-18, and Ausubel, Chapters 1, 9, 13, 15, and 16.

[0139] The expression vector preferably includes, but is optional, at least one selectable cell surface marker for isolating cells modified by the compositions and methods of the present disclosure. The selectable cell surface markers of the present disclosure consist of surface proteins, glycoproteins, or groups of proteins that distinguish a cell or subset of cells from another defined subset of cells. Preferably, the selectable cell surface marker distinguishes cells modified by the compositions or methods of the present disclosure from cells that have not been modified by the compositions or methods of the present disclosure. Examples of such cell surface markers include, but are not limited to, “cluster designation” or “classification determinant” proteins (often abbreviated as “CD”) such as cleaved or full-length forms of CD19, CD271, CD34, CD22, CD20, CD33, CD52, or combinations thereof. Cell surface markers include the suicide gene marker RQR8 (Philip B et al. Blood. 2014 Aug 21;124(8):1277-87).

[0140] The expression vector preferably contains at least one selectable drug resistance marker for isolating cells modified by the compositions and methods of the present disclosure, but this is optional. The selectable drug resistance markers of the present disclosure may include wild-type or mutant Neo, DHFR, TYMS, FRANCE, RAD51C, GCS, MDR1, ALDH1, NKX2.2, or any combination thereof.

[0141] Those skilled in the art will be familiar with the numerous expression systems available for expressing the nucleic acids encoding the proteins of the Disclosure. Alternatively, the nucleic acids of the Disclosure can be expressed in host cells by being (operatedly) turned on in host cells containing endogenous DNA encoding the protein backbone of the Disclosure. Such methods are well known in the art, for example, as described in U.S. Patents 5,580,734, 5,641,670, 5,733,746 and 5,733,761, which are fully incorporated herein by reference.

[0142] Examples of cell cultures useful for generating protein backbone, specific parts thereof, or variants include bacteria, yeast, and mammalian cells known in the art. Mammalian cell lines often exist in the form of a single layer of cells, but mammalian cell suspensions or bioreactors can also be used. Several suitable host cell lines capable of expressing intact glycosylated proteins have been developed in the art, including COS-1 (e.g., ATCC CRL 1650), COS-7 (e.g., ATCC CRL-1651), HEK293, BHK21 (e.g., ATCC CRL-10), CHO (e.g., ATCC CRL 1610), and BSC-1 (e.g., ATCC CRL-26) cell lines, Cos-7 cells, CHO cells, hepG 2 cells, P3X63Ag8.653, SP2 / 0-Ag14, 293 cells, and HeLa cells, which are readily available, for example, from the American Type Culture Collection, Manassas, Va (www.atcc.org). Preferred host cells include lymphoid cells such as myeloma cells and lymphoma cells. Particularly preferred host cells are P3X63Ag8.653 cells (ATCC accession number CRL-1580) and SP2 / 0-Ag14 cells (ATCC accession number CRL-1851). In a preferred embodiment, the recombinant cells are P3X63Ab8.653 or SP2 / 0-Ag14 cells.

[0143] These cell expression vectors may contain, but are not limited to, one or more of the following expression regulatory sequences: origins of replication, promoters (e.g., late or early SV40 promoter, CMV promoter (US Patent No. 5,168,062, 5,385,839), HSV tk promoter, pgk (phosphoglycerin kinase) promoter, EF-1 alpha promoter (US Patent No. 5,266,491), at least one human promoter, enhancer, and / or processing information sites (e.g., ribosome binding sites, RNA splice sites, polyadenylation sites (e.g., SV40 large T Ag poly-A addition site)), and transcription terminator sequences. See, for example, Ausubel et al. and Sambrook et al. above. Other cells useful for generating the nucleic acids or proteins of this disclosure are known and / or can be found, for example, in the American Type Culture Collection Catalogue of Cell Lines and It is available from Hybridoma (www.atcc.org) or other known or commercial sources.

[0144] When eukaryotic host cells are used, polyadenylated or transcriptional terminator sequences are typically incorporated into the vector. An example of a terminator sequence is a polyadenylated sequence derived from the bovine growth hormone gene. In some embodiments, the polyA sequence is the SV40 polyA sequence.

[0145] Sequences for precise splicing of the transcript can also be included. An example of a splicing sequence is the VP1 intron derived from SV40 (Sprague, et al., J. Virol. 45:773-781 (1983)). Furthermore, as is known in the art, gene sequences for controlling replication in host cells can be incorporated into the vector.

[0146] Plasmid constructs described herein may be used to deliver nucleic acids encoding transposase domains or fusion proteins described herein to cells.

[0147] The transposase domains and fusion proteins described herein may also be delivered to cells using mRNA constructs. Thus, in one embodiment, mRNA sequences encoding the transposase domains or fusion proteins described herein are provided herein. Such mRNA sequences may be delivered to cells using nanoparticles, such as lipid nanoparticles. Examples of lipid nanoparticles are described, for example, in International Patent Applications PCT / US2021 / 055876, PCT / US2022 / 017570, U.S. Provisional Applications 63 / 397,268, 63 / 301,855 and 63 / 348,614, each of which is incorporated herein by reference in whole with respect to examples of lipid nanoparticles that may be used to deliver the mRNA constructs encoding the fusion proteins or transposase domains described herein. The mRNA constructs may also be delivered to cells by electroporation or nucleofection. The mRNA may be capped or otherwise modified.

[0148] Cells and modified cells The dual CRD transposases and fusion proteins described herein may be used in conjunction with transposons, such as transposons containing symmetric ITRs, to modify cells. The transposons may be piggyBac®(PB) transposons. In some embodiments where the transposon is a PB transposon, the transposase is a piggyBac®(PB) transposase, a piggyBac-like(PBL) transposase, or a Super piggyBac®(SPB) transposase. Non-limiting examples of PB transposons are described in detail in U.S. Patents 6,218,182, 6,962,810, 8,399,643 and International Publication No. 2010 / 099296, each of which is incorporated herein by reference in whole with respect to examples of transposons that may be used in connection with the transposases and methods described herein. Transposons may include nucleic acids encoding therapeutic proteins or therapeutic agents. Examples of therapeutic proteins are disclosed in International Publication No. 2019 / 173636 and International Patent Application PCT / US2019 / 049816.

[0149] Accordingly, modified cells comprising one or more transposons and one or more bi-CRD transposases or fusion proteins as described herein are provided herein. The cells and modified cells of this disclosure may be mammalian cells. Preferably, the cells and modified cells are human cells.

[0150] Cells modified using the dual CRD transposase or fusion protein described herein may be germ cells or somatic cells. The cells and modified cells of this disclosure may include immune cells, such as lymphoid progenitor cells, natural killer (NK) cells, T lymphocytes (T cells), and stem memory T cells (T SCM T cells), central memory T cells (T CMModified cells may include stem cell-like T cells, B lymphocytes (B cells), antigen-presenting cells (APCs), cytokine-induced killer (CIK) cells, myeloid progenitor cells, neutrophils, basophils, eosinophils, monocytes, macrophages, platelets, erythrocytes, red blood cells (RBCs), megakaryocytes, or osteoclasts. Modified cells may be differentiated, undifferentiated, or immortalized. Modified undifferentiated cells may be stem cells. Modified undifferentiated cells may be induced pluripotent stem cells. Modified cells may be T cells, hematopoietic stem cells, natural killer cells, macrophages, dendritic cells, monocytes, megakaryocytes, or osteoclasts. Modified cells may be modified while quiescent, activated, interphase, prophase, metaphase, anaphase, or telophase. Modified cells may be fresh cells, cryopreserved cells, bulk cells, cells classified into subpopulations, cells from whole blood, cells from leukocyte apheresis, or cells from immortalized cell lines. Detailed descriptions for isolating cells from leukocyte apheresis products or blood are disclosed in International Publication No. 2019 / 173636 and International Patent Application No. PCT / US2019 / 049816.

[0151] The method of the present disclosure can modify and / or produce a population of modified T cells, wherein at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, or any percentage between these, are stem memory T cells (T SCM ) or T SCMThe cells express one or more cell surface markers, where one or more cell surface markers include CD45RA and CD62L. The cell surface markers may include one or more of CD62L, CD45RA, CD28, CCR7, CD127, CD45RO, CD95, CD95, and IL-2Rβ. The cell surface markers may include one or more of CD45RA, CD95, IL-2Rβ, CCR7, and CD62L.

[0152] This disclosure provides a method for expressing a CAR on the surface of a cell. The method comprises (a) obtaining a cell population, (b) contacting the cell population with a composition containing a CAR or a sequence encoding a CAR under conditions sufficient to transfer the CAR across the cell membrane of at least one cell in the cell population, thereby generating a modified cell population, (c) culturing the modified cell population under conditions suitable for incorporating the sequence encoding a CAR, and (d) growing and / or selecting at least one cell from the modified cell population that expresses a CAR on its cell surface. A more detailed description of the method for expressing a CAR on the surface of a cell is disclosed in International Publication No. 2019 / 049816 and International Patent Application PCT / US2019 / 049816.

[0153] This disclosure provides cells or a population of cells comprising a composition comprising (a) an inducible transgene construct comprising a sequence encoding an inducible promoter and a sequence encoding a transgene, and (b) a receptor construct comprising a sequence encoding a constitutive promoter and a sequence encoding an exogenous receptor such as a CAR, wherein when constructs (a) and (b) are incorporated into the genomic sequence of cells, the exogenous receptor is expressed, and when the exogenous receptor binds to a ligand or antigen, it transmits an intracellular signal that directly or indirectly targets the inducible promoter that modifies the expression of the inducible transgene (a), thereby modifying gene expression.

[0154] This disclosure further provides compositions comprising modified, amplified, and selected cell populations of the methods described herein.

[0155] The modified cells of this disclosure (e.g., CAR T cells) can be further modified to enhance their therapeutic potential. Alternatively, or in addition to this, the modified cells can be further modified to reduce their sensitivity to immunological and / or metabolic checkpoints, for example, by blocking and / or diluting specific checkpoint signals that are naturally delivered to the cell within the tumor immunosuppressive microenvironment (e.g., checkpoint inhibition).

[0156] The modified cells of this disclosure (e.g., CAR T cells) may be further modified to silence or reduce the expression of (i) one or more genes encoding receptors for inhibitory checkpoint signals, (ii) one or more genes encoding intracellular proteins involved in checkpoint signaling, (iii) one or more genes encoding transcription factors that interfere with therapeutic efficacy, (iv) one or more genes encoding receptors for cell death or apoptosis, (v) one or more genes encoding metabolic sensing proteins, (vi) one or more genes encoding proteins that confer sensitivity to cancer therapies such as monoclonal antibodies, and / or (vii) one or more genes encoding growth advantage factors. Non-limiting examples of genes that may be modified to silence or reduce their expression or suppress their function include, but are not limited to, exemplary inhibitory checkpoint signals, intracellular proteins, transcription factors, receptors for cell death or apoptosis, metabolic sensing proteins, proteins that confer sensitivity to cancer therapies, and growth advantage factors disclosed in International Publication No. 2019 / 173636.

[0157] The modified cells of this disclosure (e.g., CAR T cells) can be further modified to express modified / chimeric checkpoint receptors. Modified / chimeric checkpoint receptors may include null receptors, decoy receptors, or dominant-negative receptors. Exemplary null, decoy, or dominant-negative intracellular receptors / proteins include, but are not limited to, downstream signaling components of inhibitory checkpoint signals, transcription factors, cytokines or cytokine receptors, chemokines or chemokine receptors, cell death or apoptosis receptors / ligands, metabolic sensing molecules, proteins that confer sensitivity to cancer treatment, and oncogenes or tumor suppressor genes. Non-exclusive examples of cytokines, cytokine receptors, chemokines, and chemokine receptors are disclosed in International Publication No. 2019 / 173636.

[0158] Genome modification may involve introducing nucleic acid sequences, transgenes, and / or genome editing constructs into cells ex vivo, in vivo, in vitro, or in situ to stably incorporate nucleic acid sequences, transiently incorporate nucleic acid sequences, induce site-directed integration of nucleic acid sequences, or induce biased integration of nucleic acid sequences. A nucleic acid sequence may be a transgene.

[0159] Stable chromosome integration can be random, site-specific, or biased. While we do not wish to be bound by theory, the addition of DNA-binding domains to transposases as described herein is thought to improve the site specificity of the transposases.

[0160] Site-specific integration can occur at safe harbor sites. Genomic safe harbor sites can provide a place for the integration of new genetic material in a way that ensures the newly inserted genetic element functions reliably (e.g., is expressed at therapeutically effective levels) and does not cause harmful changes in the host genome that pose a risk to the host organism. Non-limiting examples of potential genomic safe harbors include the intron sequence of the human albumin gene, adeno-associated virus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19, the site of the chemokine (CC motif) receptor 5 (CCR5) gene, and the site of the human ortholog at the mouse Rosa26 locus.

[0161] Site-directed transgene integration can occur at sites that disrupt the expression of a target gene. This disruption can occur through site-directed integration at introns, exons, promoters, gene elements, enhancers, suppressors, start codons, stop codons, and response elements. Non-exclusive examples of target genes targeted by site-directed integration include TRAC, TRAB, PD1, any immunosuppressive genes, and genes involved in allorejection.

[0162] Site-directed transgene integration can occur at sites that result in enhanced expression of the target gene. Enhanced target gene expression can occur through site-directed integration at introns, exons, promoters, gene elements, enhancers, suppressors, start codons, stop codons, and response elements.

[0163] Site-specific transgene integration sites can be unstable chromosomal insertions. Unstable integration can be transient non-chromosomal integration, semi-stable non-chromosomal integration, semi-persistent non-chromosomal insertion, or unstable chromosomal insertion. Transient non-chromosomal insertions can be epichromosomal or cytoplasmic. In one embodiment, transient non-chromosomal insertions of transgenes are not integrated into the chromosome, and the modified genetic material is not replicated during cell division.

[0164] Site-specific transgene integration sites may be modified binding sites for DNA targeting domains in transposon domains, fusion proteins, or tandem dimers as described herein. For example, the TTAA target DNA integration site for SPB may be modified to insert an adjacent DNA binding site to a DNA targeting domain containing three zinc finger motifs (e.g., a DNA targeting domain containing or consisting of the sequence of Sequence ID No. 74, or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto). For example, the DNA targeting domain containing three zinc finger motifs encoded by Sequence ID No. 74 is thought to bind to the DNA sequence GCGTGGGCG. Therefore, the introduction of two copies of the target DNA sequence adjacent to the TTAA target integration site of SPB is thought to improve site-specific integration of the SPB transposase domain containing the DNA targeting domain with three zinc finger motifs. The two copies of the target sequence are inverted (5') and complementary (3') orientations.

[0165] In some embodiments, polynucleotides are provided herein that include, in some embodiments, an inverted complementary sequence of the target site for the DNA targeting domain in the order of 5' to 3', a first spacer, a TTAA target integration site for the SPB, a second spacer, and the sequence of the target site for the DNA targeting domain. In some embodiments, the first and second spacers are of the same length. In some embodiments, the first and / or second spacers are 3 bp long. In some embodiments, the first and / or second spacers are 4 bp long. In some embodiments, the first and / or second spacers are 5 bp long. In some embodiments, the first and / or second spacers are 6 bp long. In some embodiments, the first and / or second spacers are 7 bp long. In some embodiments, the first and / or second spacers are 8 bp long. In some embodiments, the first and / or second spacers are 9 bp long. In some embodiments, the first and / or second spacers are 10 bp long.

[0166] Exemplary polynucleotide sequences, including the inverted complement of the target site sequence of the DNA targeting domain containing three zinc finger motifs in the order of 5' to 3', a first spacer, the TTAA target integration site of the SPB, a second spacer, and the target site sequence of the DNA targeting domain containing three zinc finger motifs, are shown in SEQ ID NOs: 94-97. The lengths of the first and second spacers in SEQ ID NOs: 94-97 are 8 bp, 7 bp, 6 bp, and 5 bp, respectively. The inverted target site for the DNA targeting domain and its complement are underlined, and the TTAA sequence is shown in bold. TIFF2026517869000019.tif7170 (Sequence ID 94) TIFF2026517869000020.tif7170 (Sequence ID 95) TIFF2026517869000021.tif7170 (Sequence ID 96) TIFF2026517869000022.tif7170 (Sequence ID 97)

[0167] The modified target site can be introduced into a cell or cell line to facilitate targeted genomic manipulation. For example, a cell line manipulated to include a modified target site for an SPB or PBx provided herein can be transfected with the SPB or PBx, as well as a transposon containing the donor DNA, so that the donor DNA is inserted into the modified target site. In some embodiments, the cell line is a T cell line. In some embodiments, the modified target sequence is introduced into a high-expression genomic region. In certain embodiments, a cell line is provided herein in which a nucleic acid sequence is stably incorporated into its genomic sequence, comprising, in a specific embodiment, an inverted complement of the sequence of the target site of a DNA targeting domain containing three zinc finger motifs in a 5' to 3' order, a first spacer, the TTAA target integration site of the SPB, a second spacer, and the sequence of the target site of the DNA targeting domain containing three zinc finger motifs. In some embodiments, the cell line includes one of sequence numbers 94-97 stably incorporated into its genome. In some embodiments, the cells are in vitro cells, for example, cells in cell culture.

[0168] In the case of DNA-binding domains containing TALENs, the target site is determined by the TALEN sequence. Those skilled in the art will be able to modify the TALEN sequence to achieve desired target specificity.

[0169] Genome modification can result in the unstable chromosomal integration of introduced genes. Integrated introduced genes may be silenced, removed, excised, or further modified.

[0170] In some embodiments, the transposase domains, fusion proteins, and tandem dimer complexes provided herein have better transposase efficacy than their wild-type counterparts. Transposase activity can be measured by any suitable assay known in the art or described herein, such as the Split GFP assay. For example, the transposase domains, fusion proteins, and tandem dimer complexes provided herein may have on-target genome integration activity equivalent to their wild-type counterparts, but with reduced off-target genome integration activity compared to their wild-type counterparts.

[0171] In some embodiments, transposase domains comprising an N-terminal deletion and a DNA-targeting domain provided herein have an on-target activity to off-target activity ratio of at least 50 times, at least about 100 times, at least about 150 times, at least about 200 times, at least about 250 times, at least about 300 times, at least about 350 times, at least about 400 times, at least about 450 times, at least about 500 times, at least about 550 times, at least about 600 times, at least about 650 times, at least about 700 times, at least about 750 times, at least about 800 times, at least about 850 times, at least about 900 times, at least about 950 times, or at least about 1000 times compared to a wild-type transposase domain.

[0172] In some embodiments, a transposase domain provided herein, comprising a DNA-targeting domain inserted into the N-terminal region of the transposase domain, has an on-target activity to off-target activity ratio of at least 50 times, at least about 100 times, at least about 150 times, at least about 200 times, at least about 250 times, at least about 300 times, at least about 350 times, at least about 400 times, at least about 450 times, at least about 500 times, at least about 550 times, at least about 600 times, at least about 650 times, at least about 700 times, at least about 750 times, at least about 800 times, at least about 850 times, at least about 900 times, at least about 950 times, or at least about 1000 times compared to a wild-type transposase domain.

[0173] In certain embodiments, modified cells are used therapeutically in adoptive cell therapy.

[0174] A adoptive cell composition that is "universally" safe for administration to any patient (not just the patient from whom the adoptive cell composition originates) requires a significant reduction or elimination of alloreactivity. For this purpose, the cells of the Disclosure (e.g., allogeneic cells) can be modified to inhibit the expression or function of a class of T cell receptors (TCRs) and / or major histocompatibility complexes (MHCs). TCRs mediate graft-versus-host (GvH) responses, and MHCs mediate host-versus-graft (HvG) responses. In a preferred embodiment, any expression and / or function of TCRs is eliminated to prevent T cell-mediated GvH that could cause death of the subject. Thus, in a preferred embodiment, the Disclosure provides a pure TCR-negative allogeneic T cell composition (e.g., each cell in the composition expresses TCRs at undetectable or non-existent levels).

[0175] The expression and / or function of MHC class I (MHC-I, specifically HLA-A, HLA-B, and HLA-C) is reduced or eliminated to prevent HvG and consequently improve cell engraftment in the target. Improved engraftment leads to longer cell persistence and, consequently, a broader therapeutic range for the target. Specifically, the expression and / or function of beta-2-microglobulin (B2M), a structural component of MHC-I, is reduced or eliminated. Non-limiting examples of guide RNAs (gRNAs) for targeting and deleting MHC activators are disclosed in International Patent Application PCT / US2019 / 049816.

[0176] A detailed description of genetic modifications of endogenous sequences encoding the naturally occurring chimeric stimulatory receptors, TCR-alpha (TCR-α), TCR-beta (TCR-β), and / or beta-2-microglobulin (β2M), as well as naturally occurring polypeptides including the HLA class I histocompatibility antigen, alpha-E (HLA-E) polypeptide, is disclosed in International Patent Application No. PCT / US2019 / 049816.

[0177] Under normal conditions, complete T cell activation depends on the involvement of the TCR in conjunction with a second signal mediated by one or more co-stimulatory receptors (e.g., CD28, CD2, 4-1BBL) that enhance the immune response. However, in the absence of a TCR, T cell proliferation is significantly reduced upon stimulation using standard activating / stimulating reagents such as agonist anti-CD3 mAbs. Accordingly, this disclosure provides a naturally occurring chimeric stimulating receptor (CSR) comprising: (a) an extracellular domain containing an activating component, the activating component being isolated or induced from a first protein; (b) a transmembrane domain; and (c) an endodomain containing at least one signaling domain, the at least one signaling domain being isolated or induced from a second protein, wherein the first and second proteins are not identical.

[0178] The activating component may include one or more parts of components of T cell receptors (TCRs), TCR complexes, TCR coreceptors, TCR costimulatory proteins, TCR inhibitory proteins, cytokine receptors, and chemokine receptors to which the agonist of the activating component binds. The activating component may also include the extracellular domain of CD2 or a part thereof to which the agonist binds.

[0179] The signaling domain may include one or more components of human signaling domains, T cell receptors (TCRs), TCR complexes, TCR coreceptors, TCR costimulatory proteins, TCR inhibitory proteins, cytokine receptors, and chemokine receptors. The signaling domain may include the CD3 protein or a portion thereof. The CD3 protein may include the CD3ζ protein or a portion thereof.

[0180] The endodomain may further contain a cytoplasmic domain. The cytoplasmic domain can be isolated or derived from a third protein. The first and third proteins may be identical. The extracellular domain may further contain a signal peptide. The signal peptide can be derived from a fourth protein. The first and fourth proteins may be identical. The transmembrane domain can be isolated or derived from a fifth protein. The first and fifth proteins may be identical.

[0181] This disclosure also provides a naturally occurring chimeric stimulating receptor (CSR) whose extracellular domain includes modifications. The modifications may include mutations or cleavage of the amino acid sequence of the activator or the first protein compared to the wild-type sequence of the activator or the first protein. Mutations or cleavage of the amino acid sequence of the activator may include mutations or cleavage of the CD2 extracellular domain to which the agonist binds or a portion thereof. Mutations or cleavage of the CD2 extracellular domain can reduce or eliminate binding to naturally occurring CD58.

[0182] This disclosure provides nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides transposons or vectors comprising nucleic acid sequences encoding any CSR disclosed herein.

[0183] This disclosure provides cells containing any CSR disclosed herein. This disclosure provides cells containing nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides cells containing vectors containing nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides cells containing transposons containing nucleic acid sequences encoding any CSR disclosed herein.

[0184] In certain embodiments, the cells of this disclosure are modified to recombinantly express dihydrofolate reductase (DHFR), thereby making the cells favorably resistant to methotrexate (MTX). MTX-resistant cells may be used in combination with subsequent MTX administration in methods of treating targets requiring it, to eliminate activated T cells and NK cells that target the modified cells or therapeutic cells, thereby increasing the in vivo persistence and efficacy of the modified cells.

[0185] This disclosure provides compositions comprising any CSR disclosed herein. This disclosure provides compositions comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising vectors comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising transposons comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising modified cells disclosed herein or compositions comprising a plurality of modified cells disclosed herein.

[0186] Methods for site-directed gene insertion are also provided herein. The dual CRD transposase domains and fusion proteins provided herein may be used to deliver a transgene to a cell and to insert the transgene into a target site. The target site may be, for example, a genome-safe harbor, i.e., a genomic site into which the transgene can be inserted in such a way that it is unlikely to cause harmful changes to the cell's gene expression profile, and that the transgene will function predictably. In some embodiments, the target site is a repeating element, such as a LINE-1 or ALU sequence. The repeating element does not encode an essential gene product, thus reducing the likelihood that the insertion will cause harmful changes to the cell's gene expression profile. One, two, or more target sites may be present within a single repeating element. In some embodiments, the target site is located within an intron (e.g., an intron of a PAH gene).

[0187] Site-directed integration can be used in vitro or in vivo. One example of in vivo application is gene therapy, which involves the delivery of a transgene to the genomic DNA of a cell.

[0188] Formulation, dosage, and method of administration This disclosure provides formulations, dosages, and methods for administering the compositions and cells described herein. In one embodiment, a pharmaceutical composition comprising a tandem dimer transposase or fusion protein described herein and a pharmaceutically acceptable carrier is provided herein. In another embodiment, a pharmaceutical composition comprising modified cells described herein and a pharmaceutically acceptable carrier is provided herein.

[0189] The disclosed compositions and pharmaceutical compositions may, but are not limited to, contain at least one of any suitable adjuvants, such as diluents, binders, stabilizers, buffers, salts, lipophilic solvents, preservatives, and adjuvants. Pharmaceutically acceptable adjuvants are preferred. Non-limiting examples of such sterile solutions and methods for their preparation are well known in the art, for example, but are not limited to, Gennaro, Ed., Remington's Pharmaceutical Sciences, 18th Edition, Mack Publishing Co. (Easton, Pa.) 1990 and "Physician's Desk Reference", 52nd ed., Medical Economics (Montvale, NJ) 1998. Pharmaceutically acceptable carriers suitable for the administration method, solubility, and / or stability of protein backbone, fragment, or variant compositions can be typically selected, as are well known in the art or as described herein.

[0190] Non-limiting examples of suitable pharmaceutical excipients and additives include proteins, peptides, amino acids, lipids, and carbohydrates (e.g., monosaccharides, sugars including di-, tri-, tetra-, and oligosaccharides, derivatized sugars such as alditol, aldonic acid, and esterified sugars, and polysaccharides or sugar polymers), which can exist alone or in combination, and which together constitute 1 to 99.99% by weight or volume. Non-limiting examples of protein excipients include serum albumins such as human serum albumin (HSA), recombinant human albumin (rHA), gelatin, and casein. Representative amino acid / protein components that can also function as buffers include alanine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, and aspartame. One preferred amino acid is glycine.

[0191] Non-limiting examples of carbohydrate excipients suitable for use include monosaccharides such as fructose, maltose, galactose, glucose, D-mannose, and sorbose; disaccharides such as lactose, sucrose, trehalose, and cellobiose; polysaccharides such as raffinose, meletitose, maltodextrin, dextran, and starch; and algitols such as mannitol, xylitol, maltitol, lactitol, xylitol sorbitol (glucitol), and myo-inositol. Preferably, the carbohydrate excipient is mannitol, trehalose, and / or raffinose.

[0192] This composition may also contain a buffer or pH adjuster, typically a salt prepared from an organic acid or base. Typical buffers include organic acid salts such as citric acid, ascorbic acid, gluconic acid, carbonate, tartaric acid, succinic acid, acetic acid, and phthalic acid salts, as well as Tris, tromethamine hydrochloride, and phosphate buffer. Preferred buffers are organic acid salts such as citrate.

[0193] Furthermore, the disclosed compositions may include polymer excipients / additives, such as polyvinylpyrrolidone, Ficol (polymer sugar), dextrose (e.g., cyclodextrin, e.g., 2-hydroxypropyl-β-cyclodextrin), polyethylene glycol, flavoring agents, antimicrobial agents, sweeteners, antioxidants, antistatic agents, surfactants (e.g., polysorbates, e.g., "TWEEN20" and "TWEEN80"), lipids (e.g., phospholipids, fatty acids), steroids (e.g., cholesterol), and chelating agents (e.g., EDTA).

[0194] Many known and developed methods can be used to administer a therapeutically effective amount of the composition or pharmaceutical composition disclosed herein. Non-limiting examples of administration methods include bolus, oral, intradrip, intra-articular, intra-bronchial, intraperitoneal, intra-articular capsule, intra-cartilage, intracavitary, intra-abdominal, intracerebral, intracolon, intracervical, intragastric, intrahepatic, intrafocal, intramuscular, intramyocardial, intranasal, intraocular, intraosseous, intraosteal, intrapelvic, intrapericardial, intraperitoneal, intrapleural, intraprostatic, intrapulmonary, intrarectal, intranephrological, intraretinal, intraspinal, synovial, intrathoracic, intrauterine, intratumoral, intravenous, intrabladder, oral, parenteral, rectal, sublingual, subcutaneous, percutaneous, or vaginal methods. In preferred embodiments, the composition comprising the modified cells described herein is administered intravenously, for example, by intravenous infusion.

[0195] The compositions of this disclosure can be prepared for use in parenteral administration (subcutaneous, intramuscular, or intravenous) or any other administration in particular in the form of a liquid solution or suspension. For parenteral administration, the compositions disclosed herein may be formulated together with a pharmaceutically acceptable parenteral vehicle as a solution, suspension, emulsion, particles, powder, or lyophilized powder, or may be provided separately from the parenteral vehicle. Formulations for parenteral administration may contain, as common excipients, sterile water or saline, polyalkylene glycols such as polyethylene glycol, plant-derived oils, hydrogenated naphthalene, etc. Aqueous or oily suspensions for injection may be prepared according to known methods using appropriate emulsifiers or humectants and suspending agents. Drugs for injection or infusion may be non-toxic, orally unadministerable diluents such as aqueous solutions, sterile injection solutions, or suspensions in solvents. Usable vehicles or solvents include water, Ringer's solution, isotonic saline, etc., and sterile non-volatile oils may be used as common solvents or suspension solvents. For these purposes, any kind of non-volatile oils and fatty acids can be used, such as natural, synthetic, or semi-synthetic fatty oils or fatty acids, and natural, synthetic, or semi-synthetic mono-, di-, or tri-glycerides. Parenteral administration is known in the art and includes, but is not limited to, conventional injection methods, gas-pressurized needleless injection devices such as those described in U.S. Patent No. 5,851,198, and laser puncture devices such as those described in U.S. Patent No. 5,839,446.

[0196] It may be desirable to deliver the disclosed compound to the subject in a single dose over a long period, for example, from one week to one year. Various sustained-release, depot, and implant formulations can be used. For example, the dosage form may include pharmaceutically acceptable non-toxic salts of compounds with low solubility in body fluids, e.g., (a) acid addition salts with polybasic acids, e.g., phosphoric acid, sulfuric acid, citrate, tartaric acid, tannic acid, pamoic acid, alginic acid, polyglutamic acid, naphthalene mono- or disulfonic acid, polygalacturonic acid, etc., (b) salts with polyvalent metal cations, e.g., zinc, calcium, bismuth, barium, magnesium, aluminum, copper, cobalt, nickel, cadmium, etc., or salts with organic cations formed from, for example, N,N'-dibenzyl ethylenediamine or ethylenediamine, or (c) a combination of (a) and (b), e.g., zinc tannate salt. Furthermore, the disclosed compounds or preferably relatively insoluble salts, such as those described above, can be formulated into gels suitable for injection, such as aluminum monostearate gel containing sesame oil. Particularly preferred salts include zinc salts, zinc tannate salts, and pamoate salts. Another type of sustained-release depot formulation for injection contains the compound or salt dispersed for encapsulation in a rapidly degrading, non-toxic, non-antigenic polymer, such as polylactic acid / polyglycolic acid polymer, as described, for example, in U.S. Patent No. 3,773,919. The compound or preferably relatively insoluble salts, such as those described above, can also be formulated into cholesterol matrix silastic pellets, particularly for use in animals. Further sustained-release formulations, depot formulations, or implant formulations, such as gaseous or liquid liposomes, are known in the literature (U.S. Patent No. 5,770,222 and “Sustained and Controlled Release Drug Delivery Systems”, JR Robinson ed. Marcel Dekker, Inc., NY, 1978).

[0197] Treatment method In another embodiment, a method for treating a disease or disorder of interest is provided herein, comprising administering a composition comprising the modified cells described herein to the subject. The terms “subject” and “patient” are used interchangeably herein. In a preferred embodiment, the patient is human.

[0198] Modified cells can be allogeneic or autologous to the patient. In some preferred embodiments, modified cells are allogeneic cells. In some embodiments, modified cells are autologous T cells or modified autologous CAR T cells. In some preferred embodiments, modified cells are allogeneic T cells or modified allogeneic CAR T cells.

[0199] In some embodiments, the disease or disorder treated according to the method described herein is cancer. In some embodiments, the treatment method described herein may delay the progression of cancer and / or reduce the tumor burden.

[0200] In some embodiments, the disease or disorder treated according to the methods described herein is an autoimmune disorder. In some embodiments, the autoimmune disorders are autoimmune neutropenia, Guillain-Barré syndrome, epilepsy, autoimmune encephalitis, Isaac syndrome, nevus syndrome, pemphigus vulgaris, pemphigus deciduousis, bullous pemphigoid, acquired epidermolysis bullosa, pemphigoid of pregnancy, mucosal pemphigoid, antiphospholipid syndrome, autoimmune anemia, myasthenia gravis, autoimmune Graves' disease, thyroid eye disease (TED), Goodpasture syndrome, multiple sclerosis, rheumatoid arthritis, lupus, idiopathic thrombocytopenic purpura (ITP), warm autoimmune hemolytic anemia (WAIHA), chronic inflammatory demyelinating polyneuropathy (CIDP), lupus nephritis, or membranous nephropathy.

[0201] The dosage of the pharmaceutical composition administered to the subject may vary depending on known factors such as the pharmacodynamic properties of the specific drug, its mode and route of administration, the recipient's age, health status and weight, the nature and severity of symptoms, the type of concomitant therapy, the frequency of treatment, and the desired effect.

[0202] In an embodiment where the composition to be administered to the subject that requires it is the modified cell disclosed herein, about 1×10 3 to about 1×10 4 cells, about 1×10 4 to about 1×10 5 cells, about 1×10 5 to about 1×10 6 cells, about 1×10 6 to about 1×10 7 cells, about 1×10 7 to about 1×10 8 cells, about 1×10 8 to about 1×10 9 cells, about 1x10 9 to about 1x10 10 cells, about 1x10 10 to about 1x10 11 cells, about 1x10 11 to about 1x10 12 cells, about 1x10 12 to about 1x10 13 cells, about 1x10 13 to about 1x10 14 cells, about 1x10 14 to about 1x10 15 cells, about 1x10 15 to about 1x10 16 cells, about 1x10 16 to about 1x10 17 cells, about 1x10 17 to about 1x10 18 cells, about 1x10 18 to about 1x10 19 cells, or about 1×10 19 to about 1×10 20 cells may be administered. In some embodiments, the cells are administered at a dose of about 5×10 6 to about 25×10 6 cells.

[0203] In other embodiments, the dosage of the cells may depend on the body weight of the human, and per kg of the subject's body weight, for example, about 1×10 3 to about 1×10 4 cells, about 1×10 4 to about 1×10 5Each cell is approximately 1 x 10⁻⁶ 5 ~Approx. 1×10 6 Each cell is approximately 1 x 10⁻⁶ 6 ~Approx. 1×10 7 Each cell is approximately 1 x 10⁻⁶ 7 ~Approx. 1×10 8 Each cell is approximately 1 x 10⁻⁶ 8 ~Approx. 1×10 9 Each cell, approximately 1 x 10⁶ 9 ~approximately 1x10 10 Each cell, approximately 1 x 10⁶ 10 ~approximately 1x10 11 Each cell, approximately 1 x 10⁶ 11 ~approximately 1x10 12 Each cell, approximately 1 x 10⁶ 12 ~approximately 1x10 13 Each cell, approximately 1 x 10⁶ 13 ~approximately 1x10 14 Each cell, approximately 1 x 10⁶ 14 ~approximately 1x10 15 Each cell, approximately 1 x 10⁶ 15 ~approximately 1x10 16 Each cell, approximately 1 x 10⁶ 16 ~approximately 1x10 17 Each cell, approximately 1 x 10⁶ 17 ~approximately 1x10 18 Each cell, approximately 1 x 10⁶ 18 ~approximately 1x10 19 A single cell, or approximately 1 × 10⁻⁶ cells. 19 ~Approx. 1×10 20 Individual cells may be administered.

[0204] A more detailed description of the disclosed compositions and pharmaceutically acceptable excipients, formulations, dosages and methods of administration of the pharmaceutical compositions is disclosed in International Publication No. 2019 / 049816.

[0205] The transposon domains and fusion proteins provided herein may be used to deliver gene therapy. Gene therapy typically involves the delivery of a transgene into the genomic DNA of a cell. Typically, the transgene replaces a gene that is mutated or otherwise not properly expressed within the cell. The fusion proteins, transposase domains, and complexes described herein may be used to deliver therapeutic transgenes to cells and to incorporate the transgenes into target sites. In some embodiments, the treatment method involves introducing the fusion proteins and transposons described herein into cells, wherein the transposon comprises a 5'ITR, a transgene, and a 3'ITR in the order of 5' to 3'.

[0206] kit In another embodiment, a kit is provided herein comprising a cell line engineered to contain a modified target site of an SPB or PBx provided herein within its genome, preferably within a highly expressed genomic region. The kit may further comprise a composition comprising one or more SPB or PBx transposase domains or fusion proteins described herein. In some embodiments, the cell line is a T cell line.

[0207] definition When used throughout this disclosure, the singular forms "a," "an," and "the" include multiple references unless the context explicitly indicates otherwise. Thus, for example, a reference to "a method" includes multiple such methods, and a reference to "a dose" includes one or more doses and equivalents known to those skilled in the art.

[0208] The terms “about” or “approximately” mean that a particular value is within an acceptable margin of error, as determined by those skilled in the art, and this depends to some extent on the method by which the value is measured or determined, e.g., on the limitations of the measurement system. For example, “about” means within one or more standard deviations. Alternatively, “about” may mean a range of up to 20%, or up to 10%, or up to 5%, or up to 1% of a given value. Or, particularly with respect to biological systems or processes, the term may mean within one order of magnitude of the value, preferably up to five times, and more preferably up to two times. Where a particular value is described in this application and claims, unless otherwise specified, the term “about” should be assumed to mean within an acceptable margin of error of that particular value.

[0209] This disclosure provides isolated or substantially purified polynucleotide or protein compositions. “Isolated” or “purified” polynucleotides or proteins, or their biologically active portions, substantially or essentially contain no components that would normally accompany or interact with polynucleotides or proteins found in their naturally occurring environments. Therefore, isolated or purified polynucleotides or proteins, if produced by recombinant technology, substantially contain no other cellular material or culture medium, and if chemically synthesized, substantially contain no chemical precursors or other chemicals. Optimally, an “isolated” polynucleotide does not contain sequences naturally adjacent to it in the genomic DNA of the organism from which it originates (i.e., sequences located at the 5' and 3' ends of the polynucleotide (optimally, protein-coding sequences)). For example, in various embodiments, an isolated polynucleotide may contain about 5 kb, about 4 kb, about 3 kb, about 2 kb, about 1 kb, about 0.5 kb, or less than about 0.1 kb of nucleotide sequences naturally adjacent to it in the genomic DNA of the cell from which it originates. Substantially cellular protein-free proteins include protein preparations containing approximately 30%, approximately 20%, approximately 10%, approximately 5%, or less than approximately 1% (dry weight) of contaminating protein. When the proteins of this disclosure or their biologically active portions are recombinantly produced, the culture medium optimally contains approximately 30%, approximately 20%, approximately 10%, approximately 5%, or less than approximately 1% (dry weight) of chemical precursors or non-protein-of-interest chemicals.

[0210] This disclosure provides disclosed DNA sequence fragments and variants, and proteins encoded by these DNA sequences. As used throughout this disclosure, the term “fragment” refers to a portion of a DNA sequence, or a portion of an amino acid sequence, and by extension, the protein encoded thereby. A DNA sequence fragment consisting of a coding sequence may encode a protein fragment that retains the biological activity of a native protein and therefore retains DNA recognition or binding activity to a target DNA sequence, as described herein. Alternatively, DNA sequence fragments useful as hybridization probes generally do not encode a protein that retains biological activity or promoter activity. Therefore, DNA sequence fragments may range from at least about 20 nucleotides, about 50 nucleotides, about 100 nucleotides to the full-length polynucleotides of this disclosure.

[0211] The nucleic acids or proteins of this disclosure can be constructed by a modular approach, which involves pre-assembling monomer units and / or repeating units in a target vector and then assembling them into a final target vector. The polypeptides of this disclosure can be constructed by a modular approach, which involves pre-assembling repeating units in a target vector that can be composed of repeating monomers of this disclosure and then assembled into a final target vector. This disclosure provides polypeptides produced by this method and nucleic acid sequences encoding these polypeptides. This disclosure provides host organisms and cells containing nucleic acid sequences encoding polypeptides produced by this modular approach.

[0212] The term “comprising” is intended to mean that a compound, composition, or method includes the elements described but does not exclude others. “Consisting essentially of,” when used to define a composition or method, means excluding other elements that are essentially important to the combination for the purposes described. Thus, a composition essentially consisting of the components defined herein does not exclude trace amounts of contaminants or inert carriers. “Consisting of” means excluding elements and substantial method steps that are more than trace amounts of other components. The embodiments defined by each of these transitional terms are within the scope of this disclosure.

[0213] As used herein, “expression” refers to the process by which a polynucleotide is transcribed into mRNA, and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression includes the splicing of mRNA in eukaryotic cells.

[0214] "Gene expression" is the process of converting the information contained in a gene into a gene product. A gene product can be a direct transcript of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, shRNA, microRNA, structural RNA, or other types of RNA) or a protein produced by the translation of mRNA. Gene products also include RNA modified by processes such as capping, polyadenylation, methylation, and editing, as well as proteins modified by processes such as methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristoylation, and glycosylation.

[0215] The "regulation" or "control" of gene expression refers to a change in gene activity. Regulation of expression includes, but is not limited to, gene activation and gene repression.

[0216] The term "operatively linked" or its synonym (e.g., "linked operatively") means that two or more molecules are positioned relative to each other so that they can interact in a way that influences the function of one or both of the molecules or a combination thereof. In relation to nucleic acids, a promoter may be operatively linked to a nucleotide sequence encoding a transposase domain or fusion protein as described herein, resulting in the expression of the nucleotide sequence under the control of the promoter.

[0217] Components linked by non-covalent bonds, and methods for producing and using non-covalently linked components are disclosed. Various components can take on a variety of different forms, as described herein. For example, non-covalently linked (i.e., operably linked) proteins may be used to enable transient interactions that circumvent one or more problems in the art. The ability of non-covalently linked components, such as proteins, to associate and dissociate allows for functional associations only under circumstances where such association is required for the desired activity, or primarily functional associations. The linkage only needs to last long enough to achieve the desired effect.

[0218] A method for inducing a protein to a specific gene locus in the genome of an organism is disclosed. This method may include a step of providing a DNA localization component and a step of providing an effector molecule, the DNA localization component and the effector molecule being operablely linked via non-covalent linkage.

[0219] A "target site" or "target sequence" is a nucleic acid sequence that defines the portion of the nucleic acid to which a binding molecule will bind, provided that sufficient conditions for binding are present.

[0220] The terms “nucleic acid,” “oligonucleotide,” or “polynucleotide” refer to at least two nucleotides linked by a covalent bond. A single-stranded description also defines the sequence of a complementary strand. Thus, a nucleic acid may encompass a complementary strand of the described single-stranded strand. The nucleic acids of this disclosure also encompass substantially identical nucleic acids and their complements that retain the same structure or encode the same protein.

[0221] The nucleic acids of this disclosure may be single-stranded or double-stranded. The nucleic acids of this disclosure may contain double-stranded sequences even if the majority of the molecule is single-stranded. The nucleic acids of this disclosure may contain single-stranded sequences even if the majority of the molecule is double-stranded. The nucleic acids of this disclosure may include genomic DNA, cDNA, RNA, or hybrids thereof. The nucleic acids of this disclosure may include combinations of deoxyribonucleotides and ribonucleotides. The nucleic acids of this disclosure may include combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. The nucleic acids of this disclosure may be synthesized to include non-natural amino acid modifications. The nucleic acids of this disclosure may be obtained by chemical synthesis or recombinant methods.

[0222] The nucleic acids disclosed may have their entire base sequence or any part thereof that does not exist in nature. The nucleic acids disclosed may contain one or more mutations, substitutions, deletions, or insertions that do not exist in nature, and the entire nucleic acid sequence may not exist in nature. The nucleic acids disclosed may contain one or more duplicate, inverted, or repeat sequences that do not exist in nature, and as a result, the entire nucleic acid sequence may not exist in nature. The nucleic acids disclosed may contain modified nucleotides, artificial nucleotides, or synthetic nucleotides that do not exist in nature, and the entire nucleic acid sequence may not exist in nature.

[0223] When the genetic code contains redundancy, multiple nucleotide sequences can encode a particular protein. All such nucleotide sequences are assumed herein.

[0224] As used throughout this disclosure, the term “promoter” refers to a synthetic or naturally occurring molecule that can confer, activate, or enhance the expression of nucleic acids within a cell. A promoter may include one or more specific transcriptional regulatory sequences to further enhance expression and / or alter its spatial and / or temporal expression. A promoter may also include distal enhancer or repressor elements located thousands of base pairs away from the transcription start site. Promoters may originate from viruses, bacteria, fungi, plants, insects, animals, and the like. Promoters can constitutively or differentially control the expression of gene components with respect to the cell, the tissue or organ in which expression occurs, or the developmental stage in which expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions, or inducers. Representative examples of promoters include the bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, EF-1 alpha promoter, CAG promoter, SV40 early promoter or SV40 late promoter, and CMV IE promoter.

[0225] As used throughout this disclosure, the term “vector” refers to a nucleic acid sequence containing an origin of replication. Examples of vectors include viral vectors, bacteriophages, bacterial artificial chromosomes, and yeast artificial chromosomes. A vector may be either a DNA vector or an RNA vector. A vector may be a self-replicating extrachromosomal vector, preferably a DNA plasmid. A vector consists of amino acids and a DNA sequence, an RNA sequence, or a combination of both DNA and RNA sequences.

[0226] Conservative substitutions of amino acids, i.e., substitution of an amino acid with a different amino acid having similar properties (e.g., hydrophilicity, degree and distribution of charged regions), are typically recognized in the art as involving only minor changes. These small changes can be partially identified by considering the hydrophobicity index of amino acids, as understood in the art. (Kyte et al., J.Mol.Biol.157:105-132 (1982)). The hydrophobicity index of an amino acid takes into account its hydrophobicity and charge. Substitution with amino acids having similar hydrophobicity indices can maintain protein function. In some embodiments, amino acids with hydrophobicity indices of ±2 are substituted. The hydrophilicity of amino acids can also be used to identify substitutions that maintain the biological function of a protein. By considering the hydrophilicity of amino acids in the context of peptides, it is possible to calculate the maximum local mean hydrophilicity of the peptide, which is a useful indicator that has been reported to correlate well with antigenicity and immunogenicity. U.S. Patent No. 4,554,101 is fully incorporated herein by reference.

[0227] By substituting amino acids with similar hydrophilicity values, peptides with preserved biological activity, such as immunogenicity, can be obtained. Substitutions can be performed with amino acids whose hydrophilicity values ​​are within ±2 of each other. Both the hydrophobicity index and hydrophilicity value of an amino acid are influenced by its specific side chain. Consistent with this observation, it is understood that amino acid substitutions suitable for biological function depend on the relative similarity of the amino acids, particularly their side chains, as revealed by their hydrophobicity, hydrophilicity, charge, size, and other properties.

[0228] As used herein, “conservative” amino acid substitutions may be defined as shown in Tables 4, 5, and 6 below. In some embodiments, fusion polypeptides and / or nucleic acids encoding such fusion polypeptides include conservative substitutions introduced by modifying the polynucleotides encoding the polypeptides of this disclosure. Amino acids can be classified by their physical properties and their contribution to the secondary and tertiary structures of proteins. A conservative substitution is the replacement of one amino acid with another amino acid having similar properties. Exemplary conservative substitutions are shown in Table 4. [Table 4]

[0229] Alternatively, conserved amino acids can be classified as shown in Table 5, as presented by Lehninger (Biochemistry, Second Edition, Worth Publishers, Inc. NY, NY (1975), pp. 71-77). [Table 5]

[0230] Alternatively, exemplary conservative substitutions are shown in Table 6. [Table 6]

[0231] The polypeptides and proteins of this disclosure may have sequences, either whole or in part, that do not exist in nature. The polypeptides and proteins of this disclosure may contain one or more mutations, substitutions, deletions, or insertions that do not exist in nature, and the entire amino acid sequence may not exist in nature. The polypeptides and proteins of this disclosure may contain one or more duplicate, inverted, or repeat sequences, and as a result, the sequence may not exist in nature, and the entire amino acid sequence may not exist in nature. The polypeptides and proteins of this disclosure may contain modified amino acids, artificial amino acids, or synthetic amino acids that do not exist in nature, and the entire amino acid sequence may not exist in nature.

[0232] When used throughout this disclosure, the identity between two sequences may be determined using a standalone executable BLAST engine program (bl2seq) for blasting two sequences, which can be obtained from the National Center for Biotechnology Information (NCBI) ftp site using default parameters (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250, which is incorporated herein by reference in its entirety). When used in the context of two or more nucleic acid or polypeptide sequences, the terms “identical” or “same” refer to a specific percentage of the same residues across a particular region of each sequence. In some embodiments, sequence identity is determined across the entire length of the sequences. The percentage can be calculated by optimally aligning the two sequences, comparing the two sequences over a given region, determining the number of positions where identical residues appear in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the given region, and multiplying the result by 100 to obtain the percentage of sequence identity. If the two sequences have different lengths, or if alignment generates sequences with one or more ends shifted, and the specified comparison region contains only a single sequence, the residues of the single sequence are included in the denominator but not in the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be performed manually or using computer sequencing algorithms such as BLAST or BLAST 2.0.

[0233] In certain embodiments, if a sequence has a specific sequence identity with respect to a particular sequence number (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%), then the sequence and the sequence of that sequence number have the same length. In certain embodiments, if a sequence has a specific sequence identity with respect to a particular sequence number (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%), then the sequence and the sequence of that sequence number differ solely due to conservative amino acid substitutions.

[0234] As used throughout this disclosure, the term “endogenous” refers to a nucleic acid or protein sequence that is naturally associated with the target gene or the host cell into which it is introduced.

[0235] As used throughout this disclosure, the term “exogenous” means nucleic acids or protein sequences that are not naturally associated with the target gene or the host cell into which it is introduced, and includes multiple copies of naturally occurring nucleic acids that are not naturally occurring, such as DNA sequences, or naturally occurring nucleic acid sequences located at genomic locations that are not naturally occurring.

[0236] This disclosure provides a method for introducing a polynucleotide construct containing a DNA sequence into a host cell. “Introducing” means presenting the polynucleotide construct to the cell in a manner that allows access to the interior of the host cell. The method of this disclosure does not depend on a specific method for introducing the polynucleotide construct into a host cell, but only on the polynucleotide construct accessing the interior of a single host cell. Methods for introducing polynucleotide constructs into bacteria, plants, fungi, and animals are known in the art, but are not limited to stable transformation methods, transient transformation methods, and virus-mediated methods. [Examples]

[0237] The examples in this section are provided for illustrative purposes only and are not intended to limit the invention.

[0238] Example 1 - Construction of Super piggyBac transposase containing tandem cysteine-rich domain (CRD) This example demonstrates the construction of an exemplary Super piggyBac (SPB) transposase containing a double C-terminal CRD.

[0239] Using the amino acid sequence of SPB transposase containing the N-terminal nuclear localization sequence (NLS, SEQ ID NO: 1), an SPB transposase sequence containing an additional C-terminal CRD was constructed. Amino acid residues 542-594 (SEQ ID NO: 2) of SPB transposase containing the SPB CRD domain were ligated to the C-terminus of the SPB transposase sequence using the AGGG linker (SEQ ID NO: 27) to generate an SPB dual CRD transposase (SEQ ID NO: 3) containing the N-terminal NLS. As shown in Example 3, the integration and excision activities of SPB transposase containing the dual CRD were compared with those of wild-type SPB transposase using wild-type and symmetry-inverted terminal repeats (ITRs).

[0240] Example 2 - Construction of Transposons Containing Symmetry-Inverted Terminal Repeat (ITR) Mutants This example shows the construction of transposons containing symmetry ITRs for use with SPB transposase and site-specific SPB / PBx transposase containing dual CRDs.

[0241] SPB transposase binds asymmetrically to the 35 bp LE ITR (SEQ ID NO: 4) and 63 bp RE ITR (SEQ ID NO: 5) of the transposon, while SPB transposase containing a double CRD symmetrically recognizes and binds to a transposon having LE ITR sequences at both ends of the transposon. Therefore, the 63 bp RE ITR of the transposon was replaced with a second copy of the 35 bp LE ITR (SEQ ID NO: 6), called symmetric ITR or symmetric ITR 0 bp. Considering that all four CRD domains of two double CRD SPB dimers are in close proximity during ITR binding, several versions of symmetric ITR were designed to examine possible steric hindrance by shifting the CRD binding site 1 bp, 2 bp, or 3 bp from the DDBD binding site to generate symmetric ITR 1 bp (SEQ ID NO: 7), symmetric ITR 2 bp (SEQ ID NO: 8), and symmetric ITR 3 bp (SEQ ID NO: 9). Since the LE ITR 10 bp DDBD binding site (SEQ ID NO: 11) has a one-base pair difference from the RE ITR 10 bp DDBD binding site (SEQ ID NO: 12), a second version of the symmetric ITR 3 bp version was constructed (symmetric ITR 3 bp SNP, SEQ ID NO: 10) to examine the sequence difference between the LE binding site and the RE DDBD binding site. The integration and excision activities of transposons containing wild-type ITR or symmetric ITR were analyzed using a double excision / integration luciferase reporter assay.

[0242] Example 3 - Method for Measuring the Excision and Integration Activities of SPB Transposase Containing a Double CRD Domain and Transposons Containing Symmetric Inverted Terminal Repeat (ITR) Mutants This example shows a method for measuring the excision activity and integration activity of SPB transposase containing a double CRD using a transposon containing wild-type or symmetric ITR.

[0243] A double excision / integration luciferase reporter (SEQ ID NO: 13) was used to test SPB transposases containing double CRDs of transposons, including wild-type or symmetric ITRs (see Figure 1). Figure 1 shows schematic diagrams of the double reporter plasmid designs used to confirm excision and integration rates for each mutant transposon. Using the H-2kk GFP transposon reporter (Reporter 1), increased H2kk expression was observed with increased transposon excision. Using Reporter 2, increased GFP expression was observed with increased transposon integration. In an alternative design for Reporter 2, increased firefly luciferase expression was observed with increased transposon excision, and increased NanoLuc was observed with increased transposon integration. As schematically shown in Figure 2, the wild-type 63bp RE ITR sequence of the transposon in the reporter was replaced with each of the symmetric ITRs (SEQ ID NOs: 6-10). K562 cells were nucleofected using 20 μl of SF buffer and programmed FF-120 according to the manufacturer's instructions. Each reactant contained 50 ng of a dual luciferase reporter and 450 ng of a dual CRD (SEQ ID NO: 3) or SPB (SEQ ID NO: 1) expression plasmid. As a negative control, each dual luciferase reporter was transfected without a transposase. One day after transfection, the luciferase signal was measured using Promega's dual luciferase reagent and plate reader. The results are shown in Table 7. [Table 7]

[0244] As shown in Table 7, a strong luciferase signal indicating transposon excision and integration was detected with SPB transposases combined with reporters containing wild-type ITR sequences, but not with reporters containing any of the symmetric ITR sequences. SPB transposases containing a dual CRD domain yielded a luciferase signal in combination with both wild-type ITRs and all symmetric ITR designs, but expression was detected at lower levels than with SPB transposases. The symmetric ITR 2bp version (SEQ ID NO: 8) yielded the highest signal among all symmetric ITR designs using SPB transposases containing a dual CRD domain. Furthermore, both the LE ITR 10bp DDBD binding site (SEQ ID NO: 11) and the RE ITR 10bp DDBD binding site (SEQ ID NO: 12) were functional for transposons containing symmetric ITRs, as observed for symmetric ITR 3bp (SEQ ID NO: 9) and symmetric ITR 3bp SNP (SEQ ID NO: 10) samples. In the absence of the transposase, the reporter construct resulted in only background-level luciferase expression.

[0245] Example 4 - Construction and analysis of a TAL array-piggyBac double CRD transposase (ss-PBx-CRD) composition (TAL-PBx-CRD) designed for site-directed transposition in a specific gene. This example illustrates the construction of a TAL array-Super piggyBac PBx transposase containing a dual CRD domain fusion protein composition (TAL-PBx-CRD) useful in achieving site-directed transposition at a specific target gene locus.

[0246] TAL-PBx transposases targeting LINE1 repeat element LINE1 L2 TAL-PBx (SEQ ID NO: 15) and LINE1 R2.2 TAL-PBx (SEQ ID NO: 16), which contains a (delta 1-85) PBx fusion site (d85+73) with the C-terminus and N-terminus of +73 TAL, were pre-constructed (see, for example, International Publication No. 2023 / 060089). Two LINE1 TAL-PBx (d85+73) constructs were modified by adding an additional CRD domain to the C-terminus of the PBx transposase sequence. In one example, LINE L2 TAL-PBx(d85+73) (SEQ ID NO: 15) and LINE R2.2 TAL-PBx(d85+73) (SEQ ID NO: 16) were modified by adding AGGG linker (SEQ ID NO: 27) and amino acids 547-594 (SEQ ID NO: 2), which contain the SPB transposase CRD domain, to create LINE L2 TAL-PBx(d85+73) double CRD (SEQ ID NO: 17) and LINE R2.2 TAL-PBx(d85+73) double CRD (SEQ ID NO: 18). Furthermore, a second left-right pair of LINE1 was constructed, and amino acids 542-594 of the SPB containing the CRD domain (sequence number 14) were directly added to LINE L2 TAL-PBx(d85+73) (sequence number 15) and LINE R2.2 TAL-PBx(d85+73) (sequence number 16) to create LINE L2 TAL-PBx(d85+73) bilayer CRD (sequence number 19) and LINE R2.2 TAL-PBx(d85+73) bilayer CRD (sequence number 20).

[0247] Using a disrupted GFP reporter system, the LINE1 TAL-PBx-double CRD transposase was tested in combination with a transposon containing a symmetric ITR. Briefly, this reporter system contains an EF1a promoter (SEQ ID NO: 22) that drives the expression of a GFP CDS reporter (SEQ ID NO: 23), followed by an SV40 polyadenylated sequence (SEQ ID NO: 24). The GFP reporter is disrupted at the TTAA sequence by the transposon and fragmented into a first GFP moiety (SEQ ID NO: 25) and a second GFP moiety (SEQ ID NO: 26). Upon transposase-mediated excision of the transposon, the reporter can be seamlessly repaired, resulting in the recovery of a full-length GFP CDS and subsequent GFP expression as a substitute for excision activity. The wild-type 63bp RE ITR sequence in the reporter was replaced with a 0bp symmetric ITR (SEQ ID NO: 6), a symmetric ITR with a 1bp spacer, and a symmetric ITR 1bp SNP (SEQ ID NO: 21), which is the RE version of the 10bp DDBD binding site.

[0248] Each reporter construct was co-transfected into HEK293T cells using TAL-ssSPB-CRD transposase. For control, the TAL-ssSPB-CRD transposase was replaced with either PiggyBac(PBx) excised only, or with the original wild-type ssSPB without the transposase. Briefly, one day prior to transfection, 120,000 HEK293T cells were seeded in 500 μL of DMEM + 10% FBS using a 24-well plate. 50 ng of transposase expression vector pairs were combined with 450 ng of reporter and transfected using 1 μL of JetPrime transfection reagent. Two days after transfection, the percentage of GFP-positive cells was measured by flow cytometry. The results are shown in Table 8. [Table 8]

[0249] As shown in Table 8, both TAL-PBx-CRD designs yielded higher resection than TAL-PBx transposases containing a single CRD version and exhibited comparable resection activity to PBx transposases. Furthermore, PBx and single-CRD ssSPB transposases yielded higher resection activity using wild-type ITR reporters than those containing symmetric ITR reporters, although the TAL-PBx version containing a dual CRD domain yielded higher resection than the wild-type ITR reporter for symmetric ITR reporters. Negative controls without transposases had background levels of GFP signaling. Genomic DNA was collected from the transfected cells described above, and site-specific integration of transposons into the LINE1 target site was quantified by ddPCR. Table 9 shows the forward transposon integration events per haploid genome. [Table 9]

[0250] As shown in Table 9, using wild-type TAL-PBx transposase containing a single CRD, approximately 3–5 integration events were detected per haploid genome, with slightly more integration events occurring for transposons containing wild-type ITRs than for symmetric ITRs. Using TAL-PBx transposase containing a dual CRD domain resulted in more integrations, leading to 6–9 transposon integrations per haploid genome with wild-type ITRs, and with symmetric ITRs, the number of transposon integrations increased to 9–13 per haploid genome. Two versions of the TAL-PBx dual CRD fusion protein yielded similar numbers of integrations.

[0251] Example 5 - Construction of Super PiggyBac transposase containing tandem cysteine-rich domain (CRD) and N-terminal deletion. This example demonstrates the construction of an exemplary Super piggyBac (SPB) transposase including a double C-terminal CRD and an N-terminal deletion (NTD) of the 74 amino acids at the very N-terminus of SPB.

[0252] Using the amino acid sequence of an SPB transposase containing an N-terminal nuclear localization sequence (NLS, SEQ ID NO: 1), an SPB transposase sequence containing an additional C-terminal CRD and a 74-amino acid NTD was constructed. Amino acid residues 542-594 (SEQ ID NO: 2) of the SPB transposase containing the SPB CRD domain were ligated to the C-terminus of the SPB transposase sequence using an AGGG linker to generate an SPB double CRD transposase containing the N-terminal NLS. As shown in Example 3, the integration and cleavage activity of the SPB transposase containing the double CRD was compared with that of the wild-type SPB transposase using wild-type and symmetric inverted-end repeat sequences (ITRs).

[0253] Example 6: Identification of piggyBac thermal stability mutations A computational algorithm was used to generate a list of point mutations that could potentially improve the protein thermal stability of PiggyBac transposase. For example, chain C of the Piggybac cryoEM structure 6X67 was used as input to the algorithm FireProt2.0 (https: / / loschmidt.chemi.muni.cz / fireprotweb / ). As output, the amino acid at each position predicted to yield the highest thermal stability was calculated. Table 10 lists the positions where the prediction of the most thermally stable amino acid differs from the wild-type sequence. [Table 10]

[0254] Each point mutation was individually cloned into a TAL-PBx transposase targeting the LINE1 R2.2 TAL-PBx (sequence number 16), which is a LINE1 repeat element with the +73 TALC terminus and N terminus deleted, generating sequences 99 to 129.

[0255] To test the ability of the new variant to catalyze site-specific translocation, an "all-in-one site-specific excision / insertion episomal reporter" system was constructed. This episomal reporter system includes a plasmid that contains the piggyBac transposon donor together with the transposon integration site on the same plasmid. The transposon in this plasmid disrupts the open reading frame of GFP preceded by the EF1a promoter and followed by a polyadenylation signal sequence. The vector also contains, in the opposite orientation, a polyA and transcription termination site, a TTAA integration site adjacent to the LINE1 R2.2 right target sequence (SEQ ID NO: 130) and a 13bp spacer, followed by a PEST destabilized mScarlet reporter gene and a polyadenylation signal sequence. This "all-in-one site-specific excision / insertion episomal reporter" (SEQ ID NO: 131) does not express GFP and hardly or does not express mScarlet when transfected only into cells. Upon transposon excision catalyzed by SPB, PBx or ssSPB, the GFP coding sequence is restored and GFP is expressed. When the CMV promoter containing the transposon is site-specifically integrated into its target site upstream of the mScarlet gene, mScarlet is expressed at levels above background. The design of the reporter is described in more detail in International Publication No. WO 2023 / 060089.

[0256] Each TAL-PBx SSM mutant expression vector was co-transfected into HEK293T cells with an all-in-one site-specific excision / integrated episomal reporter. Briefly, a transfection mix containing 50 ng of mutant TAL-PBx, 50 ng of reporter plasmid, and 0.3 μl of Transit2020 transfection reagent was prepared in 20 pl serum-free OptiMEM medium. Approximately 60,000 HEK293T cells in 180 μl of DMEM medium supplemented with 10% FBS were added to this mix, and then 80 μl of this transfection mixture was seeded in double rows into a clear-bottom 96-well plate and incubated at 37°C, 5% CO2. As a control, benchmark TAL-ssSPB (SEQ ID NO: 16), TAL-ssSPB not specific to the LINE1 target, or catalytically killed transposase were transfected instead of mutant TAL-ssSPB. GFP and mScarlet fluorescence were detected using an Incucyte live cell analyzer. The fluorescence cell percentages for the excision (GFP) reporter and the site-specific integration (mScarlet) reporter are shown in Figure 3 and Table 11, respectively. [Table 11]

[0257] As shown in Figure 3 and Table 11, the benchmark TAL-ssSPB yielded excision activity and site-directed integration activity in approximately 40% and 24% of cells, respectively. Off-target and catalytically inactive controls showed no site-directed integration activity at all. The M298L mutant of TAL-ssSPB exhibited higher excision and site-directed integration activity than the benchmark TAL-ssSPB construct, while the integration activity of all other mutants was similar to or lower than the benchmark.

[0258] Example 7: Construction and evaluation of ssSPB with highly active and thermally stable mutations. In this example, the effects of combining highly active and thermally stable mutations with the bipolar CRD TAL-ssSPB were investigated. The LINE L2 TAL-ssSPB(d85+73) bipolar CRD (SEQ ID NO: 17) and LINE R2.2 TAL-ssSPB(d85+73) bipolar CRD (SEQ ID NO: 18) from Example 4 were modified to incorporate the PBx highly active mutation into the ssSPB. Specifically, the R372H mutation was introduced into each TAL-ssSPB to create SEQ ID NOs: 142 and 143, the S103P / S509G / N571S mutation was introduced to create SEQ ID NOs: 144 and 145, and the S103P / S509G / N571S / R372H mutation was introduced to create SEQ ID NOs: 146 and 147.

[0259] Each mutant pair of TAL-ssSPB was co-transfected into HEK293T cells with a transposon donor containing a symmetric ITR and PuroR-2A-GFP cargo (SEQ ID NO: 152). As a control, the mutant TAL-ssSPB-CRD transposase was replaced with the original wild-type ssSPB or PBx. Briefly, one day before transfection, 120,000 HEK293T cells or HepG2 cells were seeded in 500 μL of DMEM + 10% FBS using a 24-well plate. 50 ng of transposase expression vector pairs were combined with 450 ng of reporter and transfected using 1 μL of JetPrime transfection reagent. Two days after transfection, genomic DNA was collected from the cells, and site-specific integration of forward and reverse transposons at the LINE1 target was detected and quantified by ddPCR. The results are shown in Table 12. [Table 12]

[0260] As shown in Table 12, increased site-specific integration was observed using the R372H and S103P / S509G / N571S mutants compared to the benchmark TAL-ssSPB. The highest site-specific integration was observed when all mutants were combined (S103P / S509G / N571S / R372H).

[0261] In another experiment, the LINE L2 TAL-ssSPB(d85+73) double CRD (SEQ ID NO: 17) and LINE R2.2 TAL-ssSPB(d85+73) double CRD (SEQ ID NO: 18) from Example 4 were modified to incorporate PBx high-activity mutations and thermal-stable mutations into the ssSPB. Specifically, the M298L mutation was introduced into each TAL-ssSPB to create SEQ ID NOs: 148 and 149, and the S103P / S509G / N571S / R372H / M298L mutation was introduced to create SEQ ID NOs: 150 and 151.

[0262] In K562 and HEK293T cells, the site-directed integration activity of dual CRD TAL-ssSPBs (SEQ ID NOs. 150 and 151) containing five mutations (S103P / S509G / N571S / R372H / M298L) was compared with that of dual CRD TAL-ssSPBs lacking the S103P / S509G / N571S / R372H / M298L mutations (SEQ ID NOs. 17-18) and single CRD TAL-ssSPBs without the S103P / S509G / N571S / R372H / M298L mutations (SEQ ID NOs. 15-16). Integration per haploid genome is shown in Table 13 as the sum of forward and reverse integrations. [Table 13]

[0263] As shown in Table 13, increased site-specific integration was observed when using dual CRDs compared to single CRDs (TAL-ssSPB). The highest level of site-specific integration was observed using dual CRDs containing five mutations (S103P / S509G / N571S / R372H / M298L).

[0264] Example 8: Construction and evaluation of TAL-ssSPB using shorter and longer TAL arrays In this experiment, a dual TAL-ssSPB construct containing five mutations (S103P / S509G / N571S / R372H / M298L) was constructed to target TTAA in the human B2M gene at the sequence: CAGGCAGGATGAATCTGTGCTCTGATCCCTGAGGCATTTAATATGTTCTTATTATT AGAAGCTCAGATGCAAAGAGCT (SEQ ID NO: 153). Specifically, the TAL array was designed to bind to a sequence 13 bp or 14 bp away from the TTAA integration site. Upstream of TTAA, the TAL array was designed to bind to targets of 9 bp, 10 bp, 13 bp, and 14 bp length. Downstream of TTAA, the TAL array was designed to bind to targets of 9 bp, 10 bp, 11 bp, 12 bp, 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, and 20 bp length. Table 14 lists TAL target sequences including the 5'T preceding the target, along with the length of the TAL target (bp) and the distance from the TTAA integration site (bp), and is shown in Figure 4. [Table 14]

[0265] HEK293T cells were co-transfected with TAL-ssSPBs in pairs (one upstream, one downstream) along with transposon donors containing symmetric ITRs. Briefly, 120,000 HEK293T cells were seeded in 500 μL of DMEM + 10% FBS in a 24-well plate one day prior to transfection. A 50 ng transposase expression vector pair was combined with 450 ng of reporter and transfected using 1 μL of JetPrime transfection reagent. In one experiment, either a 9 bp or 10 bp upstream TAL-ssSPB was co-transfected with each of the downstream TAL-ssSPBs. In another experiment, either a 13 bp or 14 bp upstream TAL-ssSPB was co-transfected with each of the 11 bp-20 bp downstream TAL-ssSPBs. As a negative control, TAL-ssSPBs were replaced with PBx. Two days after transfection, genomic DNA was recovered from the cells, and site-specific transposon integration in a single orientation at the LINE1 target was detected and quantified by ddPCR. The results are shown in Tables 15 and 16. [Table 15] [Table 16]

[0266] As shown in Tables 15 and 16, all constructed TAL-ssSPB pairs, including TAL arrays of various lengths, catalyzed site-specific integration into B2M genomic regions at a higher level than PBx controls.

[0267] Example 9: In vivo editing of mouse liver with TAL-ssSPB having a dual CRD, a highly active mutation, and a thermal stability mutation. In this example, the LINE1 locus was edited in vivo in mouse hepatocytes using TAL-ssSPB, which has a dual CRD and highly active and thermally stable mutations. The TTAA within open reading frame 1 of the L1Mda family of the mouse LINE1 sequence was selected as the target site. A 10 bp TAL binding site was identified on both sides of the TTAA, separated from the TTAA by a 13 bp spacer. The target site was: The sequence number is TIFF2026517869000036.tif7170 (SEQ ID NO: 175), with the TAL binding site underlined and TTAA shown in bold. A TAL array (SEQ ID NO: 170) that binds to the left site (CAAGAAGGAC) and a TAL array (SEQ ID NO: 171) that binds to the inverted complement of the right site (TAGCAGTGCT) were designed. As described above, the TAL arrays were constructed in a TAL-ssSPB expression construct having a double CRD and high-activity and thermally stable mutations (S103P / S509G / N571S / R372H / M298L) to generate mLINE1 left TAL-ssSPB (SEQ ID NO: 172) and mLINE1 right TAL-ssSPB (SEQ ID NO: 173). A nanoplasmid consisting of a promoter that drives the expression of the firefly luciferase gene was constructed as an in vivo transposon donor (SEQ ID NO: 174). The firefly luciferase gene was interrupted by a piggyBac transposon with symmetric ITR and PCR primer binding sites to prevent functional luciferase expression in its initial state. During transposase-catalyzed excision of the transposon and seamless repair of its skeleton, the luciferase gene was repaired, resulting in quantifiable luciferase expression via luminescence readout. Site-specific integration of the excised transposon can then be detected and quantified using a PCR approach.

[0268] To perform in vivo site-directed transposition, 25 ug of plasmid DNA was delivered to 10- to 12-week-old female Balb / C mice by hydrodynamic delivery (HDD). Briefly, the HDD injection consisted of 25 ug of plasmid DNA in 2 mL of Trans-ITEE buffer (Mirus Bio) injected into the lateral tail vein over 3-5 seconds. The 25 ug dose of plasmid DNA consisted of 12.5 ug of in vivo transposon donor DNA, 7.5 ug of left mLINE1 TAL-ssSPB expression construct, and 7.5 ug of right mLINE1 TAL-ssSPB expression construct. As negative controls, mice in other groups were injected with buffer alone (PBS) or a mutant expression vector (PBx) consisting of 12.5 ug of transposon donor DNA and 12.5 ug of piggyBac excision alone. Three days after delivery, in vivo bioluminescence imaging (BLI) was performed to quantify the amount of transposon excised from donor DNA catalyzed by transposase (TAL-ssSPB or PBx). Mice were injected with 50 μl of IVISBrite D-luciferin (Revvity Health Sciences), and after 6 minutes, the luciferase signal was measured as total photon flux / second using an IVIS imager (Perkin Elmer) with an automated exposure time. After BLI, the mice were euthanized, and the left lobe of the liver was collected. Genomic DNA was extracted from the collected liver tissue for quantification of site-specific integration into the mouse LINE1 locus by digital droplet PCR (ddPCR). PCR amplicons generated from primers within transposons and primers adjacent to the insertion site of the LINE1 genomic DNA target were used to measure site-specific integration. Whole genomic DNA introduced into the PCR reaction was quantified using a second PCR amplicon generated from two primers in an untargeted genomic region. Table 17 shows the site-specific integration frequencies, expressed as integrations per 100 haploid genomes, along with BLI measurements. As seen in Table 17, transposon excision was detected at similar levels in the PBx and TAL-ssSPB groups. As expected, site-specific transposon integration into the mouse LINE1 target site was detected only in the TAL-ssSPB group. Table 17

Claims

1. A fusion protein comprising, in N-terminus to C-terminus, a nuclear localization signal (NLS), a DNA targeting domain, and a first transposase domain containing one of the amino acid sequences of SEQ ID NOs. 28-69.

2. The fusion protein according to claim 1, wherein the first transposase domain comprises any of the amino acid sequences of SEQ ID NOs. 28 to 48.

3. The fusion protein according to claim 1, wherein the first transposase domain comprises any of the amino acid sequences of SEQ ID NOs: 49 to 69.

4. The fusion protein according to claim 2, wherein the first transposase domain comprises the amino acid sequence of SEQ ID NO: 30 or 38.

5. The fusion protein according to claim 3, wherein the first transposase domain comprises the amino acid sequence of SEQ ID NO: 51 or 59.

6. The fusion protein according to any one of claims 1 to 5, wherein the DNA targeting domain comprises one or more zinc finger motifs.

7. The fusion protein according to claim 1 or 2, wherein the DNA targeting domain includes the sequence of Sequence ID No.

74.

8. The fusion protein according to any one of claims 1 to 5, wherein the DNA targeting domain comprises one or more TAL domains.

9. The fusion protein according to claim 8, wherein the DNA targeting domain binds to a nucleic acid sequence encoding GFP, a LINE1 repeat element, or a zinc finger 268 (ZFM268) binding site.

10. It is a fusion protein, (a) TAL array and (b) A modified Super piggyBac transposase ("SPB") comprising an N-terminal deletion and a modified SPB transposase comprising a second cysteine-rich domain (CRD) fused to the C-terminus of the SPB transposase, A fusion protein in which the C-terminus of the TAL array is fused to the N-terminal amino acid of the N-terminal deletion SPB to produce a TAL array-N-terminal deletion SPB fusion protein.

11. The fusion protein according to claim 10, further comprising a GS or GGGS linker positioned between the TAL array and the N-terminal deletion SPB.

12. The fusion protein according to claim 10 or 11, wherein the SPB includes an N-terminal deletion including a deletion of amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103.

13. It is a fusion protein, (a) TAL array and (b) A modified Super piggyBac transposase ("SPB") comprising an N-terminal deletion, one or more embedded deletion PBx mutations, and a second cysteine-rich domain (CRD) fused to the C-terminus of the SPB transposase, A fusion protein in which the C-terminus of the TAL array is fused to the N-terminal amino acid of the N-terminal deletion SPB to produce a TAL array-N-terminal deletion SPB fusion protein.

14. The fusion protein according to claim 13, further comprising a GS or GGGS linker positioned between the TAL array and the N-terminal deletion SPB.

15. The fusion protein according to claim 13 or 14, wherein the modified Super piggyBac transposase comprises any of the amino acid sequences of SEQ ID NOs. 28 to 69.

16. A fusion protein comprising, in order from the N-terminus to the C-terminus, a DNA targeting domain and a first transposase domain containing the sequence shown in SEQ ID NO: 1 or 3, wherein the first transposase domain contains a deletion of the 83 to 103 most N-terminal amino acids of SEQ ID NO: 1 or 73.

17. The fusion protein according to claim 16, wherein the DNA targeting domain comprises one or more zinc finger motifs.

18. The fusion protein according to claim 16, wherein the DNA targeting domain comprises one or more TAL domains.

19. The fusion protein according to any one of claims 16 to 18, wherein the DNA targeting domain is bound to a nucleic acid sequence encoding GFP, zinc finger 268 (ZFM268), phenylalanine hydroxylase (PAH), beta-2-microglobulin (B2M), or a LINE1 repeat element.

20. The fusion protein according to any one of claims 16 to 19, wherein the first transposase domain and the DNA targeting domain are linked by a linker.

21. The fusion protein according to claim 20, wherein the linker comprises the sequence GGGGS.

22. The fusion protein according to any one of claims 16 to 21, wherein the first transposase domain includes an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103.

23. The fusion protein according to any one of claims 16 to 22, wherein the first transposase domain comprises (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K and D201R, or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E and R504D for SEQ ID NO: 1 or 73, and the numbering begins at the 12th residue of SEQ ID NO: 1 and the 1st residue of SEQ ID NO:

73.

24. The fusion protein according to any one of claims 16 to 23, further comprising a second transposase domain at the C-terminus of the first transposase domain, wherein the second transposase domain comprises the sequence shown in SEQ ID NO: 1 or 73.

25. The fusion protein according to claim 24, wherein the second transposase domain includes a deletion of N-terminal amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103 of SEQ ID NO: 1 or 73.

26. The fusion protein according to any one of claims 16 to 25, wherein the second transposase domain comprises (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K and D201R, or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E and R504D for SEQ ID NO: 1 or 73, and the numbering begins at the 12th residue of SEQ ID NO: 1 and the 1st residue of SEQ ID NO:

73.

27. A polynucleotide comprising a nucleic acid sequence encoding a fusion protein according to any one of claims 1 to 26.

28. A transposon comprising a symmetrical left-end (LE) inverted-end repeat sequence (ITR) and a right-end (RE) inverted-end repeat sequence (ITR), wherein the nucleotide sequences of the LE ITR and RE ITR are each sequence number 6.

29. A transposon comprising a symmetrical left-end (LE) inverted-end repeat sequence (ITR) and a right-end (RE) inverted-end repeat sequence (ITR), wherein the nucleotide sequence of the LE ITR includes SEQ ID NO: 6, and the nucleotide sequence of the RE ITR includes SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 21, or SEQ ID NO:

98.

30. A transposon according to claim 28 or 29, comprising a nucleotide sequence encoding a therapeutic protein.

31. The transposon according to claim 30, comprising a promoter sequence that controls the expression of the therapeutic protein.

32. A method for site-specific insertion of a therapeutic gene into one or more genomic loci of a cell, comprising simultaneously introducing a transposon according to any one of claims 28 to 30 and a polynucleotide according to claim 27 into the cell.

33. A method for modifying the genome of a cell, comprising providing the cell with a fusion protein according to any one of claims 1 to 26, wherein the cell includes a modified binding site comprising, in the order 5' to 3', an inverted sequence of the target site relative to the DNA targeting domain, a first spacer, a TTAA target integration site for the SPB, a second spacer, and a complement of the sequence of the target site relative to the DNA targeting domain.

34. A fusion protein comprising, in order from the N-terminus to the C-terminus, a nuclear localization signal (NLS), a DNA targeting domain, and a first transposase domain containing one of the amino acid sequences of SEQ ID NOs. 132 to 141.

35. The fusion protein according to claim 34, wherein the DNA targeting domain comprises one or more TAL domains.

36. The fusion protein according to claim 35, wherein the DNA targeting domain binds to a nucleic acid sequence encoding GFP, a LINE1 repeat element, or a zinc finger 268 (ZFM268) binding site.

37. A fusion protein according to claim 35 or 36, comprising the sequences shown in SEQ ID NOs: 142-151.