Transposase and its use

A fusion protein with a transposase and DNA targeting domain addresses the need for site-specific gene editing in the LPA gene, enabling efficient and targeted integration for gene therapy applications.

JP2026515653APending Publication Date: 2026-05-19POSEIDA THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
POSEIDA THERAPEUTICS INC
Filing Date
2024-04-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

The need for site-specific transposases for gene editing is not adequately met by existing methods, particularly for targeting the lipoprotein (LPA) gene, which has multiple copies of target sites and is highly expressed in hepatocytes, making it a potential target for gene therapy.

Method used

A fusion protein comprising a transposase domain and a DNA targeting domain, specifically designed to bind to the LPA gene, is used to introduce transgenes site-specifically into the LPA gene, utilizing a linker and incorporating a detectable marker like GFP.

Benefits of technology

The fusion protein enables efficient and targeted integration of transgenes into the LPA gene, enhancing gene editing capabilities and providing a method for generating engineered cells with stable genomic modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515653000001_ABST
    Figure 2026515653000001_ABST
Patent Text Reader

Abstract

This disclosure generally relates to a fusion protein comprising a transposase domain and a DNA targeting domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 494,306, filed on April 5, 2023, which is hereby incorporated by reference in its entirety.

[0002] Reference to Electronically Submitted Sequence Listing This application includes a sequence listing submitted in XML format via Patent Center, which is hereby incorporated by reference in its entirety. The aforementioned XML copy, created on March 19, 2024, is named "POTH - 083_001WO_SeqList" and has a size of 237,142 bytes.

[0003] Field The present disclosure generally relates to fusion proteins comprising a transposase domain and a DNA targeting domain. Methods of using the fusion proteins for site - specific translocation are also provided.

Background Art

[0004] Background Transposases can be used to introduce non - endogenous DNA sequences into genomic DNA and are advantageous in many respects over other gene editing methods. However, the need for, for example, site - specific transposases for use in gene editing is not met.

[0005] The lipoprotein (a), or LPA, gene evolved from a replication event of the adjacent plasminogen (PLG) gene. This replication event occurred during primate evolution approximately 40 million years ago. Both genes contain loop structures known as kringle domains. In LPA, the kringle domains are replicated segmentally, so that each copy of LPA can contain up to 50 copies of kringle domains (Schmidt et al., J Lipid Res. 2016 Aug;57(8):1339-59). At the genomic DNA level, each kringle domain repeat extends to approximately 5.5 kb of DNA and consists of two exons and two introns. The intronic portion of the repeat contains several potential target sites for integration via site-specific transposases.

[0006] LPA is an attractive target for site-directed transposition because it contains multiple copies of the same target site and is likely to incorporate at least one transposon. Furthermore, LPA is highly expressed in hepatocytes, which means it likely has an open chromosomal landscape suitable for supporting editing and high expression of the incorporated transgene. It is a non-essential gene, and knockout is associated with lower cholesterol levels. These traits combined make it a potential target site for gene therapy. [Overview of the project]

[0007] overview In one embodiment, a fusion protein is provided herein comprising a DNA targeting domain and a transposase domain containing the sequence shown in SEQ ID NO: 4, wherein the DNA targeting domain binds to a nucleic acid sequence encoding an LPA repeat element. In some embodiments, the DNA targeting domain comprises one, two, or three zinc finger motifs. In some embodiments, the DNA targeting domain comprises one or more TAL domains. In some embodiments, the TAL domains contain the sequence shown in any one of SEQ ID NOs: 35-38. In some embodiments, the DNA targeting domain binds to a nucleic acid sequence encoding a kringle domain repeat element or intron adjacent to the sequence encoding the kringle domain repeat element in the LPA gene.

[0008] In some embodiments, the transposase domain and the DNA targeting domain are linked by a linker. In some embodiments, the linker includes GGGGS (SEQ ID NO: 181).

[0009] In some embodiments, the DNA targeting domain is inserted into the N-terminus of the transposase domain at a position after the 82nd amino acid and before the 105th amino acid of SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes one or more amino acids in the transposase domain between (and including) the 83rd and 105th amino acids of SEQ ID NO: 4. In some embodiments, the transposase domain includes an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103.

[0010] In some embodiments, the transposase domain includes the sequence shown in any one of sequence numbers 7 to 27. In some embodiments, the transposase domain includes (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R, or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D.

[0011] In another embodiment, polynucleotides comprising nucleic acid sequences encoding the fusion protein described herein are provided herein.

[0012] In another embodiment, a vector comprising a polynucleotide described herein is provided herein.

[0013] In another embodiment, a method for incorporating a transgene into a cellular genomic target site is provided herein, the method comprising introducing a fusion protein and transposon described herein into a cell, wherein the transposon comprises a 5'ITR, a transgene, and a 3'ITR in the order of 5' to 3'. In some embodiments, the transposon further comprises an exogenous promoter between the 5'ITR and the transgene. In some embodiments, the transgene encodes a detectable marker. In some embodiments, the detectable marker is GFP. In some embodiments, the transgene is (a) a gene that is not expressed by the cell prior to the introduction of the fusion protein and transposon, or (b) a gene that exhibits reduced, insufficient, and / or altered expression by the cell prior to the introduction of the fusion protein and transposon.

[0014] In some embodiments, the genomic target site is located on the LPA gene. In some embodiments, the genomic target site is located on a repeat element. In some embodiments, the repeat element is an LPA repeat element. In some embodiments, the genomic target site is located on an intron of the gene. In some embodiments, the genomic target site is located on an intron of the LPA gene. In some embodiments, the cells are in vivo.

[0015] In another embodiment, a method for modifying the genome of a cell is provided herein, the method comprising providing a fusion protein described herein to a cell, wherein the cell provides a fusion protein comprising, in the order 5' to 3', a sequence of a target site for a DNA targeting domain, a first spacer, a TTAA target integration site for an SPB, a second spacer, and a modified binding site comprising the reverse complement of the sequence of the target site for the DNA targeting domain. In some embodiments, the target integration site comprises the sequence TTAA. In some embodiments, the target integration site comprises the nucleic acid sequence shown in any one of SEQ ID NOs: 81 to 88.

[0016] In another embodiment, an integration cassette for site-directed transfer of a nucleic acid to the genome of a cell is provided herein, comprising a nucleic acid comprising or consisting thereof a central transposon ITR integration site TTAA sequence flanked by an upstream TAL array target sequence and a downstream TAL array target sequence, wherein each of the upstream TAL array target sequence and the downstream TAL array target sequence is 12 or 13 base pairs away from the TTAA sequence. In some embodiments, the integration site comprises the sequence TTAA. In some embodiments, the integration site comprises the nucleic acid sequence described in any one of SEQ ID NOs: 81-88. In some embodiments, each of the upstream TAL array target site sequence and the downstream TAL array target site sequence is the same. In some embodiments, each of the upstream TAL array target site sequence and the downstream TAL array target site sequence is different. In some embodiments, each of the upstream TAL array target site and the downstream TAL array target site targets a 7-30 bp sequence of the LPA repeat element.

[0017] In another embodiment, cells comprising an embedded cassette described herein, which is stably incorporated into the cell's genome, are provided herein.

[0018] In another embodiment, a method for site-directed transfer of a DNA molecule to the genome of a cell is provided herein, comprising introducing into a cell containing the embedded cassette described herein a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed in the cell, and a DNA molecule containing a transposon, wherein the expressed fusion protein, by site-directed transfer, incorporates the transposon into the TTAA embedding site of the stably embedded embedded cassette.

[0019] In another embodiment, a method for generating engineered cells by site-directed transposition is provided herein, comprising introducing into a cell containing the embedded cassette described herein a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed in the cell, and a DNA molecule containing a transposon, wherein the expressed fusion protein, by site-directed transposition, incorporates the transposon into the TTAA embedding site of the stably embedded embedded cassette, thereby generating engineered cells. In some embodiments, the sequence TTAA. In some embodiments, the embedding site comprises the nucleic acid sequence shown in any one of sequence numbers 81-88. [Brief explanation of the drawing]

[0020] [Figure 1A-1D] This paper demonstrates the introduction of a DNA-binding domain into a transposase using essential heterodimers.

[0021] [Figure 2] This is a schematic diagram showing a Split GFP splicing site-specific reporter.

[0022] [Figure 3] This is a schematic diagram showing a catalytic ssSPB dimer that binds to excised transposons and recognizes their genomic integration target sites. [Modes for carrying out the invention]

[0023] Detailed explanation A fusion protein comprising a transposase domain and a DNA targeting domain is provided herein. In particular, the DNA targeting domain may target the lipoprotein A (LPA) gene. Methods for constructing the transposase domain and the fusion protein, cells modified using the fusion protein provided herein, and methods for treating such cells are also provided.

[0024] In some embodiments, fusion proteins comprising an SPB or PBx domain and a DNA targeting domain are provided herein. The DNA targeting domain is further described below.

[0025] Transposase domain In one embodiment, a fusion protein comprising one or more transposase domains is provided herein. In some embodiments, the transposase domain is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a hyperfunctional piggyBac transposase domain. In preferred embodiments, the transposase domain is a Super piggyBac® transposase domain (SPB). Non-limiting examples of SPB transposases are described in detail in U.S. Patents 6,218,182, 6,962,810, 8,399,643 and PCT Publication No. WO2010 / 099296, each of which is incorporated herein by reference in whole for examples of transposase domains that may be used in the fusion proteins described herein.

[0026] In some embodiments, the transposase domain is a Super PiggyBac transposase (SPB) domain. The SPB contains one or more hyperactivity mutations compared to the wild-type piggyBac transposase. An exemplary wild-type SPB sequence, including the nuclear localization sequence (NLS), is shown in SEQ ID NO: 1, with the NLS in italics and the hyperactivity mutation in bold. For the purpose of describing deletions and mutations, the sequence numbering of the SPB transposase domain begins at residue 12 of SEQ ID NO: 1.

[0027] (Sequence ID 1).

[0028] In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 1 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 1.

[0029] An exemplary sequence of wild-type SPB transposase lacking the NLS domain is shown in Sequence ID No. 2. For the purpose of describing deletions and mutations, the numbering of SPB transposase domain sequences begins at residue 5 of Sequence ID No. 2. In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in Sequence ID No. 2. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in Sequence ID No. 2 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in Sequence ID No. 2.

[0030] The transposase domains used in the fusion proteins described herein can be isolated or derived from insects, vertebrates, crustaceans, or tunicates, as described in detail in PCT Publication No. WO2019 / 173636 and International Publication No. WO2020 / 051374. In preferred embodiments, the SPB transposase domain is isolated or derived from the insect Trichoplusia ni (GenBank accession no. AAA87375), the silkworm Bombyx mori (GenBank accession no. BAD11135), or Macdunnoughia crassisigna (GenBank accession no. ABZ85926.1).

[0031] In some embodiments, the transposase domain is a combined-deficient type. A combined-deficient transposase domain is a transposase that can excise the corresponding transposon but incorporates the excisen transposon at a lower frequency than the corresponding wild-type transposase. Examples of combined-deficient transposases are disclosed in U.S. Patents 6,218,185, 6,962,810, 8,399,643 and International Publication No. 2019 / 173636, each of which is incorporated herein by reference in its entirety as an example of a transposase domain that may be used in the fusion proteins described herein. A list of combined-deficient amino acid substitutions is disclosed in U.S. Patent 10,041,077, which is incorporated herein by reference in its entirety as an example of an amino acid substitution that may be introduced into the transposase domains described herein.

[0032] Wild-type SPB can be made into a combined knockout type by introducing mutations, such as K93A, R372A, K375A, R376A and / or D450N (numbering starting from residue 5 for SEQ ID NO: 2). The introduction of mutations R372A, K375A, R376A and D450N is thought to make the transposase a combined knockout type but retain its cleavage function. An exemplary sequence of the combined knockout transposase domain is PBx containing the NLS shown in SEQ ID NO: 3. In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 3, which has one, two, three, four, or five conserved amino acid substitutions.

[0033] The sequence of the integrated-deficient PBx transfer domain, which does not contain NLS, is shown in Sequence ID 4: (Sequence ID 4).

[0034] In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 4. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 4 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 4.

[0035] Transposase domain including N-terminal deletion In some embodiments, transposase domains (e.g., SPB transposase domains or PBx transposase domains) containing a partial deletion of the amino terminus (also known as the "N terminus," "N-terminal domain," or "NTD") of the transposase domain are provided herein. SPB transposase domains or PBx transposase domains containing an N-terminal domain deletion were previously described in International Patent Application Publication PCT / US2022 / 77549, the whole of which is incorporated herein by reference for examples of transposase domains that may be used in the fusion proteins described herein.

[0036] Exemplary sequences of the SPB transposase domain and the PBx transposase domain, respectively, with the N-terminal amino acids 1-93 deleted, are shown in SEQ ID NOs: 5 and 6.

[0037] NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFL IRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDN WFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (Sequence ID 5).

[0038] In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 5. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 5 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 5.

[0039] NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFL IRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDN WFTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (Sequence ID 6).

[0040] In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 6. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 6 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 6.

[0041] Other exemplary sequences of the PBx transposase domain, including the N-terminal deletion, are shown in SEQ ID NOs: 7-27 in Table 1.

[0042] In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in any one of SEQ ID NOs: 7-27. In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence shown in any one of SEQ ID NOs: 7-27 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence shown in any one of SEQ ID NOs: 7-27. [Table 1] TIFF2026515653000003.tif255170TIFF2026515653000004.tif253170TIFF2026515653000005.tif255170TIFF2026515653000006.tif47170

[0043] DNA targeting domain The transposase domains and fusion proteins provided herein may further comprise one or more DNA targeting domains. The DNA targeting domains may be bound to the C-terminus or N-terminus of the transposase domain or fusion protein. In some embodiments, the DNA targeting domain is bound to the N-terminus of a transposase domain, such as a transposase domain containing an N-terminal deletion. While we do not wish to be bound by theory, it is believed that the addition of a DNA targeting domain to a transposase domain improves site-specific transposase activity by targeting the transposase fused to the DNA targeting domain to a target site. In some embodiments, insertion of a DNA targeting domain improves site-specific transposase activity by at least twofold, at least threefold, at least fourfold, or at least fivefold compared to the same transposase domain without the DNA targeting domain.

[0044] Any DNA targeting domain known in the art, including but not limited to CRISPR, zinc finger motifs, TALE, and transcription factors, may be used in connection with the transposase domains, fusion proteins, and tandem dimeric transposases described herein. In some embodiments, the DNA targeting domain comprises one, two, or three zinc finger motifs. In some embodiments, the DNA targeting domain comprises three zinc finger motifs. In some embodiments, the three zinc finger motifs are flanked by a GGGGS (SEQ ID NO: 181) linker. In some embodiments, the three zinc finger motifs flanked by the GGGGS (SEQ ID NO: 181) linker contain cumulative sequences such as the sequence shown in SEQ ID NO: 28: GGGGSERPYACPVESCDRRFSRSDELTRHIRIHTGQKPFQCRICMRNFSRSDHLTTHIRTHTGEKPFACDICGRKFARSDERKRHTKIHLRQKDGGGGS (SEQ ID NO: 28) or sequences having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto.

[0045] In certain embodiments, fusion proteins comprising an N-terminal deletion, an NLS, and a transposase domain containing three zinc finger motifs are provided herein. In some embodiments, the NLS comprises or consists of the sequence shown in SEQ ID NO: 29.

[0046] In some embodiments, the DNA targeting domain is a TAL array. Xanthomonas-derived TALEs (transcription activator-like effectors) typically contain a 288-amino acid N-terminus, followed by a variable number of approximately 34-amino acid repeat sequences, and then a 278-amino acid C-terminus (SEQ ID NO: 30). However, truncated versions have been documented in the literature (see, e.g., Miller et al., Nat Biotechnol 29, 143-148 (2011)). TALs fused to FokI nucleases (called TALENs) most often contain N-terminal and C-terminal truncations. For example, the first 152 amino acids of the N-terminus are often removed (referred to as delta-152; SEQ ID NO: 31), and the C-terminus is often truncated, leaving 63 amino acids (referred to as +63; SEQ ID NO: 32).

[0047] TAL contains an array of 34 amino acids repeated a variable number of times. Two amino acids at positions 12 and 13 are altered to determine which nucleotide the TAL repeat recognizes. This feature allows for programming the TAL array to bind to a specific DNA sequence. Amino acids NG recognize T, NI recognize A, NN recognize G or A, HD recognize C, NK recognize G, and NS recognize A, C, G, or T. Other amino acids within the 34-residue repeat can also be altered. For example, position 11 is often changed to N for a repeat that recognizes G. Also, positions 4 and 32 are often altered to reduce the repeatability of the array, but not to determine binding specificity. The number of 34-amino acid repeats in the array determines the length of the recognized DNA sequence (one protein repeat binds to one DNA bp). Furthermore, the last bp is recognized by a "half array" which consists of 20 amino acids instead of 34.

[0048] Furthermore, the N-terminal domain of TAL (e.g., SEQ ID NO: 31) recognizes and requires a T located very close to the 5' position of the target DNA sequence. Mutations of the TAL N-terminal domain that no longer require a 5'T have been documented in the literature (Lamb et al., Nucleic Acids Res. 2013 Nov;41(21):9779-85). For example, the NT-G variant requires a 5'G instead of a 5'T (SEQ ID NO: 33), while the NT-βN variant does not require any specific 5' nucleotide (SEQ ID NO: 34). These mutant N-terminal domain sequences can be used to provide further sequence options that can be targeted using TAL arrays.

[0049] Generally, each TAL array contains nine 34-amino acid repeats followed by a 20-amino acid "half" repeat. TAL arrays can be synthesized using adjacent BsmBI-type IIS restriction sites. In one embodiment, individual TAL modules containing 34-amino acid or 20-amino acid "half" repeats can be designed and synthesized by flanking them with BsmBI-type IIS restriction sites. The entire TAL module set includes four modules capable of recognizing either A, C, G, or T at each 10bp position (40 modules / 10bp target), and one TAL half-repeat module. Exemplary TAL modules are shown in Sequence IDs 35-38, where X is any amino acid: TAL module version 1: LTPDQVVAIAXXXGGKQALETVQRLLPVLCQDHG (Sequence ID 35) TAL module version 2: LTPEQVVAIAXXXGGKQALETVQRLLPVLCQAHG (Sequence ID 36) ·TAL module version 3''LTPDQVVAIAXXXGGKQALETVQRLLPVLCQAHG (Sequence ID 37) • TAL module version 4: LTPAQVVAIAXXXGGKQALETVQRLLPVLCQDHG (sequence number 38).

[0050] An exemplary TAL half-module is shown in SEQ ID NO: 39, where X is any amino acid: LTPEQVVAIAXXXGGRPALE(SEQ ID NO: 39).

[0051] Pairs of TAL arrays targeting sequences within desired genes can be designed, corresponding modules selected, and each TAL array can be assembled in-frame by pooling them together using "Golden Gate Assembly". The DNA sequences encoding the TAL arrays generated herein can be further codon-optimized using the GeneArt algorithm (Thermo Fisher).

[0052] When designing left and right TAL arrays, each containing an N-terminal domain that recognizes T and a TAL C-terminal domain that fuses to an N-terminal deletion transposase sequence (i.e., TAL-ssSPB or TAL-PBx; described below), one TAL array recognizes sequence 5' of TTAA, and the other TAL array recognizes sequence 3' of TTAA. Since sequence 5' of TTAA is almost always different from sequence 3' of TTAA in the genomic DNA target, TAL-ssSPB is most often used as a heterodimer consisting of two different TAL domains that recognize two different DNA sequences. Furthermore, the sequence recognized by the TAL array is not directly adjacent to TTAA. Instead, it is separated from TTAA by a spacer of a given bp length, e.g., a 12bp, 13bp, or 14bp spacer.

[0053] TAL arrays can target any desired DNA sequence (e.g., genomic DNA sequence). It will be apparent to those skilled in the art that any left-side TAL array of a given target can be combined with any right-side TAL array of the same target.

[0054] In some embodiments, the TAL array targets green fluorescent protein (GFP). A TAL-piggyBac transposase fusion protein containing an N-terminal deletion piggyBac transposase sequence and a combined deletion N-terminal piggyBac transposase targeting GFP is described in the jointly owned international patent application publication PCT / 2022 / 22549.

[0055] In some embodiments, the TAL array targets the LPA gene repeat element. Exemplary sequences of left-hand TAL arrays targeting the LPA repeat element are shown in SEQ ID NOs: 116, 118, 121, 124, 125, 127, 129, 131, 133, 135, 137, 139, and 141. Exemplary sequences of right-hand TAL arrays targeting LPA are shown in SEQ ID NOs: 117, 119, 120, 122, 123, 126, 128, 130, 132, 134, 136, 138, 140, and 142. In some embodiments, the left-hand TAL array targeting the LPA repeat element binds to nucleic acid molecules containing the sequences shown in SEQ ID NOs: 89, 91, 94, 97, 98, 100, 102, 104, 106, 108, 110, 112, and 114. In some embodiments, right-side TAL arrays targeting LPA repeat elements bind to nucleic acid molecules containing the sequences shown in SEQ ID NOs: 90, 92, 93, 95, 96, 99, 101, 103, 105, 107, 109, 111, 113, and 115. It will be apparent to those skilled in the art that any left-side TAL array disclosed herein can be combined with any right-side TAL array disclosed herein. Exemplary genomic target sites for LPA repeat elements are shown in SEQ ID NOs: 81-88.

[0056] This disclosure provides fusion proteins comprising a DNA targeting domain bound to a transposase domain in different ways. In some embodiments, the DNA targeting domain may be fused to or ligated to the N-terminus of a transposase domain, including an N-terminal deletion. For example, the DNA targeting domain may be inserted into the transposase domain at an appropriate location in the N-terminal region of the transposase domain. In some embodiments, the DNA targeting domain may substitute one or more amino acids in the N-terminal region of the transposase domain. In some embodiments, the DNA targeting domain is inserted into the transposase domain at an appropriate location in the N-terminal region of the transposase domain without amino acid substitution.

[0057] The DNA targeting domain can be inserted at the N-terminus of the transposase domain. For example, the DNA targeting domain is inserted at the N-terminus of the transposase domain after the 82nd amino acid and before the 105th amino acid in SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 82nd and 83rd amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 83rd and 84th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 84th and 85th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 85th and 86th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between amino acids 86 and 87 of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between amino acids 87 and 88 of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between amino acids 88 and 89 of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between amino acids 89 and 90 of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4.In some embodiments, the DNA targeting domain is inserted between the 90th and 91st amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 91st and 92nd amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 92nd and 93rd amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 93rd and 94th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 94th and 95th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 95th and 96th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 96th and 97th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 97th and 98th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 98th and 99th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4.In some embodiments, the DNA targeting domain is inserted between the 99th and 100th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 100th and 101st amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 101st and 102nd amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 102nd and 103rd amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 103rd and 104th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain is inserted between the 104th and 105th amino acids of SEQ ID NO: 2 or 3 (which have numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain includes the sequence of SEQ ID NO: 28, or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto. The transposase domain may further include an NLS, for example, the NLS of SEQ ID NO: 29.

[0058] The DNA targeting domain may substitute one or more amino acids in the N-terminal region of the transposase domain. For example, the DNA targeting domain may substitute one or more amino acids in the transposase domain between (and including) the 83rd and 105th amino acids of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 83rd amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 84th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 85th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 86th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 87th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 88th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 89th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 90th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4.In some embodiments, the DNA targeting domain substitutes the 91st amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 92nd amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 93rd amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 94th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 95th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 96th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 97th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 98th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 99th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 100th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4.In some embodiments, the DNA targeting domain substitutes the 101st amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 102nd amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 103rd amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 104th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain substitutes the 105th amino acid of SEQ ID NO: 2 or 3 (each having numbering starting from the 5th or 12th amino acid) or SEQ ID NO: 4. In some embodiments, the DNA targeting domain includes the sequence of SEQ ID NO: 28, or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto. The transposase domain may further include an NLS, for example, the NLS of SEQ ID NO: 29.

[0059] An exemplary sequence of a fusion protein containing a transposase domain with a 93-amino acid N-terminal deletion, an NLS, and three zinc finger motifs flanked by a GGGGS (SEQ ID NO: 181) linker is shown in SEQ ID NO: 40, with the NLS in italics, the sequence containing the three zinc finger motifs and the GGGGS linker underlined, and the transposase domain with the 93-amino acid N-terminal deletion in bold:

[0060] MAPKKKRKV GGGGSERPYACPVESCDRRFSRSDELTRHIRIHTGQKPFQCRICMRNFSRSDHLTTHIRTHTGEKPFACDICGRKFARSDERKRHTKIHLRQKDGGGGSNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFL IRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDN WFTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (Sequence ID 40)

[0061] In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 40. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 40 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 40. An exemplary sequence of a fusion protein containing a combined deletion transposase domain with a 93-amino acid N-terminal deletion, the NLS, and three zinc finger motifs flanked by the GGGGS (SEQ ID NO: 181) linker is shown in SEQ ID NO: 180, with the NLS in italics, the sequence containing the three zinc finger motifs and the GGGGS linker underlined, and the transposase domain with the 93-amino acid N-terminal deletion in bold:

[0062] MAPKKKRKV GGGGSERPYACPVESCDRRFSRSDELTRHIRIHTGQKPFQCRICMRNFSRSDHLTTHIRTHTGEKPFACDICGRKFARSDERKRHTKIHLRQKDGGGGSNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFL IRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNW FTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (Sequence ID 180).

[0063] In some embodiments, the fusion proteins described herein include a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 180. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 180 having one, two, three, four, or five conserved amino acid substitutions. In some embodiments, the fusion proteins described herein include a transposase domain comprising the amino acid sequence shown in SEQ ID NO: 180.

[0064] Nuclear localization signals In some embodiments, the transposase domains and fusion proteins provided herein may include an in-frame nuclear localization sequence (NLS). Examples of transposases fused to nuclear localization signals are disclosed in U.S. Patents 6,218,185, 6,962,810, 8,399,643 and International Publication No. 2019 / 173636, each of which is incorporated herein in its entirety as an example of a transposase domain that may be used in the fusion proteins described herein. In some embodiments, the NLS includes the sequence PKKKRKV (SEQ ID NO: 29). In certain embodiments, the in-frame NLS is located upstream (N-terminus) of a transposase domain that includes an N-terminal deletion.

[0065] Generally, the NLS is preferably located at the N-terminus of the fusion protein. In some embodiments, the NLS is fused to or ligated to the N-terminus of the transposase domain. In some embodiments, the NLS is fused to or ligated to the N-terminus of the DNA targeting domain.

[0066] In certain embodiments, the in-frame NLS is directly fused to the amino terminus of the transposase domain containing the N-terminal deletion. In some embodiments, the NLS is bound to the N-terminus of the transposase domain containing the N-terminal deletion via a linker (e.g., a GGGGS linker or a GGS linker).

[0067] In some embodiments, the initiating methionine is introduced before the NLS. In some embodiments, additional alanine residues are introduced before and / or after the NLS to ensure in-frame translation. Thus, the residue numbering in SEQ ID NOs: 1 and 3 begins at the 12th residue of SEQ ID NOs: 1 and 3 for the purpose of identifying deleted and mutated residues. In SEQ ID NOs: 2, which is a sequence of SPB without an NLS, the residue numbering begins at the 5th residue for the purpose of identifying deleted and mutated residues. In SEQ ID NOs: 4, the numbering begins at the first residue for the purpose of identifying deleted and mutated residues.

[0068] In some embodiments, the fusion protein includes an NLS and a transposase domain containing a 93-amino acid N-terminal deletion. In some embodiments, the fusion protein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the fusion protein includes the amino acid sequence shown in SEQ ID NO: 5.

[0069] Essential heterodimers and tandem dimers In another embodiment, a tandem dimer transposase comprising two fusion proteins is provided herein, wherein each fusion protein comprises a transposase domain, and one or both fusion proteins further comprise a DNA targeting domain. In some embodiments, both fusion proteins comprise the DNA targeting domain. In some embodiments, both fusion proteins comprise the DNA targeting domain, which targets a DNA sequence adjacent to the DNA sequence that is the insertion site targeted by the transposase. In some embodiments, only one of the two fusion proteins in the tandem dimer transposase comprises the DNA targeting domain. The DNA targeting domain may be bound to the C-terminus or N-terminus of the fusion protein.

[0070] Accordingly, in some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a first transposase domain and a first DNA targeting domain, and (b) a second fusion protein comprising a first transposase domain and a second DNA targeting domain, wherein the first DNA targeting domain and the second DNA targeting domain are different, and the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposite charges that enable the two fusion proteins to form a complex.

[0071] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a first transposase domain comprising a first NLS, a first DNA targeting domain, and an N-terminal deletion in the order of N-terminus to C-terminus, and (b) a second fusion protein comprising a second NLS, a second DNA targeting domain, and a second transposase domain comprising an N-terminal deletion in the order of N-terminus to C-terminus, wherein the transposase domains of the first fusion protein and the transposase domain of the second fusion protein have opposite charges that enable the two fusion proteins to form a complex. In some embodiments, the first and / or second transposase domains are SPB domains. In some embodiments, the first and / or second transposase domains are PBx transposase domains. In some embodiments, the first and / or second transposase domains include an N-terminal deletion at amino acids 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103. In some embodiments, the first and second transposase domains include the sequence of SEQ ID NO: 5 or 6. In some embodiments, the first and / or second DNA targeting domains include one, two, or three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domains include the sequence of SEQ ID NO: 28. In some embodiments, the first and / or second DNA targeting domains include a TAL motif.

[0072] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a first transposase domain comprising a first NLS and the sequence of SEQ ID NO: 2, 3, or 4 in N-terminal to C-terminal order, and (b) a second fusion protein comprising a second NLS and the sequence of SEQ ID NO: 2, 3, or 4 in N-terminal to C-terminal order, wherein the first and second transposase domains comprise DNA targeting domains, and the transposase domains of the first fusion protein and the transposase domain of the second fusion protein have opposite charges that enable the two fusion proteins to form a complex. In some embodiments, the first and / or second DNA targeting domains comprise one, two, or three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domains comprise the sequence of SEQ ID NO: 28. In some embodiments, the first and / or second DNA targeting domains comprise a TAL motif. In some embodiments, the first DNA targeting domain substitutes one or more amino acids between (and including) the 83rd and 105th amino acids of the first transposase domain, each having a numbering that begins with residue 5 or 12 of SEQ ID NO: 2 or 3. In some embodiments, the first DNA targeting domain substitutes residues 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103 of the first transposase domain with a numbering that begins with residue 5 or 12 of SEQ ID NO: 2 or 3. In some embodiments, the first DNA targeting domain substitutes one or more amino acids between (and including) the 83rd and 105th amino acids of the second transposase domain, each having a numbering that begins with residue 5 or 12 of SEQ ID NO: 2 or 3, respectively.In some embodiments, the second DNA targeting domain substitutes residues 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103 of the second transposase domain, with the numbering beginning at residues 5 or 12 of sequence number 2 or 3, respectively.

[0073] In another embodiment, transposase domain-containing fusion proteins are provided herein that can form an essential heterodimer with another fusion protein containing a transposase domain. While not wishing to be bound by theory, it is conceivable that two such fusion proteins assemble into a dimeric structure held together by a combination of charge interactions, hydrogen bonds, π-cation pairs, and hydrophobic interactions. Thus, each essential heterodimer complex contains two transposase domains. In some embodiments, the two fusion proteins provided herein form a complex, the aforementioned complex comprising (a) a first fusion protein containing a transposase domain and (b) a second fusion protein containing a transposase domain, wherein the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposite charges, enabling the two fusion proteins to form a complex. In non-limiting examples, the assembled complex may be a single dimer (two protein molecules) or a dimer of dimers (four protein molecules, or a tetramer).

[0074] By introducing charged residues to amino acids that contribute to dimerization with the second fusion protein, it is possible to design a pair of fusion proteins that can only associate with a predetermined tandem dimer configuration. By introducing mutations that enable only one configuration of the tandem dimer, it becomes possible to introduce a DNA targeting domain into the fusion protein, thus increasing the specificity of the transposase domain. This is shown in Figures 1A and 1B for SPB and in Figures 1C and 1D for PBx: When a DNA targeting domain is introduced into a fusion protein that can dimerize in any configuration, including homodimerization, the tandem dimer transposase will have four DNA targeting domains. However, only two of the DNA targeting domains interact with DNA, while the other two potentially sterically interfere with the transposase-DNA interaction. Any suitable DNA targeting domain described herein or known in the art may be used in the fusion proteins described herein.

[0075] Mutations in the transposase domain that confer positive or negative charge can be determined by those skilled in the art. For fusion proteins containing first and second transposase domains, the crystal structure published by Chen et al. (Nat Commun 11, 3446 (2020)) can be used to identify residue pairs within the transposase domain adjacent to the tandem dimer formed by the two such fusion proteins. Creating positively charged and negatively charged transposase domains by altering the charge of such residue pairs can be achieved using standard techniques such as site-directed mutagenesis.

[0076] For example, one or more of M185, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593, and / or F594 can be mutated with an SPB transposase domain (e.g., an SPB described in SEQ ID NO: 1 or 2, having numbering starting at the 12th residue of SEQ ID NO: 1 and the 5th residue of SEQ ID NO: 2) to generate an SPB- or SPB+ transposase domain. Similarly, one or more of M185, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593 and / or F594 can be mutated in the PBx transposase domain (for example, the PBx transposase domain of SEQ ID NO: 3, where the numbering begins at the 12th residue of SEQ ID NO: 3, or the PBx transposase domain of SEQ ID NO: 4) to generate PBx- (minus) or PBx+ (plus) transposase domains.

[0077] In some embodiments, the fusion proteins described herein may comprise (i) one SPB+ transposase domain, or (ii) one SPB- transposase domain.

[0078] To achieve the formation of essential heterodimers, a pair of mutations can be introduced into the fusion protein or transposase domain to generate positively and negatively charged fusion protein or transposase domains that can then interact to form a heterodimer. In some embodiments, the mutated residue pairs are shown in Table 2. For example, one or more mutations listed in the column labeled "Protein 1" may be introduced into the first SPB or PBx domain, and one or more corresponding mutations listed in the column labeled "Protein 2" may be introduced into the second SPB or PBx domain. In some embodiments, members of the residue pair are mutated to have opposite charges. [Table 2]

[0079] To introduce a positive charge, amino acids with uncharged side chains, such as methionine, or amino acids with negatively charged side chains, such as aspartic acid, can be replaced with positively charged amino acids, such as lysine or arginine. To introduce a negative charge, amino acids with positively charged side chains, such as arginine or lysine, or amino acids with hydrophobic side chains, such as leucine, can be replaced with negatively charged amino acids, such as aspartic acid or glutamic acid.

[0080] In certain embodiments, to generate an SPB+ fusion protein, one or more of the following mutations are introduced into the SPB transposase domain of the fusion protein provided herein (for example, the SPB described in SEQ ID NO: 1 or 2, having numbering beginning at the 12th residue of SEQ ID NO: 1 and the 5th residue of SEQ ID NO: 2): M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, the SPB+ transposase domain includes the M185R mutation and the D198K mutation. In some embodiments, the SPB+ transposase domain includes the M185R mutation and the D201R mutation. In some embodiments, the SPB+ transposase domain includes the D197K mutation and the D201R mutation. In some embodiments, the SPB+ transposase domain includes the D198K mutation and the D201R mutation. In some embodiments, the SPB+ transposase domain includes the M185R mutation, the D198K mutation, and the D201R mutation.

[0081] In certain embodiments, to generate a PBx+ fusion protein, one or more of the following mutations are introduced into the PBx transposase domain of the fusion protein provided herein (for example, the PBx transposase domain of SEQ ID NO: 3, whose numbering begins at the 12th residue of SEQ ID NO: 3, or the PBx transposase domain of SEQ ID NO: 4): M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, the PBx+ transposase domain includes the M185R mutation and the D198K mutation. In some embodiments, the PBx+ transposase domain includes the M185R mutation and the D201R mutation. In some embodiments, the PBx+ transposase domain includes the D197K mutation and the D201R mutation. In some embodiments, the SPB+ transposase domain includes the D198K mutation and the D201R mutation. In some embodiments, the PBx+ transposase domain includes the M185R mutation, the D198K mutation, and the D201R mutation.

[0082] In certain embodiments, one or more of the following mutations are introduced into the SPB transposase domain of the fusion protein provided herein (for example, the SPB described in SEQ ID NO: 1 or 2, having numbering beginning at the 12th residue of SEQ ID NO: 1 and the 5th residue of SEQ ID NO: 2) to generate SPB fusion proteins: L204D, L204E, K500D, K500E, R504E, and R504D. In some embodiments, the SPB-transposase domain includes the L204E mutation and the K500D mutation. In some embodiments, the SPB-transposase domain includes the L204E mutation and the R504D mutation. In some embodiments, the SPB-transposase domain includes the K500 mutation and the R504D mutation. In some embodiments, the SPB-transposase domain includes the L204E mutation, the K500D mutation, and the R504D mutation.

[0083] In certain embodiments, to generate a PBx fusion protein, one or more of the following mutations are introduced into the PBx transposase of the fusion protein provided herein (e.g., the PBx transposase domain of SEQ ID NO: 3 or SEQ ID NO: 4, whose numbering begins at the 12th residue of SEQ ID NO: 3): L204D, L204E, K500D, K500E, R504E, and R504D. In some embodiments, the PBx-transposase domain includes the L204E mutation and the K500D mutation. In some embodiments, the PBx-transposase domain includes the L204E mutation and the R504D mutation. In some embodiments, the PBx-transposase domain includes the K500 mutation and the R504D mutation. In some embodiments, the PBx-transposase domain includes the L204E mutation, the K500D mutation, and the R504D mutation.

[0084] Exemplary sequences of SPB+ transposase domains are shown in SEQ ID NOs: 42-54. Exemplary sequences of SPB- transposase domains are shown in SEQ ID NOs: 55-64. In some embodiments, the transposase domains provided herein include the amino acid sequence shown in any one of SEQ ID NOs: 42-64. In some embodiments, the transposase domains provided herein include the amino acid sequence shown in any one of SEQ ID NOs: 42-64, further comprising one or more conserved amino acid sequences.

[0085] In some embodiments, the fusion protein described herein comprises a transposase domain comprising the amino acid sequence shown in any one of SEQ ID NOs: 42-54. In some embodiments, the transposase domain comprises the amino acid sequence shown in any one of SEQ ID NOs: 42-54, further comprising one or more conserved amino acid sequences.

[0086] In some embodiments, the fusion protein described herein comprises a transposase domain comprising the amino acid sequence shown in any one of SEQ ID NOs. 55-64. In some embodiments, the transposase domain comprises the amino acid sequence shown in any one of SEQ ID NOs. 55-64, further comprising one or more conserved amino acid sequences.

[0087] In some embodiments, a complex is provided herein that comprises (a) a first fusion protein comprising a transposase domain comprising an amino acid sequence shown in any one of SEQ ID NOs: 42-54, and (b) a second fusion protein comprising a transposase domain comprising an amino acid sequence shown in any one of SEQ ID NOs: 55-64.

[0088] SPB+, SPB-, PBx+, and PBx- fusion proteins and transposase domains may further comprise N-terminal deletions of the transposase domains described herein. Accordingly, in some embodiments, SPB+ fusion proteins comprising transposase domains comprising N-terminal deletions of approximately 20 amino acids, approximately 40 amino acids, approximately 60 amino acids, approximately 80 amino acids, approximately 100 amino acids, or approximately 115 amino acids are provided herein. In some embodiments, the transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 98 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 99 amino acids.In some embodiments, the transposase domain includes a 100-amino acid N-terminal deletion. In some embodiments, the transposase domain includes a 101-amino acid N-terminal deletion. In some embodiments, the transposase domain includes a 102-amino acid N-terminal deletion. In some embodiments, the transposase domain includes a 103-amino acid N-terminal deletion.

[0089] In some embodiments, SPB fusion proteins are provided herein that include a transposase domain containing an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 81 amino acids, about 82 amino acids, about 83 amino acids, about 84 amino acids, about 85 amino acids, about 86 amino acids, about 87 amino acids, about 88 amino acids, about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 83 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 85 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 88 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain includes a 98-amino acid N-terminal deletion.In some embodiments, the transposase domain includes an N-terminal deletion of 99 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 100 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 101 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 102 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 103 amino acids.

[0090] In some embodiments, PBx+ fusion proteins are provided herein that include a transposase domain containing an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 100 amino acids, or about 115 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 83 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 85 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 88 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 98 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 99 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 100 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 101 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 102 amino acids.In some embodiments, the transposase domain includes a 103-amino acid N-terminal deletion.

[0091] In some embodiments, PBx fusion proteins are provided herein that include a transposase domain containing an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 81 amino acids, about 82 amino acids, about 83 amino acids, about 84 amino acids, about 85 amino acids, about 86 amino acids, about 87 amino acids, about 88 amino acids, about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 83 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain contains an N-terminal deletion of 85 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 88 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 93 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain includes a 98-amino acid N-terminal deletion.In some embodiments, the transposase domain includes an N-terminal deletion of 99 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 100 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 101 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 102 amino acids. In some embodiments, the transposase domain includes an N-terminal deletion of 103 amino acids.

[0092] Built-in cassette Also provided herein are embedding cassettes for site-specific transfer of DNA molecules into the genome of a cell. In some embodiments, the embedding cassette includes an embedding site of sequence TTAA. In some embodiments, the embedding cassette for site-specific transfer of nucleic acid into the genome of a cell includes or comprises a nucleic acid including a central transposon ITR embedding site CTTAAA sequence flanked by an upstream TAL array target sequence and a downstream TAL array target sequence, each of which is 12 or 13 base pairs away from the CTTAAA sequence. In some embodiments, each of at least one upstream TAL array target site sequence and each of which is downstream TAL array target site sequences are the same. In some embodiments, each of at least one upstream TAL array target site sequence and each of which is downstream TAL array target site sequences are different. In some embodiments, each of at least one upstream TAL array target site and each of which is downstream TAL array target site sequences targets a 10 bp sequence of the LPA repeat element.

[0093] A method for site-directed transfer of a DNA molecule into the genome of a cell containing a stably integrated integration cassette is also provided, comprising introducing into the cell: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed in the cell; and b) a DNA molecule containing a transposon, wherein the expressed fusion protein, by site-directed transfer, integrates the transposon into the CTTAAA sequence of the stably integrated integration cassette.

[0094] A method for generating engineered cells by site-directed transposition is provided herein, comprising introducing into cells containing a stably incorporated integration cassette: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed in the cell; and b) a DNA molecule containing a transposon, wherein the expressed fusion protein, by site-directed transposition, incorporates the transposon into the CTTAAA sequence of the stably incorporated integration cassette, thereby generating engineered cells.

[0095] nucleic acid Polynucleotides comprising nucleic acid sequences encoding the fusion proteins described herein are also provided herein. In some embodiments, the polynucleotides are isolated.

[0096] The isolated polynucleotides of this disclosure can be prepared using (a) recombinant methods, (b) synthetic techniques, (c) purification techniques, and / or (d) a combination thereof, as is well known in the art.

[0097] Methods for constructing nucleic acids encoding transposase domains containing the N-terminal deletion described herein are either well known in the art or described herein, and are, for example, PCR-based mutagenesis methods.

[0098] The fusion products of the present invention can be produced using any suitable method known in the art or described herein.

[0099] The isolated polynucleotides of this disclosure, e.g., RNA, cDNA, genomic DNA, or any combination thereof, can be obtained from a biological source using any number of cloning methodologies known to those skilled in the art. In some embodiments, oligonucleotide probes that selectively hybridize to the polynucleotides of this disclosure under stringent conditions are used to identify desired sequences in a cDNA or genomic DNA library.

[0100] Methods for amplifying RNA or DNA are well known in the art and can be used in accordance with this disclosure without excessive experimentation, based on the teachings and guidelines presented herein. Known methods for amplifying DNA or RNA include polymerase chain reaction (PCR) and related amplification processes (e.g., U.S. Patents 4,683,195, 4,683,202, 4,800,159, 4,965,188 (Mullis et al.), 4,795,699 and 4,921,794 (Tabor et al.), 5,142,033 (Innis), 5,122,464 (Wilson et al.), 5,091,310 (Innis), and 5,066,584 (G)). Examples include, but are not limited to, the works of Ausubel et al. (Yylensten et al.), Gelfand et al. (Gelfand et al.), Silver et al. (Silver et al.), Biswas (Biswas), and Ringold (Ringold), as well as RNA-mediated amplification using antisense RNA against a target sequence as a template for double-stranded DNA synthesis (US Patent No. 5,130,238 (Malek et al.), trade name NASBA), and the entire contents of those references are incorporated herein by reference. (See, for example, Ausubel or Sambrook).

[0101] For example, polymerase chain reaction (PCR) technology can be used to directly amplify the sequences of the polynucleotides and related genes of this disclosure from a genomic DNA or cDNA library. PCR and other in vitro amplification methods are also useful for purposes such as cloning nucleic acid sequences encoding proteins to be expressed, creating nucleic acids to be used as probes to detect the presence of desired mRNA in a sample, nucleic acid sequencing, and other purposes. Examples of techniques sufficient to guide those skilled in the art through in vitro amplification methods can be found in Berger, Sambrook, Ausubel, Mullis, et al., U.S. Patent No. 4,683,202 (1987), and Innis, et al., PCR Protocols: A Guide to Methods and Applications, Eds., Academic Press Inc., San Diego, Calif. (1990). Commercial kits for genomic PCR amplification are known in the art. For example, see the Advantage-GC Genomic PCR Kit (Clontech). Furthermore, to improve the yield of long PCR products, for example, the T4 gene 32 protein (Boehringer Mannheim) can be used.

[0102] The isolated polynucleotides of this disclosure can also be prepared by direct chemical synthesis using known methods (see, for example, Ausubel et al. above). Chemical synthesis generally yields single-stranded oligonucleotides, which can be converted to double-stranded DNA by hybridization with a complementary sequence or by polymerization using DNA polymerase with the single strand as a template. Those skilled in the art will understand that while the chemical synthesis of DNA is limited to sequences of approximately 100 bases or more, longer sequences can be obtained by ligating shorter sequences.

[0103] Expression vectors and host cells This disclosure also relates to vectors containing the polynucleotides of this disclosure, host cells genetically engineered with the recombinant vectors, and the generation of at least one protein backbone by recombinant techniques, as is well known in the art. See, for example, Sambrook et al. and Ausubel et al., respectively, which are fully incorporated herein by reference.

[0104] Polynucleotides can be optionally conjugated to vectors containing selectable markers for replication in a host. Generally, plasmid vectors are introduced in precipitates such as calcium phosphate precipitates or in complexes with charged lipids. If the vector is a virus, it can be packaged in vitro using a suitable packaging cell line and transduced into host cells.

[0105] The DNA insert may be operablely ligated to a suitable promoter. In some embodiments, the promoter is the EF-1α promoter. The expression construct further includes a transcription start site, a termination site, and a ribosome binding site for translation in the transcribed region. The coding portion of the mature transcript expressed by the construct preferably includes a translation start codon (e.g., ATG) at the beginning of the mRNA to be translated and a appropriately placed stop codon (e.g., UAA, UGA, or UAG) at the end, with UAA and UAG preferred for expression in mammalian or eukaryotic cells.

[0106] The expression vector may contain at least one selection marker. Such markers include, for example, ampicillin, zeosin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / Geneticin (neo gene), DHFR (encoding Dihydrofolate Reductase and conferring resistance to Methotrexate), mycophenolic acid, or glutamine synthase (GS, U.S. Patent Nos. 5,122,464, 5,770,359, and 5,827,739), blasticidine (bsd gene), resistance genes for eukaryotic cell culture, and ampicillin, zeosin (Sh Examples of such genes include, but are not limited to, the bla gene, puromycin (pac gene), hygromycin B (hygB gene), G418 / Geneticin (neo gene), kanamycin, spectinomycin, streptomycin, carbenicillin, bleomycin, erythromycin, polymyxin B, or tetracycline resistance genes (the above patents are fully incorporated herein by reference). Suitable culture media and conditions for the above host cells are known in the art. Suitable vectors will be readily apparent to those skilled in the art. The introduction of vector constructs into host cells can be carried out by calcium phosphate transfection, DEAE-dextran-mediated transfection, cationic lipid-mediated transfection, electroporation, transduction, infection, or other known methods. Such methods are described in the art, for example, in Sambrook, Chapters 1-4 and 16-18, and in Ausubel, Chapters 1, 9, 13, 15, and 16.

[0107] An expression vector may include at least one selectable cell surface marker for isolating cells modified by the compositions and methods of the present disclosure. The selectable cell surface markers of the present disclosure consist of surface proteins, glycoproteins, or groups of proteins that distinguish a cell or subset of cells from another defined subset of cells. Preferably, the selectable cell surface marker distinguishes cells modified by the compositions or methods of the present disclosure from cells that have not been modified by the compositions or methods of the present disclosure. Examples of such cell surface markers include, but are not limited to, “cluster designation” or “classification determinant” proteins (often abbreviated as “CD”) such as truncated or full-length forms of CD19, CD271, CD34, CD22, CD20, CD33, CD52, or combinations thereof. Cell surface markers include the suicide gene marker RQR8 (Philip B et al. Blood. 2014 Aug 21;124(8):1277-87).

[0108] The expression vector may include at least one selectable drug resistance marker for isolating cells modified by the compositions and methods of the present disclosure. The selectable drug resistance markers of the present disclosure may include wild-type or mutant Neo, DHFR, TYMS, FRANCF, RAD51C, GCS, MDR1, ALDH1, NKX2.2, or any combination thereof.

[0109] Those skilled in the art are familiar with the numerous expression systems available for expressing the nucleic acids encoding the proteins of the Disclosure. Alternatively, the nucleic acids of the Disclosure can be expressed in host cells by being (operationally) turned on in host cells containing endogenous DNA encoding the protein backbone of the Disclosure. Such methods are well known in the art, for example, as described in U.S. Patents 5,580,734, 5,641,670, 5,733,746 and 5,733,761, which are fully incorporated herein by reference.

[0110] Examples of cell cultures useful for generating protein backbone, specific parts thereof, or variants include bacterial, yeast, and mammalian cells, which are known in the art. Mammalian cell lines often exist in the form of a single layer of cells, but mammalian cell suspensions or bioreactors can also be used. Several suitable host cell lines capable of expressing intact glycosylated proteins have been developed in the art, including COS-1 (e.g., ATCC CRL 1650), COS-7 (e.g., ATCC CRL-1651), HEK293, BHK21 (e.g., ATCC CRL-10), CHO (e.g., ATCC CRL 1610), and BSC-1 (e.g., ATCC CRL-26) cell lines, Cos-7 cells, CHO cells, hepG 2 cells, P3X63Ag8.653, SP2 / 0-Ag14, 293 cells, and HeLa cells, which are readily available, for example, from the American Type Culture Collection, Manassas, Va (www.atcc.org). Preferred host cells include lymphoid cells such as myeloma cells and lymphoma cells. Particularly preferred host cells are P3X63Ag8.653 cells (ATCC accession number CRL-1580) and SP2 / 0-Ag14 cells (ATCC accession number CRL-1851). In a preferred embodiment, the recombinant cells are P3X63Ab8.653 or SP2 / 0-Ag14 cells.

[0111] Expression vectors for these cells may include, but are not limited to, one or more of the following expression regulatory sequences: origins of replication, promoters (e.g., late or early SV40 promoter, CMV promoter (US Patent No. 5,168,062, 5,385,839), HSV tk promoter, pgk (phosphoglycerin kinase) promoter, EF-1 alpha promoter (US Patent No. 5,266,491), at least one human promoter, enhancers, and / or processing information sites (e.g., ribosome binding sites, RNA splice sites, polyadenylation sites (e.g., SV40 large T Ag poly-A addition site)), and transcriptional terminator sequences. See, for example, Ausubel et al. and Sambrook et al. Other cells useful for generating the nucleic acids or proteins of this disclosure are known and / or available from, for example, the American Type Culture Collection Catalogue of Cell Lines and Hybridoma (www.atcc.org) or other known or commercial sources.

[0112] When eukaryotic host cells are used, polyadenylated or transcriptional terminator sequences are typically incorporated into the vector. An example of a terminator sequence is a polyadenylated sequence derived from the bovine growth hormone gene. In some embodiments, the polyA sequence is the SV40 polyA sequence.

[0113] Sequences for precise splicing of the transcript can also be included. An example of a splicing sequence is the VP1 intron derived from SV40 (Sprague, et al., J. Virol. 45:773-781 (1983)). Furthermore, as is known in the art, gene sequences for controlling replication in host cells can be incorporated into the vector.

[0114] Plasmid constructs described herein may be used to deliver nucleic acids encoding transposase domains or fusion proteins described herein to cells.

[0115] The transposase domains and fusion proteins described herein may be delivered to cells using mRNA constructs. Accordingly, in one embodiment, mRNA sequences encoding the transposase domains or fusion proteins described herein are provided herein. Such mRNA sequences may be delivered to cells using nanoparticles, such as lipid nanoparticles. Examples of lipid nanoparticles are described, for example, in International Patent Applications PCT / US2021 / 055876, PCT / US2022 / 017570, U.S. Provisional Applications 63 / 397,268, 63 / 301,855, and 63 / 348,614, each of which is incorporated herein by reference in whole for examples of lipid nanoparticles that may be used to deliver the fusion proteins or transposase domains encoding the mRNA constructs described herein. The mRNA constructs may also be delivered to cells by electroporation or nucleofection. The mRNA may be capped or otherwise modified.

[0116] Cells and modified cells The transposases and fusion proteins described herein may be used in combination with transposons to modify cells. The transposon may be a piggyBac®(PB) transposon. In some embodiments where the transposon is a PB transposon, the transposase is a piggyBac®(PB) transposon, a piggyBac-like(PBL) transposon, or a Super piggyBac®(SPB) transposon. Non-limiting examples of PB transposons are described in detail in U.S. Patents 6,218,182, 6,962,810, 8,399,643 and PCT Publication No. WO2010 / 099296, each of which is incorporated herein by reference in whole with respect to examples of transposons that may be used in combination with the transposases and fusion proteins described herein. Transposons may include nucleic acids encoding therapeutic proteins or therapeutic agents. Examples of therapeutic proteins are disclosed in PCT publication numbers WO2019 / 173636 and WO2020 / 051374, each of which is incorporated herein by reference in its entirety as an example of a therapeutic protein that may be encoded by a transposon used in conjunction with the transposases and fusion proteins described herein.

[0117] Accordingly, modified cells comprising one or more transposons and one or more tandem dimer transposases or fusion proteins as described herein are provided herein. The cells and modified cells of this disclosure may be mammalian cells. Preferably, the cells and modified cells are human cells.

[0118] Cells modified using the site-specific transposase fusion proteins described herein may be germ cells or somatic cells. The cells and modified cells of this disclosure include immune cells, such as lymphoid progenitor cells, natural killer (NK) cells, T lymphocytes (T cells), and stem memory T cells (T SCM T cells), central memory T cells (TCM Modified cells may include stem cell-like T cells, B lymphocytes (B cells), antigen-presenting cells (APCs), cytokine-induced killer (CIK) cells, myeloid progenitor cells, neutrophils, basophils, eosinophils, monocytes, macrophages, platelets, erythrocytes, red blood cells (RBCs), megakaryocytes, or osteoclasts. Modified cells may be differentiated, undifferentiated, or immortalized. Modified undifferentiated cells may be stem cells. Modified undifferentiated cells may be induced pluripotent stem cells. Modified cells may be T cells, hematopoietic stem cells, natural killer cells, macrophages, dendritic cells, monocytes, megakaryocytes, or osteoclasts. Modified cells may be modified while the cell is quiescent, activated, resting, intermediate, prophase, metaphase, anaphase, or telophase. Modified cells may be fresh cells, cryopreserved cells, bulk cells, cells classified into subpopulations, cells from whole blood, cells from leukocyte extrusion, or cells from immortalized cell lines. Detailed descriptions for isolating cells from leukocyte apheresis products or blood are disclosed in PCT publications WO2019 / 173636 and WO2020 / 051374, each of which is incorporated herein by reference in whole.

[0119] The method of disclosure may be modified. T A population of cells can be modified and / or produced, where at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% or any percentage between them are stem memory T cells (T SCM ) or T SCMThe cells express one or more cell surface markers, and one or more of these cell surface markers include CD45RA and CD62L. The cell surface markers may include one or more of CD62L, CD45RA, CD28, CCR7, CD127, CD45RO, CD95, CD95, and IL-2Rβ. The cell surface markers may include one or more of CD45RA, CD95, IL-2Rβ, CCR7, and CD62L.

[0120] This disclosure provides a method for expressing a CAR on the surface of a cell. The method comprises (a) obtaining a cell population, (b) contacting the cell population with a composition containing a CAR or a sequence encoding a CAR under conditions sufficient to transfer the CAR across the cell membrane of at least one cell in the cell population, thereby generating a modified cell population, (c) culturing the modified cell population under conditions suitable for incorporating a sequence encoding a CAR, and (d) growing and / or selecting at least one cell from the modified cell population that expresses a CAR on its cell surface. A more detailed description of the method for expressing a CAR on the surface of a cell is disclosed in PCT publication numbers WO2019 / 049816 and WO2020 / 051374, each of which is incorporated herein by reference in whole.

[0121] This disclosure provides cells or a population of cells comprising a composition comprising (a) an inducible transgene construct comprising a sequence encoding an inducible promoter and a sequence encoding a transgene, and (b) a receptor construct comprising a sequence encoding a constitutive promoter and a sequence encoding an exogenous receptor such as a CAR, wherein when constructs (a) and (b) are incorporated into the genomic sequence of cells, the exogenous receptor is expressed, and when the exogenous receptor binds to a ligand or antigen, it transmits an intracellular signal that directly or indirectly targets an inducible promoter that modulates the expression of the inducible transgene (a), thereby modifying gene expression.

[0122] This disclosure further provides compositions comprising modified, amplified, and selected cell populations based on the methods described herein.

[0123] The modified cells of this disclosure (e.g., CAR T cells) may be further modified to enhance their therapeutic potential. Alternatively, or in addition to this, the modified cells may be further modified to reduce their sensitivity to immunological and / or metabolic checkpoints, for example, by blocking and / or diluting specific checkpoint signals (e.g., checkpoint inhibitors) that are naturally delivered to the cells within the tumor immunosuppressive microenvironment.

[0124] The modified cells of this disclosure (e.g., CAR T cells) may be further modified to silence or reduce the expression of (i) one or more genes encoding receptors for inhibitory checkpoint signals, (ii) one or more genes encoding intracellular proteins involved in checkpoint signaling, (iii) one or more genes encoding transcription factors that interfere with therapeutic efficacy, (iv) one or more genes encoding receptors for cell death or apoptosis, (v) one or more genes encoding metabolism-sensing proteins, (vi) one or more genes encoding proteins that confer sensitivity to cancer treatment, including monoclonal antibodies, and / or (vii) one or more genes encoding growth advantage factors. Non-limited examples of genes that may be modified to silence or reduce expression or suppress their function include, but are not limited to, the exemplary inhibitory checkpoint signals, intracellular proteins, transcription factors, receptors for cell death or apoptosis, metabolism-sensing proteins, proteins that confer sensitivity to cancer treatment, and growth advantage factors disclosed in PCT Publication No. WO2019 / 173636.

[0125] The modified cells of this disclosure (e.g., CAR T cells) may be further modified to express modified / chimeric checkpoint receptors. Modified / chimeric checkpoint receptors may include null receptors, decoy receptors, or dominant-negative receptors. Examples of null, decoy, or dominant-negative intracellular receptors / proteins include, but are not limited to, downstream signaling components of inhibitory checkpoint signals, transcription factors, cytokines or cytokine receptors, chemokines or chemokine receptors, cell death or apoptosis receptors / ligands, metabolic sensing molecules, proteins that confer sensitivity to cancer treatment, and oncogenes or tumor suppressor genes. Non-exclusive examples of cytokines, cytokine receptors, chemokines, and chemokine receptors are disclosed in PCT publication number WO2019 / 173636.

[0126] Genome modification may involve introducing nucleic acid sequences, transgenes, and / or genome editing constructs into cells ex vivo, in vivo, in vitro, or in situ to stably incorporate nucleic acid sequences, transiently incorporate nucleic acid sequences, induce site-directed integration of nucleic acid sequences, or induce biased integration of nucleic acid sequences. A nucleic acid sequence may be a transgene.

[0127] Stable chromosome integration can be random, site-specific, or biased. While we do not wish to be bound by theory, the addition of DNA-binding domains to tandem dimer transposases described herein is thought to improve the site specificity of the transposases.

[0128] Site-directed integration can occur at safe harbor sites. Genomic safe harbor sites can provide a place for the integration of new genetic material in a way that ensures the newly inserted genetic element functions reliably (e.g., is expressed at therapeutically effective levels) and does not cause harmful changes in the host genome that pose a risk to the host organism. Non-limiting examples of potential genomic safe harbors include the intron sequence of the human albumin gene, adeno-associated virus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19, the site of the chemokine (CC motif) receptor 5 (CCR5) gene, and the site of the human ortholog at the mouse Rosa26 locus.

[0129] Site-directed transgene integration can occur at sites that interfere with the expression of a target gene. This interference can occur through site-directed integration at introns, exons, promoters, gene elements, enhancers, suppressors, start codons, stop codons, and response elements. Non-exclusive examples of target genes targeted by site-directed integration include TRAC, TRAB, PDI, any gene encoding an immunosuppressive protein, and genes encoding proteins involved in allorejection.

[0130] Site-directed transgene integration can occur at sites that result in enhanced expression of a target gene. Enhanced target gene expression can occur through site-directed integration at introns, exons, promoters, gene elements, enhancers, suppressors, start codons, stop codons, and response elements.

[0131] Site-specific transgene insertion sites can be unstable chromosomal insertions. Unstable integration can be transient non-chromosomal integration, semi-stable non-chromosomal integration, semi-persistent non-chromosomal insertion, or unstable chromosomal insertion. Transient non-chromosomal insertions can be epichromosomes or cytoplasm. In one embodiment, a transient non-chromosomal insertion of a transgene is not integrated into a chromosome, and the modified genetic material is not replicated during cell division.

[0132] Site-directed transgene integration sites may be modified binding sites for DNA targeting domains in transposon domains, fusion proteins, or tandem dimers as described herein. For example, the TTAA target DNA integration site for SPB may be modified to insert an adjacent DNA binding site for a DNA targeting domain containing three zinc finger motifs (e.g., the sequence of SEQ ID NO: 28, or a DNA targeting domain containing or comprising a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto). For example, a DNA targeting domain containing three zinc finger motifs is thought to bind to the DNA sequence GCGTGGGCG. Therefore, the introduction of two copies of the sequence GCGTGGGCG adjacent to the TTAA target integration site of SPB is thought to improve site-directed integration of the SPB transposase domain containing the DNA targeting domain containing three zinc finger motifs. In some embodiments, the two copies of the sequence GCGTGGGCG are oriented in reverse (5') and complementary (3') orientations.

[0133] In some embodiments, polynucleotides are provided herein that include, in some embodiments, the reverse complement of the sequence of a target site for a DNA targeting domain, a first spacer, a TTAA target integration site for the SPB, a second spacer, and the sequence of the target site for the DNA targeting domain in 5' to 3' order. In some embodiments, the first spacer and the second spacer are of the same length. In some embodiments, the first and / or second spacer is 3 bp long. In some embodiments, the first and / or second spacer is 4 bp long. In some embodiments, the first and / or second spacer is 5 bp long. In some embodiments, the first and / or second spacer is 6 bp long. In some embodiments, the first and / or second spacer is 7 bp long. In some embodiments, the first and / or second spacer is 8 bp long. In some embodiments, the first and / or second spacer is 9 bp long. In some embodiments, the first and / or second spacer is 10 bp long.

[0134] The modified target site can be introduced into cells or cell lines to facilitate targeted genome engineering. For example, a cell line engineered to include a modified target site for an SPB or PBx provided herein can be transfected with an SPB or PBx, as well as a transposon containing the donor DNA, so that the donor DNA is inserted into the modified target site. In some embodiments, the cell line is a T cell line. In some embodiments, the modified target sequence is introduced into a highly expressed genomic region. In some embodiments, the cells are in vitro cells, e.g., cells in cell culture.

[0135] In the case of a DNA-binding domain containing TAL, the target site is determined by the sequence of TAL. Those skilled in the art will be able to modify the TAL sequence to achieve desired target specificity.

[0136] Genome modification can involve the unstable chromosomal integration of a transgene. The integrated transgene may be silenced, removed, excised, or further modified.

[0137] In some embodiments, the transposase domains, fusion proteins, and tandem dimer complexes provided herein have better transposase efficacy than their wild-type counterparts. Transposase activity can be measured by any suitable assay known in the art or described herein, such as the Split GFP assay. For example, the transposase domains, fusion proteins, and tandem dimer complexes provided herein may have on-target genomic integration activity equivalent to their wild-type counterparts, but with reduced off-target genomic integration activity compared to their wild-type counterparts.

[0138] In some embodiments, the transposase domain and DNA targeting domain provided herein have an on-target activity to off-target activity ratio that is increased by at least 50 times, at least about 100 times, at least about 150 times, at least about 200 times, at least about 250 times, at least about 300 times, at least about 350 times, at least about 400 times, at least about 450 times, at least about 500 times, at least about 550 times, at least about 600 times, at least about 650 times, at least about 700 times, at least about 750 times, at least about 800 times, at least about 850 times, at least about 900 times, at least about 950 times, or at least about 1000 times compared to an unmodified SPB transposase.

[0139] In some embodiments, a transposase domain provided herein, comprising a DNA targeting domain inserted into the N-terminal region of the transposase domain, has an on-target activity to off-target activity ratio that is at least 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500-fold, at least about 550-fold, at least about 600-fold, at least about 650-fold, at least about 700-fold, at least about 750-fold, at least about 800-fold, at least about 850-fold, at least about 900-fold, at least about 950-fold, or at least about 1000-fold increased.

[0140] In certain embodiments, modified cells are used therapeutically in adoptive cell therapy.

[0141] Adoptive cell compositions that are “universally” safe for administration to any patient (not just the patient from which they originate) require a significant reduction or elimination of alloreactivity. For this purpose, the cells of the Disclosure (e.g., allogeneic cells) can be modified to interrupt the expression or function of a class of T cell receptors (TCRs) and / or major histocompatibility complexes (MHCs). TCRs mediate graft-versus-host (GvH) responses, and MHCs mediate host-versus-graft (HvG) responses. In a preferred embodiment, any expression and / or function of TCRs is eliminated to prevent T cell-mediated GvH that could lead to the death of the subject. Thus, in a preferred embodiment, the Disclosure provides a pure TCR-negative allogeneic T cell composition (e.g., each cell in the composition expresses TCRs at such low levels that they are undetectable or absent).

[0142] The expression and / or function of MHC class I (MHC-I, specifically HLA-A, HLA-B, and HLA-C) is reduced or eliminated to prevent HvG and consequently improve cell engraftment in the target. Improved engraftment results in longer cell persistence and therefore a larger therapeutic area of ​​the target. Specifically, the expression and / or function of beta-2-microglobulin (B2M), a structural component of MHC-I, is reduced or eliminated. Non-limiting examples of guide RNAs (gRNAs) for targeting and deleting MHC activators are disclosed in PCT application number PCT / US2019 / 049816.

[0143] A detailed description of genetic modifications of endogenous sequences encoding the naturally occurring chimeric stimulatory receptors, TCR-alpha (TCR-α), TCR-beta (TCR-β), and / or beta-2-microglobulin (β2M), as well as naturally occurring polypeptides including the HLA class I histocompatibility antigen, alpha-E (HLA-E) polypeptide, is disclosed in its entirety in PCT application publication no. WO2020 / 051374, which is incorporated herein by reference.

[0144] Under normal conditions, complete T cell activation depends on the involvement of the TCR in conjunction with a second signal mediated by one or more co-stimulatory receptors (e.g., CD28, CD2, 4-1BBL) that enhance the immune response. However, in the absence of a TCR, T cell proliferation is significantly reduced upon stimulation using a standard activation / stimulation reagent containing an agonist anti-CD3 mAb. Accordingly, this disclosure provides a naturally occurring chimeric stimulating receptor (CSR) comprising: (a) an external domain containing an activating component, the activating component being isolated or induced from a first protein; (b) a transmembrane domain; and (c) an endodomain containing at least one signaling domain, the at least one signaling domain being isolated or induced from a second protein, wherein the first and second proteins are not identical.

[0145] The activating component may include one or more parts of components of T cell receptors (TCRs), TCR complexes, TCR coreceptors, TCR costimulatory proteins, TCR inhibitory proteins, cytokine receptors, and chemokine receptors to which the agonist of the activating component binds. The activating component may also include the extracellular domain of CD2 or a portion thereof to which the agonist binds.

[0146] The signaling domain may include one or more components of human signaling domains, T cell receptors (TCRs), TCR complexes, TCR coreceptors, TCR costimulatory proteins, TCR inhibitory proteins, cytokine receptors, and chemokine receptors. The signaling domain may include the CD3 protein or a portion thereof. The CD3 protein may include the CD3ζ protein or a portion thereof.

[0147] The endodomain may further contain a cytoplasmic domain. The cytoplasmic domain may be isolated or derived from a third protein. The first and third proteins may be identical. The external domain may further contain a signal peptide. The signal peptide may be derived from a fourth protein. The first and fourth proteins may be identical. The transmembrane domain may be isolated or derived from a fifth protein. The first and fifth proteins may be identical.

[0148] This disclosure also provides naturally occurring chimeric stimulating receptors (CSRs) whose ectodomains include modifications. Modifications may include mutations or truncations in the amino acid sequence of the activator or the first protein compared to the wild-type sequence of the activator or the first protein. Amino acid sequence mutations or truncations of the activator may include mutations or truncations of the CD2 extracellular domain or a portion thereof to which the agonist binds. Mutations or truncations of the CD2 extracellular domain may reduce or eliminate binding to naturally occurring CD58.

[0149] This disclosure provides nucleic acid sequences encoding any CSR disclosed herein. This disclosure also provides transposons or vectors comprising nucleic acid sequences encoding any CSR disclosed herein.

[0150] This disclosure provides cells containing any CSR disclosed herein. This disclosure provides cells containing nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides cells containing vectors containing nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides cells containing transposons containing nucleic acid sequences encoding any CSR disclosed herein.

[0151] This disclosure provides compositions comprising any CSR disclosed herein. This disclosure provides compositions comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising vectors comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising transposons comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising modified cells disclosed herein or compositions comprising a plurality of modified cells disclosed herein.

[0152] Methods for site-directed gene integration are also provided herein. Transposase domains and fusion proteins provided herein may be used to deliver a transgene to a cell and integrate the transgene into a target site. The target site may be, for example, a genome-safe harbor, i.e., a genomic site into which the transgene can be integrated in such a way that it is ensured to function predictably and not cause any changes in the host genomic DNA sequence. In some embodiments, the target site is a repeat element, such as an LPA sequence. One, two, or more target sites may be present within a single repeat element. In some embodiments, the target site is located within an intron (e.g., an intron of an LPA gene).

[0153] Site-directed integration can be used in vitro or in vivo. One example of in vivo application is gene therapy, which involves the delivery of a transgene to the genomic DNA of a cell.

[0154] Formulation, dosage, and method of administration This disclosure provides formulations, dosages, and methods of administration of compositions and cells as described herein. In one embodiment, a pharmaceutical composition comprising a tandem dimer transposase or fusion protein as described herein and a pharmaceutically acceptable carrier is provided herein. In another embodiment, a pharmaceutical composition comprising modified cells as described herein and a pharmaceutically acceptable carrier is provided herein.

[0155] The disclosed compositions and pharmaceutical compositions may, but are not limited to, comprise at least one of any suitable adjuvants, such as diluents, binders, stabilizers, buffers, salts, lipophilic solvents, preservatives, and adjuvants. Pharmaceutically acceptable adjuvants are preferred. Non-limiting examples of such sterile solutions and methods for their preparation are well known in the art, including, but are not limited to, Gennaro, Ed., Remington's Pharmaceutical Sciences, 18th Edition, Mack Publishing Co. (Easton, Pa.) 1990 and "Physician's Desk Reference," 52nd ed., Medical Economics (Montvale, NJ) 1998. Pharmaceutically acceptable carriers suitable for the dosage, solubility, and / or stability of protein backbone, fragment, or variant compositions can be typically selected, as are well known in the art or as described herein.

[0156] Non-limiting examples of pharmaceutical excipients and additives suitable for use include proteins, peptides, amino acids, lipids, and carbohydrates (e.g., monosaccharides, sugars including di-, tri-, tetra-, and oligosaccharides, derivatized sugars such as alditol, aldonic acid, and esterified sugars, and polysaccharides or sugar polymers), which can exist alone or in combination, and which together constitute 1 to 99.99% by weight or volume. Non-limiting examples of protein excipients include serum albumins such as human serum albumin (HSA), recombinant human albumin (rHA), gelatin, and casein. Representative amino acid / protein components that can also function as buffers include alanine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, and aspartame. One preferred amino acid is glycine.

[0157] Non-limiting examples of suitable carbohydrate excipients include monosaccharides such as fructose, maltose, galactose, glucose, D-mannose, and sorbose; disaccharides such as lactose, sucrose, trehalose, and cellobiose; polysaccharides such as raffinose, melegitose, maltodextrin, dextran, and starch; and algitols such as mannitol, xylitol, maltitol, lactitol, xylitol sorbitol (glucitol), and myo-inositol. Preferably, the carbohydrate excipient is mannitol, trehalose, and / or raffinose.

[0158] This composition may also contain a buffer or pH adjuster, typically a buffer being a salt prepared from an organic acid or base. Typical buffers include organic acid salts such as citric acid, ascorbic acid, gluconic acid, carbonate, tartaric acid, succinic acid, acetic acid, and phthalic acid salts, as well as Tris, tromethamine hydrochloride, and phosphate buffer. Preferred buffers are organic acid salts such as citrate.

[0159] Furthermore, the disclosed compositions may include polymer excipients / additives, such as polyvinylpyrrolidone, Ficol (polymer sugar), dextrose (e.g., cyclodextrin, e.g., 2-hydroxypropyl-β-cyclodextrin), polyethylene glycol, flavoring agents, antimicrobial agents, sweeteners, antioxidants, antistatic agents, surfactants (e.g., polysorbates, e.g., "TWEEN 20" and "TWEEN 80"), lipids (e.g., phospholipids, fatty acids), steroids (e.g., cholesterol), and chelating agents (e.g., EDTA).

[0160] Many known and developed methods can be used to administer a therapeutically effective amount of the compositions or pharmaceutical compositions disclosed herein. Non-limiting examples of administration methods include bolus, intrabuccal mucosal, injection, intraarticular, intrabronchial, intraperitoneal, intrasacral, intracavitary, intracelial, intracavitary, intracerebellar, intraventricular, intracolonic, intracervical, intragastric, intrahepatic, intrafocal, intramuscular, intramyocardial, transnasal, intraocular, intraosseous, intrapelvic, intrapericardial, intraperitoneal, intrapleural, intrabladder, intrapulmonary, rectal, intrarenal, intraretinal, intraspinal cord, synovial, intrathoracic, intrauterine, intratumoral, intravenous, intrabladder, oral, parenteral, rectal, sublingual, subcutaneous, transdermal, or vaginal means. In preferred embodiments, compositions comprising modified cells described herein are administered intravenously, for example, by intravenous infusion.

[0161] The compositions of this disclosure can be prepared for parenteral (subcutaneous, intramuscular, or intravenous) or any other administration, particularly in the form of liquid solutions or suspensions. For parenteral administration, the compositions disclosed herein may be formulated together with a pharmaceutically acceptable parenteral vehicle as a solution, suspension, emulsion, particles, powder, or lyophilized powder, or may be provided separately from the parenteral vehicle. Formulations for parenteral administration may contain, as common excipients, sterile water or saline, polyalkylene glycol, e.g., polyethylene glycol, vegetable oil, hydrogenated naphthalene, etc. Aqueous or oily suspensions for injection can be prepared according to known methods using appropriate emulsifiers or wetting agents and suspending agents. Injectable or infusion formulations may be non-toxic, orally unadministerable diluents such as aqueous solutions, sterile injection solutions, or suspensions in solvents. Usable vehicles or solvents include water, Ringer's solution, isotonic saline, etc., and sterile non-volatile oils can be used as common solvents or suspension solvents. For these purposes, any kind of non-volatile oils and fatty acids can be used, such as natural, synthetic, or semi-synthetic fatty oils or fatty acids, and natural, synthetic, or semi-synthetic mono-, di-, or tri-glycerides. Parenteral administration is known in the art and includes, but is not limited to, conventional injection methods, gas-pressurized needleless injection devices such as those described in U.S. Patent No. 5,851,198, and laser puncture devices such as those described in U.S. Patent No. 5,839,446.

[0162] It may be desirable to deliver the disclosed compound to the subject in a single dose over a long period, for example, from one week to one year. Various sustained-release, depot, and implant formulations can be utilized. For example, the formulation may include pharmaceutically acceptable non-toxic salts of compounds with low solubility in body fluids, e.g., (a) acid addition salts with polybasic acids, e.g., phosphoric acid, sulfuric acid, citrate, tartaric acid, tannic acid, pamoic acid, alginic acid, polyglutamic acid, naphthalene mono or disulfonic acid, polygalacturonic acid, etc., (b) salts with polyvalent metal cations, e.g., zinc, calcium, bismuth, barium, magnesium, aluminum, copper, cobalt, nickel, cadmium, etc., or salts with organic cations formed from, for example, N,N'-dibenzyl ethylenediamine or ethylenediamine, or (c) a combination of (a) and (b), e.g., zinc tannate salt. Furthermore, the disclosed compounds or preferably relatively insoluble salts, such as those described above, can be formulated into gels suitable for injection, such as aluminum monostearate gel containing sesame oil. Particularly preferred salts include zinc salts, zinc tannate salts, and pamoate salts. Another type of sustained-release depot formulation for injection contains the compound or salt dispersed for encapsulation in a slowly degradable, non-toxic, non-antigenic polymer, such as polylactic acid / polyglycolic acid polymer, as described, for example, in U.S. Patent No. 3,773,919. The compounds or preferably relatively insoluble salts, such as those described above, can also be formulated into cholesterol matrix silastic pellets, particularly for use in animals. Further sustained-release formulations, depot formulations, or implant formulations, such as gaseous or liquid liposomes, are known in the literature (U.S. Patent No. 5,770,222 and “Sustained and Controlled Release Drug Delivery Systems”, JR Robinson ed. Marcel Dekker, Inc., NY, 1978).

[0163] Treatment method In another embodiment, a method for treating a disease or disorder of interest is provided herein, comprising administering a composition comprising modified cells as described herein to a subject. The terms “subject” and “patient” are used interchangeably herein. In a preferred embodiment, the patient is human.

[0164] The modified cells may be allogeneic or autologous to the patient. In some preferred embodiments, the modified cells are allogeneic cells. In some embodiments, the modified cells are autologous T cells or modified autologous CAR T cells. In some preferred embodiments, the modified cells are allogeneic T cells or modified allogeneic CAR T cells.

[0165] In some embodiments, the disease or disorder treated according to the methods described herein is cancer. Non-limiting examples of cancer include leukemia, acute leukemia, acute lymphoblastic leukemia (ALL), acute lymphoblastic leukemia, B cell, T cell, or FAB. Examples include ALL, acute myeloid leukemia (AML), acute myeloid leukemia, chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), hairy cell leukemia, myelodysplastic syndrome (MDS), lymphoma, Hodgkin's disease, malignant lymphoma, non-Hodgkin lymphoma, Burkitt lymphoma, multiple myeloma, Kaposi's sarcoma, colorectal cancer, pancreatic cancer, nasopharyngeal cancer, malignant histiocytosis, paraneoplastic syndrome / hypercalcemia of malignant tumors, solid tumors, bladder cancer, breast cancer, colorectal cancer, endometrial cancer, head cancer, cervical cancer, hereditary nonpolyposis cancer, Hodgkin lymphoma, liver cancer, lung cancer, non-small cell lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, renal cell carcinoma, testicular cancer, adenocarcinoma, sarcoma, malignant melanoma, hemangioma, metastatic disease, cancer-related bone resorption, and cancer-related bone pain.

[0166] In some embodiments, the diseases or disorders treated according to the methods described herein are liver diseases or hepatic disorders, urea cycle disorders, metabolic hepatic disorders, or hemophilia. In some embodiments, metabolic hepatic disorders may be ornithine transcarbamylase (OTC) deficiency. In some embodiments, metabolic hepatic disorders may be methylmalonic acidemia (MMA).

[0167] In non-limiting examples, the present disclosure provides methods of treating a subject having hemophilia. In some embodiments, the hemophilia can be hemophilia A. In some embodiments, the hemophilia can be hemophilia B.

[0168] In non-limiting examples, the present disclosure provides methods of treating phenylketonuria (PKU) in a subject.

[0169] In some embodiments, the present disclosure provides methods of treating autoimmune diseases. In some embodiments, the autoimmune disease is autoimmune neutropenia, Guillain-Barré syndrome, epilepsy, autoimmune encephalitis, Isaac's syndrome, vitiligo syndrome, pemphigus vulgaris, pemphigus foliaceus, bullous pemphigoid, epidermolysis bullosa acquisita, pemphigoid gestationis, mucous membrane pemphigoid, antiphospholipid syndrome, autoimmune anemia, myasthenia gravis, autoimmune Graves' disease, thyroid eye disease (TED), Goodpasture's syndrome, multiple sclerosis, rheumatoid arthritis, lupus, idiopathic thrombocytopenic purpura (ITP), warm autoimmune hemolytic anemia (WAIHA), chronic inflammatory demyelinating polyneuropathy (CIDP), lupus nephritis or membranous nephropathy.

[0170] The dosage of the pharmaceutical composition administered to a subject can be varied according to known factors such as the pharmacodynamic properties of the particular agent, and its mode and route of administration, the age, health, and weight of the recipient, the nature and degree of the symptoms, the type of concurrent treatment, the frequency of treatment, and the desired effect.

[0171] In embodiments where the composition administered to a subject that needs it is the modified cells disclosed herein, 1×10 3 ~ about 1×10 4 cells, about 1×10 4 ~ about 1×10 5 cells, about 1×10 5 ~ about 1×10 6 cells, about 1×10 6 ~ about 1×10 7 cells, about 1×10 7 ~ about 1×10 8cells, approximately 1 x 10 8 ~Approx. 1×10 9 cells, approximately 1 x 10 9 ~Approx. 1×10 10 cells, approximately 1 x 10 10 ~Approx. 1×10 11 cells, approximately 1 x 10 11 ~Approx. 1×10 12 cells, approximately 1 x 10 12 ~Approx. 1×10 13 cells, approximately 1 x 10 13 ~Approx. 1×10 14 cells, approximately 1 x 10 14 ~Approx. 1×10 15 cells, approximately 1 x 10 15 ~Approx. 1×10 16 cells, approximately 1 x 10 16 ~Approx. 1×10 17 cells, approximately 1 x 10 17 ~Approx. 1×10 18 cells, approximately 1 x 10 18 ~Approx. 1×10 19 Cells, or approximately 1 × 10⁻⁶ 19 ~Approx. 1×10 20 Cells may be administered. In some embodiments, the cells are approximately 5 × 10 6 ~Approx. 25×10 6 It is administered in cellular doses.

[0172] In other embodiments, the cell dosage may depend on the human body weight, for example, about 1 × 10⁻⁶ 3 ~Approx. 1×10 4 cells, approximately 1 x 10 4 ~Approx. 1×10 5 cells, approximately 1 x 10 5 ~Approx. 1×10 6 cells, approximately 1 x 10 6 ~Approx. 1×10 7 cells, approximately 1 x 10 7 ~Approx. 1×10 8 cells, approximately 1 x 10 8 ~Approx. 1×10 9 cells, approximately 1 x 10 9 ~Approx. 1×10 10 cells, approximately 1 x 10 10 ~Approx. 1×10 11 cells, approximately 1 x 10 11 ~Approx. 1×10 12cells, approximately 1 x 10 12 ~Approx. 1×10 13 cells, approximately 1 x 10 13 ~Approx. 1×10 14 cells, approximately 1 x 10 14 ~Approx. 1×10 15 cells, approximately 1 x 10 15 ~Approx. 1×10 16 cells, approximately 1 x 10 16 ~Approx. 1×10 17 cells, approximately 1 x 10 17 ~Approx. 1×10 18 cells, approximately 1 x 10 18 ~Approx. 1×10 19 Cells, or approximately 1 × 10⁻⁶ 19 ~Approx. 1×10 20 Cells can be administered per kilogram of body weight of the target individual.

[0173] A more detailed description of the disclosed compositions and pharmaceutically acceptable excipients, formulations, dosages, and methods of administration of the pharmaceutical compositions is disclosed in PCT Publication No. WO2020 / 051374.

[0174] The transposase domains and fusion proteins provided herein may be used to deliver gene therapy. Gene therapy typically involves the delivery of a transgene into the genomic DNA of a cell. Typically, the transgene replaces a gene that is mutated or otherwise not properly expressed in the cell. For example, a transgene may replace a gene that exhibits reduced expression, insufficient expression, and / or altered expression in a cell. In some embodiments, such reduced, insufficient, and / or altered expression may directly or indirectly result in a disease or disorder, such as liver disease or disorder, urea cycle disorder, metabolic hepatopathy, or hemophilia. The fusion proteins, transposase domains, and complexes described herein may be used to deliver a therapeutic transgene to a cell and to incorporate the transgene into a target site. In some embodiments, a therapeutic method involves introducing the fusion protein and transposon provided herein into a cell, wherein the transposon comprises a 5'ITR, a transgene, and a 3'ITR in the order of 5' to 3'.

[0175] In some embodiments, the therapeutic transgene is a gene that is expressed at a lower level, and this lower expression results in a disease or disorder. In some embodiments, the therapeutic transgene is a gene that is expressed in a modified pattern compared to the wild-type gene, and this modified expression results in a disease or disorder. Accordingly, methods for treating diseases or disorders caused by or associated with changes in gene expression are provided herein, comprising administering the transposons and transposases described herein to a subject in need thereof.

[0176] The fusion proteins, transposase domains, and therapeutic transgenes delivered to cells by the complex described herein may encode therapeutic polypeptides. In some embodiments, the therapeutic polypeptide is a factor VIII polypeptide, a factor IX polypeptide, a phenylalanine hydroxylase (PAH), an ornithine transcarbamylase (OTC) polypeptide, or a methylmalonyl-CoA mutase (MUT1) polypeptide.

[0177] In non-limiting examples, the transposase domains and fusion proteins provided herein may be used to deliver liver-directed gene therapy. In some embodiments, liver-directed gene therapy can be used to treat ornithine transcarbamylase (OTC) deficiency, and the therapeutic polypeptide encoded by the therapeutic transgene may include ornithine transcarbamylase (OTC) polypeptide. In some embodiments, liver-directed gene therapy can be used to treat methylmalonic acidemia (MMA), and at least one therapeutic protein encoded by the therapeutic transgene may include methylmalonyl-CoA mutase (MUT1) polypeptide.

[0178] In some embodiments, liver-directed gene therapy can be used to treat hemophilia A, and at least one therapeutic protein encoded by the therapeutic transgene may include factor VIII. In some embodiments, liver-directed gene therapy can be used to treat hemophilia B, and at least one therapeutic protein encoded by the therapeutic transgene may include factor IX.

[0179] In some embodiments, liver-directed gene therapy can be used to treat phenylketonuria (PKU), and at least one therapeutic protein encoded by the therapeutic transgene may include phenylalanine hydroxylase (PAH).

[0180] kit In another embodiment, a kit is provided herein comprising a cell line engineered to contain a modified target site of an SPB or PBx provided herein within its genome, preferably within a highly expressed genomic region. The kit may further comprise a composition comprising one or more SPB or PBx transposase domains or fusion proteins described herein. In some embodiments, the cell line is a T cell line.

[0181] definition As used throughout this disclosure, the singular forms "a," "an," and "the" include multiple referents unless the context otherwise explicitly indicates otherwise. Thus, for example, a reference to "a method" includes multiple such methods, and a reference to "a dose" includes one or more doses and their equivalents that are known to those skilled in the art.

[0182] The terms “about” or “approximately” mean that a particular value is within an acceptable margin of error, as determined by those skilled in the art, and this depends to some extent on the method by which the value is measured or determined, e.g., on the limitations of the measurement system. For example, “about” means within one or more standard deviations. Alternatively, “about” may mean a range of up to 20%, or up to 10%, or up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term may mean within one order of magnitude of the value, preferably up to five times, and more preferably up to two times. Where a particular value is described in this application and claims, unless otherwise specified, the term “about” should be assumed to mean within an acceptable margin of error of that particular value.

[0183] This disclosure provides isolated or substantially purified polynucleotide or protein compositions. “Isolated” or “purified” polynucleotides or proteins, or their biologically active portions, substantially or essentially contain no components that would normally accompany or interact with polynucleotides or proteins found in their naturally occurring environments. Therefore, isolated or purified polynucleotides or proteins, if produced by recombinant technology, substantially contain no other cell material or culture medium, and if chemically synthesized, substantially contain no chemical precursors or other chemicals. Optimally, an “isolated” polynucleotide does not contain sequences naturally adjacent to it in the genomic DNA of the organism from which it originates (i.e., sequences located at the 5' and 3' ends of the polynucleotide (optimally, protein-coding sequences)). For example, in various embodiments, an isolated polynucleotide may contain approximately 5kb, 4kb, 3kb, 2kb, 1kb, 0.5kb, or less than 0.1kb of nucleotide sequences naturally adjacent to it in the genomic DNA of the cell from which it originates. Substantially cellular protein-free proteins include protein preparations containing approximately 30%, 20%, 10%, 5%, or less than 1% (dry weight) of contaminating protein. When the proteins of this disclosure or their biologically active portions are recombinantly produced, the culture medium optimally contains approximately 30%, 20%, 10%, 5%, or less than 1% (dry weight) of chemical precursors or non-protein-of-interest chemicals.

[0184] This disclosure provides disclosed DNA sequence fragments and variants, as well as proteins encoded by these DNA sequences. As used throughout this disclosure, the term “fragment” refers to a portion of a DNA sequence, or a portion of an amino acid sequence, and by extension, the protein encoded thereby. A DNA sequence fragment consisting of a coding sequence may encode a protein fragment that retains the biological activity of the native protein and therefore retains DNA recognition or binding activity to a target DNA sequence, as described herein. Alternatively, DNA sequence fragments useful as hybridization probes generally do not encode a protein that retains biological activity or promoter activity. Therefore, DNA sequence fragments may range from at least about 20 nucleotides, about 50 nucleotides, about 100 nucleotides to the full-length polynucleotides of this disclosure.

[0185] The nucleic acids or proteins of this disclosure can be constructed by a modular approach, which involves pre-assembling monomer units and / or repeat units in a target vector and then assembling them into a final target vector. The polypeptides of this disclosure can be constructed by a modular approach, which involves pre-assembling repeat units in a target vector that can be composed of repeat monomers of this disclosure and then assembled into a final target vector. This disclosure provides polypeptides produced by this method and nucleic acid sequences encoding these polypeptides. This disclosure provides host organisms and cells containing nucleic acid sequences encoding polypeptides produced by this modular approach.

[0186] The term “comprising” is intended to mean that a composition and method includes the elements described but does not exclude others. “Consisting essentially of,” when used to define a composition and method, means excluding other elements that are essentially important to the combination for the described purpose. Thus, a composition essentially consisting of the components defined herein does not exclude trace amounts of contaminants or inert carriers. “Consisting of” means excluding other components and elements that are more than trace amounts of substantial method steps. The embodiments defined by each of these transitional terms are within the scope of this disclosure.

[0187] As used herein, “expression” refers to the process by which a polynucleotide is transcribed into mRNA, and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression includes the splicing of mRNA in eukaryotic cells.

[0188] "Gene expression" is the process of converting the information contained in a gene into a gene product. A gene product can be a direct transcript of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, shRNA, microRNA, structural RNA, or other types of RNA) or a protein produced by the translation of mRNA. Gene products also include RNA modified by processes such as capping, polyadenylation, methylation, and editing, as well as proteins modified by processes such as methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristylation, and glycosylation.

[0189] The "regulation" or "control" of gene expression refers to a change in gene activity. Regulation of expression includes, but is not limited to, gene activation and gene repression.

[0190] The term "operatively linked" or its synonyms (e.g., "linked operatively") means that two or more molecules are positioned relative to each other so that they can interact in a way that influences the function of one or both of the molecules or a combination thereof. In relation to nucleic acids, a promoter may be operatively linked to a nucleotide sequence encoding a transfer domain or fusion protein as described herein, thereby placing the expression of the nucleotide sequence under the control of the promoter.

[0191] Components linked by non-covalent bonds, and methods for producing and using non-covalently linked components are disclosed. Various components can take on a variety of different forms, as described herein. For example, non-covalently linked (i.e., operably linked) proteins can be used to enable transient interactions that circumvent one or more problems in the art. The ability of non-covalently linked components, such as proteins, to associate and dissociate allows for functional association only under circumstances where such association is required for the desired activity, or primarily under such circumstances. The linkage only needs to last long enough to achieve the desired effect.

[0192] A method for inducing a protein to a specific gene locus in the genome of an organism is disclosed. This method may include a step of providing a DNA localization component and a step of providing an effector molecule, the DNA localization component and the effector molecule being operablely linked via non-covalent linkage.

[0193] A "target site" or "target sequence" is a nucleic acid sequence that defines the portion of the nucleic acid to which a binding molecule will bind, provided that sufficient conditions for binding are present.

[0194] The terms “nucleic acid,” “oligonucleotide,” or “polynucleotide” refer to at least two nucleotides linked by a covalent bond. A single-stranded description also defines the sequence of the complementary strand. Thus, a nucleic acid may encompass the complementary strand of a described single-stranded strand. The nucleic acids of this disclosure also encompass substantially identical nucleic acids and their complements that retain the same structure or encode the same protein.

[0195] The nucleic acids of this disclosure may be single-stranded or double-stranded. The nucleic acids of this disclosure may contain double-stranded sequences even if the majority of the molecule is single-stranded. The nucleic acids of this disclosure may contain single-stranded sequences even if the majority of the molecule is double-stranded. The nucleic acids of this disclosure may include genomic DNA, cDNA, RNA, or hybrids thereof. The nucleic acids of this disclosure may include combinations of deoxyribonucleotides and ribonucleotides. The nucleic acids of this disclosure may include combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. The nucleic acids of this disclosure may be synthesized to include non-natural amino acid modifications. The nucleic acids of this disclosure may be obtained by chemical synthesis or recombinant methods.

[0196] The nucleic acids disclosed herein may not exist in nature in whole or in part. The nucleic acids disclosed may contain one or more mutations, substitutions, deletions, or insertions that do not exist in nature, and the entire nucleic acid sequence may not exist in nature. The nucleic acids disclosed may contain one or more duplicate sequences, reverse sequences, or repeat sequences, and as a result, such sequences do not exist in nature, and the entire nucleic acid sequence may not exist in nature. The nucleic acids disclosed may contain modified nucleotides, artificial nucleotides, or synthetic nucleotides that do not exist in nature, and the entire nucleic acid sequence may not exist in nature.

[0197] If the genetic code contains redundancy, multiple nucleotide sequences may encode a particular protein. All such nucleotide sequences are assumed herein.

[0198] As used throughout this disclosure, the term “promoter” refers to a synthetic or naturally occurring molecule that can confer, activate, or enhance the expression of a nucleic acid within a cell. A promoter may include one or more specific transcriptional regulatory sequences to further enhance expression and / or alter its spatial and / or temporal expression. A promoter may also include distal enhancer or repressor elements located thousands of base pairs away from the transcription start site. Promoters may originate from viruses, bacteria, fungi, plants, insects, animals, and the like. A promoter may constitutively, constitutively, or differentially control the expression of a gene component with respect to the cell, tissue, or organ in which the expression occurs, or to the developmental stage in which the expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions, or inducers. Representative examples of promoters include the bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, EF-1 alpha promoter, CAG promoter, SV40 early promoter or SV40 late promoter, and CMV IE promoter.

[0199] As used throughout this disclosure, the term “vector” refers to a nucleic acid sequence containing an origin of replication. Examples of vectors include viral vectors, bacteriophages, bacterial artificial chromosomes, and yeast artificial chromosomes. A vector may be either a DNA vector or an RNA vector. A vector may be a self-replicating extrachromosomal vector, preferably a DNA plasmid. A vector may contain amino acids and a DNA sequence, an RNA sequence, or a combination of both DNA and RNA sequences.

[0200] Conservative substitutions of amino acids, i.e., substitution of an amino acid with a different amino acid having similar properties (e.g., hydrophilicity, degree and distribution of charged regions), are typically recognized in the art as involving only minor changes. These small changes can be partially identified by considering the hydrophobicity index of amino acids, as understood in the art. Kyte et al., J.Mol.Biol.157:105-132 (1982). The hydrophobicity index of an amino acid takes into account its hydrophobicity and charge. Substitution with amino acids having similar hydrophobicity indices can maintain the function of the protein. In some embodiments, amino acids with a hydrophobicity index of ±2 are substituted. The hydrophilicity of amino acids can also be used to identify substitutions that maintain the biological function of the protein. By considering the hydrophilicity of amino acids in the context of peptides, it is possible to calculate the maximum local mean hydrophilicity of the peptide, which is a useful indicator that has been reported to correlate well with antigenicity and immunogenicity. U.S. Patent No. 4,554,101 is fully incorporated herein by reference.

[0201] By substituting amino acids with similar hydrophilicity values, peptides with preserved biological activity, such as immunogenicity, can be obtained. Substitutions can be performed with amino acids whose hydrophilicity values ​​are within ±2 of each other. Both the hydrophobicity index and hydrophilicity of an amino acid are influenced by its specific side chain. Consistent with this observation, it is understood that amino acid substitutions suitable for biological function depend on the relative similarity of amino acids, particularly their side chains, as revealed by their hydrophobicity, hydrophilicity, charge, size, and other properties.

[0202] As used herein, “conservative” amino acid substitutions may be defined as shown in Tables 3, 4, and 5 below. In some embodiments, fusion polypeptides and / or nucleic acids encoding such fusion polypeptides include conservative substitutions introduced by modifying the polynucleotides encoding the polypeptides of this disclosure. Amino acids can be classified by their physical properties and their contribution to the secondary and tertiary structures of proteins. A conservative substitution is the replacement of one amino acid with another amino acid having similar properties. Exemplary conservative substitutions are shown in Table 3. [Table 3]

[0203] Alternatively, conserved amino acids can be classified as shown in Table 4, as described by Lehninger (Biochemistry, Second Edition; Worth Publishers, Inc. NY, NY (1975), pp. 71-77). [Table 4]

[0204] Alternatively, exemplary conservative substitutions are shown in Table 5. [Table 5]

[0205] The polypeptides and proteins of this disclosure may have sequences, or parts thereof, that do not exist in nature. The polypeptides and proteins of this disclosure may contain one or more mutations, substitutions, deletions, or insertions that do not exist in nature, and the entire amino acid sequence may not exist in nature. The polypeptides and proteins of this disclosure may contain one or more duplicate sequences, reverse sequences, or repeat sequences, and as a result, the sequence may not exist in nature, and the entire amino acid sequence may not exist in nature. The polypeptides and proteins of this disclosure may contain modified amino acids, artificial amino acids, or synthetic amino acids that do not exist in nature, and the entire amino acid sequence may not exist in nature.

[0206] As used throughout this disclosure, the identity between two sequences may be determined using a standalone executable BLAST engine program (bl2seq) for blasting two sequences, which can be obtained from the National Center for Biotechnology Information (NCBI) ftp site using default parameters (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250, which is incorporated herein by reference in its entirety). When used in the context of two or more nucleic acid or polypeptide sequences, the terms “identical” or “same” refer to a specific percentage of the same residues across a particular region of each sequence. In some embodiments, sequence identity is determined across the entire length of the sequences. The percentage can be calculated by optimally aligning the two sequences, comparing the two sequences over a given region, determining the number of positions where identical residues appear in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the given region, and multiplying the result by 100 to obtain the percentage of sequence identity. If the two sequences have different lengths, or if alignment generates sequences with one or more ends shifted, and the specified comparison region contains only a single sequence, the residues of the single sequence are included in the denominator but not in the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identification can be performed manually or using computer sequencing algorithms such as BLAST or BLAST 2.0.

[0207] In certain embodiments, if a sequence has a specific sequence identity with a specific sequence number (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%), the sequence and the sequence number have the same length. In certain embodiments, if a sequence has a specific sequence identity with a specific sequence number (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%), the sequence and the sequence number differ only for the sake of conservation amino acid substitutions.

[0208] As used throughout this disclosure, the term “endogenous” refers to a nucleic acid or protein sequence that is naturally associated with the target gene or the host cell into which it is introduced.

[0209] As used throughout this disclosure, the term “exogenous” means a nucleic acid or protein sequence that is not naturally associated with the target gene or the host cell into which it is introduced, and includes multiple copies of naturally occurring nucleic acids that are not naturally occurring, such as DNA sequences, or naturally occurring nucleic acid sequences located at genomic locations that are not naturally occurring.

[0210] This disclosure provides a method for introducing a polynucleotide construct containing a DNA sequence into a host cell. “Introducing” means presenting the polynucleotide construct to the cell in a manner that allows access to the interior of the host cell. The method of this disclosure does not depend on a specific method for introducing the polynucleotide construct into a host cell, but only on the polynucleotide construct accessing the interior of a single host cell. Methods for introducing polynucleotide constructs into bacteria, plants, fungi, and animals are known in the art, but are not limited to stable transformation methods, transient transformation methods, and virus-mediated methods. [Examples]

[0211] The examples in this section are provided for illustrative purposes only and are not intended to limit the invention.

[0212] Example 1: Construction of Super PiggyBac transposase with amino-terminal deletion Plasmids containing either a nucleotide sequence encoding the full-length wild-type Super PiggyBac transposase (SPB; SEQ ID NO: 2) or a nucleotide sequence encoding a combined deletion variant of Super PiggyBac transposase (PBx; SEQ ID NO: 3) with amino acid substitutions at positions R372A, K375A, and D450N were used as templates for PCR mutagenesis to generate N-terminal deletion transposase variants lacking the N-terminal 93 amino acids (SPBΔ1-93 and PBxΔ1-93, respectively).

[0213] In short, forward and reverse primers were designed to amplify portions of the SPB and PBx coding sequences corresponding to amino acids 94–594. The resulting DNA fragments encoding SPBΔ1–93 or PBxΔ1–93 were used with purchased gBlock gene fragments to construct DNA-binding domain-transposase fusion proteins via a cutting-edge two-fragment Gibson Assembly.

[0214] Additional N-terminal deletion transposase variants lacking the N-terminal 85 amino acids (SPBΔ1-85 and PBxΔ1-85, respectively) were generated as described herein.

[0215] Example 2: Design and construction of a TAL array targeting LPA This example illustrates the design and construction of a TAL array composition targeting the LPA gene, which may be used in a method for verifying the target specificity of TAL arrays. The TAL array was constructed using the design criteria shown below.

[0216] The lipoprotein A (LPA) gene contains up to 50 copies of partially duplicated elements, making it a potentially attractive target for optimizing the likelihood of site-directed metastasis events in target sequences, thereby leading to an increase in the number of metastatic cells.

[0217] TAL array pairs containing the N-terminal domain that recognizes T were designed targeting four specific 10 bp left-right pair sequences within the repeat element of the LPA gene. For three of these targets, multiple TAL array pairs were designed using either a 12 bp or 13 bp spacer.

[0218] The left and right target sequences, along with the upstream 5'T used to construct the TAL array targeting the LPA gene, are shown in Table 7. [Table 7]

[0219] Individual TAL modules containing 34-amino acid repeats or 20-amino acid "half" repeats were synthesized by flanking them with BsmBI type IIS restriction sites. The entire module set contained four modules capable of recognizing either A, C, G, or T for each 10bp position within the target sequence (40 modules / 10bp target). Pairs of TAL arrays targeting sequences within the LPA gene were designed, corresponding modules were selected, and they were pooled together using "Golden Gate Assembly" and assembled in-frame to create each LPA TAL array. All coding sequences used were codon-optimized for human expression.

[0220] Using seven pairs of left and right combinations, we designed and constructed the left-side LPA TAL arrays LPAL1, LPAL2, LPAL3, LPAL4.1, and LPAL4.2 (sequences 116, 118, 121, 124, and 125, respectively) and the right-side LPA TAL arrays LPAR1, LPAR2.1, LPAR2.2, LPA3.1, LPAR3.2, and LPAR4 (sequences 117, 119, 120, 122, 123, and 126, respectively).

[0221] Example 3: Construction and analysis of TAL arrays - piggyBac transposase (ss-SPB) composition (TAL-PBxs) designed for site-directed transposition in the LPA gene This example illustrates the construction of a TAL array Super piggyBac transposase fusion protein composition (TAL-ssSPB), which is useful in methods for achieving site-directed transposition at specific target gene loci.

[0222] An expression plasmid was synthesized containing the following TAL-PBx fusion constructs: 5' to 3' direction: CMV promoter, T7 promoter, Kozak sequence, 3× Flag tag (SEQ ID NO: 65), SV40 NLS (SEQ ID NO: 66), Delta 152 TAL N-terminal domain (SEQ ID NO: 31), two BsmBI type IIS restriction enzyme sites, +63 TAL C-terminal domain (SEQ ID NO: 32), GGGS linker, delta 1-93 PBx (containing N-terminal 93 amino acid deletions and mutations at R372A, K375A, D450N in the Super piggyBac transposer zecodon sequence; SEQ ID NO: 6), and bGH polyadenylated sequence.

[0223] Cloning a left- or right-sided TAL array flanked by BsmBIs to the BsmBI site of an expression plasmid results in in-frame fusion of the TAL array and PBx coding sequence via a linker sequence that generates a full-length TAL-PBx construct. All coding sequences used were codon-optimized for human expression using the GeneArt algorithm (Thermo Fisher).

[0224] Eleven TAL arrays designed and constructed in Example 2, flanked at the BsmBI termini, were cloned into the BsmBI restriction site of the expression plasmid described above to produce eleven TAL-PBx constructs: LPAL1, LPAL2, LPAL3, LPAL4.1, and LPAL4.2 left-side TAL-PBxs (sequences 143, 145, 148, 151, and 152, respectively) and LPAR1, LPAR2.1, LPAR2.2, LPA3.1, LPAR3.2, and LPAR4 right-side TAL-PBxs (sequences 144, 146, 147, 149, 150, and 153, respectively).

[0225] Example 4: Demonstration of site-directed transposition using a TAL array - piggyBac transposase (ss-SPB) composition (TAL-PBxs) and episomal split GFP splicing reporter system This example describes exemplary compositions and methods for demonstrating site-directed transposition at specific episomal loci using a TAL array-SPB transposase fusion protein.

[0226] The episomal split GFP splicing reporter system was used to evaluate the site-directed transposition efficiency of various TAL array-SPB transposase fusion proteins constructed in Example 3. The reporter system consists of two plasmids. A first plasmid, the "reporter," was constructed, containing the EF1a promoter (SEQ ID NO: 67), Kozak sequence, the first portion of the GFP open reading frame (SEQ ID NO: 68), a splice donor (SEQ ID NO: 69), and two BsaI type IIS restriction enzyme sites in the 5' to 3' direction. The BsaI sites allow cloning of target TTAA sequences adjacent to a variable-length spacer flanked by target recognition sequences of the TAL array. A second plasmid, the "donor," was constructed containing the TTAA sequence, a 35 bp PiggyBac minimum 5' ITR (SEQ ID NO: 70), a splice acceptor site (SEQ ID NO: 71), the second part of the GFP open reading frame (SEQ ID NO: 72), a synthetic polyadenylated sequence (SEQ ID NO: 73), a 63 bp PiggyBac minimum 3' ITR (SEQ ID NO: 74), and the TTAA sequence in the 5' to 3' direction. A schematic diagram of the Split GFP reporter plasmid is shown in Figure 2.

[0227] Four different naturally occurring LPA target sequences (SEQ ID NOs. 81-84) were cloned into the episomal reporter plasmid described above. Complementary oligonucleotides containing the LPA genomic DNA sequences (SEQ ID NOs. 81-84) were synthesized. The complementary oligonucleotides contained a 4 bp overhang that matched the overhang created in the split GFP splicing reporter after digestion with BsaI. The oligonucleotides were annealed and ligated into the digested vector to create reporters that matched each LPA TAL-PBx pair constructed in Example 3.

[0228] TAL arrays were designed and constructed to produce heterodimer pairs of TAL-ssSPB (i.e., one left-right TAL array-PBx). Each TAL-PBx construct pair was co-transfected into HEK293T cells with its corresponding reporter plasmid and donor plasmid. As a negative control, each TAL-PBx construct pair was co-transfected into HEK293T cells with mismatched reporter plasmids (i.e., TAL-PBx pair 1 had reporter 2, TAL-PBx pair 2 had reporter 3, TAL-PBx pair 3 had reporter 4, and TAL-PBx pair 4 had reporter 1) and donor plasmids. A transfection mixture containing 26 ng of TAL-ssSPB expression vector, 170 ng of reporter plasmid, 117 ng of donor plasmid, and 0.78 μl of Transit-2020 transfection reagent was assembled in a total volume of 26 μl of Serum-Free OptiMem medium. 95,000 HEK293T cells were added to 250 μl of DMEM medium supplemented with 10% FBS, the transfection mixture was seeded into 48-well plates, incubated at 37°C with 5% CO2 for 4 days, and the cells were divided into a 1:3 ratio on day 2.

[0229] When reporter and donor plasmids are co-transfected into cells with TAL-PBx, TAL-PBx catalyzes the excision of transposons from the donor plasmid and their site-specific integration into the TTAA target site of the reporter plasmid. Figure 3 is a schematic diagram showing a catalytic ssSPB dimer that binds to the excised transposon and recognizes its genomic integration target site. Following site-specific transposition, transcription, splicing, and translation, a reconstructed GFP coding sequence is generated (DNA, SEQ ID NO: 75; amino acids; SEQ ID NO: 76), which can be detected for fluorescence. The percentage of on-target site-specific transposition-positive cells for various TAL-PBx pairs was determined by FACS analysis, and the results are shown in Table 8. [Table 8]

[0230] As shown in Table 8, all TAL-ssSPBs catalyzed site-specific translocation for their respective on-target reporters, but reporters containing mismatched off-targets did not. Furthermore, the highest translocation was observed at target 1, the only target with a TTTAAA integration site.

[0231] Example 5: Determination of optimal adjacent 5' and 3' nucleotides directly adjacent to the TTAA integration site The aforementioned example demonstrates that target 1, the target site with the most robust integration, contains 5'T and 3'A nucleotides directly adjacent to the TTAA target site, generating a TTTAAA integration site. This example illustrates additional compositions and methods for preparing optimal target sites for site-directed transposition by determining the optimal adjacent 5' and 3' nucleotides directly adjacent to the TTAA integration site.

[0232] The episomal split GFP splicing reporter described in Example 4 was used to evaluate the site-directed transposition efficiency of various TAL-PBx fusion proteins targeting the green fluorescent protein (GFP) gene. SPB transposase fusion proteins GFP1 right-side TAL-PBx and GFP1 left-side TAL-PBx, targeting specific 10 bp right-side and 10 bp left-side sequences within the coding region of the TAL array-GFP gene, were prepared as described in Examples 14 and 18 of International Patent Application Publication PCT / US2022 / 77549, the contents of which are incorporated in their entirety by reference.

[0233] To construct a reporter plasmid compatible with GFP1 right-sided TAL-PBx, a complementary oligonucleotide was synthesized containing the target site of GFP1 right-sided TAL downstream of T, followed by a 12bp spacer, then TTAA, followed by another 12bp spacer, then the reverse complement of the TAL target site, followed by A (SEQ ID NO: 172). The spacer sequence was such that the nucleotide immediately 5' of TTAA was C and the nucleotide immediately 3' of TTAA was C. The complementary oligonucleotide contained a 4bp overhang that matched the overhang created in the split GFP splicing reporter after digestion with BsaI. The oligonucleotide was annealed and ligated to the digested vector to construct a reporter compatible with GFP1 right-sided TAL-PBx. Similar oligos were synthesized using a 12bp modified spacer sequence to generate the TTTAAA integration site (SEQ ID NO: 173) by mutating the adjacent 5' and 3' nucleotides directly adjacent to the TTAA integration sequence to T and A, respectively, or to generate the CTTAAA integration site (SEQ ID NO: 174) by mutating them to C and A, respectively. Similar oligos were synthesized (SEQ ID NO: 175) containing the target site of GFP1 Right TAL downstream of T, followed by a 13bp spacer, followed by TTAA, followed by another 13bp spacer, followed by the reverse complement of the TAL target site, followed by A. The spacer sequence was such that the nucleotide immediately 5' of TTAA was C, and the nucleotide immediately 3' of TTAA was C. Similarly, similar oligos were synthesized using a modified 13bp spacer sequence to generate the TTTAAA integration site (SEQ ID NO: 176) by mutating the adjacent 5' and 3' nucleotides directly adjacent to the TTAA integration sequence to T and A, respectively; the CTTAAA integration site (SEQ ID NO: 177) by mutating them to C and A, respectively; the TTTAAG integration site (SEQ ID NO: 178) by mutating them to T and G, respectively; or the CTTAAG integration site (SEQ ID NO: 179) by mutating them to C and G, respectively.

[0234] Each reporter plasmid and donor plasmid were co-transfected into HEK293T cells. The GFP1 Right TAL-PBx expression plasmid (SEQ ID NO: 77) was used. As a negative control, the GFP1 Left TAL-PBx expression plasmid (SEQ ID NO: 78), which does not recognize the GFP1 Right target sequence, was transfected instead of the GFP1 Right TAL-PBx expression plasmid. HEK293T cells were seeded into 24-well plates in 500 μL of DMEM medium supplemented with 10% FBS. The following day, a transfection mixture was assembled in 50 μL of JetPrime buffer containing 50 ng of TAL-ssSPB expression vector, 225 ng of reporter plasmid, 225 ng of donor plasmid, and 1 μL of JetPrime transfection reagent. The mixture was added to HEK293T cells, and the cells were incubated at 37°C and 5% CO2 for 4 days. On day 1, the cells were divided into 1:6. The proportion of on-target site-specific metastasis-positive cells in various constructs was determined by FACS analysis on day 4.

[0235] When reporter and donor plasmids are co-transfected into cells with TAL-PBx, TAL-PBx catalyzes the excision of transposons from the donor plasmid and its site-specific integration into the TTAA target site of the reporter plasmid. Following site-specific transposition, transcription, splicing, and translation, a reconstructed GFP coding sequence is generated (DNA SEQ ID NO: 75; amino acid SEQ ID NO: 76), which can be detected for fluorescence. The percentage of on-target site-specific transposition-positive cells for various spacer length constructs was determined by FACS analysis, and the results are shown in Table 9.

[0236] As shown in Table 9, GFP1 Right TAL-PBx catalyzed site-directed translocation, resulting in GFP signals above background levels at all target sites. The TTTAAA target site produced a stronger GFP signal than the CTTAAC and CTTAAG target sites. The CTTAAA and TTTAAG target sites produced the largest GFP signals. GFP1 left-side TAL-PBx did not produce a GFP signal above background levels using a GFP1 right-specific reporter. [Table 9]

[0237] Example 6: Demonstration of site-directed transposition using a TAL array - piggyBac transposase (ss-SPB) composition (TAL-PBxs) Based on the results of Example 5, a second set of four different LPA target sequences naturally found in genomic DNA (SEQ ID NOs. 85-88) was cloned into the episomal reporter plasmid described in Example 4. Similar to the first target set evaluated in Example 4, each target sequence in the second set has a 10 bp TAL binding site and either a 12 bp or 13 bp spacer on either side of TTAA. Furthermore, each target sequence in the second set contains a spacer sequence such that the nucleotide immediately 5' of TTAA is T and the nucleotide immediately 3' of TTAA is A, generating a TTTAAA integration site, or contains a spacer sequence such that the nucleotide immediately 5' of TTAA is C and the nucleotide immediately 3' of TTAA is A, generating a CTTAAA integration site. In addition, since thymidine is not immediately 5' of all LPA target sites, the TAL N-terminal domain was mutated so that it does not require any particular nucleotide 5' of the binding site. These mutations were introduced into the wild-type TAL sequence by substituting YH into the amino acid sequence QWS at positions 79-81 of SEQ ID NO: 31, thereby creating the NT-βN variant (SEQ ID NO: 34).

[0238] The TAL arrays were constructed using the design criteria described herein or to target these TAL binding sites as shown below.

[0239] TAL array pairs were designed targeting four specific 10 bp left-right pair sequences within a second set of four LPA target sites. For each target, multiple TAL array pairs were designed using either a 12 bp or 13 bp spacer.

[0240] Table 10 shows the left and right target sequences, along with the 5' nucleotides used to construct the TAL array targeting the LPA gene. [Table 10]

[0241] Using eight left-right pair combinations, the left-side LPA TAL arrays LPAL5.1, LPAL5.2, LPAL6.1, LPAL6.2, LPAL7.1, LPAL7.2, LPAL8.1, and LPAL8.2 (sequences 127, 129, 131, 133, 135, 137, 139, and 141, respectively) and the right-side LPA TAL arrays LPAR5.1, LPAR5.2, LPAR6.1, LPAR6.2, LPAR7.1, LPAR7.2, LPAR8.1, and LPAR8.2 (sequences 128, 130, 132, 134, 136, 138, 140, and 142, respectively) were designed and constructed as described in Example 2.

[0242] The TAL-PBx fusion construct was prepared as follows: 5' to 3' direction: CMV promoter, T7 promoter, Kozak sequence, 3xFlag tag (SEQ ID NO: 65), SV40 NLS (SEQ ID NO: 66), Delta 152 TAL N-terminal domain of the TAL NT-BN variant (SEQ ID NO: 34), two BsmBI type IIS restriction enzyme sites, +73 TAL C-terminal domain (SEQ ID NO: 79), GGGS linker, delta 1-85 PBx (containing an N-terminal 85 amino acid deletion and mutations at R372A, K375A, D450N in the Super piggyBac transposase codon sequence; SEQ ID NO: 9) and an expression plasmid containing the bGH polyadenylation sequence was synthesized.

[0243] Cloning of the left or right TAL array flanked by BsmBI into the BsmBI site of the expression plasmid results in an in-frame fusion of the TAL array and the PBx coding sequence via a linker sequence that generates the full-length TAL-PBx construct. All coding sequences used were codon-optimized for human expression using the GeneArt algorithm (Thermo Fisher).

[0244] Sixteen TAL arrays flanked by BsmBI ends were cloned into the BsmBI restriction site of the above expression plasmid to generate sixteen TAL-PBx constructs: LPAL5.1, LPAL5.2, LPAL6.1, LPAL6.2, LPAL7.1, LPAL7.2, LPAL8.1 and LPAL8.2 left TAL-PBxs (SEQ ID NOs: 154, 156, 158, 160, 162, 164, 166 and 168 respectively) and LPAR5.1, LPAR5.2, LPAR6.1, LPAR6.2, LPAR7.1, LPAR7.2, LPAR8. I and LPAR8.2 right TAL-PBxs (SEQ ID NOs: 155, 157, 159, 161, 163, 165, 167 and 169 respectively) were prepared as described in Example 3.

[0245] Furthermore, the TAL arrays LPAL1 (SEQ ID NO: 116) and LPAR1 (SEQ ID NO: 117) described in Example 2 were cloned into the expression plasmids described above in this Example 6 to generate the TAL-PBx constructs LPAL1 v2 (SEQ ID NO: 170) and LPAR2 v2 (SEQ ID NO: 171).

[0246] A. Episomal target site-specific translocation The activities of the new variant TAL-PBx fusions were determined using the respective episomal split GFP splicing reporters. Briefly, each reporter plasmid and donor plasmid were co-transfected into HEK293T cells along with the corresponding TAL-PBx expression plasmid. Approximately 120,000 HEK293T cells were seeded into 24-well plates in 500 μl of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture containing 50 ng of the TAL-PBx expression vector, 225 ng of the reporter plasmid, 225 ng of the donor plasmid and 1 μl of JetPrime transfection reagent in a total volume of 50 μl of JetPrime buffer was assembled. This mixture was added to the HEK293T cells and they were incubated at 37 °C, 5% CO2 for 4 days, and the cells were split 1:6 on day 1. The percentage of GFP-positive cells was determined for each sample. The results are shown in Table 11.

Table 11

[0247] As seen in Table 11, the TAL-ssSPBs targeting targets 1, 6 and 8 resulted in the highest translocations. Furthermore, the TAL-ssSPBs utilizing the 13bp spacer resulted in higher editing than those utilizing the 12bp spacer.

[0248] B. Genomic target site-specific translocation After confirming that the newly designed LPA TALs were functional and recognized their target sequences, the TAL-PBx constructs were used to edit endogenous genomic LPA targets in the immortalized hepatocyte line Huh7. Briefly, 100,000 cells were seeded in 24-well plates in RPMI medium + 10% FBS the day before transfection. The following day, 0.5ug or 1ug of mRNA encoding LPA Target 1 TAL-ssSPB pairs (SEQ ID NOs. 143 and 144) was mixed with 0.5μl or 1μl of MessengerMax reagent, respectively, to generate ssSPB-mRNA-lipid complexes. Simultaneously, 450ng of transposon donor vector (SEQ ID NOs. 80) was mixed with 0.5μl of P3000 reagent and 1μl of lipofectamine 3000 to generate DNA-lipid complexes. 50 μl of ssSPB mRNA lipid complex and 50 μl of DNA lipid complex were delivered to cells and incubated at 37°C.

[0249] To evaluate site-specific integration of transposon donors into the LPA locus, genomic DNA was extracted from transfected cells two days after transfection and analyzed by digital droplet PCR (ddPCR) using a probe-based detection scheme. One primer that binds within the transposon was paired with a primer that binds to LPA genomic DNA near the TTAA integration site. Therefore, amplicons should only be generated after site-specific transposition to the LPA locus. Since integration is not directional, two assays were designed for each LPA target to detect forward and reverse transposon integration. As a negative control, genomic DNA was extracted from untransfected cells and used as a template for the ddPCR reaction to demonstrate the specificity of the primer / probe set. The results are shown in Table 12. [Table 12]

[0250] As shown in Table 12, amplicons corresponding to forward and / or reverse transposon integration were detected from genomic DNA isolated using cells transfected with the LPA TAL-PBx construct along with transposons, providing direct evidence of genomic integration at the LPA locus.

Claims

1. A fusion protein comprising a DNA targeting domain and a transposase domain containing the sequence shown in Sequence ID No. 4, wherein the DNA targeting domain binds to a nucleic acid sequence encoding an LPA repeat element.

2. The fusion protein according to claim 1, wherein the DNA targeting domain comprises one, two, or three zinc finger motifs.

3. The fusion protein according to claim 1, wherein the DNA targeting domain comprises one or more TAL domains.

4. The method according to claim 3, wherein the TAL domain includes the sequence shown in any one of sequence numbers 35 to 38.

5. The fusion protein according to any one of claims 1 to 4, wherein the DNA targeting domain binds to a nucleic acid sequence encoding a kringle domain repeat element in the LPA gene or to an intron adjacent to a sequence encoding a kringle domain repeat element.

6. The fusion protein according to any one of claims 1 to 5, wherein the transposase domain and the DNA targeting domain are linked by a linker.

7. The fusion protein according to claim 6, wherein the linker contains the sequence GGGGS (Sequence ID 181).

8. The fusion protein according to any one of claims 1 to 7, wherein the DNA targeting domain is inserted at the N-terminus of the transposase domain at a position after the 82nd amino acid and before the 105th amino acid of SEQ ID NO:

4.

9. The fusion protein according to any one of claims 1 to 7, wherein the DNA targeting domain substitutes one or more amino acids in the transposase domain between (including) the 83rd and 105th amino acids of SEQ ID NO:

4.

10. The fusion protein according to any one of claims 1 to 9, wherein the transposase domain contains an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103.

11. The fusion protein according to any one of claims 1 to 10, wherein the transposase domain comprises the sequence shown in any one of sequence numbers 7 to 27.

12. The fusion protein according to any one of claims 1 to 11, wherein the transposase domain comprises (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R; or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D.

13. A polynucleotide comprising a nucleic acid sequence encoding a fusion protein according to any one of claims 1 to 12.

14. A vector comprising a polynucleotide as described in claim 13.

15. A method for incorporating a transgene into a genomic target site of a cell, wherein the method comprises introducing a fusion protein and a transposon according to any one of claims 1 to 12 into the cell, wherein the transposon comprises a 5'-ITR, the transgene, and a 3'-ITR in the order of 5' to 3'.

16. The method according to claim 15, wherein the transposon further comprises an exogenous promoter between the 5' ITR and the transgene.

17. The method according to claim 15 or 16, wherein the introduced gene encodes a detectable marker.

18. The method according to claim 17, wherein the detectable marker is GFP.

19. The method according to claim 15 or 16, wherein the introduced gene is (a) not expressed by the cell prior to the introduction of the fusion protein and the transposon, or (b) a gene that shows reduced expression, deficiency and / or alteration of expression by the cell prior to the introduction of the fusion protein and the transposon.

20. The method according to any one of claims 15 to 19, wherein the genomic target site is located on the LPA gene.

21. The method according to any one of claims 15 to 19, wherein the genome target site is located within a repeat element.

22. The method according to claim 21, wherein the repeat element is an LPA repeat element.

23. The method according to any one of claims 15 to 19, wherein the genome target site is located in an intron of a gene.

24. The method according to claim 23, wherein the genome target site is located in an intron of the LPA gene.

25. The method according to any one of claims 15 to 24, wherein the cells are in vivo.

26. A method for modifying the genome of a cell, the method comprising providing the cell with a fusion protein according to any one of claims 1 to 12, wherein the cell provides a fusion protein comprising, in the order 5' to 3', a sequence of a target site for the DNA targeting domain, a first spacer, a TTAA target integration site for the SPB, a second spacer, and a modified binding site comprising the reverse complement of the sequence of the target site for the DNA targeting domain.

27. The method according to claim 26, wherein the target integration site includes the sequence TTAA.

28. The method according to claim 26, wherein the target integration site includes a nucleic acid sequence shown in any one of sequence numbers 81 to 88.

29. An implantation cassette for site-directed transfer of nucleic acid to a cell's genome, comprising a nucleic acid comprising or consisting of a central transposon ITR integration site TTAA sequence sandwiched between an upstream TAL array target sequence and a downstream TAL array target sequence, wherein each of the upstream TAL array target sequence and the downstream TAL array target sequence is 12 or 13 base pairs away from the TTAA sequence.

30. The embedded cassette according to claim 29, wherein the embedded portion includes the array TTAA.

31. The embedded cassette according to claim 29, wherein the embedded portion includes a nucleic acid sequence shown in any one of sequence numbers 81 to 88.

32. The embedded cassette according to any one of claims 29 to 31, wherein the upstream TAL array target site sequence and the downstream TAL array target site sequence are the same.

33. The embedded cassette according to any one of claims 29 to 31, wherein the upstream TAL array target site sequence and the downstream TAL array target site sequence are each different.

34. The embedded cassette according to any one of claims 29 to 33, wherein each of the upstream TAL array target region and the downstream TAL array target a 7 to 30 bp sequence of LPA repeat elements.

35. A cell comprising an embedded cassette according to any one of claims 29 to 34, which is stably incorporated into the cell's genome.

36. A method for site-specific transfer of DNA molecules to the genome of a cell, wherein the cell according to claim 35, a) A nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed within the cell, and b) A DNA molecule containing a transposon, wherein the expressed fusion protein incorporates the transposon into the TTAA integration site of the stably incorporated integration cassette by site-directed transposition, and A method that includes introducing [something].

37. A method for generating manipulated cells by site-directed metastasis, The cells described in claim 31, a) A nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed within the cell, and b) A DNA molecule containing a transposon, wherein the expressed fusion protein incorporates the transposon into the TTAA integration site of the stably incorporated integration cassette by site-directed transposition, thereby generating the manipulated cell. A method that includes introducing [something].

38. The method according to claim 36 or 37, wherein the incorporated portion includes the sequence TTAA.

39. The method according to claim 36 or 37, wherein the integration site includes a nucleic acid sequence shown in any one of sequence numbers 81 to 88.