Transposases and their uses
Patent Information
- Application Number
- JP2024519988
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-29
- Filing Date
- 2022-10-04
- Publication Date
- 2025-10-10
AI Technical Summary
There is a need for site-specific transposases for gene editing that can efficiently introduce non-endogenous DNA sequences into genomic DNA.
The development of fusion proteins comprising transposase domains with N-terminal deletions and linkers, which form essential heterodimers, enhancing site-specific transposition and integration of DNA into target sites.
The fusion proteins demonstrate improved site-specific transposition and integration efficiency, reducing off-target effects and increasing the precision of genetic modifications.
Smart Images

Figure 00000114_0000 
Figure 00000114_0001 
Figure 00000114_0002
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 252,028, filed October 4, 2021, No. 63,312,928, filed February 23, 2022, and No. 63 / 369,863, filed July 29, 2022, each of which is incorporated by reference in its entirety herein.
[0002] Reference to an Electronically Submitted Sequence Listing This application contains a Sequence Listing that was submitted via the Patent Center in XML format and is incorporated herein by reference in its entirety. A copy of said XML, created on October 3, 2022, is named "POTH-069-001WO-SeqList_ST26" and is 787,153 bytes in size.
[0003] The present disclosure generally relates to a transposase domain, in particular a transposase domain that includes an N-terminal deletion, as well as a transposase domain that forms an essential heterodimer, and a fusion protein that includes a transposition matrix domain and a DNA targeting domain. Also provided is a method of using the fusion protein for site-specific transposition. [Background technology]
[0004] Transposase can be used to introduce non-endogenous DNA sequences into genomic DNA, and transposase has many advantages over other gene editing methods.However, the demand for site-specific transposase for use in gene editing remains unmet. Summary of the Invention
[0005] In one aspect, provided herein is a fusion protein comprising a first transposase domain, a linker, and a second transposase domain, wherein (a) the first and second transposase domains are identical, or (b) the first and second transposase domains are identical except that the second transposase domain comprises an N-terminal deletion. In some embodiments, the first transposase domain is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a hyperactive piggyBac transposase domain. In some embodiments, the first transposase domain is a Super PiggyBac (SPB) transposase domain. In some embodiments, the second transposase domain is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a hyperactive piggyBac transposase domain. In some embodiments, the second transposase domain is a Super PiggyBac transposase domain. In some embodiments, the first transposase domain and the second transposase domain are piggyBac transposase domains. In some embodiments, the first piggyBac transposase domain and the second piggyBac transposase domain are hyperactive piggyBac transposase domains. In some embodiments, the first transposase domain is an SPB transposase domain. In some embodiments, the first transposase domain and the second transposase domain are SPB transposase domains.
[0006] In some embodiments, the N-terminal deletion of the second transposase domain comprises amino acids 1-20. In some embodiments, the amino-terminal deletion of the second transposase domain comprises amino acids 1-40. In some embodiments, the amino-terminal deletion of the second transposase domain comprises amino acids 1-60. In some embodiments, the amino-terminal deletion of the second transposase domain comprises amino acids 1-80. In some embodiments, the amino-terminal deletion of the second transposase domain comprises amino acids 1-100. In some embodiments, the amino terminus of the second transposase domain comprises amino acids 1-115. In some embodiments, the first transposase domain further comprises an in-frame nuclear localization signal (NLS).
[0007] In some embodiments, the linker is juxtaposed between the C-terminus of the first transposase domain and the N-terminus of the second transposase domain. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 16.
[0008] In some embodiments, the fusion protein comprises an amino acid sequence of any one of SEQ ID NOs: 8-14. In some embodiments, the fusion protein further comprises a mutation in one or both of the transposase domains. In some embodiments, the mutation is selected from the group consisting of (a) M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R, or (b) L204D, L204E, K500D, K500E, R504E, and R504D. In some embodiments, the fusion protein comprises two or three of the mutations selected from the group consisting of M185R, D198K, and D201R in one or both of the transposase domains. In some embodiments, the fusion protein comprises two or three of the mutations selected from the group consisting of L204E, K500D, and R504D in one or both of the transposase domains.
[0009] In another aspect, provided herein is a transposase domain comprising a sequence selected from any one of SEQ ID NOs: 31-53. In some embodiments, the transposase domain comprises a sequence selected from any one of SEQ ID NOs: 31-53, and further comprises one or more conserved amino acid sequences.
[0010] In another aspect, provided herein is a fusion protein comprising a first transposase domain, a linker, and a second transposase domain, wherein the first transposase domain and / or the second transposase domain comprise an identical sequence selected from any one of SEQ ID NOs: 31 to 43. In another aspect, provided herein is a fusion protein comprising a first transposase domain, a linker, and a second transposase domain, wherein the first transposase domain and / or the second transposase domain comprise an identical sequence selected from any one of SEQ ID NOs: 44 to 53.
[0011] In some embodiments, the fusion protein provided herein further comprises a DNA targeting domain. In some embodiments, the DNA targeting domain is attached to the N-terminus of the fusion protein. In some embodiments, the DNA targeting domain is attached to the C-terminus of the fusion protein. In some embodiments, the DNA targeting domain is selected from the group consisting of CRISPR, zinc finger, TALE, and transcription factor.
[0012] In another aspect, provided herein is a transposase domain that comprises an N-terminal deletion compared to the sequence set forth in SEQ ID NO:1 or SEQ ID NO:55 (numbering beginning at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NO:55). In some embodiments, the transposase domain is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a hyperactive piggyBac transposase domain. In some embodiments, the transposase domain is an SPB transposase domain. In some embodiments, the N-terminal deletion comprises amino acids 1-20. In some embodiments, the N-terminal deletion comprises amino acids 1-40. In some embodiments, the N-terminal deletion comprises amino acids 1-60. In some embodiments, the N-terminal deletion comprises amino acids 1-80. In some embodiments, the N-terminal deletion comprises amino acids 1-100. In some embodiments, the N-terminal deletion comprises amino acids 1-115.
[0013] In some embodiments, the transposase domain further comprises an in-frame nuclear localization signal (NLS). In some embodiments, the in-frame NLS is fused to the amino terminus of the transposase domain. In some embodiments, the transposase domain comprises the amino acid sequence of any one of SEQ ID NOs: 2-7.
[0014] In another aspect, the present invention provides a nucleic acid molecule comprising a nucleotide sequence encoding the fusion protein described herein.In some embodiments, the nucleic acid molecule further comprises a promoter operably linked to the nucleotide sequence encoding the fusion protein.In some embodiments, the nucleic acid molecule further comprises a polyA sequence located downstream of the nucleotide sequence encoding the second transposase domain.
[0015] In another aspect, provided herein is a nucleic acid molecule comprising a nucleotide sequence encoding a transposase domain as described herein.In some embodiments, the nucleic acid molecule further comprises a promoter operably linked to the nucleotide sequence encoding the transposase domain.In some embodiments, the nucleic acid molecule further comprises a polyA sequence located downstream of the nucleotide sequence encoding the transposase domain.
[0016] In another aspect, provided herein is a cell comprising the nucleic acid molecule described herein. In some embodiments, the cell is derived from a patient. In some embodiments, the cell further comprises a chimeric antigen receptor (CAR). In some embodiments, the cell is an immune cell. In some embodiments, the cell is a T cell.
[0017] In another aspect, provided herein is a method of treating a disease or disorder in a patient, the method comprising administering to the patient a cell described herein. In some embodiments, the cell is autologous. In some embodiments, the cell is allogeneic. In some embodiments, the disease or disorder is cancer.
[0018] In another aspect, a first fusion protein comprising: (a) a first transposase domain, a linker, a second transposase domain, and a first DNA targeting domain, wherein (i) the first and second transposase domains are identical, or (ii) the first and second transposase domains are identical except that the second transposase domain comprises an N-terminal deletion; and a second fusion protein comprising: (a) a first transposase domain, a linker, a second transposase domain, and a second DNA targeting domain, wherein (i) the first and second transposase domains are identical, except that the second transposase domain comprises an N-terminal deletion. Provided herein is a complex comprising (i) a second fusion protein, wherein the first and second transposase domains are identical, or (ii) a second fusion protein, wherein the first and second transposase domains are identical except that the second transposase domain comprises an N-terminal deletion, and wherein the first DNA targeting domain and the second DNA targeting domain are different, and the transposase domain of the first fusion protein and the transposition matrix domain of the second fusion protein have opposite charges that enable the two fusion proteins to form a complex.
[0019] In some embodiments, the transposase domain of the first fusion protein comprises at least one mutation and the transposition matrix domain of the second fusion protein comprises at least one mutation that confers an opposite charge. In some embodiments, the second transposase domain of the first fusion protein and the second transposase domain of the second fusion protein are SPB transposase domains. In some embodiments, the at least one mutation is selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, the at least one mutation is selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D.
[0020] In some embodiments, the N-terminal deletion includes amino acids 1-20. In some embodiments, the N-terminal deletion includes amino acids 1-40. In some embodiments, the N-terminal deletion includes amino acids 1-60. In some embodiments, the N-terminal deletion includes amino acids 1-80. In some embodiments, the N-terminal deletion includes amino acids 1-100. In some embodiments, the N-terminal deletion includes amino acids 1-115.
[0021] In some embodiments, the first DNA targeting domain is attached to the C-terminus of the first fusion protein, and the second DNA targeting domain is attached to the C-terminus of the second fusion protein. In some embodiments, the first DNA targeting domain is attached to the N-terminus of the first fusion protein, and the second DNA targeting domain is attached to the N-terminus of the second fusion protein. In some embodiments, the DNA targeting domain is selected from the group consisting of CRISPR, zinc finger, TALE, and transcription factor.
[0022] In another aspect, provided herein is a complex comprising: (a) a first fusion protein comprising a first transposase domain, a linker, a second transposase domain, and a first DNA targeting domain, wherein the first and / or the second transposase domain of the first fusion protein comprise an identical amino acid sequence set forth in any one of SEQ ID NOs: 31-43; and (b) a second fusion protein comprising a first transposase domain, a linker, a second transposase domain, and a second DNA targeting domain, wherein the first and / or the second transposase domain of the second fusion protein comprise an identical amino acid sequence set forth in any one of SEQ ID NOs: 44-53. In some embodiments, the DNA targeting domain is selected from the group consisting of CRISPR, zinc finger, TALE, and transcription factor.
[0023] In another aspect, provided herein is a fusion protein comprising, from N-terminus to C-terminus, a nuclear localization signal (NLS), a DNA targeting domain, and a first transposase domain comprising the sequence of SEQ ID NO: 65 or 55. In some embodiments, the fusion protein further comprises a protein stabilization domain (PSD). In some embodiments, the PSD comprises SEQ ID NO: 68. In some embodiments, the DNA targeting domain comprises three zinc finger motifs. In some embodiments, the DNA targeting domain comprises the sequence of SEQ ID NO: 57. In some embodiments, the DNA targeting domain comprises one or more TAL domains. In some embodiments, the DNA targeting domain binds to a nucleic acid sequence encoding GFP, ZFM268, phenylalanine hydroxylase (PAH), beta-2-microglobulin (B2M), or a LINE1 repeat element.
[0024] In some embodiments, the transposase domain comprises at least one mutation selected from the group consisting of: (a) M92R, M92K, D104K, D104R, D105K, D105R, D108K, and D108R, or (b) L111D, L111E, K407D, K407E, R411E, and R411D. In some embodiments, the fusion protein comprises the sequence of SEQ ID NO: 67 or 69.
[0025] In some embodiments, the fusion protein further comprises a second transposase domain. In some embodiments, the second transposase domain comprises the sequence of SEQ ID NO: 55 or 56. In some embodiments, the second transposase domain is connected to the C-terminus of the first transposase domain via a linker.
[0026] In another aspect, provided herein is a fusion protein comprising: (a) a TAL array; and (b) a Super piggyBac transposase ("SPB") comprising an N-terminal deletion, wherein the TAL array and a polynucleotide encoding the N-terminal deleted SPB are fused in-frame to encode a TAL array-N-terminal deleted SPB fusion protein. In some embodiments, the fusion protein further comprises an in-frame GS or GGGGS linker located between the TAL array and the N-terminal deleted SPB. In some embodiments, the SPB comprises an N-terminal deletion comprising a deletion of amino acids 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103. In some embodiments, the fusion protein further comprises one or more mutations in SPB at amino acids R372A, K375A, or D450N. In some embodiments, the SPB comprises a sequence set forth in SEQ ID NOs: 81-106. In some embodiments, the SPB is an integrase-deficient SPB (PBx).
[0027] In another aspect, provided herein is a complex comprising: (a) a first fusion protein comprising, from N-terminus to C-terminus, a first NLS, a first DNA targeting domain, a first transposase domain comprising the sequence of SEQ ID NO: 65 or 66, a linker, and a second transposase domain; and (b) a second fusion protein comprising, from N-terminus to C-terminus, a second NLS, a second DNA targeting domain, a third transposase domain comprising the sequence of SEQ ID NO: 65 or 66, a linker, and a fourth transposase domain, wherein the transposase domain of the first fusion protein and the transposition matrix domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex.
[0028] In some embodiments, the second and / or fourth transposase domain is a SPB domain. In some embodiments, the second and / or fourth transposase domain is a PBx transposase domain. In some embodiments, the second and / or fourth transposase domain comprises the sequence of SEQ ID NO: 55. In some embodiments, the second and / or fourth transposase domain comprises the sequence of SEQ ID NO: 56.
[0029] In some embodiments, the first transposase domain comprises at least one mutation selected from the group consisting of M92R, M92K, D104K, D104R, D105K, D105R, D108K, and D108R. In some embodiments, the second transposase domain comprises at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, the third transposase domain comprises at least one mutation selected from the group consisting of L111D, L111E, K407D, K407E, R411E, and R411D. In some embodiments, the fourth transposase domain comprises at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, and R504E, R504D.
[0030] In some embodiments, the first fusion protein further comprises a first PSD between the first (NLS) and the first DNA-targeting domain, and / or the second fusion protein further comprises a second PSD between the second NLS and the second DNA-targeting domain. In some embodiments, the first and / or second PSD comprises the sequence of SEQ ID NO:68.
[0031] In some embodiments, the first and / or second DNA targeting domain comprises three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domain comprises the sequence of SEQ ID NO:57.
[0032] In another aspect, provided herein is a polynucleotide comprising a nucleic acid sequence encoding a fusion protein provided herein. In another aspect, provided herein is a vector comprising a polynucleotide provided herein.
[0033] In another aspect, provided herein is a cell comprising the polynucleotide or vector provided herein. In some embodiments, the cell further comprises a chimeric antigen receptor (CAR). In some embodiments, the cell is an immune cell.
[0034] In another aspect, provided herein is a pharmaceutical composition comprising a cell provided herein and a pharma- ceutically acceptable carrier.
[0035] In another aspect, provided herein is a method of treating a disease or disorder in a patient, comprising administering to the patient a cell or pharmaceutical composition provided herein. In some embodiments, the cell is allogeneic. In some embodiments, the disease or disorder is cancer.
[0036] In another aspect, provided herein is a method of modifying the genome of a cell, the method comprising providing to the cell a fusion protein comprising, from N-terminus to C-terminus, an NLS, a PSD, a DNA targeting domain, and a transposase domain comprising a sequence of SEQ ID NO: 65 or 66, the cell comprising a modified binding site comprising, in 5' to 3' order, a reverse sequence of the target site for the DNA targeting domain, a first spacer, a TTAA target integration site for the SPB, a second spacer, and a complement of the sequence of the target site for the DNA targeting domain. In some embodiments, the DNA targeting domain comprises a sequence of SEQ ID NO: 57. In some embodiments, the fusion protein comprises a sequence of SEQ ID NO: 67. In some embodiments, the first spacer and the second spacer are each 7 bp in length. In some embodiments, the modified binding site comprises a sequence of any one of SEQ ID NOs: 61-64.
[0037] In another aspect, provided herein is an integration cassette for site-specific transfer of a DNA molecule into the genome of a cell. In one embodiment, an integration cassette for site-specific transfer of a nucleic acid into the genome of a cell comprises a nucleic acid comprising or consisting of a central transposon ITR integration site TTAA sequence flanked by at least one upstream zinc finger motif DNA binding domain binding site ("ZFM-DBD") and at least one downstream ZFM-DBD, each of which is separated from the TTAA sequence by 7 base pairs. In one embodiment, each of the at least one upstream and downstream ZFM-DBD sites is a ZFM268 binding site. In one embodiment, each of the ZFM268 binding sites comprises SEQ ID NO:60. In one embodiment, the integration cassette comprises or consists of SEQ ID NO:62.
[0038] In another aspect, provided herein is an integration cassette for site-specific transposition of a nucleic acid into the genome of a cell, comprising or consisting of a nucleic acid comprising or consisting of a central transposon ITR integration site TTAA sequence flanked by an upstream TAL array target sequence and a downstream TAL array target sequence, each of which is separated from the TTAA sequence by 12 to 14 base pairs.
[0039] In another aspect, provided herein is an integration cassette for site-specific transposition of a nucleic acid into the genome of a cell, comprising a nucleic acid comprising a central transposon ITR integration site TTTAAA sequence flanked by an upstream TAL array target sequence and a downstream TAL array target sequence, each of the upstream and downstream TAL array target sequences being separated by 12 base pairs from the TTTAAA sequence.
[0040] In one embodiment, each of the at least one upstream and downstream TAL array target site sequences are identical. In one embodiment, each of the at least one upstream and downstream TAL array target site sequences are different. In one embodiment, each of the at least one upstream and downstream TAL array target sites targets a 7-30 bp (e.g., 10 bp) sequence of the beta-2-microglobulin gene ("B2M"), the phenylalanine hydroxylase gene ("PAH"), or a LINE1 repeat element. In one embodiment, at least one upstream TAL array target sequence and at least one downstream TAL array target sequence bind to a nucleic acid comprising the sequence GCGTGGGCG. In one embodiment, the integration cassette comprises SEQ ID NO:62.
[0041] In certain embodiments, cells are provided that contain an integration cassette for site-specific transposition of a DNA molecule provided herein stably integrated into the genome of the cell.
[0042] In a particular aspect, a method for site-specific transposition of a DNA molecule into the genome of a cell containing a stably integrated integration cassette is provided, comprising integrating into said cell a) a nucleic acid encoding a fusion protein comprising a DNA binding domain and a transposase, said fusion protein being expressed in said cell, and b) a DNA molecule comprising a transposon, wherein said expressed fusion protein integrates said transposon into the TTAA sequence of said stably integrated integration cassette by site-specific transposition.
[0043] In a particular aspect, a method for generating a modified cell by site-specific transposition is provided, comprising introducing into a cell containing a stably integrated integration cassette: a) a nucleic acid encoding a fusion protein comprising a DNA binding domain and a transposase, wherein the fusion protein is expressed in the cell; and b) a DNA molecule comprising a transposon, wherein the expressed fusion protein integrates the transposon by site-specific transposition into a TTAA sequence of the stably integrated integration cassette, thereby generating the recombinant cell.
[0044] In another aspect, provided herein is a fusion protein comprising, from N-terminus to C-terminus, a DNA targeting domain and a first transposase domain comprising the sequence set forth in SEQ ID NO:544, wherein the first transposase domain comprises a deletion of most of the N-terminal amino acids from 83 to 103 of SEQ ID NO:544.
[0045] In some embodiments, the DNA targeting domain comprises three zinc finger motifs. In some embodiments, the DNA targeting domain comprises one or more TAL domains. In some embodiments, the TAL domain comprises a sequence set forth in any one of SEQ ID NOs: 107-110. In some embodiments, the DNA targeting domain binds to a nucleic acid sequence encoding GFP zinc finger 268 (ZFM268), phenylalanine hydroxylase (PAH), beta-2-microglobulin (B2M), or a LINE1 repeat element.
[0046] In some embodiments, the first transposase domain and the DNA targeting domain are connected by a linker. In some embodiments, the linker comprises the sequence GGGGS.
[0047] In some embodiments, the first transposase domain comprises an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103. In some embodiments, the transposase domain comprises a sequence set forth in any one of SEQ ID NOs:86-106.
[0048] In some embodiments, the first transposase domain comprises (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R, or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D.
[0049] In some embodiments, the fusion protein further comprises a second transposase domain C-terminal to the first transposase domain, wherein the second transposase domain comprises a sequence set forth in SEQ ID NO: 544. In some embodiments, the second transposase domain comprises a deletion of N-terminal amino acids 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103 of SEQ ID NO:544. In some embodiments, the second transposase domain comprises (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R, or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D.
[0050] In another aspect, provided herein is a polynucleotide comprising a nucleic acid sequence encoding a fusion protein provided herein. Also provided herein is a vector comprising a polynucleotide provided herein.
[0051] In another aspect, provided herein is a method of integrating a transgene into a genomic target site of a cell, the method comprising introducing into the cell a fusion protein provided herein and a transposon, the transposon comprising, in 5' to 3' order, a 5' ITR, a transgene, and a 3' ITR. In some embodiments, the transposon further comprises an exogenous promoter between the 5' ITR and the transgene. In some embodiments, the transgene encodes a detectable marker. In some embodiments, the detectable marker is GFP. In some embodiments, the transgene is a gene that is not expressed by the cell prior to introduction of the fusion protein and the transposon.
[0052] In some embodiments, the genomic target site is located on chromosome 17 or 21. In some embodiments, the genomic target site is located within the B2M gene. In some embodiments, the genomic target site is located within a repetitive element. In some embodiments, the repetitive element is a LINE element. In some embodiments, the genomic target site is located within an intron of a gene. In some embodiments, the genomic target site is located within an intron of a PAH gene. In some embodiments, the cell is in vivo.
[0053] In another aspect, provided herein is a method of modifying the genome of a cell, the method comprising providing a fusion protein provided herein to the cell, the cell comprising a modified binding site comprising, in 5' to 3' order, a reverse sequence of the target site for the DNA targeting domain, a first spacer, a TTAA target incorporation site for the SPB, a second spacer, and the complement of the sequence of the target site for the DNA targeting domain.
[0054] In another aspect, provided herein is an integration cassette for site-specific transposition of a nucleic acid into the genome of a cell comprising a nucleic acid comprising or consisting of a central transposon ITR integration site TTAA sequence flanked by at least one upstream zinc finger motif DNA-binding domain binding site ("ZFM-DBD") and at least one downstream ZFM-DBD, wherein each of the upstream and downstream ZFM-DBDs are separated from the TTAA sequence by 7 base pairs.
[0055] In another aspect, provided herein is an integration cassette for site-specific transposition of a nucleic acid into the genome of a cell, comprising or consisting of a nucleic acid comprising or consisting of a central transposon ITR integration site TTAA sequence flanked by an upstream TAL array target sequence and a downstream TAL array target sequence, each of which is separated from the TTAA sequence by 12 to 14 base pairs.
[0056] In another aspect, provided herein is an integration cassette for site-specific transposition of a nucleic acid into the genome of a cell, comprising a nucleic acid comprising a central transposon ITR integration site TTTAAA sequence flanked by an upstream TAL array target sequence and a downstream TAL array target sequence, each of the upstream and downstream TAL array target sequences being separated by 12 base pairs from the TTTAAA sequence. In some embodiments, at least one upstream and downstream TAL array target site sequence is identical. In some embodiments, each of the at least one upstream and downstream TAL array target site sequence is different. In some embodiments, each of the at least one upstream and downstream TAL array target site targets a 10 bp sequence of the beta-2-microglobulin gene ("B2M"), the phenylalanine hydroxylase gene ("PAH"), or a LINE1 repeat element. In some embodiments, at least one upstream TAL array target sequence and at least one downstream TAL array target sequence bind to a nucleic acid comprising the sequence GCGTGGGCG.
[0057] In another aspect, provided herein is a cell comprising an integration cassette provided herein stably integrated into the genome of the cell. In another aspect, provided herein is a method for site-specific transposition of a DNA molecule into the genome of a cell, comprising introducing into the cell provided herein a nucleic acid encoding a fusion protein comprising a DNA binding domain and a transposase, wherein the fusion protein is expressed in the cell and the DNA molecule comprises a transposon, and wherein the expressed fusion protein integrates the transposon into the TTAA sequence of the stably integrated integration cassette by site-specific transposition.
[0058] In another aspect, provided herein is a method for generating a modified cell by site-specific transposition, comprising introducing into a cell provided herein a nucleic acid encoding a fusion protein comprising a DNA binding domain and a transposase, wherein the fusion protein is expressed in the cell and a DNA molecule comprising a transposon, wherein the expressed fusion protein integrates the transposon by site-specific transposition into a TTAA sequence of the stably integrated integration cassette, thereby generating the recombinant cell. [Brief description of the drawings]
[0059] [Figure 1] A shows a schematic diagram of an SPB construct containing an N-terminal deletion as described herein, and B shows a schematic diagram of an SPB construct with an inserted DNA-binding domain. [Figure 2A] 1 shows the introduction of a DNA-binding domain into a transposase using an essential heterodimer. [Figure 2B] 1 shows the introduction of a DNA-binding domain into a transposase using an essential heterodimer. [Figure 2C] 1 shows the introduction of a DNA-binding domain into a transposase using an essential heterodimer. [Figure 2D]1 shows the introduction of a DNA-binding domain into a transposase using an essential heterodimer. [Diagram 3] 1 shows the results of a truncation reporter assay demonstrating activity of wild-type transposase domains and transposase domains containing N-terminal deletions, such as "-20aa," indicating an N-terminal deletion of 20, 40, 60, 80, or 115 amino acids. [Figure 4A] 1 shows the results of cleavage and integration reporter assays, respectively, demonstrating the cleavage or integration activity of a wild-type SPB domain and a fusion protein ("tdSPB") that contains either two wild-type SPB transposase domains or one wild-type SPB transposase domain and one transposase domain that contains an N-terminal deletion, such as "-20aa" indicating an N-terminal deletion of 20, 40, 60, 80, or 115 amino acids within the second transposase domain. [Figure 4B] 1 shows the results of cleavage and integration reporter assays, respectively, demonstrating the cleavage or integration activity of a wild-type SPB domain and a fusion protein ("tdSPB") that contains either two wild-type SPB transposase domains or one wild-type SPB transposase domain and one transposase domain that contains an N-terminal deletion, such as "-20aa" indicating an N-terminal deletion of 20, 40, 60, 80, or 115 amino acids within the second transposase domain. [Figure 5A] Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 5B]Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 5C] Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 5D] Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 5E] Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 5F] Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 5G]Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 5H] Figure 1 is a series of graphs showing the results of cleavage and integration activity for various SPB transposase homodimers and heterodimers. K562 cells were nucleofected with a dual luciferase reporter and SPB expression plasmid. One day after transfection, luciferase signal was measured as a proxy for cleavage or integration activity. [Figure 6A] Schematic diagram of the dual reporter plasmid design used to confirm cleavage and integration rates with each mutant transposon. With the H-2kk GFP transposon reporter (reporter 1), an increase in H2kk expression is observed if there is increased cleavage of the transposon. With reporter 2, an increase in GFP expression is observed if there is increased integration of the transposon. In an alternative design of reporter 2, an increase in firefly luciferase expression is observed if there is increased cleavage of the transposon, and an increase in Nanoluc is observed if there is increased integration of the transposon. [Figure 6B] Schematic diagram of the H-2kk GFP transposon reporter (reporter 1). Both the circular and linear maps show the structural features of the transposon. When there is increased excision of the transposon, increased H2kk expression is observed, and when there is increased integration of the transposon, increased GFP is observed. [Figure 6C] Schematic diagram of the firefly luciferase Nanoluc transposon reporter. Both circular and linear maps show the structural features of the transposon. When there is increased excision of the transposon, firefly luciferase expression is observed, and when there is increased integration of the transposon, an increase in Nanoluc is observed. [Figure 7] FIG. 1 is a schematic diagram showing a split GFP splice site specific reporter. [Figure 8] We show integration and cleavage activity by wild-type SPB, an SPB with an N-terminal deletion of 93 amino acids and a DNA-targeting domain containing three zinc finger motifs (ZFM-SPB), and an integration-deleted SPB (PBx) with an N-terminal deletion of 93 amino acids and a DNA-targeting domain containing three zinc finger motifs (ZFM-PBx) at modified target sites containing spacers of various lengths between the SPB target site and the ZFM target site. [Figure 9A] The off-target genomic integration activity, on-target episomal integration activity, and the ratio of on-target activity to off-target activity by SPB, ZFM-SPB, and ZFM-PBx, respectively, are shown. [Figure 9B] The off-target genomic integration activity, on-target episomal integration activity, and the ratio of on-target activity to off-target activity by SPB, ZFM-SPB, and ZFM-PBx, respectively, are shown. [Figure 9C] The off-target genomic integration activity, on-target episomal integration activity, and the ratio of on-target activity to off-target activity by SPB, ZFM-SPB, and ZFM-PBx, respectively, are shown. [Figure 10A] The cleavage and integration activities of ZFM-PBx and ZFM-PBx-NTD are shown. [Figure 10B] The cleavage and integration activities of ZFM-PBx and ZFM-PBx-NTD are shown. [Figure 10C] The cleavage and integration activities of ZFM-PBx and ZFM-PBx-NTD are shown. [Figure 11] A schematic diagram of the GFP cleavage-only reporter is shown. [Figure 12] Figure 1 shows the sequence specificity of GFP TALENs using a single stranded annealing (SSA) assay. L and R indicate the left and right TAL arrays, respectively. [Figure 13] 1 shows the sequence specificity of PAH TALENs using a single stranded annealing (SSA) assay. L and R indicate the left and right TAL arrays, respectively. [Figure 14] Shows sequence specificity of PAH TALENs using an episomal split GFP splice site specific reporter assay. [Figure 15] Shows sequence specificity of PAH TALENs including on-target and off-target array pairs using an episomal split GFP splice site specific reporter assay. [Figure 16] Figure 1 shows the percentage of site-specific transposition into genomic DNA at six TTAA target sites in the LINE1 repetitive element as detected by ddPCR. Transposon integration was measured against a reference gene and is reported as the percentage of site-specific transposition per haploid genome. [Figure 17] ddPCR data showing site-specific transposition into genomic DNA for four TTAA sites within the B2M gene. Droplets with high amplitude along the Y-axis contain edited genomic DNA templates. [Figure 18] Figure 1 shows the integration activity of various PBx-ZFN fusion constructs as measured by split GFP assay. [Figure 19] Figure 1 shows the integration activity of TAL-PBx fusion constructs with various truncations of the PBx N-terminal domain, as measured by split-GFP assay, using reporters in which the TAL binding site was separated from the TTAA integration site by a 11-bp, 12-bp, 13-bp, or 14-bp spacer. [Figure 20] Schematic diagram of various TAL-PBx fusion constructs. A series of TAL C-terminal domain truncations retaining 13, 23, 33, 43, 54, 64, or 73 amino acids were fused in combination with PBx N-terminally truncated by 85, 88, 93, 99, or 103 amino acids. [Figure 21]Figure 21 shows the integration activity of various TAL-PBx fusion constructs shown in Figure 20 as measured by split GFP assay. TAL-PBx fusions were tested using target sites in which the TAL binding site was separated from the TTAA integration site by a 11 bp, 12 bp, 13 bp, or 14 bp spacer. [Figure 22] Schematic diagram of the "all-in-one site-specific cleavage / integration episomal reporter". This episomal reporter system contains a plasmid containing a transposon donor along with all of the transposon integration sites on the same plasmid. The transposon contains a CMV promoter. The transposon in this plasmid is initiated by the EF1a promoter and is followed by a polyadenylation signal sequence, disrupting the GFP open reading frame. The vector also contains, in reverse orientation, a polyA and transcription pause site, a TTAA integration site flanked by the target sequence and spacer, followed by a PEST-destabilized mScarlet reporter and a polyadenylation signal sequence. This "all-in-one site-specific cleavage / integration episomal reporter" should not express GFP and should not or hardly express mScarlet when transfected into cells alone. Upon transposon cleavage catalyzed by SPB, PBx, or ssSPB, GFP should be expressed. Upon site-specific integration of the transposon-containing CMV promoter into the target site, expression is conferred upstream of mScarlet. [Figure 23] Cleavage and site-specific integration activity of various TAL-PBx constructs containing mutations at positions 372 or 375 are shown. [Figure 24] Using an episomal split GFP splice site-specific reporter assay, we demonstrate the sequence specificity of ZF-PBx, designed to recognize ZF268, chr17, and chr21, with on-target and off-target array pairs. [Figure 25A]The site-specific integration activity of ZF268-PBx and ZF268-tdPBx is shown at target sites having ZF268 binding sites on both sides of the TTAA or on one side of the TTAA, as measured using an episomal split GFP splicing site-specific reporter assay. [Figure 25B] 1 shows the cleavage and site-specific integration activities of PAH2 or PAH3 TAL-PBx and TAL-tdPBX tested as pairs or as individual left or right fusion proteins measured using an episomal split GFP splice site-specific reporter assay. [Figure 25C] 1 shows the cleavage and site-specific integration activities of PAH2 or PAH3 TAL-PBx and TAL-tdPBX tested as pairs or as individual left or right fusion proteins measured using an episomal split GFP splice site-specific reporter assay. [Figure 26A] 1 shows the site-specific integration activity of TAL-PBx at chr17 target sites cloned into an episomal split GFP splice site-specific reporter. [Figure 26B] 1 shows the site-specific integration activity of TAL-PBx at the chr17 target in genomic DNA measured by ddPCR. Droplets with high amplitude along the y-axis contain edited genomic DNA templates. Droplets with high amplitude along the x-axis contain genomic DNA reference gene templates in the bottom plot. [Figure 26C] 1 shows the site-specific integration activity of TAL-PBx at the chr17 target in genomic DNA measured by ddPCR. Droplets with high amplitude along the y-axis contain edited genomic DNA templates. Droplets with high amplitude along the x-axis contain genomic DNA reference gene templates in the bottom plot. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0060] Provided herein are fusion proteins that contain transposase domains, and in particular transposase domains that contain N-terminal deletions. The fusion proteins that contain the transposase domains may be further mutated to form the required heterodimer. Also provided are methods for producing the transposase domains and fusion proteins, cells modified with the fusion proteins provided herein, and methods of treatment using such cells.
[0061] The transposase domains provided herein can be, for example, a wild-type transposase domain or an integration-deficient (truncated only) transposase domain.
[0062] Also provided herein are fusion proteins comprising one or more transposase domains and a DNA targeting domain. In some embodiments, the fusion protein further comprises a protein stabilization domain.
[0063] Transposase domain and fusion proteins containing the transposase domain In one aspect, provided herein are transposase domains and fusion proteins comprising the same (e.g., comprising a first and a second transposase domain). In some embodiments, the transposase domain is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a hyperactive piggyBac transposase domain. In preferred embodiments, the transposase domain is a Super piggyBac™ transposase domain (SPB). Non-limiting examples of SPB transposases are described in detail in U.S. Pat. Nos. 6,218,182; 6,962,810; 8,399,643, and PCT International Publication WO 2010 / 099296.
[0064] In some embodiments, the transposase domain is a Super PiggyBac transposase (SPB) domain. An exemplary wild-type SPB sequence, including the nuclear localization sequence (NLS), is shown in SEQ ID NO:1, with the NLS in italics, the hyperactivity mutations in bold, and the cysteine-rich domain (CRD) underlined. Numbering of the SPB transposase domain sequence to describe deletions and mutations begins at residue 12 of SEQ ID NO:1. [Table 1]
[0065] An exemplary sequence of a wild-type SPB transposase lacking the NLS domain is set forth in SEQ ID NO: 55. Numbering of the sequence of the SPB transposase domain to describe deletions and mutations begins at residue 5 of SEQ ID NO:55.
[0066] The transposase domains used in the fusion proteins described herein can be isolated or derived from insects, vertebrates, crustaceans, or tunicates, as described in more detail in PCT Publication Nos. WO2019 / 173636 and PCT / US2019 / 049816. In a preferred embodiment, the SPB transposase domain is isolated or derived from the insects Trichoplusia ni (GenBank Accession No. AAA87375) or Bombyx mori (GenBank Accession No. BAD11135).
[0067] In some embodiments, the transposase domain is an integration deletion type. An integration deletion type transposase domain is a transposase that can cleave its corresponding transposon, but integrates the cleaved transposon at a lower frequency than the corresponding wild-type transposase. Examples of integration deletion type transposases are disclosed in U.S. Patent No. 6,218,185; U.S. Patent No. 6,962,810, U.S. Patent No. 8,399,643, and WO 2019 / 17363. A list of integration deletion type amino acid substitutions is disclosed in U.S. Patent No. 10,041,077. A wild-type SPB can be made integration deletion type by introducing mutations such as K93A, R372A, K375A, R376A, and / or D450N (numbering starts at residue 5, relative to SEQ ID NO:55). The introduction of the mutations R372A, K375A, R376A, and D450N is believed to confer a transposase integration deletion, but maintain cleavage function. An exemplary sequence of an integration deletion transposase domain is PBx containing the NLS set forth in SEQ ID NO: 56. The sequence of an integration deletion PBx transposition matrix domain without the NLS is set forth in SEQ ID NO: 544. GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIV KWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILM MCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (sequence number 544).
[0068] Transposase domain containing N-terminal deletion In some embodiments, provided herein are transposase domains (e.g., SPB transposase domains or PBx transposase domains) that include a deletion of a portion of the amino terminus (also referred to as the "N-terminus" or "N-terminal domain" or "NTD") of the transposase domain. Without wishing to be bound by theory, it is believed that in the context of a tandem dimeric transposase (or a dimer comprising two fusion proteins described herein), the N-terminal domain of the transposase (e.g., SPB) may introduce steric hindrance between the two dimers of the tandem dimer or between the dimer and the DNA.
[0069] In some embodiments, the deleted portion of the N-terminus is about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 100 amino acids, or about 115 amino acids. In some embodiments, the deleted portion of the N-terminus is about 15-25 amino acids, about 25-35 amino acids, about 35-45 amino acids, about 45-55 amino acids, about 55-65 amino acids, about 65-75 amino acids, about 75-85 amino acids, about 85-95 amino acids, about 95-105 amino acids, or about 105-120 amino acids.
[0070] In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-20 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-40 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-60 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-80 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-83 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-84 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-85 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-86 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544.In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-87 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-88 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-89 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-90 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-91 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-92 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-93 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-94 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544.In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-95 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-96 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-97 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-98 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-99 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-100 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-101 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-102 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544.In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-103 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544. In some embodiments, the transposase domain comprises a deletion of N-terminal amino acids 1-115 relative to SEQ ID NO:1, 55, or 56 (numbering begins at residue 12 of SEQ ID NO:1, and residue 5 of SEQ ID NOs:55 and 56), or relative to SEQ ID NO:544.
[0071] Exemplary sequences of an SPB transposase domain with amino acids 1-93 deleted from the N-terminus and a PBx transposase domain with amino acids 1-93 deleted from the N-terminus are shown in SEQ ID NOs: 65 and 66, respectively: NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFD FLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNI TCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 65) NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFD FLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNI TCDNWFTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 66)
[0072] Other exemplary sequences of SPB transposition matrix domains containing N-terminal deletions are set forth in SEQ ID NOs: 2 to 7. Exemplary sequences of PBx transposase domains containing N-terminal deletions are set forth in Table 1, SEQ ID NOs: 86 to 106. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6]
[0073] Fusion proteins containing a transposase domain Also provided herein are fusion proteins comprising one or more transposase domains described herein.
[0074] In some embodiments, provided herein is a fusion protein comprising an SPB or PBx domain and a DNA targeting domain. The DNA targeting domain is further described below. In some embodiments, provided herein is a fusion protein comprising an SPB or PBx domain, a DNA targeting domain, and a protein stabilizing domain (PSD). The PSD is further described below.
[0075] In some embodiments, the fusion proteins provided herein comprise, from N-terminus to C-terminus, a PSD, a DNA targeting domain, and a transposase domain comprising an N-terminal deletion.
[0076] In some embodiments, the fusion protein comprises two transposase domains, e.g., SPB or PBx. In some embodiments, provided herein is a fusion protein comprising a first transposase domain and a second transposase domain, wherein the first transposase domain is a full-length transposase domain (e.g., SPB set forth in SEQ ID NO: 1 or 55, or PBx set forth in SEQ ID NO: 56 (numbering begins at residue 12 of SEQ ID NO: 1 and residue 5 of SEQ ID NOs: 55 and 56), or PBx set forth in SEQ ID NO: 544), and the second transposase domain is identical to the first transposase domain, except that the second transposase domain comprises an N-terminal deletion. In certain aspects, both the first and second transposase domains are piggyBac transposase domains. In certain aspects, the second transposase domain comprises an N-terminal deletion and is a hyperactive piggyBac transposase domain. In certain aspects, the second transposase domain comprises an N-terminal deletion and is a hyperactive piggyBac transposase domain. In certain aspects, the second transposase domain comprises an N-terminal deletion and is a PBx transposase domain. In certain aspects, the second transposase domain comprises an N-terminal deletion and is SPB. In certain aspects, both the first and second transposase domains are hyperactive piggyBac transposase domains. In some embodiments, the first and / or second transposase domains are PBx transposase domains. A schematic diagram illustrating an exemplary fusion protein construct is shown in FIG. 1A.
[0077] In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 81 amino acids, about 82 amino acids, about 83 amino acids, about 84 amino acids, about 85 amino acids, about 86 amino acids, about 87 amino acids, about 88 amino acids, about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises an N-terminal deletion of about 15-25 amino acids, about 25-35 amino acids, about 35-45 amino acids, about 45-55 amino acids, about 55-65 amino acids, about 65-75 amino acids, about 75-85 amino acids, about 85-95 amino acids, about 95-105 amino acids, or about 105-120 amino acids. In certain aspects, the first full-length transposase domain further comprises an in-frame nuclear localization sequence (NLS). In certain aspects, the in-frame NLS is located upstream (i.e., N-terminal) of the nucleotide sequence encoding the first transposase domain. In some embodiments, the NLS comprises or consists of the sequence of SEQ ID NO:15.
[0078] In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-20 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-40 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-60 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-80 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-81 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-82 at the N-terminus.In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-83 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-84 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-85 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-86 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-87 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-88 at the N-terminus.In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-89 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-90 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-91 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-92 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-93 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-94 at the N-terminus.In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-95 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-96 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-97 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-98 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-99 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain comprises a deletion of amino acids 1-100 at the N-terminus.In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain except that the second transposase domain comprises a deletion of amino acids 1-101 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain except that the second transposase domain comprises a deletion of amino acids 1-102 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain except that the second transposase domain comprises a deletion of amino acids 1-103 at the N-terminus. In some embodiments, the first transposase domain of the fusion protein is a full-length transposase domain and the second transposase domain of the fusion protein is identical to the first transposase domain, except that the second transposase domain includes a deletion of amino acids 1-115 at the N-terminus.
[0079] In certain aspects, the amino terminus of the second transposase domain of the fusion protein is fused to the C-terminus of the first transposase domain via a linker sequence. In some embodiments, the linker is 10-15 amino acids in length. In some embodiments, the linker is 13 amino acids in length. In some embodiments, the linker comprises, consists of, or consists essentially of the amino acid sequence ARLAKLGGGAPAVGGGPKAADKGLP (SEQ ID NO: 16).
[0080] In certain aspects, provided herein are fusion proteins that comprise, in an N-terminal to C-terminal orientation, an in-frame NLS, a first hyperactive piggyBac full-length transposase domain, a linker, and a second transposase domain that comprises an N-terminal deletion. Exemplary sequences for such fusion proteins are set forth in SEQ ID NOs: 8-14, although it will be apparent to one of skill in the art that any of the transposase domains set forth in SEQ ID NOs: 1-7, 55, 56, 58, 59, 65-67, 80-106, or 544 can be freely combined in any order and in any orientation within the context of the fusion proteins provided herein.
[0081] An exemplary sequence of a fusion protein comprising a full-length transposase domain is set forth in SEQ ID NO: 8. In some embodiments, the fusion proteins provided herein comprise a sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the sequence set forth in SEQ ID NO:8.
[0082] In some embodiments, the fusion proteins provided herein comprise two transposase domains (e.g., an SPB transposase domain set forth in SEQ ID NO:1 or 55 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NO:55), or a PBx transposase domain set forth in SEQ ID NO:544), each of which comprises an N-terminal deletion compared to a wild-type transposase domain. The two transposase domains can have identical sequences, or they can have different sequences. For example, each of the two transposase domains that comprise an N-terminal deletion can comprise any one of SEQ ID NOs:2-7, or a sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to a sequence set forth in any one of SEQ ID NOs:2-7. In some embodiments, each of the two transposase domains comprising an N-terminal deletion comprises any one of SEQ ID NOs: 86-106, or a sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 86-106.
[0083] In certain embodiments, the fusion proteins provided herein comprise a first full-length transposase domain (e.g., SPB set forth in SEQ ID NO:1 or 55 (numbering begins at residue 12 of SEQ ID NO:1 and residue 5 of SEQ ID NO:55), or PBx set forth in SEQ ID NO:544), and a second transposase domain, wherein the first transposase domain and the second transposase domain are N-terminally deleted (e.g., a transposase domain comprising a sequence set forth in any one of SEQ ID NOs:2-7, or a PBx set forth in any one of SEQ ID NOs:2-7). or a transposase domain comprising a sequence set forth in any one of SEQ ID NOs: 86-106, or a transposase domain comprising a sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to a sequence set forth in any one of SEQ ID NOs: 86-106.
[0084] In some embodiments, the fusion protein comprises a first full-length transposase domain and a second transposase domain, where the first transposase domain and the second transposase domain are identical except that the second transposase domain comprises an N-terminal deletion of about 20 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO:9. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO:9.
[0085] In some embodiments, the fusion protein comprises a first full-length transposase domain and a second transposase domain, where the first transposase domain and the second transposase domain are identical except that the second transposase domain comprises an N-terminal deletion of about 40 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 10. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 10.
[0086] In some embodiments, the fusion protein comprises a first full-length transposase domain and a second transposase domain, where the first transposase domain and the second transposase domain are identical except that the second transposase domain comprises an N-terminal deletion of about 60 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO:11. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO:11.
[0087] In some embodiments, the fusion protein comprises a first full-length transposase domain and a second transposase domain, where the first transposase domain and the second transposase domain are identical except that the second transposase domain comprises an N-terminal deletion of about 80 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 12. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 12.
[0088] In some embodiments, the fusion protein comprises a first full-length transposase domain and a second transposase domain, where the first transposase domain and the second transposase domain are identical except that the second transposase domain comprises an N-terminal deletion of about 100 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 13. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 13.
[0089] In some embodiments, the fusion protein comprises a first full-length transposase domain and a second transposase domain, where the first transposase domain and the second transposase domain are identical except that the second transposase domain comprises an N-terminal deletion of about 115 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 14. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 14.
[0090] DNA targeting domain The transposase domain and fusion protein provided herein may further comprise one or more DNA targeting domains. The DNA targeting domain can be attached to the C-terminus or N-terminus of the transposase domain or fusion protein. In a preferred embodiment, the DNA targeting domain is attached to the N-terminus of the transposase domain, for example, the transposase domain that includes an N-terminal deletion. Without being bound by theory, it is believed that adding a DNA targeting domain to the transposase domain improves the site-specific transposase activity by targeting the transposase fused to the DNA targeting domain to the targeting site. In some embodiments, inserting a DNA targeting domain improves the site-specific transposase activity by at least 2-fold, at least 3-fold, at least 4-fold, or at least 5-fold compared to the same transposase domain without the DNA targeting domain.
[0091] Any DNA targeting domain known in the art can be used in the context of the transposase domains, fusion proteins, and tandem dimer transposases described herein, including, but not limited to, CRISPR, zinc finger motifs, TALEs, and transcription factors. In some embodiments, the DNA targeting domain comprises three zinc finger motifs. In some embodiments, the three zinc finger motifs are flanked by a GGGGS linker. In some embodiments, the three zinc finger motifs flanked by a GGGGs linker cumulatively comprise the sequence set forth in SEQ ID NO: 57: GGGGSERPYACPVESCDRRFSRSDELTRHIRIHTGQKPFQCRICMRNFSRSDHLTTHIRTHTGEKPFACDICGRKFARSDERKRHTKIHLRQKDGGGGS (SEQ ID NO: 57), or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto.
[0092] In certain embodiments, provided herein is a fusion protein comprising a transposase domain comprising an N-terminal deletion, an NLS, and three zinc finger motifs. In some embodiments, the NLS comprises or consists of the sequence set forth in SEQ ID NO: 15.
[0093] In some embodiments, the DNA targeting domain is a TAL array. TALEs (transcription activation-like effectors) from Xanthomonas typically contain a 288 amino acid N-terminus followed by a variable number of arrays of about 34 amino acid repeats followed by a 278 amino acid C-terminus (SEQ ID NO: 77), although truncated versions have been described in the literature (see, for example, Miller et al., Nat Biotechnol 29, 143-148 (2011)). TALs fused to FokI nucleases (called TALENs) most often contain N- and C-terminal truncations. For example, the first 152 amino acids at the N-terminus are often removed (called delta 152; SEQ ID NO: 73), and the C-terminus is often pilfered leaving 63 amino acids (called +63; SEQ ID NO: 76).
[0094] TALs contain an array of 34 amino acids that are repeated a variable number of times. Two amino acids at positions 12 and 13 are varied and determine the nucleotide that the TAL repeat recognizes. This property allows the TAL array to be programmed to bind to a specific DNA sequence. Amino acids NG recognize T, NI recognize A, NN recognize G or A, HD recognize C, NK recognize G, and NS recognize A, C, G, or T. Other amino acids within the 34 residue repeat may also be varied. For example, position 11 is often changed to N for a repeat that recognizes G. Positions 4 and 32 are also often varied to reduce the repetitiveness of the array but not to determine binding specificity. The number of 34 amino acid repeats in the array determines the length of the DNA sequence recognized (one protein repeat binds one DNA bp). Additionally, the last bp is recognized by a "half array" that is 20 amino acids instead of 34.
[0095] Furthermore, the N-terminal domain of TAL (e.g., SEQ ID NO: 73) recognizes and requires a T located immediately 5' of the target DNA sequence. Mutations in the TAL N-terminal domain have been described in the literature that no longer require a 5' T (Lamb et al., Nucleic Acids Res. 2013 Nov; 41(21): 9779-85. doi: 10.1093 / nar / gkt754. Epub 2013 Aug 26. PMID: 23980031; PMCID: PMC3834825). For example, the NT-G mutant requires a 5' G (SEQ ID NO: 74) instead of a 5' T, while NT-βN does not require any specific 5' nucleotide (SEQ ID NO: 75). These mutated N-terminal domain sequences can be used to provide additional sequence options that can be targeted using TAL arrays.
[0096] Each TAL array contains a 34 amino acid repeat followed by nine 20 amino acid "half" repeats and was synthesized adjacent to a BsmBI type IIS restriction enzyme cleavage site. In one embodiment, individual TAL molecules containing either the 34 amino acid or the 20 amino acid "half" repeat can be designed and synthesized adjacent to a BsmBI type IIS restriction enzyme cleavage site. The entire TAL module set contains four molecules capable of recognizing either A, C, G, or T for each of the 10 bp positions (40 molecules / 10 bp target) and one TAL half-repeat molecule. Exemplary TAL molecules are set forth in SEQ ID NOs: 107-110, where X is any of the following amino acids: TAL module version 1: LTPDQVVAIAXXXGGKQALETVQRLLPVLCQDHG (sequence number 107) TAL module version 2: LTPEQVVAIAXXXGGKQALETVQRLLPVLCQAHG (sequence number 108) TAL module version 3'' LTPDQVVAIAXXXGGKQALETVQRLLPVLCQAHG (sequence number 109) TAL module version 4: LTPAQVVAIAXXXGGKQALETVQRLLPVLCQDHG (SEQ ID NO: 110).
[0097] An exemplary TAL half-module is set forth in SEQ ID NO:111, where X is any of the following amino acids: LTPEQVVAIAXXXGGRPALE.
[0098] Pairs of TAL array targeting sequences can be designed within desired genes, and corresponding molecules can be selected and pooled using "Golden Gate Assembly" to assemble each TAL array in frame. The DNA sequences encoding the TAL arrays generated herein can be further codon-optimized using the GeneArt algorithm (Thermo Fisher).
[0099] In designing left and right TAL arrays, which contain an N-terminal domain that recognizes T and a TAL C-terminal domain that can be fused to an N-terminal deleted transposase sequence (i.e., TAL-ssSPB, or TAL-PBx; see below), one TAL array recognizes a sequence 5' to TTAA, and the other TAL array recognizes a sequence 3' to TTAA. Since the sequence 5' to TTAA is most often different from the sequence 3' to TTAA in the genomic DNA target, TAL-ssSPB is most often used as a heterodimer consisting of two different TAL domains that recognize two different DNA sequences. Furthermore, the sequence recognized by the TAL array is not directly adjacent to the TTAA. Instead, the sequence is separated from the TTAA by a spacer of a given bp length, for example, a spacer of 12bp, 13bp, or 14bp.
[0100] TAL arrays can target any DNA sequence of interest (e.g., genomic DNA sequences). It will be apparent to one skilled in the art that any left TAL sequence for a given target can be combined with any right TAL array for the same target.
[0101] In some embodiments, the TAL array targets green fluorescent protein (GFP). Exemplary sequences of the left TAL array targeting GFP are set forth in SEQ ID NOs: 113 and 115. Exemplary sequences of the right TAL array targeting GFP are set forth in SEQ ID NOs: 114 and 116. In some embodiments, the left TAL array targeting GFP binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NO: 240 or 242, and in some embodiments, the right TAL array targeting GFP binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NO: 241 or 243.
[0102] In some embodiments, the TAL array targets ZFN268. An exemplary sequence of a TAL array targeting ZFN268 serving as a left and right array is set forth in SEQ ID NO: 112. In some embodiments, the TAL array targeting ZFN268 binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NO: 239.
[0103] In some embodiments, the TAL array targets phenylalanine hydroxylase (PAH). Exemplary sequences of the left TAL array targeting PAH are set forth in SEQ ID NOs: 117, 119, 121, 123, 125, and 127. Exemplary sequences of the right TAL array targeting PAH are set forth in SEQ ID NOs: 118, 120, 122, 124, 126, and 128. In some embodiments, the left TAL array targeting PAH binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NOs: 244, 246, 248, 250, 252, or 254. In some embodiments, the left TAL array targeting PAH binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NOs: 245, 247, 249, 251, 253, or 255. Exemplary genomic target sites for PAH are set forth in SEQ ID NOs: 360-365.
[0104] In some embodiments, the TAL array targets a LINE1 repetitive element. Exemplary sequences of the left TAL array targeting a LINE1 repetitive element are set forth in SEQ ID NOs: 129, 131, 134, 136, 137, 139, and 141. Exemplary sequences of the right TAL array targeting a LINE1 repetitive element are set forth in SEQ ID NOs: 130, 132, 133, 135, 138, 140, 142, and 143. In some embodiments, the left TAL array targeting a LINE1 repetitive element binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NOs: 256, 258, 261, 263, 264, 266, or 268. In some embodiments, the right TAL array targeting a LINE1 repetitive element binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NOs: 257, 259, 260, 262, 265, 267, 269, or 270. Exemplary genomic target sites for LINE1 elements are set forth in SEQ ID NOs: 36-374.
[0105] In some embodiments, the TAL array targets the beta-2-microglobulin gene (B2M). Exemplary sequences of the left TAL array targeting B2M are set forth in SEQ ID NOs: 144, 146, 148, 150, 152, 154, 156, 518, and 520. Exemplary sequences of the right TAL array targeting B2M are set forth in SEQ ID NOs: 145, 147, 149, 151, 153, 155, 157, 519, and 521. In some embodiments, the left TAL array targeting B2M binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NOs: 271, 273, 275, 277, 279, 281, 283, 514, or 516. In some embodiments, the right TAL array targeting B2M binds to a nucleic acid molecule comprising a sequence set forth in SEQ ID NO: 272, 274, 276, 278, 280, 282, 284, 515, or 517. Exemplary genomic target sites for B2M are set forth in SEQ ID NOs: 375-381.
[0106] The DNA targeting domain can be fused to or linked to the N-terminus of the transposase domain that contains the N-terminal deletion. For example, the DNA targeting domain can be inserted into the transposase domain at a suitable position within the N-terminal region of the transposase domain.
[0107] The DNA targeting domain can be inserted at the N-terminus of the transposase domain. In some embodiments, the DNA targeting domain is inserted between the 82nd and 83rd amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 83rd and 84th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 84th and 85th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 85th and 86th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 86th and 87th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 87th and 88th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 88th and 89th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 89th and 90th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 90th and 91st amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 91st and 92nd amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544.In some embodiments, the DNA targeting domain is inserted between the 92nd and 93rd amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 93rd and 94th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 94th and 95th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 95th and 96th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 96th and 97th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 97th and 98th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 98th and 99th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 99th and 100th amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 100th and 101st amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between the 101st and 102nd amino acids of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO:544. In some embodiments, the DNA targeting domain is inserted between amino acids 102 and 103 of SEQ ID NO: 55 or 56 (numbering starting at the 5th amino acid), or SEQ ID NO: 544.In some embodiments, the DNA targeting domain is inserted between amino acids 103 and 104 of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO: 544. In some embodiments, the DNA targeting domain is inserted between amino acids 104 and 105 of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or SEQ ID NO: 544. In some embodiments, the DNA targeting domain comprises a sequence of SEQ ID NO: 57 or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto. The transposase domain may further comprise an NLS, for example, the NLS of SEQ ID NO: 15.
[0108] In some embodiments, the DNA targeting domain replaces the 83rd amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 84th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 85th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 86th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 87th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 88th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 89th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 90th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 91st amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 92nd amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain replaces the 93rd amino acid of SEQ ID NO:55 or 56 (numbering starting at the 5th amino acid), or of SEQ ID NO:544.In some embodiments, the DNA targeting domain replaces the 94th amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 95th amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 96th amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 97th amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 98th amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 99th amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 100th amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 101st amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 102nd amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 103rd amino acid of SEQ ID NO:55 or 56 (numbering starting from the 5th amino acid) or of SEQ ID NO:544. In some embodiments, the DNA targeting domain replaces the 104th amino acid of SEQ ID NO:55 or 56 (numbering starting at the 5th amino acid), or of SEQ ID NO:544.In some embodiments, the DNA targeting domain replaces the 105th amino acid of SEQ ID NO: 55 or 56 (numbering starting from the 5th amino acid), or of SEQ ID NO: 544. In some embodiments, the DNA targeting domain comprises a sequence of SEQ ID NO: 57, or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto. The transposase domain may further comprise an NLS, for example, the NLS of SEQ ID NO: 15.
[0109] An exemplary sequence of a fusion protein containing a transposase domain containing a 93 amino acid N-terminal deletion, an NLS, and three zinc finger motifs flanked by GGGGS linkers is shown in SEQ ID NO: 58, where the NLS is shown in italics, the sequence containing the three zinc finger motifs and the GGGGS linker is underlined, and the transposase domain containing the 93 amino acid N-terminal deletion is shown in bold. [Table 3]
[0110] An exemplary sequence of a fusion protein comprising an integrated deleted transposase domain containing a 93 amino acid N-terminal deletion, an NLS, and three zinc finger motifs flanked by GGGGS linkers is set forth in SEQ ID NO: 59, where the NLS is shown in italics, the sequence containing the three zinc finger motifs and the GGGGS linker is underlined, and the transposase domain containing the 93 amino acid N-terminal deletion is shown in bold. [Table 4]
[0111] Protein Stabilization Domains In some embodiments, the fusion protein provided herein may further comprise a protein stabilization domain (PSD).When present, the PSD is preferably attached to the N-terminus of the DNA targeting domain.Without wishing to be bound by theory, it is believed that adding the PSD can improve the protein stability or stability of transposase tetramer-DNA complex.
[0112] The PSD can be approximately the same size as the N-terminal deletion in the transposase domain, for example, in some embodiments, the N-terminal deletion of the transposase domain includes amino acids 1-93 and the PSD includes 92 amino acids.
[0113] In some embodiments, the PSD comprises amino acids 1-90 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-90 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-91 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-91 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-92 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-92 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-92 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-93 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-93 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-94 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-94 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-95 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-95 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-96 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-96 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-97 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-97 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-98 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55).In some embodiments, the PSD comprises amino acids 1-98 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-99 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-99 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56). In some embodiments, the PSD comprises amino acids 1-100 of SEQ ID NO:55 (numbering begins at residue 5 of SEQ ID NO:55). In some embodiments, the PSD comprises amino acids 1-100 of SEQ ID NO:56 (numbering begins at residue 5 of SEQ ID NO:56).
[0114] In some embodiments, the PSD comprises the sequence GSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRG (SEQ ID NO: 68).
[0115] Thus, provided herein is a fusion protein that comprises, from N-terminus to C-terminus, a nuclear localization signal (NLS), a PSD, a DNA targeting domain, and a transposase domain that comprises an N-terminal deletion compared to the sequence set forth in SEQ ID NO:55 or 56 (numbering beginning at residue 5 of SEQ ID NO:55 or 56).
[0116] Exemplary sequences of fusion proteins comprising a PSD, an NLS, a DNA targeting domain, and a transposase domain comprising an N-terminal deletion are shown in SEQ ID NOs: 67 (PBx transposase domain) and 69 (SPB transposase domain), with the NLS (here, PKKKRKV) in italics, the NTD in bold and underlined, the DNA targeting domain (here, three zinc finger motifs flanked by GGGGS linkers) underlined, and the N-terminal deleted transposase domain (here, PBx) in bold: [Table 5]
[0117] Nuclear localization signal In some embodiments, the transposase domain and fusion protein provided herein may comprise an in-frame nuclear localization sequence (NLS). Examples of transposases fused to nuclear localization signals are disclosed in U.S. Patent Nos. 6,218,185; 6,962,810, 8,399,643, and WO 2019 / 17363. In some embodiments, the NLS comprises the sequence of PKKKRKV (SEQ ID NO: 15). In certain aspects, the in-frame NLS is located upstream (N-terminal) of the transposase domain that comprises the N-terminal deletion.
[0118] Generally, NLS is preferably located at the N-terminus of fusion protein.In some embodiments, NLS is fused or linked to the N-terminus of transposase domain.In some embodiments, NLS is fused or linked to the N-terminus of PSD.In some embodiments, NLS is fused or linked to the N-terminus of PSD.
[0119] In certain aspects, the in-frame NLS is fused directly to the amino terminus of the transposase domain comprising the N-terminal deletion. In some embodiments, the NLS is linked to the N-terminus of the transposase domain comprising the N-terminal deletion via a linker (e.g., a GGGGS linker or a GGS linker).
[0120] In some embodiments, an initiating methionine is introduced before the NLS. In some embodiments, an additional alanine residue is introduced before and / or after the NLS to ensure in-frame translation. As such, the numbering of residues in SEQ ID NO:1 starts at residue 12 of SEQ ID NO:1 to identify the deleted and mutated residues. In SEQ ID NOs:55 and 56, which are sequences of SPB and PBx, respectively, without the NLS, the numbering of residues starts at residue 5 to identify the deleted and mutated residues. In SEQ ID NO:544, the numbering starts at residue 1 to identify the deleted and mutated residues.
[0121] In some embodiments, the fusion protein comprises an NLS and a transposase domain comprising an N-terminal deletion of 20 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 2. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 2.
[0122] In some embodiments, the fusion protein comprises an NLS and a transposase domain comprising an N-terminal deletion of 40 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 3.
[0123] In some embodiments, the fusion protein comprises an NLS and a transposase domain comprising an N-terminal deletion of 60 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 4.
[0124] In some embodiments, the fusion protein comprises an NLS and a transposase domain comprising an N-terminal deletion of 80 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 5. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 5.
[0125] In some embodiments, the fusion protein comprises an NLS and a transposase domain comprising an N-terminal deletion of 100 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO:6. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO:6.
[0126] In some embodiments, the fusion protein comprises an NLS and a transposase domain comprising an N-terminal deletion of 115 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 7.
[0127] In some embodiments, the fusion protein comprises an NLS and a transposase domain comprising an N-terminal deletion of 93 amino acids. In some embodiments, the fusion protein comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 65. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 65.
[0128] Obligatory heterodimers and tandem dimers In another aspect, provided herein is a tandem dimeric transposase comprising two fusion proteins, each fusion protein comprising a first and a second transposase domain, and one or both fusion proteins further comprising a DNA targeting domain. In some embodiments, both fusion proteins comprise a DNA targeting domain. In some embodiments, both fusion proteins comprise a DNA targeting domain, and the DNA targeting domain targets a DNA sequence adjacent to the DNA sequence that is the insertion site targeted by the transposase. In some embodiments, only one of the two fusion proteins in the tandem dimeric transposase comprises a DNA targeting domain. The DNA targeting domain can be attached to the C-terminus or N-terminus of the fusion protein.
[0129] Thus, in some embodiments, a first fusion protein comprising: (a) a first transposase domain, a linker, a second transposase domain, and a first DNA targeting domain, where (i) the first and second transposase domains are identical, or (ii) the first and second transposase domains are identical except that the second transposase domain comprises an N-terminal deletion; and a second fusion protein comprising: and (i) the first and second transposase domains are identical, or (ii) the first and second transposase domains are identical except for the second transposase domain comprising an N-terminal deletion; wherein the first DNA targeting domain and the second DNA targeting domain are different, and the transposase domain of the first fusion protein and the transposition matrix domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex.
[0130] In some embodiments, provided herein is a complex comprising: (a) a first fusion protein comprising, from N-terminus to C-terminus, a first NLS, a first DNA targeting domain, a first transposase domain comprising an N-terminal deletion, a linker, and a second transposase domain; and (b) a second fusion protein comprising, from N-terminus to C-terminus, a second NLS, a second DNA targeting domain, a third transposase domain comprising an N-terminal deletion, a linker, and a fourth transposase domain, wherein the transposase domain of the first fusion protein and the transposition matrix domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex. In some embodiments, the first, second, third, and / or fourth transposase domain is a SPB domain. In some embodiments, the first, second, third, and / or fourth transposase domain is a PBx transposase domain. In some embodiments, the first and / or third transposase domain comprises an N-terminal deletion of 83, 84, 85, 86, 87, 88, 89, 90, 91, 21, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103 amino acids. In some embodiments, the first and third transposase domain comprises a sequence of SEQ ID NO: 65 or 66. In some embodiments, the second and fourth transposase domain comprises a sequence of SEQ ID NO: 55, 56, or 544. In some embodiments, the first and / or second DNA targeting domain comprises three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domain comprises a sequence of SEQ ID NO: 57. In some embodiments, the first and / or second DNA targeting domain comprises a TAL motif.
[0131] In some embodiments, provided herein is a complex comprising: (a) a first fusion protein comprising, from N-terminus to C-terminus, a first NLS, a first PSD, a first DNA targeting domain, a first transposase domain comprising an N-terminal deletion, a linker, and a second transposase domain; and (b) a second fusion protein comprising, from N-terminus to C-terminus, a second NLS, a second PSD, a second DNA targeting domain, a third transposase domain comprising an N-terminal deletion, a linker, and a fourth transposase domain, wherein the transposase domain of the first fusion protein and the transposition matrix domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex. In some embodiments, the first, second, third, and / or fourth transposase domain is a SPB domain. In some embodiments, the first, second, third, and / or fourth transposase domain is a PBx transposase domain. In some embodiments, the first and third transposase domains comprise the sequence of SEQ ID NO: 65 or 66. In some embodiments, the second and fourth transposase domains comprise the sequence of SEQ ID NO: 55, 56, or 544. In some embodiments, the first and / or second PSD comprises the sequence of SEQ ID NO: 68. In some embodiments, the first and / or second DNA targeting domain comprises three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domain comprises the sequence of SEQ ID NO: 57.
[0132] In some embodiments, provided herein is a complex comprising: (a) a first fusion protein comprising, from N-terminus to C-terminus, a first transposase domain comprising a first NLS, the sequence of SEQ ID NO: 55, 56, or 544, a linker, a second transposase domain; and (b) a second fusion protein comprising, from N-terminus to C-terminus, a third transposase domain comprising a second NLS, the sequence of SEQ ID NO: 55, 56, or 544, a linker, and a fourth transposase domain, wherein the first and third transposase domains comprise a DNA targeting domain, and the transposase domain of the first fusion protein and the transposition matrix domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex. In some embodiments, the second and / or fourth transposase domain is a SPB domain. In some embodiments, the second and / or fourth transposase domain is a PBx transposase domain. In some embodiments, the second and fourth transposase domains comprise the sequence of SEQ ID NO: 55, 56, or 544. In some embodiments, the first and / or second PSDs comprise the sequence of SEQ ID NO: 68. In some embodiments, the first and / or second DNA targeting domains comprise three zinc finger motifs. In some embodiments, the first and / or second DNA targeting domains comprise the sequence of SEQ ID NO: 57. In some embodiments, the first DNA targeting domain replaces residues 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103 of the first transposase domain, with numbering starting at residue 5 of SEQ ID NO: 55 or 56.In some embodiments, the second DNA targeting domain replaces residues 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103 of the third transposase domain, where numbering begins at residue 5 of SEQ ID NO: 55 or 56.
[0133] In some embodiments, provided herein is a complex comprising: (a) a first fusion protein comprising a first transposase domain, a linker, a second transposase domain, and a first DNA targeting domain, wherein the first and / or the second transposase domain of the first fusion protein comprise an identical amino acid sequence set forth in any one of SEQ ID NOs: 31-43; and (b) a second fusion protein comprising a first transposase domain, a linker, a second transposase domain, and a second DNA targeting domain, wherein the first and / or the second transposase domain of the second fusion protein comprise an identical amino acid sequence set forth in any one of SEQ ID NOs: 44-53.
[0134] In another aspect, provided herein is a fusion protein comprising a first transposase domain and a second transposase domain, capable of forming the requisite heterodimer with another fusion protein comprising the first transposase domain and the second transposase domain. Without wishing to be bound by theory, it is believed that such two fusion proteins assemble into a tandem dimer structure, held together by a combination of charge interactions, hydrogen bonds, π-cation pairs, and hydrophobic interactions. Such tandem dimer structures are referred to herein as "tandem dimer transposases". Thus, each tandem dimer comprises four transposase domains. In some embodiments, two fusion proteins provided herein form a complex comprising (a) a first fusion protein comprising a first transposase domain, a linker, and a second transposase domain, and (b) a second fusion protein comprising a first transposase domain, a linker, and a second transposase domain, wherein the transposase domain of the first fusion protein and the transposition matrix domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex.
[0135] In some embodiments, the first fusion protein comprises a first transposase domain of SEQ ID NO:55 and / or a second transposase domain of SEQ ID NO:65. In some embodiments, the second fusion protein comprises a first transposase domain of SEQ ID NO:55 and / or a second transposase domain of SEQ ID NO:65. In some embodiments, the first fusion protein comprises a first transposase domain of SEQ ID NO:55 and / or a second transposase domain of SEQ ID NO:66. In some embodiments, the second fusion protein comprises a first transposase domain of SEQ ID NO:55 and / or a second transposase domain of SEQ ID NO:66. In some embodiments, the first fusion protein comprises a first transposase domain of SEQ ID NO:55 and / or a second transposase domain of SEQ ID NO:67. In some embodiments, the second fusion protein comprises a first transposase domain of SEQ ID NO:55 and / or a second transposase domain of SEQ ID NO:67.
[0136] In some embodiments, the first fusion protein comprises a first transposase domain of SEQ ID NO:56 and / or a second transposase domain of SEQ ID NO:65. In some embodiments, the second fusion protein comprises a first transposase domain of SEQ ID NO:56 and / or a second transposase domain of SEQ ID NO:65. In some embodiments, the first fusion protein comprises a first transposase domain of SEQ ID NO:56 and / or a second transposase domain of SEQ ID NO:66. In some embodiments, the second fusion protein comprises a first transposase domain of SEQ ID NO:56 and / or a second transposase domain of SEQ ID NO:66. In some embodiments, the first fusion protein comprises a first transposase domain of SEQ ID NO:56 and / or a second transposase domain of SEQ ID NO:67. In some embodiments, the second fusion protein comprises a first transposase domain of SEQ ID NO:56 and / or a second transposase domain of SEQ ID NO:67.
[0137] Introducing charged residues into amino acids that contribute to dimerization with a second fusion protein allows for the design of a pair of fusion proteins that can only associate with each other to form a tandem dimer of a given structure. Introducing mutations that allow only one structure of the tandem dimer allows the introduction of a DNA targeting domain into the fusion protein, thereby increasing the specificity of the transposase domain. This is shown in Figures 2A and 2B for SPB and Figures 2C and 2D for PBx. Introducing a DNA targeting domain into a fusion protein that can dimerize in any structure, including homodimerization, results in four DNA targeting domains present in the tandem dimer transposase. However, only two DNA targeting domains interact with DNA, while the other two potentially remain sterically hindered from transposase-DNA interaction. Any suitable DNA targeting domain described herein or known in the art can be used in the fusion proteins described herein.
[0138] One skilled in the art would readily determine mutations in the transposase domain that confer positive or negative charge. In the case of a fusion protein containing a first and a second transposase domain, the crystal structure published in Chen et al. (Nat Commun 11, 3446 (2020)) can be used to identify residue pairs in the transposase domain that are adjacent to the tandem dimer formed by two such fusion proteins. Altering the charge of such residue pairs to generate positively and negatively charged transposase domains can be accomplished using standard techniques such as site-directed mutagenesis.
[0139] For example, one or more of M185, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593, and / or F594 can be mutated within an SPB transposase domain (e.g., an SPB as set forth in SEQ ID NO:1 or 55 (numbering begins at residue 12 of SEQ ID NO:1 and at residue 5 of SEQ ID NO:55) to generate an SPB- or SPB+ transposase domain. Similarly, M1 One or more of 85, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593, and / or F594 can be mutated in a PBx transposase domain (e.g., a PBx transposase domain of SEQ ID NO: 56 (numbering begins at the 5th residue of SEQ ID NO: 56) or a PBx transposase domain of SEQ ID NO: 544) to generate a PBx- or a PBx+ transposase domain.
[0140] The fusion proteins described herein can contain (i) one or two SPB+ transposase domains, or (ii) one or two SPB- transposase domains.
[0141] To achieve the requisite heterodimer formation, pairs of mutations can be introduced into the fusion protein or transposase domain to generate positively and negatively charged fusion proteins or transposase domains that can then interact and form heterodimers. In some embodiments, the pairs of residues that are mutated are those listed in Table 2. For example, one or more of the mutations listed in the column labeled "Protein 1" can be introduced into a first SPB or PBx domain, and the corresponding one or more mutations listed in the column labeled "Protein 2" can be introduced into a second SPB or PB domain. In some embodiments, members of a residue pair are mutated to have opposite charges. [Table 6]
[0142] To introduce a positive charge, an amino acid with an uncharged side chain, e.g., methionine, or an amino acid with a negatively charged side chain, e.g., aspartic acid, can be changed to a positively charged amino acid, e.g., lysine or arginine. To introduce a negative charge, an amino acid with a positively charged side chain, e.g., arginine or lysine, or an amino acid with a hydrophobic side chain, e.g., leucine, can be changed to a negatively charged amino acid, e.g., aspartic acid or glutamic acid.
[0143] In certain embodiments, one or more of the following mutations are introduced into one or both SPB transposase domains (e.g., the SPB set forth in SEQ ID NO: 1 or 55 (numbering begins at residue 12 of SEQ ID NO: 1 and residue 5 of SEQ ID NO: 55) of the fusion proteins provided herein to generate an SPB+ fusion protein. In some embodiments, the SPB+ transposase domain is In some embodiments, the SPB+ transposase domain comprises a M185R mutation, and a D198K mutation. In some embodiments, the SPB+ transposase domain comprises a M185R mutation, and a D201R mutation. In some embodiments, the SPB+ transposase domain comprises a D197K mutation, and a D201R mutation. In some embodiments, the SPB+ transposase domain comprises a D198K mutation, and a D201R mutation. In some embodiments, the SPB+ transposase domain comprises a M185R mutation, a D198K mutation, and a D201R mutation.
[0144] In certain embodiments, one or more of the following mutations are introduced into one or both PBx transposase domains of a fusion protein provided herein (e.g., the PBx transposase domain of SEQ ID NO: 56 (numbering begins at the 5th residue of SEQ ID NO: 56) or the PBx transposase domain of SEQ ID NO: 544) to generate a PBx+ fusion protein. In some embodiments, the PBx+ transposase domain comprises an M185R mutation and a D198K mutation. In some embodiments, the PBx+ transposase domain comprises an M185R mutation and a D201R mutation. In some embodiments, the PBx+ transposase domain comprises a D197K mutation and a D201R mutation. In some embodiments, the SPB+ transposase domain comprises a D198K mutation and a D201R mutation. In some embodiments, the PBx+ transposase domain comprises an M185R mutation, a D198K mutation, and a D201R mutation.
[0145] In certain embodiments, one or more of the following mutations are introduced into one or both SPB transposase domains (e.g., an SPB as set forth in SEQ ID NO: 1 or 55 (numbering begins at residue 12 of SEQ ID NO: 1 and residue 5 of SEQ ID NO: 55) of a fusion protein provided herein to generate an SPB-fusion protein. In some embodiments, the SPB-transposase domain comprises a L204E mutation and a K500D mutation. In some embodiments, the SPB-transposase domain comprises a L204E mutation and a R504D mutation. In some embodiments, the SPB-transposase domain comprises a K500 mutation and a R504D mutation. In some embodiments, the SPB-transposase domain comprises a L204E mutation, a K500D mutation, and a R504D mutation.
[0146] In certain embodiments, one or more of the following mutations are introduced into one or both PBx transposases of the fusion proteins provided herein (e.g., the PBx transposase domain of SEQ ID NO: 56 (numbering starts at the 5th residue of SEQ ID NO: 56) or the PBx transposase domain of SEQ ID NO: 544) to generate a PBx-fusion protein. In some embodiments, the PBx-transposase domain comprises an L204E mutation and a K500D mutation. In some embodiments, the PBx-transposase domain comprises an L204E mutation and a R504D mutation. In some embodiments, the PBx-transposase domain comprises a K500 mutation and a R504D mutation. In some embodiments, the PBx-transposase domain comprises an L204E mutation, a K500D mutation, and a R504D mutation.
[0147] Exemplary sequences of SPB+ transposase domains are set forth in SEQ ID NOs: 31-43. Exemplary sequences of SPB- transposase domains are set forth in SEQ ID NOs: 44-53. In some embodiments, the transposase domains provided herein comprise an amino acid sequence set forth in any one of SEQ ID NOs: 31-53. In some embodiments, the transposase domains provided herein comprise an amino acid sequence set forth in any one of SEQ ID NOs: 31-53, further comprising one or more conserved amino acid sequences.
[0148] In some embodiments, the fusion proteins described herein comprise a first transposase domain and a second transposase domain, wherein both the first and second transposase domains comprise an amino acid sequence set forth in any one of SEQ ID NOs: 31-43. In some embodiments, the first and second transposase domains comprise identical sequences. In some embodiments, the first and second transposase domains comprise different sequences. In some embodiments, both the first and second transposase domains comprise an amino acid sequence set forth in any one of SEQ ID NOs: 31-43, further comprising one or more conservative amino acid sequences.
[0149] In some embodiments, the fusion proteins described herein comprise a first transposase domain and a second transposase domain, wherein both the first and second transposase domains comprise an amino acid sequence set forth in any one of SEQ ID NOs: 44-53. In some embodiments, the first and second transposase domains comprise identical sequences. In some embodiments, the first and second transposase domains comprise different sequences. In some embodiments, both the first and second transposase domains comprise an amino acid sequence set forth in any one of SEQ ID NOs: 44-54, further comprising one or more conservative amino acid sequences.
[0150] In some embodiments, provided herein is a complex comprising: (a) a first fusion protein comprising a first transposase domain, a linker, and a second transposase domain, wherein the first and / or the second transposase domain of the first fusion protein comprise an identical amino acid sequence set forth in any one of SEQ ID NOs: 31-43; and (b) a second fusion protein comprising a first transposase domain, a linker, and a second transposase domain, wherein the first and / or the second transposase domain of the second fusion protein comprise an identical amino acid sequence set forth in any one of SEQ ID NOs: 44-53.
[0151] The SPB+, SPB-, PBx+, and PBx- fusion proteins and transposase domains can further comprise an N-terminal deletion of a second transposase domain as described herein. Thus, in some embodiments, provided herein are SPB+ fusion proteins comprising a first and a second SPB+ transposase domain, the first and second SPB+ transposase domains being identical except that the second transposase domain comprises an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 100 amino acids, or about 115 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 93 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 94 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 95 amino acids.In some embodiments, the second transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 98 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 99 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 103 amino acids.
[0152] In some embodiments, an SPB-fusion protein comprises a first and a second SPB-transposase domain, wherein the first and second SPB-transposase domains are a SPB-fusion protein comprising a first SPB-transposase domain, the second transposase domain being a SPB-fusion protein comprising a first SPB-transposase domain, the second transposase domain being a SPB-fusion protein comprising a first SPB-transposase domain, the second transposase domain being a SPB-fusion protein comprising a first SPB-transposase domain, the second transposase domain being a SPB-fusion protein comprising a first SPB-transposase domain, the second transposase domain being a SPB-fusion protein comprising a first SPB-transposase domain, the Provided herein are the same SPB-fusion proteins except that the second transposase domain comprises an N-terminal deletion of about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 93 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 94 amino acids.In some embodiments, the second transposase domain comprises an N-terminal deletion of 95 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 98 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 99 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 103 amino acids.
[0153] In some embodiments, provided herein is a PBx+ fusion protein comprising a first and a second PBx+ transposase domain, the first and second PBx+ transposase domains being identical except that the second transposase domain comprises an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 100 amino acids, or about 115 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 93 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 94 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 95 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 98 amino acids.In some embodiments, the second transposase domain comprises an N-terminal deletion of 99 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 103 amino acids.
[0154] In some embodiments, a PBx-fusion protein comprises a first and a second PBx-transposase domain, wherein the first and second PBx-transposase domains are selected from the group consisting of a first PBx transposase domain, a second PBx transposase domain, a third PBx transposase domain, a fourth PBx transposase domain, a fifth PBx transposase domain, a sixth PBx transposase domain, a sixth PBx transposase domain, a seventh ... Provided herein are PBx-fusion proteins identical thereto except that the second transposase domain comprises an N-terminal deletion of about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 93 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 94 amino acids.In some embodiments, the second transposase domain comprises an N-terminal deletion of 95 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 98 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 99 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the second transposase domain comprises an N-terminal deletion of 103 amino acids.
[0155] In some embodiments, provided herein are complexes comprising: (a) a first fusion protein comprising a first transposase domain, a linker, and a second transposase domain, where (i) the first and second transposase domains are identical, or (ii) the first and second transposase domains are identical, except that the second transposase domain comprises an N-terminal deletion; and (b) a second fusion protein comprising a first transposase domain, a linker, and a second transposase domain, where (i) the first and second transposase domains are identical, or (ii) the first and second transposase domains are identical, except that the second transposase domain comprises an N-terminal deletion.
[0156] The transposon domain sequences provided herein can be freely combined. Thus, in some embodiments, the present specification provides a fusion protein comprising a first transposon domain and a second transposon domain, wherein the first transposon domain comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 31 to 53, and the second transposon domain comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1 to 7. In some embodiments, provided herein is a fusion protein comprising a first transposon domain and a second transposon domain, wherein the first transposon domain comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to a sequence set forth in any one of SEQ ID NOs: 31-53, and the second transposon domain comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to a sequence set forth in any one of SEQ ID NOs: 1-7.
[0157] Built-in Cassette Also provided herein is an integration cassette for site-specific transfer of a DNA molecule into the genome of a cell. In some embodiments, the integration cassette for site-specific transfer of a nucleic acid into the genome of a cell comprises a nucleic acid consisting of an upstream zinc finger motif DNA binding domain binding site ("ZFM-DBD") and a central transposon ITR integration site TTAA sequence flanked by a downstream ZFM-DBD, each of which is separated by 7 base pairs from the TTAA sequence. In some embodiments, each of at least one of the upstream and downstream ZFM-DBD sites is a ZFM268 binding site. In some embodiments, each of the ZFM268 binding sites comprises SEQ ID NO:60. In some embodiments, the integration cassette comprises or consists of SEQ ID NO:62.
[0158] Also provided herein is a cell comprising an integration cassette for site-specific transposition of a DNA molecule stably integrated into the genome of the cell. In some embodiments, the integration cassette comprises or consists of SEQ ID NO:62.
[0159] Also provided is a method for site-specific transposition of a DNA molecule into the genome of a cell containing a stably integrated integration cassette, comprising integrating into the cell a) a nucleic acid encoding a fusion protein comprising a DNA binding domain and a transposase, the fusion protein being expressed in the cell, and b) a DNA molecule comprising a transposon, the expressed fusion protein integrating the transposon into the TTAA sequence of the stably integrated integration cassette by site-specific transposition. In some embodiments of the method, the integration cassette comprises or consists of SEQ ID NO:62.
[0160] Also provided is a method of generating a modified cell by site-specific transposition, comprising introducing into a cell containing a stably integrated integration cassette a) a nucleic acid encoding a fusion protein comprising a DNA binding domain and a transposase, said fusion protein being expressed in said cell, and b) a DNA molecule comprising a transposon, said expressed fusion protein integrating said transposon into a TTAA sequence of said stably integrated integration cassette by site-specific transposition, thereby generating said recombinant cell. In some embodiments of the method, the integration cassette comprises or consists of SEQ ID NO:62.
[0161] nucleic acid Also provided herein is a polynucleotide comprising a nucleic acid sequence encoding a fusion protein described herein. In some embodiments, the polynucleotide is isolated.
[0162] The isolated polynucleotides of the disclosure can be produced using (a) recombinant methods, (b) synthetic techniques, (c) purification techniques, and / or (d) combinations thereof, all of which are well known in the art.
[0163] Methods for constructing nucleic acids encoding transposase domains containing N-terminal deletions as described herein, such as PCR-based mutagenesis, are well known in the art or are described herein. Exemplary primers that can be used to construct transposase domains containing N-terminal deletions are shown in Table 3. [Table 7]
[0164] The fusions of the present invention can be produced using any suitable method known in the art or described herein.
[0165] The isolated polynucleotides of the present disclosure, e.g., RNA, cDNA, genomic DNA, or any combination thereof, can be obtained from biological sources using any number of cloning methodologies known to those of skill in the art. In some embodiments, oligonucleotide probes that selectively hybridize to the polynucleotides of the present disclosure under stringent conditions are used to identify the desired sequence within a cDNA or genomic DNA library.
[0166] Methods for amplifying RNA or DNA are well known in the art and can be used in accordance with the present disclosure without undue experimentation, based on the teachings and guidance provided herein. Known methods of DNA or RNA amplification include the polymerase chain reaction (PCR) and related amplification processes (see, e.g., U.S. Pat. Nos. 4,683,195; 4,683,202; 4,800,159; 4,965,188 (Mullis); 4,795,699 and 4,921,794 (Tabor, et al.); 5,142,033 (Innis); 5,122,464 (Wilson, et al.); 5,091,310 (Innis); 5,066,584 (Gyllensten, et al.); 4,889,818 (Gelfand, et al.); 4,994,370 (Silver, et al.)). Nos. 4,766,067 (Biswas); 4,656,134 (Ringold)), and RNA-mediated amplification (U.S. Pat. No. 5,130,238 (Malek, et al.), trade name NASBA), in which antisense RNA against a target sequence is used as a template for double-stranded DNA synthesis, the contents of which are incorporated herein by reference in their entireties (see, e.g., Ausubel, supra, or Sambrook, supra).
[0167] For example, the sequences of the polynucleotides of the present disclosure and related genes can be directly amplified from genomic DNA or cDNA libraries using polymerase chain reaction (PCR) technology. PCR and other in vitro amplification methods may also be useful, for example, to clone nucleic acid sequences encoding proteins to be expressed, to allow nucleic acids to be used as probes to detect the presence of desired mRNA in a sample for nucleic acid sequencing, or for other purposes. Examples of techniques sufficient to instruct the skilled artisan in in vitro amplification methods can be found in Berger (supra), Sambrook (supra), and Ausubel (supra), as well as Mullis, et al., U.S. Patent No. 4,683,202 (1987); and Innis, et al., PCR Protocols A Guide to Methods and Applications, Eds., Academic Press Inc., San Diego, Calif (1990). Commercially available kits for genomic PCR amplification are well known in the art. See, for example, Advantage-GC Genomic PCR Kit (Clontech). Additionally, for example, the T4 gene 32 protein (Boehringer Mannheim) can be used to improve yields of long PCR products.
[0168] The polynucleotides of the present disclosure can also be prepared by direct chemical synthesis according to known methods (see, for example, Ausubel, et al., supra). Chemical synthesis generally produces a single-stranded oligonucleotide, which can be converted into double-stranded DNA by hybridization with a complementary sequence or by polymerization with a DNA polymerase using the single strand as a template. Those skilled in the art will understand that chemical synthesis of DNA may be limited to sequences of about 100 bases or more, while longer sequences can be obtained by ligating shorter sequences.
[0169] Expression vectors and host cells The present disclosure also relates to vectors comprising the polynucleotides of the present disclosure, host cells genetically engineered with the recombinant vectors, and the production of at least one protein scaffold by recombinant techniques well known in the art, see, e.g., Sambrook, et al. (supra); Ausubel, et al., each of which is incorporated herein by reference in its entirety.
[0170] Polynucleotide can be optionally linked to a vector that contains a selectable marker for propagation in host.Generally, plasmid vector is introduced into precipitate such as calcium phosphate precipitate or into a complex with charged lipid.If vector is virus, it can be packaged in vitro using suitable packaging cell line and then transduced into host cell.
[0171] The DNA insert must be operably linked to a suitable promoter. In some embodiments, the promoter is the EF-1α promoter. The expression construct further contains sites for transcription initiation, termination, and, within the transcribed region, a ribosome binding site for translation. The coding portion of the mature transcript expressed by the construct preferably includes a translation initiation codon and a termination codon (e.g., UAA, UGA, or UAG) appropriately positioned at the beginning and end of the mRNA to be translated, with UAA and UAG being preferred for mammalian or eukaryotic cell expression.
[0172] The expression vector preferably, but optionally, contains at least one selectable marker, such as, for example, ampicillin, zeocin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / geneticin (neo gene), DHFR (encoding dihydrofolate reductase and conferring resistance to methotrexate), mycophenolic acid, or glutamine synthetase (GS, U.S. Pat. Nos. 5,122,464; 5,770,359; 5,827,739), blasticidin (bsd gene), resistance genes for eukaryotic cell culture, as well as ampicillin, zeocin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / geneticin (neo gene), DHFR (encoding dihydrofolate reductase and conferring resistance to methotrexate), mycophenolic acid, or glutamine synthetase (GS, U.S. Pat. Nos. 5,122,464; 5,770,359; 5,827,739), blasticidin (bsd gene), resistance genes for culturing E. coli and other bacteria or prokaryotes. Examples of suitable host cells include, but are not limited to, bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / geneticin (neo gene), kanamycin, spectinomycin, streptomycin, carbenicillin, bleomycin, erythromycin, polymyxin B, or tetracycline resistance genes (the above patents are incorporated herein by reference in their entireties). Appropriate media and conditions for the above-mentioned host cells are well known in the art. Suitable vectors will be readily apparent to those skilled in the art. Introduction of the vector construct into the host cell can be accomplished by calcium phosphate transfection, DEAE-dextran mediated transfection, cationic lipid mediated transfection, electroporation, transduction, infection, or other well-known methods. Such methods are described in the art, such as Sambrook (supra), chapters 1-4 and 16-18; Ausubel (supra), chapters 1, 9, 13, 15, 16.
[0173] The expression vector preferably includes at least one selectable cell surface marker for isolating cells modified by the compositions and methods of the present disclosure, but this is optional. The selectable cell surface markers of the present disclosure include surface proteins, glycoproteins, or groups of proteins that distinguish one cell or a subset of cells from another defined subset of cells. Preferably, the selectable cell surface marker distinguishes those cells modified by the compositions or methods of the present disclosure from cells that are not modified by the compositions or methods of the present disclosure. Such cell surface markers include, but are not limited to, for example, "cluster of designation" or "classification determinant" proteins (often abbreviated as "CD"), such as, for example, truncated or full-length forms of CD19, CD271, CD34, CD22, CD20, CD33, CD52, or any combination thereof. Cell surface markers further include the suicide gene marker RQR8 (Philip B et al. Blood. 2014 Aug 21;124(8):1277-87).
[0174] The expression vector preferably, but optionally, includes at least one selectable drug resistance marker for isolating cells modified by the compositions and methods of the present disclosure. The selectable drug resistance markers of the present disclosure can include wild-type or mutant Neo, DHFR, TYMS, FRANCF, RAD51C, GCS, MDR1, ALDH1, NKX2.2, or any combination thereof.
[0175] Those skilled in the art are familiar with the many expression systems available for expressing the nucleic acid encoding the protein of the present disclosure.Alternatively, the nucleic acid of the present disclosure can be expressed in a host cell by turning on (by manipulation) in a host cell containing endogenous DNA encoding the protein scaffold of the present disclosure.Such methods are well known in the art, for example, as described in U.S. Patent Nos. 5,580,734, 5,641,670, 5,733,746, and 5,733,761, the entirety of which is incorporated herein by reference.
[0176] Exemplary cell cultures useful for producing protein scaffolds, specified portions or variants thereof, are bacterial, yeast, and mammalian cells, which are well known in the art. Mammalian cell systems are often in the form of monolayers of cells, although suspensions or bioreactors of mammalian cells can also be used. A number of suitable host cell lines capable of expressing intact glycosylated proteins have been developed in the art and include COS-1 (e.g., ATCC CRL 1650), COS-7 (e.g., ATCC CRL-1651), HEK293, BHK21 (e.g., ATCC CRL-10), CHO (e.g., ATCC CRL 1610) and BSC-1 (e.g., ATCC CRL-26) cell lines, Cos-7 cells, CHO cells, hep G2 cells, P3X63Ag8.653, SP2 / 0-Ag14, 293 cells, HeLa cells, etc., which are readily available, for example, from the American Type Culture Collection (Manassas, Va.) (www.atcc.org). Preferred host cells include cells of lymphoid origin, such as myeloma and lymphoma cells. Particularly preferred host cells are P3X63Ag8.653 cells (ATCC Deposit No. CRL-1580) and SP2 / 0-Ag14 cells (ATCC Deposit No. CRL-1851). In a preferred embodiment, the recombinant cell is a P3X63Ab8.653 or SP2 / 0-Ag14 cell.
[0177] Expression vectors for these cells can include, for example, but are not limited to, one or more of the following expression control sequences: an origin of replication; a promoter, such as a late or early SV40 promoter, a CMV promoter (U.S. Pat. Nos. 5,168,062 and 5,385,839), an HSV tk promoter, a pgk (phosphoglycerate kinase) promoter, an EF-1α promoter (U.S. Pat. No. 5,266,491), at least one human promoter; an enhancer and / or a processing information site, such as a ribosome binding site, an RNA splice site, a polyadenylation site (e.g., an SV40 large T Ag polyA addition site), and a transcription termination sequence. See, e.g., Ausubel et al. (supra); Sambrook, et al. (supra). Other cells useful for producing the nucleic acids or proteins of the disclosure are known and / or available, for example, from the American Type Culture Collection's Catalog of Cell Lines and Hybridomas (www.atcc.org), or other known or commercial sources.
[0178] When using eukaryotic host cells, polyadenylation or transcription terminator sequence is typically incorporated into the vector.An example of terminator sequence is the polyadenylation sequence from bovine growth hormone gene.In some embodiments, the polyA sequence is SV40 polyA sequence.
[0179] Sequences for accurate splicing of the transcript can also be included. An example of a splicing sequence is the VP1 intron from SV40 (Sprague, et al., J. Virol. 45:773-781 (1983)). Additionally, gene sequences for controlling replication in host cells can be incorporated into the vector, as is well known in the art.
[0180] The plasmid constructs described herein can be used to deliver nucleic acids encoding the transposase domains or fusion proteins described herein to cells.
[0181] The transposase domains and fusion proteins described herein can also be delivered to cells using mRNA constructs. Thus, in one embodiment, an mRNA sequence encoding the transposase domains or fusion proteins described herein is provided herein. Such mRNA sequences can be delivered to cells using nanoparticles, e.g., lipid nanoparticles. Examples of lipid nanoparticles are described, for example, in International Patent Application Nos. PCT / US2021 / 055876, PCT / US2022 / 017570, U.S. Provisional Patent Application Nos. 63 / 397,268, 63 / 301,855, and 63 / 348,614, each of which is incorporated by reference in its entirety for lipid nanoparticles that can be used to deliver, for example, mRNA constructs encoding the fusion proteins or transposase domains described herein. The mRNA constructs can also be delivered to cells by electroporation or nucleofection. The mRNA can be capped or otherwise modified.
[0182] Cells and modified cells The tandem dimer transposases and fusion proteins described herein can be used in conjunction with transposons to modify cells. The transposon can be a piggyBac™ (PB) transposon. In some embodiments, when the transposon is a PB transposon, the transposase is a piggyBac™ (PB) transposase, a piggyBac-like ((PBL)) transposase, or a Super piggyBac™ (SPB) transposase. Non-limiting examples of PB transposons are described in detail in U.S. Patent Nos. 6,218,182; 6,962,810; 8,399,643, and PCT International Publication WO 2010 / 099296. The transposon can include a nucleic acid encoding a therapeutic protein or agent. Examples of therapeutic proteins include those disclosed in PCT International Publication No. WO 2019 / 173636 and PCT / US2019 / 049816.
[0183] Thus, provided herein are modified cells comprising one or more transposons and one or more tandem dimer transposases or fusion proteins as described herein. The cells and modified cells of the present disclosure can be mammalian cells. Preferably, the cells and modified cells are human cells.
[0184] The cells modified with the tandem dimer transposases described herein can be germline cells or somatic cells. The cells and modified cells of the present disclosure can be immune cells, such as lymphoid progenitor cells, natural killer (NK) cells, T lymphocytes (T cells), stem memory T cells (T SCM cells), central memory T cells (T CM), stem cell-like T cells, B lymphocytes (B cells), antigen presenting cells (APCs), cytokine-induced killer (CIK) cells, bone marrow progenitor cells, neutrophils, basophils, eosinophils, monocytes, macrophages, platelets, erythrocytes, red blood cells (RBCs), megakaryocytes, or osteoclasts. The modified cells can be differentiated, undifferentiated, or immortalized. The modified undifferentiated cells can be stem cells. The modified undifferentiated cells can be induced pluripotent stem cells. The modified cells can be T cells, hematopoietic stem cells, natural killer cells, macrophages, dendritic cells, monocytes, megakaryocytes, or osteoclasts. The modified cells can be modified while the cells are quiescent in an activated state and resting in interphase, prophase, metaphase, anaphase, or telophase. The modified cells can be sorted in bulk, fresh, lyophilized, into subpopulations, from whole blood, leukapheresis, or from immortalized cell lines. Detailed descriptions for isolating cells from leukapheresis products or blood are disclosed in PCT International Publication No. WO 2019 / 173636 and PCT / US2019 / 049816.
[0185] The disclosed methods can modify and / or produce a population of modified T cells, where at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, or any percentage therebetween, of the modified T cells in the population are stem memory T cells (T SCM ) or T SCMThe one or more cell surface markers of the CD45-like cells may include one or more of CD62L, CD45RA, CD28, CCR7, CD127, CD45RO, CD95, CD95, and IL-2Rβ. The cell surface markers may include one or more of CD45RA, CD95, IL-2Rβ, CCR7, and CD62L.
[0186] The present disclosure provides a method for expressing a CAR on the surface of a cell. The method includes: (a) obtaining a cell population; (b) contacting the cell population with a composition comprising a CAR or a sequence encoding a CAR under conditions sufficient to translocate the CAR through the cell membrane of at least one cell of the cell population, thereby generating a modified cell population; (c) culturing the modified cell population under conditions suitable for the incorporation of the sequence encoding the CAR; and (d) expanding and / or selecting at least one cell from the modified cell population that expresses the CAR on the cell surface. A more detailed description of methods for expressing a CAR on the surface of a cell is disclosed in PCT International Publication WO 2019 / 049816 and PCT / US2019 / 049816.
[0187] The disclosure provides a cell, or population of cells, the cells comprising a composition comprising (a) an inducible transgene construct comprising a sequence encoding an inducible promoter and a sequence encoding a transgene, and (b) a receptor construct comprising a sequence encoding a constitutive promoter and a sequence encoding an exogenous receptor, e.g., a CAR, wherein upon integration of the (a) construct and the (b) construct into the genomic sequence of the cell, the exogenous receptor is expressed, and upon binding of the exogenous receptor to a ligand or antigen, the cell, or population of cells, transduces an intracellular signal that directly or indirectly targets the inducible promoter that controls expression of the inducible transgene (a) and alters gene expression.
[0188] The disclosure further provides compositions comprising the modified, expanded, and selected cell populations of the methods described herein.
[0189] The modified cells (e.g., CAR T cells) of the present disclosure can be further modified to improve their therapeutic capabilities. Alternatively, or in addition, the modified cells can be further modified to render them less sensitive to immune and / or metabolic checkpoints, for example, by blocking and / or diluting certain checkpoint signals (e.g., checkpoint inhibition) that are naturally delivered to the cells within the tumor immunosuppressive microenvironment.
[0190] The modified cells (e.g., CAR T cells) of the present disclosure can be further modified to silence or reduce expression of: (i) one or more gene(s) encoding a receptor(s) for an inhibitory checkpoint signal; (ii) one or more gene(s) encoding an intracellular protein involved in checkpoint signaling; (iii) one or more gene(s) encoding a transcription factor that interferes with the effectiveness of a therapeutic therapy; (iv) one or more gene(s) encoding a cell death or cellular apoptosis receptor; (v) one or more gene(s) encoding a metabolic sensing protein; (vi) one or more gene(s) encoding a protein that confers sensitivity to cancer therapies, including monoclonal antibodies; and / or (vii) one or more gene(s) encoding a growth advantage. Non-limiting examples of genes that can be modified to silence or reduce expression or suppress function include, but are not limited to, exemplary inhibitory checkpoint signals, intracellular proteins, transcription factors, cell death or cell apoptosis receptors, metabolic sensing proteins, proteins that confer sensitivity to cancer therapeutics, and growth advantage factors disclosed in PCT International Publication WO 2019 / 173636.
[0191] The modified cells (e.g., CAR T cells) of the present disclosure can be further modified to express modified / chimeric checkpoint receptors. Modified / chimeric checkpoint receptors can include null, decoy, or dominant-negative receptors. Exemplary null, decoy, or dominant-negative intracellular receptors / proteins include, but are not limited to, inhibitory checkpoint signals, transcription factors, cytokines or cytokine receptors, chemokines or chemokine receptors, cell death or apoptosis receptors / ligands, metabolic sensing molecules, proteins that confer sensitivity to cancer therapy, and signaling components downstream of oncogenes or tumor suppressor genes. Non-limiting examples of cytokines, cytokine receptors, chemokines, and chemokine receptors are disclosed in PCT International Publication WO 2019 / 173636.
[0192] Genome modification can include introducing a nucleic acid sequence, a transgene, and / or a genome editing construct into a cell ex vivo, in vivo, in vitro, or in situ to stably integrate the nucleic acid sequence, to transiently integrate the nucleic acid sequence, to create site-specific integration of the nucleic acid sequence, or to create biased integration of the nucleic acid sequence. The nucleic acid sequence can be a transgene.
[0193] Stable chromosomal integration can be random, site-specific, or biased. Without wishing to be bound by theory, it is believed that adding a DNA-binding domain to the tandem dimer transposase described herein improves the site specificity of the transposase.
[0194] Site-specific integration can occur at safe harbor sites. Genomic safe harbor sites can accommodate the integration of new genetic material in a manner that ensures that the newly inserted genetic element functions (e.g., is expressed at therapeutically effective expression levels) and does not cause deleterious changes to the host genome that pose a risk to the host organism. Non-limiting examples of potential genomic safe harbor sites include intronic sequences of the human albumin gene, adeno-associated virus site 1 (AAVS1), the naturally occurring site of AAV viral integration on chromosome 19, the site of the chemokine (CC motif) receptor 5 (CCR5) gene, and the human ortholog of the mouse Rosa26 locus.
[0195] Site-specific transgene integration can occur at a site that disrupts expression of a target gene. Disruption of target gene expression can occur by site-specific integration at introns, exons, promoters, genetic elements, enhancers, suppressors, start codons, stop codons, and response elements. Non-limiting examples of target genes targeted by site-specific integration include TRAC, TRAB, PDI, any immunosuppressive gene, and genes involved in allorejection.
[0196] Site-specific transgene integration can occur at sites that result in enhanced expression of the target gene. Enhancement of target gene expression can occur by site-specific integration at introns, exons, promoters, genetic elements, enhancers, suppressors, start codons, stop codons, and response elements.
[0197] The site-specific transgene integration site can be a non-stable chromosomal integration. The non-stable integration can be a transient non-chromosomal integration, a semi-stable non-chromosomal integration, a semi-persistent non-chromosomal integration, or a non-stable chromosomal integration. The transient non-chromosomal integration can be extrachromosomal or cytoplasmic. In one embodiment, the transient non-chromosomal integration of the transgene is not integrated into a chromosome and the modified genetic material is not replicated during cell division.
[0198] The site-specific transgene integration site can be a modified binding site for the DNA targeting domain in the transposon domain, fusion protein, or tandem dimer described herein.For example, the TTAA target DNA integration site for SPB can be modified to insert an adjacent DNA binding site for the DNA targeting domain that comprises three zinc finger motifs (e.g., the DNA targeting domain comprises or consists of the sequence of SEQ ID NO:57 or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto).For example, the DNA targeting domain that comprises three zinc finger motifs is believed to bind to the DNA sequence GCGTGGGCG (SEQ ID NO:60).Therefore, it is believed that introducing two copies of SEQ ID NO:60 adjacent to the TTAA target integration site for SPB improves the site-specific integration of the SPB transposase domain that comprises the DNA targeting domain that comprises three zinc finger motifs. The two sequences of SEQ ID NO:60 are oriented in reverse (5') and complementary (3') orientation.
[0199] In some embodiments, provided herein is a polynucleotide that includes, from 5' to 3', a reverse sequence of the target site relative to the DNA targeting domain, a first spacer, a TTAA target incorporation site relative to the SPB, a second spacer, and a complement of the sequence of the target site relative to the DNA targeting domain. In some embodiments, the first spacer and the second spacer have the same length. In some embodiments, the first and / or second spacer are 3 bp long. In some embodiments, the first and / or second spacer are 4 bp long. In some embodiments, the first and / or second spacer are 5 bp long. In some embodiments, the first and / or second spacer are 6 bp long. In some embodiments, the first and / or second spacer are 7 bp long. In some embodiments, the first and / or second spacer are 8 bp long. In some embodiments, the first and / or second spacer are 9 bp long. In some embodiments, the first and / or second spacer are 10 bp long.
[0200] Exemplary sequences of polynucleotides comprising, from 5' to 3', the reverse sequence of the target site for the DNA targeting domain comprising three zinc finger motifs, a first spacer, a TTAA target incorporation site for SPB, a second spacer, and the complement of the sequence of the target site for the three zinc finger motif DNA targeting domain are set forth in SEQ ID NOs: 61-64. The lengths of the first and second spacers in SEQ ID NOs: 61-64 are 8 bp, 7 bp, 6 bp, and 5 bp, respectively, with the reverse and complement of the target site for the DNA targeting domain underlined and the TTAA sequence in bold: [ka]
[0201] The modified target site can be introduced into a cell or cell line to facilitate targeted genomic recombination. For example, a cell line that has been engineered to contain a modified target site for an SPB or PBx provided herein can be transfected with a transposon that contains the SPB or PBx, plus donor DNA, such that the donor DNA is inserted into the modified target site. In some embodiments, the cell line is a T cell line. In some embodiments, the modified target sequence is introduced into a highly expressed genomic region. In certain embodiments, a cell line is provided herein that comprises a nucleic acid sequence stably integrated into its genome sequence that includes, from 5' to 3', the reverse sequence of the target site for the DNA targeting domain containing three zinc finger motifs, a first spacer, a TTAA target integration site for the SPB, a second spacer, and the complement of the sequence of the target site for the three zinc finger motif DNA targeting domain. In some embodiments, the cell line comprises a sequence of any one of SEQ ID NOs: 61-64, stably integrated into its genome. In some embodiments, the cells are in vitro cells, eg, cells in cell culture.
[0202] For the DNA binding domain that comprises TALEN, the target site is determined by the sequence of TALEN.Those skilled in the art can modify TALEN sequence to achieve desired target specificity.Methods for engineered zinc finger nucleases that bind to specific targets are described, for example, in Sander et al., Nat Methods. 2011 Jan;8(1): 67-69.
[0203] The genomic modification can be a non-stable chromosomal integration of the transgene. The integrated transgene can be silenced, removed, truncated or further modified.
[0204] In some embodiments, the transposase domains, fusion proteins, and tandem dimers provided herein have better transposase efficacy than their wild-type counterparts. Transposase activity can be measured by any suitable assay known in the art, such as split GFP assay, as described herein. For example, the transposase domains, fusion proteins, and tandem dimer complexes provided herein can have on-target genome integration activity comparable to their wild-type counterparts, but can have reduced off-target genome integration activity compared to their wild-type counterparts.
[0205] In some embodiments, a transposase domain comprising an N-terminal deletion and a DNA targeting domain provided herein has a ratio of on-target activity to off-target activity that is at least 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500-fold, at least about 550-fold, at least about 600-fold, at least about 650-fold, at least about 700-fold, at least about 750-fold, at least about 800-fold, at least about 850-fold, at least about 900-fold, at least about 950-fold, or at least about 1000-fold greater than a wild-type transposase domain.
[0206] In some embodiments, a transposase domain comprising a DNA targeting domain inserted into the N-terminal region of a transposase domain provided herein has a ratio of on-target activity to off-target activity that is at least 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500-fold, at least about 550-fold, at least about 600-fold, at least about 650-fold, at least about 700-fold, at least about 750-fold, at least about 800-fold, at least about 850-fold, at least about 900-fold, at least about 950-fold, or at least about 1000-fold greater than a wild-type transposase domain.
[0207] In certain embodiments, the modified cells are used therapeutically in adoptive cell therapy.
[0208] An adoptive cell composition that is "universally" safe for administration to any patient (not just the patient of origin) requires significant reduction or elimination of alloreactivity. To this end, the cells (e.g., allogeneic cells) of the present disclosure can be modified to prevent expression or function of T cell receptors (TCRs) and / or certain classes of major histocompatibility complexes (MHCs). TCRs mediate graft-versus-host (GvH) reactions, while MHCs mediate host-versus-graft (HvG) reactions. In preferred embodiments, any expression and / or function of TCRs is eliminated to prevent T cell-mediated GvH, which can cause the death of the subject. Thus, in preferred embodiments, the present disclosure provides a pure TCR-negative allogeneic T cell composition (e.g., each cell of the composition expresses at a level that is either undetectable or so low as to be nonexistent).
[0209] The expression and / or function of MHC class I (MHC-I, specifically HLA-A, HLA-B, and HLA-C) is reduced or eliminated, preventing HvG and thus improving the engraftment of cells in the subject. Improved engraftment results in longer persistence of cells and therefore a longer therapeutic window from the subject. Specifically, the expression and / or function of the structural element of MHC-I, beta-2-microglobulin (B2M), is reduced or eliminated. Non-limiting examples of guide RNAs (gRNAs) for targeting and deleting MHC activators are disclosed in PCT Application No. PCT / US2019 / 049816.
[0210] A detailed description of genetic modifications of endogenous sequences encoding non-naturally occurring chimeric stimulating receptors, TCR-alpha (TCR-α), TCR-beta (TCR-β), and / or beta-2-microglobulin (β2M), and non-naturally occurring polypeptides, including HLA class I histocompatibility antigens, alpha chain E (HLA-E) polypeptides, is disclosed in PCT Application No. PCT / US2019 / 049816.
[0211] Under normal conditions, full T cell activation depends on the engagement of the TCR in combination with a second signal mediated by one or more costimulatory receptors (e.g., CD28, CD2, 4-1BBL) to boost the immune response. However, in the absence of the TCR, T cell proliferation is severely reduced upon stimulation with standard activation / stimulation reagents, including agonistic anti-CD3 mAbs. Thus, the present disclosure provides a non-naturally occurring chimeric stimulatory receptor (CSR) comprising: (a) an ectodomain comprising an activation component, said activation component being isolated or derived from a first protein; (b) a transmembrane domain; and (c) an endodomain comprising at least one signal transduction domain, said at least one signal transduction domain being isolated or derived from a second protein, said first protein and said second protein being non-identical.
[0212] The activation component can include a portion of one or more of a T cell receptor (TCR), a TCR complex, a TCR co-receptor, a TCR co-stimulatory protein, a TCR inhibitory protein, a cytokine receptor, and a chemokine receptor to which an agonist of the activation component binds. The activation component can include the CD2 extracellular domain, or a portion thereof, to which an agonist binds.
[0213] The signal transduction domain may comprise one or more of a human signal transduction domain component, a T cell receptor (TCR), a component of a TCR complex, a component of a TCR co-receptor, a component of a TCR co-stimulatory protein, a component of a TCR inhibitory protein, a cytokine receptor, and a chemokine receptor. The signal transduction domain may comprise a CD3 protein, or a portion thereof. The CD3 protein may comprise a CD3 zeta protein, or a portion thereof.
[0214] The endodomain can further comprise a cytoplasmic domain. The cytoplasmic domain can be isolated from or derived from a third protein. The first protein and the third protein can be identical. The ectodomain can further comprise a signal peptide. The signal peptide can be derived from a fourth protein. The first protein and the fourth protein can be identical. The transmembrane domain can be isolated from or derived from a fifth protein. The first protein and the fifth protein can be identical.
[0215] The present disclosure also provides a non-naturally occurring chimeric stimulating receptor (CSR) in which the ectodomain comprises a modification. The modification can include a mutation or truncation of the amino acid sequence of the activating component or the first protein when compared to the wild-type sequence of the activating component or the first protein. The mutation or truncation of the amino acid sequence of the activating component can include a mutation or truncation of the CD2 extracellular domain or a portion thereof to which the agonist binds. The mutation or truncation of the CD2 extracellular domain can reduce or eliminate binding to naturally occurring CD58.
[0216] The present disclosure provides a nucleic acid sequence encoding any of the CSRs disclosed herein.The present disclosure provides a transposon or vector comprising a nucleic acid sequence encoding any of the CSRs disclosed herein.
[0217] The present disclosure provides a cell comprising any of the CSRs disclosed herein.The present disclosure provides a cell comprising a nucleic acid sequence encoding any of the CSRs disclosed herein.The present disclosure provides a cell comprising a vector comprising a nucleic acid sequence encoding any of the CSRs disclosed herein.The present disclosure provides a cell comprising a transposon comprising a nucleic acid sequence encoding any of the CSRs disclosed herein.
[0218] The present disclosure provides a composition comprising any of the CSRs disclosed herein. The present disclosure provides a composition encoding a nucleic acid sequence encoding any of the CSRs disclosed herein. The present disclosure provides a composition comprising a vector comprising a nucleic acid sequence encoding any of the CSRs disclosed herein. The present disclosure provides a composition comprising a transposon comprising a nucleic acid sequence encoding any of the CSRs disclosed herein. The present disclosure provides a composition comprising a modified cell disclosed herein or a composition comprising a plurality of modified cells disclosed herein.
[0219] Methods for site-specific gene integration are also provided herein. The transposon domains and fusion proteins provided herein can be used to deliver transgenes to cells and integrate the transgenes into target sites. The target site can be, for example, a genomic safe harbor, i.e., a genomic site where the transgene can be integrated to ensure that the transgene functions as expected and does not cause changes in the host genome DNA sequence. In some embodiments, the target site is a repetitive element, such as a LINE-1 or ALU sequence. Repetitive elements do not code for gene products, making insertion less likely to cause deleterious changes in gene expression profile in cells. Within one repetitive element, there can be one, two, or more target sites. In some embodiments, the target site is located within an intron (e.g., an intron of the PAH gene).
[0220] Site-specific integration can be used in vitro or in vivo. One example of an in vivo application is gene therapy, which involves delivering a transgene into the genomic DNA of a cell.
[0221] Formulations, Dosages, and Modes of Administration The present disclosure provides formulations, doses, and methods for administering the compositions and cells described herein. In one aspect, provided herein is a pharmaceutical composition comprising a tandem dimer transposase or fusion protein described herein and a pharma- ceutically acceptable carrier. In another aspect, provided herein is a pharmaceutical composition comprising a modified cell described herein and a pharma- ceutically acceptable carrier.
[0222] The disclosed compositions and pharmaceutical compositions may include at least one of any suitable auxiliary agent, such as, but not limited to, diluents, binders, stabilizers, buffers, salts, lipophilic solvents, preservatives, adjuvants, etc. Pharmaceutically acceptable auxiliary agents are preferred. Non-limiting examples of such sterile solutions and methods for preparing such sterile solutions are well known in the art, such as, but not limited to, Gennaro, Ed., Remington's Pharmaceutical Sciences, 18th Edition, Mack Publishing Co. (Easton, Pa.) 1990, and "Physician's Desk Reference", 52nd ed., Medical Economics (Montvale, NJ) 1998. Pharmaceutically acceptable carriers suitable for the mode of administration, solubility and / or stability of the protein scaffold, fragment or variant composition, as known in the art or described herein, can be routinely selected.
[0223] Non-limiting examples of pharmaceutical excipients and additives suitable for use include proteins, peptides, amino acids, lipids, and carbohydrates (e.g., sugars including monosaccharides, disaccharides, trisaccharides, tetrasaccharides, and oligosaccharides; derivatized sugars such as alditols, aldonic acids, esterified sugars; and polysaccharides or sugar polymers), which may be present alone or in combination and may comprise 1-99.99% by weight or volume, alone or in combination. Non-limiting examples of protein excipients include serum albumins, such as human serum albumin (HSA), recombinant human albumin (rHA), gelatin, casein, and the like. Exemplary amino acids / protein components that may also function in a buffering capacity include alanine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, aspartame, and the like. One preferred amino acid is glycine.
[0224] Non-limiting examples of carbohydrate excipients suitable for use include monosaccharides such as fructose, maltose, galactose, glucose, D-mannose, sorbose, etc.; disaccharides such as lactose, sucrose, trehalose, cellobiose, etc.; polysaccharides such as raffinose, melezitose, maltodextrin, dextran, starch, etc.; and alditols such as mannitol, xylitol, maltitol, lactitol, xylitol, sorbitol (glucitol), myo-inositol, etc. Preferably, the carbohydrate excipient is mannitol, trehalose, and / or raffinose.
[0225] The composition can also include buffer or pH adjuster, and typically buffer is a salt prepared from organic acid or base.Exemplary buffer includes organic acid, such as citric acid, ascorbic acid, gluconic acid, carbonic acid, tartaric acid, succinic acid, acetic acid, or phthalic acid salt; Tris, tromethamine hydrochloride, or phosphate buffer.Preferred buffer is organic acid salt, such as citrate.
[0226] Additionally, the disclosed compositions can include polymeric excipients / additives such as polyvinylpyrrolidone, Ficoll (polymeric sugar), dextrates (e.g., cyclodextrins, such as 2-hydroxypropyl-β-cyclodextrin), polyethylene glycol, flavoring agents, antimicrobial agents, sweeteners, antioxidants, antistatic agents, surfactants (e.g., polysorbates, e.g., TWEEN® 20 and TWEEN® 80), lipids (e.g., phospholipids, fatty acids), steroids (e.g., cholesterol), and chelating agents (e.g., EDTA).
[0227] Many well-known and developed modes can be used to administer a therapeutically effective amount of a composition or pharmaceutical composition disclosed herein. Non-limiting examples of modes of administration include bolus, buccal, injection, intra-articular, intrabronchial, intraabdominal, intravesical, intracartilaginous, intracavitary, intraperitoneal, intracerebellar, intraventricular, intracolonic, intracervical, intragastric, intrahepatic, intralesional, intramuscular, intramyocardial, intranasal, intraocular, intraosseous, intraosteal, intrapelvic, intrapericardial, intraperitoneal, intrapleural, intraprostatic, intrapulmonary, intrarectal, intrarenal, intraretinal, intraspinal, intrasynovial, intrathoracic, intrauterine, intratumoral, intravenous, intravesical, oral, parenteral, rectal, sublingual, subcutaneous, transdermal, or intravaginal means. In a preferred embodiment, the composition comprising the modified cells described herein is administered intravenously, for example, by intravenous injection.
[0228] The compositions of the present disclosure can be prepared for parenteral (subcutaneous, intramuscular, or intravenous) or any other administration, particularly in the form of a solution or suspension. For parenteral administration, the compositions disclosed herein can be formulated as a solution, suspension, emulsion, particles, powder, or lyophilized powder, either in combination with a pharma- ceutically acceptable parenteral vehicle or provided separately. Preparations for parenteral administration may contain sterile water or saline, polyalkylene glycols, such as polyethylene glycol, vegetable oils, hydrogenated naphthalene, and the like, as common excipients. Aqueous or oily suspensions for injection can be prepared by using suitable emulsifiers or wetting agents and suspending agents according to well-known methods. Agents for injection or infusion can be non-toxic, parenterally administrable diluents, such as aqueous solutions, sterile injectable solutions, or suspensions in solvents. Vehicles or solvents suitable for use include water, Ringer's solution, isotonic saline, and the like, and sterile non-volatile oils can be used as common solvents or suspending agents. For these purposes, any type of non-volatile oil and fatty acid can be used, including natural, synthetic or semi-synthetic fatty oils or fatty acids; natural, synthetic or semi-synthetic mono-, di- or triglycerides. Parenteral administration is well known in the art and includes, but is not limited to, conventional injection means, the gas pressure needleless injection device described in U.S. Patent No. 5,851,198, and the laser perforator device described in U.S. Patent No. 5,839,446.
[0229] It may be desirable to deliver the disclosed compounds to a subject over an extended period of time, for example, from one week to one year from a single dose. A variety of sustained release, depot, or implant dosage forms can be utilized. For example, the dosage form can contain a pharma- ceutically acceptable, non-toxic salt of a compound that has low solubility in body fluids, such as (a) an acid addition salt with a polybasic acid, such as phosphoric acid, sulfuric acid, citric acid, tartaric acid, tannic acid, pamoic acid, alginic acid, polyglutamic acid, naphthalene mono- or disulfonic acid, polygalacturonic acid, and the like; (b) a salt with a polyvalent metal cation, such as zinc, calcium, bismuth, barium, magnesium, aluminum, copper, cobalt, nickel, cadmium, and the like, or with an organic cation formed, for example, from N,N'-dibenzyl-ethylenediamine or ethylenediamine; or (c) a combination of (a) and (b), such as zinc tannate salt. Additionally, the disclosed compounds, or preferably relatively insoluble salts, such as those just mentioned, can be formulated in gels suitable for injection, such as aluminum monostearate gels, for example with sesame oil. Particularly preferred salts are zinc salts, zinc tannate salts, pamoate salts, and the like. Another type of injectable sustained release depot formulation contains the compound or salt dispersed for encapsulation in a slowly degrading non-toxic, non-antigenic polymer, such as polylactic acid / polyglycolic acid polymers, as described, for example, in U.S. Pat. No. 3,773,919. The compounds, or preferably relatively insoluble salts, such as those just mentioned, can also be formulated in cholesterol matrix silastic pellets, particularly for use in animals. Additional sustained release depot or implant formulations, such as gas or liquid liposomes, are well known in the literature (U.S. Pat. No. 5,770,222 and "Sustained and Controlled Release Drug Delivery Systems", JR Robinson ed., Marcel Dekker, Inc., NY, 1978).
[0230] Treatment In another aspect, provided herein is a method of treating a disease or disorder in a subject, the method comprising administering to the subject a composition comprising modified cells as described herein. The terms "subject" and "patient" are used interchangeably herein. In a preferred embodiment, the patient is a human.
[0231] The modified cells may be allogeneic or autologous to the patient. In some preferred embodiments, the modified cells are allogeneic cells. In some embodiments, the modified cells are autologous T cells or modified autologous CAR T cells. In some preferred embodiments, the modified cells are allogeneic T cells or modified allogeneic CAR T cells.
[0232] In some embodiments, the disease or disorder treated according to the methods described herein is cancer. In some embodiments, the treatment methods described herein can slow the progression of cancer and / or reduce tumor burden.
[0233] The dose of the pharmaceutical composition that may be administered to a subject may vary depending on known factors, such as the pharmacodynamic characteristics of the particular agent and its mode and route of administration; the age, health, and weight of the recipient; the nature and extent of the symptoms, type of concurrent treatment, frequency of treatment, and the desired effect.
[0234] In embodiments where the composition that can be administered to a subject in need thereof is the modified cells disclosed herein, about 1×10 3 ~Approx. 1×10 4 cells; approximately 1 x 10 4 ~Approx. 1×10 5 cells; approximately 1 x 10 5 ~Approx. 1×10 6 cells; approximately 1 x 10 6 ~Approx. 1×10 7 cells; approximately 1 x 10 7 ~Approx. 1×10 8 cells; approximately 1 x 10 8 ~Approx. 1×10 9cells; approximately 1 x 10 9 ~Approx. 1×10 10 cells; approximately 1 x 10 10 ~Approx. 1×10 11 cells; approximately 1 x 10 11 ~Approx. 1×10 12 cells; approximately 1 x 10 12 ~Approx. 1×10 13 cells; approximately 1 x 10 13 ~Approx. 1×10 14 cells; approximately 1 x 10 14 ~Approx. 1×10 15 cells; approximately 1 x 10 15 ~Approx. 1×10 16 cells; approximately 1 x 10 16 ~Approx. 1×10 17 cells; approximately 1 x 10 17 ~Approx. 1×10 18 cells; approximately 1 x 10 18 ~Approx. 1×10 19 cells; or approximately 1 x 10 19 ~Approx. 1×10 20 In some embodiments, the cells are administered in an amount of about 5×10 6 ~Approx. 25×10 6 The cells are administered in doses of 100 cells.
[0235] In other embodiments, the dose of cells may depend on the body weight of the human, e.g., about 1×10 cells per kg of subject body weight. 3 ~Approx. 1×10 4 cells; approximately 1 x 10 4 ~Approx. 1×10 5 cells; approximately 1 x 10 5 ~Approx. 1×10 6 cells; approximately 1 x 10 6 ~Approx. 1×10 7 cells; approximately 1 x 10 7 ~Approx. 1×10 8 cells; approximately 1 x 10 8 ~Approx. 1×10 9 cells; approximately 1 x 10 9 ~Approx. 1×10 10 cells; approximately 1 x 10 10 ~Approx. 1×10 11 cells; approximately 1 x 10 11 ~Approx. 1×1012 cells; approximately 1 x 10 12 ~Approx. 1×10 13 cells; approximately 1 x 10 13 ~Approx. 1×10 14 cells; approximately 1 x 10 14 ~Approx. 1×10 15 cells; approximately 1 x 10 15 ~Approx. 1×10 16 cells; approximately 1 x 10 16 ~Approx. 1×10 17 cells; approximately 1 x 10 17 ~Approx. 1×10 18 cells; approximately 1 x 10 18 ~Approx. 1×10 19 cells; or approximately 1 x 10 19 ~Approx. 1×10 20 The cells can be administered.
[0236] A more detailed description of the pharma- ceutically acceptable excipients, formulations, dosages, and methods of administration of the disclosed compositions and pharmaceutical compositions is disclosed in PCT International Publication WO 2019 / 049816.
[0237] The transposon domains and fusion proteins provided herein can be used to deliver gene therapy. Gene therapy typically involves the delivery of a transgene into the genomic DNA of a cell. Typically, the transgene replaces a mutated or otherwise inappropriately expressed cell. The fusion proteins, transposase domains, and complexes described herein can be used to deliver therapeutic transgenes to cells and integrate the transgene at a target site. In some embodiments, the method of treatment includes introducing into a cell a fusion protein and a transposon according to any one of claims 1 to 13, the transposon including, from 5' to 3', a 5' ITR, a transgene, and a 3' ITR.
[0238] kit In another aspect, provided herein is a kit comprising a cell line engineered to comprise a modified target site for SPB or PBx provided herein in its genome, preferably in a highly expressed genomic region.The kit can further comprise a composition comprising one or more SPB or PBx transposase domains or fusion proteins described herein.In some embodiments, the cell line is a T cell line.
[0239] definition As used throughout this disclosure, unless the context clearly indicates otherwise, the singular forms "a," "and," and "the" include plural referents. Thus, for example, reference to "a method" includes a plurality of such methods, reference to "a dose" includes reference to one or more doses and equivalents thereof known to those skilled in the art, and so forth.
[0240] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, and depends in part on the method of measuring or determining the value, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more standard deviations. Alternatively, "about" can mean a range of up to 20%, or up to 10%, or up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2-fold, of the value. Unless otherwise stated, when a particular value is described in this application and claims, the term "about" is deemed to mean within an acceptable error range for the particular value.
[0241] The present disclosure provides isolated or substantially isolated polynucleotide or protein compositions. An "isolated" or "purified" polynucleotide or protein, or a biologically active portion thereof, is substantially or essentially free of components that normally accompany or normally interact with the polynucleotide or protein found in its natural environment. Thus, an isolated or purified polynucleotide or protein is substantially free of other cellular materials or culture media when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. Optionally, an "isolated" polynucleotide is free of sequences (optimally protein-encoding sequences) that naturally flank the polynucleotide in the genomic DNA of the organism from which the polynucleotide is derived (i.e., sequences located at the 5' and 3' ends of the polynucleotide). For example, in various embodiments, an isolated polynucleotide can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequence that naturally flanks the polynucleotide in the genomic DNA of the cell from which the polynucleotide is derived. Proteins that are substantially free of cellular material include preparations that have less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of contaminating proteins. When the proteins of the present disclosure, or biologically active portions thereof, are recombinantly produced, optimal media exhibit less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of chemical precursors or non-protein chemicals of interest.
[0242] The present disclosure provides fragments and variants of the disclosed DNA sequences and proteins encoded by these DNA sequences. As used throughout this disclosure, the term "fragment" refers to a portion of a DNA sequence or a portion of an amino acid sequence and thus the protein encoded thereby. A fragment of a DNA sequence, including a coding sequence, can encode a protein fragment that maintains the biological activity of the native protein and thus the DNA recognition or binding activity to a target DNA sequence as described herein. Alternatively, a fragment of a DNA sequence useful as a hybridization probe generally does not encode a protein that maintains biological activity or does not maintain promoter activity. Thus, a fragment of a DNA sequence can range from at least about 20 nucleotides, about 50 nucleotides, about 100 nucleotides, up to the full-length polynucleotide of the present disclosure.
[0243] The nucleic acids or proteins of the disclosure can be constructed by a modular approach, including pre-assembling monomeric and / or repeating units in a target vector that can then be assembled into a final destination vector. The polypeptides of the disclosure can be constructed by a modular approach by pre-assembling repeating units in a target vector that can include repeating monomers of the disclosure and can then be assembled into a final destination vector. The disclosure provides polypeptides produced by the methods, as well as nucleic acid sequences encoding these polypeptides. The disclosure provides host organisms and cells comprising nucleic acid sequences encoding the polypeptides produced by the modular approach.
[0244] The term "comprising" is intended to mean that the compositions and methods include the recited element, but do not exclude other elements. When used to define compositions and methods, "consisting essentially of" shall mean excluding other elements of any essential importance to the combination when used for the intended purpose. Thus, a composition consisting essentially of the elements defined herein does not exclude trace amounts of contaminants, or inert carriers. "Consisting of" shall mean excluding more than trace elements of other ingredients and substantial method steps. The embodiments defined by each of these transitional terms are within the scope of this disclosure.
[0245] As used herein, "expression" refers to the process by which a polynucleotide is transcribed into mRNA and / or the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression may also include splicing of the mRNA in a eukaryotic cell.
[0246] "Gene expression" refers to the conversion of the information contained in a gene into a gene product. A gene product can be the direct transcription product of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, shRNA, microRNA, structural RNA, or any other type of RNA) or a protein produced by translation of an mRNA. Gene products also include RNA that is modified by processes such as capping, polyadenylation, methylation, and editing, as well as proteins that are modified, for example, by methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristyrylation, and glycosylation.
[0247] "Modulation" or "regulation" of gene expression refers to a change in the activity of a gene. Modulation of expression can include, but is not limited to, gene activation and gene repression.
[0248] The term "operatively linked" or its equivalents (e.g., "operably linked" means that two or more molecules are positioned relative to one another such that they can interact to affect a function attributable to one or both molecules, or a combination thereof. In the context of nucleic acids, a promoter can be operably linked to a nucleotide sequence encoding a translocation matrix domain or fusion protein described herein, and results in expression of the nucleotide sequence under the control of the promoter.
[0249] Non-covalently linked components and methods of making and using non-covalently linked components are disclosed. The various components can take a wide variety of different forms as described herein. For example, non-covalently linked (i.e., operably linked) proteins can be used to allow for temporary interactions that avoid one or more problems in the art. The ability of non-covalently linked components, such as proteins, to associate and dissociate allows for functional association only or primarily under circumstances where such association is necessary for the desired activity. The association can have a sufficient duration to allow for the desired effect.
[0250] Disclosed is a method for targeting a protein to a specific locus within the genome of an organism. The method can include providing a DNA localization component and providing an effector molecule, wherein the DNA localization component and the effector molecule can be operably linked via a non-covalent bond.
[0251] A "target site" or "target sequence" is a nucleic acid sequence that defines a portion of a nucleic acid to which a binding molecule binds when sufficient conditions for binding are present.
[0252] The term "nucleic acid" or "oligonucleotide" or "polynucleotide" refers to at least two nucleotides covalently linked to each other. The depiction of a single strand also defines the sequence of the complementary strand. Thus, a nucleic acid can also encompass the complementary strand of a depicted single strand. The nucleic acids of the present disclosure also encompass substantially identical nucleic acids and their complements that retain the same structure or encode the same protein.
[0253] The nucleic acids of the present disclosure can be single-stranded or double-stranded. The nucleic acids of the present disclosure can contain double-stranded sequences even when the majority of the molecule is single-stranded. The nucleic acids of the present disclosure can contain single-stranded sequences even when the majority of the molecule is double-stranded. The nucleic acids of the present disclosure can include genomic DNA, cDNA, RNA, or hybrids thereof. The nucleic acids of the present disclosure can contain combinations of deoxyribonucleotides and ribonucleotides. The nucleic acids of the present disclosure can contain combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine hypoxanthine, isocytosine, and isoguanine. The nucleic acids of the present disclosure can be synthesized to include non-natural amino acid modifications. The nucleic acids of the present disclosure can be obtained by chemical synthesis or by recombinant methods.
[0254] The nucleic acid of the present disclosure may be non-naturally occurring, either in its entirety or in any part thereof. The nucleic acid of the present disclosure may contain one or more mutations, substitutions, deletions, or insertions that are not naturally occurring, making the entire nucleic acid sequence non-naturally occurring. The nucleic acid of the present disclosure may contain one or more duplicated, inverted, or repeated sequences that result in a sequence that is not naturally occurring, making the entire nucleic acid sequence non-naturally occurring. The nucleic acid of the present disclosure may contain modified, artificial, or synthetic nucleotides that are not naturally occurring, making the entire nucleic acid sequence non-naturally occurring.
[0255] Given the degeneracy of the genetic code, multiple nucleotide sequences may encode any particular protein, and all such nucleotide sequences are contemplated herein.
[0256] As used throughout this disclosure, the term "operably linked" refers to the expression of a gene under the control of a spatially connected promoter. The promoter may be located 5' (upstream) or 3' (downstream) of the gene under its control. The distance between the promoter and the gene may be approximately the same as the distance between the promoter and the gene it controls in the gene from which it is derived. Variation in the distance between the promoter and the gene can be accommodated without loss of promoter function.
[0257] As used throughout this disclosure, the term "promoter" refers to a synthetic or naturally derived molecule capable of conferring, activating, or enhancing expression of a nucleic acid in a cell. A promoter may further include one or more specific transcription control sequences that enhance its expression and / or alter its spatial and / or temporal expression. A promoter may include distal enhancer or repressor elements, which can be located thousands of base pairs from the transcription start site. Promoters can be derived from sources including viruses, bacteria, fungi, plants, insects, and animals. A promoter can constitutively or specifically control the expression of a gene with respect to the cell, tissue, or organ in which expression occurs, or with respect to the developmental stage in which expression occurs, or in response to an external stimulus such as a physiological stress, a pathogen, a metal ion, or an inducer. Representative examples of promoters include the bacteriophage T7 promoter, the bacteriophage T3 promoter, the SP6 promoter, the lac operator-promoter, the tac promoter, the SV40 late promoter, the SV40 early promoter, the RSV-LTR promoter, the CMV IE promoter, the EF-1α promoter, the CAG promoter, the SV40 early promoter or the SV40 late promoter, and the CMV IE promoter.
[0258] As used throughout this disclosure, the term "vector" refers to a nucleic acid sequence that contains an origin of replication. The vector may be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. The vector may be a DNA or RNA vector. The vector may be a self-replicating extrachromosomal vector, preferably a DNA plasmid. The vector may contain a combination of amino acids with DNA sequences, RNA sequences, or both DNA and RNA sequences.
[0259] Conservative substitutions of amino acids, i.e., replacing one amino acid with another having similar properties (hydrophilicity, degree and distribution of charged regions), are recognized in the art as typically involving minor changes. These minor changes can be identified, in part, by considering the hydrophilicity index of amino acids, as understood in the art. Kyte et al., J. Mol. Biol. 157: 105-132 (1982). The hydrophilicity index of an amino acid is based on a consideration of its hydrophobicity and charge. Amino acids with similar hydrophilicity indexes can be substituted and still retain protein function. In one embodiment, amino acids with hydrophilicity indexes of ±2 are substituted. The hydrophilicity of amino acids can also be used to identify substitutions that result in proteins that retain biological function. Considering the hydrophilicity of amino acids in a peptide allows the calculation of the maximum local average hydrophilicity of the peptide, a useful index that has been reported to correlate well with antigenicity and immunogenicity. U.S. Patent No. 4,554,101, incorporated herein by reference in its entirety.
[0260] Substitution of amino acids with similar hydrophilicity values can result in peptides that retain biological activity, e.g., immunogenicity. Substitutions can be made with amino acids with hydrophilicity values within ±2 of each other. Both the hydrophobicity index and hydrophilicity value of an amino acid are influenced by the specific side chain of that amino acid. Consistent with that observation, it is understood that amino acid substitutions that are compatible with biological function depend on the relative similarity of the amino acids, particularly their side chains, as revealed by hydrophobicity, hydrophilicity, charge, size, and other properties.
[0261] As used herein, "conservative" amino acid substitutions can be defined as set forth in Table 4, Table 5, and Table 6 below. In some embodiments, the fusion polypeptide and / or nucleic acid encoding such a fusion polypeptide comprises a conservative substitution introduced by modifying the polynucleotide encoding the polypeptide of the present disclosure. Amino acids can be classified according to their physical properties and contribution to the secondary and tertiary structure of proteins. A conservative substitution is the replacement of one amino acid with another amino acid having similar properties. Exemplary conservative substitutions are shown in Table 4. [Table 8]
[0262] Alternatively, conservative amino acids can be grouped as described by Lehninger (Biochemistry, Second Edition; Worth Publishers, Inc. NY, NY (1975), pp. 71-77), as shown in Table 5. [Table 9]
[0263] Alternatively, exemplary conservative substitutions are shown in Table 6. [Table 10]
[0264] The polypeptides and proteins of the present disclosure may be non-naturally occurring, either in their entirety or in any part thereof. The polypeptides and proteins of the present disclosure may contain one or more mutations, substitutions, deletions, or insertions that are not naturally occurring, making the entire amino acid sequence non-naturally occurring. The polypeptides and proteins of the present disclosure may contain one or more duplicated, inverted, or repeated sequences, the resulting sequence of which is not naturally occurring, making the entire amino acid sequence non-naturally occurring. The polypeptides and proteins of the present disclosure may contain modified, artificial, or synthetic amino acids that are not naturally occurring, making the entire amino acid sequence non-naturally occurring.
[0265] As used throughout this disclosure, identity between two sequences can be determined by using a stand-alone executable BLAST engine to blast two sequences (bl2seq), which can be loaded from the National Center for Biotechnology Information (NCBI) ftp site using default parameters (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250; incorporated herein by reference in its entirety). When used in the context of two or more nucleic acid or polypeptide sequences, the term "identical" or "identity" refers to a certain percentage of residues that are the same over a specific region of each of the sequences. In some embodiments, sequence identity is determined over the entire length of the sequence. The percentage can be calculated by optimally aligning the two sequences, comparing the two sequences at a specific region, determining the number of positions where identical residues occur in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the specific region, and multiplying the result by 100 to obtain the percentage of sequence identity. If the two sequences are of different lengths or the alignment produces one or more staggered ends and the particular region being compared contains only a single sequence, the residues of the single sequence are included in the denominator rather than the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be performed manually or using a computer sequence algorithm such as BLAST or BLAST 2.0.
[0266] In certain embodiments, if a sequence has a certain sequence identity (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%) to a certain SEQ ID NO, the sequence and the sequence of the SEQ ID NO have the same length. In certain embodiments, if a sequence has a certain sequence identity (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%) to a certain SEQ ID NO, the sequence and the sequence of the SEQ ID NO differ only by conservative amino acid substitutions.
[0267] As used throughout this disclosure, the term "endogenous" refers to a nucleic acid or protein sequence that is naturally associated with the target gene or host cell into which it is introduced.
[0268] As used throughout this disclosure, the term "exogenous" refers to a naturally occurring nucleic acid, e.g., a DNA sequence, or a nucleic acid or protein sequence that is not naturally associated with a target gene or host cell into which it is introduced, including multiple, naturally occurring copies of a naturally occurring nucleic acid sequence located in a non-naturally occurring genomic location.
[0269] The present disclosure provides a method for introducing a polynucleotide construct comprising a DNA sequence into a host cell. By "introducing" it is intended to present the polynucleotide construct to the cell such that the construct can access the inside of the host cell. The method of the present disclosure does not depend on a particular method for introducing a polynucleotide construct into a host cell, only that the polynucleotide construct can access the inside of a certain cell of the host. Methods for introducing polynucleotide constructs into bacteria, plants, fungi, and animals are well known in the art, including but not limited to stable transformation, transient transformation, and virus-mediated methods. EXAMPLES
[0270] The examples in this section are offered by way of illustration and are not intended to limit the invention.
[0271] Example 1: A series of nested constructs of the N-terminal part of the SPB transposase domain A series of nested deletions of the N-terminal portion of SPB transposase were constructed using PCR-based mutagenesis. A plasmid containing a DNA sequence encoding wild-type SPB transposase with an N-terminal NLS (SEQ ID NO: 24) under the control of the EF-1α promoter was used as a DNA template for PCR-based mutagenesis to generate deletions of 20, 40, 60, 80, 100, or 115 amino acids at the N-terminus of the SPB transposase sequence. Briefly, a forward primer was designed complementary to the downstream sequence adjacent to the C-terminal deletion boundary (SEQ ID NOs: 17-22), and a reverse primer was designed complementary to the upstream amino-terminal NLS sequence (SEQ ID NO: 23). SPB transposase coding fragments were generated using a thermocycler and a Q5 Hotstart kit (NEB Labs) under the conditions shown in Tables 7 and 8 and according to the manufacturer's instructions. [Table 11] [Table 12]
[0272] The crude PCR products were directly treated with the KLD enzyme kit (Grainger) according to the manufacturer's procedure. The KLD enzyme mix contains kinase, ligase, and the restriction enzyme DpnI, resulting in ligated full-length fragments suitable for direct cloning into a plasmid vector. The sizes of the SPB transposase fragments were determined by gel electrophoresis, and these DNA fragments of the desired size were cloned into a plasmid vector. The resulting plasmids were transformed into Zymo DH5a MixAndGo (T3007) complement according to the manufacturer's procedure. The nucleotide sequence of each SPB construct, including the N-terminal deletion, was confirmed by direct Sanger DNA sequencing.
[0273] Example 2: Construction of fusion proteins This example illustrates an exemplary method for constructing a tandem dimeric transposase of the invention using two-fragment Gibson assembly.
[0274] Two fragments were used for Gibson assembly of the plasmid: (1) a plasmid backbone containing the EF1α promoter, NLS sequence, first SPB transposon domain, poly-A signal, key elements for plasmid replication, etc., and (2) a tandem dimer SPB expressing the L3 linker plus the second SPB full-length domain with different codon usage. This fragment is directly provided as a gene blocking fragment. To assemble the plasmid backbone, the wild type SPB plasmid (SEQ ID NO: 24) is amplified with the following primers: forward: tctagaaccggtcatggccg (SEQ ID NO: 25), reverse: GAAGCAGCTCTGGCACATG (SEQ ID NO: 26).
[0275] The insert fragment containing the second SPB transposase domain is provided directly as a double-stranded gene blocking DNA fragment. The sequence of the insert fragment is set forth in SEQ ID NO: 27. The DNA sequence of the assembled product is set forth in SEQ ID NO: 30.
[0276] The amplified region of the template fragment shares a region of complementarity after the C-terminus of the SPB coding sequence with a region located upstream of the 5' end of the second SPB coding sequence, and immediately following 5' exonuclease digestion, polymerase fills in the ins and DNA ligation results in an in-frame fusion of the first transposase domain sequence with the second transposase domain, including an intervening 13 amino acid linker, generating tdSPB.
[0277] To construct fusion proteins of the invention containing partial deletions of the amino terminus of the SPB transposase domain, tdSPB was used as a DNA template in the PCR mutagenesis assay described in Example 1 to generate fusion proteins containing amino-terminal deletions of 20 amino acids, 40 amino acids, 60 amino acids, 80 amino acids, 100 amino acids, or 115 amino acids (SEQ ID NOs: 9-14) in the second transposase domain alone. The two SPB transposase domain sequences have different codon usage in the N-terminal deletion sequence, allowing a forward primer to be designed with complementarity to the second transposase domain coding sequence. The presence of each deletion in the second transposase domain and the integrity of the coding sequence of the first transposase domain were confirmed by Sanger DNA sequencing.
[0278] Example 3: Method for measuring the cleavage activity of SPB transposase domains and fusion proteins The assay is designed to measure the cleavage activity of the transposase domain and fusion proteins containing the transposase domain. In the assay, the transposition matrix domain or fusion proteins containing the first and second transposase domains are co-administered to cells with a reporter transposon construct, in which the transposon contains a DNA nucleotide sequence encoding a non-functional GFP, whose coding sequence is interrupted by an intervening piece of DNA flanked by TTAA sequences and inverted terminal repeat (ITR) sequences of the PB transposon. A schematic diagram of the reporter (GFP cleavage only reporter) is shown in FIG. 11. The TTAA sequence and ITRs serve as recognition sites for the SPB transposase, and if the transposase domain or fusion protein has cleavage activity, the intervening DNA is cleaved, restoring the intact, full-length coding sequence of the GFP gene. Thus, transposase domains and fusion proteins with transposase activity produce GFP positive cells in this assay that can be identified and quantified by FACS.
[0279] In the first experiment, the cleavage activity of SPB transposase domains with various sizes of N-terminal deletions, as described in Example 1, was measured. On day 0, HEK293 cells were seeded in 48-well plates at a density of 70,000 cells / well, DMEM medium supplemented with 10% FBS was added to each well, and the cells were cultured at 37° C. and 5% CO2. On day 1, the medium was aspirated and the cells were resuspended in a buffer containing Jetprime transfection reagent (Polyplus Transfection) according to the manufacturer's instructions. SPB transposase domains and reporter transposon constructs were added at the concentrations per well shown in Table 9. [Table 13]
[0280] Approximately 24 hours later, the cells were resuspended in PBS supplemented with 5% FBS, and the number of cells expressing GFP was determined using flow cytometry, and the results are shown in FIG.
[0281] As shown in Figure 3, the wild-type full-length SPB transposase domain generated approximately 31% GFP-positive cells. Deletion of the first 20 amino acid residues at the N-terminus of the SPB transposase domain had little effect on the percentage of GFP-positive cells, and deletion of 40, 60, or even 80 amino acids at the N-terminus of the SPB transposase domain reduced the percentage of GFP-positive cells by only 25-50% of wild-type activity. Deletion of 100 or 115 amino acid residues further reduced SPB transposase activity, but the SPB transposase domain with a deletion of 115 amino acids (approximately 1 / 3 of the Sb transposase coding sequence) still maintained 25% of wild-type activity.
[0282] In the second experiment, HEK293 were seeded on day 0 and the cells were transfected and the number of cells expressing GFP was determined as described above in the first experiment, except that the reporter transposon construct was co-administered at the same concentration and under the same conditions with one of the fusion proteins containing one of the N-terminally deleted transposase domains prepared in Example 2. The results are shown in Figure 4A.
[0283] As shown in Figure 4A, all fusion proteins containing the wild-type SPB transposase domain linked to the N-terminal deleted SPB transposase domain ("tdSPB" in Figure 4A) maintained cleavage activity at approximately 75% of the level of the wild-type SPB transposase domain ("monomeric SPB" in Figure 4A), indicating that the N-terminal deleted fusion proteins are functional in recognizing and cleaving DNA.
[0284] Example 4: Methods for measuring the integration activity of SPB transposase domains and fusion proteins This assay is designed to measure the integration activity of a fusion protein containing two SPB transposase domains. In this assay, the fusion protein is co-administered to cells with a reporter transposon construct, in which the transposon contains a DNA nucleotide sequence encoding GFP, where the coding sequence is flanked by TTAA and ITR sequences of the PB transposon. The TTAA and ITR sequences act as recognition sites for the SPB transposase domain, and if the fusion protein has integration activity, the DNA encoding GFP will be integrated into genomic DNA, whereupon it will be expressed, producing GFP-positive cells that can be identified and quantified by FACS.
[0285] The integration activity of fusion proteins containing one wild-type transposase domain and one N-terminally deleted transposase domain was measured. On day 0, HEK293 cells were seeded into 48-well plates at a density of 70,000 cells / well, DMEM medium supplemented with 10% FBS was added to each well, and the cells were cultured at 37°C and 5% CO2. On day 1, the medium was aspirated and the cells were resuspended in Jetprime buffer containing transfection reagent (Polyplus Transfection) according to the manufacturer's instructions, and fusion proteins containing SPB transposase domain and reporter transposon construct were added at the concentrations per well shown in Table 10. [Table 14]
[0286] Approximately 24 hours later (day 2), the medium was removed and the cells were resuspended in fresh DMEM medium supplemented with 1% FBS and incubated for an additional 3 days. On day 6, the medium was removed again and the cells were resuspended in fresh DMEM medium supplemented with 1% FBS and incubated for an additional 2 days. On day 8, the cells were resuspended in PBS supplemented with 5% FBS and the number of cells expressing GFP was measured using flow cytometry. The results are shown in Figure 4B.
[0287] As shown in FIG. 4B, a fusion protein containing a wild-type SPB transposase domain fused to a second wild-type SPB transposase domain via a linker ("tdSPB" in FIG. 4B) reduces integration activity by about 33% compared to the wild-type SPB transposase domain alone ("monomeric SPB" in FIG. 4B). Fusion proteins containing a wild-type transposase domain and a N-terminal deleted transposase domain with a deletion of up to 100 amino acids at the N-terminus of the second SPB transposase domain show similar or greater activity than a fusion protein containing two wild-type SPB transposase domains. A deletion of 60 amino acids at the N-terminus of the second transposase domain, however, increases integration activity to a level equal to the wild-type SPB transposase domain alone, approximately 33% greater than the integration activity of a fusion protein containing two wild-type SPB transposase domains.
[0288] Example 5: Rational design of SPB heterodimers: SPB dimers are believed to be held together by a combination of salt bridges, hydrogen bonds, π-cation pairs, and hydrophobic interactions. Residues involved in these interactions in the SPB dimer can be identified by examining published structures of piggyBac (PB) transposase (see, for example, Structural basis of seamless excision and specific targeting by piggyBac transposase. Chen Q, Luo W, Veach RA, Hickman AB, Wilson MH, Dyda F. Nat Commun (2020) 11 p.3446). Two structures deposited at NCBI, 6X67 and 6X68, were analyzed using the "Interaction Analysis" tool in the NCBI Protein Structure 3D Viewer to find amino acids likely involved in dimerization between the two PB transposase monomers. Default settings were used, which explore potential hydrogen bonds up to 3.8 Å, salt bridges up to 6 Å, π-cation pairs up to 6 Å, and other contacts up to 4 Å. The residue pairs shown in Table 2 were identified. These residues are found within the DNA-binding and dimerization domain (DDBD) (residues 118-263, 458-535) or within the cysteine-rich C-terminal domain (CRD) (residues 554-594). Although each and every one of these residues, plus the surrounding residues, could theoretically be mutated within the SPB or PBx transposase monomers to create the obligate heterodimer, the residues within the DDBD were investigated first because the structure of the SPB dimer is more symmetric around the DDBD than around the CRD. For example, within the DDBD, D198 of monomer 1 interacts with K500 of monomer 2, and K500 of monomer 1 interacts with D198 of monomer 2. However, within the CRD, R583 of monomer 1 interacts with D588 of monomer 2, but D588 of monomer 1 does not interact with R583 of monomer 2.
[0289] Initial studies focused on two salt bridges likely involved in holding the PB dimer together, namely, between D198 and K500, and between D201 and R504. By exchanging a negatively charged residue (D) with a positively charged residue (K, R) in one SPB transposase domain, and by exchanging a positively charged residue with a negatively charged residue in the second SPB transposase domain, two new types of SPB mutants, SPB+ and SPB-, were generated. SPB+ was expected to repel SPB+, and similarly, SPB- was expected to repel SPB-. Because opposite charges attract, SPB+ was expected to heterodimerize with SPB-.
[0290] Subsequently, uncharged residues were also mutated to charged residues to generate additional charges at the dimerization interface. For example, M185 of one PB transposase monomer is located very close to L204 of the second PB transposase monomer. To add a positive charge to monomer 1, an M185K mutation was introduced, and to add a negative charge to monomer 2, an L204E mutation was introduced.
[0291] The separate point mutations that create different versions of SPB+ can be combined in any possible combination to create additional SPB+ mutants. The same is true for SPB- mutations. SPB+ and SPB- mutant monomers can be used as the transposase domain of the fusion proteins described herein.
[0292] Example 6: Testing of SPB heterodimers: The SPB+ or SPB- transposase domain mutants described in Example 5 were cloned into an expression vector driven by the EF1α promoter. Specifically, SPB mutants containing SEQ ID NO: 31, 32, 33, 34, 35, 36, 37, 38, 39, 44, 45, 46, 47, 48, 49, or 50 were tested. The nucleotide sequence of the expression vector is set forth in SEQ ID NO: 54.
[0293] Each mutant was then nucleofected into K562 cells either alone (to form homodimers, e.g., two SPB+ mutants) or together with its respective heterodimeric counterpart (e.g., an SPB+ mutant and a corresponding SPB- mutant). To assay transposition activity, cells were co-transfected with a dual cleavage / integration luciferase reporter vector. The vector was designed such that the firefly luciferase open reading frame was disrupted by the SPB transposon. Initially, firefly luciferase is not expressed, but SPB-mediated cleavage of the transposon and seamless restoration results in expression. The transposon itself expresses a destabilized Nanoluc luciferase mRNA. Nanoluc expression from the episomal vector is unstable because the mRNA lacks a polyA tail and contains a 3' destabilizing element. Integration of the transposon into genomic DNA allows the mRNA to take out the polyA and use the splice donor sequence on the transposon to excise the destabilizing element, resulting in luciferase expression. The reporter vector is shown in the lower diagram of Figure 6A.
[0294] K562 cells were nucleofected using 20 μL of SF buffer and program FF-120. Each reaction contained 50 ng of dual luciferase reporter and 500 ng of SPB expression plasmid. To test SPB homodimers, 500 ng of SPB expression plasmid was used. To test SPB as a heterodimer, 250 ng of each SPB expressing plasmid was used. One day after transfection, luciferase signals were measured using Promega's dual luciferase reagent and a plate reader. The results are shown in Figures 5A-5H. Some constructs showed little to no activity as homodimers, but certainly showed activity as heterodimers. Heterodimer activity reached 25-50% of the activity of wild-type SPB. The best transposase activity was observed with the following combinations: SPB+ D198K, and SPB- K500D, R504D; SPB+ D198K, and SPB- L204E, K500D; and SPB+ D198K, D201R, and SPB- K500D, R504D.
[0295] Example 7: Construction of amino-terminal deletions of the Super PiggyBac transposase Plasmids containing the nucleotide sequence encoding the full-length wild-type Super PiggyBac transposase (SPB; SEQ ID NO: 55) or an integrated deletion variant of the Super PiggyBac transposase (PBx; SEQ ID NO: 56) containing amino acid substitutions at positions R372A, K375A, and D450N were used as templates for PCR mutagenesis to generate N-terminal deletion transposase variants lacking the N-terminal 93 amino acids (SPBΔ1-93 and PBxΔ1-93, respectively).
[0296] Briefly, forward and reverse primers were designed to amplify portions of the SPB and PBx coding sequences corresponding to amino acids 94 to 594. The resulting DNA fragments encoding SPBΔ1-93 or PBxΔ1-93 were used together with purchased gBlock gene fragments to construct DNA-binding domain-transposase fusion proteins by state-of-the-art two-fragment Gibson assembly.
[0297] Example 8: Construction of a transposase containing a DNA-binding domain Transposases containing a DNA-binding domain were generated by fusing three zinc finger DNA-binding motifs (ZF268) in frame to the N-terminus (amino acid 94) of SPBΔ1-93 or PBxΔ1-93. Briefly, a gBlock DNA fragment encoding the ZF268 zinc finger protein binding motif flanked by GGGGS linkers (SEQ ID NO:57) was assembled with a DNA fragment encoding SPBΔ1-93 or PBxΔ1-93 from Example 7 and cloned into an expression vector containing in-frame initiation methionine and alanine codons followed by the SV40 nuclear localization sequence (NLS).
[0298] Expression plasmids for ZFM-SPB (SPB with a 93 amino acid N-terminal deletion and a DNA targeting domain containing three zinc finger motifs ZF268) or ZFM-PBx (PBx with a 93 amino acid N-terminal deletion and a DNA targeting domain containing three zinc finger motifs ZF268) were assembled using Gibson assembly. The reactions were carried out under isothermal conditions using three enzyme activities: 5' exonuclease to generate long overhangs, polymerase to fill gaps in annealed single-stranded regions, and DNA ligase to seal nicks in annealed and filled gaps and assemble DNA fragments in the correct order.
[0299] The resulting expression plasmid encodes the transposase ZFM-SPB (SEQ ID NO: 58) containing the full-length DNA binding domain, and ZFM-PBx (SEQ ID NO: 59) containing the N-terminal NLS. Expression of ZFM-SPB and ZFM-PBx is under the control of the EF1α promoter, and each coding sequence is followed by a C-terminal polyadenylation signal.
[0300] Example 9: Design of target integration sequences flanking the TTAA integration site The TTAA target DNA integration site for SPB was modified to insert an adjacent DNA binding site for the zinc finger protein ZF268. ZF268 binds to the 9 nucleotide DNA sequence GCGTGGGCG (SEQ ID NO: 60). A series of four constructs were prepared in which the distance between the TTAA site and the ZF268 binding site was varied by 8, 7, 6, or 5 bp (SEQ ID NOs: 61-64, respectively). The four constructs were individually cloned into a split GFP site-specific integration reporter plasmid to measure the relative differences in linker length in transposase-based integration. A schematic diagram of the split GFP reporter plasmid is shown in Figure 7.
[0301] Example 10: Effect of the linker length between the TTAA incorporation site and the adjacent DNA binding domain site on incorporation and cleavage activity The four targeted TTAA integration site constructs containing various linker lengths generated in Example 9 were tested for transposase integration and cleavage activity. The reporter systems used to test integration or cleavage are shown in Figures 6A-6C (dual cleavage / integration reporter) and Figure 7 (split GFP splice site specific reporter). Figure 6A shows a schematic of the assay, and Figures 6B and 6C show vector maps of the plasmids used.
[0302] Intrinsic activity The integration activity of transposases containing DNA-binding domains was measured using a site-specific TTAA integration GFP reporter plasmid. If the transposases containing DNA-binding domains maintain integration activity, integration of the transposon into the site-specific TTAA integration site by the functional transposase restores the full-length GFP coding sequence, resulting in the expression of GFP, from which positive GFP cells can be identified and quantified. Results are shown as the percentage of positive GFP cells per cell population.
[0303] On day 0, 60,000 HEK293 cells were seeded in a 48-well plate. On day 1, 25 ng of a plasmid encoding a transposase (e.g., wt-SPB, ZFM-SPB, or ZFM-PBx), 112.5 ng of a transposon donor plasmid, and 112.5 ng of a site-specific integration reporter plasmid with one of the different linker lengths were delivered to designated wells of the 48-well plate, and the cells were co-transfected using jetPrime reagent (Polyplus) according to the manufacturer's instructions. On day 4, the transfected cells were analyzed by flow cytometry to measure the percentage of GFP-positive cells.
[0304] Cleavage activity The cleavage activity of transposase containing DNA binding domain was measured using a transposon donor plasmid containing a nucleotide sequence encoding H2Kk gene containing an integrated transposon that interrupts the H2Kk coding sequence, inactivating the expression of functional H2Kk protein.If transposase containing DNA binding domain maintains cleavage activity, the expressed fusion protein will cleave the integrated transposon and restore the full-length H2Kk coding sequence.H2Kk is a cell surface protein, and its expression can be detected on the cell surface using fluorescent anti-H2Kk antibody.
[0305] On day 0, 60,000 HEK293 cells were seeded in a 48-well plate. On day 1, 25 ng of a plasmid encoding a transposase (e.g., wild-type SPB, ZFM-SPB, or ZFM-PBx) and 112.5 ng of a transposon donor plasmid were delivered to each well of a 48-well plate, and cells were co-transfected using jetPrime reagent (Polyplus) according to the manufacturer's instructions. On day 2, cells were treated with fluorescent anti-H2Kk antibody and analyzed by flow cytometry to measure the percentage of H2Kk-positive cells.
[0306] result As shown in Figure 8, wild-type SPB, lacking the DNA-binding domain, showed high levels of integration and cleavage activity regardless of linker length, while ZMF-SPB showed reduced but similar cleavage activity for all linker lengths compared to wild-type SPB and reduced but varying levels of integration activity compared to wild-type SPB, with the highest level of integration activity detected with the 7 bp linker (approximately 50% of WT SPB) and the next highest level detected with the 8 bp linker.
[0307] ZFM-PBx showed reduced but similar cleavage activity for all linker lengths compared to wild-type SPB, but slightly higher than ZFM-SPB. In contrast, however, ZFM-PBx showed widely varying levels of integration activity compared to wild-type SPB and ZFM-SPB. ZFM-PBx showed reduced integration activity at linker lengths of 5, 6, and 8 compared to ZFM-SPB, and was highly reduced compared to wild-type SPB. For targeted TTAA integration sites containing a linker length of 7 bp, ZFM-PBx showed integration levels that exceeded those of wild-type SPB, which were nearly twice as high as those of ZFM-SPB. The combined integration activity results suggest that a 7 bp linker between the TTAA integration site and the adjacent DNA binding site is optimal for the integration activity of transposases containing a DNA binding domain as described in Example 8.
[0308] Example 11: Random genomic integration activity for wild-type SPB, ZFM-SPB, and ZFM-PBx To measure the cleavage and random genome integration activities of wild-type SPB, ZFM-SPB, and ZFM-PBx, transposons containing the EF1α promoter and full-length GFP coding sequence were used. When the transposon is cleaved from the donor plasmid by a transposase (e.g., wild-type SPB, ZFM-SPB, or ZFM-PBx), integration occurs at random genome TTAA positions. Random genome integration activity is expressed as the percentage of GFP-positive cells.
[0309] As shown in FIG. 9A, wild-type SPB exhibits the highest level of random off-target genome integration activity. In comparison, ZFM-SPB exhibited reduced cleavage activity as well as random genome integration activity. The reduced overall activity of ZFM-SPB is likely due to the truncated N-terminus of SPB. In particular, the cleavage activity of ZFM-PBx was significantly higher than that of ZFM-SPB. This is likely because ZFM-PBx contains the D450N mutation, which is known to boost the cleavage activity of piggyBac transposase. Importantly, the random genome integration activity of ZFM-PBx was dramatically reduced, which is likely because the fusion protein is based on the integration-deficient PBx. This elimination of random genome integration is believed to be the key to achieving a greater on-to-off integration ratio for ZFM-PBx.
[0310] Example 12: Ratio of on-target to off-target integration activity for wild-type SPB, ZFM-SPB, and ZFM-PBx A split GFP site-specific episomal reporter plasmid containing TTAA integration sites adjacent to ZF268 binding sites with optimal 7 bp linkers was used as a reporter to test on-target episomal integration with wild-type SPB, ZFM-SPB, and ZFM-PBx transposases. Integration of the transposon at the site-specific TTAA target site restores functional GFP activity. Site-specific integration activity for wild-type SPB, ZFM-SPB, and ZFM-PBx was measured as described in Example 10 and is shown in Figure 9B.
[0311] The on-target to off-target integration ratios for ZFM-SPB and ZFM-PBx were calculated by dividing the on-target integration activity by the corresponding random genomic integration activity. The on-target to off-target integration ratios for ZFM-SPB and ZFM-PBx were then normalized to wild-type SPB.
[0312] The results are shown in Figure 9C. As shown in Figure 9C, the ratio of on-target to off-target activity of ZFM-SPB is 3.5-fold compared to wild-type SPB. This result suggests that the zinc finger binding motif indeed preferred integration at the on-target TTAA site. However, this 3.5-fold improvement is only a modest improvement because ZFM-SPB maintains the ability to randomly integrate into genomic TTAA sites even with the zinc finger binding motif. In contrast, the ratio of on-target to off-target activity of ZFM-PBx is 383-fold compared to wild-type, more than 100-fold higher than ZFM-SPB, indicating improved on-target site-specific transfer and reduced off-target site-specific transfer.
[0313] Example 13: Off- and On-Target Activities of ZFM-PBx with Intact N-Termini The cleavage activity and random genome integration activity of SPB, ZFM-PBx, and ZFM-PBx containing PSD (NTD-ZFM-PBx, SEQ ID NO: 67) were measured as described in Example 10 above. The results are shown in FIG. 10A. Both cleavage and integration activities were increased in NTD-ZFM-PBx compared to ZFM-PBx. FIG. 10B shows that both ZFM-PBx and NTD-ZFM-PBx showed reduced off-target activity compared to SPB, while on-target activity was increased in NTD-ZFM-PBx. FIG. 10C shows that the specificity of ND-ZFM-PBx for SPB is increased compared to the specificity of ZFM-PBx for SPB.
[0314] Example 14: Design and construction of TAL arrays targeting specific genes This example illustrates the design and construction of exemplary gene-targeted TAL array compositions that can be used in methods to validate the target specificity of TAL arrays.
[0315] Using the design criteria described herein or below, TAL arrays were constructed targeting the following genes: GFP, zinc finger 268 (ZFN268), phenylalanine hydroxylase (PAH), beta-2-microglobulin (B2M), and LINE1 repeat elements.
[0316] A. GFP For proof of concept, TAL array pairs containing an N-terminal domain that recognizes T were designed to target specific 10 bp right and 10 bp left pair sequences within the GFP coding region described above (see, e.g., Reyon et al., Nat Biotechnol. 2012 May;30(5):460-5. doi: 10.1038 / nbt.2170. PMID: 22484455; PMCID: PMC355894). 7 In one example, left and right TAL array pairs were designed to target TGCCACCTACG (SEQ ID NO: 240) and TGCAGATGAAC (SEQ ID NO: 241), respectively, generating the GFP1 left TAL array (SEQ ID NO: 113) and the GFP1 right TAL array (SEQ ID NO: 114).
[0317] A second series of TAL array pairs was designed containing an N-terminal domain recognizing a GFP-targeting T targeting the 10 bp GFP sequence TGGCCCACCCT (sequence number 242) and TGCACGCCGTA (sequence number 243), generating the GFP2 left TAL array (sequence number 115) and the GFP2 right TAL array (sequence number 116).
[0318] B. Zinc Finger 268 A TAL array containing an N-terminal domain that recognizes T was designed to target a specific 10 bp sequence in the ZFM268 target site. A TAL array was designed to target the sequence TACGCCCACGC (SEQ ID NO: 239) of zinc finger 268 to generate the ZFM268 TAL array (SEQ ID NO: 112).
[0319] C. PAH In particular, TAL array pairs were designed that contain an N-terminal domain that recognizes T, targeting six specific 10 bp right and left pair sequences of the PAH gene present in introns 1 and 2 of the PAH gene. The TTAA site is located 24 bp downstream of the T nucleotide and 24 bp upstream of the A nucleotide, allowing for a 10 bp TAL recognition target site and a 13 bp spacer on either side of the TTAA. The left and right target sequences used to generate TAL arrays targeting the PAH gene are shown in Table 11. [Table 15]
[0320] Six left and right pair combinations were used to design and construct PAH left TAL arrays 1-6 (sequence numbers 117, 119, 121, 123, 125, and 127, respectively) and PAH right TAL arrays 1-6 (sequence numbers 118, 120, 122, 124, 126, and 128, respectively).
[0321] D. B2M TAL array pairs were designed that target seven specific 10 bp right and left paired sequences of the B2M gene, containing an N-terminal domain that recognizes T. The left and right TAL array target sequences used to design TAL arrays targeting the B2M gene are shown in Table 12. [Table 16]
[0322] Individual TAL molecules containing 34 amino acids or 20 amino acid "half" repeats were synthesized flanked by BsmBI type IIS restriction enzyme cleavage sites. The entire module set contains four molecules (40 molecules / 10 bp target) that can recognize either A, C, G, or T for each 10 bp position in the target sequence. Pairs of TAL array targeting sequences were designed within the B2M gene, and corresponding molecules were selected, pooled, and assembled in frame using "Golden Gate Assembly" to generate each B2M TAL-array. All coding sequences used were codon-optimized for human expression.
[0323] Nine left and right pair combinations were used to design and construct B2M left TAL arrays 1-7 (sequence numbers 144, 146, 148, 150, 152, 154, 156, 518, and 520, respectively) and B2M right TAL arrays 1-7 (sequence numbers 145, 147, 149, 151, 153, 155, 157, 519, and 521, respectively).
[0324] E. LINE1 repeat elements TAL array pairs were designed that contain an N-terminal domain that recognizes T, targeting six specific 10 bp right and left pair sequences of LINE-1 repetitive elements. Some of the LINE1 pairs had more than one left or right target sequence designed to the same position.
[0325] The left and right target sequences used to design TAL array pairs targeting LINE1 repetitive elements are shown in Table 13. [Table 17]
[0326] Individual TAL molecules containing 34 amino acid or 20 amino acid "half" repeats were synthesized flanked by BsmBI type IIS restriction enzyme cleavage sites. The entire module set contains four molecules (40 molecules / 10 bp target) capable of recognizing either A, C, G, or T for each 10 bp position. Pairs of TAL array targeting sequences within the LINE1 repeat were designed and corresponding molecules were selected and pooled to assemble each LRE TAL array in frame using "Golden Gate Assembly". All coding sequences used were codon-optimized for human expression.
[0327] Nine left and right paired target sequences were used to design and construct the LINE1 repeat element (LRE) left TAL arrays LREL1, LREL2, LREL3, LRE4L1, LRE4L2, LREL5, and LREL6 (SEQ ID NOs: 129, 131, 134, 136, 137, 139, and 141, respectively), and the LINE1 repeat element right TAL arrays LRE1, LRE2R1+, LRE2R2+, LRER3, LRER4, LRER5, LRE6R1+, and LRE6R2+ (SEQ ID NOs: 130, 132, 133, 135, 138, 140, 142, and 143, respectively).
[0328] Example 15: General method for designing and constructing TAL-FokI fusions (also called TALENs) This example illustrates an exemplary general method for designing and constructing TALENs that can be used in methods to validate TAL array target specificity.
[0329] The target site specificity of TAL arrays, such as the TAL array constructed in Example 14, was measured, in part, by constructing TAL-FokI fusion proteins (TALENs) that were used in subsequent assays to measure TAL-specific endonuclease activity at the locations of the designed target sites.
[0330] A TALEN expression plasmid was designed and synthesized containing, from 5' to 3', a CMV promoter, a T7 promoter, a Kozak sequence, a 3x flag tag (SEQ ID NO: 70), an SV40 NLS (SEQ ID NO: 71), a delta152 TAL N-terminal domain (SEQ ID NO: 73), two BsmBI type IIS restriction enzyme sites for insertion of left or right TAL arrays, a +63 TAL C-terminal domain (SEQ ID NO: 76), a GS linker, a FokI nuclease (SEQ ID NO: 79), and a bGH polyadenylation sequence.
[0331] Cloning of the left or right TAL array flanked by BsmBI into the BsmBI site of the expression plasmid results in an in-frame fusion of the TAL array and the FokI coding sequence via a linker generating full-length TALENs. All coding sequences used were codon-optimized for human expression using the GeneArt algorithm (Thermo Fisher).
[0332] Example 16: Construction of TAL-FokI fusions (TALENs) targeting specific genes This example demonstrates the construction of TALENs, including the TAL array designed and constructed in Example 14.
[0333] Expression vectors containing TALENs comprising each of the TAL arrays containing an N-terminal domain that recognizes T, as constructed in Example 14, were prepared as generally described in Example 15.
[0334] A. GFP DNA sequences encoding the GFP1 left or right TAL arrays of Example 14A, or the GFP2 left or right TAL arrays, containing adjacent BsmBI ends, were individually cloned into the BsmBI type IIS restriction enzyme sites of TALEN expression vectors generating GFP1 TALENs (sequence numbers 159 and 160) and GFP2 TALENs (sequence numbers 161 and 162).
[0335] B. ZFN268 The DNA sequence encoding the ZFN268 TAL array of Example 14B, containing flanking BsmBI ends, was cloned into the BsmBI type IIS restriction enzyme site of the TALEN expression vector to generate the ZFN268 TALEN (SEQ ID NO: 158).
[0336] C. PAH DNA sequences encoding the left or right TAL arrays of PAH pair numbers 1 to 6 of Example 14C, containing adjacent BsmBI ends, were individually cloned into the BsmBI type IIS restriction enzyme sites of TALEN expression vectors generating 12 PAH left and right TALENs (sequence numbers 163, 165, 167, 169, 171, and 173), and (sequence numbers 164, 166, 168, 170, 172, and 174), respectively.
[0337] D. LINE1 repetitive elements DNA sequences encoding the left or right TAL arrays of LINE1 repetitive element (LRE) pair numbers 1-6 of Example 14E containing adjacent BsmBI ends were individually cloned into the BsmBI type IIS restriction enzyme sites of the TALEN expression vectors generating 16 LRE left and right TALENs LRE1L, 2L, 3L, 4L1, 4L2, 5L, and 6L (SEQ ID NOs: 175, 177, 180, 182, 183, 185, and 187), and LRE1R, 2R1+, 2R2+, 3R, 4R, 5R, 6R1+, and 6R2+ (SEQ ID NOs: 176, 178, 179, 181, 184, 186, 188, and 189), respectively.
[0338] Example 17: Method for analyzing TAL array target site specificity using TALENs in a single strand annealing (SSA) assay This example describes an exemplary assay for measuring site-specific cleavage of a target site by TALENs, including the TAL arrays of the present invention.
[0339] The sequence specificity of TALENs (including those constructed in Example 16), including TAL arrays such as the TAL array constructed in Example 14, was measured, in part, using a single-stranded annealing (SSA) assay.
[0340] The SSA luciferase reporter plasmid was designed and synthesized as previously described (see, e.g., Juillerat A, et al., Comprehensive analysis of the specificity of transcription activator-like effector nucleases. Nucleic Acids Res. 2014 Apr;42(8):5390-402. doi: 10.1093 / nar / gku155. Epub 2014 Feb 24. PMID: 24569350; PMCID: PMC4005648). The plasmid contains, from 5' to 3', a CMV promoter, a Kozak sequence, a first N-terminal segment of firefly luciferase coding sequence (SEQ ID NO: 237), two stop codons, two BsaI type IIS restriction enzyme cleavage sites, a second C-terminal segment of firefly luciferase coding sequence (SEQ ID NO: 238), and an SV40 polyadenylation sequence. The two segments of the firefly luciferase coding sequence contain 628 bp of overlapping sequence. When the target site for TALEN is cloned into the BsaI site and the reporter construct is cut, it can be restored in cells by single-stranded annealing, which results in the expression of the full-length firefly luciferase coding sequence and firefly luciferase (SEQ ID NO: 236), indicating that TALEN recognizes its target site in a site-specific manner.
[0341] Complementary oligos were synthesized downstream of the T, containing the target site for each TAL array, followed by a 16 bp spacer, followed by the reverse complement of the TAL target site, followed by an A. Additionally, complementary oligos were synthesized containing the target site for the left TAL array, followed by a 16 bp spacer, followed by the reverse complement of the target site for the right TAL array, followed by an A. The complementary oligos contained 4 bp overhangs that were compatible with the overhangs created in the SSA reporter after BsaI digestion. The oligos were annealed and ligated into the digested vector to generate SSA reporters compatible with each TALEN.
[0342] GFP For example, a GFP1 reporter plasmid containing two left TAL array target sequences (SEQ ID NO: 287), two right TAL array target sequences (SEQ ID NO: 288), one left and one right TAL array (SEQ ID NO: 286), and a GFP2 reporter plasmid containing two left TAL array target sequences (SEQ ID NO: 290), two right TAL array target sequences (SEQ ID NO: 291), one left and one right TAL array (SEQ ID NO: 289). In addition, a ZFN268 TAL array target site (SEQ ID NO: 285) was prepared as a second target. All of these constructs were used in the subsequent SSA assay.
[0343] The cleavage activity of the six GFP TALENs (GFP1 and GFP2) and ZFM268 TALEN constructed in Example 16 was measured. A transfection mixture containing 45 ng of left TALEN, 45 ng of right TALEN, 10 ng of the corresponding reporter, and 0.3 μL of Transit-2020 transfection reagent was assembled in Serum Free OptiMem medium with a total volume of 20 μL. As a negative control, each TALEN pair was transfected simultaneously with a reporter lacking the correct target site sequence. 60,000 HEK293T cells were added to 180 μL of DMEM medium supplemented with 10% FBS, and the transfection mixture was plated in a 96-well plate and incubated at 37° C., 5% CO2 for 1 day. The next day, cell lysis buffer was added to the cells and the lysate was transferred to a white 96-well plate. A buffer containing a substrate for firefly luciferase was mixed with the cells and luciferase luminescence was detected using a plate reader. The results are shown in Table 14 and Figure 12. [Table 18]
[0344] As shown in Table 14, luciferase was rapidly detected at levels several orders of magnitude higher when the corresponding TALEN and reporter pairs were co-transfected together than in the negative control, indicating on-site targeting activity of each TALEN construct.
[0345] PAH In another experiment, SSA reporter plasmids targeting PAH were designed and constructed for the PAH TALENs constructed in Example 16C: PAH1-6 left TALENs (sequence numbers 163, 165, 167, 169, 171, and 173), and PAH1-6 right TALENs (sequence numbers 164, 166, 168, 170, 172, and 174).
[0346] SSA assay was carried out using the method described above. Briefly, two copies of each PAH target site were cloned into SSA reporter plasmid, separated by 16bp spacers, PAH1 left and right (SEQ ID NO: 292 and 293); PAH2 left and right (SEQ ID NO: 294 and 295); PAH3 left and right (SEQ ID NO: 296 and 297); PAH4 left and right (SEQ ID NO: 298 and 299); PAH5 left and right (SEQ ID NO: 300 and 301); and PAH6 left and right (SEQ ID NO: 302 and 303).
[0347] Each TALEN was co-transfected with its corresponding reporter or a reporter containing a non-targeting sequence, and luciferase was measured the next day. The results are shown in Table 15 and Figure 13. [Table 19]
[0348] LINE-1 repetitive elements In another experiment, SSA reporter plasmids containing two copies of each LINE1 target site separated by a 16 bp spacer (SEQ ID NOs: 304-318) targeting the LINE1 repeat element were designed and constructed for each constructed LINE1 TALEN in Example 16D: TALEN LRE1L, 2L, 3L, 4L1, 4L2, 5L, and 6L (SEQ ID NOs: 175, 177, 180, 182, 183, 185, and 187), and LRE1R, 2R1+, 2R2+, 3R, 4R, 5R, 6R1+, and 6R2+ (SEQ ID NOs: 176, 178, 179, 181, 184, 186, 188, and 189), respectively. The results are shown in Table 16. [Table 20]
[0349] As shown in Table 16, most of the tested TALENs produced luciferase signals over an order of magnitude higher with on-target and off-target reporters. The SSA assay indicates that the newly designed TALs can recognize their intended target sequences and induce the fused FokI nuclease to cleave the adjacent DNA, resulting in single-stranded annealing and luciferase expression.
[0350] Example 18: Construction and analysis of TAL array-piggyBac transposase (ss-SPB) compositions designed for site-specific transposition at specific genes (TAL-PBx) This example demonstrates the construction of a TAL array-Super piggyBac transposase fusion protein composition (TAL-ssSPB), useful in a method to achieve site-specific transposition at a specific target locus.
[0351] A TAL-PBx fusion construct was prepared similar to the ZFM268-PBx construct described above in Examples 14 and 16. An expression plasmid was synthesized containing, from 5' to 3', a CMV promoter, a T7 promoter, a Kozak sequence, a 3x flag tag (SEQ ID NO: 70), an SV40 NLS (SEQ ID NO: 71), a delta 152 TAL N-terminal domain (SEQ ID NO: 73), two BsmBI type IIS restriction enzyme sites, a +63 TAL C-terminal domain (SEQ ID NO: 76), a GGGS linker, a delta 1-93PBx (containing a 93 deletion at the N-terminus and mutations at R372A, K375A, and D450N in the Super piggyBac transposase codon sequence SEQ ID NO: 66), and a bGH polyadenylation sequence.
[0352] Cloning of the left or right TAL array flanked by BsmBI into the BsmBI site of the expression plasmid results in an in-frame fusion of the TAL array and the PBx coding sequence via a linker sequence generating the full-length TAL-PBx construct. All coding sequences used were codon-optimized for human expression using the GeneArt algorithm (Thermo Fisher).
[0353] A. GFP1 and 2 TAL-PBx and ZFM 268 TAL-PBx Two pairs of TAL arrays were designed that target the TAL array targeting sequence in the GFP coding sequence of Example 14A as well as a 10 base pair sequence (ACGCCCACGC downstream of the T; SEQ ID NO: 239) that contains the reverse complement of the ZFM268 target site of Example 14B. Each TAL array was synthesized with a 34 amino acid repeat followed by 9 20 amino acid "half" repeats adjacent to a BsmBI type IIS restriction enzyme cleavage site. This allows for cloning of each TAL array in frame with the remainder of the open reading frame in the expression plasmid to generate GFP1 Left TAL-PBx (SEQ ID NO: 191), GFP1 Right TAL-PBx (SEQ ID NO: 192), GFP2 Left TAL-PBx (SEQ ID NO: 193), GFP2 Right TAL-PBx (SEQ ID NO: 194), and ZFM268 TAL-PBx (SEQ ID NO: 190). All coding sequences used were codon-optimized for human expression.
[0354] GFP TAL-PBx and ZFM268 TAL-PBx constructs were used in Example 19 to determine the optimal spacer distance between TTAA incorporation sites, as well as the location of the left and right TAL target sequences relative to the TAL-PBx constructs.
[0355] B. PAH1-6 left and right TAL-PBx The PAH locus was selected as a target for site-specific transposition into genomic DNA. Six TTAA sites were selected within the first two introns that fit the motifs described herein. TAL arrays targeting these sequences were synthesized in Example 14C and cloned into the TAL-ssSPB expression vector using the method described in Example 17, thereby generating the PAH1-6 left TAL-PBx (SEQ ID NOs: 195, 197, 199, 201, 203, and 205, respectively) and PAH1-6 right TAL-PBx sequences (SEQ ID NOs: 196, 198, 200, 202, 204, and 206, respectively).
[0356] C. B2M Left and Right TAL-PBx Nine TAL arrays designed and constructed in Example 14D flanked by BsmBI termini were cloned into the BsmBI restriction enzyme cleavage sites of the expression plasmid described above to generate 18 B2M1-9 TAL-PBx constructs: B2M1-9 left TAL-PBx (sequence numbers 222, 224, 226, 228, 230, 232, 234, 522, and 524, respectively), and B2M1-9 right TAL-PBx (sequence numbers 223, 225, 227, 229, 231, 233, 235, 523, and 525, respectively).
[0357] D. LINE1 repeat element left and right TAL-PBx LINE1 repetitive elements occur thousands of times throughout the human genome, making them potentially attractive targets for optimizing the chances of site-specific transposition events at the target sequence, thereby resulting in an increase in the numbers of metastatic cells.
[0358] Fifteen TAL arrays designed and constructed in Example 14E flanked by BsmBI ends were cloned into the BsmBI restriction sites of the expression plasmids described above to generate fifteen LRE1-6 TAL-PBx constructs: LRE1L, LRE2L, LRE3L, LRE4.1L, LRE4.2L, LRE5L, and LRE6L left TAL-PBx (SEQ ID NOs: 207, 209, 212, 214, 215, 217, and 219, respectively), and LRE1R, LRE2.1R, LRE2.2R, LRE3R, LRE4R, LRE5R, LRE6.1R, and LRE6.2R right TAL-PBx (SEQ ID NOs: 208, 210, 211, 213, 216, 218, 220, and 221, respectively).
[0359] Example 19: Determination of optimal spacer length between the TTAA incorporation site and the left and right TAL target sequences using an episomal split GFP splicing reporter system This example illustrates exemplary compositions and methods for preparing optimal target sites for site-specific transposition using a TAL array-SPB transposase fusion protein.
[0360] The episomal split GFP splicing reporter system was used to evaluate different spacer lengths on site-specific transposition efficiency. The reporter system is composed of two plasmids. The first plasmid, "reporter", was constructed containing, from 5' to 3', the EF1α promoter (SEQ ID NO: 325), the Kozak sequence, the first part of the GFP open reading frame (SEQ ID NO: 326), the splice donor (SEQ ID NO: 327), and two BsaI type IIS restriction enzyme sites. The BsaI sites allow cloning of target TTAA sequences flanked by spacers of variable length flanked by target recognition sequences for the TAL array. A second plasmid, "donor," was constructed containing, from 5' to 3', a TTAA sequence, a 35 bp PiggyBac minimal 5' ITR (SEQ ID NO:319), a splice acceptor site (SEQ ID NO:321), a second part of the GFP open reading frame (SEQ ID NO:322), a synthetic polyadenylation sequence (SEQ ID NO:323), a 63 bp PiggyBac minimal 3' ITR (SEQ ID NO:320), and a TTAA sequence.
[0361] A complementary oligo was synthesized containing the target site for the GFP1 right TAL downstream of the T, followed by a 6 bp spacer, followed by TTAA, followed by a 6 bp spacer, followed by the reverse complement of the TAL target site, followed by A (SEQ ID NO: 330). The complementary oligo contained a 4 bp overhang that was compatible with the overhang created in the split GFP splicing reporter after digestion with BsaI. The oligo was annealed and ligated into the digested vector to create a reporter compatible with the GFP1 right TAL-PBx. Similar oligos were synthesized replacing the two 6 bp spacers with spacers of length 7 bp (SEQ ID NO:331), 8 bp (SEQ ID NO:332), 9 bp (SEQ ID NO:333), 10 bp (SEQ ID NO:334), 11 bp (SEQ ID NO:335), 12 bp (SEQ ID NO:336), 13 bp (SEQ ID NO:337), 14 bp (SEQ ID NO:338), and 15 bp (SEQ ID NO:339). These were cloned in the same manner to generate reporters containing spacers of variable length.
[0362] Each reporter plasmid and donor plasmid were co-transfected with the GFP1 right TAL-PBx expression plasmid into HEK293T cells. As a negative control, ZFM268 TAL-PBx expression plasmid, which does not recognize the GFP1 target sequence, was transfected instead of the GFP1 right TAL-PBx expression plasmid. A transfection mixture containing 26 ng of TAL-ssSPB expression vector, 170 ng of reporter plasmid, 117 ng of donor plasmid, and 0.78 μL of Transit-2020 transfection reagent was assembled in a total volume of 26 μL of Serum Free OptiMem medium. 95,000 HEK293T cells were added to 250 μL of DMEM medium supplemented with 10% FBS, and the transfection mixture was plated in a 48-well plate and incubated at 37°C, 5% CO2 for 4 days, and the cells were split 1:3 on the second day.
[0363] When the reporter and donor plasmids are co-transfected into cells with TAL-PBx, TAL-PBx catalyzes the excision of the transposon from the donor plasmid and its site-specific integration into the TTAA target site of the reporter plasmid. After site-specific transposition, transcription, splicing, and translation, a reconstructed GFP coding sequence (DNA, SEQ ID NO: 328; amino acid; SEQ ID NO: 329) is produced and fluorescence can be detected. The percentage of target site-specific transposition positive cells for constructs with various spacer lengths was analyzed by FACS analysis, and the results are shown in Table 17. [Table 21]
[0364] As shown in Figure 13, GFP1 right TAL-PBx catalyzed site-specific translocation and resulted in GFP signals above background levels with target sites containing 12-bp, 13-bp, and 14-bp spacers separating the TTAA incorporation site from the TAL binding site. The negative control ZFM268 TAL-PBx did not result in GFP signals above background using the GFP1 right specific reporter.
[0365] To determine whether optimal spacer lengths were consistent across TAL-ssSPBs, similar reporters were constructed containing TAL target sites for GFP1 left, GFP2 right, GFP2 left, and ZFM268 TAL-PBxs as described above. These constructs were tested using a narrower set of spacer lengths: 11 bp, 12 bp, 13 bp, 14 bp, and 15 bp constructs for GFP1 left (SEQ ID NOs: 345-349), GFP2 left (SEQ ID NOs: 350-354), GFP2 right (SEQ ID NOs: 355-359), and ZFM268 (SEQ ID NOs: 340-344).
[0366] Each reporter plasmid and donor plasmid were co-transfected with the corresponding TAL-ssSPB expression plasmid into HEK293T cells. 120,000 HEK293T cells were plated in a 24-well plate in 500uL of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture was assembled containing 50ng of TAL-ssSPB expression vector, 225ng of reporter plasmid, 225ng of donor plasmid, and 1uL of JetPrime transfection reagent in a total volume of 50uL of JetPrime buffer. The mixture was added to HEK293T cells, and the cells were incubated at 37°C, 5% CO2 for 4 days, with the cells split 1:6 on day 1. The percentage of target site-specific translocation positive cells for constructs of various spacer lengths was analyzed by FACS analysis, and the results are shown in Table 18. [Table 22]
[0367] As shown in Table 18, 12 bp and 13 bp spacers were optimal and resulted in the highest GFP expression due to site-specific transfer of the donor transposon into the reporter plasmid in the cell populations for all TAL-PBx constructs and reagents tested.
[0368] In another experiment, the donor plasmid target integration site containing the optimal 13 bp spacer was modified to mutate the adjacent 5' and 3' nucleotides immediately adjacent to the TTAA incorporation sequence to T and A, respectively, generating a TTTAAA incorporation site flanked by a 12 bp spacer between two TAL target sequences: GFP1 right (SEQ ID NO: 382); GFP2 left (SEQ ID NO: 383); GFP2 right (SEQ ID NO: 384); GFP2 left (SEQ ID NO: 385), and ZFM268 (SEQ ID NO: 386). The modified TTTAAA (13 bp v2) and TTAA (13 bp) donor plasmids were compared using an episomal split GFP splicing reporter system using GFP1 left TAL-PBx, GFP1 right TAL-PBx, GFP2 left TAL-PBx, GFP2 right, TAL-PBx, ZFM268 TAL-PBx expression plasmids as described in Example 18A.
[0369] Briefly, each reporter plasmid and donor plasmid were co-transfected with the corresponding TAL-PBx expression plasmid into HEK293T cells. Approximately 120,000 HEK293T cells were plated in a 24-well plate in 500 μL of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture was assembled containing 50 ng of GFP1 TAL-PBx or ZFM268 TAL-PBx expression vector, 225 ng of reporter plasmid, 225 ng of donor plasmid, and 1 μL of JetPrime transfection reagent in a total volume of 50 μL of JetPrime buffer. This mixture was added to HEK293T cells, which were incubated at 37°C, 5% CO2 for 4 days, and the cells were split 1:6 on day 1. The percentage of GFP positive cells was determined for each TTAA or TTTAAA integration site construct and the results are shown in Table 19. [Table 23]
[0370] As shown in Table 19, modification of the TTAA incorporation site to TTTAAA resulted in an approximately two-fold increase in the number of GFP-expressing cells within the transferred cell population for each GFP TAL-PBx, as well as for the ZFM268 TAL-PBx.
[0371] Example 20: TAL-PBx targeted site-specific translocation at specific gene loci This example demonstrates that the TAL-ssSPB (TAL-PBx) compositions of the invention can effect site-specific transposition of transposons at specific episomal and genomic loci.
[0372] A. PAH episomes and genomic targeted site-specific transfer i. episome Episomal split GFP splicing reporter constructs were designed and cloned as described above. Six PAH target sequences (SEQ ID NOs: 360-365), naturally found in genomic DNA, were cloned into episomal reporter plasmids. These plasmids were co-transfected with the TAL recognition sequence, a 13 bp spacer of optimal length, TTAA, a second 13 bp spacer of optimal length, the reverse complement of the TAL recognition sequence, and A. Arrays were designed and constructed to generate TAL-ssSPB heterodimer pairs (i.e., one left and one right, TAL array-PBx). PAH1-6-TAL-PBx construct pairs were assayed as described above, and the results are shown in Table 20 and Figure 14. [Table 24]
[0373] As shown in Table 20, split GFP splicing reporter assays demonstrate that the newly constructed PAH TAL-PBx is capable of site-specific transposition to a target sequence naturally found in genomic DNA.
[0374] In separate experiments, the reporter plasmid was also co-transfected with either the PAH left or right TAL-PBx constructs (i.e., homodimers) and assayed as described above. The results are shown in Table 21 and FIG. [Table 25]
[0375] As shown in Table 21, PAH TAL-PBx homodimers, capable of recognizing only the left or right target sequence of an integration site that contains both left and right target sequences, still resulted in site-specific translocation at the target site compared to the off-target control, albeit at lower levels than the corresponding heterodimer pairs.
[0376] ii. Genomic site-specific transposition After confirming that the newly designed PAH TAL was functional and recognized its target sequence, the PAH TAL-PBx construct was used to catalyze site-specific transposition into endogenous genomic DNA. Briefly, 120,000 HEK293T cells were plated in a 24-well plate in 500uL of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture was assembled containing 25ng of PAH left TAL-PBx expression vector, 25ng of PAH right TAL-PBx expression vector, 450ng of PiggyBac transposon donor plasmid, and 1uL of JetPrime transfection reagent in a total volume of 50uL of JetPrime buffer. The mixture was added to HEK293T cells, which were incubated at 37°C, 5% CO2 for 4 days, and the cells were split 1:6 on day 1.
[0377] The transposon donor plasmid contained, from 5' to 3', a 309 bp fragment containing the Piggybac 5' ITR (SEQ ID NO: 319) and part of the UTR, a "cargo" consisting of multiple restriction enzyme recognition sites, a 238 bp fragment containing the Piggybac 3' ITR (SEQ ID NO: 320) and part of the UTR, and a PiggyBac containing the TTAA. As controls, transfections were also performed using Super PiggyBac transposase (SPB; SEQ ID NO: 80) or no transposase instead of PAH TAL-PBx to examine random or no integration of the transposon from the donor plasmid.
[0378] To investigate the site-specific integration of the transposon donor into the PAH locus, genomic DNA was extracted from transfected cells and analyzed by digital droplet PCR (ddPCR) using a probe-based detection scheme. A primer that binds within the transposon was paired with a primer that binds to the PAH genomic DNA near the TTAA integration site. Thus, an amplicon should be generated only after site-specific integration into the PAH locus. Since integration is not directional, two assays were designed to detect transposon integration in the forward and reverse directions for each PAH target.
[0379] We detected amplicons corresponding to forward and / or reverse transposon integration in genomic DNA isolated from cells transfected with the PAH1 TAL-PBx, PAH2 TAL-PBx, and PAH3 TAL-PBx constructs, providing direct evidence of genomic integration at the PAH locus. With SPB transposase, a small number of amplicons were detected, likely due to low levels of random integration events, while no amplicons were detected in the absence of transposase, suggesting that site-specific transposition in PAH1, PAH2, and PAH3 targets sequences only in the presence of the TAL-PBx constructs.
[0380] B. Episomal-targeted site-specific transfer of LINE1 repetitive elements Nine different LINE1 repetitive element genomic sequences derived from the LINE1 Ta1d consensus sequence (sequence numbers 366-374) were selected as target sequences for episomal site-specific transfer using a pair of LRE1-6 TAL-PBx constructs.
[0381] LRE1-6 left and right TAL-PBx in Example 18D constructed respectively: LRE1-6 TAL-PBx constructs: LRE1L, LRE2L, LRE3L, LRE4.1L, LRE4.2L, LRE5L, and LRE6L left TAL-PBx (SEQ ID NOs: 207, 209, 212, 214, 215, 217, and 219, respectively), and episomal split GFP splicing reporter constructs for LRE1R, LRE2.1R, LRE2.2R, LRE3R, LRE4R, LRE5R, LRE6.1R, and LRE6.2R right TAL-PBx (208, 210, 211, 213, 216, 218, 220, and 221, respectively) were designed and cloned as described above.
[0382] Episomal split GFP splicing assays were performed as described above. Briefly, each LINE1 genomic target site (SEQ ID NOs: 366-374) was cloned into the reporter.
[0383] Each TAL-PBx construct was co-transfected with its corresponding reporter or a reporter containing a non-targeting sequence, and GFP was measured the next day. The results are shown in Table 22. [Table 26]
[0384] ii. Genomic site-specific transposition After confirming that the newly designed LINE1 TALs were functional and recognized their target sequences, the LINE1 TAL-PBx constructs were used to catalyze site-specific transposition into endogenous genomic DNA. Briefly, 120,000 HEK293T cells were plated in 24-well plates in 500uL of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture was assembled containing 25ng of LINE1 left TAL-PBx expression vector, 25ng of LINE1 right TAL-PBx expression vector, 225ng of PiggyBac transposon donor plasmid, and 1uL of JetPrime transfection reagent in a total volume of 50uL of JetPrime buffer. The mixture was added to HEK293T cells, which were incubated at 37°C, 5% CO2 for 3 days, and the cells were split 1:6 on day 1.
[0385] The transposon donor Nanoplasmid contained, from 5' to 3', a TTAA, a 309 bp fragment containing part of the Piggybac 5' ITR and UTR, a "cargo" consisting of the EF1α promoter, a puromycin resistance gene, a 2A peptide, and a GFP reporter, followed by a 238 bp fragment containing part of the Piggybac 3' ITR and UTR, and a PiggyBac containing a TTAA. As controls, transfections were also performed with PBx transposase (SEQ ID NO: 56) instead of LINE1 TAL-PBx, or with no transposase, to examine random or no integration of the transposon from the donor plasmid.
[0386] To investigate the site-specific integration of the transposon donor into the LINE1 locus, genomic DNA was extracted from the transfected cells 3 days after transfection and analyzed by digital droplet PCR (ddPCR) using a probe-based detection scheme. A primer that binds within the transposon was paired with a primer that binds to the LINE1 genomic DNA near the TTAA integration site. Thus, an amplicon should only be generated after site-specific integration into the LINE1 locus. Since integration is not directional, two assays were designed to detect transposon integration in the forward and reverse directions for each LINE1 target. The results are shown in Figure 16 and Table 23. [Table 27]
[0387] As shown in FIG. 16 and Table 23, amplicons corresponding to forward and / or reverse transposon integration were detected from genomic DNA isolated with cells transfected with the LINE1 TAL-PBx construct, providing direct evidence of genomic integration at the LINE1 locus. Higher levels of transposition were detected for targets 2, 4, and 6 than for targets 1, 3, and 5. In the absence of the TAL-PBx construct, no high levels of amplicons were detected, suggesting that site-specific transposition in LINE1 targets sequences only in the presence of the TAL-PBx construct. An additional primer set detecting a reference single copy gene was used to measure the number of genomes represented per ddPCR reaction. This allowed quantification of the percentage (average) of genomes containing the edited LINE1 locus.
[0388] The target sites with the most robust integration, targets 2, 4, and 6, all contain the TTTAAA integration site shown in Figure 16. These data are consistent with the data shown in Example 19 and Table 19, indicating that the TAL-PBx fusion composition prefers TTTAAA integration sites over TTAA integration sites.
[0389] C. B2M episomal and genomic targeted site-specific transfer i. episome Genomic sequences derived from the first intron of the B2M gene (SEQ ID NOs: 375-381) were selected as target sequences for episomal site-specific transfer using the B2M1-7 TAL-PBx construct pair (SEQ ID NOs: 222-235). The B2M genomic sequences (SEQ ID NOs: 375-381) were cloned into an episomal split GFP reporter vector and episomal split GFP splicing assays were performed as described above. Briefly, each B2M TAL-PBx pair was co-transfected with the corresponding reporter and GFP was measured 4 days after transfection. The results are shown in Table 24. [Table 28]
[0390] As shown in Table 24, four of the seven B2M TAL-PBx pairs (pairs 4, 5, 6, and 7) catalyzed site-specific transposition with measurable frequency.
[0391] ii. Genomic site-specific transposition After confirming that the newly designed B2M TALs were functional and recognized their target sequences, the active B2M TAL-PBx constructs were used to catalyze site-specific transposition into endogenous genomic DNA. Briefly, 120,000 HEK293T cells were plated in 24-well plates in 500uL of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture was assembled containing 25ng of B2M left TAL-PBx expression vector, 25ng of B2M right TAL-PBx expression vector, 225ng of PiggyBac transposon donor plasmid, and 1uL of JetPrime transfection reagent in a total volume of 50uL of JetPrime buffer. The mixture was added to HEK293T cells, which were incubated at 37°C, 5% CO2 for 5 days, and the cells were split 1:8 on day 1.
[0392] The transposon donor Nano plasmid contained, from 5' to 3', a 309 bp fragment containing the Piggybac 5' ITR (SEQ ID NO: 319) and part of the UTR, a "cargo" consisting of the EF1α promoter, a puromycin resistance gene, a 2A peptide, and a GFP reporter, followed by a 238 bp fragment containing the Piggybac 3' ITR (SEQ ID NO: 320) and part of the UTR, and a PiggyBac containing the TTAA. As controls, transfections were also performed with PBx transposase (SEQ ID NO: 56), or B2M TAL-PBx transposase (SEQ ID NO: 56), or no transposase, to investigate random or no integration of the transposon from the donor plasmid.
[0393] To investigate the site-specific integration of the transposon donor into the B2M locus, genomic DNA was extracted from transfected cells 5 days after transfection and analyzed by digital droplet PCR (ddPCR) using a probe-based detection scheme. A primer that binds within the transposon was paired with a primer that binds to the B2M genomic DNA near the TTAA integration site. Therefore, an amplicon should be generated only after site-specific integration into the B2M locus. The results are shown in Figure 17.
[0394] As shown in Figure 17, amplicons corresponding to transposon integration were detected from genomic DNA isolated with cells transfected with the B2M TAL-PBx construct, providing direct evidence of genomic integration at the B2M locus. In the absence of the TAL-PBx construct, no amplicons were detected at high levels, suggesting that site-specific transposition at B2M targets sequences only in the presence of the TAL-PBx construct.
[0395] Example 21: Construction of PBx fusion proteins A zinc finger domain (SEQ ID NO: 57) flanked by GGGGS linkers at both the N- and C-termini was inserted into the SV40 NLS PBx replacing one of the various positions between P86-S99 (ZF-ssSPB fusion points shown in Table 25). Thus, the construct maintained the N-terminus of the PBx upstream of the zinc finger domain. The sequences of the constructs are set forth in SEQ ID NOs: 67, and 387-399. These sequences were used to investigate integration activity using the split-GFP reporter shown in Figure 7 with targets shown in SEQ ID NOs: 61-64. The results are shown in Figure 18 and Table 25. [Table 29]
[0396] Example 22: Construction of TALEN and TAL-PBx fusions that recognize alternative nucleotides other than thymidine 5' of the target binding site A.TALEN The wild-type TAL sequence, which most efficiently recognizes the target sequence immediately 3' of the T, was mutated to recognize a 5'G instead of a 5'T (NT-G mutant; SEQ ID NO: 74) or a mutant that does not require any specific 5' nucleotide (NT-βN; SEQ ID NO: 75). These mutations were introduced into the GFP1 right TALEN (SEQ ID NO: 160; Example 16) by mutating the amino acid sequence QW located at positions 119-120 to the amino acid sequence SR to generate the NT-G variant, or by replacing the amino acid sequence QWS at positions 119-121 with YH to generate the NT-βN variant, to generate the GFP1 right TALEN NT-G (SEQ ID NO: 401) and the GFP1 right TALEN NT-βN (SEQ ID NO: 402).
[0397] The design of TALENs NT-G and NT-βN was tested using single stranded annealing reporters (Example 17). The target site corresponding to the GFP1 right TALEN (SEQ ID NO: 288) was modified to replace the T at the 5' of the target site with either A, C, or G to generate SEQ ID NOs: 403-405. Transfection mixtures containing 90 ng of each TALEN, 10 ng of the corresponding reporter, and 1.5 μL of Transit-2020 transfection reagent were assembled in a total volume of 20 μL of Serum Free OptiMem medium. As negative controls, TALENs or reporters were transfected alone. An aliquot of 30,000 HEK293T cells was added to 180 μL of DMEM medium supplemented with 10% FBS, and the transfection mixture was plated in a 96-well plate and incubated at 37°C, 5% CO2 for 1 day. The next day, cell lysis buffer was added to the cells and the lysates were transferred to a white 96-well plate. A buffer containing a substrate for firefly luciferase was mixed with the cells and the luminescence of luciferase was detected using a plate reader. The results are shown in Table 26. [Table 30]
[0398] As shown in Table 25, the WT TALENs provided the highest degree of cleavage of targets containing 5'T, but the NT-G and NT-βN versions were also capable of similar cleavage of targets containing 5'-A, C, G, or T.
[0399] B.TAL-PBx fusion NT-G and NT-βN mutations were introduced into the GFP1 right TAL-PBx fusion (SEQ ID NO: 192; Example 18) to generate the GFP1 right NT-G TAL-PBx fusion (SEQ ID NO: 406) and the GFP1 right NT-βN TAL-PBx fusion (SEQ ID NO: 407). The novel TAL-PBx fusion designs were tested using the episomal split GFP splicing reporter system (Example 19). The GFP1 right target site with a 13 bp spacer (SEQ ID NO: 337) was modified to replace the T 5' of the target site with either A, C, or G to generate SEQ ID NOs: 408-410.
[0400] The activity of the novel mutant TAL-PBx fusions was measured using their corresponding episomal split GFP splicing reporters. Briefly, each reporter plasmid and donor plasmid were co-transfected with the corresponding TAL-PBx expression plasmid into HEK293T cells. Approximately 120,000 HEK293T cells were plated in a 24-well plate in 500 μL of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture was assembled containing 50 ng of TAL-PBx expression vector, 225 ng of reporter plasmid, 225 ng of donor plasmid, and 1 μL of JetPrime transfection reagent in a total volume of 50 μL of JetPrime buffer. This mixture was added to HEK293T cells, which were incubated at 37° C., 5% CO2 for 4 days, with the cells split 1:6 on day 1. The percentage of GFP-positive cells was measured for each sample. The results are shown in Table 27. [Table 31]
[0401] As shown in Table 26, the WT TAL-PBx fusions showed the greatest integration percentage in targets with a 5'T, as did the corresponding TALEN versions, while the mutant NT-G was able to integrate similarly in targets with a 5'G and in targets with 5'-A, C, G, or T, indicating that these alternative target sites can be effectively targeted and modified using the TALEN and TAL-PBx fusion compositions of the present disclosure.
[0402] Example 23: Construction of TAL-PBx fusions containing N-terminal deletions of PBx of various sizes A 93 amino acid N-terminal deletion of PBx (SEQ ID NO:66; Example 7) was used to construct the first exemplary TAL-PBx fusion. To further explore the location of the deletion site, 10 amino acids of the PBx sequence were added back in increments of 1 amino acid to generate PBxdelta83-PBxdelta92 (SEQ ID NOs:86-95). In addition, 10 additional amino acids were deleted in increments of 1 amino acid to generate PBxdelta94-PBxdelta103 (SEQ ID NOs:97-106). These 20 newly truncated PBx sequences were used to replace PBxdelta93 in GFP1RightTAL-PBx (SEQ ID NO:192) to generate GFP1RightTal-PBxdelta83-92 (SEQ ID NOs:450-459) and GFP1RightTal-PBxdelta94-103 (SEQ ID NOs:460-469).
[0403] These corresponding episomal split GFP splicing reporters described in Example 19 were used to test the novel mutant GFP1 right TAL-PBx fusions. Briefly, the site-specific reporter plasmids and donor plasmids were co-transfected with the corresponding GFP1 right TAL-PBx expression plasmids into HEK293T cells. As a benchmark control, the original GFP1 right TAL-PBx fusion with a 93 amino acid truncation of PBx was transfected (SEQ ID NO: 192). As a negative control, non-targeted (GFP1 left TAL-PBx) was transfected (SEQ ID NO: 191). The reporter plasmid contained two targeted GFP1 right targeting sites (downstream of 5'T) flanked by a 13 bp spacer with a central TTAA insertion site (SEQ ID NO: 470). The experiment was repeated using reporters with spacers containing an 11 bp spacer (SEQ ID NO: 335), a 12 bp spacer (SEQ ID NO: 336), and a 14 bp spacer (SEQ ID NO: 338). To perform the transfection, approximately 120,000 HEK293T cells were plated in a 24-well plate in 500 μL of DMEM medium supplemented with 10% FBS. The next day, a transfection mixture was assembled containing 50 ng of TAL-PBx expression vector, 225 ng of reporter plasmid, 225 ng of donor plasmid, and 1 μL of JetPrime transfection reagent in a total volume of 50 μL of JetPrime buffer. This mixture was added to the HEK293T cells, which were incubated for 4 days at 37° C. and 5% CO2, with the cells split 1:6 on day 1. Four days after transfection, the percentage of GFP-positive cells was measured for each sample. The results are shown in FIG. [Table 32]
[0404] As shown in Figure 19 and Table 27, all of the new constructs were able to catalyze site-specific transposition above background levels of 12bp, 13bp, and 14bp spacer targets at several levels for some TAL-PBx constructs, outperforming the benchmarks. The broad range of activity across a wide range of deletions and various spacer lengths allows flexibility in designing TAL-PBx fusion constructs that can target a diverse set of genomic targets with various spacing and TAL-PBx designs.
[0405] Example 24: Construction of TAL-PBx fusions of various sizes, including deletions of the TAL C-terminal domain Naturally occurring TALs contain a 278 amino acid C-terminal domain (SEQ ID NO: 77). The first exemplary TAL-PBx fusion constructed contained a truncated C-terminal domain that maintained 63 amino acids (SEQ ID NO: 76). To explore the role of the size of the C-terminal domain, alternative truncations of the TAL C-terminal domain were designed. Truncated TAL C-terminal domains that maintained 13, 23, 33, 43, 53, or 73 amino acids were constructed (SEQ ID NOs: 471-476). These C-terminal domain deletions were used to replace the 63 amino acid TAL C-terminal domain in GFP1 right TAL-PBx (sequence number 192) to generate GFP1 right TAL-PBx+13 (sequence number 477), GFP1 right TAL-PBx+23 (sequence number 478), GFP1 right TAL-PBx+33 (sequence number 479), GFP1 right TAL-PBx+43 (sequence number 480), GFP1 right TAL-PBx+53 (sequence number 481), and GFP1 right TAL-PBx+73 (sequence number 482).
[0406] To test the effect of a GGGGS linker sequence located between the TAL and PBx sequences, a second set of constructs was generated containing the 13, 23, 33, 43, 53, 63, and 73 amino acid C-terminal domains of the TAL lacking the GGGGS linker to generate GFP1 right TAL-PBx+13-GGGGS linker (SEQ ID NO: 483), GFP1 right TAL-PBx+23-GGGGS (SEQ ID NO: 484), GFP1 right TAL-PBx+33-GGGGS (SEQ ID NO: 485), GFP1 right TAL-PBx+43-GGGGS (SEQ ID NO: 486), GFP1 right TAL-PBx+53-GGGGS (SEQ ID NO: 487), GFP1 right TAL-PBx+63-GGGGS (SEQ ID NO: 488), and GFP1 right TAL-PBx+73-GGGGS (SEQ ID NO: 489). Additionally, an array of truncated TAL C-terminal domains was used in combination with some of the alternative PBx N-terminal variants constructed in Example 23. The 63 amino acid TAL C-terminal domain in GFP1 Right TAL-PBx delta85 (SEQ ID NO: 452) was replaced with alternative TAL C-terminal domain truncations to generate GFP1 Right TAL-PBx delta85+13 (SEQ ID NO: 490), GFP1 Right TAL-PBx delta85+23 (SEQ ID NO: 491), GFP1 Right TAL-PBx delta85+33 (SEQ ID NO: 492), GFP1 Right TAL-PBx delta85+43 (SEQ ID NO: 493), GFP1 Right TAL-PBx delta85+53 (SEQ ID NO: 494), GFP1 Right TAL-PBx delta85+73 (SEQ ID NO: 495). The 63 amino acid TAL C-terminal domain in GFP1 Right TAL-PBx delta 88 (sequence number 455) was replaced with alternative TAL C-terminal domain truncations to generate GFP1 Right TAL-PBx delta 88+13 (sequence number 496), GFP1 Right TAL-PBx delta 88+23 (sequence number 497), GFP1 Right TAL-PBx delta 88+33 (sequence number 498), GFP1 Right TAL-PBx delta 88+43 (sequence number 499), GFP1 Right TAL-PBx delta 88+53 (sequence number 500), GFP1 Right TAL-PBx delta 88+73 (sequence number 501).The 63 amino acid TAL C-terminal domain in GFP1 Right TAL-PBx delta 99 (sequence number 465) was replaced with alternative TAL C-terminal domain truncations to generate GFP1 Right TAL-PBx delta 99+13 (sequence number 502), GFP1 Right TAL-PBx delta 99+23 (sequence number 503), GFP1 Right TAL-PBx delta 99+33 (sequence number 504), GFP1 Right TAL-PBx delta 99+43 (sequence number 505), GFP1 Right TAL-PBx delta 99+53 (sequence number 506), GFP1 Right TAL-PBx delta 99+73 (sequence number 507). The 63 amino acid TAL C-terminal domain in GFP1Right TAL-PBxdelta103 (SEQ ID NO:469) was replaced with alternative TAL C-terminal domain truncations to generate GFP1Right TAL-PBxdelta103+13 (SEQ ID NO:508), GFP1Right TAL-PBxdelta103+23 (SEQ ID NO:509), GFP1Right TAL-PBxdelta103+33 (SEQ ID NO:510), GFP1Right TAL-PBxdelta103+43 (SEQ ID NO:511), GFP1Right TAL-PBxdelta103+53 (SEQ ID NO:512), GFP1Right TAL-PBxdelta103+73 (SEQ ID NO:513). These constructs are shown diagrammatically in FIG.
[0407] Site-specific integration (percentage of GFP positive cells) was measured for each construct and the results are shown in FIG. 21 and Table 29. [Table 33-1] [Table 33-2]
[0408] As shown in Figure 21 and Table 29, 88 and 89 amino acid N-terminal deletions of PBx were often superior to 93, 99, and 103 amino acid truncations. Furthermore, TAL C-terminal domains of 73, 63, 53, and 43 amino acids in length were often superior to TAL C-terminal domains of 33, 23, and 13 amino acids. Various combinations outperformed the benchmarks for different target spacer lengths, allowing flexibility in designing TAL-PBx fusion constructs to target diverse genomic loci.
[0409] Example 25: Site-saturation mutagenesis and relative integration-deletion activity of PBx R372A and K372A mutations The mutations R372A and K375A within the integration domain of the PiggyBac transposase amino acid sequence confer an integration-deficient form of the transposase while maintaining cleavage function. It has been proposed that converting positively charged lysine and arginine residues to neutrally charged alanines reduces the affinity of the transposase for the negatively charged DNA backbone adjacent to its TTAA integration site.
[0410] As a strategy to increase site-specific transposition, further mutations at these "PBx" positions 372 and 375 were explored as a way to titrate PBx transposase affinity for DNA. Site saturation mutagenesis (or SSM) is a technique in which an amino acid at a given position is mutated to all 19 other amino acids. SSM was performed at position 372 in the context of a TAL-PBx fusion containing a K375A mutation. Additionally, SSM was performed at position 375 in the context of a TAL-PBx fusion containing a R372A mutation. Specifically, SSM was performed in a GFP1 right TAL-PBx fusion (SEQ ID NO: 192). In the context of this TAL-PBx fusion, PBx positions 372 and 375 correspond to positions 849 and 852 of TAL-PBx. SSM resulted in a 372 mutant at position 19 (SEQ ID NO: 411-429) and a 375 mutant at position 19 (SEQ ID NO: 430-448).
[0411] An "all-in-one site-specific cleavage / integration episomal reporter" system was developed to test the ability of the novel mutants to catalyze site-specific transposition (Figure 22). This episomal reporter system contains a plasmid containing a transposon donor along with all the transposon integration sites on the same plasmid. The transposon is composed of, from 5' to 3', a TTAA sequence, a 35 bp PiggyBac minimal 5' ITR (SEQ ID NO: 319), a CMV promoter, a 63 bp PiggyBac minimal 3' ITR (SEQ ID NO: 320), and a TTAA sequence. The transposon in this plasmid is initiated by the EF1a promoter and is followed by a polyadenylation signal sequence, disrupting the GFP open reading frame. The vector also contains, in reverse orientation, a polyA and transcription pause site, a GFP1 right target sequence, and a TTAA integration site flanked by a 13 bp spacer, followed by a PEST-destabilized mScarlet reporter and a polyadenylation signal sequence. This "all-in-one site-specific cleavage / integration episomal reporter" (SEQ ID NO: 449) should not express GFP and should not or should not express mScarlet when transfected into cells alone. Upon transposon cleavage (catalyzed by SPB, PBx, or ssSPB), GFP should be expressed. Upon site-specific integration of the CMV promoter containing transposon into its target site upstream of mScarlet, mScarlet should be expressed above background levels (Figure 22).
[0412] Each of the TAL-PBx SSM mutant expression vectors was co-transfected with the all-in-one site-specific cleavage / integration episomal reporter into HEK293T. Briefly, a transfection mix was assembled containing 50ng of mutant TAL-PBx, 50ng of reporter plasmid, and 0.3μL of Transit2020 transfection reagent in a total volume of 20μL of serum-free OptiMem medium. To this, approximately 60,000 HEK293T cells were added in 180μL of DMEM medium supplemented with 10% FBS, and then 80μL of this transfection mix was plated in duplicate in a clear-bottom 96-well plate and incubated at 37℃, 5% CO2. As a control, the original R372A, K375A TAL-PBx, plus SPB, were transfected instead of the SSM mutant TAL-PBx. GFP and mScarlet fluorescence were detected using an Incucyte live cell analysis instrument. The percentage of fluorescent cells for the cleavage (GFP) and site-specific integration (mScarlet) reporters, respectively, is shown in Figure 23 and Table 30. [Table 34]
[0413] As shown in Figures 23A and B and Table 29, several of the SSM mutants conferred site-specific integration similar to or greater than the benchmark R372A, K375A TAL-PBx fusion, indicating that the integration / cleavage activity of the PBx sequence is titratable depending on the position of the amino acids at positions 372 and 375.
[0414] Example 26: Integration of TTAA genomic sites suitable for site-specific integration and design of zinc finger motif-PBx fusions targeting specific TTAA genomic locations Zinc finger motif PBx (ZFM-PBx) fusion proteins require precise spacing (6bp, 7bp, or 8bp) between the zinc finger binding site and the TTAA incorporation site for efficient site-specific integration. ZFM-PBx fusions also require two zinc finger binding sites adjacent to the target TTAA incorporation site, facilitating greater activity. We developed a tailored software program that selects zinc finger targetable TTAAs along the genome, taking into account the published CoDA zinc finger library, as well as the spacing requirements between the zinc finger motif binding site and the TTAA. We selected three TTAA target sites on the human genome (SEQ ID NOs: 526-528). A total of six zinc finger PBx fusions were generated to target these three sites (Table 31). [Table 35]
[0415] Two sites are located on chromosome 21 (referred to as chr21-1, chr21-2) (SEQ ID NO: 526-527) and one site is located on chromosome 17 (referred to as chr17-1) (SEQ ID NO: 528), as shown in Table 31. A total of six ZFM-PBx fusions targeting these three endogenous sites were generated by Gibson assembly.
[0416] To determine whether the newly generated ZFM-PBx fusions are functional and capable of site-specific integration, an episomal site-specific integration assay was performed using a split GFP reporter syste...
Claims
1. A fusion protein comprising, from N- to C-terminus, a DNA targeting domain and a first transposase domain comprising the sequence set forth in SEQ ID NO:544, wherein said first transposase domain comprises a deletion of most of the N-terminal amino acids 83-103 of SEQ ID NO:
544.
2. The fusion protein of claim 1 , wherein the DNA targeting domain comprises three zinc finger motifs.
3. 2. The fusion protein of claim 1, wherein the DNA targeting domain comprises one or more TAL domains.
4. The fusion protein of claim 3, wherein the TAL domain comprises a sequence set forth in any one of SEQ ID NOs: 107 to 110.
5. 2. The fusion protein of claim 1, wherein the DNA targeting domain binds to a nucleic acid sequence encoding GFP zinc finger 268 (ZFM268), phenylalanine hydroxylase (PAH), beta-2-microglobulin (B2M), or a LINE1 repeat element.
6. 2. The fusion protein of claim 1, wherein the first transposase domain and the DNA targeting domain are connected by a linker.
7. The fusion protein of claim 6 , wherein the linker comprises the sequence GGGGS.
8. 2. The fusion protein of claim 1, wherein the first transposase domain comprises an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103.
9. The fusion protein of claim 1, wherein the transposase domain comprises a sequence set forth in any one of SEQ ID NOs: 86 to 106.
10. 2. The fusion protein of claim 1, wherein the first transposase domain comprises at least one mutation selected from the group consisting of: (a) M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R; or (b) L204D, L204E, K500D, K500E, R504E, and R504D.
11. 2. The fusion protein of claim 1, further comprising a second transposase domain C-terminal to the first transposase domain, wherein the second transposase domain comprises the sequence set forth in SEQ ID NO:
544.
12. 12. The fusion protein of claim 11, wherein the second transposase domain comprises a deletion of N-terminal amino acids 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103 of SEQ ID NO:
544.
13. 12. The fusion protein of claim 11, wherein the second transposase domain comprises at least one mutation selected from the group consisting of: (a) M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R; or (b) L204D, L204E, K500D, K500E, R504E, and R504D.
14. A polynucleotide comprising a nucleic acid sequence encoding the fusion protein of claim 1.
15. A vector comprising the polynucleotide of claim 14.
16. 1. A method for integrating a transgene into a genomic target site in a cell, comprising: The method comprises introducing into the cell the fusion protein of claim 1 and a transposon; The method as described above, wherein the transposon comprises, from 5' to 3', a 5' ITR, the transgene, and a 3' ITR.
17. 17. The method of claim 16, wherein the transposon further comprises an exogenous promoter between the 5' ITR and the transgene.
18. 17. The method of claim 16, wherein the transgene encodes a detectable marker.
19. 19. The method of claim 18, wherein the detectable marker is GFP.
20. 17. The method of claim 16, wherein the transgene is a gene that is not expressed by the cell prior to introduction of the fusion protein and the transposon.
21. 17. The method of claim 16, wherein the genomic target site is located on chromosome 17 or 21.
22. 17. The method of claim 16, wherein the genomic target site is located within the B2M gene.
23. 17. The method of claim 16, wherein the genomic target site is located within a repetitive element.
24. 24. The method of claim 23, wherein the repeating elements are LINE elements.
25. 17. The method of claim 16, wherein the genomic target site is located within an intron of a gene.
26. 26. The method of claim 25, wherein the genomic target site is located within the intron of the PAH gene.
27. 17. The method of claim 16, wherein the cell is in vivo.
28. 1. A method for modifying the genome of a cell, comprising: The method comprises providing the cell with the fusion protein of claim 1; The method, wherein the cell comprises a modified binding site comprising, from 5' to 3', the reverse sequence of the target site for the DNA targeting domain, a first spacer, a TTAA target incorporation site for SPB, a second spacer, and the complement of the sequence of the target site for the DNA targeting domain.
29. 1. An integration cassette for site-specific transposition of a nucleic acid into the genome of a cell comprising a nucleic acid comprising or consisting of at least one upstream zinc finger motif DNA-binding domain binding site ("ZFM-DBD") and a central transposon ITR integration site TTAA sequence flanked by at least one downstream ZFM-DBD, The integration cassette, wherein each of the upstream and downstream ZFM-DBDs is separated from the TTAA sequence by 7 base pairs.
30. 1. An integration cassette for site-specific transposition into the genome of a cell, comprising or consisting of a nucleic acid comprising or consisting of a central transposon ITR integration site TTAA sequence flanked by an upstream TAL array target sequence and a downstream TAL array target sequence of the nucleic acid, The integration cassette, wherein each of the upstream and downstream TAL array target sequences is separated from the TTAA sequence by 12 to 14 base pairs.
31. A central transposon ITR integration site TTTA flanked by an upstream TAL array target sequence and a downstream TAL array target sequence An integration cassette for site-specific transfer of a nucleic acid into the genome of a cell, comprising a nucleic acid comprising an AA sequence, The upstream and downstream TAL array target sequences are each separated by 12 base pairs from the TTTAAA sequence.
32. 31. The integration cassette of claim 30, wherein each of the at least one upstream and downstream TAL array target site sequences is identical.
33. 31. The integration cassette of claim 30, wherein each of the at least one upstream and downstream TAL array target site sequences is different.
34. 31. The integration cassette of claim 30, wherein each of the at least one upstream and downstream TAL array target sites targets a 10 bp sequence of the beta-2-microglobulin gene ("B2M"), the phenylalanine hydroxylase gene ("PAH"), or a LINE1 repeat element.
35. 33. The integration cassette of claim 32, wherein the at least one upstream TAL array target sequence and the at least one downstream TAL array target sequence bind to a nucleic acid comprising the sequence GCGTGGGCG.
36. 30. A cell comprising the integration cassette of claim 29 stably integrated into the genome of said cell.
37. 37. A method for site-specific transfer of a DNA molecule into the genome of a cell, comprising administering to the cell of claim 36: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed in the cell; and b) a DNA molecule containing a transposon wherein the expressed fusion protein integrates the transposon into the TTAA sequence of the stably integrated integration cassette by site-specific transposition.
38. 37. A method for producing a modified cell by site-specific transposition, comprising: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase, wherein the fusion protein is expressed in the cell; and b) a DNA molecule containing a transposon wherein the expressed fusion protein integrates the transposon by site-specific transposition into the TTAA sequence of the stably integrated integration cassette, thereby generating the modified cell.