Transposases and uses thereof

Fusion proteins with TAL Arrays and modified transposase domains provide precise and efficient site-specific transposition into rDNA repeats, addressing the need for accurate gene replacement by achieving high on-target integration fidelity for therapeutic interventions.

WO2026064550A1PCT designated stage Publication Date: 2026-03-26POSEIDA THERAPEUTICS INC
View PDF 50 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

There is a need for site-specific transposases with increased fidelity for gene replacement that can accurately target specific gene loci, as existing methods lack the necessary precision and efficiency.

Method used

Fusion proteins comprising a TAL Array targeting rDNA repeats and a transposase domain with N-terminal deletions, nucleolar trafficking sequences, and protein stabilization domains are developed, which include thermostable and hyperactive piggyBac transposase domains with dual cysteine-rich domains for site-specific transposition into rDNA repeats or LINE1 elements.

Benefits of technology

The fusion proteins achieve high transposition fidelity, integrating transgenes into rDNA target sites with a ratio of at least 89% on-target integration, enabling precise genome modification for therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025047035_26032026_PF_FP_ABST
    Figure US2025047035_26032026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure generally relates to fusion proteins comprising TAL Arrays targeting a repetitive element and a transposase domain comprising amino terminal deletions, as well as dual cysteine rich domains (CRD), for targeting site-specific transposition into ribosomal DNA (rDNA) repeats or LINE1 repetitive elements, polynucleotides and vectors encoding the fusion proteins, and methods of use therefor.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 000218-0153-WO1 TRANSPOSASES AND USES THEREOF CROSS-REFERENCE TORELATEDAPPLICATIONS

[0001] This application claims benefit of and priority to U.S. Provisional Patent Application No.63 / 696,577, filed September 19, 2024, and U.S. Provisional Patent Application No. 63 / 716,833, filed November 6, 2024, the contents of each of which are herein incorporated by reference in their entireties. SEQUENCE LISTING

[0002] The instant application contains a Sequence Listing which has been submitted in XML format and is herein incorporated by reference in its entirety. Said XML copy, created on September 18, 2025 is named “000218-0153-WO1-SL” and is 420,393 bytes in size. FIELD

[0003] This disclosure generally relates to fusion proteins targeting rDNA comprising a TAL Array and a transposase domain. Also provided are methods of use of the fusion proteins for site-specific transposition of DNA and AAV virions into rDNA target sites. BACKGROUND

[0004] Transposases may be used to introduce non-endogenous DNA sequences into genomic DNA, and are in many ways advantageous to other methods gene editing. However, there remains an unmet need for site-specific transposases for use in, e.g., gene replacement, that have increased fidelity at specific gene loci. SUMMARY

[0005] In one aspect, provided herein is a fusion protein, comprising, in N-terminal to C- terminal direction: (a) a TAL Array targeting an rDNA repeat; and (b) a transposase of SEQ ID NO: 6 comprising an N-terminal deletion. In some embodiments, the transposase comprises a N-terminal deletion comprising a deletion of amino acids 1-74, 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1- 101, 1-102 or 1-103 of SEQ ID NO: 6.

[0006] In some embodiments, the fusion protein further comprises a nucleolar trafficking sequence (NTS), wherein the C-terminus of the NTS is attached to the N-terminus of the TAL Array. In some embodiments, the NTS is an RNA polymerase I (PolI) subunit. In some embodiments, the PolI subunit is an A12.2 subunit, a A35.4 subunit, an A43 subunit or anAttorney Docket No.: 000218-0153-WO1 A49 subunit. In some embodiments, the PolI subunit comprises the nucleic acid sequence of any one of SEQ ID NO: 119-122. In some embodiments, the NTS is an R2 retrotransposon transposase sequence. In some embodiments, the R2 retrotransposon transposase sequence comprises the nucleic acid sequence of SEQ ID NO: 123 or 124.

[0007] In some embodiments, the fusion protein further comprises a protein stabilization domain (PSD). In some embodiments, the PSD comprises the amino acid sequence of SEQ ID NO: 190.

[0008] In some embodiments, the fusion protein further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS sequence comprises the nucleic acid sequence of SEQ ID NO: 191.

[0009] In some embodiments, the fusion protein further comprises a GS or GGGGS (SEQ ID NO: 125) linker positioned (a) between the TAL Array and the SPB transposase; (b) between nucleolar trafficking sequence and the TAL Array; and / or (c) between the PSD and the TAL Array.

[0010] In some embodiments, the SPB transposase comprises the amino acid sequence of any of SEQ ID Nos: 8-28 or 208.

[0011] In some embodiments, the TAL Array comprises one or more TAL domains targeting an rDNA 45S repeat, an rDNA 18S repeat, an rDNA 5.8S repeat or an rDNA 28S repeat. In some embodiments, each of the one or more TAL domains targets an rDNA repeat comprising the nucleic acid sequence of any one of SEQ ID NOs: 59 - 80. In some embodiments, the TAL Array comprises the amino acid sequence of any one of SEQ ID NO: 81-114.

[0012] In some embodiments, the fusion protein comprises the amino acid sequence of any one of SEQ ID NOs: 126-159.

[0013] In another aspect, provided herein is a polynucleotide comprising a nucleic acid sequence encoding a fusion protein disclosed herein. In another aspect, provided herein is a vector comprising a polynucleotide provided herein..

[0014] In another aspect, provided herein is a pharmaceutical composition comprising (a) a fusion protein disclosed herein, a polynucleotide disclosed herein, or a vector disclosed herein (b) a pharmaceutically acceptable excipient.

[0015] In another aspect, provided herein is a method of modifying the genome of a cell, the method comprising introducing into the cell (a) a polynucleotide comprising a nucleic acid encoding a fusion protein comprising a transposase of SEQ ID NO: 6 comprising an N- terminal deletion; and a TAL Array targeting an rDNA repeat, wherein the C-terminus of theAttorney Docket No.: 000218-0153-WO1 TAL Array is fused to the N-terminus of the SPB; and (b) a DNA molecule comprising a transposon, herein the transposon comprises, in 5’ to 3’ order: a 5’ITR, the transgene, and a 3’ ITR, wherein the transgene is integrated into a TTAA site flanked by a left rDNA target site and a right rDNA target site.

[0016] In some embodiments, the transposon is a small transposon comprising symmetrical ITRs. In some embodiments, the transposon comprises the sequence of SEQ ID NO: 202. In some embodiments, the transposon further comprises an exogenous promoter operatively linked to the transgene.

[0017] In some embodiments, the transgene encodes a protein the expression of which is beneficial for treating a metabolic disorder, hemophilia, cancer or PKU.

[0018] In some embodiments, the TTAA site comprises the nucleic acid sequence of any one of SEQ ID NOs 36-58. In some embodiments, the rDNA repeat comprises the nucleic acid sequence of any one of SEQ ID NOs: 59-80. In some embodiments, the left rDNA target site comprises the sequence of SEQ ID NO: 63 and the right rDNA target site comprises the sequence of SEQ ID NO: 64. In some embodiments, the left rDNA target site comprises the sequence of SEQ ID NO: 79 and the right rDNA target site comprises the sequence of SEQ ID NO: 80. In some embodiments, the cell is in vivo. In some embodiments, the transgene is integrated into the genome of the cell with a transposon fidelity of on target to off target transposition integration ratio of at least 89%.

[0019] In another aspect, provided herein is a method for generating an engineered cell by site-specific transposition, comprising introducing into a cell (a) a nucleic acid encoding a fusion protein comprising a rDNA TAL Array and a transposase domain comprising the sequence of SEQ ID NO: 6; and (b) a DNA molecule comprising a transposon; wherein the transposon is integrated by site-specific transposition into a TTAA sequence of an rDNA repeat. In some embodiments, the DNA molecule comprising a transposon is a nanoplasmid or an AAV virion. In some embodiments, the nucleic acid encoding the fusion protein is mRNA introduced into the cell by a lipid nanoparticle. In some embodiments, the nucleic acid encoding the fusion protein and / or the DNA molecule comprising the transposon are introduced into the cell by a lipid nanoparticle. BRIEF DESCRIPTION OF DRAWINGS

[0020] FIG.1 shows a schematic illustrating a TAL-PBx Plus transposase dimer binding to a genomic target site comprising a TTAA integration site, upstream and downstream TAL binding sites, and intervening spacer sequences.Attorney Docket No.: 000218-0153-WO1

[0021] FIG.2 shows a schematic illustrating of human acrocentric chromosomes 13, 14,15, 21, and 22 comprising the 45S rDNA repeat elements that each encode the 18S, 5.8S, 28S rDNA sequences.

[0022] FIG.3 shows a schematic illustrating the addition of polI or R2 nucleolus-targeting domains to PBx Plus, TAL Array - PBx Plus.

[0023] FIG.4 shows a schematic illustrating PBx, TAL Array – PBx Plus, R2-PBx, R2- TAL Array - PBx, PolI-PBx, and PolI - TAL Array – PBx Plus fusion proteins. DETAILED DESCRIPTION

[0024] Provided herein are fusion proteins comprising TAL Arrays fused to thermostable, hyperactive, integration deficient PBx transposase domains comprising dual cysteine rich domains (CRD; “PBx Plus”) for targeting site-specific transposition into ribosomal DNA (rDNA) repeats or LINE1 repetitive elements, polynucleotides and vectors encoding the fusion proteins, and methods of use therefor. Transposase Domains Comprising Dual Cysteine Rich Domains (CRD) The transposases provided herein generally comprise a dimerization and DNA binding domain (DDBD), which allows the transposase to bind to the inverted terminal repeats (ITRs) at the ends of the transposon and is also involved in protein dimerization The transposes generally further comprise a catalytic domain, which catalyzes the transposition reaction, and an insertion domain which allows the transposase to interact with the target DNA into which the transposon integrates. At the C-terminus, a cysteine rich domain (CRD) may be attached to the rest of the transposase, for example, the CRD may be attached via a linker approximately 20 amino acids in length. In some embodiments, one or both of the CRDs comprise the sequence STEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 7)

[0025] Like the DDBD, the CRD is also believed to be involved in ITR binding and protein dimerization. Upon binding of a PiggyBac transpose to the left end (LE) and right end (RE) ITRs, protein dimerization brings the two ITRs together to form a synaptic hairpin complex. This arrangement is believed to be required for transposition to occur as the transposase cuts TTAA sequences flanking the transposon in trans.

[0026] The DDBD and the CRD bind to the ITRs in a sequence-specific manner. The DDBD interacts with about 10bp of DNA located 6bp in from the TTAA sequences flankingAttorney Docket No.: 000218-0153-WO1 the transposon. Binding of the DDBD is symmetrical, with the DDBD of one transposase monomer binding to the LE ITR and the DDBD of the second monomer binding to the RE ITR. The CRDs of the first dimer bind to a 19bp sequence of the LE ITR found immediately distal to the DDBD binding site. CRD binding is asymmetrical, with both CRDs of the first dimer interacting with the 19bp sequence on the LE ITR only. The RE ITR contains a second DDBD binding sequence followed by a 19bp CRD binding sequence starting 34bp in from the TTAA. The first dimer that binds proximal to the TTAAs catalyzes the transposition reaction.

[0027] In one aspect, provided herein are transposase domains comprising a second cysteine rich domain (CRD) in addition to the first endogenous CRD. In some embodiments, the second CRD is linked to the C-terminus of the transposase to generate transposase domains comprising dual CRDs. In some embodiments, the transposase domains comprising dual CRDs comprise an N-terminal deletion. In some embodiments, the transposase domain comprising dual CRDs is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a hyperactive piggyBac transposase domain. In preferred embodiments, the transposase domain comprising dual CRDs is a Super piggyBac® transposase domains (SPB). Non-limiting examples of SPB transposases are described in detail in U.S. Patent No.6,218,182; U.S. Patent No.6,962,810; U.S. Patent No.8,399,643 and PCT Publication No. WO 2010 / 099296, each of which is incorporated herein by reference in its entirety for examples of transposase domains that may be used in connection with the fusion proteins described herein.

[0028] In some embodiments, the transposase domains and fusion proteins provided herein may comprise an in-frame nuclear localization sequence (NLS). Examples of transposases fused to a nuclear localization signal are disclosed in U.S. Patent No.6,218,185; U.S. Patent No.6,962,810, U.S. Patent No.8,399,643 and WO 2019 / 173636. In some embodiments, the NLS comprises the sequence of PKKKRKV (SEQ ID NO: 191). In certain aspects, the in- frame NLS is located upstream (N-terminal) of the transposase domain comprising an N- terminal deletion.

[0029] In general, the NLS is preferably located at the N-terminal end of a fusion protein. In some embodiments, the NLS is fused or linked to the N-terminus of a transposase domain. In some embodiments, the NLS is fused or linked to the N-terminus of a DNA targeting domain. In some embodiments, the NLS is fused or linked to the N-terminus of a PSD.

[0030] In certain aspects, the in-frame NLS is fused directly to the amino terminus of the PBx Plus transposase domain comprising an N-terminal deletion. In some embodiments, theAttorney Docket No.: 000218-0153-WO1 NLS is attached to the N-terminus of a PBx Plus transposase domain comprising an N- terminal deletion via a linker (e.g., a GGGGS linker (SEQ ID NO: 125) or a GGS linker).

[0031] In some embodiments, an initiator methionine is introduced before the NLS. In some embodiments, additional alanine residues are introduced before and / or after the NLS to ensure in-frame translation. Transposase Domains Comprising Thermostability and / or Hyperactivity Mutations A. Hyperactivity Mutations

[0032] The transposase domains provided herein may comprise one or more hyperactivity mutations.

[0033] In some embodiments, an SPB or PBx transposase domain provided herein comprises one or more hyperactivity mutations (in additional to those present in SPB). In some embodiments, an SPB or PBx transposase domain provided herein comprises a S103P mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a R372H mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a S509G mutation. In some embodiments, an SPB transposase domain provided herein comprises a N571S mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a S103P mutation, a R372H mutation, a S509G mutation and a N571S mutation.

[0034] In some embodiments, a SPB or PBx transposase domain provided herein comprises one or more hyperactivity mutations (in additional to those present in SPB & PBx). In some embodiments, a SPB or PBx transposase domain provided herein comprises a S103P mutation. In some embodiments, a SPB or PBx transposase domain provided herein comprises a R372H mutation. In some embodiments, a SPB or PBx transposase domain provided herein comprises a S509G mutation. In some embodiments, a SPB or PBx transposase domain provided herein comprises a N571S mutation. In some embodiments, a PBx transposase domain provided herein comprises a S103P mutation, a R372H mutation, a S509G mutation, and a N571S mutation. In some embodiments, a SPB or PBx transposase domain comprising the S103P, S509G, and N571S mutations. In some embodiments, a SPB or PBx transposase domain comprising the S103P, S509G, N571S, and R372H mutations comprises the amino acid sequence set forth in SEQ ID NO: 29.Attorney Docket No.: 000218-0153-WO1

[0035] The S103P, S509G, R372H and / or N571S mutation can also be introduced into any of the truncated dual CRD PBx Plus sequences provided herein. A person of skill will appreciate that the numbering of the residues depends on the size of the truncation.

[0036] Similarly, the S103P, S509G, R372H and / or N571S mutation may be introduced into any of the fusion proteins described below. B. Thermostability mutations

[0037] The transposase domains provided herein may comprise one or more thermostability mutations.

[0038] In some embodiments, an SPB or PBx transposase domain provided herein comprises one or more thermostability mutations. In some embodiments, an SPB or PBx transposase domain provided herein comprises one or more of the following mutations: I182L,S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D, M298L, or any subgroup thereof. In some embodiments, an SPB or PBx transposase domain provided herein comprises a I182L mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a S301A mutation. In some embodiments, an SPB transposase domain provided herein comprises a C420M mutation. In some embodiments, an SPB transposase domain provided herein comprises a M185L mutations. In some embodiments, an SPB or PBx transposase domain provided herein comprises a R315K mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a D421H mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a F200W mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a Q318G mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a N427D mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a V207I mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a E331R mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a Q434E mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a M226F mutation. In some embodiments, an SPB transposase domain provided herein comprises a V336I mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a V436I mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a I231T mutation. In someAttorney Docket No.: 000218-0153-WO1 embodiments, an SPB or PBx transposase domain provided herein comprises a S373K mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a I474L mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a V240K mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a V381E mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a K500R mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a Q254N mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a T392S mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a S513P mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a A263E mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a A411N mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a K525P mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a S289A mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a S419T mutation. In some embodiments, an SPB or PBx transposase domain provided herein comprises a R567D mutation.

[0039] In some embodiments, the PBx transposase domain sequence comprises the M298L mutation comprising or consisting essentially of SEQ ID NO: 30. C. Combinations of Hyperactivity and Thermostability Mutations

[0040] The hyperactivity mutations and thermostability mutations described herein may be freely combined. Thus, a transposase domain may comprise one, two or all of the hyperactivity mutations S103P, R372H S509G and / or N571S and any or all of the thermostability mutations I182L,S301A, C420M, M185L, R315K, D421H, F200W, Q318G, N427D, V207I, E331R, Q434E, M226F, V336I, V436I, I231T, S373K, I474L, V240K, V381E, K500R, Q254N, T392S, S513P, A263E, A411N, K525P, S289A, S419T, R567D and / or M298L.

[0041] In some embodiments, an SPB or PBx transposase domain provided herein comprises an M298L thermostability mutation and one or more of the S103P, R372H S509G and / or N571S hyperactivity mutations. In some embodiments, an SPB or PBx transposase domain provided herein comprises an M298L thermostability mutation and the S103P, S509G and N571S hyperactivity mutations. In some embodiments, a PBx transposase domainAttorney Docket No.: 000218-0153-WO1 comprising the S103P, S509G, N571S, R372H, and M298L mutations comprises the amino acid sequence set forth in SEQ ID NO: 31. D. Integration-Deficiency Mutations

[0042] In some embodiments, the transposase domain is integration deficient. An integration deficient transposase domain can excise its corresponding transposon, but that integrates the excised transposon at a lower frequency than a corresponding wild type transposase. Examples of integration deficient transposases are disclosed in U.S. Patent No. 6,218,185; U.S. Patent No.6,962,810, U.S. Patent No.8,399,643 and International Patent Application Publication No. WO 2019 / 173636 each of which is incorporated herein by reference in its entirety for examples of transposase domains that may be used in connection with the fusion proteins described herein. A list of integration deficient amino acid substitutions is disclosed in US Patent No.10,041,077, which is incorporated herein by reference in its entirety for examples of mutations that may be introduced into the transposase domains described herein. A wildtype SPB may be rendered integration deficient by introducing mutations, for example, K93A, R372A, K375A, R376A and / or D450N (relative to SEQ ID NO: 1, with numbering beginning at residue 12). It is believed that the introduction of mutations R372A, K375A, R376A and D450N renders the transposase integration deficient, but retains the excision function. The amino acid sequence of an integration deficient PBx transposase domain not comprising an NLS is set forth in SEQ ID NO: 5: GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTS SGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRS QRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEI YAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIR PTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKY GIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDN WFTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVS YKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSR KTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRK RLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCK KCKKVICREHNIDMCQSCF (SEQ ID NO: 5).Attorney Docket No.: 000218-0153-WO1 Illustrative Transposase Domain Sequences

[0043] An exemplary wildtype SPB sequence with an NLS and a CRD is shown in SEQ ID NO: 1 with the NLS shown in italics, hyperactive mutations shown in bold underline, and the CRD underlined. The numbering of sequence of the SPB transposase domain for the purpose of describing deletions and mutations begins at residue 12 of SEQ ID NO: 1: MAPKKKRKVGGGGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEA FIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRR SRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTS ATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIR CLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPF RVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPV HGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSM FCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLD QMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLY MSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSK IRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 1)

[0044] An exemplary wildtype SPB sequence comprising dual CRDs with the second, C- terminal CRD attached via an AGGG peptide linker sequence (SEQ ID NO: 2) to the first, endogenous CRD is shown in SEQ ID NO: 3 with the NLS shown in italics, hyperactivemutations shown in bold, the linker sequence shown in lowercase font and each of the dualCRDs underlined. The numbering of sequence of the SPB transposase domain for the purpose of describing deletions and mutations begins at residue 12 of SEQ ID NO: 3: MAPKKKRKVGGGGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEA FIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRR SRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTS ATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIR CLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPF RVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPV HGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSM FCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLD QMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLY MSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSK IRRKANASCKKCKKVICREHNIDMCQSCFagggSTEEPVMKKRTYCTYCPSKIRRKAN ASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 3)Attorney Docket No.: 000218-0153-WO1

[0045] The amino acid sequence of an SPB transposase domain comprising dual CRDs not comprising an NLS is set forth in SEQ ID NO: 4. GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTS SGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRS QRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEI YAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIR PTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKY GIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDN WFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSY KPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRK TNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKR LEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKK CKKVICREHNIDMCQSCFAGGGSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVI CREHNIDMCQSCF (SEQ ID NO: 4).

[0046] The transposase domains used in the fusion proteins described herein can be isolated or derived from an insect, vertebrate, crustacean or urochordate as described in more detail in PCT Publication No. WO 2019 / 173636 and PCT / US2019 / 049816. In preferred aspects, the SPB transposase domain is isolated or derived from the insect Trichoplusia ni (GenBank Accession No. AAA87375) or Bombyx mori (GenBank Accession No. BAD11135).

[0047] The amino acid sequence of thermostable, hyperactive, an integration deficient PBx Plus transposase domain not comprising an NLS but comprising dual CRDs with a second CRD (SEQ ID NO: 7) linked to the C-terminus of PBx sequence via an AGGG linker sequence (SEQ ID NO: 2) is set forth in SEQ ID NO: 6:

[0048] GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVH EVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKPTRRSRVSA LNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDT NEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMD DKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPN KPSKYGIKILMLCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNI TCDNWFTSIPLAKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTSMFCFDGPL TLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVM TCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMGLTSS FMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKASAttorney Docket No.: 000218-0153-WO1 ASCKKCKKVICREHNIDMCQSCFAGGGSTEEPVMKKRTYCTYCPSKIRRKANASCK KCKKVICREHNIDMCQSCF (SEQ ID NO: 6) Transposase Domains Comprising N-Terminal Deletions

[0049] In some embodiments, provided herein are dual CRD transposase domains (e.g., SPB transposase domains or PBx transposase domains) comprising a deletion of a portion of the amino terminus (also referred to as the “N-terminus” or the “N-terminal Domain,” or “NTD”) of the transposase domain.

[0050] In some embodiments, the deleted portion of the N-terminus is about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 100 amino acids or about 115 amino acids. In some embodiments, the deleted portion of the N-terminus is about 15-25 amino acids, about 25-35 amino acids, about 35-45 amino acids, about 45-55 amino acids, about 55-65 amino acids, about 65-75 amino acids, about 75-85 amino acids, about 85-95 amino acids, about 95-105 amino acids, or about 105-120 amino acids.

[0051] In some embodiments, the transposase domain, e.g., a SPB transposase, a PBx transposase or PBx Plus transposase, comprises a deletion of amino acids 1-74 of the N- terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-83 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-84 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-85 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-86 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-87 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-88 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-89 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-90 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-91 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-92 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-93 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-94 of theAttorney Docket No.: 000218-0153-WO1 N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-95 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-96 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-97 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-98 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-99 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-100 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-101 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-102 of the N-terminus relative to SEQ ID NOs: 4-6. In some embodiments, the transposase domain comprises a deletion of amino acids 1-103 of the N-terminus relative to SEQ ID NOs: 4-6.

[0052] An illustrative sequence of a M282L thermostable, S103P, S509G, N571S, and R372H hyperactive, integration-deficient, PBx transposase domain (“PBx Plus”) comprising dual CRDs with a deletion of amino acids 1-85 of the N-terminus is shown in SEQ ID NO: 10: PQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEII SEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFD RSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNY TPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMLCDSGTKYMINGMPYLGRGT QTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVHSN AREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGK PQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHN VSSKGEKVQSRKKFMRNLYMGLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSD DSTEEPVMKKRTYCTYCPSKIRRKASASCKKCKKVICREHNIDMCQSCFAGGGSTEE PVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 10)

[0053] Illustrative sequences of PBx Plus transposase domains comprising a M282L thermostable mutation, a S103P, a S509G, a N571S, and a R372H hyperactive mutations, integration-deficient mutations comprising varying N-terminal deletion lengths and a second piggyBac CRD appended via an AGGG linker are set forth in SEQ ID NOs: 8-28, and 208 in Table 1.Attorney Docket No.: 000218-0153-WO1 Table 1: Illustrative sequences of PBx Plus Thermostable, Hyperactive, N-terminally deleted, Dual CRD PBx DomainsAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Design of TAL Arrays Targeting rDNA Repeats and LINE1 Elements

[0054] The fusion proteins provided herein may comprise one or more TAL arrays that bind to a target region of interest in the genome. In some embodiments, the TAL arrays target ribosomal DNA (rDNA) or a LINE1 element. A. Description and Design of TAL Arrays

[0055] TALEs (Transcription activator-like effectors) from Xanthomonas typically contain a 288 amino acid N-terminus followed by an array of a variable number of ~34 amino acid repeats followed by a 278 amino acid C-terminus; however, truncated versions have been described in the literature (e.g., see Miller et al., Nat Biotechnol 29, 143–148 (2011). The sequence of these ~34 amino acid repeats determines to which target sequence a TAL binds. TALs fused to a FokI nuclease (called TAL effector nucleases or TALENs) most often contain truncations of the N and C terminus compared to a wildtype Fok1 nuclease. For example, the first 152 amino acids of the N-terminus are often removed and the C-terminus is often truncated leaving 63 amino acids.

[0056] TALs contain arrays of 34 amino acids repeated a variable number of times. Two amino acids at position 12 and 13 are varied and determine which nucleotide the TAL repeat will recognize. This feature allows a TAL array to be programed to bind a specific DNA sequence. The amino acids NG recognize T, NI recognize A, NN recognize G or A, HD recognize C, NK recognize G, NS recognize A, C, G or T. Other amino acids within the 34 residue repeat may also be varied. For example position 11 is often changed to an N for repeats that recognize G. Also, positions 4 and 32 are often varied to reduce the repetitivenessAttorney Docket No.: 000218-0153-WO1 of the array but not to determine the binding specificity. The number of 34 amino acid repeats in an array determines the length of the DNA sequence recognized (one protein repeat binds one DNA bp). Furthermore, the last bp is recognized by a “half array” that is 20 amino acids rather than 34.

[0057] In addition, the N-terminal domain of TALs recognizes and requires a T that is located immediately 5’ of the target DNA sequence. Mutations of TAL N-terminal domains have been described in the literature that no longer require a 5’ T (Lamb et al., Nucleic Acids Res.2013 Nov;41(21):9779-85. doi: 10.1093 / nar / gkt754. Epub 2013 Aug 26. PMID: 23980031; PMCID: PMC3834825.) For example, the NT-G mutant requires a 5’G instead of a 5’T while the NT-βN mutant does not require any specific 5’ nucleotide. These mutated N- terminal domain sequences may be used to provide additional sequence options that may be targeted using TAL Arrays.

[0058] Pairs of TAL arrays targeting sequences in the desired gene may be designed and the corresponding modules selected and pooled together using “Golden Gate Assembly,” to assemble in frame each TAL-Array. The DNA sequence encoding TAL Arrays generated herein may be further codon optimized using GeneArt algorithms (Thermo Fisher).

[0059] When designing left and right TAL Arrays comprising a N-terminal domain recognizing a T and a TAL C-terminal domain to be fused to an N-terminal deleted transposase sequence (i.e., TAL-ssSPB or TAL-PBx; described below), one TAL Array recognizes a sequence 5’ of the TTAA and the other TAL Array recognizes a sequence 3’ of the TTAA. Since the sequence 5’ of TTAA is most often different from the sequence 3’ of TTAA in genomic DNA targets, TAL-ssSPB will most often be used as a heterodimer consisting of two different TAL domains that recognize two different DNA sequences. Additionally, the sequence recognized by the TAL Array is not directly adjacent to the TTAA. Instead, it is separated from the TTAA by a spacer of a given bp length, e.g., spacers of 12bp, 13bp or 14 bp. rDNA

[0060] Ribosomal DNA (rDNA) is the genomic DNA that encodes non-coding ribosomal RNA (rRNA). rRNA is highly abundant and is a structural component of ribosomes. In eukaryotes, there are four rRNAs that make up the RNA component of the ribosome. These are the 5S, 5.8S, 18S, and 28S rRNAs.5S rRNA is encoded by 5S rDNA found on Chromosome 1 in humans.18S, 5.8S, and 28S are encoded in a single transcript, 45S, which is further processed down into the three subunits.45S rDNA exists as tandem arrays. InAttorney Docket No.: 000218-0153-WO1 humans, it’s estimated that around 200-1000 copies of the 45S rDNA are found on the short arms of the acrocentric chromosomes 13, 14, 15, 21, 22. As rRNA is essential for protein synthesis, the rDNA repeats are among the most conserved DNA sequences in the genome.

[0061] Ribosomal DNA repeats are located in the nucleolus of a cell on the short arm of acrocentric chromosomes. Fig.2 shows a schematic representation of the human acrocentric chromosomes 13, 14,15, 21, and 22 and a blow up of an exemplary 45S rDNA repeat located on chromosome 13. Each 45S ribosomal DNA repeat element comprises the 18S, 5.8S, 28S rDNA sequences. As shown in Fig.2, approximately 23 TTAA integration sequences are present across each 45S rDNA repeat. It was hypothesized that these TTAA sequences allow for rDNA to be targeted by piggyBac site-specific transposition, given that rDNA contains multiple genomic 45S rDNA repeats and 23 potential TTAA integration sites per 45S rDNA repeat.

[0062] Without wishing to be bound by theory, it is believed that an rDNA targeting TAL array may be designed to target any rDNA sequence upstream or downstream of a TTAA integration site. It will be apparent to a person of skill in the art that any left TAL array sequence for a given target provided herein can be combined with any right TAL array sequence provided herein for the same target.

[0063] In certain aspects, the rDNA TAL Array of the fusion protein comprises one or more TAL domains targeting a nucleic acid sequence flanking an TTAA integration site located in an rDNA 45S repeat. An exemplary rDNA 45S repeat nucleic acid sequence is set forth in SEQ ID NO: 32: GCTGACACGCTGTCCTCTGGCGACCTGTCGCTGGAGAGGTTGGGCCTCCGGATGC GCGCGGGGCTCTGGCCTACCGGTGACCCGGCTAGCCGGCCGCGCTCCTGCTTGAG CCGCCTGCCGGGGCCCGCGGGCCTGCTGCTCTCTCGCGCGTCCGAGCGTCCCGAC TCCCGGTGCCGGCCCGGGTCCGGGTCTCTGACCCACCCGGGGGCGGCGGGGAAG GCGGCGAGGGCCACCGTGCCCCCGTGCGCTCTCCGCTGCGGGCGCCCGGGGCGG CCGCGACAACCCCACCCCGCTGGCTCCGTGCCGTGCGTGTCAGGCGTTCTCGTCT CCGCGGGGTTGTCCGCCGCCCCTTCCCCGGAGTGGGGGGTTGGCCGGAGCCGAT CGGCTCGCTGGCCGGCCGGCCGGCCTCCGCTCCCGGGGGGCTCTTCGTGATCGAT GTGGTGACGTCGTGCTCTCCCGGGCCGGGTCCGAGCCGCGACGGGCGAGGGGCG GACGTTCGTGGCGAACGGGACCGTCCTTCTCGCTCCGCCCCGCGGGGGTCCCCTC GTCTCTCCTCTCCCCGCCCGCCGGCGGTGCGTGTGGGAAGGCGTGGGGTGCGGAC CCCGGCCCGACCTCGCCGTCCCGCCCGCCGCCTTCTGCGTCGCGGGTGCGGGCCG GCGGGGTCCTCTGACGCGGCAGACAGCCCTCGCTGTCGCCTCCAGTGGTTGTCGAAttorney Docket No.: 000218-0153-WO1 CTTGCGGGCGGCCCCCCTCCGCGGCGGTGGGGGTGCCGTCCCGCCGGCCCGTCGT GCTGCCCTCTCGGGGGGTTTGCGCGAGCGTCGGCTCCGCCTGGGCCCTTGCGGTG CTCCTGGAGCGCTCCGGGTTGTCCCTCAGGTGCCCGAGGCCGAACGGTGGTGTGT CGTTCCCGCCCCCGGCGCCCCCTCCTCCGGTCGCCGCCGCGGTGTCCGCGCGTGG GTCCTGAGGGAGCTCGTCGGTGTGGGGTTCGAGGCGGTTTGAGTGAGACGAGAC GAGACGCGCCCCTCCCACGCGGGGAAGGGCGCCCGCCTGCTCTCGGTGAGCGCA CGTCCCGTGCTCCCCTCTGGCGGGTGCGCGCGGGCCGTGTGAGCGATCGCGGTGG GTTCGGGCCGGTGTGACGCGTGCGCCGGCCGGCCGCCGAGGGGCTGCCGTTCTG CCTCCGACCGGTCGTGTGTGGGTTGACTTCGGAGGCGCTCTGCCTCGGAAGGAAG GAGGTGGGTGGACGGGGGGGCCTGGTGGGGTTGCGCGCACGCGCGCACCGGCCG GGCCCCCGCCCTGAACGCGAACGCTCGAGGTGGCCGCGCGCAGGTGTTTCCTCGT ACCGCAGGGCCCCCTCCCTTCCCCAGGCGTCCCTCGGCGCCTCTGCGGGCCCGAG GAGGAGCGGCTGGCGGGTGGGGGGAGTGTGACCCACCCTCGGTGAGAAAAGCCT TCTCTAGCGATCTGAGAGGCGTGCCTTGGGGGTACCGGATCCCCCGGGCCGCCGC CTCTGTCTCTGCCTCCGTTATGGTAGCGCTGCCGTTAGCGACCCGCTCGCAGAGG ACCCTCCTCCGCTTCCCCCTCGACGGGGTTGGGGGGGAGAAGCGAGGGTTCCGC CGGCCACCGCGGTGGTGGCCGAGTGCGGCTCGTCGCCTACTGTGGCCCGCGCCTC CCCCCTTCCGAGTCGGGGGAGGATCCCGCCGGGCCGGGCCCGGCGTTCCCAGCG GGTTGGGACGCGGCGGCCGGCGGGCGGTGGGTGTGCGCGCCCGGCGCTCTGTCC GGCGCGTGACCCCCTCCGCCGCGAGTCGGCTCTCCGCCCGCTCCCGTGCCGAGTC GTGACCGGTGCCGACGACCGCGTTTGCGTGGCACGGGGTCGGGCCCGCCTGGCC CTGGGAAAGCGTCCCACGGTGGGGGCGCGCCGGTCTCCCGGAGCGGGACCGGGT CGGAGGATGGACGAGAATCACGAGCGACGGTGGTGCGGGCGTGTCGGGTTCGTG GCTGCGGTCGCTCCGGGGCCCCCGGTGGCGGGGCCCCGGGGCTCGCGAGGCGGT TCTCGGTGGGGGCCGAGGGCCGTCCGGCGTCCCAGGCGGGGCGCCGCGGGACCG CCCTCGTGTCTGTGGCGGTGGGATCCCGCGGCCGTGTTTTCCTGGTGGCCCGGCC GTGCCTGAGGTTTCTCCCCGAGCCGCCGCCTCTGCGGGCTCCCGGGTGCCCTTGC CCTCGCGGTCCCCGGCCCTCGCCCGTCTGTGCCCTCTTCCCCGCCCGCCGCCCGCC GATCCTCTTCTTCCCCCCGAGCGGCTCACCGGCTTCACGTCCGTTGGTGGCCCCG CCTGGGACCGAACCCGGCACCGCCTCGTGGGGCGCCGCCGCCGGCCGCTGATCG GCCCGGCGTCCGCGTCCCCCGGCGCGCGCCTTGGGGACCGGGTCGGTGGCGCCC CGCGTGGGGCCCGGTGGGCTTCCCGGAGGGTTCCGGGGGTCGGCCTGCGGCGCG TGCGGGGGAGGAGACGGTTCCGGGGGACCGGCCGCGACTGCGGCGGCGGTGGT GGGGGGAGCCGCGGGGATCGCCGAGGGCCGGTCGGCCGCCCCGGGTGCCGCGCAttorney Docket No.: 000218-0153-WO1 GGTGCCGCCGGCGGCGGTGAGGCCCCGCGCGTGTGTCCCGGCTGCGGTCGGCCG CGCTCGCGGGGTCCCCGTGGCGTCCCCTTCCCCGCCGGCCGCCTTTCTCGCGCCTT CCCCGTCGCCCCGGCCTCGCCCGTGGTCTCTCGTCTTCTCCCGGCCCGCTCTTCCG AACCGGGTCGGCGCGTCCCCCGGGTGCGCCTCGCTTCCCGGGCCTGCCGCGGCCC TTCCCCGAGGCGTCCGTCCCGGGCGTCGGCGTCGGGGAGAGCCCGTCCTCCCCGC GTGGCGTCGCCCCGTTCGGCGCGCGCGTGCGCCCGAGCGCGGCCCGGTGGTCCCT CCCGGACAGGCGTTCGTGCGACGTGTGGCGTGGGTCGACCTCCGCCTTGCCGGTC GCTCGCCCTCTCCCCGGGTCGGGGGGTGGGGCCCGGGCCGGGGCCTCGGCCCCG GTCGCGGTCCCCCGTCCCGGGCGGGGGCGGGCGCGCCGGCCGGCCTCGGTCGGC CCTCCCTTGGCCGTCGTGTGGCGTGTGCCACCCCTGCGCCCGCGCCCGCCGGCGG GGCTCGGAGCCGGGCTTCGGCCGGGCCCCGGGCCCTCGACCGGACCGGTGCGCG GGCGCTGCGGCCGCACGGCGCGACTGTCCCCGGGCCGGGCACCGAGGTCCGCCT CTCGCTCGCCGCCCGGACGTCGGGGCCGCCCCGCGGGGCGGGCGGAGCGCCGTC CCCGCCTCGCCGCCGCCCGCGGGCGCCGGCCGCGCGCGCGCGCGTGGCCGCCGG TCCCTCCCGGCCGCCGGGCGCGGGTCGGGCCGTCCGCCTCCTCGCGGGCGGGCG CGACGAAGAAGCGTCGCGGGTCTGTGGCGCGGGGCCCCGGTGGTCGTGTCGCGT GGGGGGCGGGTGGTTGGGGCGTCCGGTTCGCCGCGCCCCGCCCCGGCCCCACCG GTCCCGGCCGCCGCCCCCGCGCCCGCTCGCTCCCTCCCGTCCGCCCGTCCGCGGC CCGTCCGTCCGTCCGTCGTCCTCCTCGCTTGCGGGGCGCCGGGCCCGTCCTCGCG AGGCCCCCCGGCCGGCCGTCCGGCCGCGTCGGGGCCTCGCCGCGCTCTACCTTAC CTACCTGGTTGATCCTGCCAGTAGCATATGCTTGTCTCAAAGATTAAGCCATGCA TGTCTAAGTACGCACGGCCGGTACAGTGAAACTGCGAATGGCTCATTAAATCAG TTATGGTTCCTTTGGTCGCTCGCTCCTCTCCTACTTGGATAACTGTGGTAATTCTA GAGCTAATACATGCCGACGGGCGCTGACCCCCTTCGCGGGGGGGATGCGTGCAT TTATCAGATCAAAACCAACCCGGTCAGCCCCTCTCCGGCCCCGGCCGGGGGGCG GGCGCCGGCGGCTTTGGTGACTCTAGATAACCTCGGGCCGATCGCACGCCCCCCG TGGCGGCGACGACCCATTCGAACGTCTGCCCTATCAACTTTCGATGGTAGTCGCC GTGCCTACCATGGTGACCACGGGTGACGGGGAATCAGGGTTCGATTCCGGAGAG GGAGCCTGAGAAACGGCTACCACATCCAAGGAAGGCAGCAGGCGCGCAAATTA CCCACTCCCGACCCGGGGAGGTAGTGACGAAAAATAACAATACAGGACTCTTTC GAGGCCCTGTAATTGGAATGAGTCCACTTTAAATCCTTTAACGAGGATCCATTGG AGGGCAAGTCTGGTGCCAGCAGCCGCGGTAATTCCAGCTCCAATAGCGTATATT AAAGTTGCTGCAGTTAAAAAGCTCGTAGTTGGATCTTGGGAGCGGGCGGGCGGT CCGCCGCGAGGCGAGCCACCGCCCGTCCCCGCCCCTTGCCTCTCGGCGCCCCCTCAttorney Docket No.: 000218-0153-WO1 GATGCTCTTAGCTGAGTGTCCCGCGGGGCCCGAAGCGTTTACTTTGAAAAAATTA GAGTGTTCAAAGCAGGCCCGAGCCGCCTGGATACCGCAGCTAGGAATAATGGAA TAGGACCGCGGTTCTATTTTGTTGGTTTTCGGAACTGAGGCCATGATTAAGAGGG ACGGCCGGGGGCATTCGTATTGCGCCGCTAGAGGTGAAATTCTTGGACCGGCGC AAGACGGACCAGAGCGAAAGCATTTGCCAAGAATGTTTTCATTAATCAAGAACG AAAGTCGGAGGTTCGAAGACGATCAGATACCGTCGTAGTTCCGACCATAAACGA TGCCGACCGGCGATGCGGCGGCGTTATTCCCATGACCCGCCGGGCAGCTTCCGG GAAACCAAAGTCTTTGGGTTCCGGGGGGAGTATGGTTGCAAAGCTGAAACTTAA AGGAATTGACGGAAGGGCACCACCAGGAGTGGAGCCTGCGGCTTAATTTGACTC AACACGGGAAACCTCACCCGGCCCGGACACGGACAGGATTGACAGATTGATAGC TCTTTCTCGATTCCGTGGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATT TGTCTGGTTAATTCCGATAACGAACGAGACTCTGGCATGCTAACTAGTTACGCGA CCCCCGAGCGGTCGGCGTCCCCCAACTTCTTAGAGGGACAAGTGGCGTTCAGCC ACCCGAGATTGAGCAATAACAGGTCTGTGATGCCCTTAGATGTCCGGGGCTGCA CGCGCGCTACACTGACTGGCTCAGCGTGTGCCTACCCTACGCCGGCAGGCGCGG GTAACCCGTTGAACCCCATTCGTGATGGGGATCGGGGATTGCAATTATTCCCCAT GAACGAGGAATTCCCAGTAAGTGCGGGTCATAAGCTTGCGTTGATTAAGTCCCTG CCCTTTGTACACACCGCCCGTCGCTACTACCGATTGGATGGTTTAGTGAGGCCCT CGGATCGGCCCCGCCGGGGTCGGCCCACGGCCCTGGCGGAGCGCTGAGAAGACG GTCGAACTTGACTATCTAGAGGAAGTAAAAGTCGTAACAAGGTTTCCGTAGGTG AACCTGCGGAAGGATCATTAACGGAGCCCGGAGGGCGAGGCCCGCGGCGGCGC CGCCGCCGCCGCGCGCTTCCCTCCGCACACCCACCCCCCCACCGCGACGCGGCGC GTGCGCGGGCGGGGCCCGCGTGCCCGTTCGCTCGCTCGCTCGTTCGTTCGCCGCC CGGCCCCGCCGGCCGCGAGAGCCGGAGAACTCGGGAGGGAGACGGGGGAGAGA GAGAGAGAGAGAGAGAGAGAGAGAGAGAAAGAAGGGCGTGTCGTTGGTGTGCG CGTGTCGTGGGGCCGGCGGGCGGCGGGGAGCGGTCCCCGGCCGCGGCCCCGACG GCGTGGGTGTCGGCGGGCGCGGGGGCGGTTCTCGGCGGCGTCGCGGCGGGTCTG GGGGGTCTCGGTGCCCTCCTCCCCGCCGGGGCCCGTCGTCCGGCCCCGCCGCGCC GGCTCCCCGTCTTCGGGGCCGGCCGGATTCCCGTCGCCTCCGCCGCGCCGCTCCG CGCCGCCGGGCACGGCCCCGCTCGCTCTCCCCGGCCTTCCCGCTAGGGCGTCTCG AGGGTCGGGGGCCGGACGCCGGTCCCCTCCCCCGCCTCCTCGTCCGCCCCCCCGC CGTCCAGGTACCTAGCGCGTTCCGGCGCGGAGGTTTAAAGACCCCTTGGGGGGA TCGCCCGTCCGCCCGTGGGTCGGGGGCGGTGGTGGGCCCGCGGGGGAGTCCCGT CGGGAGGGGCCCGGCCCCTCCCGCGCCTCCACCGCGGACTCCGCTCCCCGGCCGAttorney Docket No.: 000218-0153-WO1 GGGCCGCGCCGCCGCCGCCGCCGCGGCGGCCGTCGGGTGGGGGCTTTACCCGGC GGCCGTCGCGCGCCTGCCGCGCGTGTGGCGTGCGCCCCGCGCCGTGGGGGCGGG AACCCCCGGGCGCCTGTGGGGTGGTGTCCGCGCTCGCCCCCGCGTGGGCGGCGC GCGCCTCCCCGTGGTGTGAAACCTTCCGACCCCTCTCCGGAGTCCGGTCCCGTTT GCTGTCTCGTCTGGCCGGCCTGAGGCAACCCCCTCTCCTCTTGGGCGGGGTTGGG GGACGTGCCGCGCCAGGAAGGGCCTCCTCCCGGTGCGTCGTCGGGAGCGCCCTC GCCAAATCGACCTCGTACGACTCTTAGCGGTGGATCACTCGGCTCGTGCGTCGAT GAAGAACGCAGCTAGCTGCGAGAATTAATGTGAATTGCAGGACACATTGATCAT CGACACTTCGAACGCACTTGCGGCCCCGGGTTCCTCCCGGGGCTACGCCTGTCTG AGCGTCGCTTGCCGATCAATCGCCCCCGGGGGTGCCTCCGGGCTCCTCGGGGTGC GCGGCTGGGGGTTCCCTCGCAGGGCCCGCCGGGGGCCCTCCGTCCCCCTAAGCG CAGACCCGGCGGCGTCCGCCCTCCTCTTGCCGCCGCGCCCGCCCCTTCCCCCTCC CCCCGCGGGCCCTGCGTGGTCACGCGTCGGGTGGCGGGGGGGAGAGGGGGGCGC GCCCGGCTGAGAGAGACGGGGAGGGCGGCGCCGCCGCCGCCCGCGAAGACGGA GAGGGAAAGAGAGAGCCGGCTCGGGCCGAGTTCCCGTGGCCGCCGCCTGCGGTC CGGGTTCCTCCCTCGGGGGGCTCCCTCGCGCCGCGCGCGGCTCGGGGTTCGGGGT TCGTCGGCCCCGGCCGGGTGGAAGGTCCCGTGCCCGTCGTCGTCGTCGTCGCGCG TCGTCGGCGGTGGGGGCGTGTTGCGTGCGGTGTGGTGGTGGGGGAGGAGGAAGG CGGGTCCGGAAGGGGCAGGGTGCCGGCGGGGAGAGAGGGTCGGGGGAGCGCGT CCCGGTCGCCGCGGTTCGCCGCCCGCCCCCGGTGGCGGCCCGGCGTCCGGCCGA CCGCCGCTCCCGCGCCCCTCCTCCTCCCCGCCGCCCCTCCTCCGAGGCCCCGCCC GTCCTCCTCGCCCTCCCCGCGCGTACGCGCGCGCGCCCGCCCGCCCGGCTCGCCT CGCGGCGCGTCGGCCGGGGCCGGGAGCCCGCCCCGCGGCCCGCCCGGCCGCGCC CGTGGCCGCGGCGCCGGGGTTCGCGTGTCCCCGGCGGCGACCCGCGGGACGCCG CGGTGTCGTCCGCCGTCGCGCGCCCGCCTCCGGCTCGCGGCCGCGCCGCGCCGCG CCGGGGCCCCGTCCCGAGCTTCCGCGTCGGGGCGGGGCGGCTCCGCCGCCGCGT CCTCGGACCCGTCCCCCCGACCTCCGCGGGGGAGACGGGTCGGGGCGTGCGGCG CCCGTCCCGCCCCCGGCCCGTGCCCCTCCCTCCGGTCGTCCCGCTCCGGCGGGGC GGCGCGGGGGCGCCGTCGGCCGCGCGCTCTCTCTCCCGTCGCCTCTCCCCCTCGC CGGGCCCGTCTCCCGACGGAGCGTCGGGCGGGCGGTCGGGCCGGCGCGATTCCG TCCGTCCGTCCGCCGAGCGGCCCGTCCCCCTCCGAGACGCGACCTCAGATCAGAC GTGGCGACCCGCTGAATTTAAGCATATTAGTCAGCGGAGGAAAAGAAACTAACC AGGATTCCCTCAGTAACGGCGAGTGAACAGGGAAGAGCCCAGCGCCGAATCCCC GCCCCGCGGCGGGGCGCGGGACATGTGGCGTACGGAAGACCCGCTCCCCGGCGCAttorney Docket No.: 000218-0153-WO1 CGCTCGTGGGGGGCCCAAGTCCTTCTGATCGAGGCCCAGCCCGTGGACGGTGTG AGGCCGGTAGCGGCCCCCGGCGCGCCGGGCCCGGGTCTTCCCGGAGTCGGGTTG CTTGGGAATGCAGCCCAAAGCGGGTGGTAAACTCCATCTAAGGCTAAATACCGG CACGAGACCGATAGTCAACAAGTACCGTAAGGGAAAGTTGAAAAGAACTTTGAA GAGAGAGTTCAAGAGGGCGTGAAACCGTTAAGAGGTAAACGGGTGGGGTCCGC GCAGTCCGCCCGGAGGATTCAACCCGGCGGCGGGTCCGGCCGTGTCGGCGGCCC GGCGGATCTTTCCCGCCCCCCGTTCCTCCCGACCCCTCCACCCGCCCTCCCTTCCC CCGCCGCCCCTCCTCCTCCTCCCCGGAGGGGGCGGGCTCCGGCGGGTGCGGGGG TGGGCGGGCGGGGCCGGGGGTGGGGTCGGCGGGGGACCGTCCCCCGACCGGCG ACCGGCCGCCGCCGGGCGCATTTCCACCGCGGCGGTGCGCCGCGACCGGCTCCG GGACGGCTGGGAAGGCCCGGCGGGGAAGGTGGCTCGGGGGGCCCCGTCCGTCCG TCCGTCCGTCCTCCTCCTCCCCCGTCTCCGCCCCCCGGCCCCGCGTCCTCCCTCGG GAGGGCGCGCGGGTCGGGGCGGCGGCGGCGGCGGCGGTGGCGGCGGCGGCGGC GGCGGCGGGACCGAAACCCCCCCCGAGTGTTACAGCCCCCCCGGCAGCAGCACT CGCCGAATCCCGGGGCCGAGGGAGCGAGACCCGTCGCCGCGCTCTCCCCCCTCC CGGCGCCCACCCCCGCGGGGAATCCCCCGCGAGGGGGGTCTCCCCCGCGGGGGC GCGCCGGCGTCTCCTCGTGGGGGGGCCGGGCCACCCCTCCCACGGCGCGACCGC TCTCCCACCCCTCCTCCCCGCGCCCCCGCCCCGGCGACGGGGGGGGTGCCGCGCG CGGGTCGGGGGGCGGGGCGGACTGTCCCCAGTGCGCCCCGGGCGGGTCGCGCCG TCGGGCCCGGGGGAGGTTCTCTCGGGGCCACGCGCGCGTCCCCCGAAGAGGGGG ACGGCGGAGCGAGCGCACGGGGTCGGCGGCGACGTCGGCTACCCACCCGACCCG TCTTGAAACACGGACCAAGGAGTCTAACACGTGCGCGAGTCGGGGGCTCGCACG AAAGCCGCCGTGGCGCAATGAAGGTGAAGGCCGGCGCGCTCGCCGGCCGAGGTG GGATCCCGAGGCCTCTCCAGTCCGCCGAGGGCGCACCACCGGCCCGTCTCGCCC GCCGCGCCGGGGAGGTGGAGCACGAGCGCACGTGTTAGGACCCGAAAGATGGT GAACTATGCCTGGGCAGGGCGAAGCCAGAGGAAACTCTGGTGGAGGTCCGTAGC GGTCCTGACGTGCAAATCGGTCGTCCGACCTGGGTATAGGGGCGAAAGACTAAT CGAACCATCTAGTAGCTGGTTCCCTCCGAAGTTTCCCTCAGGATAGCTGGCGCTC TCGCAGACCCGACGCACCCCCGCCACGCAGTTTTATCCGGTAAAGCGAATGATTA GAGGTCTTGGGGCCGAAACGATCTCAACCTATTCTCAAACTTTAAATGGGTAAGA AGCCCGGCTCGCTGGCGTGGAGCCGGGCGTGGAATGCGAGTGCCTAGTGGGCCA CTTTTGGTAAGCAGAACTGGCGCTGCGGGATGAACCGAACGCCGGGTTAAGGCG CCCGATGCCGACGCTCATCAGACCCCAGAAAAGGTGTTGGTTGATATAGACAGC AGGACGGTGGCCATGGAAGTCGGAATCCGCTAAGGAGTGTGTAACAACTCACCTAttorney Docket No.: 000218-0153-WO1 GCCGAATCAACTAGCCCTGAAAATGGATGGCGCTGGAGCGTCGGGCCCATACCC GGCCGTCGCCGGCAGTCGAGAGTGGACGGGAGCGGCGGGGGCGGCGCGCGCGC GCGCGCGTGTGGTGTGCGTCGGAGGGCGGCGGCGGCGGCGGCGGCGGCGGGGG TGTGGGGTCCTTCCCCCGCCCCCCCCCCCCACGCCTCCTCCCCTCCTCCCGCCCAC GCCCCGCTCCCCGCCCCCGGAGCCCCGCGGACGCTACGCCGCGACGAGTAGGAG GGCCGCTGCGGTGAGCCTTGAAGCCTAGGGCGCGGGCCCGGGTGGAGCCGCCGC AGGTGCAGATCTTGGTGGTAGTAGCAAATATTCAAACGAGAACTTTGAAGGCCG AAGTGGAGAAGGGTTCCATGTGAACAGCAGTTGAACATGGGTCAGTCGGTCCTG AGAGATGGGCGAGCGCCGTTCCGAAGGGACGGGCGATGGCCTCCGTTGCCCTCG GCCGATCGAAAGGGAGTCGGGTTCAGATCCCCGAATCCGGAGTGGCGGAGATGG GCGCCGCGAGGCGTCCAGTGCGGTAACGCGACCGATCCCGGAGAAGCCGGCGGG AGCCCCGGGGAGAGTTCTCTTTTCTTTGTGAAGGGCAGGGCGCCCTGGAATGGGT TCGCCCCGAGAGAGGGGCCCGTGCCTTGGAAAGCGTCGCGGTTCCGGCGGCGTC CGGTGAGCTCTCGCTGGCCCTTGAAAATCCGGGGGAGAGGGTGTAAATCTCGCG CCGGGCCGTACCCATATCCGCAGCAGGTCTCCAAGGTGAACAGCCTCTGGCATGT TGGAACAATGTAGGTAAGGGAAGTCGGCAAGCCGGATCCGTAACTTCGGGATAA GGATTGGCTCTAAGGGCTGGGTCGGTCGGGCTGGGGCGCGAAGCGGGGCTGGGC GCGCGCCGCGGCTGGACGAGGCGCCGCCGCCCCCCCCACGCCCGGGGCACCCCC CTCGCGGCCCTCCCCCGCCCCACCCCGCGCGCGCCGCTCGCTCCCTCCCCGCCCC GCGCCCTCTCTCTCTCTCTCTCTCCCCCGCTCCCCGTCCTCCCCCCTCCCCGGGGG AGCGCCGCGTGGGGGCGGCGGCGGGGGGAGAGAAGGGTCGGGGCGGCAGGGGC CGGCGGCGGCCCGCCGCGGGGCCCCGGCGGCGGGGGCACGGTCCCCCGCGAGG GGGGCCCGGGCACCCGGGGGGCCGGCGGCGGCGGCGACTCTGGACGCGAGCCG GGCCCTTCCCGTGGATCGCCCCAGCTGCGGCGGGCGTCGCGGCCGCCCCCGGGG AGCCCGGCGGGCGCCGGCGCGCCCCCCCCACCCCCACCCCACGTCTCGTCGCGC GCGCGTCCGCTGGGGGCGGGGAGCGGTCGGGCGGCGGCGGTCGGCGGGCGGCG GGGCGGGGCGGTTCGTCCCCCCGCCCTACCCCCCCGGCCCCGTCCGCCCCCCGTT CCCCCCTCCTCCTCGGCGCGCGGCGGCGGCGGCGGCGGCGGCAGGCGGCGGAGG GGCCGCGGGCCGGTCCCCCCCGCCGGGTCCGCCCCCGGGGCCGCGGTTCCGCGC GGCGCCTCGCCTCGGCCGGCGCCTAGCAGCCGACTTAGAACTGGTGCGGACCAG GGGAATCCGACTGTTTAATTAAAACAAAGCATCGCGAAGGCCCGCGGCGGGTGT TGACGCGATGTGATTTCTGCCCAGTGCTCTGAATGTCAAAGTGAAGAAATTCAAT GAAGCGCGGGTAAACGGCGGGAGTAACTATGACTCTCTTAAGGTAGCCAAATGC CTCGTCATCTAATTAGTGACGCGCATGAATGGATGAACGAGATTCCCACTGTCCCAttorney Docket No.: 000218-0153-WO1 TACCTACTATCCAGCGAAACCACAGCCAAGGGAACGGGCTTGGCGGAATCAGCG GGGAAAGAAGACCCTGTTGAGCTTGACTCTAGTCTGGCACGGTGAAGAGACATG AGAGGTGTAGAATAAGTGGGAGGCCCCCGGCGCCCCCCCGGTGTCCCCGCGAGG GGCCCGGGGCGGGGTCCGCCGGCCCTGCGGGCCGCCGGTGAAATACCACTACTC TGATCGTTTTTTCACTGACCCGGTGAGGCGGGGGGGCGAGCCCCGAGGGGCTCTC GCTTCTGGCGCCAAGCGCCCGGCCGCGCGCCGGCCGGGCGCGACCCGCTCCGGG GACAGTGCCAGGTGGGGAGTTTGACTGGGGCGGTACACCTGTCAAACGGTAACG CAGGTGTCCTAAGGCGAGCTCAGGGAGGACAGAAACCTCCCGTGGAGCAGAAG GGCAAAAGCTCGCTTGATCTTGATTTTCAGTACGAATACAGACCGTGAAAGCGG GGCCTCACGATCCTTCTGACCTTTTGGGTTTTAAGCAGGAGGTGTCAGAAAAGTT ACCACAGGGATAACTGGCTTGTGGCGGCCAAGCGTTCATAGCGACGTCGCTTTTT GATCCTTCGATGTCGGCTCTTCCTATCATTGTGAAGCAGAATTCACCAAGCGTTG GATTGTTCACCCACTAATAGGGAACGTGAGCTGGGTTTAGACCGTCGTGAGACA GGTTAGTTTTACCCTACTGATGATGTGTTGTTGCCATGGTAATCCTGCTCAGTACG AGAGGAACCGCAGGTTCAGACATTTGGTGTATGTGCTTGGCTGAGGAGCCAATG GGGCGAAGCTACCATCTGTGGGATTATGACTGAACGCCTCTAAGTCAGAATCCC GCCCAGGCGGAACGATACGGCAGCGCCGCGGAGCCTCGGTTGGCCTCGGATAGC CGGTCCCCCGCCTGTCCCCGCCGGCGGGCCGCCCCCCCCTCCACGCGCCCCGCGC GCGCGGGAGGGCGCGTGCCCCGCCGCGCGCCGGGACCGGGGTCCGGTGCGGAGT GCCCTTCGTCCTGGGAAACGGGGCGCGGCCGGAAAGGCGGCCGCCCCCTCGCCC GTCACGCACCGCACGTTCGTGGGGAACCTGGCGCTAAACCATTCGTAGACGACC TGCTTCTGGGTCGGGGTTTCGTACGTAGCAGAGCAGCTCCCTCGCTGCGATCTAT TGAAAGTCAGCCCTCGACACAAGGGTTTGTCCGCGCGCGCGCGCGCGCGTGCGT GCGGGGGGCCCGGCGGGGCGTGCGCGTCCGGCGCCGTCCGTCCTTCCGTTCGTCT TCCTCCCTCCCGGCCTCTCCCGCCGACCGCGGGCGTGGTGGTGGGGGTGGGGGG GGGGAGGGCGCGCGACCCCGGTCGGCGCGCCCCGCTTCTTCGGTTCCCGCCTCCT CCCCGTTCACCGCCGGGGCGGCTCGTCCGCTCCGGGCCGGGACGGGGTCCGGGG AGCGTGGTTTGGGAGCCGCGGAGGCGGCCGCGCCGAGCCGGGCCCGTGGCCCGC CGGTCCCCGTCCCGGGGGTTGGCCGCGCGGGCCCCGGTGGGGCGGCCACCCGGG GTCCCGGCCCTCGCG (SEQ ID NO: 32).

[0064] In certain aspects, the rDNA TAL Array of the fusion protein comprises one or more TAL domain targeting a nucleic acid sequence flanking an TTAA integration site located in the rDNA 18S gene. An exemplary rDNA 18S repeat nucleic acid sequence is set forth in SEQ ID NO: 33:Attorney Docket No.: 000218-0153-WO1 TACCTGGTTGATCCTGCCAGTAGCATATGCTTGTCTCAAAGATTAAGCCATGCAT GTCTAAGTACGCACGGCCGGTACAGTGAAACTGCGAATGGCTCATTAAATCAGT TATGGTTCCTTTGGTCGCTCGCTCCTCTCCTACTTGGATAACTGTGGTAATTCTAG AGCTAATACATGCCGACGGGCGCTGACCCCCTTCGCGGGGGGGATGCGTGCATT TATCAGATCAAAACCAACCCGGTCAGCCCCTCTCCGGCCCCGGCCGGGGGGCGG GCGCCGGCGGCTTTGGTGACTCTAGATAACCTCGGGCCGATCGCACGCCCCCCGT GGCGGCGACGACCCATTCGAACGTCTGCCCTATCAACTTTCGATGGTAGTCGCCG TGCCTACCATGGTGACCACGGGTGACGGGGAATCAGGGTTCGATTCCGGAGAGG GAGCCTGAGAAACGGCTACCACATCCAAGGAAGGCAGCAGGCGCGCAAATTACC CACTCCCGACCCGGGGAGGTAGTGACGAAAAATAACAATACAGGACTCTTTCGA GGCCCTGTAATTGGAATGAGTCCACTTTAAATCCTTTAACGAGGATCCATTGGAG GGCAAGTCTGGTGCCAGCAGCCGCGGTAATTCCAGCTCCAATAGCGTATATTAA AGTTGCTGCAGTTAAAAAGCTCGTAGTTGGATCTTGGGAGCGGGCGGGCGGTCC GCCGCGAGGCGAGCCACCGCCCGTCCCCGCCCCTTGCCTCTCGGCGCCCCCTCGA TGCTCTTAGCTGAGTGTCCCGCGGGGCCCGAAGCGTTTACTTTGAAAAAATTAGA GTGTTCAAAGCAGGCCCGAGCCGCCTGGATACCGCAGCTAGGAATAATGGAATA GGACCGCGGTTCTATTTTGTTGGTTTTCGGAACTGAGGCCATGATTAAGAGGGAC GGCCGGGGGCATTCGTATTGCGCCGCTAGAGGTGAAATTCTTGGACCGGCGCAA GACGGACCAGAGCGAAAGCATTTGCCAAGAATGTTTTCATTAATCAAGAACGAA AGTCGGAGGTTCGAAGACGATCAGATACCGTCGTAGTTCCGACCATAAACGATG CCGACCGGCGATGCGGCGGCGTTATTCCCATGACCCGCCGGGCAGCTTCCGGGA AACCAAAGTCTTTGGGTTCCGGGGGGAGTATGGTTGCAAAGCTGAAACTTAAAG GAATTGACGGAAGGGCACCACCAGGAGTGGAGCCTGCGGCTTAATTTGACTCAA CACGGGAAACCTCACCCGGCCCGGACACGGACAGGATTGACAGATTGATAGCTC TTTCTCGATTCCGTGGGTGGTGGTGCATGGCCGTTCTTAGTTGGTGGAGCGATTTG TCTGGTTAATTCCGATAACGAACGAGACTCTGGCATGCTAACTAGTTACGCGACC CCCGAGCGGTCGGCGTCCCCCAACTTCTTAGAGGGACAAGTGGCGTTCAGCCAC CCGAGATTGAGCAATAACAGGTCTGTGATGCCCTTAGATGTCCGGGGCTGCACG CGCGCTACACTGACTGGCTCAGCGTGTGCCTACCCTACGCCGGCAGGCGCGGGT AACCCGTTGAACCCCATTCGTGATGGGGATCGGGGATTGCAATTATTCCCCATGA ACGAGGAATTCCCAGTAAGTGCGGGTCATAAGCTTGCGTTGATTAAGTCCCTGCC CTTTGTACACACCGCCCGTCGCTACTACCGATTGGATGGTTTAGTGAGGCCCTCG GATCGGCCCCGCCGGGGTCGGCCCACGGCCCTGGCGGAGCGCTGAGAAGACGGTAttorney Docket No.: 000218-0153-WO1 CGAACTTGACTATCTAGAGGAAGTAAAAGTCGTAACAAGGTTTCCGTAGGTGAA CCTGCGGAAGGATCATTA (SEQ ID NO: 33).

[0065] In certain aspects, the rDNA TAL Array of the fusion protein comprises one or more TAL domain targeting a nucleic acid sequence flanking an TTAA integration site located in the rDNA 5.8S gene. An exemplary rDNA 5.8S repeat nucleic acid sequence is set forth in SEQ ID NO: 34: GACTCTTAGCGGTGGATCACTCGGCTCGTGCGTCGATGAAGAACGCAGCTAGCT GCGAGAATTAATGTGAATTGCAGGACACATTGATCATCGACACTTCGAACGCAC TTGCGGCCCCGGGTTCCTCCCGGGGCTACGCCTGTCTGAGCGTCG (SEQ ID NO: 34).

[0066] In certain aspects, the rDNA TAL Array of the fusion protein comprises one or more TAL domains targeting a nucleic acid sequence flanking an TTAA integration site located in the rDNA 28S gene. An exemplary rDNA 28S repeat nucleic acid sequence is set forth in SEQ ID NO: 35: CGACCTCAGATCAGACGTGGCGACCCGCTGAATTTAAGCATATTAGTCAGCGGA GGAAAAGAAACTAACCAGGATTCCCTCAGTAACGGCGAGTGAACAGGGAAGAG CCCAGCGCCGAATCCCCGCCCCGCGGCGGGGCGCGGGACATGTGGCGTACGGAA GACCCGCTCCCCGGCGCCGCTCGTGGGGGGCCCAAGTCCTTCTGATCGAGGCCCA GCCCGTGGACGGTGTGAGGCCGGTAGCGGCCCCCGGCGCGCCGGGCCCGGGTCT TCCCGGAGTCGGGTTGCTTGGGAATGCAGCCCAAAGCGGGTGGTAAACTCCATC TAAGGCTAAATACCGGCACGAGACCGATAGTCAACAAGTACCGTAAGGGAAAGT TGAAAAGAACTTTGAAGAGAGAGTTCAAGAGGGCGTGAAACCGTTAAGAGGTAA ACGGGTGGGGTCCGCGCAGTCCGCCCGGAGGATTCAACCCGGCGGCGGGTCCGG CCGTGTCGGCGGCCCGGCGGATCTTTCCCGCCCCCCGTTCCTCCCGACCCCTCCA CCCGCCCTCCCTTCCCCCGCCGCCCCTCCTCCTCCTCCCCGGAGGGGGCGGGCTC CGGCGGGTGCGGGGGTGGGCGGGCGGGGCCGGGGGTGGGGTCGGCGGGGGACC GTCCCCCGACCGGCGACCGGCCGCCGCCGGGCGCATTTCCACCGCGGCGGTGCG CCGCGACCGGCTCCGGGACGGCTGGGAAGGCCCGGCGGGGAAGGTGGCTCGGG GGGCCCCGTCCGTCCGTCCGTCCGTCCTCCTCCTCCCCCGTCTCCGCCCCCCGGCC CCGCGTCCTCCCTCGGGAGGGCGCGCGGGTCGGGGCGGCGGCGGCGGCGGCGGT GGCGGCGGCGGCGGCGGCGGCGGGACCGAAACCCCCCCCGAGTGTTACAGCCCC CCCGGCAGCAGCACTCGCCGAATCCCGGGGCCGAGGGAGCGAGACCCGTCGCCG CGCTCTCCCCCCTCCCGGCGCCCACCCCCGCGGGGAATCCCCCGCGAGGGGGGTC TCCCCCGCGGGGGCGCGCCGGCGTCTCCTCGTGGGGGGGCCGGGCCACCCCTCCAttorney Docket No.: 000218-0153-WO1 CACGGCGCGACCGCTCTCCCACCCCTCCTCCCCGCGCCCCCGCCCCGGCGACGGG GGGGGTGCCGCGCGCGGGTCGGGGGGCGGGGCGGACTGTCCCCAGTGCGCCCCG GGCGGGTCGCGCCGTCGGGCCCGGGGGAGGTTCTCTCGGGGCCACGCGCGCGTC CCCCGAAGAGGGGGACGGCGGAGCGAGCGCACGGGGTCGGCGGCGACGTCGGC TACCCACCCGACCCGTCTTGAAACACGGACCAAGGAGTCTAACACGTGCGCGAG TCGGGGGCTCGCACGAAAGCCGCCGTGGCGCAATGAAGGTGAAGGCCGGCGCGC TCGCCGGCCGAGGTGGGATCCCGAGGCCTCTCCAGTCCGCCGAGGGCGCACCAC CGGCCCGTCTCGCCCGCCGCGCCGGGGAGGTGGAGCACGAGCGCACGTGTTAGG ACCCGAAAGATGGTGAACTATGCCTGGGCAGGGCGAAGCCAGAGGAAACTCTGG TGGAGGTCCGTAGCGGTCCTGACGTGCAAATCGGTCGTCCGACCTGGGTATAGG GGCGAAAGACTAATCGAACCATCTAGTAGCTGGTTCCCTCCGAAGTTTCCCTCAG GATAGCTGGCGCTCTCGCAGACCCGACGCACCCCCGCCACGCAGTTTTATCCGGT AAAGCGAATGATTAGAGGTCTTGGGGCCGAAACGATCTCAACCTATTCTCAAAC TTTAAATGGGTAAGAAGCCCGGCTCGCTGGCGTGGAGCCGGGCGTGGAATGCGA GTGCCTAGTGGGCCACTTTTGGTAAGCAGAACTGGCGCTGCGGGATGAACCGAA CGCCGGGTTAAGGCGCCCGATGCCGACGCTCATCAGACCCCAGAAAAGGTGTTG GTTGATATAGACAGCAGGACGGTGGCCATGGAAGTCGGAATCCGCTAAGGAGTG TGTAACAACTCACCTGCCGAATCAACTAGCCCTGAAAATGGATGGCGCTGGAGC GTCGGGCCCATACCCGGCCGTCGCCGGCAGTCGAGAGTGGACGGGAGCGGCGGG GGCGGCGCGCGCGCGCGCGCGTGTGGTGTGCGTCGGAGGGCGGCGGCGGCGGCG GCGGCGGCGGGGGTGTGGGGTCCTTCCCCCGCCCCCCCCCCCCACGCCTCCTCCC CTCCTCCCGCCCACGCCCCGCTCCCCGCCCCCGGAGCCCCGCGGACGCTACGCCG CGACGAGTAGGAGGGCCGCTGCGGTGAGCCTTGAAGCCTAGGGCGCGGGCCCGG GTGGAGCCGCCGCAGGTGCAGATCTTGGTGGTAGTAGCAAATATTCAAACGAGA ACTTTGAAGGCCGAAGTGGAGAAGGGTTCCATGTGAACAGCAGTTGAACATGGG TCAGTCGGTCCTGAGAGATGGGCGAGCGCCGTTCCGAAGGGACGGGCGATGGCC TCCGTTGCCCTCGGCCGATCGAAAGGGAGTCGGGTTCAGATCCCCGAATCCGGA GTGGCGGAGATGGGCGCCGCGAGGCGTCCAGTGCGGTAACGCGACCGATCCCGG AGAAGCCGGCGGGAGCCCCGGGGAGAGTTCTCTTTTCTTTGTGAAGGGCAGGGC GCCCTGGAATGGGTTCGCCCCGAGAGAGGGGCCCGTGCCTTGGAAAGCGTCGCG GTTCCGGCGGCGTCCGGTGAGCTCTCGCTGGCCCTTGAAAATCCGGGGGAGAGG GTGTAAATCTCGCGCCGGGCCGTACCCATATCCGCAGCAGGTCTCCAAGGTGAA CAGCCTCTGGCATGTTGGAACAATGTAGGTAAGGGAAGTCGGCAAGCCGGATCC GTAACTTCGGGATAAGGATTGGCTCTAAGGGCTGGGTCGGTCGGGCTGGGGCGCAttorney Docket No.: 000218-0153-WO1 GAAGCGGGGCTGGGCGCGCGCCGCGGCTGGACGAGGCGCCGCCGCCCCCCCCAC GCCCGGGGCACCCCCCTCGCGGCCCTCCCCCGCCCCACCCCGCGCGCGCCGCTCG CTCCCTCCCCGCCCCGCGCCCTCTCTCTCTCTCTCTCTCCCCCGCTCCCCGTCCTCC CCCCTCCCCGGGGGAGCGCCGCGTGGGGGCGGCGGCGGGGGGAGAGAAGGGTC GGGGCGGCAGGGGCCGGCGGCGGCCCGCCGCGGGGCCCCGGCGGCGGGGGCAC GGTCCCCCGCGAGGGGGGCCCGGGCACCCGGGGGGCCGGCGGCGGCGGCGACT CTGGACGCGAGCCGGGCCCTTCCCGTGGATCGCCCCAGCTGCGGCGGGCGTCGC GGCCGCCCCCGGGGAGCCCGGCGGGCGCCGGCGCGCCCCCCCCACCCCCACCCC ACGTCTCGTCGCGCGCGCGTCCGCTGGGGGCGGGGAGCGGTCGGGCGGCGGCGG TCGGCGGGCGGCGGGGCGGGGCGGTTCGTCCCCCCGCCCTACCCCCCCGGCCCC GTCCGCCCCCCGTTCCCCCCTCCTCCTCGGCGCGCGGCGGCGGCGGCGGCGGCGG CAGGCGGCGGAGGGGCCGCGGGCCGGTCCCCCCCGCCGGGTCCGCCCCCGGGGC CGCGGTTCCGCGCGGCGCCTCGCCTCGGCCGGCGCCTAGCAGCCGACTTAGAACT GGTGCGGACCAGGGGAATCCGACTGTTTAATTAAAACAAAGCATCGCGAAGGCC CGCGGCGGGTGTTGACGCGATGTGATTTCTGCCCAGTGCTCTGAATGTCAAAGTG AAGAAATTCAATGAAGCGCGGGTAAACGGCGGGAGTAACTATGACTCTCTTAAG GTAGCCAAATGCCTCGTCATCTAATTAGTGACGCGCATGAATGGATGAACGAGA TTCCCACTGTCCCTACCTACTATCCAGCGAAACCACAGCCAAGGGAACGGGCTTG GCGGAATCAGCGGGGAAAGAAGACCCTGTTGAGCTTGACTCTAGTCTGGCACGG TGAAGAGACATGAGAGGTGTAGAATAAGTGGGAGGCCCCCGGCGCCCCCCCGGT GTCCCCGCGAGGGGCCCGGGGCGGGGTCCGCCGGCCCTGCGGGCCGCCGGTGAA ATACCACTACTCTGATCGTTTTTTCACTGACCCGGTGAGGCGGGGGGGCGAGCCC CGAGGGGCTCTCGCTTCTGGCGCCAAGCGCCCGGCCGCGCGCCGGCCGGGCGCG ACCCGCTCCGGGGACAGTGCCAGGTGGGGAGTTTGACTGGGGCGGTACACCTGT CAAACGGTAACGCAGGTGTCCTAAGGCGAGCTCAGGGAGGACAGAAACCTCCCG TGGAGCAGAAGGGCAAAAGCTCGCTTGATCTTGATTTTCAGTACGAATACAGAC CGTGAAAGCGGGGCCTCACGATCCTTCTGACCTTTTGGGTTTTAAGCAGGAGGTG TCAGAAAAGTTACCACAGGGATAACTGGCTTGTGGCGGCCAAGCGTTCATAGCG ACGTCGCTTTTTGATCCTTCGATGTCGGCTCTTCCTATCATTGTGAAGCAGAATTC ACCAAGCGTTGGATTGTTCACCCACTAATAGGGAACGTGAGCTGGGTTTAGACC GTCGTGAGACAGGTTAGTTTTACCCTACTGATGATGTGTTGTTGCCATGGTAATC CTGCTCAGTACGAGAGGAACCGCAGGTTCAGACATTTGGTGTATGTGCTTGGCTG AGGAGCCAATGGGGCGAAGCTACCATCTGTGGGATTATGACTGAACGCCTCTAA GTCAGAATCCCGCCCAGGCGGAACGATACGGCAGCGCCGCGGAGCCTCGGTTGGAttorney Docket No.: 000218-0153-WO1 CCTCGGATAGCCGGTCCCCCGCCTGTCCCCGCCGGCGGGCCGCCCCCCCCTCCAC GCGCCCCGCGCGCGCGGGAGGGCGCGTGCCCCGCCGCGCGCCGGGACCGGGGTC CGGTGCGGAGTGCCCTTCGTCCTGGGAAACGGGGCGCGGCCGGAAAGGCGGCCG CCCCCTCGCCCGTCACGCACCGCACGTTCGTGGGGAACCTGGCGCTAAACCATTC GTAGACGACCTGCTTCTGGGTCGGGGTTTCGTACGTAGCAGAGCAGCTCCCTCGC TGCGATCTATTGAAAGTCAGCCCTCGACACAAGGGTTTGT (SEQ ID NO: 35).

[0067] The nucleic acid sequence of the identified 23 TTAA integration sites and 50 bp of flanking upstream and downstream DNA sequences are shown in Table 2. Table 2: rDNA TTAA Integration Site SequencesAttorney Docket No.: 000218-0153-WO1

[0068] Based on the flanking nucleic acid sequences, TAL Arrays can be designed against the rDNA upstream and downstream flanking sequences for TTAA Sites 3, 4, 5, 6, 9, 11, 15, 17, 19, 22 & 23 as shown in Table 3. Table 3: rDNA Upstream (L) and Downstream (R) Target SequencesAttorney Docket No.: 000218-0153-WO1

[0069] Based on the flanking sequences shown in SEQ ID NOs: 36 – 58, additional TAL Arrays targeting alternative sequences may be readily designed using the methods disclosed herein and known to those skilled in the art. TAL Arrays Targeting rDNA TTAA Integration Site Flanking Sequences

[0070] In some embodiments, a TAL array targets an rDNA repeat element. Specific nucleic acid sequences upstream of (“left”) or downstream of (“right”) the selected TTAA integrations were chosen to design TAL Arrays for TTAA Sites 3, 4, 5, 6, 9, 11, 15, 17, 19, 22 & 23. The amino acid sequences for the left and right TAL Arrays targeting these TTAA sites are shown as SEQ ID NOs: 81-102 in Table 4. Table 4: rDNA TAL Array SequencesAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1

[0071] As shown in Table 4, illustrative sequences of left TAL Arrays targeting a rDNA repeat element target site TTAA are set forth in SEQ ID NOs: 81, 83, 85, 87, 89, 91, 93, 95, 97, 99 and 101. Illustrative sequences of right TAL Arrays targeting a rDNA repeat element target site TTAA are set forth in SEQ ID NOs: 82, 84, 86, 88, 90, 92, 94, 96, 98, 100 and 102. CpG-Modified rDNA Targeting TAL Arrays

[0072] In certain aspects, provided are modified rDNA-targeting TAL Arrays that have been modified to design around methylated CpG dinucleotides.

[0073] As discussed above, TAL modules contain a two amino acid region, called the repeat variable diresidue (RVD), that interacts with 1bp of DNA in a sequence specific manner. The amino acid sequence histidine-aspartate (HD) binds to the cytosine (C) DNA base for example. In mammalian cells, many cytosines become methylated (5mC) when followed by a guanine base (referred to as CpG methylation). Methylation of CpG sequences within TAL binding sites reduces binding of the TAL; however, swapping the RVD region of the affected module, for example to a single asparagine (N), can restore binding. CpG sequences are abundant in rDNA. The amino acid sequence of TAL module arrays designed with a N in place of the HD RVD for sites that contain a CpG (CpG rDNA TAL module arrays) are listed in Table 5 (SEQ ID NOs: 103-114).Attorney Docket No.: 000218-0153-WO1 Table 5: Illustrative CpG-Modified rDNA Targeting TAL ArraysAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1

[0074] Illustrative sequences of CpG-modified left TAL Arrays target a rDNA repeat element target site TTAA are set forth in SEQ ID NOs: 103, 105, 106, 108, 110, 112 and 114. Illustrative sequences of right TAL Arrays targeting a rDNA repeat element target site TTAA are set forth in SEQ ID NOs: 104, 107, 109, 111, and 113. TAL Arrays Targeting LINE1 Repeat Elements

[0075] In some embodiments, a TAL Array targets a LINE1 repeat element.

[0076] In certain aspects, the LINE1 TAL Arrays left and right TAL Arrays target the genomic LINE1 sequences shown in Table 6. Table 6: LINE1 TAL Array Target Sequences

[0077] Based on the LINE1 target sequence, a left and a right TAL Array targeting the 2L and 2.2R target site can be designed. The amino acid sequences of the LINE1 left and right TAL Arrays are shown in Table 7.Attorney Docket No.: 000218-0153-WO1 Table 7: LINE1 Left and Right TAL Arrays

[0078] In some embodiments, the left TAL array targeting a LINE1 repeat element comprises the sequence set forth in SEQ ID NO: 117. In some embodiments, the right TAL array targeting a LINE1 repeat element comprises the sequence set forth in SEQ ID NO: 118. Nucleolar Trafficking Sequences

[0079] In certain aspects, the fusion proteins described herein further comprise a nucleolar trafficking sequence. In certain embodiments, the nucleolar trafficking sequence is an unique RNA polymerase I (PolI) subunit. In certain embodiments, the nucleolar trafficking sequence is an R2 retrotransposon transposase sequence.

[0080] In certain embodiments, the nucleolar trafficking sequence is a unique RNA PolI subunit. PolI exclusively transcribes rDNA into rRNA. PolI is a complex and several subunits are shared among RNA polymerases (PolII and PolIII). Other subunits though, are unique to PolI and include A12.2 (POLR1H), A34.5 (POLR1G), A43 (POLR1F), and A49 (POLR1E).Attorney Docket No.: 000218-0153-WO1 The amino acid sequences of the unique A12.2 subunit, the A34.5 subunit, the A43 subunit and the A49 subunit are shown in Table 8. Table 8: Unique PolI Subunit Sequences

[0081] In certain embodiments, the nucleolar trafficking sequence is a portion of the R2 retrotransposon transposase. R2 retrotransposons integrate at a specific location within the 45S rDNA sequence. R2 retrotransposons are abundant in arthropods, but have also been found in other organisms, such as birds.Attorney Docket No.: 000218-0153-WO1

[0082] In certain embodiments, the R2 retrotransposon transposase sequence is a Bombyx mori R2 retrotransposase. In certain embodiments, the Bombyx mori R2 retrotransposase comprises, or consists entirely of, the amino acid sequence: MMASTALSLMGRCNPDGCTRGKHVTAAPMDGPRGPSSLAGTFGWGLAIPAGEPCG RVCSPATVGFFPVAKKSNKENRPEASGLPLESERTGDNPTVRGSAGADPVGQDAPG WTCQFCERTFSTNRGLGVHKRRAHPVETNTDAAPMMVKRRWHGEEIDLLARTEAR LLAERGQCSGGDLFGALPGFGRTLEAIKGQRRREPYRALVQAHLARFGSQPGPSSGG CSAEPDFRRASGAEEAGEERCAEDAAAYDPSAVGQMSPDAARVLSELLEGAGRRRA CRAMRPKTAGRRNDLHDDRTASAHKTSRQKRRAEYARVQELYKKCRSRAAAEVID GACGGVGHSLEEMETYWRPILERVSDAPGPTPEALHALGRAEWHGGNRDYTQLWK PISVEEIKASRFDWRTSPGPDGIRSGQWRAVPVHLKAEMFNAWMARGEIPEILRQCR TVFVPKVERPGGPGEYRPISIASIPLRHFHSILARRLLACCPPDARQRGFICADGTLENS AVLDAVLGDSRKKLRECHVAVLDFAKAFDTVSHEALVELLRLRGMPEQFCGYIAHL YDTASTTLAVNNEMSSPVKVGRGVRQGDPLSPILFNVVMDLILASLPERVGYRLEME LVSALAYADDLVLLAGSKVGMQESISAVDCVGRQMGLRLNCRKSAVLSMIPDGHR KKHHYLTERTFNIGGKPLRQVSCVERWRYLGVDFEASGCVTLEHSISSALNNISRAPL KPQQRLEILRAHLIPRFQHGFVLGNISDDRLRMLDVQIRKAVGQWLRLPADVPKAYY HAAVQDGGLAIPSVRATIPDLIVRRFGGLDSSPWSVARAAAKSDKIRKKLRWAWKQ LRRFSRVDSTTQRPSVRLFWREHLHASVDGRELRESTRTPTSTKWIRERCAQITGRDF VQFVHTHINALPSRIRGSRGRRGGGESSLTCRAGCKVRETTAHILQQCHRTHGGRILR HNKIVSFVAKAMEENKWTVELEPRLRTSVGLRKPDIIASRDGVGVIVDVQVVSGQRS LDELHREKRNKYGNHGELVELVAGRLGLPKAECVRATSCTISWRGVWSLTSYKELR SIIGLREPTLQIVPILALRGSHMNWTRFNQMTSVMGGGVG (SEQ ID NO: 123).

[0083] In certain embodiments, the R2 retrotransposon transposase sequence is a catalytically dead R2 retrotransposase. In certain embodiments, the catalytically dead endonuclease retrotransposase is a Bombyx mori R2 retrotransposase comprising, or consisting entirely of, the amino acid sequence: MMASTALSLMGRCNPDGCTRGKHVTAAPMDGPRGPSSLAGTFGWGLAIPAGEPCG RVCSPATVGFFPVAKKSNKENRPEASGLPLESERTGDNPTVRGSAGADPVGQDAPG WTCQFCERTFSTNRGLGVHKRRAHPVETNTDAAPMMVKRRWHGEEIDLLARTEAR LLAERGQCSGGDLFGALPGFGRTLEAIKGQRRREPYRALVQAHLARFGSQPGPSSGG CSAEPDFRRASGAEEAGEERCAEDAAAYDPSAVGQMSPDAARVLSELLEGAGRRRA CRAMRPKTAGRRNDLHDDRTASAHKTSRQKRRAEYARVQELYKKCRSRAAAEVID GACGGVGHSLEEMETYWRPILERVSDAPGPTPEALHALGRAEWHGGNRDYTQLWKAttorney Docket No.: 000218-0153-WO1 PISVEEIKASRFDWRTSPGPDGIRSGQWRAVPVHLKAEMFNAWMARGEIPEILRQCR TVFVPKVERPGGPGEYRPISIASIPLRHFHSILARRLLACCPPDARQRGFICADGTLENS AVLDAVLGDSRKKLRECHVAVLDFAKAFDTVSHEALVELLRLRGMPEQFCGYIAHL YDTASTTLAVNNEMSSPVKVGRGVRQGDPLSPILFNVVMDLILASLPERVGYRLEME LVSALAYADDLVLLAGSKVGMQESISAVDCVGRQMGLRLNCRKSAVLSMIPDGHR KKHHYLTERTFNIGGKPLRQVSCVERWRYLGVDFEASGCVTLEHSISSALNNISRAPL KPQQRLEILRAHLIPRFQHGFVLGNISDDRLRMLDVQIRKAVGQWLRLPADVPKAYY HAAVQDGGLAIPSVRATIPDLIVRRFGGLDSSPWSVARAAAKSDKIRKKLRWAWKQ LRRFSRVDSTTQRPSVRLFWREHLHASVDGRELRESTRTPTSTKWIRERCAQITGRDF VQFVHTHINALPSRIRGSRGRRGGGESSLTCRAGCKVRETTAHILQQCHRTHGGRILR HNKIVSFVAKAMEENKWTVELEPRLRTSVGLRKPAIIASRDGVGVIVAVQVVSGQRS LDELHREKRNKYGNHGELVELVAGRLGLPKAECVRATSCTISWRGVWSLTSYKELR SIIGLREPTLQIVPILALRGSHMNWTRFNQMTSVMGGGVG (SEQ ID NO: 124). Fusion Proteins Comprising Transposase Domain and TAL Arrays

[0084] Also provided herein are fusion proteins comprising one or more transposase domains described herein and a TAL array. A schematic illustration of a fusion protein binding to a rDNA TTAA target is shown in FIG.1.

[0085] In some embodiments, provided herein is a fusion protein comprising an rDNA- targeting TAL Array, and an SPB, PBx or PBx Plus transposase domain comprising a second C-terminal CRD in addition to the first endogenous CRD. Any TAL Array comprising one or more TAL domain targeting rDNA repeats or LINE1 repeat elements described herein may be combined with any transposase domain provided herein.

[0086] In some embodiments, a fusion protein provided herein comprises, in N-terminal to C-terminal order, a TAL Array targeting a rDNA repeat or a LINE1 repeat element, and a PBx Plus transposase domain comprising an N-terminal deletion and dual CRDs.

[0087] In some embodiments, a fusion protein provided herein comprises, in N-terminal to C-terminal order, a NLS, a nucleolar trafficking sequence, a TAL Array targeting a rDNA repeat or a LINE1 repeat element, and a PBx Plus transposase domain comprising an N- terminal deletion and dual CRDs.

[0088] In some embodiments, a fusion protein provided herein comprises, in N-terminal to C-terminal order, a Flag Tag sequence, a NLS, a protein stabilization domain (PSD), a TAL Array targeting a rDNA repeat or a LINE1 repeat element, and a PBx Plus transposase domain comprising an N-terminal deletion and dual CRDs.Attorney Docket No.: 000218-0153-WO1

[0089] In some embodiments, a fusion protein provided herein comprises, in N-terminal to C-terminal order, a nucleolar trafficking sequence, a PSD, a TAL Array targeting a rDNA repeat or a LINE1 repeat element, and a PBx Plus transposase domain comprising an N- terminal deletion and dual CRDs.

[0090] Schematic representations of the various fusion proteins design and configurations are shown in Fig.4. The fusion proteins described herein were configured using similar designs features and modular sequences. Briefly, all of the fusion proteins illustrated in FIG. 4 comprise a thermostable, hyperactive, an N-terminal deleted PBx transposase domain comprising a M298L thermostable mutation, S103P, a R372H, a S509G and a N571S hyperactivity mutations and dual CRDs. The sequence of the thermostable, hyperactive, N-terminally deleted PBx transposase domain comprising dual CRDs may be one of those in Table 1 (SEQ ID NOs: 8-28, or 208). In most instances, the TAL Array targeting, for example rDNA or a LINE1 repeat element, is flanked with linker sequences and the C-terminus of the TAL Array is fused to the N-terminus of the PBx transposase domain via the linker. A nucleic acid sequence comprising nuclear localization sequence (NLS) and optionally a Flag tag sequence is attached to the N-terminus of the TAL Array via the linker sequence. A nucleic acid comprising a nucleolar trafficking sequence, e.g., an unique PolI subunit, flanked by linker sequences may be attached upstream of the NLS or may be attached directly to the N-terminus of the TAL Array.

[0091] In certain embodiments, the fusion proteins comprise a protein stabilizing domain (PSD) flanked by linkers positioned between the NLS sequence and the N-terminus of the TAL Array sequence or the nucleolar trafficking sequence and the N-terminus of the TAL Array sequence. The PSD may comprise the N-terminal deleted nucleic acid sequence from the PBx Plus transposase domain. In one embodiment, the fusion protein comprises a 1-85 of the N-terminal deleted PBx Plus transposase comprising dual CRD proteins (SEQ ID NO: 10) and a PSD consisting of amino acids 1-85 of the PBx Plus transposase (SEQ ID NO: 190). In some embodiments, the fusion protein comprises a GG linker between the NLS and the PSD.

[0092] In certain embodiments, the fusion protein comprises in the N-terminal to C-terminal direction, a flag tag sequence, a NLS, a rDNA-targeting TAL Array and a PBx Plus transposase domain comprising dual CRDs. In certain embodiments, the fusion protein comprises in the N-terminal to C-terminal direction, a NLS, a rDNA-targeting TAL Array and a PBx Plus transposase domain comprising dual CRDs.Attorney Docket No.: 000218-0153-WO1

[0093] An exemplary sequence of a fusion protein comprising an N-terminal Flag tag (SEQ ID NO: 192) an NLS (SEQ ID NO: 191), a GG linker, a PSD (SEQ ID NO: 190), an rDNA-targeting 3L TAL Array (underlined; SEQ ID NO: 81) flanked by GGGGS linkers (SEQ ID NO: 125) and an N-terminal deletion of 1-85 amino acids of a PBx Plus transposase domain comprising dual CRDs(SEQ ID NO: 10) is shown in SEQ ID NO: 126, where the N-terminal Flag tag is shown in bold, underlined font, the NLS is shown in italics, the PSD is shown in italic, underlined font, the sequence comprising the 3L TAL Array and GGGGS linkers is underlined, and the transposase domain comprising an N-terminal deletion of 85 amino acid is shown in bold, the CRD linker sequence in bold lowercase font and the second CRD (SEQ ID NO: 7) in bold italics. MDYKDHDGDYKDHDIDYKDDDDKMAPKKKRKVGGGGSSLDDEHILSALLQSDDELV GEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILT LGGGGSVDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVALSQHPAA LGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPPLQLDTG QLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQALETVQRL LPVLCQDHGLTPEQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIG GKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLT PDQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQR LLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIASH DGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDH GLTPEQVVAIANNNGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNGGGKQALET VQRLLPVLCQAHGLTPAQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIA SNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQA HGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPEQVVAIANNNGGRPALE SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVKKGLPHAPALIKRTNRRIPERT SHRVADHAQVVRVLGGGGGSPQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQR GPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDE IYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDK SIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIP NKPSKYGIKILMLCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHG SCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTS MFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGV DTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRK KFMRNLYMGLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKAttorney Docket No.: 000218-0153-WO1 KRTYCTYCPSKIRRKASASCKKCKKVICREHNIDMCQSCFagggSTEEPVMKKRTY CTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 126) A person of skill in the art will appreciate that the N-terminal Flag-tag is present to facilitate protein purification and is not necessarily present in the constructs used to modify cells.

[0094] In certain embodiments, fusion proteins comprising the rRNA-targeting left and right TAL Arrays shown in Table 3 (SEQ ID NOs: 81-102) were similarly constructed as shown above. The fusion proteins comprising an rDNA-targeting TAL Array- 1-85 N- terminal deleted PBx Plus transposase domains are shown in Table 9 (SEQ ID NOs: 126- 147). Table 9: rDNA-targeting TAL Array – PBx Fusion Plus ProteinsAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1

[0095] In certain embodiments, fusion proteins were constructed comprising the modified CpG rDNA-targeting TAL Arrays shown in Table 5 using methods analogous for the construction of the rDNA TAL Array – PBx Plus fusion proteins in Table 9. Fusion proteins comprising the modified CpG rDNA-targeting TAL Arrays (SEQ ID NOs: 103-114) are shown in Table 10. Table 10: Modified CpG TAL Array -PBx Fusion ProteinsAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1LINE1 Repeat Element TAL Array – PBx Plus Fusion Proteins

[0096] In certain embodiments, fusion proteins were designed and constructed targeting LINE1 repeat elements. The fusion proteins comprise the LINE1 repeat element TAL Arrays in Table 4. The amino acid sequence of LINE1-targeting TAL Array – PBx Plus fusion proteins are shown in Table 11. Table 11: LINE1-targeting TAL Array – PBx Fusion ProteinsAttorney Docket No.: 000218-0153-WO1

[0097] In certain embodiments, the fusion proteins further comprise a nucleolar trafficking sequence. In certain embodiments, the nucleolar trafficking sequence is a PolI subunit. The presence of the nucleolar trafficking sequence facilitates entry from the nucleus to the nucleolus, where the rDNA is located (e.g., see Fig.3).

[0098] In certain embodiments, the fusion proteins comprise a nucleolar trafficking sequence attached directly to the N-terminus of a PBx Plus transposase domain comprising dual CRDs.Attorney Docket No.: 000218-0153-WO1

[0099] An exemplary sequence of a fusion protein comprising in the N-terminal to C- terminal direction: a sequence comprising a Flag tag sequence (SEQ ID NO: 192) and a NLS (SEQ ID NO: 191), a PolI A12.2 nucleolar trafficking sequence (SEQ ID NO: 119) flanked by GGGGS linkers (SEQ ID NO: 125); and an N-terminal deleted of 1-85 amino acids of a PBx Plus transposase domain comprising dual CRDs(SEQ ID NO: 10) separated by an AGGG linker (SEQ ID NO: 2), is shown in SEQ ID NO: 204, where the N-terminal Flag tag is shown in bold, underlined font, the NLS is shown in italics, the A12.2 sequence flanked by GGGGS linkers is underlined, and the PBx Plus transposase domain comprising an N- terminal deletion of 85 amino acid is shown in bold, the CRD linker sequence in bold lowercase font and the second CRD (SEQ ID NO: 7) in bold italics. MDYKDHDGDYKDHDIDYKDDDDKMAPKKKRKVGGGGSMSVMDLANTCSSFQSD LDFCSDCGSVLPLPGAQDTVTCIRCGFNINVRDFEGKVVKTSVVFHQLGTAMPMSVE EGPECQGPVVDRRCPRCGHEGMAYHTRQMRSADEGQTVFYTCTNCKFQEKEDSGG GGSPQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFK LFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDN HMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRK IWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMLCDS GTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPL AKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPK PAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRK TNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMGLTSSFM RKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKA SASCKKCKKVICREHNIDMCQSCFagggSTEEPVMKKRTYCTYCPSKIRRKANASCK KCKKVICREHNIDMCQSCF (SEQ ID NO: 204).

[0100] Illustrative PolI subunit – PBx Plus fusion proteins comprising the A12.2, A34.5, A43 and A49 subunits are shown in Table 12. Table 12. PolI subunit – PBx fusion proteinsAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1

[0101] An exemplary sequence of a fusion protein comprising an NLS (SEQ ID NO: 191), a PolI A12.2 nucleolar trafficking sequence (SEQ ID NO: 119); a sequence comprising an rDNA-targeting 5L TAL Array (SEQ ID NO: 85) flanked by GGGGS linkers (SEQ ID NO: 125), and an N-terminal deletion of 1-85 amino acids of a PBx Plus transposase domain comprising dual CRDs (SEQ ID NO: 10) separated by an AGGG linker (SEQ ID NO: 2), is shown in SEQ ID NO: 162, where the NLS is shown in italics, the A12.2 sequence flanked by GGGGS linkers is in italics and underlined, the rDNA-targeting 5L TAL Array with a C-terminal GGGGS linker is underlined and the PBx Plus transposase domain comprising an N-terminal deletion of 85 amino acid is shown in bold, the CRD linker sequence in bold lowercase font and the second CRD (SEQ ID NO: 7) in bold italics. MAPKKKRKVGGGGSMSVMDLANTCSSFQSDLDFCSDCGSVLPLPGAQDTVTCIRCGFNI NVRDFEGKVVKTSVVFHQLGTAMPMSVEEGPECQGPVVDRRCPRCGHEGMAYHTRQM RSADEGQTVFYTCTNCKFQEKEDSGGGGSVDLRTLGYSQQQQEKIKPKVRSTVAQHHAttorney Docket No.: 000218-0153-WO1 EALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGA RALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTP DQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLL PVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNGG GKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLT PEQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPDQVVAIASNIGGKQALETVQRL LPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHD GGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAAL TNDHLVALACLGGRPALDAVKKGLPHAPALIKRTNRRIPERTSHRVADHAQVVRVL GGGGGSPQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQRGPTRMCRNIYDPLL CFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVR KDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTP VRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILML CDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTS IPLAKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSY KPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTC SRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMGLTS SFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIR RKASASCKKCKKVICREHNIDMCQSCFagggSTEEPVMKKRTYCTYCPSKIRRKANA SCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 162).

[0102] In certain embodiments, fusion proteins were designed and constructed that comprise the A12.2, A 34.5, A43 and A49 PolI subunits, the TAL Arrays targeting rDNA TTAA sites 5 or 23 (5L / 5R or 23L / 23R) and the PBx Plus transposase comprising dual CRDs. The amino acid sequences for the PolI – rDNA TAL Array – PBx Plus transposase fusion proteins are shown in Table 13. Table 13: PolI – rDNA TAL Array – PBx Plus Fusion ProteinsAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1

[0103] In certain embodiments, LINE1-targeting TAL Array PBx Plus transposase constructs comprise a nucleolar trafficking sequence. In certain embodiments, the nucleolar trafficking sequence is a unique PolI subunit.

[0104] An exemplary sequence of a fusion protein comprising an NLS (SEQ ID NO: 191), a PolI A12.2 nucleolar trafficking sequence (SEQ ID NO: 119); a sequence comprising an LINE1-targeting 2L TAL Array (underlined; SEQ ID NO: 117) flanked by GGGGS linkers (SEQ ID NO: 125) and an N-terminal deletion of 1-85 amino acids of a PBx Plus transposase domain comprising dual CRDs(SEQ ID NO: 10) is show in SEQ ID NO: 178, where the NLS is shown in italics, the A12.2 sequence flanked by GGGGS linkers is in italics and underlined, the LINE12L TAL Array with a C-terminal GGGGS linker is underlined and the PBx Plus transposase domain comprising an N-terminal deletion of 85 amino acid is shown in bold, the CRD linker sequence in bold lowercase font and the second CRD (SEQ ID NO: 7) in bold italics.Attorney Docket No.: 000218-0153-WO1 MAPKKKRKVGGGGSMSVMDLANTCSSFQSDLDFCSDCGSVLPLPGAQDTVTCIRCGFNI NVRDFEGKVVKTSVVFHQLGTAMPMSVEEGPECQGPVVDRRCPRCGHEGMAYHTRQM RSADEGQTVFYTCTNCKFQEKEDSGGGGSVDLRTLGYSQQQQEKIKPKVRSTVAQHH EALVGHGFTHAHIVALSQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGA RALEALLTVAGELRGPPLQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTP DQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPEQVVAIASHDGGKQALETVQRL LPVLCQAHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPAQVVAIASNIG GKQALETVQRLLPVLCQDHGLTPDQVVAIANNNGGKQALETVQRLLPVLCQDHGLT PEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQR LLPVLCQAHGLTPAQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASH DGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAA LTNDHLVALACLGGRPALDAVKKGLPHAPALIKRTNRRIPERTSHRVADHAQVVRV LGGGGGSPQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQRGPTRMCRNIYDPL LCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAV RKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFT PVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILM LCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFT SIPLAKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSY KPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTC SRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMGLTS SFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIR RKASASCKKCKKVICREHNIDMCQSCFagggSTEEPVMKKRTYCTYCPSKIRRKANA SCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 178)

[0105] The amino acid sequences of LINE1-targeting fusion proteins comprising other LINE1 TAL Arrays and additional PolI subunits are shown Table 14. Table 14: PolI-LINE1 TAL Array – PBx Plus Fusion ProteinsAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1

[0106] In certain embodiments, the nucleolar trafficking sequence is an R2 retrotransposon transposase.

[0107] An exemplary sequence of a fusion protein comprising, a Flag tag sequence (SEQ ID NO: 192), a NLS (SEQ ID NO: 191), a wild type R2 retrotransposon nucleolar trafficking sequence (SEQ ID NO: 123); a PSD (comprising residues 4-85 of SEQ ID NO: 190); a sequence comprising an rDNA-targeting 23L TAL Array (underlined; SEQ ID NO: 101) flanked by GGGGS linkers (SEQ ID NO: 125) and an N-terminal deletion of 1-85 amino acids of a PBx Plus transposase domain comprising dual CRDs (SEQ ID NO: 10) is show in SEQ ID NO: 186, where the N-terminal Flag tag is shown in bold, underlined font, the NLS is shown in italics, the R2 retrotransposon sequence flanked by GGGGS linkers is in italics and underlined, the PSD is in plain lowercase font; the rDNA-targeting 23L TAL Array flanked by GGGGS linkers is underlined and the PBx Plus transposase domain comprising an N-terminal deletion of 85 amino acid is shown in bold, the CRD linker sequence in bold lowercase font and the second CRD (SEQ ID NO: 7) in bold italics. MDYKDHDGDYKDHDIDYKDDDDKMAPKKKRKVGGGGSTALSLMGRCNPDGCTRG KHVTAAPMDGPRGPSSLAGTFGWGLAIPAGEPCGRVCSPATVGFFPVAKKSNKENRPEAS GLPLESERTGDNPTVRGSAGADPVGQDAPGWTCQFCERTFSTNRGLGVHKRRAHPVETN TDAAPMMVKRRWHGEEIDLLARTEARLLAERGQCSGGDLFGALPGFGRTLEAIKGQRRRAttorney Docket No.: 000218-0153-WO1 EPYRALVQAHLARFGSQPGPSSGGCSAEPDFRRASGAEEAGEERCAEDAAAYDPSAVGQ MSPDAARVLSELLEGAGRRRACRAMRPKTAGRRNDLHDDRTASAHKTSRQKRRAEYARV QELYKKCRSRAAAEVIDGACGGVGHSLEEMETYWRPILERVSDAPGPTPEALHALGRAEW HGGNRDYTQLWKPISVEEIKASRFDWRTSPGPDGIRSGQWRAVPVHLKAEMFNAWMARG EIPEILRQCRTVFVPKVERPGGPGEYRPISIASIPLRHFHSILARRLLACCPPDARQRGFICA DGTLENSAVLDAVLGDSRKKLRECHVAVLDFAKAFDTVSHEALVELLRLRGMPEQFCGYI AHLYDTASTTLAVNNEMSSPVKVGRGVRQGDPLSPILFNVVMDLILASLPERVGYRLEMEL VSALAYADDLVLLAGSKVGMQESISAVDCVGRQMGLRLNCRKSAVLSMIPDGHRKKHHYL TERTFNIGGKPLRQVSCVERWRYLGVDFEASGCVTLEHSISSALNNISRAPLKPQQRLEILR AHLIPRFQHGFVLGNISDDRLRMLDVQIRKAVGQWLRLPADVPKAYYHAAVQDGGLAIPS VRATIPDLIVRRFGGLDSSPWSVARAAAKSDKIRKKLRWAWKQLRRFSRVDSTTQRPSVRL FWREHLHASVDGRELRESTRTPTSTKWIRERCAQITGRDFVQFVHTHINALPSRIRGSRGR RGGGESSLTCRAGCKVRETTAHILQQCHRTHGGRILRHNKIVSFVAKAMEENKWTVELEP RLRTSVGLRKPDIIASRDGVGVIVDVQVVSGQRSLDELHREKRNKYGNHGELVELVAGRLG LPKAECVRATSCTISWRGVWSLTSYKELRSIIGLREPTLQIVPILALRGSHMNWTRFNQMTS VMGGGVGGGGSslddehilsallqsddelvgedsdsevsdhvseddvqsdteeafidevhevqptssgseildeqnvieq pgsslasnriltlGGGGSVDLRTLGYSQQQQEKIKPKVRSTVAQHHEALVGHGFTHAHIVAL SQHPAALGTVAVKYQDMIAALPEATHEAIVGVGKQWSGARALEALLTVAGELRGPP LQLDTGQLLKIAKRGGVTAVEAVHAWRNALTGAPLNLTPDQVVAIASHDGGKQAL ETVQRLLPVLCQDHGLTPEQVVAIASNIGGKQALETVQRLLPVLCQAHGLTPDQVVA IASHDGGKQALETVQRLLPVLCQAHGLTPAQVVAIANNNGGKQALETVQRLLPVLC QDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPEQVVAIASNGGGKQA LETVQRLLPVLCQAHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQAHGLTPAQV VAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPV LCQDHGLTPEQVVAIASNGGGKQALETVQRLLPVLCQAHGLTPDQVVAIASHDGGK QALETVQRLLPVLCQAHGLTPEQVVAIASNGGGRPALESIVAQLSRPDPALAALTND HLVALACLGGRPALDAVKKGLPHAPALIKRTNRRIPERTSHRVADHAQVVRVLGGG GGSPQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFK LFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDN HMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRK IWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMLCDS GTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPL AKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPK PAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKAttorney Docket No.: 000218-0153-WO1 TNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMGLTSSFM RKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKA SASCKKCKKVICREHNIDMCQSCFagggSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 186).

[0108] Illustrative amino acid sequences of rDNA-targeting fusion proteins comprising other rDNA TAL Arrays and additional R2 retrotransposon sequences are shown Table 15. Table 15: R2 Retrotransposon – rDNA TAL Array – PBx Plus Fusion ProteinsAttorney Docket No.: 000218-0153-WO1Attorney Docket No.: 000218-0153-WO1

[0109] In certain embodiments, the C-terminus of the R2 retrotransposon sequence is attached via a linker sequence to the N-terminus of the N-terminal deleted PBx Plus sequence (e.g., delta 1-85). An illustrative amino acid sequence of an R2 retrotransposon sequence – 1- 85 PBx Plus fusion protein comprising an N-terminal flag tag (bold: SEQ ID NO; 192), a NLS (italics; SEQ ID NO: 191) and PBx Plus 1-85 transposase domain comprising dual CRDs (SEQ ID NO: 10; bold / italics) is shown in SEQ ID NO: 188: MDYKDHDGDYKDHDIDYKDDDDKMAPKKKRKVGGGGSTALSLMGRCNPDGCTR GKHVTAAPMDGPRGPSSLAGTFGWGLAIPAGEPCGRVCSPATVGFFPVAKKSNKEN RPEASGLPLESERTGDNPTVRGSAGADPVGQDAPGWTCQFCERTFSTNRGLGVHKR RAHPVETNTDAAPMMVKRRWHGEEIDLLARTEARLLAERGQCSGGDLFGALPGFG RTLEAIKGQRRREPYRALVQAHLARFGSQPGPSSGGCSAEPDFRRASGAEEAGEERC AEDAAAYDPSAVGQMSPDAARVLSELLEGAGRRRACRAMRPKTAGRRNDLHDDRT ASAHKTSRQKRRAEYARVQELYKKCRSRAAAEVIDGACGGVGHSLEEMETYWRPIL ERVSDAPGPTPEALHALGRAEWHGGNRDYTQLWKPISVEEIKASRFDWRTSPGPDGI RSGQWRAVPVHLKAEMFNAWMARGEIPEILRQCRTVFVPKVERPGGPGEYRPISIASI PLRHFHSILARRLLACCPPDARQRGFICADGTLENSAVLDAVLGDSRKKLRECHVAV LDFAKAFDTVSHEALVELLRLRGMPEQFCGYIAHLYDTASTTLAVNNEMSSPVKVG RGVRQGDPLSPILFNVVMDLILASLPERVGYRLEMELVSALAYADDLVLLAGSKVGAttorney Docket No.: 000218-0153-WO1 MQESISAVDCVGRQMGLRLNCRKSAVLSMIPDGHRKKHHYLTERTFNIGGKPLRQV SCVERWRYLGVDFEASGCVTLEHSISSALNNISRAPLKPQQRLEILRAHLIPRFQHGFV LGNISDDRLRMLDVQIRKAVGQWLRLPADVPKAYYHAAVQDGGLAIPSVRATIPDLI VRRFGGLDSSPWSVARAAAKSDKIRKKLRWAWKQLRRFSRVDSTTQRPSVRLFWRE HLHASVDGRELRESTRTPTSTKWIRERCAQITGRDFVQFVHTHINALPSRIRGSRGRR GGGESSLTCRAGCKVRETTAHILQQCHRTHGGRILRHNKIVSFVAKAMEENKWTVE LEPRLRTSVGLRKPDIIASRDGVGVIVDVQVVSGQRSLDELHREKRNKYGNHGELVE LVAGRLGLPKAECVRATSCTISWRGVWSLTSYKELRSIIGLREPTLQIVPILALRGSHM NWTRFNQMTSVMGGGVGGGGSPQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQRG PTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAF FGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTL RENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGI KILMLCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWF TSIPLAKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKP KPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNR WPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMGLTSSFMRKRLEA PTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKASASCKKCKKV ICREHNIDMCQSCFAGGGSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREH NIDMCQSCF (SEQ ID NO: 188).

[0110] In certain embodiments, the C-terminus of the R2 catalytic dead endonuclease retrotransposon sequence is attached via a linker sequence to the N-terminus of the N- terminal deleted PBx Plus sequence (e.g., delta 1-85). An illustrative amino acid sequence of an R2 retrotransposon sequence – 1-85 PBx Plus fusion protein comprising an N-terminal flag tag (bold: SEQ ID NO; 192), a NLS (italics; SEQ ID NO: 191) and PBx Plus 1-85 transposase domain comprising dual CRDs (SEQ ID NO: 10; bold / italics) is shown in SEQ ID NO: 189. MDYKDHDGDYKDHDIDYKDDDDKMAPKKKRKVGGGGSTALSLMGRCNPDGCTR GKHVTAAPMDGPRGPSSLAGTFGWGLAIPAGEPCGRVCSPATVGFFPVAKKSNKEN RPEASGLPLESERTGDNPTVRGSAGADPVGQDAPGWTCQFCERTFSTNRGLGVHKR RAHPVETNTDAAPMMVKRRWHGEEIDLLARTEARLLAERGQCSGGDLFGALPGFG RTLEAIKGQRRREPYRALVQAHLARFGSQPGPSSGGCSAEPDFRRASGAEEAGEERC AEDAAAYDPSAVGQMSPDAARVLSELLEGAGRRRACRAMRPKTAGRRNDLHDDRT ASAHKTSRQKRRAEYARVQELYKKCRSRAAAEVIDGACGGVGHSLEEMETYWRPIL ERVSDAPGPTPEALHALGRAEWHGGNRDYTQLWKPISVEEIKASRFDWRTSPGPDGIAttorney Docket No.: 000218-0153-WO1 RSGQWRAVPVHLKAEMFNAWMARGEIPEILRQCRTVFVPKVERPGGPGEYRPISIASI PLRHFHSILARRLLACCPPDARQRGFICADGTLENSAVLDAVLGDSRKKLRECHVAV LDFAKAFDTVSHEALVELLRLRGMPEQFCGYIAHLYDTASTTLAVNNEMSSPVKVG RGVRQGDPLSPILFNVVMDLILASLPERVGYRLEMELVSALAYADDLVLLAGSKVG MQESISAVDCVGRQMGLRLNCRKSAVLSMIPDGHRKKHHYLTERTFNIGGKPLRQV SCVERWRYLGVDFEASGCVTLEHSISSALNNISRAPLKPQQRLEILRAHLIPRFQHGFV LGNISDDRLRMLDVQIRKAVGQWLRLPADVPKAYYHAAVQDGGLAIPSVRATIPDLI VRRFGGLDSSPWSVARAAAKSDKIRKKLRWAWKQLRRFSRVDSTTQRPSVRLFWRE HLHASVDGRELRESTRTPTSTKWIRERCAQITGRDFVQFVHTHINALPSRIRGSRGRR GGGESSLTCRAGCKVRETTAHILQQCHRTHGGRILRHNKIVSFVAKAMEENKWTVE LEPRLRTSVGLRKPAIIASRDGVGVIVAVQVVSGQRSLDELHREKRNKYGNHGELVE LVAGRLGLPKAECVRATSCTISWRGVWSLTSYKELRSIIGLREPTLQIVPILALRGSHM NWTRFNQMTSVMGGGVGGGGSPQRTIRGKNKHCWSTSKPTRRSRVSALNIVRSQRG PTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAF FGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTL RENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGI KILMLCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWF TSIPLAKNLLQEPYKLTIVGTVHSNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKP KPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNR WPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMGLTSSFMRKRLEA PTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKASASCKKCKKV ICREHNIDMCQSCFAGGGSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREH NIDMCQSCF (SEQ ID NO: 189). Protein Stabilization Domains

[0111] In some embodiments, a fusion protein provided herein may further comprise a protein stabilization domain (PSD). The PSD is preferably attached to the N-terminus of the TAL Array targeting domain, if present. Without wishing to be bound by theory, it is believed that the addition of a PSD can enhance protein stability or enhanced stability of the transposase dimer or tetramer – DNA complex.

[0112] The PSD may be of approximately the same size as the N-terminal deletion in the transposase domain. For example, in some embodiments, the N-terminal deletion of transposase domain comprises amino acids 1-85, and the PSD comprises 85 amino acids.Attorney Docket No.: 000218-0153-WO1

[0113] In some embodiments, the PSD comprises the sequence GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTS SGSEILDEQNVIEQPGSSLASNRILTL (SEQ ID NO: 190). In some embodiments, when a GGGGS linker (SEQ ID NO: 125) is used immediately upstream of the PSD, the first three amino acid residues (GGS) of the PSD are removed to prevent a repetitive sequence. In such embodiments, the PSD comprises amino acid residues 4 to 85 of SEQ ID NO: 190.

[0114] Illustrative fusion proteins targeting rDNA and LINE1 elements comprising a PSD are set forth, for example, in SEQ ID NOs: 126-147, 148-161. Nucleic Acids

[0115] Also provided herein are polynucleotides comprising nucleic acid sequences encoding the fusion proteins described herein. In some embodiments, the polynucleotides are isolated.

[0116] The disclosure provides isolated or substantially purified polynucleotide or protein compositions. An "isolated" or "purified" polynucleotide or protein, or biologically active portion thereof, is substantially or essentially free from components that normally accompany or interact with the polynucleotide or protein as found in its naturally occurring environment. Thus, an isolated or purified polynucleotide or protein is substantially free of other cellular material or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. Optimally, an "isolated" polynucleotide is free of sequences (optimally protein encoding sequences) that naturally flank the polynucleotide (i.e., sequences located at the 5' and 3' ends of the polynucleotide) in the genomic DNA of the organism from which the polynucleotide is derived. For example, in various aspects, the isolated polynucleotide can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequence that naturally flank the polynucleotide in genomic DNA of the cell from which the polynucleotide is derived. A protein that is substantially free of cellular material includes preparations of protein having less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of contaminating protein. When the protein of the disclosure or biologically active portion thereof is recombinantly produced, optimally culture medium represents less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of chemical precursors or non-protein-of-interest chemicals.

[0117] The disclosure provides fragments and variants of the disclosed DNA sequences and proteins encoded by these DNA sequences. As used throughout the disclosure, the term "fragment" refers to a portion of the DNA sequence or a portion of the amino acid sequenceAttorney Docket No.: 000218-0153-WO1 and hence protein encoded thereby. Fragments of a DNA sequence comprising coding sequences may encode protein fragments that retain biological activity of the native protein and hence DNA recognition or binding activity to a target DNA sequence as herein described. Alternatively, fragments of a DNA sequence that are useful as hybridization probes generally do not encode proteins that retain biological activity or do not retain promoter activity. Thus, fragments of a DNA sequence may range from at least about 20 nucleotides, about 50 nucleotides, about 100 nucleotides, and up to the full-length polynucleotide of the disclosure.

[0118] Nucleic acids or proteins of the disclosure can be constructed by a modular approach including preassembling monomer units and / or repeat units in target vectors that can subsequently be assembled into a final destination vector. Polypeptides of the disclosure may comprise repeat monomers of the disclosure and can be constructed by a modular approach by preassembling repeat units in target vectors that can subsequently be assembled into a final destination vector. The disclosure provides polypeptide produced by this method as well nucleic acid sequences encoding these polypeptides. The disclosure provides host organisms and cells comprising nucleic acid sequences encoding polypeptides produced this modular approach.

[0119] Nucleic acids of the disclosure may be single- or double-stranded. Nucleic acids of the disclosure may contain double-stranded sequences even when the majority of the molecule is single-stranded. Nucleic acids of the disclosure may contain single-stranded sequences even when the majority of the molecule is double-stranded. Nucleic acids of the disclosure may include genomic DNA, cDNA, RNA, or a hybrid thereof. Nucleic acids of the disclosure may contain combinations of deoxyribo- and ribo-nucleotides. Nucleic acids of the disclosure may contain combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine hypoxanthine, isocytosine and isoguanine. Nucleic acids of the disclosure may be synthesized to comprise non-natural amino acid modifications. Nucleic acids of the disclosure may be obtained by chemical synthesis methods or by recombinant methods.

[0120] Nucleic acids of the disclosure, either their entire sequence, or any portion thereof, may be non-naturally occurring. Nucleic acids of the disclosure may contain one or more mutations, substitutions, deletions, or insertions that do not naturally-occur, rendering the entire nucleic acid sequence non-naturally occurring. Nucleic acids of the disclosure may contain one or more duplicated, inverted or repeated sequences, the resultant sequence of which does not naturally-occur, rendering the entire nucleic acid sequence non-naturallyAttorney Docket No.: 000218-0153-WO1 occurring. Nucleic acids of the disclosure may contain modified, artificial, or synthetic nucleotides that do not naturally-occur, rendering the entire nucleic acid sequence non- naturally occurring.

[0121] Given the redundancy in the genetic code, a plurality of nucleotide sequences may encode any particular protein. All such nucleotides sequences are contemplated herein.

[0122] The isolated polynucleotides of the disclosure can be made using (a) recombinant methods, (b) synthetic techniques, (c) purification techniques, and / or (d) combinations thereof, as well-known in the art.

[0123] Methods of constructing nucleic acids encoding the PBx Plus transposase domains comprising an N-terminal deletion described herein are well known in the art or described herein, for example, PCR-based mutagenesis. Exemplary primers that may be used to construct transposase domains comprising an N-terminal deletion are shown in Table 16. Table 16: Exemplary Primer Sequences

[0124] The fusion of the present invention can be generated using any suitable method known in the art or described herein.

[0125] The isolated polynucleotides of this disclosure, such as RNA, cDNA, genomic DNA, or any combination thereof, can be obtained from biological sources using any number of cloning methodologies known to those of skill in the art. In some aspects, oligonucleotide probes that selectively hybridize, under stringent conditions, to the polynucleotides of the present disclosure are used to identify the desired sequence in a cDNA or genomic DNA library.

[0126] Methods of amplification of RNA or DNA are well known in the art and can be used according to the disclosure without undue experimentation, based on the teaching and guidance presented herein. Known methods of DNA or RNA amplification include, but are not limited to, polymerase chain reaction (PCR) and related amplification processes (see, e.g., U.S. Pat. Nos.4,683,195, 4,683,202, 4,800,159, 4,965,188, to Mullis, et al.; 4,795,699 andAttorney Docket No.: 000218-0153-WO1 4,921,794 to Tabor, et al; 5,142,033 to Innis; 5,122,464 to Wilson, et al.; 5,091,310 to Innis; 5,066,584 to Gyllensten, et al; 4,889,818 to Gelfand, et al; 4,994,370 to Silver, et al; 4,766,067 to Biswas; 4,656,134 to Ringold) and RNA mediated amplification that uses anti- sense RNA to the target sequence as a template for double-stranded DNA synthesis (U.S. Pat. No.5,130,238 to Malek, et al, with the tradename NASBA), the entire contents of which references are incorporated herein by reference. (See, e.g., Ausubel, supra; or Sambrook, supra.)

[0127] For instance, polymerase chain reaction (PCR) technology can be used to amplify the sequences of polynucleotides of the disclosure and related genes directly from genomic DNA or cDNA libraries. PCR and other in vitro amplification methods can also be useful, for example, to clone nucleic acid sequences that code for proteins to be expressed, to make nucleic acids to use as probes for detecting the presence of the desired mRNA in samples, for nucleic acid sequencing, or for other purposes. Examples of techniques sufficient to direct persons of skill through in vitro amplification methods are found in Berger, supra, Sambrook, supra, and Ausubel, supra, as well as Mullis, et al., U.S. Pat. No.4,683,202 (1987); and Innis, et al., PCR Protocols A Guide to Methods and Applications, Eds., Academic Press Inc., San Diego, Calif. (1990). Commercially available kits for genomic PCR amplification are known in the art. See, e.g., Advantage-GC Genomic PCR Kit (Clontech). Additionally, e.g., the T4 gene 32 protein (Boehringer Mannheim) can be used to improve yield of long PCR products.

[0128] The polynucleotides of the disclosure can also be prepared by direct chemical synthesis by known methods (see, e.g., Ausubel, et al., supra). Chemical synthesis generally produces a single-stranded oligonucleotide, which can be converted into double-stranded DNA by hybridization with a complementary sequence, or by polymerization with a DNA polymerase using the single strand as a template. One of skill in the art will recognize that while chemical synthesis of DNA can be limited to sequences of about 100 or more bases, longer sequences can be obtained by the ligation of shorter sequences. Expression Vectors and Host Cells

[0129] The disclosure also relates to vectors that include polynucleotides of the disclosure, host cells that are genetically engineered with the recombinant vectors, and the production of at least one protein scaffold by recombinant techniques, as is well known in the art. See, e.g., Sambrook, et al., supra; Ausubel, et al., supra, each entirely incorporated herein by reference.Attorney Docket No.: 000218-0153-WO1

[0130] The polynucleotides can optionally be joined to a vector containing a selectable marker for propagation in a host. Generally, a plasmid vector is introduced in a precipitate, such as a calcium phosphate precipitate, or in a complex with a charged lipid. If the vector is a virus, it can be packaged in vitro using an appropriate packaging cell line and then transduced into host cells.

[0131] The DNA insert should be operatively linked to an appropriate promoter. In some embodiments, the promoter is an EF-1α promoter. The expression constructs will further contain sites for transcription initiation, termination and, in the transcribed region, a ribosome binding site for translation. The coding portion of the mature transcripts expressed by the constructs will preferably include a translation initiating at the beginning and a termination codon (e.g., UAA, UGA or UAG) appropriately positioned at the end of the mRNA to be translated, with UAA and UAG preferred for mammalian or eukaryotic cell expression.

[0132] As used throughout the disclosure, the term "operably linked" refers to the expression of a gene that is under the control of a promoter with which it is spatially connected. A promoter can be positioned 5' (upstream) or 3' (downstream) of a gene under its control. The distance between a promoter and a gene can be approximately the same as the distance between that promoter and the gene it controls in the gene from which the promoter is derived. Variation in the distance between a promoter and a gene can be accommodated without loss of promoter function.

[0133] As used throughout the disclosure, the term "promoter" refers to a synthetic or naturally-derived molecule which is capable of conferring, activating or enhancing expression of a nucleic acid in a cell. A promoter can comprise one or more specific transcriptional regulatory sequences to further enhance expression and / or to alter the spatial expression and / or temporal expression of same. A promoter can also comprise distal enhancer or repressor elements, which can be located as much as several thousand base pairs from the start site of transcription. A promoter can be derived from sources including viral, bacterial, fungal, plants, insects, and animals. A promoter can regulate the expression of a gene component constitutively or differentially with respect to cell, the tissue or organ in which expression occurs or, with respect to the developmental stage at which expression occurs, or in response to external stimuli such as physiological stresses, pathogens, metal ions, or inducing agents. Representative examples of promoters include the bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator-promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, EF-1 Alpha promoter, CAG promoter, SV40 early promoter or SV40 late promoter and the CMVAttorney Docket No.: 000218-0153-WO1 IE promoter. A vector can be a viral vector, bacteriophage, bacterial artificial chromosome or yeast artificial chromosome. A vector can be a DNA or RNA vector. A vector can be a self- replicating extrachromosomal vector, and preferably, is a DNA plasmid. A vector may comprise a combination of an amino acid with a DNA sequence, an RNA sequence, or both a DNA and an RNA sequence.

[0134] Expression vectors will preferably but optionally include at least one selectable marker. Such markers include, e.g., but are not limited to, ampicillin, zeocin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / Geneticin (neo gene), DHFR (encoding Dihydrofolate Reductase and conferring resistance to Methotrexate), mycophenolic acid, or glutamine synthetase (GS, U.S. Pat. Nos.5,122,464; 5,770,359; 5,827,739), blasticidin (bsd gene), resistance genes for eukaryotic cell culture as well as ampicillin, zeocin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / Geneticin (neo gene), kanamycin, spectinomycin, streptomycin, carbenicillin, bleomycin, erythromycin, polymyxin B, or tetracycline resistance genes for culturing in E. coli and other bacteria or prokaryotes (the above patents are entirely incorporated hereby by reference). Appropriate culture mediums and conditions for the above-described host cells are known in the art. Suitable vectors will be readily apparent to the skilled artisan. Introduction of a vector construct into a host cell can be effected by calcium phosphate transfection, DEAE-dextran mediated transfection, cationic lipid-mediated transfection, electroporation, transduction, infection or other known methods. Such methods are described in the art, such as Sambrook, supra, Chapters 1-4 and 16-18; Ausubel, supra, Chapters 1, 9, 13, 15, 16.

[0135] Expression vectors will preferably but optionally include at least one selectable cell surface marker for isolation of cells modified by the compositions and methods of the disclosure. Selectable cell surface markers of the disclosure comprise surface proteins, glycoproteins, or group of proteins that distinguish a cell or subset of cells from another defined subset of cells. Preferably the selectable cell surface marker distinguishes those cells modified by a composition or method of the disclosure from those cells that are not modified by a composition or method of the disclosure. Such cell surface markers include, e.g., but are not limited to, “cluster of designation” or “classification determinant” proteins (often abbreviated as “CD”) such as a truncated or full length form of CD19, CD271, CD34, CD22, CD20, CD33, CD52, or any combination thereof. Cell surface markers further include the suicide gene marker RQR8 (Philip B et al. Blood.2014 Aug 21; 124(8):1277-87).

[0136] Expression vectors will preferably but optionally include at least one selectable drug resistance marker for isolation of cells modified by the compositions and methods of theAttorney Docket No.: 000218-0153-WO1 disclosure. Selectable drug resistance markers of the disclosure may comprise wild-type or mutant Neo, DHFR, TYMS, FRANCF, RAD51C, GCS, MDR1, ALDH1, NKX2.2, or any combination thereof.

[0137] Those of ordinary skill in the art are knowledgeable in the numerous expression systems available for expression of a nucleic acid encoding a protein of the disclosure. Alternatively, nucleic acids of the disclosure can be expressed in a host cell by turning on (by manipulation) in a host cell that contains endogenous DNA encoding a protein scaffold of the disclosure. Such methods are well known in the art, e.g., as described in U.S. Pat. Nos. 5,580,734, 5,641,670, 5,733,746, and 5,733,761, entirely incorporated herein by reference.

[0138] Illustrative of cell cultures useful for the production of the protein scaffolds, specified portions or variants thereof, are bacterial, yeast, and mammalian cells as known in the art. Mammalian cell systems often will be in the form of monolayers of cells although mammalian cell suspensions or bioreactors can also be used. A number of suitable host cell lines capable of expressing intact glycosylated proteins have been developed in the art, and include the COS-1 (e.g., ATCC CRL 1650), COS-7 (e.g., ATCC CRL-1651), HEK293, BHK21 (e.g., ATCC CRL-10), CHO (e.g., ATCC CRL 1610) and BSC-1 (e.g., ATCC CRL- 26) cell lines, Cos-7 cells, CHO cells, hep G2 cells, P3X63Ag8.653, SP2 / 0-Ag14, 293 cells, HeLa cells and the like, which are readily available from, for example, American Type Culture Collection, Manassas, Va. (www.atcc.org). Preferred host cells include cells of lymphoid origin, such as myeloma and lymphoma cells. Particularly preferred host cells are P3X63Ag8.653 cells (ATCC Accession Number CRL-1580) and SP2 / 0-Ag14 cells (ATCC Accession Number CRL-1851). In a preferred aspect, the recombinant cell is a P3X63Ab8.653 or an SP2 / 0-Ag14 cell.

[0139] Expression vectors for these cells can include one or more of the following expression control sequences, such as, but not limited to, an origin of replication; a promoter (e.g., late or early SV40 promoters, the CMV promoter (U.S. Pat. Nos.5,168,062; 5,385,839), an HSV tk promoter, a pgk (phosphoglycerate kinase) promoter, an EF-1 alpha promoter (U.S. Pat. No.5,266,491), at least one human promoter; an enhancer, and / or processing information sites, such as ribosome binding sites, RNA splice sites, polyadenylation sites (e.g., an SV40 large T Ag poly A addition site), and transcriptional terminator sequences. See, e.g., Ausubel et al., supra; Sambrook, et al., supra. Other cells useful for production of nucleic acids or proteins of the present disclosure are known and / or available, for instance, from the American Type Culture Collection Catalogue of Cell Lines and Hybridomas (www.atcc.org) or other known or commercial sources.Attorney Docket No.: 000218-0153-WO1

[0140] When eukaryotic host cells are employed, polyadenylation or transcription terminator sequences are typically incorporated into the vector. An example of a terminator sequence is the polyadenylation sequence from the bovine growth hormone gene. In some embodiments, the polyA sequence is an SV40 polyA sequence.

[0141] Sequences for accurate splicing of the transcript can also be included. An example of a splicing sequence is the VP1 intron from SV40 (Sprague, et al., J. Virol.45:773-781 (1983)). Additionally, gene sequences to control replication in the host cell can be incorporated into the vector, as known in the art.

[0142] The plasmid constructs described herein may be used to deliver nucleic acids encoding the transposase domains or fusion proteins described herein to a cell.

[0143] The transposase domains and fusion proteins described herein may also be delivered to a cell using mRNA constructs. Thus, in one embodiment, provided herein is an mRNA sequence encoding a transposase domain or a fusion protein described herein. Such mRNA sequences may be delivered to a cell using a nanoparticle, for example, a lipid nanoparticle. Examples of lipid nanoparticles are described in, e.g., International Patent Applications No. PCT / US2021 / 055876, No. PCT / US2022 / 017570, U.S. Provisional Application No.63 / 397,268, U.S. Provisional Application No.63 / 301,855 and U.S. Provisional Application No.63 / 348,614, each of which is incorporated herein by reference in its entirety for examples of lipid nanoparticles that may be used to deliver mRNA constructs encoding the fusion proteins or transposase domains described herein. An mRNA construct may also be delivered to a cell by electroporation or nucleofection. The mRNA may be capped or otherwise modified. Cells and Modified Cells

[0144] The fusion proteins described herein are useful for site-specific genome modification of cells. Genome modification can comprise introducing a nucleic acid sequence, transgene and / or a genomic editing construct into a cell ex vivo, in vivo, in vitro or in situ to stably integrate a nucleic acid sequence, transiently integrate a nucleic acid sequence, produce site-specific integration of a nucleic acid sequence, or produce a biased integration of a nucleic acid sequence. The nucleic acid sequence can be a transgene.

[0145] The site-specific integration can occur at a rDNA repeat

[0146] The site-specific transgene integration can occur at a site that results in enhanced expression of a target gene. Enhancement of target gene expression can occur by site-specificAttorney Docket No.: 000218-0153-WO1 integration at introns, exons, promoters, genetic elements, enhancers, suppressors, start codons, stop codons, and response elements.

[0147] For fusion proteins comprising TAL Arrays, the target site is determined by the sequence of one or more TAL binding domain. A person of skill in the art using the disclosure herein and elsewhere will be able to modify the TALEN sequences to achieve the desired target specificity.

[0148] The genome modification can be a non-stable chromosomal integration of a transgene. The integrated transgene can become silenced, removed, excised, or further modified.

[0149] In some embodiments, the fusion proteins provided herein have better transposase fidelity than their wildtype equivalents (e.g., compared to piggyBac or SPB transposases). Transposase fidelity is a measure of the accuracy of the genetic modification induced by a transposase. A transposase with higher transposase fidelity may, for example, have less of- target integration activity compared to a transposase with lower fidelity. Alternatively or additionally, a transposase with higher transposase fidelity may have higher on-target activity compared to a transposase with lower fidelity. On-target and off-target activity may be measured by any suitable assay known in the art or described herein, for example, a Split GFP assay

[0150] In some embodiments, the fusion proteins provided herein have improved transposase activity compared to wildtype transposes (e.g., compared to piggyBac or SPB transposases). Transposase activity may be measured by any suitable assay known in the art or described herein, for example, a Split GFP assay. For example, the fusion proteins provided herein may enhanced on-target genome integration activity to their wildtype counterparts, but have decreased off-target genome integration activity compared to their wildtype counterparts. The rDNA targeting fusion proteins have a transposition integration ratio of at least 89%, at least 90:1, at least 91:1, at least 92:1, at least 93:1, at least 94:1, at least 95, at least 96:1, 97:1, at least 98:1 or at least 99:1.

[0151] In some embodiments, a rDNA TAL Array fusion proteins comprising an N- terminal deleted, thermostable, hyperactive, integration deficient PBx Plus transposase domain provided herein has a ratio of on-target to off-target activity of at least 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450- fold, at least about 500-fold, at least about 550-fold, at least about 600-fold, at least about 650-fold, at least about 700-fold, at least about 750-fold, at least about 800-fold, at leastAttorney Docket No.: 000218-0153-WO1 about 850-fold, at least about 900-fold, at least about 950-fold, or at least about 1000-fold compared to the wildtype transposase domain (e.g., a piggyBac transposase domain).

[0152] Also provided herein are methods site-specific gene integration. The transposon domains and fusion proteins provided herein may be used to deliver a transgene to a cell and integrate the transgene into a target site. In some embodiments, the target site is a repetitive element, such as a rDNA repeat or LINE1 repetitive element. There may be one, two or more target sites within one repetitive element.

[0153] The site-specific integration may be used in vitro or in vivo. An example of an in vivo application is gene therapy, which involves the delivery of a transgene to the genomic DNA of a cell. Formulations, Dosages and Modes of Administration

[0154] The present disclosure provides formulations, dosages and methods for administration of the compositions or modified cells described herein.

[0155] In one aspect, provided herein is a pharmaceutical composition comprising a nucleic acid molecule comprising a sequence encoding a fusion protein described herein and a pharmaceutically acceptable carrier. In one aspect, the nucleic acid molecule comprising a sequence encoding a fusion protein is comprised in a lipid nanoparticle. In one aspect, the nucleic acid molecule comprising a sequence encoding a fusion protein and a DNA molecule comprising a transposon are comprised in a lipid nanoparticle. The lipid nanoparticle may be targeted to a specific cell type, for example, a liver cell.

[0156] LNPs that may be used to encapsulate the compositions disclosed herein may comprise a structural lipid, a phospholipid, and / or a PEGylated lipid. LNPs that may be used to encapsulate the compositions disclosed herein may further comprise a targeting ligand. LNPs that may be used to encapsulate the compositions disclosed herein may further comprise a polyphenol.

[0157] In some aspects, a structural lipid can be a steroid. In some aspects, a structural lipid can be a sterol. In some aspects, a structural lipid can comprise cholesterol. In some aspects, a structural lipid can comprise ergosterol. In some aspects, a structural lipid can be a phytosterol.

[0158] The phospholipid may be any amphiphilic molecule that comprises a polar (hydrophilic) headgroup comprising phosphate and two hydrophobic fatty acid chains. In some aspects, a phospholipid can comprise 1,2-Dioleoyl-sn-glycero-3-phosphocholine (DOPC).Attorney Docket No.: 000218-0153-WO1

[0159] The PEGylated lipid may be any lipid that is modified (e.g. covalently linked to) at least one polyethylene glycol molecule. In some aspects, a PEGylated lipid can comprise 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000, hereafter referred to as DMG-PEG2000 or PEG-DMG.

[0160] In some embodiments, the targeting ligand comprises GalNac. In some embodiments, the targeting ligand comprising GalNac is GalNac-PEG.

[0161] In some embodiments, a polyphenol is added to the LNP used to encapsulate the compositions disclosed herein. In some embodiments the polyphenol is pentagalloylglucose.

[0162] In some embodiments, the LNP that may be used to encapsulate the compositions disclosed herein comprises the compound HBC365. The structure of HBC365 is

[0163] Compositions and methods for preparing HBC365-comprising LNPs are disclosed in International Patent Application No: PCT / US2024 / 012245 (published as WO 2024 / 155938, referring to HBC365 as “Compound No.37”), which is incorporated herein by reference for examples of LNPs that may be used to deliver the compositions described herein.

[0164] In some aspects, a lipid nanoparticle can comprise lipid and nucleic acid at a specified ratio (weight / weight). Suitable ratios of nucleic acid to lipid are described in PCT / US2024 / 012245 (published as WO 2024 / 155938).

[0165] In some embodiments, the LNP comprises 50% HBC365 lipid; 38.5% Cholesterol; 10% DOPC and 1.5% DMG-PEG2000; Lipid: Nucleic acid ratio 50:1 (H2-8 formulation).

[0166] In some embodiments, the LNP comprises 50% HBC365 lipid; 41% Cholesterol; 7.5% DOPC and 1% DMG-PEG2000; 0.5% GalNac-PEG; 0.15% pentagalloylglucose (Lipid:Attorney Docket No.: 000218-0153-WO1 Nucleic acid ratio 60:1) (H3-5 formulation). Compositions and methods for preparing polyphenol-comprising LNPs are disclosed in PCT / US2024 / 012245 (published as WO 2024 / 155938) and US Provisional Application No: 63 / 678,021.

[0167] Without wishing to be bound by theory, it is believed that the compositions provided herein can be used to modify cells in vivo. In vivo modifications may include, for example, the delivery of a transgene to a cell. Thus, in some embodiments, provided herein is a method of integrating a transgene into a rDNA genomic target site of a cell, the method comprising introducing into the cell a rDNA TAL Array SPB fusion protein provided herein and a transposon, wherein the transposon comprises, in 5’ to 3’ order: a 5’ITR, the transgene, and a 3’ ITR. The improved fidelity of the fusion proteins described herein may allow for effective and safe cell modification in vivo. For example, in some embodiments, the fusion protein integrates the transposon with a fidelity of on target to off target transposition integration ratio of at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.

[0168] Any suitable transposon may be used or modified to deliver the transgene by a method described herein. In some embodiments, is a small transposon comprising symmetrical ITRs. In some embodiments, the transposon further comprises an exogenous promoter between the 5’ ITR and the transgene. In some embodiments, the transposon does not comprise symmetrical ITRS.

[0169] Any suitable transgene may be delivered to a cell using the methods described herein. In some embodiments, the transgene encodes a selectable marker or a reporter gene. In some embodiments, the selectable marker is puromycin and the reporter gene is GFP. In some embodiments, the transgene is a therapeutic transgene, for example a transgene encoding a protein the expression of which is useful for treating a metabolic disorder, hemophilia, cancer or PKU.

[0170] In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes a methylmalonyl-CoA mutase (MUT1) polypeptide. In some aspects, a transgene sequence comprises a nucleic acid sequence that encodes a MUT1 polypeptide, wherein the MUT1 polypeptide comprises, consists essentially of or consist of an amino acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to the sequence set forth in EQ ID NO: 209, 210, 232 or 233. In some aspects, a transgene sequence comprises a nucleic acid sequence that encodes a MUT1 polypeptide, wherein the MUT1 polypeptide comprises, consists essentially of or consist ofAttorney Docket No.: 000218-0153-WO1 the amino acid sequence set forth in EQ ID NO: 209, 210, 232 or 233. In some aspects, a nucleic acid sequence that encodes a MUT1 polypeptide comprises, consists essentially of, or consists of a nucleic acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to any of the sequences set forth in any one of SEQ ID NOs: 211-222. In some aspects, a nucleic acid sequence that encodes a MUT1 polypeptide comprises, consists essentially of, or consists of the nucleic acid sequences set forth in any one of SEQ ID NOs: 211-222.

[0171] In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes an ornithine transcarbamylase (OTC) polypeptide. In some aspects, a transgene sequence comprises a nucleic acid sequence that encodes an OTC polypeptide, wherein the OTC polypeptide comprises, consists essentially of or consist of an amino acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 234, 235, 238 or 239. In some aspects, a transgene sequence comprises a nucleic acid sequence that encodes an OTC polypeptide, wherein the OTC polypeptide comprises, consists essentially of or consist of the amino acid sequence set forth in SEQ ID NO: 234, 235, 238 or 239. In some aspects, a nucleic acid sequence that encodes an OTC polypeptide comprises, consists essentially of, or consists of a nucleic acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to any of the sequences set forth in SEQ ID NOs: 236, 237, 240, or 241. In some aspects, a nucleic acid sequence that encodes an OTC polypeptide comprises, consists essentially of, or consists of the nucleic acid sequence set forth in SEQ ID NOs: 236, 237, 240, and 241.

[0172] In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes an iCAS9 polypeptide. In some aspects, a transgene sequence comprises a nucleic acid sequence that encodes an iCAS9 polypeptide, wherein the iCAS9 polypeptide comprises, consists essentially of or consist of an amino acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 223 or 225. In some aspects, a transgene sequence comprises a nucleic acid sequence that encodes an iCAS9 polypeptide, wherein the iCAS9 polypeptide comprises, consists essentially of or consist of the amino acid set forth in SEQ ID NO: 223 or 225. In some aspects, a nucleic acid sequence that encodes an iCAS9 polypeptide comprises, consists essentially of, or consists of a nucleic acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to the sequence set forth in SEQ ID NOs: 224 or 226. In some aspects, a nucleic acid sequence thatAttorney Docket No.: 000218-0153-WO1 encodes an iCAS9 polypeptide comprises, consists essentially of, or consists of the nucleic acid sequence set forth in SEQ ID NOs: 224 or 226.

[0173] In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes for a human phenylalanine hydroxylase (hPAH) polypeptide. In some aspects, a nucleic acid sequence that encodes a hPAH polypeptide comprises, consists essentially of, or consists of a nucleic acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 227. In some aspects, a nucleic acid sequence that encodes a hPAH polypeptide comprises, consists essentially of, or consists of the nucleic acid sequence set forth in SEQ ID NO: 227.

[0174] In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes a Factor VIII (FVIII) polypeptide. In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes a FVIII polypeptide, wherein the FVIII polypeptide comprises, consists essentially of or consist of an amino acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 230. In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes a FVIII polypeptide, wherein the FVIII polypeptide comprises, consists essentially of or consist of the amino acid sequence set forth in SEQ ID NO: 230. In some aspects, a nucleic acid sequence that encodes a FVIII polypeptide comprises, consists essentially of, or consists of a nucleic acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 228. In some aspects, a nucleic acid sequence that encodes a FVIII polypeptide comprises, consists essentially of, or consists of the nucleic acid sequence set forth in SEQ ID NO: 228.

[0175] In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes a Factor IX (FIX) polypeptide. In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes a FIX polypeptide, wherein the FIX polypeptide comprises, consists essentially of or consist of an amino acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO: 231. In some aspects, a transgene sequence can comprise a nucleic acid sequence that encodes a FIX polypeptide, wherein the FIX polypeptide comprises, consists essentially of or consist of the amino acid sequence set forth in SEQ ID NO: 231. In some aspects, a nucleic acid sequence that encodes a FIX polypeptide comprises, consists essentially of, or consists of a nucleic acid sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% (or any percentage in between) identical to SEQ ID NO:Attorney Docket No.: 000218-0153-WO1 229. In some aspects, a nucleic acid sequence that encodes a FIX polypeptide comprises, consists essentially of, or consists of the nucleic acid sequence set forth in SEQ ID NO: 229.

[0176] The disclosed compositions and pharmaceutical compositions can comprise at least one of any suitable auxiliary, such as, but not limited to, diluent, binder, stabilizer, buffers, salts, lipophilic solvents, preservative, adjuvant or the like. Pharmaceutically acceptable auxiliaries are preferred. Non-limiting examples of, and methods of preparing such sterile solutions are well known in the art, such as, but limited to, Gennaro, Ed., Remington's Pharmaceutical Sciences, 18th Edition, Mack Publishing Co. (Easton, Pa.) 1990 and in the “Physician's Desk Reference”, 52nd ed., Medical Economics (Montvale, N.J.) 1998. Pharmaceutically acceptable carriers can be routinely selected that are suitable for the mode of administration, solubility and / or stability of the protein scaffold, fragment or variant composition as well known in the art or as described herein.

[0177] Non-limiting examples of pharmaceutical excipients and additives suitable for use include proteins, peptides, amino acids, lipids, and carbohydrates (e.g., sugars, including monosaccharides, di-, tri-, tetra-, and oligosaccharides; derivatized sugars, such as alditols, aldonic acids, esterified sugars and the like; and polysaccharides or sugar polymers), which can be present singly or in combination, comprising alone or in combination 1-99.99% by weight or volume. Non-limiting examples of protein excipients include serum albumin, such as human serum albumin (HSA), recombinant human albumin (rHA), gelatin, casein, and the like. Representative amino acid / protein components, which can also function in a buffering capacity, include alanine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, aspartame, and the like. One preferred amino acid is glycine.

[0178] Non-limiting examples of carbohydrate excipients suitable for use include monosaccharides, such as fructose, maltose, galactose, glucose, D-mannose, sorbose, and the like; disaccharides, such as lactose, sucrose, trehalose, cellobiose, and the like; polysaccharides, such as raffinose, melezitose, maltodextrins, dextrans, starches, and the like; and alditols, such as mannitol, xylitol, maltitol, lactitol, xylitol sorbitol (glucitol), myoinositol and the like. Preferably, the carbohydrate excipients are mannitol, trehalose, and / or raffinose.

[0179] The compositions can also include a buffer or a pH-adjusting agent; typically, the buffer is a salt prepared from an organic acid or base. Representative buffers include organic acid salts, such as salts of citric acid, ascorbic acid, gluconic acid, carbonic acid, tartaric acid, succinic acid, acetic acid, or phthalic acid; Tris, tromethamine hydrochloride, or phosphate buffers. Preferred buffers are organic acid salts, such as citrate.Attorney Docket No.: 000218-0153-WO1

[0180] Additionally, the disclosed compositions can include polymeric excipients / additives, such as polyvinylpyrrolidones, ficolls (a polymeric sugar), dextrates (e.g., cyclodextrins, such as 2-hydroxypropyl-β-cyclodextrin), polyethylene glycols, flavoring agents, antimicrobial agents, sweeteners, antioxidants, antistatic agents, surfactants (e.g., polysorbates, such as “TWEEN 20” and “TWEEN 80”), lipids (e.g., phospholipids, fatty acids), steroids (e.g., cholesterol), and chelating agents (e.g., EDTA).

[0181] Many known and developed modes can be used for administering therapeutically effective amounts of the compositions or pharmaceutical compositions disclosed herein. Non- limiting examples of modes of administration include bolus, buccal, infusion, intrarticular, intrabronchial, intraabdominal, intracapsular, intracartilaginous, intracavitary, intracelial, intracerebellar, intracerebroventricular, intracolic, intracervical, intragastric, intrahepatic, intralesional, intramuscular, intramyocardial, intranasal, intraocular, intraosseous, intraosteal, intrapelvic, intrapericardiac, intraperitoneal, intrapleural, intraprostatic, intrapulmonary, intrarectal, intrarenal, intraretinal, intraspinal, intrasynovial, intrathoracic, intrauterine, intratumoral, intravenous, intravesical, oral, parenteral, rectal, sublingual, subcutaneous, transdermal or vaginal means. In preferred embodiments, a composition comprising a modified cell described herein is administered intravenously, e.g., by intravenous infusion.

[0182] A composition of the disclosure can be prepared for use for parenteral (subcutaneous, intramuscular or intravenous) or any other administration particularly in the form of liquid solutions or suspensions. For parenteral administration, a composition disclosed herein can be formulated as a solution, suspension, emulsion, particle, powder, or lyophilized powder in association, or separately provided, with a pharmaceutically acceptable parenteral vehicle. Formulations for parenteral administration can contain as common excipients sterile water or saline, polyalkylene glycols, such as polyethylene glycol, oils of vegetable origin, hydrogenated naphthalenes and the like. Aqueous or oily suspensions for injection can be prepared by using an appropriate emulsifier or humidifier and a suspending agent, according to known methods. Agents for injection or infusion can be a non-toxic, non- orally administrable diluting agent, such as aqueous solution, a sterile injectable solution or suspension in a solvent. As the usable vehicle or solvent, water, Ringer's solution, isotonic saline, etc. are allowed; as an ordinary solvent or suspending solvent, sterile involatile oil can be used. For these purposes, any kind of involatile oil and fatty acid can be used, including natural or synthetic or semisynthetic fatty oils or fatty acids; natural or synthetic or semisynthtetic mono- or di- or tri-glycerides. Parental administration is known in the art and includes, but is not limited to, conventional means of injections, a gas pressured needle-lessAttorney Docket No.: 000218-0153-WO1 injection device as described in U.S. Pat. No.5,851,198, and a laser perforator device as described in U.S. Pat. No.5,839,446.

[0183] It can be desirable to deliver the disclosed compounds to the subject over prolonged periods of time, for example, for periods of one week to one year from a single administration. Various slow release, depot or implant dosage forms can be utilized. For example, a dosage form can contain a pharmaceutically acceptable non-toxic salt of the compounds that has a low degree of solubility in body fluids, for example, (a) an acid addition salt with a polybasic acid, such as phosphoric acid, sulfuric acid, citric acid, tartaric acid, tannic acid, pamoic acid, alginic acid, polyglutamic acid, naphthalene mono- or di- sulfonic acids, polygalacturonic acid, and the like; (b) a salt with a polyvalent metal cation, such as zinc, calcium, bismuth, barium, magnesium, aluminum, copper, cobalt, nickel, cadmium and the like, or with an organic cation formed from e.g., N,N′-dibenzyl- ethylenediamine or ethylenediamine; or (c) combinations of (a) and (b), e.g., a zinc tannate salt. Additionally, the disclosed compounds or, preferably, a relatively insoluble salt, such as those just described, can be formulated in a gel, for example, an aluminum monostearate gel with, e.g., sesame oil, suitable for injection. Particularly preferred salts are zinc salts, zinc tannate salts, pamoate salts, and the like. Another type of slow release depot formulation for injection would contain the compound or salt dispersed for encapsulation in a slow degrading, non-toxic, non-antigenic polymer, such as a polylactic acid / polyglycolic acid polymer for example as described in U.S. Pat. No.3,773,919. The compounds or, preferably, relatively insoluble salts, such as those described above, can also be formulated in cholesterol matrix silastic pellets, particularly for use in animals. Additional slow release, depot or implant formulations, e.g., gas or liquid liposomes, are known in the literature (U.S. Pat. No. 5,770,222 and “Sustained and Controlled Release Drug Delivery Systems”, J. R. Robinson ed., Marcel Dekker, Inc., N.Y., 1978). Methods of Treatment

[0184] In another aspect, provided herein are methods of treating a disease or disorder in a subject, the method comprising administering to the subject a composition comprising a DNA molecule encoding a transgene and a DNA molecule encoding one or more of the rDNA targeting TAL-Array – PBx Plus transposase comprising dual CRDs fusion proteins described herein. The terms “subject” and “patient” are used interchangeably herein. In preferred embodiments, the patient is human.Attorney Docket No.: 000218-0153-WO1

[0185] Thus, in some embodiments, provided herein is a method of treating a disease or disorder in a subject in need thereof, comprising administering to the subject a rDNA TAL array fusion protein provided herein and a transposon, wherein the transposon comprises, in 5’ to 3’ order: a 5’ITR, the transgene, and a 3’ ITR. In some embodiments, the disease is a metabolic disorder. In some embodiments, the disease is cancer. In some embodiments, the disease is phenylketonuria (PKU). In some embodiments, the disease is hemophilia.

[0186] In certain embodiments, the DNA molecule comprising the transgene is a nanoplasmid comprising a piggyBac transposon comprising the transgene coding sequence. In certain embodiments, the transgene is expressed by a heterologous promoter operably associated with the transgene sequence. In certain embodiments, the transposon is a small transposon comprising symmetrical inverted terminal repeats (ITRs). An illustrative example of a small transposon comprising symmetrical ITRs is the sequence set forth in SEQ ID NO: 202. In certain embodiments, the transposon may further comprise a nucleic acid sequence encoding a selectable marker and / or a reporter gene. An illustrative example of a small transposon comprising symmetrical ITRs and a puromycin selectable marker and a GFP reporter gene is shown in the sequence set forth in SEQ ID NO: 203.

[0187] In certain embodiments, the DNA molecule comprising the transgene is an AAV virion comprising a sequence encoding the transgene flanked by piggyBac ITR sequences for site-specific transposition of the encoded transgene (an “AAV piggyBac virion”). Illustrative examples of compositions and methods for preparing an administering AAV piggyBac vectors for delivering transgenes to treat diseases and disorders are disclosed in co-owned International PCT Patent Application Publication Nos: PCT / US2021 / 020929 (metabolic liver disorders, including urea cycle disorders); PCT / US2022 / 018976 (hemophilia A & B); and PCT / US2024 / 016637 (phenylketonuria; PKU).

[0188] In certain embodiments, the DNA molecule comprising the rDNA or LINE1 targeting TAL Array - PBx Plus transposase comprising dual CRDs coding sequence is a nanoplasmid.

[0189] In certain embodiments, the nucleic acid molecule comprising the rDNA or LINE1 targeting TAL Array - PBx Plus transposase comprising dual CRDs coding sequence is a mRNA.

[0190] Any suitable transgene may be delivered to a cell using the methods described herein. In some embodiments, the transgene is a therapeutic transgene, for example a transgene encoding a protein the expression of which is useful for treating a metabolic disorder, hemophilia, cancer or PKU.Attorney Docket No.: 000218-0153-WO1

[0191] The fusion proteins provided herein may be used to deliver a gene therapy. Gene therapy usually involves the delivery of a transgene to the genomic DNA of a cell, particularly a rDNA repeat sequence. Usually, the transgene replaces a gene that is mutated or otherwise not expressed properly in the cell. The fusion proteins may be used to deliver a therapeutic transgene to a cell and integrate the transgene into a target site. In some embodiments, a method of treatment comprises introducing into the cell the fusion protein or a nucleic acid sequence encoding a fusion protein and a transposon, wherein the transposon comprises, in 5’ to 3’ order: a 5’ITR, the transgene, and a 3’ ITR. In certain embodiments, the transposon is a small transposon comprising symmetrical 5’ and 3’ ITRs (SEQ ID NO: 202).

[0192] In certain embodiments, the disease or disorder treated in accordance with the methods described herein is a metabolic liver disorder, hemophilia A, hemophilia B, or PKU by specifically integrating the corrective transgene with high fidelity in a rDNA repeat or into a LINE1 repetitive element to restore function of the dysfunctional gene sequence.

[0193] Metabolic liver disorders can include, but are not limited to, urea cycle disorders, N-Acetylglutamate Synthetase (NAGS) Deficiency, Carbamoylphosphate Synthetase I Deficiency (CPSI Deficiency), Ornithine Transcarbamylase (OTC) Deficiency, Argininosuccinate Synthetase Deficiency (ASSD) (Citrullinemia I), Citrin Deficiency (Citrullinemia II), Argininosuccinate Lyase Deficiency (Argininosuccinic Aciduria), Arginase Deficiency (Hyperargininemia), Ornithine Translocase Deficiency (HHH Syndrome) methylmalonic acidemia (MMA), progressive familial intrahepatic cholestasis type 1 (PFIC1), progressive familial intrahepatic cholestasis type 1 (PFIC2), and progressive familial intrahepatic cholestasis type 1 (PFIC3).

[0194] Other diseases and disorders may be treated by administering Factor XIII (HemA), Factor IX (HemB) or phenylalanine hydroxylase (PAH) transgenes.

[0195] In some embodiments, the disease or disorder treated in accordance with the methods described herein is a cancer. In some embodiments, a method of treatment described herein may delay cancer progression and / or reduce tumor burden. In certain embodiments, the integrated transgene encodes a chimeric antigen receptor (CAR) for in vivo CAR-T therapy.

[0196] In certain embodiments, the DNA molecule comprising the transgene, e.g., a nanoplasmid or AAV virion, and the nanoplasmid or mRNA encoding the TAL Array - PBx Plus transposase comprising dual CRDs are administered to the subject sequential or co- administered via any method known to those of skill in the art.Attorney Docket No.: 000218-0153-WO1

[0197] The methods and compositions described herein a capable of delivering and integrating transgene sequences by site specific transposition with remarkable fidelity of on target transposition integrations at rDNA repeats versus off target transposition integrations elsewhere in the genome of the target cell. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 89:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 90:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 91:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 92:1 In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 93:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 94:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 95:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 96:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 97:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 98:1. In certain embodiments, the transposition integration ratio of on target to off target integrations is at least 99:1.

[0198] The dosage of a pharmaceutical composition to be administered to a subject can vary depending upon known factors, such as the pharmacodynamic characteristics of the particular agent, and its mode and route of administration; age, health, and weight of the recipient; nature and extent of symptoms, kind of concurrent treatment, frequency of treatment, and the effect desired.

[0199] A more detailed description of pharmaceutically acceptable excipients, formulations, dosages and methods of administration of the disclosed compositions and pharmaceutical compositions is disclosed in PCT Publication No. WO 2019 / 049816. Kits

[0200] In another aspect, provided herein is a kit comprising a DNA molecule or mRNA encoding a rDNA targeting or LINE1 targeting TAL-Array – PBx Plus fusion protein provided herein and a DNA molecule comprising a desired transgene sequence for site- specific transposition into a target cell in vitro or in vivo, and instructions for using the kit.Attorney Docket No.: 000218-0153-WO1 Definitions

[0201] As used throughout the disclosure, the singular forms “a,” “and,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a method” includes a plurality of such methods and reference to “a dose” includes reference to one or more doses and equivalents thereof known to those skilled in the art, and so forth.

[0202] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” can mean within 1 or more standard deviations. Alternatively, “about” can mean a range of up to 20%, or up to 10%, or up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0203] The term "comprising" is intended to mean that the compositions and methods include the recited elements, but do not exclude others. "Consisting essentially of” when used to define compositions and methods, shall mean excluding other elements of any essential significance to the combination when used for the intended purpose. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants or inert carriers. "Consisting of shall mean excluding more than trace elements of other ingredients and substantial method steps. Aspects defined by each of these transition terms are within the scope of this disclosure.

[0204] As used herein, "expression" refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently being translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.

[0205] A conservative substitution of an amino acid, i.e., replacing an amino acid with a different amino acid of similar properties (e.g., hydrophilicity, degree and distribution of charged regions) is recognized in the art as typically involving a minor change. These minor changes can be identified, in part, by considering the hydropathic index of amino acids, as understood in the art. Kyte et al., J. Mol. Biol.157: 105-132 (1982). The hydropathic index of an amino acid is based on a consideration of its hydrophobicity and charge. Amino acids ofAttorney Docket No.: 000218-0153-WO1 similar hydropathic indexes can be substituted and still retain protein function. In an aspect, amino acids having hydropathic indexes of ±2 are substituted. The hydrophilicity of amino acids can also be used to reveal substitutions that would result in proteins retaining biological function. A consideration of the hydrophilicity of amino acids in the context of a peptide permits calculation of the greatest local average hydrophilicity of that peptide, a useful measure that has been reported to correlate well with antigenicity and immunogenicity. U.S. Patent No.4,554,101, incorporated fully herein by reference.

[0206] Substitution of amino acids having similar hydrophilicity values can result in peptides retaining biological activity, for example immunogenicity. Substitutions can be performed with amino acids having hydrophilicity values within ±2 of each other. Both the hyrophobicity index and the hydrophilicity value of amino acids are influenced by the particular side chain of that amino acid. Consistent with that observation, amino acid substitutions that are compatible with biological function are understood to depend on the relative similarity of the amino acids, and particularly the side chains of those amino acids, as revealed by the hydrophobicity, hydrophilicity, charge, size, and other properties.

[0207] As used herein, “conservative” amino acid substitutions may be defined as set out in Table 17, Table 18, and Table 19 below. In some aspects, fusion polypeptides and / or nucleic acids encoding such fusion polypeptides include conservative substitutions have been introduced by modification of polynucleotides encoding polypeptides of the disclosure. Amino acids can be classified according to physical properties and contribution to secondary and tertiary protein structure. A conservative substitution is a substitution of one amino acid for another amino acid that has similar properties. Exemplary conservative substitutions are set out in Table 17. Table 17: Conservative Substitutions I

[0208] Alternately, conservative amino acids can be grouped as described in Lehninger, (Biochemistry, Second Edition; Worth Publishers, Inc. NY, N.Y. (1975), pp.71-77) as set forth in Table 18.Attorney Docket No.: 000218-0153-WO1 Table 18: Conservative Substitutions II

[0209] Alternately, exemplary conservative substitutions are set out in Table 19. Table 19: Conservative Substitutions III

[0210] Polypeptides and proteins of the disclosure, either their entire sequence, or any portion thereof, may be non-naturally occurring. Polypeptides and proteins of the disclosure may contain one or more mutations, substitutions, deletions, or insertions that do not naturally-occur, rendering the entire amino acid sequence non-naturally occurring. Polypeptides and proteins of the disclosure may contain one or more duplicated, inverted or repeated sequences, the resultant sequence of which does not naturally-occur, rendering theAttorney Docket No.: 000218-0153-WO1 entire amino acid sequence non-naturally occurring. Polypeptides and proteins of the disclosure may contain modified, artificial, or synthetic amino acids that do not naturally- occur, rendering the entire amino acid sequence non-naturally occurring.

[0211] As used throughout the disclosure, identity between two sequences may be determined by using the stand-alone executable BLAST engine program for blasting two sequences (bl2seq), which can be retrieved from the National Center for Biotechnology Information (NCBI) ftp site, using the default parameters (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250; which is incorporated herein by reference in its entirety). The terms "identical" or "identity" when used in the context of two or more nucleic acids or polypeptide sequences, refer to a specified percentage of residues that are the same over a specified region of each of the sequences. In some embodiments, the sequence identify is determined over the entire length of a sequence. The percentage can be calculated by optimally aligning the two sequences, comparing the two sequences over the specified region, determining the number of positions at which the identical residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the specified region, and multiplying the result by 100 to yield the percentage of sequence identity. In cases where the two sequences are of different lengths or the alignment produces one or more staggered ends and the specified region of comparison includes only a single sequence, the residues of single sequence are included in the denominator but not the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be performed manually or by using a computer sequence algorithm such as BLAST or BLAST 2.0.

[0212] In certain embodiments, if a sequence has a certain sequence identity (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%) to a certain SEQ ID NO, the sequence and the sequence of the SEQ ID NO have the same length. In certain embodiments, if a sequence has a certain sequence identity (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%) to a certain SEQ ID NO, the sequence and the sequence of the SEQ ID NO only differ due to conservative amino acid substitutions.

[0213] As used throughout the disclosure, the term "endogenous" refers to nucleic acid or protein sequence naturally associated with a target gene or a host cell into which it is introduced.

[0214] As used throughout the disclosure, the term "exogenous" refers to nucleic acid or protein sequence not naturally associated with a target gene or a host cell into which it is introduced, including non-naturally occurring multiple copies of a naturally occurring nucleicAttorney Docket No.: 000218-0153-WO1 acid, e.g., DNA sequence, or naturally occurring nucleic acid sequence located in a non- naturally occurring genome location.

[0215] The disclosure provides methods of introducing a polynucleotide construct comprising a DNA sequence into a host cell. By "introducing" is intended presenting to the cell the polynucleotide construct in such a manner that the construct gains access to the interior of the host cell. The methods of the disclosure do not depend on a particular method for introducing a polynucleotide construct into a host cell, only that the polynucleotide construct gains access to the interior of one cell of the host. Methods for introducing polynucleotide constructs into bacteria, plants, fungi and animals are known in the art including, but not limited to, stable transformation methods, transient transformation methods, and virus-mediated methods. Exemplary Embodiments

[0216] Particular embodiments of the disclosure are set forth in the following numbered paragraphs:1. A fusion protein, comprising, in N-terminal to C-terminal direction:(a) a TAL Array targeting an rDNA repeat; and (b) a transposase of SEQ ID NO: 6 comprising an N-terminal deletion.2. The fusion protein of paragraph 1, wherein the transposase comprises a N-terminaldeletion comprising a deletion of amino acids 1-74, 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102 or 1-103 of SEQ ID NO: 6.3. The fusion protein of paragraph 1 or 2, further comprising a nucleolar traffickingsequence (NTS), wherein the C-terminus of the NTS is attached to the N-terminus of the TAL Array.4. The fusion protein of paragraph 3, wherein the NTS is an RNA polymerase I (PolI)subunit.5. The fusion protein of paragraph 4, wherein the PolI subunit is an A12.2 subunit, aA35.4 subunit, an A43 subunit or an A49 subunit.6. The fusion protein of paragraph 4 or 5, wherein the PolI subunit comprises the aminoacid sequence of any one of SEQ ID NO: 119-122.Attorney Docket No.: 000218-0153-WO17. The fusion protein of paragraph 3, wherein the NTS is an R2 retrotransposontransposase sequence.8. The fusion protein of paragraph 7, wherein the R2 retrotransposon transposasesequence comprises the amino acid sequence of SEQ ID NO: 123 or 124.9. The fusion protein of any one of paragraphs 1-8, further comprising a proteinstabilization domain (PSD).10. The fusion protein of paragraph 9, wherein the PSD comprises the amino acidsequence of SEQ ID NO: 190.11. The fusion protein of any one of paragraphs 1-10, further comprising a nuclearlocalization sequence (NLS).12. The fusion protein of paragraph 11, wherein the NLS sequence comprises the aminoacid sequence of SEQ ID NO: 191.13. The fusion protein of any of paragraphs 1-12, further comprising an GS or GGGGS(SEQ ID NO: 125) linker positioned (a) between the TAL Array and the SPB transposase; (b) between nucleolar trafficking sequence and the TAL Array; and / or (c) between the PSD and the TAL Array.14. The fusion protein of any one of paragraphs 1-13, wherein the SPB transposasecomprises the amino acid sequence of any of SEQ ID Nos: 8-28 or 208.15. The fusion protein of any one of paragraphs 1-14, wherein the TAL Array comprisesone or more TAL domains targeting an rDNA 45S repeat, an rDNA 18S repeat, an rDNA 5.8S repeat or an rDNA 28S repeat.16. The fusion protein of paragraph 15, wherein each of the one or more TAL domainstargets an rDNA repeat comprising the nucleic acid sequence of any one of SEQ ID NOs: 59 - 80.17. The fusion protein of paragraph 15 or 16, wherein the TAL Array comprises theamino acid sequence of any one of SEQ ID NO: 81-114.Attorney Docket No.: 000218-0153-WO118. The fusion protein of any one of paragraphs 1-17, wherein the fusion proteincomprises the amino acid sequence of any one of SEQ ID NOs: 126-159.19. A polynucleotide comprising a nucleic acid sequence encoding the fusion protein ofany one of paragraphs 1-18.20. A vector comprising the polynucleotide of paragraph 19.21. A pharmaceutical composition comprising (a) the fusion protein of any one ofparagraphs 1-18, the polynucleotide of paragraph 19, or the vector of paragraph 20 and (b) a pharmaceutically acceptable excipient.22. A method of modifying the genome of a cell, the method comprisingintroducing into the cell a. a polynucleotide comprising a nucleic acid encoding a fusion proteincomprising a transposase of SEQ ID NO: 6 comprising an N-terminal deletion; and a TAL Array targeting an rDNA repeat, wherein the C- terminus of the TAL Array is fused to the N-terminus of the SPB; and b. a DNA molecule comprising a transposon, wherein the transposoncomprises, in 5’ to 3’ order: a 5’ITR, the transgene, and a 3’ ITR, wherein the transgene is integrated into a TTAA site flanked by a left rDNA target site and a right rDNA target site.23. The method of paragraph 22, wherein the transposon is a small transposon comprisingsymmetrical ITRs.24. The method of paragraph 22 or 23, wherein the transposon comprises the sequence ofSEQ ID NO: 202.25. The method of any one of paragraphs 22-24, wherein the transposon furthercomprises an exogenous promoter operatively linked to the transgene.26. The method of any one of paragraphs 22-25, wherein the transgene encodes a proteinthe expression of which is beneficial for treating a metabolic disorder, hemophilia, cancer or PKU.Attorney Docket No.: 000218-0153-WO127. The method of any one of paragraphs 22-26, wherein the TTAA site comprises thenucleic acid sequence of any one of SEQ ID NOs 36-58.28. The method of any one of paragraphs 22-27, wherein the rDNA repeat comprises thenucleic acid sequence of any one of SEQ ID NOs: 59-80.29. The method of any one of paragraphs 22-28, wherein the left rDNA target sitecomprises the sequence of SEQ ID NO: 63 and the right rDNA target site comprises the sequence of SEQ ID NO: 64.30. The method of any one of paragraphs 22-28, wherein the left rDNA target sitecomprises the sequence of SEQ ID NO: 79 and the right rDNA target site comprises the sequence of SEQ ID NO: 80.31. The method of any one of paragraphs 22-30, wherein the cell is in vivo.32. The method of any one of paragraphs 22-31, wherein the transgene is integrated intothe genome of the cell with a transposon fidelity of on target to off target transposition integration ratio of at least 89%.33. The method of any one of paragraphs 22-32, wherein the cell is a liver cell.34. A method for generating an engineered cell by site-specific transposition, comprisingintroducing into a cell (a) a nucleic acid encoding a fusion protein comprising a rDNA TAL Array and a transposase domain comprising the sequence of SEQ ID NO: 6; and (b) a DNA molecule comprising a transposon; wherein the transposon is integrated by site-specific transposition into a TTAA sequence of an rDNA repeat.35. The method of paragraph 34, wherein the DNA molecule comprising a transposon is ananoplasmid or an AAV virion.36. The method of any one of paragraphs 34-35, wherein the nucleic acid encoding thefusion protein is mRNA introduced into the cell by a lipid nanoparticle.37. The method of any of paragraphs 34-36, wherein the nucleic acid encoding the fusionprotein and / or the DNA molecule comprising the transposon are introduced into the cell by a lipid nanoparticle.Attorney Docket No.: 000218-0153-WO1 EXAMPLES

[0217] The Examples in this section are provided for illustration and are not intended to limit the invention. Example 1: Design of TAL Array - PBx Plus Fusion Proteins for Site-specific Transposition into 45S rDNA Repeats

[0218] The general design of Xanthomonas TAL proteins that recognize given DNA target sequences was performed essentially described in co-owned patent application PCT / US2022 / 077549 (published as WO 2023 / 060089). TALs often bind to DNA sequences downstream of a 5’T. Furthermore, for optimal site-specific transposition, the TAL binding sites were properly positioned using a spacer, such as a 13bp spacer, separating the TAL binding site from the TTAA target sequence.

[0219] TAL pairs were designed to target 11 of the 23 TTAA sequences within the 45S rDNA repeats. The binding sites of these rDNA TALs, including the upstream 5’T, are set forth in SEQ ID NOs: 59-80. Additionally, as a non-rDNA control, a pair of TALs targeting the LINE1 repeat sequence were designed. The binding sites for the LINE1 TALs, including the upstream 5’T, are set forth in SEQ ID NOs: 115 & 116. TAL module arrays were designed to bind to each of the rDNA and LINE1 DNA target sites. The amino acid sequence of the rDNA TAL module arrays are set forth in SEQ ID NOs: 81-102. The amino acid sequence of the LINE1 TAL module arrays are set forth in SEQ ID NOs: 117 & 118.

[0220] A second set of TAL module arrays were designed with a N in place of the HD RVD for sites that contain a CpG. It’s been reported that methylation of CpG sequences within TAL binding sites, reduces binding of the TAL. However, it has also been reported that swapping the RVD region of the affected module, for example to a single asparagine (N), restores binding. CpG sequences are abundant in rDNA. The amino acid sequence of the CpG rDNA TAL module arrays are listed in SEQ ID NOs: 103-114.

[0221] The TAL module arrays contain constant N-terminal and C-terminal regions. For example, 136aa of the Xanthomonas N-terminal region (after deletion of the 152 most N- terminal residues) makes up the N-terminus of the TAL module array. The sequence of the delta152 TAL N-terminus is set forth as SEQ ID NO: 200. Also, a portion of the Xanthomonas C-terminal region, such as the 73 amino acid +73 region makes up the C- terminus of the TAL module array. The sequence of the +73 TAL C-terminus is set forth as SEQ ID NO: 201.Attorney Docket No.: 000218-0153-WO1

[0222] The designed TALs were next incorporated into a site-specific piggyBac transposase by fusing with the integration deficient PBx. In this example a PBx with a dual CRD, hyperactive mutants, and thermostability mutants was utilized. The sequence of this PBx transposase is set forth in SEQ ID NO: 6. The TALs were fused with PBx by inserting the entire TAL sequence, flanked on both sides by GGGGS linkers (SEQ ID NO: 125), into the PBx sequence between residues 85 and 86 of PBx. Additionally, a detection tag, such as the FLAG tag (SEQ ID NO: 192), and nuclear localization sequence, such as the SV40 NLS (SEQ ID 191), can be added to the N-terminus of the protein. The full amino acid sequence of the resulting rDNA TAL-PBxs with FLAG Tag and NLS are given as SEQ ID NOs: 126- 147. The full amino acid sequence of the resulting LINE1 TAL-ssSPBs with NLS are given as SEQ ID Nos: 160 & 161. The full amino acid sequence of the resulting CpG rDNA TAL- PBxs with FLAG Tag and NLS are set forth in SEQ ID NOs: 148-159.

[0223] For mammalian expression, each ssSPB was human codon optimized using Thermo’s GeneArt algorithm and cloned into the expression pcDNA3.1 expression vector downstream of the CMV promoter and Kozak sequence and upstream of the BGH polyadenylation signal sequence. Example 2: Methods for Site-specific Transposition into 45S rDNA Repeats in Cell Lines Using rDNA TAL Array – PBx Plus Fusion Proteins

[0224] The rDNA TAL Array - PBx Plus fusion proteins prepared in Example 1 were utilized for site-specific transposition into the 45S rDNA repeats in HEK293T and HepG2 cell lines. Briefly, each pair of expression vectors harboring the rDNA TAL Array - PBx constructs (SEQ ID NOs: 126-147) for rDNA target sites 3, 4, 5, 6, 9, 11, 15, 17,19, 22, and 23 were transfected into cells. For targets containing a CpG a second transfection was performed in which the CpG TAL-PBx (SEQ ID NOs: 148-159) was swapped in place of the non-CpG TAL-PBx. As a negative control, a vector expressing PBx without a fusion to TAL was used in place of TAL Array-PBx (SEQ ID NO: 6).

[0225] As a transposon donor, a plasmid harboring a small transposon with symmetrical ITRs was co-transfected with the TAL-PBx Plus left and right pairs. Described briefly, the transposon consists in 5’ to 3’ direction of a TTAA, the piggyBac 35bp 5’ ITR, 274bp of the piggyBac 5’ UTR, 55bp of stuffer, 175bp of the piggyBac 3’UTR, the reverse compliment of a second copy of the piggyBac 35bp 5’ ITR, and a second TTAA. The full sequence of the small transposon with symmetrical ITRs is given in SEQ ID NO: 202.Attorney Docket No.: 000218-0153-WO1

[0226] For testing in HEK293T cells, 140,000 cells were plated in collagen coated 24- well plates in 0.5 ml DMEM+10% FBS one day prior to transfection. The following day, the cells were transfected with 25 ng of each TAL-PBx Plus expression vector and 450 ng of the transposon donor vector using 1 µl of JetPrime transfection reagent.

[0227] For testing in HepG2 cells, 140,000 cells were plated in collagen coated 24-well plates in 0.5 ml of MEM+10%FBS one day prior to transfection. The following day, the cells were transfected with 25 ng of each TAL-PBx Plus expression vector and 450 ng of the transposon donor vector using 0.75 µl of Lipofectamine 3000 transfection reagent.

[0228] Two days after the transfection, genomic DNA was harvested from the cells. Site specific transposition of the transposon into the 45S rDNA repeats in the forward orientation was detected and quantified by ddPCR using one primer that binds within the transposon and a second primer that binds within the genome flanking the integration site. A second primer flanking the other side of the integration site was used to detect integrations of the transposon in the reverse orientation. The number of integrations per haploid genome for each TAL-PBx pair are given in Table 20 for HEK293T cells and in Table 21 for HepG2 cells. Table 20. Site-specific transposition into 45S rDNA repeats in HEK293T cells with TAL-PBxAttorney Docket No.: 000218-0153-WO1Table 21. Site-specific transposition into 45S rDNA repeats in HepG2 cells with rDNA TAL Array -PBx Fusion ProteinsAttorney Docket No.: 000218-0153-WO1

[0229] As shown in Tables 20 and Table 21, robust site-specific integration of the transposon was detected across all tested 45S rDNA targets in samples transfected with TAL- ssSPB but not in samples transfected with the PBx negative control. Example 3: Fidelity of Site-specific Transposition at 45S rDNA repeats in HepG2 Cells with TAL-ssSPB

[0230] The fidelity of site-specific integrations catalyzed by rDNA TAL Array - PBx Plus fusion proteins into the 45S rDNA repeats in HepG2 cells was assessed. Briefly, the pairs of expression vectors harboring the TAL Array - PBx Plus constructs targeting rDNA site 5 (SEQ ID NOs: 130 & 131) and rDNA site 23 (SEQ ID NOs: 146 & 147) were transfected into cells.

[0231] As a transposon donor, a nanoplasmid harboring a transposon with symmetrical ITRs with an EF1a promoter driving expression of a puromycin resistance gene and a GFP reporter was co-transfected with the rDNA TAL Array -PBx Plus pairs. The full sequence of the PuroR-2A-GPF transposon with symmetrical ITRs is set forth in SEQ ID NO: 203.

[0232] For testing in HepG2 cells, 140,000 cells were plated in collagen coated 24-well plates in 0.5 ml of MEM+10% FBS one day prior to transfection. The following day, the cells were transfected with 25 ng of each TAL Array-PBx expression vector and 450 ng of the transposon donor vector using 0.75 µl of Lipofectamine 3000 transfection reagent.

[0233] In a first experiment, HepG2 cells transfected with TAL-ssSPB pairs targeting rDNA sites 5 and 23 were harvested at days 6 and 9 post transfection. In a second experimentAttorney Docket No.: 000218-0153-WO1 HepG2 cells transfected with TAL Array-PBx pairs targeting rDNA sites 5 were harvested at days 7, 10, 17, and 20 post transfection. Additionally, cells grown in the presence of 3 ug / ml puromycin were harvested at days 17 post transfection.

[0234] Genomic DNA was extracted from the harvested cells and the sites of transposon integrations into the genome were determined using a linker mediated PCR method (LM- PCR). Briefly, genomic DNA was tagmented using Tn5 transposase loaded with oligos (linkers) containing adapter sequences and unique molecular identifiers (UMI). Nested PCR was performed using forward primers that bind within the transposon and reverse primers that bind within the linker added by tagmentation.

[0235] The adapter sequences on the linker enable next-generation sequencing (NGS) on an Illumina sequencing platform and the UMIs allow one to distinguish PCR duplications from unique integration events.

[0236] Integration events detected by LM-PCR were mapped to the reference human genome and classified as on-target (occurring at the desired rDNA TTAA site) or off-target (occurring at a non-rDNA TTAA site). The percentage of on-target integration events are shown in Table 22. Table 22: Fidelity of rDNA site-specific integrations in HepG2 cells Using rDNA – TAL Array - PBx Plus Fusion Proteins

[0237] As shown in Table 22, the percentage of on-target integrations at rDNA catalyzed by rDNA – TAL Array -PBx Plus fusion proteins is approximately 90% in HepG2 cells and changes little with time or with puromycin selection.Attorney Docket No.: 000218-0153-WO1 Example 4: In Vivo Site-specific Integration at 45S rDNA Repeats in Mouse Liver by Hydrodynamic Delivery of rDNA – TAL Array – PBx Plus Fusion Proteins

[0238] Site-specific integrations catalyzed by rDNA - TAL Array - PBx Plus into the 45S rDNA repeats was performed in vivo in mouse liver. Briefly, the pairs of expression vectors harboring the TAL Array - PBx Plus constructs targeting rDNA site 5 (SEQ ID NOs: 130 & 131) and rDNA site 23 (SEQ ID NOs: 146 & 147) were delivered to mouse liver by hydrodynamic delivery through the tail vein. As a non-rDNA control, LINE1 - TAL Array - PBx constructs targeting LINE1 repeat elements (SEQ ID NOs: 160 & 161) were delivered and as negative controls, a PBx expression vector lacking a TAL fusion (SEQ ID NO: 6) or PBS (buffer only) were delivered.

[0239] As a transposon donor, a nanoplasmid harboring a small transposon with symmetrical ITRs was co-delivered. The nanoplasmid was designed such that the small transposon disrupts the open reading frame of a CAGG-firefly luciferase reporter. Upon transposon mediated excision of the transposon and seamless repair of the donor, luciferase expression is restored. Site-specific integration of the transposon into the genome can be detected and quantified by ddPCR as described above.

[0240] Approximately 12.5 µg of each TAL Array - PBx pair and 12.5 µg of the transposon donor were diluted in 2 ml of Mirus Transit-EE buffer and injected into the tail vein of 10–12-week-old Balb / C mice. Luciferase expression, marking transposon excision, was quantified three days post-delivery. Following luciferase measurements, the mice were sacrificed, and genomic DNA was extracted from the liver. Site-specific integration into 45S rDNA or into LINE1 was quantified by ddPCR as described above. The excision activity, measured by luciferase signal in photons / second, and the site-specific integration, measured by ddPCR as the sum of forward and reverse integrations, are given in Table 23. Table 23. In vivo site-specific transposition in mouse liver by hydrodynamic delivery of rDNA or LINE1 – TAL Array - PBx Fusion ProteinsAttorney Docket No.: 000218-0153-WO1

[0241] As shown in Table 23, rDNA – TAL Array – PBx, and LINE1 – TAL Array - PBx fusion proteins as well as PBx transposase alone catalyzed robust excision of the integrated transposon. Site-specific integration into LINE1 and rDNA targets, however, was detected only in the presence of LINE1 – TAL Array -PBx Plus or rDNA – TAL Array – PBx Plus fusion proteins but not for PBx transposase domain lacking a TAL Array targeting LINE1 elements or rDNA. Example 5: In Vivo site-specific integration at 45S rDNA repeats in mouse liver by hybrid delivery of TAL Array – PBx Plus Fusion Proteins

[0242] Site-specific integrations catalyzed by rDNA – TAL Array – PBx Plus into the 45S rDNA repeats were performed in vivo in mouse liver using hybrid delivery. Unlike hydrodynamic delivery where the transposase and transposon are delivered in plasmid format, with hybrid delivery the transposase is delivered by lipid nanoparticle (LNP) in mRNA format and the transposon is delivered in adeno-associated virus (AAV) format.

[0243] Briefly, the rDNA – TAL Array – PBx constructs targeting rDNA site 5 (SEQ ID NOs: 130 & 131) were cloned downstream into an in vitro transcription vector harboring a T7 promoter, stabilizing UTR sequences, and a polyA tail. mRNA was produced with a Cleancap M6 and 5-methylcytosine base modifications. mRNA was encapsulated into lipid nanoparticles (LNPs) using an H2-8 formulation (50% HBC365 lipid; 38.5% Cholesterol; 10% DOPC and 1.5% DMG-PEG2000; Lipid: Nucleic acid ratio 50:1). Compositions and methods for preparing HBC365-comprising LNPs are disclosed in co-owned International Patent Application No: PCT / US2024 / 012245. As negative controls, a PBx transposase mRNA (SEQ ID NO: 6) lacking a TAL Array fusion was formulated or PBS (buffer only) was delivered.Attorney Docket No.: 000218-0153-WO1

[0244] As a transposon donor, a disrupted CAGG-firefly luciferase reporter described above was cloned between AAV ITR sequences. From this, AAV particles of AAV9 serotype were produced by VectorBuilder. The titer was determined by ddPCR as genome copies (GC) per µl. Upon transposon mediated excision of the transposon and seamless repair of the donor, luciferase expression is restored. Site-specific integration of the transposon into the genome can be detected and quantified by ddPCR as described above.

[0245] Three-week-old juvenile C57BL / 6 mice were injected with 3 mg / kg of LNP- mRNA transposase and 3E13 GC / kg of AAV transposon donor. Luciferase expression, marking transposon excision, was quantified four days post-delivery. Following luciferase measurements, the mice were sacrificed on days 4, 7, 14, 21, and 32 post injections and genomic DNA was extracted from the liver. Site-specific integration into 45S rDNA was quantified by ddPCR as described above. The excision activity, measured by luciferase signal in photons / second, and the site-specific integration, measured by ddPCR as the sum of forward and reverse integrations, are provided in Table 24. Table 24. In vivo site-specific transposition in mouse liver by hybrid delivery of rDNA – TAL Array -PBx Fusion ProteinsAttorney Docket No.: 000218-0153-WO1

[0246] As shown in Table 24, rDNA – TAL Array – PBx fusion proteins and PBx transposase catalyzed robust excision of the transposon. Site-specific integration into rDNA targets, however, was detected only in the presence of Site 5 rDNA – TAL Array – PBx fusion protein but not with PBx transposase alone, and was durable to at least 32 days post injection. Example 6: Fidelity of In Vivo site-specific integration at 45S rDNA repeats in mouse liver by hybrid delivery of rDNA - TAL Array – PBx fusion proteins

[0247] The fidelity of site-specific integrations at 45S rDNA catalyzed by Site 5 rDNA – TAL Array -PBx Plus following hybrid delivery was determined. Mouse liver samples from Example 5 were used to map the on-target and off-target integration events. Specifically, the rDNA site 5 samples harvested at Days 7 and 32 post injection were used. As negative controls, the PBS and PBx transposase only groups were used.

[0248] Genomic DNA was extracted from the harvested cells and the sites of transposon integrations into the genome were determined using a linker mediated PCR method (LM- PCR). Briefly, genomic DNA was tagmented using Tn5 transposase loaded with oligos (linkers) containing adapter sequences and unique molecular identifiers (UMI). Nested PCR was performed using forward primers that bind within the transposon and reverse primers that bind within the linker added by tagmentation.

[0249] This strategy selectively amplifies fragments containing the junction site between an integrated transposon and its genomic integration site. The adapter sequences on the linker enable next-generation sequencing (NGS) on an Illumina sequencing platform and the UMIs allow one to distinguish PCR duplications from unique integration events.

[0250] Integration events detected by LM-PCR were mapped to the reference human genome and classified as on-target (occurring at the desired rDNA TTAA site) or off-target (occurring at a non-rDNA TTAA site). LM-PCR was performed in triplicate and the sum of replicates shown. The UMI count is an upper estimate of the number of unique integration events detected as UMIs can become inflated by linker swapping during PCR. The fragmentation abundance gives a lower estimate of the number of unique integration events detected as fragment length can quickly become saturated as the number of integration events increases. The percentage of on-target integration events, UMI counts, and fragmentation abundance are shown in Table 25. Table 25. Fidelity of in vivo site-specific transposition in mouse liver by hybrid delivery of rDNA – TAL Array – PBx fusion proteinsAttorney Docket No.: 000218-0153-WO1

[0251] As shown in Table 25, site-specific integration events catalyzed by rDNA TAL Array – PBx fusion protein were detected at rDNA site 5 with greater than 99% of the mapped integrations being on-target as quantified by LM-PCR. Example 7: Trafficking TAL-ssSPB to 45S rDNA repeats with R2 retrotransposase

[0252] R2 retrotransposon integrates into 45S rDNA using a target-primed reverse transcription mechanism. Bombyx mori R2 retrotransposase (SEQ ID NO: 123) is fused to the N-terminus of a PBx transposase domain or to rDNA – TAL Array- PBx fusion protein as a means of trafficking piggyBac transposon to rDNA and increase site-specific integration into rDNA. R2 mediated integrations can be prevented by introducing D996A and D1009A mutations to the endonuclease domain, i.e., a catalytically dead retrotransposase (SEQ ID NO: 124).

[0253] In a first experiment, a dual CRD PBx lacking the first 85 amino acids (SEQ ID NO: 10) was fused down stream of either wild type R2 sequence (SEQ ID NO: 123) or a catalytically dead R2 sequence with mutated endonuclease (SEQ ID NO: 124) with a GGGS linker (SEQ ID NO: 125). A FLAG tag (SEQ ID NO: 192) and an SV40 NLS (SEQ ID NO: 191) were added to the N-terminus to generate R2-PBx (SEQ ID NO: 188) and R2- endonuclease dead-PBx fusions (SEQ ID NO: 189).

[0254] For mammalian expression, each R2-PBx was human codon optimized using Thermo’s GeneArt algorithm and cloned into the expression pcDNA3.1 expression vector downstream of the CMV promoter and Kozak sequence and upstream of the BGH polyadenylation signal sequence.

[0255] As a transposon donor, a plasmid harboring a small transposon with symmetrical ITRs was co-transfected. Described briefly, the transposon consists in 5’ to 3’ direction of aAttorney Docket No.: 000218-0153-WO1 TTAA, the piggyBac 35bp 5’ ITR, 274bp of the piggyBac 5’ UTR, 55bp of stuffer, 175bp of the piggyBac 3’UTR, the reverse compliment of a second copy of the piggyBac 35bp 5’ ITR, and a second TTAA. The full sequence of the small transposon with symmetrical ITRs is given in SEQ ID NO: 202.

[0256] For testing in HEK293T cells, 120,000 cells were plated in 24-well plates in 0.5ml DMEM+10% FBS one day prior to transfection. The following day, the cells were transfected with 50 ng of a transposase expression vector mixture and 450 ng of the transposon donor vector using 1 µl of JetPrime transfection reagent. The transposase expression vector mixture consisted of either R2-PBx alone (SEQ ID NOs 188 or 189), the TAL Array – PBx Plus pair alone in a 1:1 ratio, R2-PBx plus the TAL Array - PBx Plus pair in a 1:1:1 ratio, or PBx alone as a negative control. TAL Array -PB Plus pairs targeting rDNA sites 5 (SEQ ID NOs: 130 & 131), rDNA site 22 (SEQ ID NOs: 144 & 145), and rDNA site 23 (SEQ ID NOs: 146 & 147), along with TAL Array -PBx Plus pair targeting LINE1 (SEQ ID NOs: 160 & 161) were used.

[0257] Two days after the transfection, genomic DNA was harvested from the cells. Site specific transposition of the transposon into the 45S rDNA repeats in the forward orientation was detected and quantified by ddPCR using one primer that binds within the transposon and a second primer that binds within the genome flanking the integration site. A second primer flanking the other side of the integration site was used to detect integrations of the transposon in the reverse orientation. The number of integrations per haploid genome for each rDNA TAL Array – PBx Plus, LINE1 - TAL Array - PBx Plus and / or R2-PBx are shown in Table 26. Table 26. Site-specific transposition into 45S rDNA repeats Using rDNA TAL Array – PBx, LINE1 - TAL Array – PBx Plus fusion proteins and R2-PBx fusion proteinAttorney Docket No.: 000218-0153-WO1

[0258] As shown in Table 26, no site-specific integration was observed at any target site with the control PBx transposase alone. Using the R2-PBx or R2 endonuclease mutant-PBx alone, low site-specific integration was observed at rDNA targets but not at LINE1 targets. Using the rDNA TAL Array – PBx, LINE1 - TAL Array - PBx Plus fusion proteins, robust site-specific integration was observed at the respective target site. R2- rDNA TAL Array – PBx Plus, R2- LINE1 - TAL Array - PBx Plus fusions further boosted site-specific integration specifically at the rDNA targets but reduced site-specific integration at the non- rDNA LINE1 target, indicative of R2 mediated trafficking to rDNA in the nucleolus.

[0259] In a second experiment, the wild type R2 retrotransposon transposase sequence was fused directly to rDNA TAL Array – PBx Plus fusion protein on the N-terminus immediately following the sequence comprising the FLAG-tag (SEQ ID NO: 192) and SV40 NLS (SEQ ID NO: 191). Specifically, the pair of rDNA – TAL Array – PBx fusion proteins for rDNA site 23 (SEQ ID NOs: 146 & 147) were fused to the C-terminus of the R2 transposase sequence (SEQ ID NO: 123) to create R2-rDNA 23L TAL Array – PBx Plus fusion protein (SEQ ID NO: 187) and R2-rDNA 23R TAL Array - PBx (SEQ ID NO: 188).

[0260] The covalent R2-rDNA 23 TAL Array - PBx Plus fusion proteins were compared to R2-PBx co-delivered with rDNA 23 TAL Array - PBx Plus fusion proteins as described above in Example 2. Additionally, the constructs were delivered to K562 cells by nucleofection. Briefly, 200,000 K562 cells were nucleofected in a 20 µl volume with 700 ng of transposon donor and 300 ng of the transposase mixture. Two days later, genomic DNAAttorney Docket No.: 000218-0153-WO1 was harvested and analyzed as described above. The number of integrations per haploid genome for each rDNA 23 TAL Array - PBx and / or R2-PBx are shown in Table 27. Table 27. Site-specific transposition into 45S rDNA repeats with R2-rDNA-TAL Array-PBx Plus fusion proteins at rDNA site 23

[0261] As shown in Table 27, low site-specific integration at rDNA site 23 was observed with R2-PBx alone, while robust site-specific integration was observed with the TAL Array PBx Plus pair. Co-delivery of R2-PBx with the rDNA – 23 TAL Array - PBx Plus pair or direct fusion of R2 to the N-terminus of the rDNA – 23 TAL Array -PBx Plus fusion proteins both increased site-specific integration at rDNA site 23 in HEK293T but reduced site-specific integration in K562 cells. Example 8: Trafficking rDNA – TAL Array – PBx Plus fusion proteins to 45S rDNA repeats with RNA Polymerase PolI subunits

[0262] RNA PolI exclusively transcribes rDNA. PolI is a complex and several subunits are shared among polymerases (PolII and PolIII). Other subunits though, are unique to PolI and include A12.2 (POLR1H), A34.5 (POLR1G), A43 (POLR1F), and A49 (POLR1E). The amino acid sequences of the A12.2, A34.5, A43, and A49 subunits are set forth in SEQ ID NOs: 119-122.

[0263] In this example, fusion of PolI subunits to a transposase provides a means of shuttling the transposase and transposon specifically to rDNA. The sequences of a dual CRDAttorney Docket No.: 000218-0153-WO1 PBx Plus truncated by the first 85 amino acids (SEQ ID NO: 10) fused directly downstream of the PolI subunits and further comprising an N-terminal FLAG-tag (SEQ ID NO: 192) and SV40 NLS (SEQ ID NOs: 191) are set forth in SEQ ID NOs: 204 - 207.

[0264] In addition, the A12.2, A34.5, A43, and A49 PolI subunit sequences were fused directly to the N-terminus of TAL Array -PBx Plus fusion protein replacing the FLAG-tag and the N-terminal 85 amino acids of PBx. The rDNA TAL Array -PBx Plus pair for rDNA site 5 (SEQ ID NOs: 130 & 131) were fused to each PolI subunit to create SEQ ID NOs:162 - 169. The rDNA TAL Array - PBx Plus pair for rDNA site 23 (SEQ ID NOs: 145 and 146) was fused to each PolI subunit to create (SEQ ID NOs: 170 - 177). The LINE1 TAL Array - PBx pair (SEQ ID NOs: 160 and 161) was fused to each PolI subunit to create A12.2, A34.5, A43 and A49 fusions (SEQ ID NOs: 178 - 185). For mammalian expression, each of the R2- PBx fusion proteins was human codon optimized using Thermo’s GeneArt algorithm and cloned into the expression pcDNA3.1 expression vector downstream of the CMV promoter and Kozak sequence and upstream of the BGH polyadenylation signal sequence.

[0265] As a transposon donor, a plasmid harboring a small transposon with symmetrical ITRs was co-transfected with the rDNA or LINE1 - TAL-Array PBx Plus pairs. Described briefly, the transposon consists in 5’ to 3’ direction of a TTAA, the piggyBac 35bp 5’ ITR, 274bp of the piggyBac 5’ UTR, 55bp of stuffer, 175bp of the piggyBac 3’UTR, the reverse compliment of a second copy of the piggyBac 35bp 5’ ITR, and a second TTAA. The full sequence of the small transposon with symmetrical ITRs is given in SEQ ID NO: 202.

[0266] For testing in HEK293T cells, 120,000 cells were plated in 24-well plates in 0.5ml DMEM+10% FBS one day prior to transfection. The following day, the cells were transfected with 50 ng of the rDNA or LINE1 – TAL Array -PBx Plus pair or the PolI- rDNA / LINE1 TAL Array -PBx Plus pair of expression vectors and 450 ng of the transposon donor vector using 1 µl of JetPrime transfection reagent. As a negative control PBx transposase (SEQ ID NO: 6) was used in place of rDNA or LINE1 – TAL Array -PBx Plus fusion proteins.

[0267] For testing in K562 cells, 200,000 cells were nucleofected in a 20 µl volume with 500 ng of transposon donor and 500 ng of the rDNA or LINE1 – TAL Array -PBx Plus pair or the PolI-rDNA / LINE1 TAL Array -PBx Plus pair of expression vectors. As a negative control PBx transposase (SEQ ID NO: 6) was used in place of rDNA / LINE1 TAL Array - PBx Plus pair.

[0268] Two days after the transfection, genomic DNA was harvested from the cells. Site specific transposition of the transposon into the 45S rDNA repeats in the forward orientationAttorney Docket No.: 000218-0153-WO1 was detected and quantified by ddPCR using one primer that binds within the transposon and a second primer that binds within the genome flanking the integration site. The number of integrations in the forward direction per haploid genome for each rDNA / LINE1 TAL Array - PBx and PolI-rDNA / LINE1 TAL Array -PBx are shown in Table 28. Table 28. Site-specific transposition with TAL Array -PBx Plus and PolI-TAL array -PBx Plus Fusion ProteinsAttorney Docket No.: 000218-0153-WO1

[0269] As shown in Table 26, rDNA and LINE1 – TAL Array -PBx Plus fusion proteins catalyzed robust site-specific integration into its respective target site. Fusions of PolI subunits to rDNA and LINE1 – TAL Array -PBx Plus fusion protein increased site-specific integration at rDNA targets with A43 subunit having the strongest effect. In cells that express low levels of fusion proteins (K562), PolI-comprising fusion proteins trafficked away from the non-rDNA target LINE1, resulting in lower site-specific integrations at LINE1, demonstrating specificity for targeting rDNA in the nucleolus. Example 9: In Vivo site-specific integration at 45S rDNA repeats in mouse liver by hybrid delivery of PolI-rDNA-TAL Array – PBx Plus Fusion Proteins

[0270] Site-specific integrations catalyzed by PolI-rDNA –TAL Array – PBx Plus into the 45S rDNA repeats were performed in vivo in mouse liver using hybrid delivery. With hybrid delivery the transposase is delivered by lipid nanoparticle (LNP) in mRNA format and the transposon is delivered in adeno-associated virus (AAV) format.

[0271] Briefly, the A43 subunit PolI-rDNA – TAL Array – PBx constructs targeting rDNA site 5 (SEQ ID NOs:166 & 167) and the rDNA – TAL Array – PBx constructs targeting rDNA site 5 (SEQ ID NOs: 130 & 131) were cloned downstream into an in vitro transcription vector harboring a T7 promoter, stabilizing UTR sequences, and a polyA tail. In a first experiment, mRNA was produced with a Cleancap M6 and 5-methylcytosine base modifications. In a second experiment, mRNA was produced with either Cleancap M6 or Cleancap and with either 5-methylcytosine or n1-methyl-pseudouridine base modifications, and with or without an additional enzymatic polyadenylation step with E.coli poly(A) polymerase. mRNA was encapsulated into lipid nanoparticles (LNPs) using an H2-8 formulation (50% HBC365 lipid; 38.5% Cholesterol; 10% DOPC and 1.5% DMG-PEG2000; Lipid: Nucleic acid ratio 50:1). Compositions and methods for preparing HBC365 (Compound No.37)-comprising LNPs are disclosed in co-owned International Patent Application No: PCT / US2024 / 012245. As negative controls, a PBx transposase mRNA (SEQ ID NO: 6) lacking a TAL Array fusion was formulated or PBS (buffer only) was delivered.Attorney Docket No.: 000218-0153-WO1

[0272] As a transposon donor, a disrupted CAGG-firefly luciferase reporter described above was cloned between AAV ITR sequences. From this, AAV particles of AAV9 serotype were produced by VectorBuilder. The titer was determined by ddPCR as genome copies (GC) per µl. Upon transposon mediated excision of the transposon and seamless repair of the donor, luciferase expression is restored. Site-specific integration of the transposon into the genome can be detected and quantified by ddPCR as described above.

[0273] Three-week-old juvenile C57BL / 6 mice were injected with 3 mg / kg of LNP- mRNA transposase and 3E13 GC / kg of AAV transposon donor. Luciferase expression, marking transposon excision, was quantified four days post-delivery. Following luciferase measurements, the mice were sacrificed and genomic DNA was extracted from the liver. Site-specific integration into 45S rDNA was quantified by ddPCR as described above. The excision activity, measured by luciferase signal in photons / second, and the site-specific integration, measured by ddPCR as the sum of forward and reverse integrations, are provided in Table 27. Table 27. In vivo site-specific transposition in mouse liver by hybrid delivery of PolI-rDNA – TAL Array -PBx & rDNA – TAL Array -PBx Fusion ProteinsAttorney Docket No.: 000218-0153-WO1

[0274] As shown in Table 27, PolI-rDNA – TAL Array – PBx fusion proteins, rDNA – TAL Array – PBx fusion proteins, and PBx transposase catalyzed robust excision of the transposon. Site-specific integration into rDNA targets, however, was detected only in the presence of Site 5 PolI-rDNA – TAL Array – PBx and rDNA – TAL Array – PBx fusion proteins but not with PBx transposase or PBS buffer alone. Consistent with example 5, rDNA – TAL Array – PBx fusion proteins resulted in less than 1 site-specific integrations per 100 haploid genomes. PolI-rDNA – TAL Array – PBx consistently resulted in approximately 1 to 6 integrations per 100 haploid genomes across both experiments. PolI-rDNA – TAL Array – PBx mRNA produced with enzymatic polyadenylation resulted in higher editing than mRNA produced without enzymatic polyadenylation. PolI-rDNA – TAL Array – PBx mRNA produced with cleancap M6 resulted in higher editing than mRNA produced cleancap AG. PolI-rDNA – TAL Array – PBx mRNA produced with n1-methyl-pseudouridine base modifications resulted in higher editing than mRNA produced 5-methylcytosine. Example 10: In Vivo site-specific integration at 45S rDNA repeats in mouse liver by full non-viral delivery of PolI-rDNA-TAL Array – PBx Plus Fusion Proteins

[0275] Site-specific integrations catalyzed by PolI-rDNA –TAL Array – PBx Plus into the 45S rDNA repeats were performed in vivo in mouse liver using full non-viral delivery.Attorney Docket No.: 000218-0153-WO1 With full non-viral delivery the transposase is delivered by lipid nanoparticle (LNP) in mRNA format and the transposon is delivered by (LNP) in nanoplasmid format.

[0276] Briefly, the A43 subunit PolI-rDNA – TAL Array – PBx constructs targeting rDNA site 5 (SEQ ID NOs:166 & 167) were cloned downstream into an in vitro transcription vector harboring a T7 promoter, stabilizing UTR sequences, and a polyA tail. mRNA was produced with a Cleancap M6 and 5-methylcytosine base modifications. mRNA was encapsulated into lipid nanoparticles (LNPs) using an H2-8 formulation (50% HBC365 lipid; 38.5% Cholesterol; 10% DOPC and 1.5% DMG-PEG2000; Lipid: Nucleic acid ratio 50:1). As a negative control, PBS (buffer only) was delivered.

[0277] As a transposon donor, a nanoplasmid harboring a small transposon with symmetrical ITRs was co-delivered. The nanoplasmid was designed such that the small transposon disrupts the open reading frame of a CAGG-firefly luciferase reporter. Upon transposon mediated excision of the transposon and seamless repair of the donor, luciferase expression is restored. Site-specific integration of the transposon into the genome can be detected and quantified by ddPCR as described above. Nanoplasmid was encapsulated into lipid nanoparticles (LNPs) using an H3-5 formulation (50% HBC365 lipid; 41% Cholesterol; 7.5% DOPC and 1% DMG-PEG2000; 0.5% GalNac-PEG; 0.15% pentagalloylglucose; Lipid: Nucleic acid ratio 60:1). Compositions and methods for preparing HBC365 (Compound No. 37) and polyphenol-comprising LNPs are disclosed in co-owned International Patent Application No: PCT / US2024 / 012245 (published as WO 2024 / 155938) and US Provisional Application No: 63 / 678,021.

[0278] Three-week-old juvenile C57BL / 6 mice were treated with 8 mg / kg dexamethasone two hours prior to injection with 3 mg / kg of LNP-mRNA transposase and either 0.15 mg / kg, 0.33 mg / kg, 0.66 mg / kg, or 1.00 mg / kg of LNP-transposon donor. Luciferase expression, marking transposon excision, was quantified three days post-delivery. Following luciferase measurements, the mice were sacrificed and genomic DNA was extracted from the liver. Site-specific integration into 45S rDNA was quantified by ddPCR as described above. The excision activity, measured by luciferase signal in photons / second, and the site-specific integration, measured by ddPCR as the sum of forward and reverse integrations, are provided in Table 28. Table 28. In vivo site-specific transposition in mouse liver by full non-viral delivery of PolI- rDNA – TAL Array -PBx Fusion ProteinsAttorney Docket No.: 000218-0153-WO1As shown in Table 28, PolI-rDNA – TAL Array – PBx fusion proteins catalyzed robust excision of the transposon. Increasing site-specific integration events into rDNA targets were detected with increasing dose of Site 5 PolI-rDNA – TAL Array – PBx.

Claims

Attorney Docket No.: 000218-0153-WO1 CLAIMS What is claimed is:

1. A fusion protein, comprising, in N-terminal to C-terminal direction:(a) a TAL Array targeting an rDNA repeat; and (b) a transposase of SEQ ID NO: 6 comprising an N-terminal deletion.

2. The fusion protein of claim 1, wherein the transposase comprises a N-terminaldeletion comprising a deletion of amino acids 1-74, 1-83, 1-84, 1-85, 186, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102 or 1-103 of SEQ ID NO: 6.

3. The fusion protein of claim 1 or 2, further comprising a nucleolar trafficking sequence(NTS), wherein the C-terminus of the NTS is attached to the N-terminus of the TAL Array.

4. The fusion protein of claim 3, wherein the NTS is an RNA polymerase I (PolI)subunit.

5. The fusion protein of claim 4, wherein the PolI subunit is an A12.2 subunit, a A35.4subunit, an A43 subunit or an A49 subunit.

6. The fusion protein of claim 4 or 5, wherein the PolI subunit comprises the amino acidsequence of any one of SEQ ID NO: 119-122.

7. The fusion protein of claim 3, wherein the NTS is an R2 retrotransposon transposasesequence.

8. The fusion protein of claim 7, wherein the R2 retrotransposon transposase sequencecomprises the amino acid sequence of SEQ ID NO: 123 or 124.

9. The fusion protein of any one of claims 1-8, further comprising a protein stabilizationdomain (PSD).

10. The fusion protein of claim 9, wherein the PSD comprises the amino acid sequence ofSEQ ID NO: 190.

11. The fusion protein of any one of claims 1-10, further comprising a nuclearlocalization sequence (NLS).Attorney Docket No.: 000218-0153-WO112. The fusion protein of claim 11, wherein the NLS sequence comprises the amino acidsequence of SEQ ID NO: 191.

13. The fusion protein of any of claims 1-12, further comprising an GS or GGGGS (SEQID NO: 125) linker positioned (a) between the TAL Array and the SPB transposase; (b) between nucleolar trafficking sequence and the TAL Array; and / or (c) between the PSD and the TAL Array.

14. The fusion protein of any one of claims 1-13, wherein the SPB transposase comprisesthe amino acid sequence of any of SEQ ID Nos: 8-28 or 208.

15. The fusion protein of any one of claims 1-14, wherein the TAL Array comprises oneor more TAL domains targeting an rDNA 45S repeat, an rDNA 18S repeat, an rDNA 5.8S repeat or an rDNA 28S repeat.

16. The fusion protein of claim 15, wherein each of the one or more TAL domains targetsan rDNA repeat comprising the nucleic acid sequence of any one of SEQ ID NOs: 59 - 80.

17. The fusion protein of claim 15 or 16, wherein the TAL Array comprises the aminoacid sequence of any one of SEQ ID NO: 81-114.

18. The fusion protein of any one of claims 1-17, wherein the fusion protein comprisesthe amino acid sequence of any one of SEQ ID NOs: 126-159.

19. A polynucleotide comprising a nucleic acid sequence encoding the fusion protein ofany one of claims 1-18.

20. A vector comprising the polynucleotide of claim 19.

21. A pharmaceutical composition comprising (a) the fusion protein of any one of claims1-18, the polynucleotide of claim 19, or the vector of claim 20 and (b) a pharmaceutically acceptable excipient.

22. A method of modifying the genome of a cell, the method comprisingintroducing into the cell a. a polynucleotide comprising a nucleic acid encoding a fusion proteincomprising a transposase of SEQ ID NO: 6 comprising an N-terminalAttorney Docket No.: 000218-0153-WO1 deletion; and a TAL Array targeting an rDNA repeat, wherein the C- terminus of the TAL Array is fused to the N-terminus of the SPB; and b. a DNA molecule comprising a transposon, wherein the transposoncomprises, in 5’ to 3’ order: a 5’ITR, the transgene, and a 3’ ITR, wherein the transgene is integrated into a TTAA site flanked by a left rDNA target site and a right rDNA target site.

23. The method of claim 22, wherein the transposon is a small transposon comprisingsymmetrical ITRs.

24. The method of claim 22 or 23, wherein the transposon comprises the sequence ofSEQ ID NO: 202.

25. The method of any one of claims 22-24, wherein the transposon further comprises anexogenous promoter operatively linked to the transgene.

26. The method of any one of claims 22-25, wherein the transgene encodes a protein theexpression of which is beneficial for treating a metabolic disorder, hemophilia, cancer or PKU.

27. The method of any one of claims 22-26, wherein the TTAA site comprises the nucleicacid sequence of any one of SEQ ID NOs 36-58.

28. The method of any one of claims 22-27, wherein the rDNA repeat comprises thenucleic acid sequence of any one of SEQ ID NOs: 59-80.

29. The method of any one of claims 22-28, wherein the left rDNA target site comprisesthe sequence of SEQ ID NO: 63 and the right rDNA target site comprises the sequence of SEQ ID NO: 64.

30. The method of any one of claims 22-28, wherein the left rDNA target site comprisesthe sequence of SEQ ID NO: 79 and the right rDNA target site comprises the sequence of SEQ ID NO: 80.

31. The method of any one of claims 22-30, wherein the cell is in vivo.Attorney Docket No.: 000218-0153-WO132. The method of any one of claims 22-31, wherein the transgene is integrated into thegenome of the cell with a transposon fidelity of on target to off target transposition integration ratio of at least 89%.

33. The method of any one of claims 22-32, wherein the cell is a liver cell.

34. A method for generating an engineered cell by site-specific transposition, comprisingintroducing into a cell (a) a nucleic acid encoding a fusion protein comprising a rDNA TAL Array and a transposase domain comprising the sequence of SEQ ID NO: 6; and (b) a DNA molecule comprising a transposon; wherein the transposon is integrated by site-specific transposition into a TTAA sequence of an rDNA repeat.

35. The method of claim 34, wherein the DNA molecule comprising a transposon is ananoplasmid or an AAV virion.

36. The method of any one of claims 34-35, wherein the nucleic acid encoding the fusionprotein is mRNA introduced into the cell by a lipid nanoparticle.

37. The method of any of claims 34-36, wherein the nucleic acid encoding the fusionprotein and / or the DNA molecule comprising the transposon are introduced into the cell by a lipid nanoparticle.

Citation Information

Patent Citations

  • DNA vectors, transposons and transposases for eukaryotic genome modification

    US10041077B2

  • Data backup and recovery method for mobile terminal and mobile terminal

    US20140046903A1

  • Polylactide-drug mixtures

    US3773919A

  • Identification and preparation of epitopes on antigens and allergens on the basis of hydrophilicity

    US4554101A

  • Gene amplification in eukaryotic cells

    US4656134A