Crispr-associated transposon systems and methods

The engineered CRISPR-associated transposon systems address integration challenges in eukaryotic cells by using modified Cas and transposon proteins, achieving efficient and specific nucleic acid integration for therapeutic applications.

WO2025235884A1PCT designated stage Publication Date: 2025-11-13THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/028640
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-05
Filing Date
2025-05-09
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems are limited in their ability to efficiently integrate and modify nucleic acids in eukaryotic cells, particularly due to challenges in targeting and integration specificity, and lack of efficient non-cleavage functions for therapeutic applications.

Method used

Engineering of CRISPR-associated transposon (CAST) systems using modified Cas proteins (Cas5, Cas6, Cas7, Cas8) and transposon-associated proteins (TnsA, TnsB, TnsC, TniQ) with specific amino acid substitutions and nuclear localization sequences, combined with guide RNAs, to enhance nucleic acid integration and modification in both prokaryotic and eukaryotic cells.

Benefits of technology

The engineered CAST systems demonstrate improved integration efficiency and specificity in eukaryotic cells, enabling therapeutic applications such as gene correction and disease treatment by integrating donor nucleic acids at desired genomic loci with reduced off-target effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025028640_13112025_PF_FP_ABST
    Figure US2025028640_13112025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CAST) systems, components thereof, and methods for nucleic acid modification using the systems or components.
Need to check novelty before this filing date? Find Prior Art

Description

CRISPR-ASSOCIATED TRANSPOSON SYSTEMS AND METHODSFIELD

[0001] The present disclosure relates to Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CAST) systems and components thereof, for example, Cas proteins and transposon-associated proteins.SEQUENCE LISTING STATEMENT

[0002] The content of the electronic sequence listing titled COLUM_43200_601_SequenceListing.xml (Size: 579,286 bytes; and Date of Creation: May 8, 2025) is herein incorporated by reference in its entirety.CROSS REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Application Nos. 63 / 644,883, filed May 9, 2024, 63 / 709,110, filed October 18, 2024, and 63 / 767,039, filed March 5, 2025, the contents of which are herein incorporated by reference in their entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0004] This invention was made with government support under HG011650, EB031935, HG009490, EB027793, EB031172, GM118062, and AI142756 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND

[0005] In bacteria and archaea, CRISPR / Cas systems provide immunity by incorporating fragments of invading phage, virus, and plasmid DNA into CRISPR loci and using corresponding CRISPR RNAs (“crRNAs”) to guide the degradation of homologous sequences. Transcription of a CRISPR locus produces a “pre-crRNA,” which is processed to yield crRNAs containing spacer-repeat fragments that guide effector nuclease complexes to cleave dsDNA sequences complementary to the spacer. Several different types of CRISPR systems are known, (e.g., type I, type II, or type III), and classified based on the Cas protein type and the use of a proto-spacer-adjacent motif (PAM) for selection of proto-spacers in invading DNA.

[0006] Although RNA-guided targeting typically leads to endonucleolytic cleavage of the bound substrate, recent studies have uncovered a range of noncanonical pathways in which CRISPR protcin-RNA effector complexes have been naturally repurposed for alternativefunctions. For example, some Type I (Cascade) and Type II (Cas9) systems leverage truncated guide RNAs to achieve potent transcriptional repression without cleavage and other Type I (Cascade) and Type V (Cas 12) systems lie inside unusual bacterial Tn7-like transposons and lack nuclease components altogether.SUMMARY

[0007] Provided herein are engineered Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CAST) systems and methods utilizing thereof.

[0008] In some embodiments, the engineered CAST system comprises: a) one or more Cas proteins selected from: Cas5, Cas6, Cas7, Cas8, and combinations thereof; and b) one or more transposon-associated proteins selected from TnsA, TnsB, TnsC, TniQ, and combinations thereof.

[0009] In some embodiments, the one or more Cas proteins comprise a Cas8-Cas5 fusion protein comprising an amino acid sequence having amino acid substitutions at positions 125, 244, and 410 relative to SEQ ID NO: 5 and / or a Cas7 protein comprising an amino acid sequence having an amino acid substitutions at position 347 relative to SEQ ID NO: 6. In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having amino acid substitutions N125D, A244N, and A410R relative to SEQ ID NO: 5. In some embodiments, the Cas7 protein comprises an amino acid sequence having an amino acid substitution A347K relative to SEQ ID NO: 6.

[0010] In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 225. In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence of SEQ ID NO: 225.

[0011] In some embodiments, the Cas7 protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 224. In some embodiments, the Cas7 protein comprises an amino acid sequence of SEQ ID NO: 224.

[0012] In some embodiments, the one or more Cas proteins further comprise a Cas6 protein comprising an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, orat least 99%) identity to SEQ ID NO: 7. In some embodiments, the one or more Cas proteins further comprise a Cas6 protein comprising an amino acid sequence of SEQ ID NO: 7.[00131 In some embodiments, the one or more transposon-associated proteins comprise aTnsA protein comprising an amino acid sequence having amino acid substitutions at positions 88, 147, 170, 180, and 182 relative to SEQ ID NO: 1, a TnsB protein comprising an amino acid sequence having amino acid substitutions at positions 43, 349, 352, 390, 396, 410, 464, 526, 549, and 594 relative to SEQ ID NO: 2, and / or a TnsC protein comprising an amino acid sequence having amino acid substitutions at positions 197 and 314 relative to SEQ ID NO: 3. In some embodiments, the TnsA protein comprises an amino acid sequence having amino acid substitutions P88T, I147V, V170L, F180L, and F182L relative to SEQ ID NO: 1. In some embodiments, the TnsB protein comprises an amino acid sequence having amino acid substitutions F43S, Y349N, P352T, A390V, D396N, Q410K, H464R, V526E, Q549R, and Q594L relative to SEQ ID NO: 2. In some embodiments, the TnsC protein comprises an amino acid sequence having amino acid substitutions R197I and N314K relative to SEQ ID NO: 3.

[0014] In some embodiments, the TnsA protein comprises protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 226. In some embodiments, the TnsA protein comprises an amino acid sequence of SEQ ID NO: 226.

[0015] In some embodiments, the TnsB protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 227. In some embodiments, the TnsB protein comprises an amino acid sequence of SEQ ID NO: 227.

[0016] In some embodiments, the TnsC protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 228. In some embodiments, the TnsC protein comprises an amino acid sequence of SEQ ID NO: 228.

[0017] In some embodiments, the one or more transposon-associated proteins further comprise a TniQ protein comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 4. In some embodiments, the one or more transposon-associated proteins further comprise a TniQ protein comprising an amino acid sequence of SEQ ID NO: 4.

[0018] In some embodiments, the TnsA and TnsB proteins are linked in a TnsA-TnsB fusion protein. In some embodiments, the TnsA-TnsB fusion protein further comprises an amino acid linker between TnsA and TnsB. In some embodiments, the linker is a flexible linker. In some embodiments, the linker comprises a nuclear localization sequence.

[0019] In some embodiments, one or more Cas proteins and transposon-associated proteins are part of a single fusion protein. In some embodiments, each of the Cas proteins and transposon-associated proteins are part of a single fusion protein.

[0020] In some embodiments, any or all of the one or more Cas proteins and the one or more transposon-associated proteins comprise at least one nuclear localization sequence (NLS). In some embodiments, one or more of the one or more Cas proteins and the one or more transposon-associated proteins comprise two or more nuclear- localization sequences.

[0021] In some embodiments, the Cas7 protein comprises a single nuclear localization sequence. In some embodiments, the Cas7 protein comprises a single nuclear localization sequence at the N terminus of the Cas7 protein sequence. In some embodiments, the TnsA protein and the TnsB protein each comprise a single nuclear localization sequence. In some embodiments, the TnsA-TnsB fusion protein comprises a single nuclear localization sequence. In some embodiments, the TnsA-TnsB fusion protein comprises a single nuclear localization sequence in the linker region between the TnsA protein and the TnsB protein. In some embodiments, the Cas6 protein, the Cas8-Cas5 fusion protein, and / or the TniQ protein each comprise two nuclear localization sequences. In select embodiments, the Cas6 protein, the Cas8- Cas5 fusion protein, and the TniQ protein comprise two nuclear localization sequences at the N- terminus of the protein (e.g., TniQ, Cas6, Cas8-Cas5 fusion protein) sequence. In some embodiments, the TnsC protein comprises two or more nuclear localization sequences. In some embodiments, the TnsC protein comprises three nuclear localization sequences. In some embodiments, the TnsC protein comprises two or more (e.g., three) nuclear localization sequences at the Cterminus of the protein.

[0022] In some embodiments, the Cas8-Cas5 fusion protein is encoded by a nucleic acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 10; the Cas7 protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 9; the Cas6 protein is encoded by a nucleic acid sequence having at least70% identity to SEQ ID NO: 8; the TnsA protein and the TnsB protein are provided as a TnsA- TnsB fusion protein encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 12; the TnsC protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 13; and / or the TniQ protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 11.

[0023] In some embodiments, the Cas8-Cas5 fusion protein is encoded by a nucleic acid sequence of SEQ ID NO: 10; the Cas7 protein is encoded by a nucleic acid sequence of SEQ ID NO: 9; the Cas6 protein is encoded by a nucleic acid sequence of SEQ ID NO: 8; the TnsA protein and the TnsB protein are provided as a TnsA-TnsB fusion protein encoded by a nucleic acid sequence of SEQ ID NO: 12; the TnsC protein is encoded by a nucleic acid sequence of SEQ ID NO: 13; and / or the TniQ protein is encoded by a nucleic acid sequence of SEQ ID NO: 11.[0024| In some embodiments, the engineered CAST system comprises: a Cas8-Cas5 fusion protein is encoded by a nucleic acid sequence of SEQ ID NO: 10; a Cas7 protein is encoded by a nucleic acid sequence of SEQ ID NO: 9; a Cas6 protein is encoded by a nucleic acid sequence of SEQ ID NO: 8; a TnsA protein and a TnsB protein is provided as a TnsA-TnsB fusion protein encoded by a nucleic acid sequence of SEQ ID NO: 12; a TnsC protein is encoded by a nucleic acid sequence of SEQ ID NO: 13; and a TniQ protein is encoded by a nucleic acid sequence of SEQ ID NO: 11.

[0025] In some embodiments, the one or more Cas proteins are encoded by a single nucleic acid. In some embodiments, the one or more transposon-associated proteins are encoded by a single nucleic acid. In some embodiments, the one or more Cas proteins and the one or more transposon-associated proteins are encoded on a single nucleic acid. In some embodiments, the one or more Cas proteins and the one or more transposon-associated proteins are encoded by different nucleic acids. In some embodiments, the one or more nucleic acids comprise one or more messenger RNAs, one or more vectors, or a combination thereof.

[0026] In some embodiments, the system further comprises at least one guide RNA (gRNA) complementary to at least a portion of a target nucleic acid, or at least one nucleic acid encoding thereof. In some embodiments, the one or more Cas protein, the one or more transposon- associated protein, and the at least one gRNA are encoded by different nucleic acids. In someembodiments, at least one of the one or more Cas protein and the one or more transposon- associated protein, and the at least one gRNA arc encoded by a single nucleic acid.

[0027] In some embodiments, the at least one gRNA is a non-naturally occurring gRNA. In some embodiments, the at least one gRNA is encoded in a CRISPR RNA (crRNA) array. In some embodiments, at least one of the one or more Cas protein is part of a ribonucleoprotein complex with the at least one gRNA.

[0028] In some embodiments, the system further comprises a donor nucleic acid, wherein the donor nucleic acid comprises a cargo nucleic acid sequence flanked by at least one transposon end sequence. In some embodiments, the system further comprises a target nucleic acid.

[0029] In some embodiments, the system is a cell-free system.

[0030] Also provided are compositions and cells comprising the disclosed systems. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell, a human cell).

[0031] Additionally provided are methods for nucleic acid modification and integration. In some embodiments, the methods comprise contacting a target nucleic acid with a system, or composition thereof, as disclosed herein.

[0032] In some embodiments, the target nucleic acid sequence is in a cell. In some embodiments, contacting a target nucleic acid sequence comprises introducing the system into the cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell, a human cell).

[0033] In some embodiments, introducing the system into the cell comprises administering the system to a subject. In some embodiments, administering comprises in vivo administration. In some embodiments, the administering comprises transplantation of ex vivo treated cells comprising the system.

[0034] Also provided are methods for treating a disease or disorder in a subject comprising administering to the subject in need thereof a system, or composition thereof, as described herein. In some embodiments, the subject is human. In some embodiments, the system or composition comprises a donor nucleic acid encoding a therapeutic gene product or a wild-type or corrected version of a disease-associated gene.

[0035] Further provided are methods for inactivating a microbial gene, the method comprising introducing into one or more cells a system, or a composition thereof, as describedherein. In some embodiments, the gRNA is specific for a target site that is proximal to the microbial gene and the system or composition modifies the microbial gene. In some embodiments, the system or composition inserts a donor nucleic acid within the microbial gene. In some embodiments, the microbial gene is a bacterial antibiotic resistance gene, a virulence gene, or a metabolic gene. In some embodiments, the one or more cells are bacterial cells.

[0036] Additionally provided are methods for modifying a target nucleic acid in a plant cell comprising providing to the plant, or a plant cell, seed, fruit, plant part, or propagation material of the plant a system, or a composition thereof, as described herein. In some embodiments, the system or composition inserts a donor nucleic acid within the target nucleic acid. In some embodiments, the donor nucleic acid comprises a gene product.

[0037] Other aspects and embodiments of the disclosure will be apparent in light of the following detailed description.BRIEF DESCRIPTION OF THE DRAWINGS[0038| FIGS. 1A-1E show integration of donor nucleic acids at therapeutically relevant genomic loci in human cells using the disclosed engineered systems (eeCAST, referred to in remaining figures as evoCAST) as compared to the corresponding wild-type system. FIG. 1A is a table showing the substitutions in TnsA, TnsB, TnsC, Cas7 and Cas8 in the disclosed systems, relative to the wild-type (WT) sequence. FIG. IB shows integration of a Ikb donor nucleic acid at various loci in HEK 293T cells. FIG. 1C shows the integration of a Ikb donor nucleic acid at safe harbor loci, as indicated. FIG. ID shows the integration of a Ikb donor nucleic acid in CAR- T engineering. FIG. IE shows integration of a Ikb donor nucleic acid at loci associated with loss of function diseases, as indicated.

[0039] FIGS. 2A-2G show development and characterization of evoCAST. FIG. 2A shows a schematic and genotypes of P4-15 TnsB and an engineered system (referred to in FIGS. 2-11 and Examples 2-6 as evoCAST) components. EvoCAST also contains optimized NLS architectures for Cas6, Cas8, and TniQ. FIG. 2B is a graph of 1-kb transposon integration efficiencies by evoCAST compared to P4-15 TnsB and wild-type (WT) PseCAST at four genomic sites in HEK293T cells. FIG. 2C is a graph of integration of varying DNA payload sizes (measured as the distance between the 3' end of the transposon right end and 5' end of the transposon left end) by WT PseCAST and evoCAST in HEK293T cells. Donor DNA transfected was normalized by mass. FIG. 2D is HTS analysis of the distance between the 3' end of the target site and 5' end ofthe transposon integration site for wild-type (WT) PseCAST and evoCAST across four genomic sites in HEK293T cells. FIG. 2E shows a comparison of indcl formation across untreated cells, wild-type (WT) PseCAST, and evoCAST at four genomic sites in HEK293T cells. Indels were quantified across a 40-bp window centered at the predicted insertion site for all unintegrated reads (see materials and methods). An unpaired, two-sided t-test was performed to determine statistical significance, with “ns” indicating a p-value > 0.05. FIG. 2F shows the relative frequencies of integration in the T-RL or T-LR orientation for evoCAST across four genomic sites in HEK293T cells, determined by ddPCR using probes specific to either T-RL or T-LR integration events. FIG. 2G shows genome-wide integration events for evoCAST (top) and a negative control (bottom) in which only pDonor was transfected, detected via a modified UDiTaS workflow. Integration events are measured by the number of unique molecular identifiers (UMIs) identified at a single integration site. The on-target genomic site (AA VSI) is indicated with a red triangle. The dotted line corresponds to a single detected integration event. Shown here is one of two replicates. Data in (FIGS. 2B-2F) are shown as mean+s.d. for n=3 independent biological replicates.

[0040] FIGS. 3A-3K show evoCAST mediates efficient DNA integration at therapeutically relevant endogenous genomic loci in multiple human cell types. FIG. 3 A is a graph of 1-kb transposon integration by wild-type (WT) PseCAST and evoCAST at 14 genomic loci in HEK293T cells. Each locus was targeted using a top-performing crRNA identified in FIG. 8, except HEK3, which did not undergo crRNA spacer optimization. FIG. 3B is the integration at AAVS1 in HEK293T cells using 1-kb transposons encoding end sequences engineered to be compatible with in-frame insertion into protein-coding genes. Engineered ends maintain open reading frames (ORFs), compared to the wild-type transposon end which contains stop codons in all three possible translation frames (FIG. 9). X-axis denotes the wild-type (WT) transposon end and the transposon end variants. Stop codons in TnsB binding sites were mutated based on previous studies of Tn6677 transposon ends and transposon end sequence conservation for Type I-F CASTs. Stop codons outside of TnsB binding sites were mutated to serine, which required a single point mutation and thus was thought to be a less perturbative sequence change. FIGS. 3C- 3E are schematics depicting evoCAST applications for integrating a F9 cDNA at ALB intron 1 (FIG. 3C), a CD19-targeted chimeric antigen receptor (CAR) at the 5' UTR of TRAC (FIG. 3D), and a cDNA encoding a healthy gene copy (Aexon 1) into intron 1 of a gene associated withpathogenic loss-of-function (FIG. 3E). FIGS. 3F-3H show the integration by wild-type (WT) PseCAST and cvoCAST of F9 cDNA into ALB intron 1 in HuH7 cells (FIG. 3F), CD19-CAR into the 5' UTR of TRAC in HEK293T cells (FIG. 3G), and wild-type cDNAs (Aexon 1) into intron 1 of their corresponding endogenous locus in HEK293T cells (FIG. 3H). FIG. 31 shows 1-kb transposon integration by wild-type (WT) eCAST and evoCAST at two genomic loci in HeLa and K562 cells. FIG. 3J shows 1-kb transposon integration by wild-type (WT) PseCAST and evoCAST at two genomic loci in primary human fibroblast cells isolated from an RDEB patient. FIG. 3K is a graph of the fold-change in integration efficiencies upon co-transfection with a plasmid expressing E. coli ClpX. The dotted line represents no change upon ClpX expression. Each datapoint represents the average fold-change in 1-kb transposon integration efficiency at one genomic target across three independent biological replicates. Data in FIGS. 3A, 3B, and 3F-3J are shown as mean+s.d. for n=3 independent biological replicates.[00411 FIG. 4 shows EvoCAST with plasmid and linear donor transposon topologies.EvoCAST integration of a 1-kb transposon at two genomic sites in HEK293T cells, transfected either as plasmid or linearized DNA. Linearized DNA was generated by PCR using primers containing four phosphorothioate linkages between the first five nucleotides at the 5' ends. Standard transfection conditions used a plasmid that encodes both the transposon and crRNA cassette (blue). The effects of linearizing both the transposon and the crRNA cassette (purple) and the transposon alone (red) were assessed. Encoding the transposon on linear DNA separately from the plasmid-encoded crRNA (red) resulted in lower integration efficiency, which may be due to increasing the number of components to be delivered for CAST activity. Encoding the crRNA cassette and transposon on the same DNA sequence may provide desired editing efficiency. Data are shown as mean+s.d for n=3 independent biological replicates.

[0042] FIG. 5 shows assessment of evoCAST formation of substitution mutations in HEK293T cells. Quantification of substitution mutations for evoCAST insertions across four genomic target sites within a 30-bp window at the genome-transposon right end junction. TSD, target site duplication. Data are shown as mean±s.d for n=3 independent biological replicates.

[0043] FIGS. 6A-6C show long-read sequencing of evoCAST product amplicons. FIG. 6A shows a previously established workflow to detect and quantify the distribution of simple and cointegrate integration events in human cells. Cells were transfected with a plasmid target containing the target sequence, and parallel enrichment PCRs were performed for both simpleinsertions (primer pair Pl and P2) and cointegrate insertions (primer pair Pl and P3). PCRs were then pooled, and long-read Nanoporc sequencing was performed. FIG. 6B is a standard curve of control transfections. Multiple ratios of mock simple and cointegrate insertion plasmids were transfected and analyzed as described in FIG. 6A. FIG. 6C is an assessment of wild-type (WT) PseCAST and evoCAST. Nuclease-dead TnsA (D71A mutant) conditions were included to generate elevated cointegrate formation. Values were calculated using the linear regression shown in FIG. 6B. Data in FIG. 6B are shown for n=2 independent biological replicates. Data in FIG. 6C are shown as mean±s.d for n=3 independent biological replicates.

[0044] FIGS. 7A-7H show additional characterization of evoCAST genome-wide integration specificity. FIG. 7A shows a modified UDiTaS workflow to detect genome-wide evoCAST integration events (see materials and methods). In FIG. 7B, CAST components were iteratively removed to determine the necessary components for off-target integration in HEK293T cells. A preliminary UDiTaS protocol was used, which led to an increased frequency of PCR artifacts detected as genomic integration events, as shown for the pDonor only condition (grey). All conditions in red were tested with Nl-5 TnsC. Off-target integration involves TnsA, TnsB, and TnsC, but not QCascade. FIG. 7C shows a comparison of off-target formation by wild-type (WT) and evolved TnsC tested with P4-15 TnsB, wild-type TnsA, and wild-type QCascade. As in FIG. 7B, the preliminary UDiTaS protocol led to increased background events, as shown for the pDonor control (grey). Off-target formation does not depend on TnsC identity. FIG. 7D shows UDiTaS-based detection of CAST integration events after HEK293T cells were passaged for one month with drug selection following transfection. FIG. 7E shows UDiTaS-based detection of CAST integration events in lysate from host E. coli encoding wild-type TnsA, TnsC, Cas6, Cas7, Cas8 and TniQ on a complementary plasmid to limit evolution to TnsB which was encoded as P4-15 TnsB on a selection phase either in a phage-assisted continuous evolution (PACE) lagoon or overnight Phage-assisted noncontinuous evolution (PANCE). Since both PACE and PANCE selections only expose E. coli to CAST activity on a timescale of hours, the high rate of on-target formation in these conditions suggests that the off-target events detected in FIG. 4G are the result of persistent expression of CAST components in HEK239T cells. FIG. 7F shows the average ATAC-seq signal for the 55 UDiTaS -detected evoCAST off-targets in HEK293T cells plotted on a histogram showing the distribution of average ATAC-seq signals for 20,000 randomly sampled sets of 55 genomic sites (see materials and methods). The averageATAC-seq signal for evoCAST off-target sites is higher than that of the majority of randomly sampled sets (one-tailed - value =0.13). ATAC-seq data was obtained from a previously published dataset for HEK293 cells. FIG. 7G is WebLogos generated from the UDiTaS-detected evoCAST off-targets in HEK293T cells for two independent replicates (n denotes the number of off-targets detected in each replicate). The genomic sequences upstream of the right end (RE) of the transposon were analyzed, revealing no enrichment of AT 15 rich sequences, which was previously found for type V-K off-target events. FIG. 7H is a graph assessing off-target formation in HEK293T cells by a previously established fluorescent reporter assay. Cells were transfected with evoTnsABC (P3-37 TnsA, P4-15 TnsB, and Nl-5 TnsC) and a pDonor encoding an mCherry expression cassette. Off-target integration yields mCherry positive cells, assessed by flow cytometry on day 14 post-transfection. EeBxbl with an atrP-containing donor was used as a positive control for off-target integration. Data shown in FIGS. 7B and 7C are shown mean for n=2 independent biological replicates. Data in FIGS. 7E and 7H are shown for n=3 independent biological replicates.

[0045] FIGS. 8A-8E show optimization of crRNA sequences for evoCAST applications in human cells. FIG. 8A is a preliminary assessment of crRNA spacer sequences across five genomic loci, measuring 1-kb transposon integration by evolved TnsB and TnsC (P4-15 TnsB and Nl-5 TnsC) or wild-type (WT) PseCAST. Spacer sequences were selected to avoid off- target human genomic sites with <5 mismatches to the target sequence, ignoring every sixth base which is flipped out of the crRNA:DNA heteroduplex by Cas7. While all other experiments in used the typical crRNA repeat structure, here the atypical crRNA repeat was used, based on a report finding that atypical crRNA sequence marginally improved integration efficiency for wild-type PseCAST with ClpX. Top-performing spacer sequences, nominated here via HTS- based quantification, were selected for follow-up experiments in FIG. 8B that were quantified via ddPCR. FIG. 8B shows a comparison between a typical (SEQ ID NO: 21) and an atypical (SEQ ID NO: 22) crRNA repeat in guiding 1-kb transposon integration at seven genomic sites in HEK293T cells using evolved TnsB and TnsC (P4-15 TnsB and Nl-5 TnsC). Typical crRNAs enabled higher editing than atypical crRNAs across all sites tested. It is suspected that the higher dynamic range afforded by the efficiencies of evolved CASTs enabled more significant differences between crRNA repeat architectures to be deduced than what had previously been observed with wild-type PseCAST. Asterisks below the x-axis indicate the crRNA protospacersequences that enabled the most efficient integration at each locus initially identified in FIG. 8A. FIG. 8C shows a comparison between 32 and 33-nt crRNA spacer sequences for guiding 1-kb transposon integration at five genomic sites in HEK293T cells using evolved TnsB and TnsC (P4-15 TnsB and Nl-5 TnsC). A previous study reported a marginal improvement in integration efficiency when using 33-nt spacers instead of 32-nt spacers for wild-type PseCAST with ClpX. Here, 32-nt spacers, used for all experiments, enabled equivalent or higher integration efficiencies than 33-nt spacers across all sites tested. FIG. 8D shows a preliminary assessment of crRNA spacer sequences across eight additional genomic loci, measuring 1-kb transposon integration by evolved TnsABC (P3-37 TnsA, P4-15 TnsB, and Nl-5 TnsC) or wild-type (WT) P.scCAST. Spacer sequences were selected as in FIG. 8A, to avoid off-target human genomic sites with <5 mismatches to the target sequence, ignoring every sixth base. Typical repeat, 32-nt spacer crRNAs were used. FIG. 8E shows ddPCR quantification of lysate from the experiment shown in FIG. 8D for top-performing crRNAs, nominated via HTS-based quantification, for evolved TnsABC. Asterisks indicate the crRNA protospacer sequences that enabled the most efficient integration at each locus. Data in FIGS. 8A and 8B are shown as mean for n=2 independent biological replicates, data in FIGS. 8C-8E are shown as mean±s.d for n-3 independent biological replicates. Integration efficiencies in FIGS. 8 A and 8D were determined via HTS quantification.

[0046] FIGS. 9A and 9B show a split schematic of the open reading frames (ORFs) encoded by the Tn7016 transposon right (FIG. 9A) and left (FIG. 9B) end sequences reading into the transposon cargo. All ORFs natively contain multiple stop codons (shown in grey), each of which was engineered via single point mutations (shown in red) to permit translation read- through. Stop codons within TnsB binding sites (highlighted in light blue) were mutated based on previous studies of Tn6677 transposon ends and transposon end sequence conservation for type I-F CASTs. Stop codons outside of TnsB binding sites were mutated to serine, which required a single point mutation and thus was thought to be a less perturbative sequence change. Transposon right end sequence is SEQ ID NO:23, with ORF 1 having amino acid sequences of SEQ ID NOs: 25-27, ORF 2 having amino acid sequences of SEQ ID NOs: 28-30 and HN, and ORF 3 having amino acid sequences of SEQ ID NOs: 31-32 and HKA. Transposon left end sequence is SEQ ID NO:24, with ORF 1 having amino acid sequences of SEQ ID NOs: 33-35and AYQ, ORF 2 having amino acid sequences of SEQ ID NOs: 36-39, and ORF 3 having amino acid sequences of SEQ ID NOs: 40-42 and KL and R.

[0047] FIGS. 10A-10F show persistence of evoCAST-edited HEK293T cells and generation of clonally integrated HEK293T cell lines. FIG. 10A is a graph measuring bulk editing efficiencies for evoCAST and eePASSIGE-treated HEK293T cells for 12 days post-transfection. FIG. 1 OB is a graph of the impact of E. coli ClpX on bulk editing efficiencies over time for evoCAST targeting two genomic loci. FIG. 10C shows bulk editing efficiencies after selecting for integration events using puromycin. HEK293T cells were treated with puromycin at various time points following transfection, and then harvested at day 14. FIG. 10D is a graph measuring bulk editing efficiencies for evoCAST-treated HEK293T cells for 34 days post-transfection. Integration of a splice acceptor-pwroP cassette at the A A VS1 site enables selection for cells with on-target integration. Following 30 days of selection, HEK293T cells from the experiment shown in FIG. 10D were split into parallel cultures with and without puromycin selection (FIG. 10E). After four days, cells were harvested to quantify bulk editing efficiencies. FIG. 10F shows analysis of integrated clonal lines isolated via sorting bulk HEK293T cell populations after puromycin selection. Each datapoint represents a colony that showed detectable integration via ddPCR. The number of colonies with detected integration, as well as the average observed editing efficiency, are marked above each biological replicate transfection. The dashed line represents 33%, corresponding to the expected efficiencies if a single allele in a triploid HEK293T genome contained an integrated transposon. Data in FIGS. 10A-10E are shown as mean for n=2 independent biological replicates.

[0048] FIGS. 11A-1 IB show detection of evoCAST-integrated transgene expression in human cells. ddPCR of cDNA generated from HuH7 cell lysate four days post-transfection comparing wild-type (WT) A’.scCAST and evoCAST integrating F9 cDNA (Aexon 1) into intron 1 of ALB. Gene expression was determined via a primer pair / probe specific to the ALB exon 1- F9 exon 2 junction (FIG. 11 A). Transgene expression was normalized to TBP expression. ddPCR of cDNA generated from HEK293T cell lysate four days post-transfection comparing wild-type (WT) PseCAST and evoCAST integrating MECP2 cDNA (Aexon 1) into intron 1 of MECP2. Gene expression was determined via a primer pair / probe specific to the MECP2 exon 1- exon 2 junction, with exon 2 of the integrated transgene recoded to prevent detection ofendogenous MECP2 expression (FIG. 1 IB). Transgene was normalized to TBP expression. Data in FIGS. 11A-1 IB arc shown as mcan±s.d for n=3 independent biological replicates.

[0049] FIG. 12 shows the relationship between evoCAST integration efficiency and chromatin accessibility in HEK293T cells. Correlation between evoCAST 1-kb transposon integration efficiencies and chromatin accessibility (determined via ATAC-seq) in HEK293T cells. Each data point represents a target site from FIG. 3A, with 14 target sites shown in total. Integration efficiency for each target site is the average of three independent biological replicates shown in FIG. 3A. Chromatin accessibility for each target site was determined by averaging the normalized ATAC-seq read density for the 1-kb window centered at the target site. Pearson’s correlation was used to determine the strength and significance of the relationship.

[0050] FIGS. 13A-13C show the impact of ClpX on integration efficiencies of wild-type CAST and evoCAST in multiple human cell types. FIGS. 13A and 13B show the results of 1-kb transposon integration at two genomic loci with or without ClpX co-delivery for wild-type (WT) PseCAST (FIG. 13A) and evoCAST (FIG. 13B). FIG. 13C shows bright-field microscopy images of primary human fibroblast cells taken 48 hours post-electroporation. While evoCAST without ClpX induced no overt cytotoxicity detected above the GFP control, ClpX co-delivery caused substantially reduced cell viability (loss of typical spindle-like morphology). Data in FIGS. 13A and 13B are shown as mean+s.d. for n=3 independent biological replicates.

[0051] FIGS. 14A-14H show the comparison of evoCAST and eePASSIGE in HEK293T cells. FIG. 14A results of a 2-kb DNA cargo integration across six genomic sites in HEK293T cells by evoCAST and eePASSIGE. Target sites are not directly matched between editing strategies due to the two techniques having different DNA targeting machineries. The DNA cargo sequences are identical except for the requisite recognition elements (transposon ends for CASTs, attB site for eePASSIGE). FIG. 14B shows the high-throughput sequencing (HTS) methods for determining CAST (left) and PASSIGE (right) product purity. An initial tagging step with unique molecular identifiers (UMIs) was employed for PASSIGE product detection to mitigate amplification bias from differing PCR amplicon 10 sizes. FIGS. 14C and 14D show product purities, detected via HTS, at four genomic sites in HEK293T cells for evoCAST (FIG. 14C) and eePASSIGE (FIG. 14D), using lysate from the experiment quantified in FIG. 14A. FIG. 14E is a schematic depicting integration of a linear donor substrate by CAST (top) and PASSIGE (bottom). PASSIGE-integrated products contain genomic double-strand breaks(DSBs), while CASTs do not. FIG. 14F shows HTS of a -400 bp amplicon spanning the predicted location of the PASSIGE-gcncratcd double-strand break following linear donor integration. Both evoCAST and eePASSIGE were used to integrate a 2-kb DNA cargo at TRAC in HEK293T cells. X-axis indicates the donor DNA topology used. Linear donor DNA was generated by PCR using primers containing four phosphorothioate linkages between the first five nucleotides at the 5' ends. To sequence the entire integration products of evoCAST and eePASSIGE, amplicons were generated for long-read Nanopore sequencing (FIG. 14G). Fl / Rl primers were used for amplicon 1, F2 / R2 primers were used for amplicon 2. FIG. 14H shows nanopore sequencing of evoCAST and eePASSIGE integration products at TRAC in HEK293T cells using the PCR amplicon strategy indicated in FIG. 14G, reported as indels detected above background (determined by sequencing of synthetic, mock-integrated fragments). Asterisks above indel peaks for evoCAST reflect integrations at the minor insertion sites (peaks appear upstream and downstream the integrated product due to the 5-bp target site duplication). The dashed line for eePASSIGE indicates the predicted double-strand break location following linear donor integration. Data in FIGS. 14A, 14C, 14D, and 14F are shown as mean+s.d. for n=3 independent biological replicates, data in FIG. 14H are shown as mean for n-3 independent biological replicates.

[0052] FIGS. 15A-15C show phage-assisted continuous evolution (PACE) of CRISPR- associated transposases (CASTs). FIG. 15A is an overview of RNA-guided DNA integration by type I-F CAST. DNA targeting is mediated by the CRISPR effector complex Cascade, comprising Cas6, Cas7, Cas8, and a CRISPR RNA (crRNA) complexed with the transposition protein TniQ (together referred to as QCascade). Target DNA-bound QCascade recruits the AAA+ ATPase TnsC, which subsequently recruits the heteromeric TnsA-TnsB transposase to catalyze excision of the transposon DNA and integration of the transposon at the target locus. FIG. 15B is an overview of PACE for CAST evolution. Selection phage (SP) encodes evolving CAST proteins. Host E. coli encode a selection circuit that links CAST integration to gill expression, which produces the essential phage protein pill. Production of pill enables SPs encoding active CAST proteins to replicate. PACE occurs in a fixed volume vessel (the ‘lagoon’) under constant dilution with fresh host E. coli, such that only SPs propagating faster than the rate of dilution can persist and evolve. FIG. 15C is a schematic of the anatomy of the initial CAST PACE selection circuit. SP encodes evolving transposase proteins TnsA-TnsB (an artificialfusion) and TnsC, while non-evolving CAST components are encoded on a complementary plasmid (CPI). Integration of a transposon provided on a second complementary plasmid (CP2) into a crRNA-specified target site on the accessory plasmid (AP) installs a promoter upstream of gill, resulting in gill expression and SP propagation. Replicating SPs accumulate mutations induced by a mutagenesis plasmid (MP), such that progeny SPs encode new CAST protein variants for selection in subsequent generations.

[0053] FIGS. 16A-16F show continuous evolution of TnsABC. FIG. 16A is a summary of TnsABC evolution campaign. Whether evolution segments were conducted using PANCE or PACE is specified, with PANCE passages or PACE hours indicated. Circuit architectures are described in FIG. 18. FIG. 16B shows the results of overnight phage propagation assays with wild-type (WT) TnsABC SP, pooled evolved SPs from each evolution segment, and glll- expressing phage (positive control for propagation). X-axes indicate host E. coli variants encoding circuit 1.0. Host A was used for PANCE Nl. Hosts B and C are of increased selection stringency, manipulated by reducing the promoter strength in the transposon on CP2 and reducing the ribosome binding site strength upstream of gill on the AP. Host NT A is host A with a non-targeting crRNA. The left graph shows phage propagation levels (output phage titer divided by input titer). The right graph shows transposon integration efficiencies at the AP target site in E. coli following overnight propagation, measured by qPCR. FIG. 16C is a chart of genotypes of a subset of evolved TnsABC variants. Variants Nl-1, Pl-3, and N2-1 showed the highest integration activity among the variants emerging from their respective PANCE or PACE experiments at two tested genomic sites in HEK293T cells. Variants P2-2, P2-7, and P2-11 are representative of the genotypes that emerged from P2. FIG. 16D shows 1-kb transposon integration at two genomic loci in HEK293T cells using wild-type (WT) and evolved TnsABC variants specified in FIG. 16C. FIGS. 16E and 16F are graphs assessing the contributions of P2- derived TnsAB and TnsC subunits to overnight phage propagation levels on P2 host E. coli (FIG. 16E) and 1-kb transposon integration efficiency in HEK293T cells (FIG. 16F). Data in FIGS. 16B and 16D-16F are shown as mean±s.d. for n=3 independent biological replicates.

[0054] FIGS. 17A-17F show that TnsAB- and TnsB-focused evolution generated transposase variants that support robust integration in human cells. FIG. 17A is a summary of the evolution campaign that yielded the evolved TnsB variant, P4-15, with the highest activity in HEK293T cells. Whether evolution segments were conducted in PANCE or PACE is specified, withPANCE passages or PACE hours indicated. Circuit architectures are described in FIG. 18. FIG. 17B is a chart of genotypes of top-performing TnsB variants from each evolution segment. FIG. 17C shows 1-kb transposon integration in HEK293T cells at two genomic sites by TnsB valiants shown in FIG. 17B. FIG. 17D shows fold-change in integration efficiencies upon co-transfection with a plasmid expressing E. coli ClpX. The dotted line represents no change upon ClpX expression. FIG. 17E shows he mutated residues in the P4-15 TnsB variant mapped onto an AlphaFold3 -predicted structure of a P.seTns AB tetramer complexed with a DNA substrate that mimics the product of TnsB transesterification. TnsA structures are omitted for clarity. Each transposon end (green) contains one full TnsB binding site that is joined to the 5' end of target DNA (blue). Eow-confidence unstructured C-termini of TnsB monomers (containing residues with pEDDT < 70) are not shown. The left image shows all mutated P4-15 residues in red, with the catalytic metal-coordinating DDE residues in TnsB.l and TnsB.3 shown in orange. The upper right image shows the mutated Y349 residue predicted to contact transposon DNA. The bottom right image shows multiple predicted TnsB’TnsB interfaces that contain mutated residues. FIG. 17F shows the mutated Q594 residue (red) in the P4-15 TnsB variant mapped onto an AlphaFold3-predicted structure of the P.ycTnsB C-terminal ‘hook’ domain in complex with a PseTnsC heptamer. Data in FIGS. 17C and 17D are shown as mean+s.d. for n=3 independent biological replicates. AlphaFold3-predicted structures in FIG. 17E and 17F are available on Zenodo.

[0055] FIGS. 18A-18F are schematics of the CAST PACE circuit architectures. FIG. 18A is PACE selection circuit 1.0 for TnsABC evolution. FIG. 18B is PACE selection circuit 1.1 for TnsABC evolution, designed to prevent SP from acquiring full-length gill during evolution. Circuit 1.1 introduces a target site on CPI and requires integration at both the AP and CPI target sites to produce full-length pill. Additionally, the crRNA cassette is moved from CPI to CP2 so undesired integration at the crRNA spacer (self-targeting) does not inhibit integration at the target site via target immunity. FIG. 18C is PACE selection circuit 1.2 for TnsABC evolution, designed to reduce selection stringency to enable evolution in PACE instead of PANCE. Circuit 1.2 introduces a signal amplification step on CPI such that integration at the CPI target site activates T7 RNA polymerase (T7 RNAP), which in turn transcribes the C-terminal segment of gill. Signal amplification was added to CPI instead of the AP because CPI is a lower copy plasmid (SC101 origin) than the AP (pl5A origin), thus the CPI-encoded gill segment wasassumed to be limiting for full-length pill production. FIG. 18D is PACE selection circuit 2.0 for TnsAB evolution, which encodes wild-type TnsC on CPI to restrict evolution to TnsAB. The AP size is increased to 10 kb (previously 3.5 kb) to prevent gill acquisition via AP cointegration or recombination into the SP genome. The transposon left end on CP2 contains a mutated binding site (denoted by an asterisk) for integration host factor to mitigate evolution of potential integration host factor-dependent fitness. FIG. 18E is PACE selection circuit 2.1 for TnsAB evolution, designed to more efficiently select for TnsAB variants that are highly active in human cells. Circuit 2.1 splits the artificial TnsA-TnsB fusion into its native monomeric forms to improve translatability from PACE fitness to human cell fitness. CPI encodes an evolved TnsC variant (Nl-5) identified as enabling the highest integration efficiencies in human cells among all tested TnsC variants. The AP size is increased to 15 kb to further prevent gill acquisition. The transposon size in CP2 is increased to 5 kb to introduce a new selection stringency by requiring mobilization of a larger DNA cargo. The crRNA cassette is encoded on CP2 instead of CPI to prevent self-targeting at the crRNA spacer. FIG. 18F is PACE selection circuit 3.0 for TnsB evolution, which encodes wild-type TnsA on CPI to limit evolution to TnsB.DETAILED DESCRIPTION

[0056] In bacteria and archaea, CRISPR / Cas systems provide immunity by incorporating fragments of invading phage, virus, and plasmid DNA into CRISPR loci and using corresponding CRISPR RNAs (“crRNAs”) to guide the degradation of homologous sequences. Transcription of a CRISPR locus produces a “pre-crRNA,” which is processed to yield crRNAs containing spacer-repeat fragments that guide effector nuclease complexes to cleave dsDNA sequences complementary to the spacer. Several different types of CRISPR systems are known, (e.g., type I, type II, or type III), and classified largely based on the Cas protein type and the use of a proto-spacer-adjacent motif (PAM) for selection of proto-spacers in invading DNA.

[0057] Although RNA-guided targeting typically leads to endonucleolytic cleavage of the bound substrate, recent studies have uncovered a range of noncanonical pathways in which CRISPR protein-RNA effector complexes have been naturally repurposed for alternative functions. For example, some Type I (Cascade) and Type II (Cas9) systems leverage truncated guide RNAs to achieve potent transcriptional repression without cleavage, and other Type I (Cascade) and Type V (Cas 12) systems lie inside unusual bacterial Tn7-like transposons and lack nuclease components altogether.

[0058] The present disclosure provides engineered CAST systems and methods for nucleic acid modification (c.g., RNA-guidcd DNA integration) utilizing engineered CRISPR-transposon systems comprising one or more engineered or evolved transposon-associated and / or Cas proteins. The disclosed systems can achieve approximately 10-30% targeted gene integration efficiency without enrichment. Across 14 genomic targets in human cells, the disclosed engineered CAST system efficiency represents a 420-fold average improvement over the corresponding wild-type system. An exemplary engineered CAST system showed enhanced activity across diverse cell types, notably enabling up to 38% editing in unsorted primary human fibroblasts. Overall, the DNA integration efficiencies (average = 12.7%) are sufficient in principle to treat a variety of genetic disorders. The engineered CAST systems provide distinct advantages as a genome editing platform, such as its facile reprogramming, high product purity, and avoidance of genomic double-strand break formation. As such, the disclosed engineered CAST systems enable efficient and targeted integration of therapeutically relevant genes at many genomic loci, paving the way for mutation-agnostic therapies for loss-of-function genetic diseases.

[0059] Section headings as used in this section and the entire disclosure herein are merely for organizational purposes and are not intended to be limiting.Definitions

[0060] The terms “comprise(s),” “include(s),” “having,” “has,” “can,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. As used herein, comprising a certain sequence or a certain SEQ ID NO usually implies that at least one copy of said sequence is present in recited peptide or polynucleotide. However, two or more copies are also contemplated. The singular' forms “a,” “and,” and “the” include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments “comprising,” “consisting of,” and “consisting essentially of,” the embodiments or elements presented herein, whether explicitly set forth or not.

[0061] For the recitation of numeric ranges herein, each intervening number there between with the same degree of precision is explicitly contemplated. For example, for the range of 6-9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0-7.0, the number 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly contemplated.

[0062] Unless otherwise defined herein, scientific, and technical terms used in connection with the present disclosure shall have the meanings that arc commonly understood by those of ordinary skill in the art. For example, any nomenclature used in connection with, and techniques of cell and tissue culture, molecular biology, genetics and protein and nucleic acid chemistry and hybridization described herein are those that are well known and commonly used in the art. The meaning and scope of the terms should be clear; in the event, however of any latent ambiguity, definitions provided herein take precedent over any dictionary or extrinsic definition. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular.

[0063] The term “contacting” as used herein refers to bring or put in contact, to be in or come into contact. The term “contact” as used herein refers to a state or condition of touching or of immediate or local proximity. Contacting a composition to a target destination, such as, but not limited to, an organ, tissue, cell, or tumor, may occur by any means of administration known to the skilled artisan.

[0064] The term “gene” refers to a DNA sequence that comprises control and coding sequences necessary for the production of an RNA having a non-coding function (e.g., a ribosomal or transfer RNA), a polypeptide, or a precursor of any of the foregoing. The RNA or polypeptide can be encoded by a full length coding sequence or by any portion of the coding sequence so long as the desired activity or function is retained. Thus, a “gene” refers to a DNA or RNA, or portion thereof, that encodes a polypeptide or an RNA chain that has functional role to play in an organism. For the purpose of this disclosure, it may be considered that genes include regions that regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites, and locus control regions.

[0065] A cell has been “genetically modified,” “transformed,” or “transfected” by exogenous DNA, e.g., a recombinant expression vector, when such DNA has been introduced inside the cell. The presence of the exogenous DNA results in permanent or transient genetic change. The transforming DNA may or may not be integrated (covalently linked) into the genome of the cell. For example, the transforming DNA may be maintained on an episomal element such as aplasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming DNA has become integrated into a chromosome so that it is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones that comprise a population of daughter cells containing the transforming DNA. A “clone” is a population of cells derived from a single cell or common ancestor by mitosis. A “cell line” is a clone of a primary cell that is capable of stable growth in vitro for many generations.

[0066] As used herein, the term “hybridization” is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization (e.g., the strength of the association between the nucleic acids) is influenced by such factors as the degree of complementary between the nucleic acids, stringency of the conditions involved, and the Tmof the formed hybrid. Hybridization methods involve the annealing of one nucleic acid to another, complementary nucleic acid, e.g., a nucleic acid having a complementary nucleotide sequence. The ability of two polymers of nucleic acid containing complementary sequences to find each other and “anneal” or “hybridize” through base pairing interaction is a well-recognized phenomenon. The initial observations of the “hybridization” process by Marmur and Lane, Proc. Natl. Acad. Sci. USA, 46: 453 (1960) and Doty et al., Proc. Natl. Acad. Sci. USA, 46: 461 (1960), have been followed by the refinement of this process into an essential tool of modern biology. For example, hybridization and washing conditions are now well known and exemplified in Sambrook et al., supra. The conditions of temperature and ionic strength determine the “stringency” of the hybridization.

[0067] As used herein, “nucleic acid” or “nucleic acid sequence” refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (See Albert L. Lehninger, Principles of Biochemistry, at 793- 800 (Worth Pub. 1982)). The present technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases, and the like. The polymers or oligomers may be heterogenous or homogenous in composition and may be isolated from naturally occurring sources or may be artificially or synthetically produced. In addition, the nucleic acids may be DNA or RNA, or a mixture thereof, and may exist permanently or transitionally in single- stranded or double- stranded form, including homoduplex, heteroduplex,and hybrid states. In some embodiments, a nucleic acid or nucleic acid sequence comprises other kinds of nucleic acid structures such as, for instance, a DNA / RNA helix, peptide nucleic acid (PNA), morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41(14): 4503-4510 (2002)) and U.S. Pat. No. 5,034,506), locked nucleic acid (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 97: 5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122: 8595-8602 (2000)), and / or a ribozyme. Hence, the term “nucleic acid” or “nucleic acid sequence” may also encompass a chain comprising non-natural nucleotides, modified nucleotides, and / or non- nucleotide building blocks that can exhibit the same function as natural nucleotides (e.g., “nucleotide analogs”); further, the term “nucleic acid sequence” as used herein refers to an oligonucleotide, nucleotide or polynucleotide, and fragments or portions thereof, and to DNA or RNA of genomic or synthetic origin, which may be single or doublestranded, and represent the sense or antisense strand. The terms “nucleic acid,” “polynucleotide,” “nucleotide sequence,” and “oligonucleotide” are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof.

[0068] Nucleic acid or amino acid sequence “identity,” as described herein, can be determined by comparing a nucleic acid or amino acid sequence of interest to a reference nucleic acid or amino acid sequence. A number of mathematical algorithms for obtaining the optimal alignment and calculating identity between two or more sequences are known and incorporated into a number of available software programs. Examples of such programs include CLUSTAL-W, T- Coffee, and ALIGN (for alignment of nucleic acid and amino acid sequences), BLAST programs (e.g., BLAST 2.1, BL2SEQ, and later versions thereof) and FASTA programs (e.g., FASTA3x, FAS™, and SSEARCH) (for sequence alignment and sequence similarity searches). Sequence alignment algorithms also are disclosed in, for example, Altschul et al., J. Molecular Biol., 215(3): 403-410 (1990), Beigert et al., Proc. Natl. Acad. Sci. USA, 706(10): 3770-3775 (2009), Durbin et al., eds., Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press, Cambridge, UK (2009), Soding, Bioinformatics, 27(7): 951- 960 (2005), Altschul et al., Nucleic Acids Res., 25(17): 3389-3402 (1997), and Gusfield, Algorithms on Strings, Trees and Sequences, Cambridge University Press, Cambridge UK (1997)).

[0069] The terms “non-naturally occurring,” “engineered,” and “synthetic” are used interchangeably and indicate the involvement of the hand of man. The terms, when referring to nucleic acid molecules or polypeptides mean that the nucleic acid molecule or the polypeptide is at least substantially free from at least one other component with which they are naturally associated in nature and as found in nature.

[0070] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein, and refer to a polymer of amino acid residues linked together by peptide bonds. The terms refer to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofamesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide may also be a single molecule or may be a multi-molecular complex. A protein, peptide, or polypeptide may be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, engineered, or synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.

[0071] As used herein, the terms “providing,” “administering,” and “introducing,” are used interchangeably herein and refer to the placement of the systems of the disclosure into a cell, organism, or subject by a method or route which results in at least partial localization of the system to a desired site. The systems can be administered by any appropriate route which results in delivery to a desired location in the cell, organism, or subject.

[0072] A “subject” or “patient” may be human or non-human and may include, for example, animal strains or species used as “model systems” for research purposes, such a mouse model asdescribed herein. Likewise, patient may include either adults or juveniles (e.g., children). Moreover, patient may mean any living organism, preferably a mammal (e.g., human or nonhuman) that may benefit from the administration of compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, nonhuman primates such as chimpanzees, and other apes and monkey species; farm animals such as cattle, horses, sheep, goats, swine; domestic animals such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice, guinea pigs, and the like. Examples of nonmammals include, but are not limited to, birds, fish, and the like. In one embodiment of the methods and compositions provided herein, the mammal is a human.

[0073] A “vector” or “expression vector” is a replicon, such as plasmid, phage, virus, or cosmid, to which another DNA segment, e.g., an “insert,” may be attached or incorporated so as to bring about the replication of the attached segment in a cell.[00741 Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in practice or testing of the present disclosure. All publications, patent applications, patents and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and not intended to be limiting.CAST Systems

[0075] Disclosed herein are systems for DNA integration into a target nucleic acid sequence comprising: an engineered Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CAST) system or one or more nucleic acids encoding the engineered CAST system. The engineered CAST system comprises at least one or both of: a) one or more Cas proteins selected from: Cas5, Cas6, Cas7, and Cas8 and b) one or more transposon- associated proteins selected from TnsA, TnsB, TnsC, and TniQ. In some embodiments, the engineered CAST system comprises a) one or more Cas proteins selected from: Cas5, Cas6, Cas7, and Cas8 and b) one or more transposon-associated proteins selected from TnsA, TnsB, TnsC, and TniQ. In some embodiments, the engineered CAST system comprises Cas5, Cas6, Cas7, and Cas8 and TnsA, TnsB, TnsC, and TniQ.

[0076] In some embodiments, the engineered CAST system comprises a Cas8-Cas5 fusion protein. In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having amino acid substitutions at positions 125, 244, and 410 relative to SEQ ID NO: 5. Insome embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having amino acid substitutions N125D, A244N, and A410R relative to SEQ ID NO: 5.

[0077] In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 225. In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence of SEQ ID NO: 225.

[0078] In some embodiments, the engineered CAST system comprises a Cas7 protein. In some embodiments, the Cas7 protein comprises an amino acid sequence having an amino acid sequence having an amino acid substitutions at position 347 relative to SEQ ID NO: 6. In some embodiments, the Cas7 protein comprises an amino acid sequence having an amino acid substitution A347K relative to SEQ ID NO: 6.[0079| In some embodiments, the Cas7 protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 224. In some embodiments, the Cas7 protein comprises an amino acid sequence of SEQ ID NO: 224.

[0080] In some embodiments, the engineered CAST system comprises a Cas8-Cas5 fusion protein and a Cas7 protein. In some embodiments, the engineered CAST system comprises a Cas8-Cas5 fusion protein, a Cas7 protein, and a Cas6 protein.

[0081] In some embodiments, the Cas6 protein comprises an amino acid sequence having at least 70% identity (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO: 7. In some embodiments, the Cas6 protein comprises an amino acid sequence of SEQ ID NO: 7.

[0082] In some embodiments, the engineered CAST system comprises TnsA, TnsB, and TnsC.

[0083] In some embodiments, the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, 180, and 182 relative to SEQ ID NO: 1. In some embodiments, the TnsA protein comprises an amino acid sequence having amino acid substitutions P88T, I147V, V170L, F180L, and F182L relative to SEQ ID NO: 1.

[0084] In some embodiments, the TnsA protein comprises protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%,at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 226. In some embodiments, the TnsA protein comprises an amino acid sequence of SEQ ID NO: 226.

[0085] In some embodiments, the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 43, 349, 352, 390, 396, 410, 464, 526, 549, and 594 relative to SEQ ID NO: 2. In some embodiments, the TnsB protein comprises an amino acid sequence having amino acid substitutions F43S, Y349N, P352T, A390V, D396N, Q410K, H464R, V526E, Q549R, and Q594L relative to SEQ ID NO: 2

[0086] In some embodiments, the TnsB protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 227. In some embodiments, the TnsB protein comprises an amino acid sequence of SEQ ID NO: 227.[0087| In some embodiments, the TnsC protein comprises an amino acid sequence having amino acid substitutions at positions 197 and 314 relative to SEQ ID NO: 3. In some embodiments, the TnsC protein comprises an amino acid sequence having amino acid substitutions R197I and N314K relative to SEQ ID NO: 3.

[0088] In some embodiments, the TnsC protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 228. In some embodiments, the TnsC protein comprises an amino acid sequence of SEQ ID NO: 228.

[0089] In some embodiments, the engineered CAST system comprises TnsA, TnsB, TnsC and TniQ. In some embodiments, the TniQ protein comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 4. In some embodiments, the TniQ protein comprises an amino acid sequence of SEQ ID NO: 4.

[0090] Any of the proteins described or referenced herein may comprise one or more additional amino acid substitutions. An amino acid “replacement” or “substitution” refers to the replacement of one amino acid at a given position or residue by another amino acid at the same position or residue within a polypeptide sequence. Amino acids are broadly grouped as “aromatic” or “aliphatic.” An aromatic amino acid includes an aromatic ring. Examples of “aromatic” amino acids include histidine (H or His), phenylalanine (F or Phe), tyrosine (Y orTyr), and tryptophan (W or Trp). Non-aromatic amino acids are broadly grouped as “aliphatic.” Examples of “aliphatic” amino acids include glycine (G or Gly), alanine (A or Ala), valine (V or Vai), leucine (L or Leu), isoleucine (I or He), methionine (M or Met), serine (S or Ser), threonine (T or Thr), cysteine (C or Cys), proline (P or Pro), glutamic acid (E or Glu), aspartic acid (D or Asp), asparagine (N or Asn), glutamine (Q or Gin), lysine (K or Lys), and arginine (R or Arg).

[0091] The amino acid replacement or substitution can be conservative, semi-conservative, or non-conservative. The phrase “conservative amino acid substitution” or “conservative mutation” refers to the replacement of one amino acid by another amino acid with a common property. A functional way to define common properties between individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms (Schulz and Schinner, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such analyses, groups of amino acids may be defined where amino acids within a group exchange preferentially with each other, and therefore resemble each other most in their impact on the overall protein structure (Schulz and Schirmer, supra). Examples of conservative amino acid substitutions include substitutions of amino acids within the sub-groups described above, for example, lysine for arginine and vice versa such that a positive charge may be maintained, glutamic acid for aspartic acid and vice versa such that a negative charge may be maintained, serine for threonine such that a free -OH can be maintained, and glutamine for asparagine such that a free -NH2 can be maintained. “Semi-conservative mutations” include amino acid substitutions of amino acids within the same groups listed above, but not within the same sub-group. For example, the substitution of aspartic acid for asparagine, or asparagine for lysine, involves amino acids within the same group, but different sub-groups. “Non-conservative mutations” involve amino acid substitutions between different groups, for example, lysine for tryptophan, or phenylalanine for serine, etc.

[0092] In the systems disclosed herein, any or all of the one or more Cas protein and the one or more transposon-associated protein comprise at least one nuclear localization sequence (NLS). For example, in some embodiments, each of the one or more Cas proteins and the one or more transposon-associated proteins (e.g., Cas6, Cas7, Cas8-Cas5 fusion protein, TnsA, TnsB (or TnsA-TnsB fusion protein described below), TnsC, and TniQ) comprise at least one NLS. In some embodiments, one or more of the Cas proteins and transposon-associated proteins comprise two or more (e.g., 2, 3, 4, 5 or more) nuclear localization sequences. In some embodiments, theCas7 protein comprises a single nuclear localization sequence. Tn some embodiments, the TnsA protein and the TnsB protein each comprise a single nuclear localization sequence. In some embodiments, the TnsA-TnsB fusion protein comprises a single nuclear localization sequence, as described further below. In some embodiments, the Cas6 protein, the Cas8-Cas5 fusion protein, and / or the TniQ protein each comprise two nuclear localization sequences. In some embodiments, the TnsC protein comprises two or more nuclear localization sequences.

[0093] The at least one nuclear localization sequence may be appended to at least one of the one or more Cas protein and the one or more transposon-associated protein at a N-terminus, a C- terminus, embedded in the protein (e.g., inserted internally within the open reading frame (ORF)), or a combination thereof. When more than one nuclear localization sequence is appended to the protein, they may be in the same or different orientation relative to the protein, (e.g., one at the N-terminus and one at the end terminus or both at the same terminus). When more than one nuclear localization sequence is appended to the same terminus of the protein, the individual nuclear localization sequences may be adjacent without separation or may be sequential, separated by one or more amino acids not part of the nuclear localization sequence (e.g., a GSG linker).

[0094] The nuclear localization sequence may comprise any amino acid sequence known in the art to functionally tag or direct a protein for import into a cell’s nucleus (e.g., for nuclear transport). Usually, a nuclear localization sequence comprises one or more positively charged amino acids, such as lysine and arginine.

[0095] In some embodiments, the NLS is a monopartite sequence. A monopartite NLS comprise a single cluster of positively charged or basic amino acids. In some embodiments, the monopartite NLS comprises a sequence of K-K / R-X-K / R, wherein X can be any amino acid. Exemplary monopartite NLSs include those from the SV40 large T-antigen, c-Myc, and TUS- proteins, as described elsewhere herein.

[0096] In some embodiments, the NLS is a bipartite sequence. Bipartite NLSs comprise two clusters of basic amino acids, separated by a spacer of about 9-12 amino acids. Exemplary bipartite NLSs include the NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK (SEQ ID NO: 17) and the NLS of EGL-13, MSRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 18). In some embodiments, the NLS comprises a bipartite SV40 NLS. In certain embodiments, the NLS comprises an amino acid sequence having at least 70% (e.g., having at least 75%, at least 80%, atleast 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) similarity to KRTADGSEFESPKKKRKV (SEQ ID NO: 19). In select embodiments, the NLS consists of an amino acid sequence of KRTADGSEFESPKKKRKV (SEQ ID NO: 19).

[0097] In some embodiments, the Cas7 protein comprises a single nuclear localization sequence. In some embodiments, the Cas7 protein comprises a nuclear localization sequence of SEQ ID NO: 19, or variant thereof. In select embodiments, the Cas7 protein comprises a nuclear localization sequence of SEQ ID NO: 19, or variant thereof, at the N terminus of the Cas7 protein sequence.

[0098] In some embodiments, the TnsA protein and the TnsB protein each comprise a single nuclear localization sequence. In some embodiments, the TnsA-TnsB fusion protein comprises a single nuclear localization sequence in the linker region between the TnsA protein and the TnsB protein. The NLS may be embedded within a linker sequence, such that it is flanked by additional amino acids. In some embodiments, the NLS is flanked on each end by at least a portion of a flexible linker. In some embodiments, the NLS is flanked on each end by a glycine rich region of the linker. In some embodiments, the TnsA-TnsB fusion protein comprises a single nuclear localization sequence of SEQ ID NO: 19 in the linker region between the TnsA protein and the TnsB protein.

[0099] In some embodiments, the Cas6 protein, the Cas8-Cas5 fusion protein, and / or the TniQ protein each comprise two nuclear localization sequences of SEQ ID NO: 19, or variant thereof. In some embodiments, the two nuclear localization sequences are separated by a linker (e.g., a GSG linker). In select embodiments, the Cas6 protein, the Cas8-Cas5 fusion protein, and the TniQ protein comprise two nuclear localization sequences of SEQ ID NO: 19, or variant thereof, at the N terminus of the protein (e.g., TniQ, Cas6, Cas8-Cas5 fusion protein) sequence.

[0100] In some embodiments, the TnsC protein comprises two or more nuclear localization sequences. In some embodiments, the TnsC protein comprises three nuclear localization sequences. In some embodiments, the nuclear localization sequences are separated by a linker (e.g., a GSG linker). In some embodiments, the TnsC protein comprises three nuclear localization sequences of SEQ ID NO: 19, or variant thereof. In select embodiments, the TnsC protein comprises three nuclear localization sequences of SEQ ID NO: 19, or variant thereof at the C terminus of the protein.

[0101] In some embodiments, the Cas8-Cas5 fusion protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 10; the Cas7 protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 9; the Cas6 protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 8; the TnsA protein and the TnsB protein are provided as a TnsA-TnsB fusion protein encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 12; the TnsC protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 13; and / or the TniQ protein is encoded by a nucleic acid sequence having at least 70% identity to SEQ ID NO: 11.

[0102] In some embodiments, the Cas8-Cas5 fusion protein is encoded by a nucleic acid sequence of SEQ ID NO: 10; the Cas7 protein is encoded by a nucleic acid sequence of SEQ ID NO: 9; the Cas6 protein is encoded by a nucleic acid sequence of SEQ ID NO: 8; the TnsA protein and the TnsB protein are provided as a TnsA-TnsB fusion protein encoded by a nucleic acid sequence of SEQ ID NO: 12; the TnsC protein is encoded by a nucleic acid sequence of SEQ ID NO: 13; and / or the TniQ protein is encoded by a nucleic acid sequence of SEQ ID NO: 11.

[0103] In some embodiments, the engineered CAST system comprises: a Cas8-Cas5 fusion protein is encoded by a nucleic acid sequence of SEQ ID NO: 10; a Cas7 protein is encoded by a nucleic acid sequence of SEQ ID NO: 9; a Cas6 protein is encoded by a nucleic acid sequence of SEQ ID NO: 8; a TnsA protein and a TnsB protein is provided as a TnsA-TnsB fusion protein encoded by a nucleic acid sequence of SEQ ID NO: 12; a TnsC protein is encoded by a nucleic acid sequence of SEQ ID NO: 13; and a TniQ protein is encoded by a nucleic acid sequence of SEQ ID NO: 11.

[0104] The protein components of the disclosed system (e.g., the Cas proteins or the transposon-associated proteins) may further comprise an epitope tag (e.g., 3xFLAG tag, an HA tag, a Myc tag, and the like). In some embodiments, the epitope tag may be adjacent, either upstream or downstream, to a nuclear localization sequence. The epitope tags may be at the N- terminus, a C-terminus, or a combination thereof of the corresponding protein.

[0105] In some embodiments at least one of the one or more Cas proteins and the one or more transposon-associated proteins are provided as a fusion protein. For example, at least one of the one or more Cas proteins and the one or more transposon-associated proteins may be in a fusion protein with another wild-type or engineered Cas protein or transposon-associated protein. Insome embodiments, at least two of the disclosed engineered Cas proteins or transposon- associated proteins may be linked in a fusion protein. In some embodiments, each of the one or more Cas proteins and the one or more transposon-associated proteins are provided as a single fusion protein.

[0106] In some embodiments, TnsA and TnsB are provided as a TnsA-TnsB fusion protein. TnsA and TnsB can be fused in any orientation: N-terminus to C-terminus; C-terminus to N- terminus; N-terminus to N-terminus; or C-terminus to C-terminus, respectively. Preferably the C-terminus of TnsA is fused to the N-terminus of TnsB.

[0107] In some embodiments, any of the fusion proteins (e.g., the TnsA-TnsB fusion) may be fused using an amino acid linker peptide of various lengths to provide greater physical separation and allow more spatial mobility between the fused portions. The linker may comprise any amino acids and may be of any length. In some embodiments, the linker may be less than about 50 (e.g., 40, 30, 20, 10, or 5) amino acid residues.

[0108] In some embodiments, the linker is a flexible linker, such that the individual proteins (e.g., TnsA and TnsB) can have orientation freedom in relationship to each other. For example, a flexible linker may include amino acids having relatively small side chains, and which may be hydrophilic. Without limitation, the flexible linker may contain a stretch of glycine and / or serine residues. In some embodiments, the linker comprises at least one glycine-rich region. For example, the glycine -rich region may comprise a sequence comprising a series of GS dipeptides, wherein the series may comprise between 1 and 10 dipeptides.

[0109] In some embodiments, the linker further comprises a nuclear localization sequence (NLS). The NLS may be embedded within a linker sequence, such that it is flanked by additional amino acids. In some embodiments, the NLS is flanked on each end by at least a portion of a flexible linker. In some embodiments, the NLS is flanked on each end by a glycine rich region of the linker. Suitable nuclear localization sequences for use with the disclosed system are described elsewhere herein and are applicable to use with the fusion proteins herein, e.g., TnsA- TnsB fusion protein.

[0110] In some embodiments, the systems may further comprise a guide RNA (gRNA) or a nucleic acid encoding a gRNA, wherein the gRNA is complementary to at least a portion of a target nucleic acid sequence. In some embodiments, one or more of the at least one Cas protein are part of a ribonucleoprotein (RNP) complex with the gRNA.

[0111] The gRNA may be a crRNA, crRNA / tracrRNA (or single guide RNA, sgRNA). The terms “gRNA,” “guide RNA,” “crRNA,” and “CRISPR guide sequence” may be used interchangeably throughout and refer to a nucleic acid comprising a sequence that determines the binding specificity of the CRISPR-Cas system. A gRNA hybridizes to (complementary to, partially or completely) a target nucleic acid sequence (e.g., the genome in a host cell). In some embodiments, the at least one gRNA is encoded in a CRISPR RNA (crRNA) array.

[0112] The gRNA or portion thereof that hybridizes to the target nucleic acid (a target site) may be any length. In some embodiments, the gRNA sequence that hybridizes to the target nucleic acid is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in length. gRNAs or sgRNA(s) used in the present disclosure can be between about 5 and 100 nucleotides long, or longer (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42,43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59 60, 61, 62, 63, 63, 64, 65, 66, 67,68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 92, 93,94, 95, 96, 97, 98, 99, or 100 nucleotides in length, or longer).

[0113] To facilitate gRNA design, many computational tools have been developed (See Prykhozhij et al. (PLoS ONE, 10(3): (2015)); Zhu et al. (PLoS ONE, 9(9) (2014)); Xiao et al. (Bioinformatics. Jan 21 (2014)); Heigwer et al. (Nat Methods, 11(2): 122-123 (2014)). Methods and tools for guide RNA design are discussed by Zhu (Frontiers in Biology, 10 (4) pp 289-296 (2015)), which is incorporated by reference herein. Additionally, there are many publicly available software tools that can be used to facilitate the design of sgRNA(s); including but not limited to, Genscript Interactive CRISPR gRNA Design Tool, WU-CRISPR, and Broad Institute GPP sgRNA Designer. There are also publicly available pre-designed gRNA sequences to target many genes and locations within the genomes of many species (human, mouse, rat, zebrafish, C. elegans), including but not limited to, IDT DNA Predesigned Alt-R CRISPR-Cas9 guide RNAs, Addgene Validated gRNA Target Sequences, and GenScript Genome- wide gRNA databases.

[0114] In addition to a sequence that binds to a target nucleic acid, in some embodiments, the gRNA may also comprise a scaffold sequence (e.g., tracrRNA). In some embodiments, such a chimeric gRNA may be referred to as a single guide RNA (sgRNA). Exemplary scaffold sequences will be evident to one of skill in the art and can be found, for example, in Jinek, et al.Science (2012) 337(6096):816-821 , and Ran, et al. Nature Protocols (2013) 8:2281-2308, incorporated herein by reference in their entireties.[0H5| In some embodiments, the gRNA sequence does not comprise a scaffold sequence and a scaffold sequence is expressed as a separate transcript. In such embodiments, the gRNA sequence further comprises an additional sequence that is complementary to a portion of the scaffold sequence and functions to bind (hybridize) the scaffold sequence.

[0116] The gRNA can comprise spacer sequence. The space sequence can be any length. In some embodiments, the space sequence is 30-40 nucleotides long (e.g., 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40).

[0117] In some embodiments, the gRNA sequence is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to a target nucleic acid. In some embodiments, the gRNA sequence is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to the 3’ end of the target nucleic acid (e.g., the last 5, 6, 7, 8, 9, or 10 nucleotides of the 3’ end of the target nucleic acid).

[0118] The gRNA may be a non-naturally occurring gRNA.

[0119] The system may further comprise a target nucleic acid. The terms “target sequence,” “target nucleic acid,” and “target site” (e.g., a “target genomic DNA sequence”) are used interchangeably herein to refer to a polynucleotide (nucleic acid, gene, chromosome, genome, etc.) to which a guide sequence (e.g., a synthetic guide RNA) is designed to have complementarity, wherein hybridization between the target sequence and a guide sequence promotes the formation of a CRISPR complex, provided sufficient conditions for binding exist. The target sequence and guide sequence need not exhibit complete complementarity, provided that there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex. A target sequence may comprise any polynucleotide, such as DNA or RNA. Suitable DNA / RNA binding conditions include physiological conditions normally present in a cell. Other suitable DNA / RNA binding conditions (e.g., conditions in a cell-free system) are known in the art.]0120] The target sequence may or may not be flanked by a protospacer adjacent motif (PAM) sequence. In certain embodiments, a nucleic acid-guided nuclease can only cleave a target sequence if an appropriate PAM is present, see, for example Doudna et al., Science, 2014,346(6213): 1258096, incorporated herein by reference. A PAM can be 5' or 3' of a target sequence. A PAM can be upstream or downstream of a target sequence. In one embodiment, the target sequence is immediately flanked on the 3' end by a PAM sequence. A PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. In certain embodiments, a PAM is between 2-6 nucleotides in length. The target sequence may or may not be located adjacent to a PAM sequence (e.g., PAM sequence located immediately 3' of the target sequence) (e.g., for Type I CRISPR / Cas systems). In some embodiments, e.g., Type I systems, the PAM is on the alternate side of the protospacer (the 5' end). Makarova et al. describes the nomenclature for all the classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology 13:722-736 (2015)). Guide structures and PAMs are described in by R. Barrangou (Genome Biol. 16:247 (2015)).

[0121] Non-limiting examples of the PAM sequences include: CC, CA, AG, GT, TA, AC, CA, GC, CG, GG, CT, TG, GA, AGG, TGG, T-rich PAMs (such as TTT, TTG, TTC, etc.), NGG, NGA, NAG, NGGNG and NNAGAAW (W=A or T), NNNNGATT, NAAR (R=A or G), NNGRR (R=A or G), NNAGAA, and NAAAAC, where N is any nucleotide. In some embodiments, the PAM may comprise a sequence of CN, in which N is any nucleotide. In select embodiments, the PAM may comprise a sequence of CC.

[0122] “Complementarity” refers to the ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence by either traditional Watson-Crick or other non-traditional types. A percent complementarity indicates the percentage of residues in a nucleic acid molecule, which can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence. Full complementarity is not necessarily required, provided there is sufficient complementarity to cause hybridization. There may be mismatches distal from the PAM.

[0123] The system may further include a donor nucleic acid. The donor nucleic acid may be a part of a bacterial plasmid, bacteriophage, a virus, autonomously replicating extra chromosomal DNA element, linear plasmid, linear DNA, linear covalently closed DNA, mitochondrial or other organellar DNA, chromosomal DNA, and the like. In some embodiments, the donor nucleic acid comprises a cargo nucleic acid sequence.

[0124] The donor nucleic acid may be flanked by at least one transposon end sequence. In some embodiments, the donor nucleic acid is flanked on the 5’ and the 3’ end with a transposon end sequence. The term “transposon end sequence” refers to any nucleic acid comprising asequence capable of forming a complex with the transposase enzymes thus designating the nucleic acid between the two ends for rearrangement. Usually, these sequences contain inverted repeats and may be about 10-150 base pairs long, however the exact sequence requirements differ for the specific transposase enzymes. Transposon end sequences are well known in the art. Transposon ends sequences may or may not include additional sequences that promote or augment transposition.

[0125] The transposon end sequences on either end may be the same or different. The transposon end sequence may be the endogenous CRISPR-transposon end sequences or may include deletions, substitutions, or insertions. The endogenous CRISPR-transposon end sequences may be truncated. In some embodiments, the transposon end sequence includes an about 40 base pair (bp) deletion relative to the endogenous CRISPR-transposon end sequence. In some embodiments, the transposon end sequence includes an about 100 base pair deletion relative to the endogenous CRISPR-transposon end sequence. The deletion may be in the form of a truncation at the distal (in relation to the cargo) end of the transposon end sequences.

[0126] The donor nucleic acid, and by extension the cargo nucleic acid, may of any suitable length, including, for example, about 50-100 bp (base pairs), about 100-1000 bp, at least or about 10 bp, at least or about 20 bp, at least or about 25 bp, at least or about 30 bp, at least or about 35 bp, at least or about 40 bp, at least or about 45 bp, at least or about 50 bp, at least or about 55 bp, at least or about 60 bp, at least or about 65 bp, at least or about 70 bp, at least or about 75 bp, at least or about 80 bp, at least or about 85 bp, at least or about 90 bp, at least or about 95 bp, at least or about 100 bp, at least or about 200 bp, at least or about 300 bp, at least or about 400 bp, at least or about 500 bp, at least or about 600 bp, at least or about 700 bp, at least or about 800 bp, at least or about 900 bp, at least or about 1 kb (kilobase pair), at least or about 2 kb, at least or about 3 kb, at least or about 4 kb, at least or about 5 kb, at least or about 6 kb, at least or about 7 kb, at least or about 8 kb, at least or about 9 kb, at least or about 10 kb, or greater.

[0127] In some embodiments, the system comprises components from or derived from different CAST systems. In some embodiments, at least one of the one or more Cas proteins and the one or more transposon-associated proteins may be derived from a homologous CAST system compared to the other protein components in the system.

[0128] In some embodiments, the system comprises two or more engineered CAST systems. Pairing of orthogonal systems with their orthogonal donor DNA substrates enables tandeminsertion of multiple distinct payloads directly adjacent to each other without any risk of repressive effects from target immunity. For example, one, two, three, four, five, or more orthogonal CAST systems may be used to integrate large tandem arrays of payload DNA. In some embodiments, multiple orthogonal RNA-guided transposases and their transposon donor DNAs may be integrated into distal regions of a given chromosome or genome, such that the lack of sequence identity between the transposon ends of the distinct transposon DNA substrates prevents genetic instability and the risk of recombination.

[0129] Sequences of exemplary Cas proteins, transposon-associated proteins, gRNAs, and transposon ends can also be found in International Patent Publications WO 2020 / 181264 and WO 2022 / 261122, incorporated herein by reference. However, the invention is not limited to the disclosed or referenced exemplary sequences. Indeed, genetic sequences can vary between different strains, and this natural scope of allelic variation is included within the scope of the invention.

[0130] The system may be a cell free system. Also disclosed is a cell comprising the system described herein. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell (e.g., a cell of a nonhuman primate or a human cell). Thus, in some embodiments, disclosed herein are systems or kits for DNA integration into a target nucleic acid sequence in a eukaryotic cell (e.g., a mammalian cell, a human cell).

[0131] The one or more nucleic acids encoding the engineered CAST system may be any nucleic acid including DNA, RNA, or combinations thereof. In some embodiments, the one or more nucleic acids comprise one or more messenger RNAs, one or more vectors, or any combination thereof.

[0132] The one or more Cas proteins, the one or more transposon-associated protein (e.g., TnsA, TnsB, TnsC, and TniQ), the at least one gRNA, and the donor nucleic acid may be on the same or different nucleic acids (e.g., vector(s)). In some embodiments, the one or more Cas proteins are encoded by a single nucleic acid. In some embodiments, the one or more transposon- associated proteins are encoded by a single nucleic acid. In some embodiments, the nucleic acid encoding the one or more Cas proteins also encodes the one or more transposon-associated proteins. In some embodiments, the one or more Cas proteins are encoded by a different nucleic acid from the one or more transposon-associated proteins.

[0133] In some embodiments, the at least one gRNA is encoded by a nucleic acid different from the nucleic acid(s) encoding the one or more Cas proteins and the one or more transposon- associated proteins. In some embodiments, the at least one gRNA is encoded by a nucleic acid also encoding at least one Cas protein, at least one transposon-associated protein, or both. In some embodiments, the one or more Cas proteins, the one or more transposon-associated proteins, and the at least one gRNA are encoded by a single nucleic acid. The gRNA may be encoded anywhere in the nucleic acid encoding the one or more Cas proteins or the one or more transposon-associated proteins. In some embodiments, the gRNA is encoded in the 3’ UTR of a protein coding nucleic acid.

[0134] In some embodiments, the nucleic acid encoding the one or more Cas proteins, the one or more transposon-associated protein, the at least one gRNA, or any combination thereof further comprises the donor nucleic acid.

[0135] The present systems may further include at least one unfoldase protein. Unfoldases are proteins that catalyze the unfolding of a native protein without affecting the primary structure. The unfoldase may be an NTP driven unfoldase. NTP driven unfoldases may include ATP- dependent proteases, including, but not limited to, ATPases, AAA proteases, or AAA+ enzymes (e.g., AAA+ enzyme). In some embodiments, the at least one unfoldase protein may comprise ClpX (caseinolytic mitochondrial matrix peptidase chaperone subunit X). In some embodiments, the at least one unfoldase protein may comprise a homolog of ClpX.

[0136] ClpX homologs may be readily screened through systematic testing and optimization of a large panel of homologs, identified through bioinformatic search strategies such as BLASTp and psi-BLASTp. In some embodiments, the unfoldase protein (e.g., ClpX) is derived from the same host organism as that of the engineered CAST system. In some embodiments, the unfoldase protein (e.g., ClpX) is derived from a different host organism as that of the engineered CAST system. As such, the at least one unfoldase protein (e.g., ClpX) is not limited from which organism it is derived. In some embodiments, the unfoldase protein (e.g., ClpX) is derived from the E. coll genome. In other embodiments, the unfoldase protein (e.g., ClpX) from the cognate strain from which the engineered CAST system is derived. For example, the unfoldase protein from Vibrio cholerae HE-45 can be used alongside RNA-guided DNA integration machinery derived from Tn6677, while unfoldase proteins from Pseudoalteromonas sp. S983 can be used alongside RNA-guided DNA integration machinery derived from Tn7016.

[0137] In some embodiments, the systems further comprise one or more additional genome engineering tools. For example, the systems may further comprise nucleases, such as zinc finger nucleases (ZFNs) and / or transcription activator like effector nucleases (TALENs); transcriptional activators, transcriptional repressors, histone-modifying proteins, integrases, and recombinases.Nucleic Acids and Delivery

[0138] The present disclosure also provides for nucleic acids encoding the systems or components thereof and vectors containing or encoding these nucleic acids. The vectors may be used to propagate the nucleic acid in an appropriate cell and / or to allow expression from the nucleic acid (e.g., an expression vector). The person of ordinary skill in the art would be aware of the various vectors available for propagation and expression of a nucleic acid sequence.

[0139] The present disclosure further provides engineered, non-naturally occurring vectors and vector systems, which can encode one or more of the peptides or components of the present systems. The vector(s) can be introduced into a cell that is capable of expressing the polypeptide encoded thereby, including any suitable prokaryotic or eukaryotic cell.

[0140] The vectors of the present disclosure may be delivered to a eukaryotic cell in a subject. Modification of eukaryotic cells via the present system can take place in a cell culture, where the method comprises isolating the eukaryotic cell from a subject prior to the modification. In some embodiments, the method further comprises returning said eukaryotic cell and / or cells derived therefrom to the subject.

[0141] Viral and non-viral based gene transfer methods can be used to introduce nucleic acids encoding the disclosed polypeptides or components of the present system into cells, tissues, or a subject. Such methods can be used to administer nucleic acids encoding the disclosed polypeptides or components of the present system to cells in culture, or in a host organism. Non- viral vector delivery systems include DNA plasmids, cosmids, RNA (e.g., a transcript of a vector described herein), a nucleic acid, and a nucleic acid complexed with a delivery vehicle. Viral vector delivery systems include DNA and RNA viruses, which have either episomal or integrated genomes after delivery to the cell. Viral vectors include, for example, retroviral, lentiviral, adenoviral, adeno-associated and herpes simplex viral vectors.

[0142] In certain embodiments, plasmids that are non-replicative, or plasmids that can be cured by high temperature may be used, such that any or all of the necessary components of the system may be removed from the cells under certain conditions. For example, this may allow forDNA integration by transforming bacteria of interest, but then being left with engineered strains that have no memory of the plasmids or vectors used for the integration. Drug selection strategics may be adopted for positively selecting for cells. A donor nucleic acid may contain one or more drug- selectable markers within the cargo. Then presuming that the original donor plasmid is removed, drug selection may be used to enrich for integrated clones. Colony screenings may be used to isolate clonal events.

[0143] A variety of viral constructs may be used to deliver the disclosed polypeptides or components of the present system (such as one or more Cas proteins and / or transposon- associated proteins, gRNA(s), donor DNA, etc.) to the targeted cells and / or a subject. Nonlimiting examples of such recombinant viruses include recombinant adeno-associated virus (AAV), recombinant adenoviruses, recombinant lentiviruses, recombinant retroviruses, recombinant herpes simplex viruses, recombinant poxviruses, phages, etc. The present disclosure provides vectors capable of integration in the host genome, such as retrovirus or lentivirus. See, e.g., Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1989; Kay, M. A., et al., 2001 Nat. Medic. 7(1 ):33-40; and Walther W. and Stein U., 2000 Drugs, 60(2): 249-71, incorporated herein by reference.

[0144] In one embodiment, a nucleic acid encoding the disclosed polypeptides or components of the present system is contained in a plasmid vector that allows expression of the disclosed polypeptides or components of the present system and subsequent isolation and purification of from the recombinant vector. Accordingly, the disclosed polypeptides or components of the present system disclosed herein can be purified following expression, obtained by chemical synthesis, or obtained by recombinant methods.

[0145] To construct cells that express the disclosed polypeptides or components of the present system, expression vectors for stable or transient expression of the disclosed polypeptides or components of the present system may be constructed via conventional methods as described herein and introduced into host cells. For example, nucleic acids encoding the components of the disclosed polypeptides or components of the present system may be cloned into a suitable expression vector, such as a plasmid or a viral vector in operable linkage to a suitable promoter. The selection of expression vectors / plasmids / viral vectors should be suitable for integration and replication in eukaryotic cells.

[0146] In certain embodiments, vectors of the present disclosure can drive the expression of one or more sequences in prokaryotic cells. Promoters that may be used include T7 RNA polymerase promoters, constitutive E. coli promoters, and promoters that could be broadly recognized by transcriptional machinery in a wide range of bacterial organisms. The system may be used with various bacterial hosts.

[0147] In certain embodiments, vectors of the present disclosure can drive the expression of one or more sequences in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, Nature (1987) 329:840, incorporated herein by reference) and pMT2PC (Kaufman, et al., EMBO J. (1987) 6:187, incorporated herein by reference). When used in mammalian cells, the expression vector's control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other suitable expression systems for both prokaryotic and eukaryotic cells see, e.g., Chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL. 2nd eds., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989, incorporated herein by reference.

[0148] Vectors of the present disclosure can comprise any of a number of promoters known to the art, wherein the promoter is constitutive, regulatable or inducible, cell type specific, tissuespecific, or species specific. In addition to the sequence sufficient to direct transcription, a promoter sequence of the invention can also include sequences of other regulatory elements that are involved in modulating transcription (e.g., enhancers, Kozak sequences and introns). Many promoter / regulatory sequences useful for driving constitutive expression of a gene arc available in the art and include, but are not limited to, for example, CMV (cytomegalovirus promoter), EFla (human elongation factor 1 alpha promoter), SV40 (simian vacuolating virus 40 promoter), PGK (mammalian phosphoglycerate kinase promoter), Ubc (human ubiquitin C promoter), human beta-actin promoter, rodent beta-actin promoter, CBh (chicken beta-actin promoter), CAG (hybrid promoter contains CMV enhancer, chicken beta actin promoter, and rabbit betaglobin splice acceptor), TRE (Tetracycline response element promoter), Hl (human polymerase III RNA promoter), U6 (human U6 small nuclear promoter), and the like. Additional promoters that can be used for expression of the components of the present system, include, without limitation, cytomegalovirus (CMV) intermediate early promoter, a viral LTR such as the Roussarcoma virus LTR, HIV-LTR, HTLV-1 LTR, Maloney murine leukemia virus (MMLV) LTR, mycoloprolifcrativc sarcoma virus (MPSV) LTR, spleen focus-forming virus (SFFV) LTR, the simian virus 40 (SV40) early promoter, herpes simplex tk virus promoter, elongation factor 1- alpha (EFl -a) promoter with or without the EFl -a intron. Additional promoters include any constitutively active promoter. Alternatively, any regulatable promoter may be used, such that its expression can be modulated within a cell.

[0149] Moreover, inducible and tissue specific expression of a RNA, transmembrane proteins, or other proteins can be accomplished by placing the nucleic acid encoding such a molecule under the control of an inducible or tissue specific promoter / regulatory sequence. Examples of tissue specific or inducible promoter / regulatory sequences which are useful for this purpose include, but arc not limited to, the rhodopsin promoter, the MMTV LTR inducible promoter, the SV40 late enhancer / promoter, synapsin 1 promoter, ET hepatocyte promoter, GS glutamine synthase promoter and many others. Various commercially available ubiquitous as well as tissue- specific promoters and tumor-specific are available, for example from InvivoGen. In addition, promoters which are well known in the art can be induced in response to inducing agents such as metals, glucocorticoids, tetracycline, hormones, and the like, are also contemplated for use with the invention. Thus, it will be appreciated that the present disclosure includes the use of any promoter / regulatory sequence known in the art that is capable of driving expression of the desired protein operably linked thereto.

[0150] The vectors of the present disclosure may direct expression of the nucleic acid in a particular cell type (e.g., tissue-specific regulatory elements are used to express the nucleic acid). Such regulatory elements include promoters that may be tissue specific or cell specific. The term “tissue specific” as it applies to a promoter refers to a promoter that is capable of directing selective expression of a nucleotide sequence of interest to a specific type of tissue (e.g., seeds) in the relative absence of expression of the same nucleotide sequence of interest in a different type of tissue. The term “cell type specific” as applied to a promoter refers to a promoter that is capable of directing selective expression of a nucleotide sequence of interest in a specific type of cell in the relative absence of expression of the same nucleotide sequence of interest in a different type of cell within the same tissue. The term “cell type specific” when applied to a promoter also means a promoter capable of promoting selective expression of a nucleotidesequence of interest in a region within a single tissue. Cell type specificity of a promoter may be assessed using methods well known in the art, c.g., immunohistochemical staining.[01511 Additionally, the vector may contain, for example, some or all of the following: a selectable marker gene, such as the neomycin gene for selection of stable or transient transfectants in host cells; enhancer / promoter sequences from the immediate early gene of human CMV for high levels of transcription; transcription termination and RNA processing signals from SV40 for mRNA stability; 5’-and 3 ’-untranslated regions for mRNA stability and translation efficiency from highly-expressed genes like a-globin or 0-globin; SV40 polyoma origins of replication and ColEl for proper episomal replication; internal ribosome binding sites (IRESes), versatile multiple cloning sites; T7 and SP6 RNA promoters for in vitro transcription of sense and antisense RNA; a “suicide switch” or “suicide gene” which when triggered causes cells carrying the vector to die (e.g., HSV thymidine kinase, an inducible caspase such as iCasp9), and reporter gene for assessing expression of the chimeric receptor. Suitable vectors and methods for producing vectors containing transgenes are well known and available in the art. Selectable markers also include chloramphenicol resistance, tetracycline resistance, spectinomycin resistance, streptomycin resistance, erythromycin resistance, rifampicin resistance, bleomycin resistance, thermally adapted kanamycin resistance, gentamycin resistance, hygromycin resistance, trimethoprim resistance, dihydrofolate reductase (DHFR), GPT; the URA3, HIS4, LEU2, and TRP1 genes of S. cerevisiae.

[0152] When introduced into the cell, the vectors may be maintained as an autonomously replicating sequence or extrachromosomal element or may be integrated into host DNA.

[0153] In one embodiment, the donor DNA may be delivered using the same gene transfer system as used to deliver the Cas protein, and / or transposon-associated proteins (included on the same vector) or may be delivered using a different delivery system. In another embodiment, the donor DNA may be delivered using the same transfer system as used to deliver gRNA(s).

[0154] In one embodiment, the present disclosure comprises integration of exogenous DNA into the endogenous gene. Alternatively, an exogenous DNA is not integrated into the endogenous gene. The DNA may be packaged into an extrachromosomal or episomal vector (such as AAV vector), which persists in the nucleus in an extrachromosomal state, and offers donor-template delivery and expression without integration into the host genome. Use ofextrachromosomal gene vector technologies has been discussed in detail by Wade-Martins R (Methods Mol Biol. 2011; 738:1-17, incorporated herein by reference).[0155| The disclosed polypeptides or components of the present system (e.g., proteins, polynucleotides encoding these proteins, donor polynucleotides and compositions comprising the proteins and / or polynucleotides described herein) may be delivered by any suitable means. In certain embodiments, the polypeptides or system is delivered in vivo. In other embodiments, the polypeptides or system is delivered to isolated / cultured cells (e.g., autologous iPS cells) in vitro to provide modified cells useful for in vivo delivery to patients afflicted with a disease or condition.

[0156] Vectors according to the present disclosure can be transformed, transfected, or otherwise introduced into a wide variety of cells. Transfection refers to the taking up of a vector by a cell whether or not any coding sequences are in fact expressed. Numerous methods of transfection are known to the ordinarily skilled artisan, for example, lipofectamine, calcium phosphate co-precipitation, electroporation, DEAE-dextran treatment, microinjection, viral infection, and other methods known in the art. Transduction refers to entry of a virus into the cell and expression (e.g., transcription and / or translation) of sequences delivered by the viral vector genome. In the case of a recombinant vector, “transduction” generally refers to entry of the recombinant viral vector into the cell and expression of a nucleic acid of interest delivered by the vector genome.

[0157] Any of the vectors comprising a nucleic acid sequence that encodes the disclosed polypeptides or components of the present system is also within the scope of the present disclosure. Such a vector may be delivered into host cells by a suitable method. Methods of delivering vectors to cells are well known in the art and may include DNA or RNA electroporation, transfection reagents such as liposomes or nanoparticles to delivery DNA or RNA; delivery of DNA, RNA, or protein by mechanical deformation (see, e.g., Sharei et al. Proc. Natl. Acad. Sci. USA (2013) 110(6): 2082-2087, incorporated herein by reference); or viral transduction. In some embodiments, the vectors are delivered to host cells by viral transduction. Nucleic acids can be delivered as part of a larger construct, such as a plasmid or viral vector, or directly, e.g., by electroporation, lipid vesicles, viral transporters, microinjection, and biolistics (high-speed particle bombardment). Similarly, the construct can be delivered by any method appropriate for introducing nucleic acids into a cell. In some embodiments, the construct or thenucleic acid encoding the disclosed polypeptides or components of the present system is a DNA molecule. In some embodiments, the nucleic acid encoding the disclosed polypeptides or components of the present system is a DNA vector and may be electroporated to cells. In some embodiments, the nucleic acid encoding the disclosed polypeptides or components of the present system is an RNA molecule, which may be electroporated to cells.

[0158] Additionally, delivery vehicles such as nanoparticle- and lipid-based mRNA or protein delivery systems can be used. Further examples of delivery vehicles include lentiviral vectors, ribonucleoprotein (RNP) complexes, lipid-based delivery system, gene gun, hydrodynamic, electroporation or nucleofection microinjection, and biolistics. Various gene delivery methods are discussed in detail by Nayerossadat et al. (Adv Biomed Res. 2012; 1: 27) and Ibraheem et al. (Int J Pharm. 2014 Jan 1; 459( 1-2) :70-83), incorporated herein by reference. In some embodiments, the system is delivered, at least in part, as a ribonucleoprotein (RNP) complex comprising any or all of the one or more Cas proteins and one or more transposon-associated proteins and a gRNA.Methods of Use

[0159] Also disclosed herein are methods for nucleic acid modification or integration utilizing the disclosed system, or a composition thereof. The methods may comprise contacting a target nucleic acid sequence with a system, or a composition thereof, disclosed herein. The descriptions and embodiments provided above for the systems are applicable to the methods described herein.

[0160] The phrase “modifying a nucleic acid sequence” or “nucleic acid modification” as used herein, refers to modifying at least one physical feature of a nucleic acid sequence of interest. Nucleic acid modifications include, for example, single or double strand breaks, deletion, or insertion of one or more nucleotides, and other modifications that affect the structural integrity or nucleotide sequence of the nucleic acid sequence.

[0161] The target nucleic acid sequence may be in a cell. In some embodiments, contacting a target nucleic acid sequence comprises introducing the system, composition, or polypeptide into the cell. As described above the system, composition, or polypeptide may be introduced into eukaryotic or prokaryotic cells by methods known in the art. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.

[0162] In some embodiments, the target nucleic acid is a nucleic acid endogenous to a target cell. In some embodiments, the target nucleic acid is a genomic DNA sequence. The term“genomic,” as used herein, refers to a nucleic acid sequence (e.g., a gene or locus) that is located on a chromosome in a cell.[01631 In some embodiments, the target nucleic acid encodes a gene or gene product. The term “gene product,” as used herein, refers to any biochemical product resulting from expression of a gene. Gene products may be RNA or protein. RNA gene products include non-coding RNA, such as tRNA, rRNA, microRNA (miRNA), and small interfering RNA (siRNA), and coding RNA, such as messenger RNA (mRNA). In some embodiments, the target nucleic acid sequence encodes a protein or polypeptide. In some embodiments, the methods can be used for in-frame tagging of a protein gene product, e.g., tagging endogenous proteins.

[0164] Polynucleotides containing the target nucleic acid sequence may include, but are not limited to, purified chromosomal DNA, total cDNA, cDNA fractionated according to tissue or expression state (e.g., after heat shock or after cytokine treatment other treatment) or expression time (after any such treatment) or developmental stage, plasmid, cosmid, BAC, YAC, phage library, etc. Polynucleotides containing the target site may include DNA from organisms such as Homo sapiens, Mas domesticus, Mus spretus, Canis domesticus, Bos, Caenorhabditis elegans, Plasmodium falciparum, Plasmodium vivax, Onchocerca volvulus, Brugia malayi, Dirofilaria immitis, Leishmania, Zea maize, Arabidopsis thaliana, Glycine max, Drosophila melanogaster, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Neurospora, Escherichia coli, Salmonella typhimurium, Bacillus subtilis, Neisseria gonorrhoeae, Staphylococcus aureus, Streptococcus pneumonia, Mycobacterium tuberculosis, Aquifex, Thermus aquaticus, Pyrococcus furiosus, Thermus littoralis, Methanobacterium thermoautotrophicum, Sulfolobus caldoaceticus, and others.

[0165] The method may comprise administering to the subject, in vivo, or by transplantation of ex vivo treated cells, an effective amount of the described system, composition, or polypeptide. In some embodiments, the vector(s) is delivered to the tissue of interest by, for example, an intramuscular, intravenous, transdermal, intranasal, oral, mucosal, or other delivery methods.

[0166] The polypeptides, composition, components of the present system, or ex vivo treated cells may be administered with a pharmaceutically acceptable carrier or excipient as a pharmaceutical composition. In some embodiments, the polypeptides, composition, or components of the present system may be mixed, individually or in any combination, with apharmaceutically acceptable carrier to form pharmaceutical compositions, which are also within the scope of the present disclosure.

[0167] In some embodiments, an effective amount of the polypeptides, components of the present system, or compositions as described herein can be administered. As used herein the term “effective amount” may be used interchangeably with the term “therapeutically effective amount” and refers to that quantity that is sufficient to result in a desired activity upon administration to a subject in need thereof. Within the context of the present disclosure, the term “effective amount” refers to that quantity of the components of the system such that successful DNA integration is achieved.

[0168] When utilized as a method of treatment, the effective amount may depend on the particular condition being treated, the severity of the condition, the individual patient parameters including age, physical condition, size, gender and weight, the duration of the treatment, the nature of concurrent therapy (if any), the specific route of administration and like factors within the knowledge and expertise of the health practitioner. In some embodiments, the effective amount alleviates, relieves, ameliorates, improves, reduces the symptoms, or delays the progression of any disease or disorder in the subject. In some embodiments, the subject is a human.

[0169] In the context of the present disclosure insofar as it relates to any of the disease conditions recited herein, the terms “treat,” “treatment,” and the like mean to relieve or alleviate at least one symptom associated with such condition, or to slow or reverse the progression of such condition. Within the meaning of the present disclosure, the term “treat” also denotes to arrest, delay the onset (e.g., the period prior to clinical manifestation of a disease) and / or reduce the risk of developing or worsening a disease. For example, in connection with cancer the term “treat” may mean eliminate or reduce a patient's tumor burden, or prevent, delay, or inhibit metastasis, etc.

[0170] The phrase “pharmaceutically acceptable,” as used in connection with compositions and / or cells of the present disclosure, refers to molecular entities and other ingredients of such compositions that are physiologically tolerable and do not typically produce untoward reactions when administered to a subject (e.g., a mammal, a human). Preferably, as used herein, the term “pharmaceutically acceptable” means approved by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia foruse in mammals, and more particularly in humans. “Acceptable” means that the carrier is compatible with the active ingredient of the composition (e.g., the nucleic acids, vectors, cells, or therapeutic antibodies) and does not negatively affect the subject to which the composition(s) are administered. Any of the pharmaceutical compositions and / or cells to be used in the present methods can comprise pharmaceutically acceptable carriers, excipients, or stabilizers in the form of lyophilized formations or aqueous solutions.

[0171] Pharmaceutically acceptable carriers, including buffers, are well known in the art, and may comprise phosphate, citrate, and other organic acids; antioxidants including ascorbic acid and methionine; preservatives; low molecular weight polypeptides; proteins, such as serum albumin, gelatin, or immunoglobulins; amino acids; hydrophobic polymers; monosaccharides; disaccharides; and other carbohydrates; metal complexes; and / or non-ionic surfactants. See, e.g., Remington: The Science and Practice of Pharmacy 20th Ed. (2000) Lippincott Williams and Wilkins, Ed. K. E. Hoover.

[0172] The methods may be used for a variety of purposes. For example, the methods may include, but are not limited to, inactivation of a microbial gene, RNA-guided DNA integration in a plant or animal cell, methods of treating a subject suffering from a disease or disorder (e.g., cancer, Duchenne muscular dystrophy (DMD), sickle cell disease (SCD), [3-thalassemia, and hereditary tyrosinemia type I (HT1)), and methods of treating a diseased cell (e.g., a cell deficient in a gene which causes cancer).

[0173] The disclosed methods may modify a target DNA sequence in a cell so as to modulate expression of the target DNA sequence, e.g., expression of the target DNA sequence is increased, decreased, or completely eliminated (e.g., via deletion of a gene). The modifications of the target sequence may lead to, for example, gene correction, gene replacement, gene tagging, transgene insertion, nucleotide deletion, gene disruption, gene mutation, gene knockdown, gene tagging, etc.

[0174] In some embodiments, the methods described herein may be used to correct one or more defects or mutations in a gene (referred to as “gene correction”). In such cases, the target sequence encodes a defective version of a gene, and the disclosed compositions and systems further comprise a donor nucleic acid molecule which encodes a wild-type or corrected version of the gene. Accordingly, in some embodiments, the methods described herein may be used to insert a gene or fragment thereof into a cell.

[0175] In another embodiment, the method of modifying a target sequence can be used to delete nucleic acids from a target sequence in a host cell by cleaving the target sequence and allowing the host cell to repair the cleaved sequence in the absence of an exogenously provided donor nucleic acid molecule. Deletion of a nucleic acid sequence in this manner can be used in a variety of applications, such as, for example, to remove disease-causing trinucleotide repeat sequences in neurons, to create gene knock-outs or knock-downs, and to generate mutations for disease models in research.

[0176] In some embodiments, the methods described herein may be used to genetically modify a plant or plant cell. As used herein, genetically modified plants include a plant into which has been introduced an exogenous polynucleotide. Genetically modified plants also include a plant that has been genetically manipulated such that endogenous nucleotides have been altered to include a mutation, such as a deletion, an insertion, a transition, a transversion, or a combination thereof. For instance, an endogenous coding region could be deleted. Such mutations may result in a polypeptide having a different amino acid sequence than was encoded by the endogenous polynucleotide. Another example of a genetically modified plant is one having an altered regulatory sequence, such as a promoter, to result in increased or decreased expression of an operably linked endogenous coding region. The genetically modified plant may promote a desired phenotypic or genotypic plant trait.

[0177] Genetically modified plants can potentially have improved crop yields, enhanced nutritional value, and increased shelf life. They can also be resistant to unfavorable environmental conditions, insects, and pesticides. The present systems and methods have broad applications in gene discovery and validation, mutational and cisgenic breeding, and hybrid breeding. The present methods may facilitate the production of a new generation of genetically modified crops with various improved agronomic traits such as herbicide resistance, herbicide tolerance, drought tolerance, male sterility, insect resistance, abiotic stress tolerance, modified fatty acid metabolism, modified carbohydrate metabolism, modified seed yield, modified oil percent, modified protein percent, resistance to bacterial disease, disease (e.g. bacterial, fungal, and viral) resistance, high yield, and superior quality. The present methods may also facilitate the production of a new generation of genetically modified crops with optimized fragrance, nutritional value, shelf-life, pigmentations (e.g., lycopene content), starch content (e.g., low- gluten wheat), toxin levels, propagation and / or breeding and growth time. See, for example,CRISPR / Cas Genome Editing and Precision Plant Breeding in Agriculture (Chen et al., Annu Rev Plant Biol. 2019 Apr 29;70:667-69), incorporated herein by reference.[0178| The present method may confer one or more of the following traits to the plant cell: herbicide tolerance, drought tolerance, male sterility, insect resistance, abiotic stress tolerance, modified fatty acid metabolism, modified carbohydrate metabolism, modified seed yield, modified oil percent, modified protein percent, resistance to bacterial disease, resistance to fungal disease, and resistance to viral disease.

[0179] The present disclosure provides for a modified plant cell produced by the present method, a plant comprising the plant cell, and a seed, fruit, plant part, or propagation material of the plant. Transformed or genetically modified plant cells of the present disclosure may be as populations of cells, or as a tissue, seed, whole plant, stem, fruit, leaf, root, flower, stem, tuber, grain, animal feed, a field of plants, and the like. The present disclosure provides a transgenic plant. The transgenic plant may be homozygous or heterozygous for the genetic modification. Also provided by the present disclosure are transformed or genetically modified plant cells, tissues, plants, and products that contain the transformed or genetically modified plant cells. The present disclosure further encompasses the progeny, clones, cell lines or cells of the transgenic plants.

[0180] The present system and method may be used to modify a plant stem cell. The present disclosure further provides progeny of a genetically modified cell, where the progeny can comprise the same genetic modification as the genetically modified cell from which it was derived. The present disclosure further provides a composition comprising a genetically modified cell.

[0181] In one embodiment, the transformed or genetically modified cells, and tissues and products comprise a nucleic acid integrated into the genome, and production by plant cells of a gene product due to the transformation or genetic modification.

[0182] Methods of introducing exogenous nucleic acids into plant cells are well known in the art. Such plant cells are considered “transformed.” DNA constructs can be introduced into plant cells by various methods, including, but not limited to PEG- or electroporation-mediated protoplast transformation, tissue culture or plant tissue transformation by biolistic bombardment, or the Agrobacterium-mediated transient and stable transformation. The transformation can be transient or stable transformation. Suitable methods also include viral infection (such as doublestranded DNA viruses), transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whiskers technology, Agrobacterium-mediated transformation, and the like. The choice of method is generally dependent on the type of cell being transformed and the circumstances under which the transformation is taking place (e.g., in vitro, ex vivo, or in vivo). Transformation methods based upon the soil bacterium Agrobacterium tumefaciens are useful for introducing an exogenous nucleic acid molecule into a vascular plant. The wild-type form of Agrobacterium contains a Ti (tumor-inducing) plasmid that directs production of tumorigenic crown gall growth on host plants. Transfer of the tumor-inducing T-DNA region of the Ti plasmid to a plant genome requires the Ti plasmid-encoded virulence genes as well as T-DNA borders, which are a set of direct DNA repeats that delineate the region to be transferred. An Agrobacterium-based vector is a modified form of a Ti plasmid, in which the tumor inducing functions are replaced by the nucleic acid sequence of interest to be introduced into the plant host.

[0183] Agrobacterium-mediated transformation generally employs cointegrate vectors or binary vector systems, in which the components of the Ti plasmid are divided between a helper vector, which resides permanently in the Agrobacterium host and carries the virulence genes, and a shuttle vector, which contains the gene of interest bounded by T-DNA sequences. A variety of binary vectors are well known in the art and are commercially available, for example, from Clontech (Palo Alto, Calif.). Methods of coculturing Agrobacterium with cultured plant cells or wounded tissue such as leaf tissue, root explants, hypocotyledons, stem pieces or tubers, for example, also are well known in the art. See., e.g., Glick and Thompson, (eds.), Methods in Plant Molecular- Biology and Biotechnology, Boca Raton, Fla.: CRC Press (1993), incorporated herein by reference.

[0184] Microprojectile-mediated transformation also can be used to produce a transgenic plant. This method, first described by Klein et al. (Nature 327:70-73 (1987), incorporated herein by reference), relies on microprojectiles such as gold or tungsten that are coated with the desired nucleic acid molecule by precipitation with calcium chloride, spermidine, or polyethylene glycol. The microprojectile particles are accelerated at high speed into an angiosperm tissue using a device such as the BIOLISTIC PD-1000 (Biorad; Hercules Calif.).

[0185] In one embodiment, the present methods may be adapted to use in plants. The vectors may be optimized for transient expression of the present system in plant protoplasts, or for stable integration and expression in intact plants via the Agrobacterium-mediated transformation.

[0186] In certain embodiments, the present methods use a monocot promoter to drive the expression of one or more components of the present systems (e.g., gRNA) in a monocot plant. In certain embodiments, the present methods use a dicot promoter to drive the expression of one or more components of the present systems (e.g., gRNA) in a dicot plant.

[0187] The present methods may be used with various microbial species, including human pathogens that are medically important, and bacterial pests that are key targets within the agricultural industry, as well as antibiotic resistant versions thereof. The method may be designed to target any gene or any set of genes, such as virulence or metabolic genes, for clinical and industrial applications in other embodiments. For example, the present methods may be used to target and eliminate virulence genes from the population, to perform in situ gene knockouts, or to stably introduce new genetic elements to the metagenomic pool of a microbiome. The present systems and methods may be used to treat a multi-drug resistance bacterial infection in a subject. The present systems and methods may be used for genomic engineering within complex bacterial consortia.

[0188] The present systems and methods may be used to inactivate microbial genes. In some embodiments, the gene is an antibiotic resistance gene. For example, the coding sequence of bacterial resistance genes may be disrupted in vivo by insertion of a DNA sequence, leading to non-selective re- sensitization to drug treatment.

[0189] The methods described here also provide for treating a disease or condition in a subject. The method may comprise administering to the subject, in vivo, or by transplantation of ex vivo treated cells (e.g., disclosed T cells), a therapeutically effective amount of the present system, polypeptides, or components thereof.

[0190] In some embodiments, the methods are used to treat a pathogen or parasite on or in a subject by altering the pathogen or parasite. In some embodiments, the methods target a “disease-associated” gene. The term “disease-associated gene,” refers to any gene or polynucleotide whose gene products are expressed at an abnormal level or in an abnormal form in cells obtained from a disease-affected individual as compared with tissues or cells obtained from an individual not affected by the disease. A disease-associated gene may be expressed at anabnormally high level or at an abnormally low level, where the altered expression correlates with the occurrence and / or progression of the disease. A disease-associated gene also refers to a gene, the mutation or genetic variation of which is directly responsible or is in linkage disequilibrium with a gene(s) that is responsible for the etiology of a disease. Examples of genes responsible for such “single gene” or “monogenic” diseases include, but are not limited to, adenosine deaminase, a-1 antitrypsin, cystic fibrosis transmembrane conductance regulator (CFTR), P-hemoglobin (HBB), oculocutaneous albinism II (0CA2), Huntingtin (HTT), dystrophia myotonica-protein kinase (DMPK), low-density lipoprotein receptor (LDLR), apolipoprotein B (APOB), neurofibromin 1 (NF1), polycystic kidney disease 1 (PKD1), polycystic kidney disease 2 (PKD2), coagulation factor VIII (F8), dystrophin (DMD), phosphate-regulating endopeptidase homologue, X-linked (PHEX), methyl-CpG-binding protein 2 (MECP2), and ubiquitin- specific peptidase 9Y, Y-linked (USP9Y). Other single gene or monogenic diseases are known in the art and described in, e.g., Chial, H. Rare Genetic Disorders: Learning About Genetic Disease Through Gene Mapping, SNPs, and Microarray Data, Nature Education 1(1): 192 (2008); Online Mendelian Inheritance in Man (OMIM); and the Human Gene Mutation Database (HGMD). In another embodiment, the target genomic DNA sequence can comprise a gene, the mutation of which contributes to a particular disease in combination with mutations in other genes. Diseases caused by the contribution of multiple genes which lack simple (i.e., Mendelian) inheritance patterns are referred to in the art as a “multifactorial” or “polygenic” disease. Examples of multifactorial or polygenic diseases include, but are not limited to, asthma, diabetes, epilepsy, hypertension, bipolar disorder, and schizophrenia. Certain developmental abnormalities also can be inherited in a multifactorial or polygenic pattern and include, for example, cleft lip / palate, congenital heart defects, and neural tube defects. In another embodiment, the target DNA sequence can comprise a cancer oncogene. The present disclosure provides for gene editing methods that can ablate a disease-associated gene (e.g., a cancer oncogene), which in turn can be used for in vivo gene therapy for patients. In some embodiments, the gene editing methods include donor nucleic acids comprising therapeutic genes.Kits

[0191] Also within the scope of the present disclosure are kits that include the components of the present system.

[0192] The kit may include instructions for use in any of the methods described herein. The instructions can comprise a description of administration to a subject to achieve the intended effect. The instructions generally include information as to dosage, dosing schedule, and route of administration for the intended treatment. The kit may further comprise a description of selecting a subject suitable for treatment based on identifying whether the subject is in need of the treatment.

[0193] The kits provided herein are in suitable packaging. Suitable packaging includes, but is not limited to, vials, bottles, jars, flexible packaging, and the like. A kit may have a sterile access port (for example, the container may be an intravenous solution bag or a vial having a stopper pierceable by a hypodermic injection needle). The container may also have a sterile access port.

[0194] The packaging may be unit doses, bulk packages (e.g., multi-dose packages) or subunit doses. Instructions supplied in the kits of the disclosure are typically written instructions on a label or package insert. The label or package insert indicates that the pharmaceutical compositions are used for treating, delaying the onset, and / or alleviating a disease or disorder in a subject.

[0195] Kits optionally may provide additional components such as buffers and interpretive information. Normally, the kit comprises a container and a label or package insert(s) on or associated with the container. In some embodiments, the disclosure provides articles of manufacture comprising contents of the kits described above.

[0196] The kit may further comprise a device for holding or administering the present system, polypeptides, or composition. The device may include an infusion device, an intravenous solution bag, a hypodermic needle, a vial, and / or a syringe.

[0197] The present disclosure also provides for kits for performing nucleic acid modification and integration in vitro. Optional components of the kit include one or more of the following: buffer constituents, control plasmid(s), sequencing primers, and cells.Example 1

[0198] CAST 7016, a system encoded by Pseudoalteromonas sp. S983, and a system comprising variant components as described in FIG. 1A were tested for integration efficiency at a variety of mammalian genomic targets. The engineered system improved integration activity at some loci, and achieved about 10-30% targeted gene integration efficiency without enrichment(FIG. IB). Across all the endogenous loci tested in HEK293T cells, the average efficiency was greater than 350-fold higher than that of WT CAST 7016 (FIGS. 1C-1E).Example 2Development from evolved and engineered components

[0199] Evolved and engineered variants of P.seCAST components that enhanced integration individually or in various combinations were combined. After evaluating evolved TnsA and TnsC variants in combination with P4-15 TnsB (FIG. 2A), a TnsABC combination was identified averaging 1.3-fold improved integration across four genomic sites in HEK293T cells compared to P4-15 TnsB with wild-type TnsA and TnsC. Structure-guided engineering and optimized nuclear localization sequences (NLS) were utilized to develop a QCascade module that supported enhanced integration activity in human cells. Through this engineering, a QCascade module was selected that combined: i) a Cas7 with a neutral DNA-contacting residue mutated to lysine; ii) a Cas8 containing an engineered PAM-interacting domain previously found to improve wild-type P eCAST activity in HEK293T cells; and iii) an additional bipartite NLS at the N-termini of TniQ, Cas6, and Cas8.

[0200] By combining PACE-evolved TnsABC with rationally engineered QCascade, a CAST system (evoCAST, referred to in Example 1 and FIG. 1 as eeCAST) was developed for activity in human cells (FIG. 2A). Across four genomic sites in HEK293T cells, evoCAST averaged 19% integration, representing an average 1.2-fold improvement over P4-15 TnsB with unoptimized non-TnsB PseCAST components and an average 540-fold improvement over wild-type PseCAST (FIG. 2B). EvoCAST supported a range of DNA payload sizes up to 15 kb, the largest size tested (FIG. 2C). EvoCAST also supported integration of both plasmid and linear transposon donor DNA topologies (FIG. 4). Together, the improvements made to all seven PseCAST protein components enabled evoCAST to serve as a platform for targeted genomic integration of gene-sized DNA cargoes in mammalian cells.Example 3 Characterization of integration products[02011 evoCAST integration products were examined in HEK293T cells. High-throughput sequencing (HTS) of genome-transposon junctions revealed that evoCAST retained an insertion site preference similar to that of wild-type P eCAST, integrating -49 bp downstream of theRNA-complementary target sequence (FIG. 2D). In contrast to nuclease-mediated end-joining or HDR methods, cvoCAST yielded no detected indcl formation at unintegrated loci (FIG. 2E), despite mediating efficient DNA integration. HTS also revealed low levels (<3%) of substitution mutations within the 5-bp target-site duplication (TSD) for evoCAST integration products (FIG. 5), which may have arisen during host repair of the 5-nt gaps generated by offset TnsB transesterifications. To assess the orientation of integrated transposons, ddPCR was performed with probes specific to either T-RL or T-LR products. evoCAST was highly biased for T-RL integration across four genomic loci tested (FIG. 2F). Long-read sequencing of insertion product amplicons indicated that >80% of evoCAST and wild-type PseCAST products were simple insertions as opposed to cointegrates, undesired byproducts commonly observed with Type V-K CASTs (FIG. 6), suggesting that evoCAST development also retained the desired cut-and-paste transposition chemistry. Collectively, these results demonstrate that evoCAST offers high product purity, with integration occurring predominantly with single-bp precision, no detected indel formation, unidirectionality, and the formation of simple insertions over cointegrate byproducts.Example 4Characterization of off-target evoCAST integration

[0202] To determine the genome-wide specificity of evoCAST in human cells, a modified Uni-Directional Targeted Sequencing (UDiTaS) approach was used (FIG. 7A), which was previously applied to recombinases in bacteria and human cells. Although Type LF CASTs show high specificity in E. coli, the extremely low levels of integration with wild-type P.seCAST, together with the cytotoxicity of ClpX in human cells, precluded investigation of genome-wide specificity of wild-type PseCAST in human cells.

[0203] The specificity of evoCAST targeting AAVS1 in HEK293T cells was assessed following one week of incubation with plasmid expression vectors. On-target integration was by far the most prevalent integration product, although integration events scattered at other locations in the human genome were also identified (FIG. 2G). UMI analysis indicated that each detected off-target represented a single integration event (FIG. 2G), without detected homology to the AAVS1 target site. Across two replicates, evoCAST averaged 36% on-target integration (Table 1 below), similar to the fraction of on-target editing events calculated from published data for Cas9nuclease (average 47%) and programmable gene integration with eePASSIGE (0.51-38%, depending on the attachment site used in the donor DNA) (Table 1 below).Table 1* Tsai et al., NBT 2015 (GUIDE-seq)** Pandey et al., BME 2024[0204| None of the off-target integration sites were reproducible across replicates, suggesting an unguided mechanism. Consistent with this hypothesis, off-target integration involved TnsC but not QCascade (FIG. 7B). Taken together, these data suggest that off-target integration by evoCAST is CRIS PR-independent and likely arises from aberrantly bound TnsABC complexes. Off-target integration events persisted with wild-type TnsA and TnsC (FIG. 7C), suggesting that off-targets may stem from the enhanced activity of evolved TnsB, which may promote integration at transiently engaged off-target substrates by the TnsABC complex. Off-target sites were generally in regions of open chromatin (FIG. 7F), which arc likely to be more accessible substrates for TnsABC. EvoCAST off-target sites did not contain disproportionately high ATcontent (FTG. 7G), suggesting a different mechanism of off-target formation than previously reported for type V-K CAST.

[0205] The stochasticity and very low abundance of off-target events nominated by UDiTaS complicates their quantification via orthogonal methods such as ddPCR. UDiTaS of evoCAST- transfected cells that were enriched for on-target formation did not detect any off-target integration events (FIG. 7D), suggesting that each edited cell likely contained a single integration event, and that cells containing on-target integration generally do not contain related ‘bystander’ off-target integration events. UDiTaS of E. coli lysate from PACE experiments following incubation with P4- 15 -encoding SP revealed that >99% of integration events were on-target (FIG. 7E), suggesting that limited exposure to evoCAST (a timescale of hours in PACE, as opposed to days in HEK293T cells) drives on-target integration prior to any accumulation of off- target integration, consistent with the behavior of other genome editing agents such as nucleases, base editors, and prime editors.[0206| TnsABC-mediated off-target integration was not detected by an orthogonal fluorescent reporter assay previously used to detect off-target formation by eeBxbl recombinase with an attP-containing donor (FIG. 7H), consistent with the UDiTaS data, above, demonstrating that evoCAST mediates lower off-target integration than this eePASSIGE configuration.

[0207] These data indicate that the proportion of on-target vs. off-target editing events for evoCAST is similar to those of other current genome editing methods. Additionally, current UDiTaS methods have limited coverage of integration events, due to a substantial fraction of sequencing reads aligning to unintegrated pDonor molecules.

[0208] Further development of off-target characterization methods for targeted gene integration may enable more precise quantification of off-target editing frequencies in mammalian cells for CASTs and other large DNA integration technologies.Example 5Application of evoCAST at target genomic sites of therapeutic interest

[0209] evoCAST was used to perform targeted gene-sized DNA integration in human cells at genomic sites relevant to gene therapies. Following CRISPR RNA (crRNA) architecture and spacer sequence optimization (FIG. 8), integration with the best-performing crRNAs was assessed at 14 human genomic loci corresponding to potential therapeutic applications of evoCAST. 14% average 1-kb transposon integration efficiency was observed at these 14 targetsites, compared to 0.22% for wild-type PseCAST (FIG. 3A). EvoCAST efficiencies were positively correlated with chromatin accessibility (FIG. 12), similar to previous reports for other genome editing methods.

[0210] Therapeutic gene integration at ALB is a promising strategy for therapeutic transgene expression in hepatocytes. ALB is highly expressed in the liver, and integration of a splice acceptor-bearing donor within intron 1 enables splicing with a secretion signal in exon 1 for subsequent secretion of the protein of interest. This strategy is currently being investigated as a potential treatment for hemophilia B, which can be rescued upon only 1% restoration of circulating human factor IX (hFIX) levels. To test the potential therapeutic utility of evoCAST, F9 cDNA encoding the hyperactive hFIX Padua variant was integrated into ALB intron 1 (FIG. 3C and 3F). EvoCAST achieved 5.7% targeted integration efficiency in a human hepatocyte cell line (HuH7), compared to 0.023% for wild-type PseCAST (FIG. 3F). Consistent with these data, evoCAST resulted in F9 expression in evoCAST-treated cells, while cells treated with wild-type PseCAST did not yield detected levels of F9 expression (FIG. 11 A).[0211 J Integration of a CAR at the T-cell receptor a constant (TRAC) locus enables uniform CAR expression, enhanced T-cell potency, and delayed T-cell exhaustion. The efficiency of CD19 CAR integration at TRAC, a strategy shown to combat refractory or relapsed B-cell malignancies, was assessed (FIG. 3D and 3G). In HEK293T cells, evoCAST mediated 13% integration of CD19 at TRAC, compared to 0.061% by wild-type P.seCAST (FIG. 3G).

[0212] The programmability of evoCAST also potentiated precise integration of wild-type cDNAs at sites of endogenous gene mutation or deletion that are associated with loss-of-function genetic diseases (FIG. 3E). This strategy in principle may enable a single evoCAST treatment to ameliorate loss-of-function diseases in an allele-agnostic manner while preserving some of the endogenous regulatory context of the target gene. Integration of transposons encoding wild-type cDNAs (Aexon 1, flanked by a 5' splice acceptor and 3' poly A signal) into intron 1 of FANCA (associated with Fanconi anemia), IL2RG (X-linked severe combined immunodeficiency), MECP2 (Rett syndrome), and PAH (phenylketonuria) was assessed (FIG. 3H). EvoCAST supported substantial targeted gene insertion efficiencies of 12-15% at these loci, compared to 0.0092-0.43% for wild-type PseCAST (FIG. 3H). Integrated transgene expression was measured at MECP2, which is expressed in HEK293T cells, via RT-ddPCR using a probe specific to therecoded exon 2 in the transgene, and evoCAST, but not wild-type PseCAST, yielded detected levels of MECP2 expression (FIG. 11B).[02131 To enable broader applications of evoCAST, transposon end sequences that do not comprise integration activity but are compatible with in-frame protein tagging were engineered (FIG. 9). If in-frame insertion is desired, a researcher should first profile the distribution of DNA insertion sites for a given crRNA targeting the region of interest. This distribution is generally centered ~49 bp downstream of the target crRNA and is highly. Once the insertion site distribution is determined, the researcher can then select which of the tagging-compatible transposon end variants enables translational read-through across the transposon end into the integrated cargo.

[0214] evoCAST-edited cells persist in bulk populations when using a selectable marker, and that clonally integrated populations can be isolated via single-cell sorting (FIG. 10).[0215| Collectively, these results demonstrate that evoCAST can be reprogrammed to integrate large, diverse DNA payloads across multiple genomic loci in human cells, enabling a range of potential applications in basic research and therapeutic science.Example 6 evoCAST in other mammalian cell lines

[0216] To extend characterization of evoCAST beyond HEK293T and HuH7 cell lines, 1-kb transposon integration was assessed at two genomic sites in two additional human cell lines, HeLa and K562 cells (FIG. 51). In HeLa cells, evoCAST averaged 4.7% editing activity, compared to 0.18% for wild-type PseCAST (FIG. 51). In K562 cells, evoCAST averaged 1.6% editing, compared to 0.038% by wild-type PseCAST (FIG. 31). The lower editing activity in these cell types compared to HEK293T cells may arise from their less efficient transfection of the multiple vectors encoding CAST components along with the donor DNA. Alternative or optimized delivery methods may further improve evoCAST efficiencies in difficult-to-transfect cell lines.

[0217] To demonstrate utility in a more therapeutically relevant cell type, evoCAST was evaluated in primary human fibroblast cells isolated from a patient with recessive dystrophic epidermolysis bullosa (RDEB), a mutationally diverse loss-of-function disease that can be treated using autologous gene-corrected fibroblasts. Strikingly, evoCAST averaged 19% efficiency of 1-kb transposon integration across two genomic loci, representing a 200-foldimprovement over wild-type PseCAST (FIG. 3J). Primary human cells can support evoCAST activity, particularly when cultured under conditions that mitigate the toxicity of exogenous DNA.

[0218] The impact of ClpX co-delivery on evoCAST efficiency was assessed at two genomic 25 loci in HEK293T, HeLa, K562, and primary human fibroblast cells (FIGS. 3K, 13A and 13B). EvoCAST did not require ClpX for maximal editing in HEK293T cells (FIG. 3K). ClpX increased evoCAST editing in other cell types, notably enabling an average 37% editing in primary fibroblasts (FIG. 13B), although with increased cytotoxicity (FIG. 13C). EvoCAST displayed substantially reduced ClpX dependence compared to wild-type PseCAST across all cell types (FIG. 3K), enabling CAST-mediated editing human cells without ClpX-associated cytotoxicity.Example 7 evoCAST compared with eePASSIGE

[0219] EvoCAST was compared with eePASSIGE, a recently described method for targeted DNA integration in human cells without requiring genomic double-strand breaks. 2-kb DNA cargo integration was assessed at six genomic loci in HEK293T cells for both strategies (FIG. 14A), though the exact target sites differed between methods due to distinct DNA targeting mechanisms. eePASSIGE exhibited 1.5-2.9-fold higher integration efficiency than evoCAST, although both methods typically achieved 40 >10% integration (FIG. 14A). Both evoCAST and eePASSIGE supported 10-40% integration efficiencies in primary human fibroblasts (FIGS. 3J and 13B). The distinct mechanisms of evoCAST and eePASSIGE result in different advantages. While more efficient on average, the eePASSIGE method requires attachment site installation by prime editing prior to DNA integration, resulting in a more heterogeneous mixture of editing outcomes than is observed with evoCAST (FIGS. 14B-14D). EvoCAST’ s high product purity makes it well-suited for applications in which byproducts such as indels or unintegrated attachment site installation must be minimized. Moreover, eePASSIGE can require testing many prime editing conditions for efficient attachment site installation. In contrast, high-performance evoCAST crRNAs were identified after testing only 5-10 constructs per genomic locus.

[0220] EePASSIGE and evoCAST also differ in their compatibility with donor DNA substrates. Recombinase-mediated integration of linear DNA yields genomic double-strand breaks, which is not expected with evoCAST (FIG. 14E). Indeed, when using a linear donor,36% of eePASSIGE integration products contained indels consistent with double-strand break formation, whereas cvoCAST products contained no detected indels above background (FIGS. 14E-14H). EvoCAST’s compatibility with linear donor topology enables PCR amplicons to be more readily used as substrates for targeted integration. This feature is also advantageous for viral delivery modalities such as AAV that deliver linear DNA. When using a circular donor DNA substrate, evoCAST cleanly integrates the desired DNA payload without flanking vector sequences. In contrast, recombinase-mediated integration installs the entire vector sequence, which can include undesired DNA elements. Taken together, these findings demonstrate the strengths of evoCAST as a novel platform for single-step, programmable gene integration in human cells.Example 8 Development of CAST PACE

[0221] PACE maps the key steps of traditional, stepwise directed evolution onto the M13 bacteriophage life cycle, accelerating the laboratory evolution of biomolecules by >100-fold with minimal researcher intervention (FIG. 15B). During PACE, a selection phage (SP) expresses an evolving gene of interest in place of gill, an essential gene for phage replication, gill is instead encoded on an accessory plasmid (AP) in host E. coli under a transcriptional circuit linking gill expression to the desired activity of the protein of interest. SP populations are mutagenized via an inducible mutagenesis plasmid (MP) and diluted with fresh cells, either continuously (PACE) or periodically (phage-assisted non-continuous evolution, PANCE), in fixed-volume ‘lagoons.’

[0222] PACE efficiently and unbiasedly explores vast sequence spaces, and it exhibits few requirements beyond the evolving protein’s ability to induce gill expression in E. coli. These aspects make PACE well-suited for evolving type I-F CASTs, which are large, multi-component systems lacking extensive structural and biochemical characterization, limiting rational engineering. Evolution was focused on the transposase module of PseCAST by encoding TnsA, TnsB, and TnsC (referred to hereafter as TnsABC) on the SP.

[0223] To evolve TnsABC for increased integration efficiency, a PACE selection linking transposition activity to phage propagation was developed (FIG. 15C). Host cells contain a complementary plasmid 1 (CPI) that expresses the PseCAST components (QCascade) promoting DNA target binding. The selection requires targeted insertion of a transposon-encoded promoter sequence, provided by complementary plasmid 2 (CP2), upstream of a promoter-less gill on theAP. An SP encoding an active TnsABC variant supports promoter transposition from CP2 to AP, activating gill expression and propagation of that SP. To increase selection stringency throughout evolution, CP2 constructs were developed with progressively weaker promoter strengths, requiring more integration events into the multi-copy AP to trigger sufficient gill expression for SP propagation before dilution out of the lagoon.

[0224] Despite high integration activity in E. coli, wild-type TnsABC did not support SP propagation, even on the least stringent selection circuit. This finding suggested that, under the conditions tested, wild-type TnsABC catalysis may be too slow to activate SP propagation, which requires gill activation within minutes to hours of infection. Overnight incubation of TnsABC SP with host E. coli yielded low but detectable RNA-dependent CP2 transposon integration at the AP (0.0036%), verifying that the PACE circuit can be triggered, albeit weakly, by wild-type PseCAST. Luciferase reporter assays indicated that CP2 transposon integration at the AP is sufficient to activate downstream gene expression, and overnight propagation assays demonstrated that PseCAST expression does not interfere with phage propagation. Collectively, these findings suggested that this CAST PACE selection (circuit 1.0, FIG. 18) can link the integration activity of an SP-encoded TnsABC to SP propagation, if kinetically enhanced TnsABC variants enable transposition on a timescale relevant for phage replication.Example 9 Evolution of TnsABC

[0225] Evolution of wild-type TnsABC was initiated using PANCE (FIG. 16A), a less stringent alternative to PACE in which dilution with fresh host cells occurs serially after overnight phage propagation. To allow weakly active SP variants to accumulate new mutations in the absence of selection, passages on the selection E. coli strain were alternated with passages on a ‘drift strain’ that provides CAST-indcpcndcnt gill expression, allowing recovery and further diversification of surviving genes. Following 13 passages on host cells (PANCE Nl), pooled SPs demonstrated ~ 106-fold improved overnight propagation on the selection strain and 320-fold improved integration at the AP (FIG. 16B). PANCE successfully linked SP propagation with the integration activity of evolving TnsABC variants.

[0226] SP from Nl propagated at levels sufficient for PACE, thus PACE (Pl) was initiated with Nl-derived SP (FIG. 16A). Following 48 hours of PACE, all evolving Pl populations were dominated by ‘cheating’ SPs that propagated independently of TnsABC activity by acquiring acopy of gITI. Sequencing revealed that Pl SPs obtained gill via aberrant integration of the entire post-transposition AP vector into the SP. To reduce the risk of undcsircd gill acquisition in future evolutions, PACE circuit 1.1 was developed, which uses a split gill with each half fused to a trans-splicing split intein in either the AP or CPI (FIG. 16B). With this design, full-length gill acquisition by the SP would require integration or recombination of both the AP and CPI into the same SP genome, which is unlikely.

[0227] With two integration events, instead of one, now driving full-length gill expression within circuit 1.1, Pl -derived SPs exhibited reduced overnight propagation compared to circuit 1.0. PANCE (N2) was performed using circuit 1.1, seeding lagoons with clonal Aglll SP from Pl (FIG. 16A). Following 20 passages of alternating selection and drift, the N2 SP pool still showed insufficient propagation for PACE. To amplify signal from integration events, circuit 1.2 was developed, which contains a modified CPI that links integration to T7 RNA polymerase expression and places the N-terminal gill half under the control of a T7 promoter (FIG. 18). Using circuit 1.2, PACE (P2) was initiated, seeding lagoons with an N2 SP pool and evolving for 144 hours (FIG. 16A).

[0228] SP variants from Nl, Pl, N2, and P2 evolution experiments exhibited increasing levels of overnight propagation and integration at the AP, indicating that PACE successfully enriched active TnsABC variants (FIGS. 16B and 18). Evolved variants contained diverse mutations across TnsA, TnsB, and TnsC, with generally little mutational convergence across independently evolving SP populations. PACE explored multiple trajectories for increasing TnsABC-mediated integration efficiency.

[0229] Evolved TnsABC variants were evaluated in HEK293T cells, focusing on bestperforming variants (through N2) and representative P2 variants (FIGS. 16C and 16D). Unless otherwise noted, all human-cell integration assays assessed 1-kb transposon integration efficiencies without ClpX supplementation, quantified via droplet digital PCR (ddPCR). Encouragingly, evolution through N2 substantially improved integration at two endogenous genomic loci, increasing from an average 0.062% for wild-type PseCAST to an average 3.6% for the best-performing TnsABC variant N2-1 (FIG. 16D). However, while the P2 SP pool supported the highest overnight propagation and AP integration (FIG. 16B), P2 TnsABC variants were substantially less active in HEK293T cells than the N2-1 variant (FIG. 16D). Thesefindings suggested that TnsABC variants evolved fitness gains in E. coli during P2 that did not result in higher human-cell activity.

[0230] To better understand the disconnect between PACE fitness and human-cell integration activity, the individual contributions of evolved TnsAB (the heteromeric transposase) and TnsC (an AAA+ ATPase regulator of transposition) to SP propagation in PACE and DNA integration were evaluated in HEK293T cells (FIGS. 16E and 16F). While P2-derived TnsAB and TnsC subunits synergized to increase SP propagation (FIG. 16E), P2-derived TnsC variants reduced integration efficiencies in HEK293Ts on average by 2.8-fold compared to efficiencies mediated by wild-type TnsC (FIG. 16F). These data suggested that TnsC acquired mutations during P2 that decreased human-cell integration activity despite improving SP fitness in PACE.

[0231] Reversion analysis of P2 TnsC valiants identified D44N / G and N316D, two highly conserved mutations among P2 variants, as the source of reduced activity in human cells. Overnight propagation assays confirmed these mutations benefited SP propagation, suggesting that TnsC-mediated determinants of PseCAST activity differ between E. coli and human cells. Based on an AlphaFold3 -predicted TnsC model, D44 is near the ATP-binding pocket, and N316 lies at the interface between adjacent TnsC monomers near the target DNA. Current models of type I-F CAST mechanism suggest that ATP binding and TnsC oligomerization are necessary for recruitment to the QCascade-bound target site. Since the DNA target search space in E. coli is much smaller than in human cells, it was speculated that PACE optimized TnsC for improved target engagement in E. coli through mutations that did not benefit activity in human cells.Example 10 Evolution of TnsAB

[0232] To more effectively evolve variants that are active in human cells, PACE circuit 2.0 was developed, which encodes TnsC on CPI instead of the SP, thereby restricting evolution to TnsAB (FIG. 18). In addition, the circuit was simplified by removing split-intein gill and instead encoding full-length gill on an enlarged AP (10 kb). Aberrant AP recombination into the SP would yield a phage genome exceeding the Ml 3 phage packaging capacity, thus reducing the risk of cheating through gill acquisition. During circuit 2.0 design, the transposon left end of type I-F CASTs contained a conserved binding site for bacterial integration host factor (IHF). IHF promotes transposition activity of some type I-F CASTs in E. coli, including PseCAST to aweak extent. Thus, to prevent the evolution of IHF-dependent fitness, which would not translate to human cells, the IHF binding site in the transposon left end was mutated in CP2 (FIG. 18). [02331 Selection circuit 2.0 yielded poor SP propagation, likely due to moving TnsC from a high-copy SP to a low-copy CPI. TnsAB variants Pl -3 and N2-1, which exhibited high activity in HEK293T cells, were evolved using selection circuit 2.0 in PANCE (N3) for 25 passages, alternating selection with drift through passage 16 (FIG. 17A). Following N3, PACE (P3) was seeded with SPs encoding N3 TnsAB variants and evolved for 140 hours (FIG. 17A) using a modified circuit 2.1 architecture (FIG. 18).

[0234] TnsAB variants emerging from N3 and P3 were isolated and characterized. In contrast to previous TnsABC evolutions, most evolved TnsAB variants showed improved editing in HEK293T cells. Mutations in TnsB alone were sufficient to achieve maximum editing levels from the top-performing variants, suggesting that TnsB-related activities, which include transposon end binding and transesterification catalysis, represent the primary bottlenecks limiting PseCAST activity in human cells.Example 11Evolution of TnsB

[0235] Given the above findings, evolution was restricted to TnsB by developing PACE circuit 3.0, which encodes TnsA on CPI instead of the SP (FIG. 18). Bacterial protein ClpX enhances type I-F CAST activity in HEK293T cells, though with considerable cytotoxicity. To prevent evolution of ClpX-dependent fitness in PACE, a host E. coli strain with endogenous clpX deleted was developed. Using circuit 3.0, this strain reduced overnight propagation of TnsB-encoding SP by ~200-fold, whereas propagation of a gill-encoding SP was unaffected. These data implicated ClpX in E. coli-based PseCAST activity and revealed an altered selection pressure when evolving TnsB in a AclpX host, all evolution experiments were performed with circuit 3.0 in this AclpX host.

[0236] Evolution was focused on P3-13 TnsB, the best performing TnsB variant from P3 (FIGS. 17E and 17F). P3-13 TnsB SP did not propagate sufficiently for PACE, thus PANCE (N4) was initiated on P3-13 TnsB, performing 18 selection passages (FIG. 17A). The evolved N4 SP pool was used to seed PACE (P4), evolving for 108 hours (FIG. 17A). Most P4 TnsB variants showed improved activity in HEK293T cells compared to P3-13, with the best performing variant, P4-15, averaging 12% integration efficiency across three genomic loci(FIGS. 17B and 17C). Although TnsAB and TnsB evolution campaigns used a transposon left end with a mutated IHF binding site, representative P4 TnsB variants performed similarly with wild-type or mutant left ends in HEK293T cells. Additionally, P4-15 TnsB, despite evolving as an unfused peptide in PACE, still mediated the highest editing in HEK293T cells when fused to TnsA via a bpNLS linker, the optimal configuration for wild-type TnsB.

[0237] Evolving P4-15 TnsB in higher stringency host cells with reduced CP2 promoter strength failed to improve human-cell integration activity, as did introducing mutations from other highly active P4 TnsB valiants into P4-15 TnsB. This plateau may indicate that the evolved P4-15 TnsB variant no longer bottlenecks integration efficiency under the conditions tested, or that new selection pressures or evolutionary trajectories need to be explored for TnsB PACE to continue improving integration activity. Overall, phage encoding P4-15 TnsB experienced a total 10322-fold dilution over 76 PANCE passages and 296 hours of PACE, corresponding to hundreds of evolutionary generations.

[0238] In-depth characterization of P4-15, a high-performing evolved TnsB variant, was performed. First, ClpX effects on integration mediated by P4-15 were assessed (FIG. 17D). While ClpX enhanced wild-type TnsB integration activity on average by 4.0-fold across three genomic sites in HEK293T cells, ClpX had no impact on P4- 15 -mediated editing at any genomic site tested, with evolutionary precursors of P4-15 also exhibiting reduced ClpX reliance (FIG. 17D). Notably, ClpX independence emerged before selection on the AclpX E. coli host (FIG. 17D), suggesting early evolution experiments enriched variants with reduced ClpX-dependence. In CAST PACE, PTC disassembly is required for SP propagation, as RNA polymerase must traverse the repaired 5-nt gap to transcribe gill.

[0239] To investigate how PACE improved TnsB activity, mutated residues in P4-15 were mapped onto two AlphaFold3-predicted structure models: a TnsB strand-transfer complex (FIG. I7E) and a TnsB C-terminal ‘hook’ domain in complex with a TnsC heptamer (FIG. 17F). Based on E. coli Tn7 biochemistry, TnsB performs multiple functions in the CAST transposition cycle, including complexing with TnsA, binding to transposon ends, binding to the target-bound TnsC, catalyzing DNA cleavage and transesterification reactions, and undergoing conformational rearrangements to allow 5-nt gap fill-in. Evolved mutations span multiple TnsB domains and predicted interfaces, including TnsB*transposon end (Y349N), TnsB«TnsB (Y349N, P352T, D396N, H464R, and V526E), and TnsB’TnsC (Q594L) (FIGS. 17E and 17F),suggesting that PACE evolved multiple TnsB functionalities to improve integration efficiency in human cells.[0240| Reversion analysis of each mutation in P4-15 revealed that all ten mutations increase activity in HEK293T cells. Integration efficiency of each revertant was unchanged upon ClpX addition. When installed individually into wild-type TnsB, all mutations improved efficiencies (though several only to a modest extent) except A390V, which may act through epistasis. P352T and D396N, acquired early in evolution (FIG. 17B) and predicted to lie at a TnsB’TnsB interface (FIG. 17E), enabled the highest integration efficiencies among the evolved mutations when tested individually. No individual mutation conferred ClpX independence.

[0241] Cell-based assays were performed to illuminate the properties of evolved TnsB. Transposon-end binding was assessed by P4-15, its evolutionary precursors, and P4-15 singlemutation revertants via an established transcriptional activation assay in HEK293T cells. P4-15 TnsB exhibited 3.2-fold improved reporter activation compared to wild-type TnsB, with the largest increase resulting from mutations in Pl -3. Accordingly, reverting A390V, a mutation descending from Pl-3, in P4-15 greatly reduced activity. Residue A390 is not proximal to transposon DNA in the AlphaFold3-predicted structure, nor is it in DNA-binding domains, suggesting a long-distance mechanism by which A390V improves activity. Western blots revealed that all tested TnsB variants have similar soluble expression in HEK293T cells, suggesting that the elevated reporter signal in the transcriptional activation assay resulted from enhanced TnsB*DNA binding.

[0242] Transposition by P4-15 TnsB and its evolutionary precursors was also monitored via an E. coli liquid culture assay, in which integration results in luciferase expression. Over the first two hours of transposase induction, all evolved variants showed enhanced activity over wild-type TnsB. Highly evolved variants exhibited faster apparent rates of luciferase expression, suggesting that PACE selected for TnsB variants with faster transposition kinetics. These results also demonstrate that evolved TnsB variants have improved activity in E. coli, in addition to human cells.

[0243] Collectively, these data suggest that PACE optimized diverse TnsB interactions with itself and other CAST components to improve human-cell integration activity. These findings highlight the advantage of using directed evolution to improve CAST activity, here identifyingten activity-enhancing mutations that would be very difficult to deduce purely through rational protein engineering.Materials and Methods

[0244] General methods Antibiotics were purchased from Gold Biotechnology and used at the following concentrations: streptomycin (50 pg / mL), chloramphenicol (25 pg / mL), carbcnicillin (50 pg / mL), spectinomycin (50 pg / mL), tetracycline (10 pg / mL), and kanamycin (25 pg / mL). PCRs were performed using Phusion U Green Multiplex PCR Master Mix (ThermoFisher Scientific) or Q5 Hot Start High-Fidelity 2x Master Mix (New England BioLabs) unless otherwise noted. DNA oligonucleotides, including FAM / Iowa Black FQ-labeled DNA oligonucleotides, were obtained from Integrated DNA Technologies. Human codon-optimized wild-type PseCAST genes were synthesized by GenScript. All plasmids used in this study were cloned using USER, Golden Gate, or Gibson assembly methods. Plasmids were cloned into Maehl (ThermoFisher Scientific) chemically competent E. coll. Unless otherwise noted, plasmid DNA was amplified using the Illustra Templiphi 100 Amplification kit (GE Healthcare Life Sceinces) prior to Sanger sequencing (Quintara Biosciences) or Nanopore sequencing (Plasmidsaurus). All plasmids for E. coli experiments were purified using QIAprep Spin Miniprep Kits (Qiagen), and all plasmids for mammalian cell experiments were purified using PlasmidPlus Midiprep Kits (Qiagen) or Plasmid Plus 96 Miniprep Kits (Qiagen). All isolated plasmid DNA were eluted in nuclease-free water and quantified using a NanoDrop ONE UV-Vis spectrophotometer (ThermoFisher Scientific).

[0245] Preparation and transformation of chemically competent cells Strain S2060 was used in all luciferase, phage propagation, and plaque assays, and in all PACE experiments, except for PACE campaigns conducted on the EclpX strain generated in this study. Chemically competent cells were prepared as described previously. Briefly, an overnight culture of bacteria was diluted 50-200-fold into 2xYT media (United States Biologicals) with appropriate antibiotics and grown at 37 °C, shaking at 230 RPM until the culture reached an optical density (ODeoo) of 0.4-0.6. Cells were centrifuged at 4 °C for 5-10 min at 4,000 g. The supernatant was discarded, and cell pellets were resuspended in ice-cold TSS solution (LB media supplemented with 5% v / v DMSO, 10% w / v PEG3350, and 20 mM MgCh). Resuspended cells were aliquoted, flash frozen on dry ice, and stored at -80 °C until use.

[0246] To transform cells, 100 pL of competent cells thawed on ice were added to a prechilled mixture of plasmid (1-2 pL each; up to 3 plasmids per transformation) in 100 pL KCM solution (100 mM KC1, 30 mM CaCh, and 50 mM MgCh in H2O) and stirred gently. The mixture was incubated on ice for >5 min, heat shocked at 42 °C for 90 s, and combined with 500 pL of SOC media (New England BioLabs). Cells recovered at 37 °C, shaking at 230 RPM for 1 h. Cells were then streaked on 2xYT media + 1.5% agar (United States Biologicals) plates containing appropriate antibiotics and incubated for 16-18 h at 37 °C.

[0247] Bacteriophage cloning Phage were cloned using USER assembly as previously described with minor modification. Briefly, a 25 pL USER assembly was transformed into 100 pL of chemically competent S2060 E. coli host cells containing pJC175e (S2208), which enables activity -independent phage propagation. Transformed S2208 were incubated overnight at 37 °C in 10 mL 2xYT media shaking at 230 RPM. The saturated culture was then centrifuged for 5 min at 4,000 g, and the phage-containing supernatant was plaqued as described below. Individual phage plaques were grown in DRM media (United States Biologicals) for 6-8 h. Following incubation, the cultures were centrifuged for 5 min at 4,000 g, and the phage-containing supernatants were filtered through a 0.22-pm PVDF ultrafree centrifugal filter (Millipore) to remove residual bacteria. Phage were then sequenced via PCR amplicon sequencing (Quintara Biosciences), with sequence-confirmed phage stored at 4 °C until use.

[0248] Plaque assay Plaquing was performed as previously described. In brief, a saturated S2208 culture was back-diluted 50-100-fold into DRM containing carbenicillin. Cells were grown at 37 °C shaking at 230 RPM to an ODeoo of 0.4-1.0, at which point they were placed on ice during preparation of phage. Phage stocks were serially diluted in water by a factor of 10, up to 106-fold. 10 pL of phage stock and dilutions (typically 102, 104, and 106-fold dilutions) were combined with 100 pL of mid- log S2208 cells in 2-mL library tubes (VWR international). 1 mL of warm top agar (3:2 mixture of 2xYT medium and molten 2xYT medium agar (1.5%, resulting in a 0.6% agar final concentration), stored at 55 °C until use) was added to the phage / bacteria solution, mixed once by pipetting, and then immediately plated on one quadrant of a 2xYT medium 1.5% agar plate containing no antibiotics and 0.08% Bluo-gal (Gold Biotechnology). The plates were left to sit for 2 min undisturbed at room temperature, and then plates were incubated, without inverting, at 37 °C overnight. Phage titers were determined by quantifyingblue plaques. For higher-throughput plaquing, the reagents were adjusted for the wells of a 12- wcll plate as follows: 450 pL of top agar, 10 pL of phage, and 100 pL of cells.[02491 Overnight propagation assay For each replicate, a single colony of an E. coli host strain was picked and grown overnight at 37 °C with shaking at 230 RPM in DRM and appropriate antibiotics. Saturated cultures were back-diluted 50-fold into DRM with appropriate antibiotics and grown for -2 h at 37 °C with shaking at 230 RPM until they reached an ODeoo of ~0.4. 1 mL of culture was then added to a 96-well deep well plate (Axygen) and infected with 1E5 total phage. This mixture was then grown overnight at 37 °C and 230 RPM, and then centrifuged for 10 min at 3400 g. Phage-containing supernatant was then collected and plaqued to determine the total number of output phage. Fold propagation was calculated by dividing the number of output phage by the number of input phage.

[0250] qPCR quantification of transposition efficiency in E. coli Quantification was performed as previously described with modification. Two primer pairs were designed: one pair specific to the AP-transposon junction generated by T-RL integration at the AP target site, the second pair specific to the AP backbone. Input E. coli lysate was prepared by resuspending a cell pellet from an overnight propagation assay in 1 mL water and incubating 50 L of this solution at 95 °C for 10 min. Standards were generated by mixing T-RL-integrated AP plasmid and unintegrated AP plasmid at varying ratios, corresponding to integration efficiencies spanning 0.0064-100%. qPCR reactions contained 10 pL Q5 Hot Start High-Fidelity 2x Master Mix (New England BioLabs), 0.50 pM of forward and reverse primer, 0.2 pL lOOx SYBR Green (Invitrogen), 4 pL of 100-fold diluted lysate or standard, and nuclease-free water to 20 pL total volume. qPCR was run on a BioRad CFX96 Real Time system with the following cycling conditions: 98 °C for 2 min; 40 cycles of 98 °C for 10 s, 60 °C for 20 s, and 72 °C for 15 s. Each sample was analyzed in two parallel reactions: one reaction with the primer pair specific to T-RL integration, the second reaction with the primer pair specific to the AP backbone. Transposition efficiency was calculated using a linear regression generated by the ACqvalues (Cqdifference between the two parallel qPCR reactions) of the standards reflecting known integration efficiencies.

[0251] Luciferase assay S2060 cells were transformed with necessary plasmids. Saturated overnight cultures of single colonies were diluted 250-fold into DRM media with appropriate antibiotics and grown for -3 h at 37 °C with shaking at 230 RPM. 100 pL of cells weretransferred to a 96-well black-walled clear-bottomed plate (Costar), then 600 nm absorbance and luminescence were read using a plate reader (Tccan). Values were reported as ODeoo-normalizcd luminescence.

[0252] General mammalian cell culture conditions HEK293T (ATCC CRL-3216), K562 (ATCC CCL-243), HeLa (ATCC CCL-2), and HuH7 (a gift from Erik Sontheimer’s group, originated from ATCC) cells were cultured and passaged in Dulbecco’s modified Eagle’s medium (DMEM) plus GlutaMAX (ThermoFisher Scientific) supplemented with 10% (v / v) fetal bovine serum (Gibco, qualified). All cell types were incubated, maintained, and cultured at 37°C with 5% CO2. Cell lines were authenticated by their respective suppliers and were negative for mycoplasma by testing with Myco Alert (Lonza Biologies).

[0253] Transfection protocol for genome editing in HEK293T cells and genomic DNA preparation HEK293T cells were seeded on 48-well poly-D-lysine coated plates (Corning) at a density of 40,000-45,000 cells per well. 16-24 h after seeding, cells were transfected at 60-80% confluency with 1.5 pL Lipofectamine 2000 (ThermoFisher Scientific). Initial CAST transfections used the following stoichiometry of components: 50 ng pCas6, 50 ng pCas7, 50 ng pCas8, 50 ng pTniQ, 150 ng pTnsAB, 150 ng pTnsC, and 300 ng pDonor-crRNA. For conditions targeting a plasmid substrate, 2 ng of pTarget was added. For these initial transfections, cells were cultured for 3 days following transfection. Following optimization of / A<?C AST transfection conditions, a new stoichiometry of components was implemented for all CAST transfections, unless otherwise stated: 50 ng pCas6, 50 ng pCas7, 50 ng pCas8, 50 ng pTniQ, 150 ng pTnsAB, 25 ng pTnsC, 300 ng pDonor-crRNA, and 20 ng pPuroR. ClpX was only delivered if specified, in which case 20 ng pClpX was added. For conditions using polycistronic TAeQCascade, 133 ng pQCascade was added. EePASSIGE experiments were performed as described previously (Pandey, S., et al., Nat. Biomed. Eng (2024)), scaled for a 48-well transfection: 250 ng prime editor plasmid, 37.5 ng of each pegRNA plasmid, 250 ng Bxbl plasmid, and 375 ng donor plasmid. At 24 h post-transfection, media was exchanged for fresh DMEM + 10% FBS containing 1 pg / mL puromycin (ThermoFisher Scientific) to select for transfected cells. Unless otherwise stated, cells were cultured for 3 days following media change, for a total of 4 days incubation post-transfection. For conditions targeting a plasmid substrate, 2 ng of pTarget was added. EePASSIGE experiments were performed as described previously (Pandey, S., et al., Nat. Biomed. Eng 9, 22-39 (2025)), scaled for a 48-well transfection: 250 ngPEmax plasmid, 37.5 ng of each pegRNA plasmid, 250 ng Bxbl plasmid, and 375 ng donor plasmid. At 24 h post-transfcction, media was exchanged for fresh DMEM +10% FBS containing 1 pg / mL puromycin (ThermoFisher Scientific) to select for transfected cells. Unless otherwise stated, cells were cultured for 3 days following media change, for a total of 4 days incubation post-transfection.

[0254] For all experiments, at time of harvest the media was removed, the cells were washed with IxPBS solution (ThermoFisher Scientific), and genomic DNA was extracted via the addition of 100 pF of freshly prepared lysis buffer (10 mM Tris-HCl, pH 8.0; 0.05% SDS; 25 ug / mE proteinase K (ThermoFisher Scientific)) directly into each well of the tissue culture plate. The genomic DNA mixture was incubated at 37 °C for >1 h, followed by an 80 °C enzyme inactivation step for 30 min. This lysed mixture containing genomic DNA was stored at -20 °C until use.[0255| High-throughput sequencing of genomic DNA samples High-throughput sequencing was used to quantify integration efficiencies as previously described, with minor modification. Following genomic DNA isolation, 1 pF of the genomic DNA extract was used as input for the first of two PCR reactions. Genomic loci were amplified in PCR1 using Q5 Hot Start High- Fidelity 2x Master Mix (New England BioEabs). PCR1 was performed as follows: 98 °C for 3 min; 25 cycles of 98 °C for 15 s, 65 °C for 20 s, and 72 °C for 30 s; 72 °C for 2 min. PCR1 products were confirmed on a 1.5% agarose gel. 1 pF of PCR1 was used as an input for PCR2 to append Illumina barcodes. PCR2 was conducted for 10 cycles of amplification using Q5 Hot Start High-Fidelity 2x Master Mix (New England BioLabs). Following PCR2, samples were pooled and gel purified in a 1.5% agarose gel using a Qiaquick Gel Extraction Kit (Qiagen).Library concentration was quantified using the Qubit High-Sensitivity Assay Kit (ThermoFisher Scientific). Samples were sequenced on an Illumina MiSeq instrument (paired-end read, read 1: 200-300 cycles, read 2: 0 cycles) using an Illumina MiSeq 300 v2 Kit (Illumina).

[0256] Sequencing reads were demultiplexed using MiSeq Reporter (Illumina). A custom Python script was used for quantification of integration efficiencies, which aligned amplicons to either the unedited sequence or integrated sequence. Integration efficiency was calculated as: percentage of (number of integrated reads) / (number of integrated reads + number of unedited reads). This analysis pipeline was also used to determine the distribution of T-RL and T-LR insertion sites downstream of the target site. For detection of indels, amplicons were aligned toreference sequences using CRISPResso2. For detection of substitutions, amplicons were analyzed using a custom Python script.

[0257] ddPCR quantification of integration efficiency Droplet digital PCR (ddPCR) quantification of integration efficiencies was performed as described previously, with modification. Primer pairs spanned the genome-transposon junction, with probes designed to be specific to the most frequent T-RL integration site, which was determined via HTS (the same was done for T-LR detection). Because probes could partially hybridize to other integration sites, the set of integration products that could be detected by each primer pair / probe was determined using mock-integrated standards synthesized as eBlocks (Integrated DNA Technologies). For quantification of integration frequencies in genomic DNA samples, 1 pL of crude genomic DNA extract was added to a 25-pL (final volume) reaction mixture containing a final concentration of IxddPCR Supermix for Probes (no dUTP) (BioRad), 900 nM of each reference primer, 900 nM of each target primer, 250 nM reference probe (HEX-labeled), 250 nM target probe (FAM- labeled), and 0.2 U / pL Hindlll (New England BioLabs). For all assays, the reference primer pair and probe targeted ACTS (BioRad, unique assay ID: dHsaCNS 141996500), except for ACTB- targeting CAST conditions, which used GAPDH (BioRad, unique assay ID: dHsaCNS794216737). Droplet generation, PCR, and droplet reading steps were performed using the BioRad QX ONE platform. PCR was performed as follows: 95°C for 10 min; 40 cycles of 94°C for 30 s and 58°C for 2 min; and 98°C for 10 min. Data were analyzed using the BioRad QX ONE software 1.3, Standard Edition, according to the manufacturer’s instructions.Integration efficiency was calculated as: percentage of (concentration of integrated molecules) / (concentration of reference molecules). All reported integration efficiencies, unless otherwise noted, were determined using primer pair / probes specific to T-RL integration, which comprised >95% of total integration events for evoCAST. Because T-RL probes often did not detect all possible T-RL integration sites, reported efficiencies are likely underestimates of true integration efficiencies. For eePASSIGE experiments, previously reported (Pandey, S., et al., Nat. Biomed. Eng 9, 22-39 (2025)) primer pairs were used, with ddPCR reactions performed and analyzed as described above.

[0258] Cell viability assay to assess ClpX toxicity HEK293T cells were seeded on 96-well poly-D-lysine coated plates (Coming) at a density of -5,000 cells per well. 18-24 h after seeding, cells were transfected at 60-80% confluency with 150 ng of plasmid expressingmCherry, ClpX, or ClpX with catalytic inactivating mutations in 1 pL Lipofectamine 2000 (ThermoFisher Scientific). Each day, starting on the day of transfection (DO), cell viability was measured with the CellTiter-Glo2.0 assay (Promega) according to the manufacturers protocol. Luminescence was measured in 96-well flat-bottomed polystyrene microplates (Corning) using a M1000 Pro microplate reader (Tecan) with a 1-s integration time.

[0259] Generation ofAclpX E. coli strain for PACE

[0260] Lambda red recombineering was performed as described previously (M. S. Morrison, et al. Nat Commun 12, 5959 (2021)), with modification, to generate a S1021-derivative E. coli strain that lacked endogenous clpX. In brief, SI 021 transformed with pKD119 were grown to OD600 ~0.6 at 30° C with shaking at 230 RPM in SOC media (New England BioLabs) with tetracycline and 2 mM arabinose. Mid-log cells were made electrocompetent via washing with ice-cold 10% glycerol and then electroporated with double-stranded donor DNA containing FRT-KanR-FRT and homology arms targeting the flanking genomic regions of clpX. Electroporated cells were allowed to recover in 1 mL SOC (New England BioLabs) at 30° C overnight with shaking at 230 RPM. This overnight culture was plated on 2xYT agar with kanamycin and incubated at 30° C overnight. An individual recombinant colony, confirmed via PCR and Sanger sequencing (Quintara Biosciences), was grown overnight in 2xYT with kanamycin at 37° C with shaking at 230 RPM to cure the temperature- sensitive pKD119 plasmid. These cells were then made chemically competent and transformed with pBAD-Flp (Addgene plasmid # 122969). Transformed cells were plated on 2xYT agar with chloramphenicol and 10 mM arabinose, and incubated at 30° C overnight to allow the KanR cassette to recombine out of the genome. An individual colony that successfully recombined out the KanR cassette was identified via PCR and was then grown overnight in 2xYT at 37° C with shaking at 230 RPM to cure the temperature- sensitive pBAD-Flp. This overnight culture was plated on 2xYT agar containing streptomycin (selecting for the strain) and incubated overnight at 37° C. Individual AclpX colonies were isolated and confirmed to be sensitive to tetracycline and chloramphenicol (confirming pKD119 and pBAD-Flp were cured, respectively). Finally, to conjugate the F plasmid into the newly generated AclpX strain, a mid-log (OD600 -0.6) culture of F' donor E. coli was diluted 1:1000 into a mid-log (OD600 -0.6) culture of the newly generated F- AclpX strain and incubated for 1.5 h at 37° C. This culture was then plated on 2xYT agar with streptomycin (selecting for strain) and tetracycline (selecting for F plasmid), and1A-incubated overnight. An individual F' AclpX colony was grown overnight at 37° C in 2xYT with tetracycline and streptomycin, and the overnight culture was used to generate a glycerol stock. [02611 HEK293T fluorescent reporter assay for transposon-end binding and flow cytometry analysis Transposon-end binding transcriptional activation assays were performed as previously described. In brief, HEK293T cells were seeded on 48-well poly-D-lysine coated plates (Coming) at a density of 40,000-45,000 cells per well. 16-24 h after seeding, cells were transfected at 60-80% confluency using 1.5 pL lipofectamine 2000 (ThermoFisher Scientific). TnsB variants, fused at the C-terminus to VP64, were individually co-transfected with a GFP transfection marker and a reporter plasmid containing a PseCAST transposon end adjacent to a minimal CMV promoter and a tdTomato marker at a ratio of 200 ng: 20 ng: 60 ng (TnsB:GFP:reporter). 48-72 h post-transfection, cells were analyzed via flow cytometry on a Novocyte Penteon. GFP positive cells were analyzed for tdTomato fluorescence. The bulk mean fluorescence intensity (MFI) was calculated for each transfection and normalized to a transfection in which no TnsB-VP64 transcriptional activator was added.

[0262] Western immunoblotting Western immunoblotting assays were performed as previously described. In brief, HEK293T cells were seeded on 48-well poly-D-lysine coated plates (Coming) at a density of 40,000-45,000 cells per well. 16-24 h after seeding, cells were transfected at 60-80% confluency using 1.5 pL lipofectamine 2000 (ThermoFisher Scientific). TnsAB variants were cloned with an internal 3xFLAG-bipartite-NLS fusion, and 200 ng of each variant was individually transfected into individual wells. 48-72 h post-transfection, cells were lysed in lysis buffer (150 mM NaCl, 0.1 % Triton X-100, 50 mM Tris-HCl (pH 8.0), cOmplete EDTA-free protease inhibitor (Roche)). Proteins were resolved by SDS-PAGE and transferred to a PVDF membrane (ThermoFisher Scientific). The membrane was then washed with TBS-T (50 mM Tris-Cl (pH 7.5), 150 mM NaCl, 0.1% Tween 20) and blocked with blocking buffer (TBS-T with 5% w / v BSA). Membranes were stained with either anti-FLAG M2 antibody (Sigma F3165, diluted 1:10000) or -Actin antibody (Cell signaling #3700, diluted 1:10000) overnight at 4 °C under gentle rotation. Membranes were then stained with HRP-conjugated secondary antibodies (ab97240, ab97250; diluted 1:10000) at room temperature for 1 h. Membranes were washed and developed with SuperSignal West Dura (ThermoFisher Scientific). Band intensities were quantified using Image Lab (BioRad), and the solubility of TnsAB valiants was determined bydividing FLAG intensities by P-Actin intensities. Solubilities were normalized to that of wildtype P.seTnsAB.[02631 Long-read sequencing of integration products Detection of cointegrates followed a previously described protocol with minor modification. HEK293T cells were transfected as described above for CAST integration assays. Approximately 96 h post-transfection, cells were lysed and DNA was harvested as previously described. Two separate PCRs were performed with equivalent volumes of input lysate, with primer sequences. The first PCR reaction contained a primer that annealed to the pTarget upstream of the target sequence (“Pl”) and a primer that annealed to the left transposon end (“P2”). The second PCR contained the same forward primer (Pl) but contained a reverse primer that annealed to the pDonor backbone downstream of the transposon (“P3”), such that only cointegrates should be amplified. PCRs were then pooled and purified via lx bead cleanup (Omega). Purified samples were then prepped for Nanopore sequencing using the Native Barcoding Kit 24 V14 (Oxford Nanopore, SQK-NBD 114.24) and loaded onto a R10.4.1 flow cell, sequencing for 18-24 h. Reads were analyzed using BBDuk from the BBTools suite (v.38.00; sourceforge.net / projects / bbmap). Reads with a minimum quality score of 8 were filtered to contain the upstream target region, the right transposon end, and the left transposon end. Filtered reads that contained the pDonor backbone sequence were considered a cointegrate sequence, while filtered reads that did not contain this sequence were considered a simple insertion. pDonor-backbone containing reads were also aligned to the expected sequence of a cointegrate and manually inspected to ensure accuracy. The frequency of cointegrates was calculated as percentage of (number of cointegrate reads ) / (number of simple insertion reads - number of cointegrate reads). The denominator corrects for the double-counting of cointegrate products as simple insertion reads, since the cointegrate product amplifies with both primer pairs. Analyses of transfections with defined ratios of plasmids containing mock simple insertion and co-integrate products generated a standard curve that was used to calculate co-integrate product frequencies for experimental conditions.

[0264] For long-read sequencing comparing evoCAST and eePASSIGE integration products, PCR reactions were first purified via 0.9x magnetic bead cleanup (Omega). Purified samples were then prepped for Nanopore sequencing using the Native Barcoding Kit (Oxford Nanopore SQK-NBD114.96). Samples were loaded onto a RIO.4.1 flow cell, sequencing for 18-24 h using the super-accurate base calling settings. Samples were processed using porechop with defaultsettings and a minimum split read length of 500 bp. Reads containing the primer binding sites used for PCR amplification (with a Hamming distance < 2) were extracted and aligned to a reference locus of either a mock CAST integrant (T-RL, 49-bp 40 insertion site) or a mock PASSIGE integrant using minimap2 and the parameters “-ax map-ont — sam-hit-only — secondary=no”. The resulting bam files were analyzed for insertions and deletions (indels) via the CIGAR string of each read alignment, and the proportion of indels was calculated relative to the total number of reads at a position. The background sequencing error frequency at each position was calculated using PCR amplicons generated from synthetic mock-integrated fragments (IDT eBlocks). This error frequency was subtracted from the calculated indel frequencies. In the region of the integration product that was covered by both PCR amplicons, the average indel proportion of both amplicons was reported.

[0265] UDiTaS sample preparation, sequencing, and computational analysis HEK293T cells were transfected as described above for CAST integration assays, except the pPuroR was omitted, and the transposon in the pDonor was modified to contain a promoter-driven puromycin cassette and an N10 UMI immediately flanking the transposon right end. Following transfection, cells were placed on puromycin selection for 7 days. Cells were then lysed as described for CAST integration assays, and the genomic DNA (gDNA) was purified via bead cleanup (Omega) and quantified using the Qubit High-Sensitivity Assay Kit (ThermoFisher Scientific).

[0266] TnY was purified as previously described (N. Liscovitch-Brauer, et al., Nat Biotechnol 39, 1270-1277 (2021)), preloaded with full-length Nextera Read 2 / Indexed oligos, and diluted to the appropriate working concentration such that 100 ng of gDNA would be tagmented into ~2 kb fragments. 100 ng of gDNA were tagmented as previously described with modification: following tagmentation, reactions were incubated with 0.4 U of Proteinase K (NEB) for 10 min at 55 °C to ensure release of the transposase from the gDNA. Reactions were then column purified with a DNA Clean & Concentrator Kit (Zymo), and eluted in 25 pL. An initial PCR1 was performed with a forward primer that anneals to the transposon cargo upstream of the N10 UMI, and a reverse primer that anneals to the P7 adapter sequence installed via tagmentation. PCR1 was performed KAPA HiFi Hotstart (Roche) as follows: 98°C for 5 min; 15 cycles of 98°C for 20 s, 55°C for 30 s, and 72 for 1 min; and 72 °C for 5 min. PCRls were purified using Omega Mag-Bind TotalPure magnetic beads (Omega Bio-Tek) at a ratio of 0.9x and eluted into 50 pL nuclease-free water. 2 pL eluted DNA was used as input for PCR2, which appendedIllumina sequencing adapters. After 15 cycles of PCR2 (same conditions as PCR 1), the reaction was resolved on a gel, and a smear corresponding to a size range of 350-800 bp was extracted. Samples were sequenced on an Element Biosciences AVITI instrument (paired-end read, read 1: 150 cycles, read 2: 150 cycles) using a Cloudbreak Freestyle Kit (Element Biosciences).

[0267] Reads were processed using a custom Python script. In brief, reads were first trimmed and quality filtered using cutadapt (v4.2, -a CTGTCTCTTATACACATCT -A CTGTCTCTTATACACATCT -minimum-length 15 -q 20; CTGTCTCTTATACACATCT is SEQ ID NO: 20). After adapter trimming, reads were then filtered to contain the right transposon end using BBDuk from the BBTools suite (v.38.00; sourceforge.net / projects / bbmap). Reads aligning to transfected plasmids with a Hamming distance <3 were discarded. UMIs were extracted prior to mapping using umitools (vl.1.4, extract — bc-pattem=NNNNNNNNNN). Flank sequences that passed filtering were then mapped to the GRCh38 reference genome using Bowtie2 (v2.4.2, —very- sensitive -no-mixed -no-discordant). Alignments were UMI processed using umitools (dedup). Final, UMI-processed alignments were manually inspected, and insertion events were defined as meeting the following criteria: >1 mapped read per UMI; <3 mismatches in the genomic alignment; paired reads mapped within 1200 bp of each other; and a primary read alignment. Insertion events were considered on-target if they occurred <100 bp downstream of the target site.

[0268] DNA from E. coli PACE host cells was prepared for UDiTaS by resuspending pelleted host cells in water, incubating at 95 °C for 10 min, and purifying DNA via bead cleanup (Omega). DNA was quantified using the Qubit High-Sensitivity Assay Kit (ThermoFisher Scientific), and 100 ng of DNA was tagmented and analyzed as described above, except in this case alignment was to a custom reference sequence containing the E. coli DH10B genome (from which the S2060 PACE strain is derived) and all PACE selection circuit plasmid sequences.

[0269] Analysis of ATAC-seq data ATAC-seq data for HEK293 cells were obtained from GSE1O8513 in BigWig (.bw) format and converted to BedGraph format using the UCSC Genome Browser utility BigWigToBedGraph. For characterization of off-target events, the 200- bp sequences surrounding each of the mapped off-target integration sites (previously aligned to the GRCh38 genome) were re-aligned to the GRCh37 genome (used in the published ATAC-seq dataset) using Bowtie2 with ‘-end-to-end’ settings, with 55 of 56 off-targets successfully aligning. Off-target locations and the 1-kb surrounding sequences were then extracted usingsamtools, and the average ATAC-seq score for each 1-kb region was fetched using the bedtools map function. The average ATAC-seq score of all off-target regions was then calculated. To perform a bootstrap analysis of randomly sampling 1-kb regions of the GRCh37 genome, the bedtools random function was used to generate 20,000 unique sets of 55 1-kb regions. The ATAC-seq scores for the 55 regions within each set were fetched using bedtools map, and the averages of each set were used to generate the histogram shown in FIG. 7F.

[0270] To assess the relationship between ATAC-seq signal and on-target evoCAST integration efficiencies, the average ATAC-seq score for each 1-kb region surrounding the target site was fetched using the bedtools map function. Pearson correlation analysis was performed using Prism 10 (GraphPad) to assess the relationship between on-target integration efficiency and the corresponding mean ATAC-seq score.

[0271] HEK293T fluorescent reporter assay for off-target integration and flow cytometry analysis A reporter assay for off-target integration in HEK293T cells was performed as described previously (Pandey, S., et al., Nat. Biomed. Eng (2024)), with minor modifications. Briefly, HEK293T cells were seeded on 48-well poly-D-lysine coated plates (Coming) at a density of 40,000-45,000 cells per well. 16-24 h after seeding, cells were transfected at 60-80% confluency using 1.5 pL Lipofectamine 2000 (ThermoFisher Scientific). TnsABC conditions contained 300 ng pDonor, 100 ng pTnsAB, and 25 ng pTnsC. Bxbl conditions contained 450 ng pDonor and 300 ng pBxbl (plasmid amounts were kept consistent with previously optimized amounts for efficient CAST and eePASSIGE activity). For all conditions, the donor contained mCherry under a cytomegalovirus (CMV) promoter. Cells were passaged for 2 weeks posttransection to dilute the pDonor. Cells were then trypsinized, resuspended in lx PBS solution, and analyzed via flow cytometry using the Sony MA900 Cell Sorter (Sony Biotechnology) at the Broad Institute flow cytometry core.

[0272] Generation of a clonal evoCAST-edited HEK293T cell line HEK293T cells were transfected according to the above protocol for CAST integration, except the pPuroR was omitted. To enable selection for edited cells, the transposon contained a splice acceptor and puromycin resistance gene, such that integration into the transcriptionally active AAVS1 locus would enable puromycin resistance. Cells were placed under selection with 1 pg / mL puromycin (ThermoFisher Scientific) 4 days post-transfection and passaged for > 1 month (this long timeline was chosen to demonstrate the durability of CAST-edited cells in a bulk transfectedpopulation). Cells were then single-cell sorted into poly-D-lysine coated 96-well plates (Coming) using a MA900 Cell Sorter (Sony) with the single cell 3-drop setting. Sorted cells were monitored after sorting, and wells with single colonies were marked for further analysis. After the cells had expanded for ~10 days, marked wells were split into two separate poly-D-lysine coated 96-well plates (Corning). After additional expansion for 3-5 days, one plate of the expanded cells was harvested for analysis of integration efficiency by ddPCR. Clonal cell lines with detectable integration events via ddPCR were further expanded for cell line generation.

[0273] Transfection ofHeLa and HuH7 cells For HeLa cell transfections, cells were seeded on 48-well poly-D-lysine coated plates (Coming) at a density of 30,000 cells per well. Between 16-24 h after seeding, cells were transfected at 60-80% confluency with 1 pL Lipofectamine 3000 (ThermoFisher Scientific). For HuH7 cell transfections, cells were seeded on 48-well poly- D-lysine coated plates (Coming) at a density of 40,000 cells per well. Between 16-24 h after seeding, cells were transfected at 60-80% confluency with 1.5 pL Lipofectamine 2000 (ThermoFisher Scientific). For both HeLa and Huh7 cells, transfections used 50 ng pCas6, 50 ng pCas7, 50 ng pCas8, 50 ng pTniQ, 150 ng pTnsAB, 25 ng pTnsC, 300 ng pDonor-crRNA, and 20 ng pPuroR. At 24 h post-transfection, media was exchanged for fresh DMEM + 10% FBS containing 1 ug / mL puromycin (ThermoFisher Scientific) to select for transfected cells. Cells were cultured for 3 days following media change, for a total of 4 days incubation posttransfection. Genomic DNA isolation was performed as described above for HEK293T transfections.

[0274] Nucleofection ofK562 cells For K562 nucleofections, 3 pg of CAST components (same ratio of components as used in HEK293T cell experiments: 215 ng pCas6, 215 ng pCas7, 215 ng pCas8, 215 ng pTniQ, 647 ng pTnsAB, 108 ng pTnsC, 1.29 pg pDonor-crRNA, and 86 ng pPuroR) were nucleofected in a final volume of 20 pL in a 16- well nucleocuvette strip (Lonza). Cells were nucleofected using the SF Cell Line 4D-Nucleofector X Kit (Lonza), with 500,000 cells per sample (program FF-120), according to the manufacturer’s protocol. At 24 h post-transfection, media was exchanged for fresh DMEM + 10% FBS containing 1 ug / mL puromycin (ThermoFisher Scientific) to select for transfected cells. Cells were cultured for 3 days following media change, for a total of 4 days incubation post-transfection. Genomic DNA isolation was performed as described above for HEK293T transfections.

[0275] Primary human fibroblast cell culture conditions and electroporation Following the Declaration of the Helsinki Principles, primary dermal fibroblast cells were obtained from a recessive dystrophic epidermolysis bullosa (RDEB) patient via a 3 mm punch biopsy. Tissue was minced and plated in a 6-well dish. Adherent cells were expanded and cultured in MEM-alpha complete media containing 20% (v / v) fetal bovine serum (Atlas Biologicals, F-0500-D), lx Glutamax, lx penicillin / streptomycin (Thermo-Fisher 15-140-122), lx Antioxidant Supplement (Sigma Aldrich, A1345), lx nonessential amino acids (Thermo Scientific, 11140050), 10 ng / mL epidermal growth factor (Sigma Aldrich, E4127), and 0.5 ng / mL fibroblast growth factor (Sigma Aldrich F3133). Fibroblasts were passaged when at 100% confluency. For evoCAST experiments, cells were cultured in a T-150 flask (Fisher Scientific 1012634) and harvested by trypsinization using Trypsin-EDTA (0.25%; Thermo Scientific 20 25200114). Cells were removed using complete media, and then washed and resuspended in Neon Buffer R.[0276| Per condition, two electroporations of 100,000 cells each were performed using the Neon NXT instrument (Thermo Fisher) with the following protocol: 1700 V, 20 ms, 1 pulse, using buffers R and E10 with 10-pL tips. Each electroporation (per 100,000 cells) contained: 200 ng pCas6, 200 ng pCas7, 200 ng pCas8, 200 ng pTniQ, 800 ng pTnsAB, 200 ng pTnsC, 582 ng linearized donor-crRNA, and 80 ng pClpX (if added). Donor-crRNA was delivered as linear DNA to mitigate total DNA mass and related DNA toxicity. Both electroporation reactions were combined into one well of a 24-well plate (thus constituting a single replicate) containing MEMalpha complete media as above, but without penicillin and streptomycin. Complete, conditioned media from parental fibroblasts was added ~4 hours post-electroporation and then replenished daily until harvest. Cells were harvested five days post-nucleofection. Media was first removed, then cells were trypsinized (Trypsin-EDTA (0.25%); Thermo Scientific 25200114) and harvested in the MEM-alpha media. Resuspended cells were pelleted by centrifugation, and genomic DNA was isolated using the Monarch Genomic DNA Purification Kit (New England BioLabs, T3010S). Maintaining high cell density post-electroporation was critical to preserving cell viability and observing higher editing, as the above electroporation protocol using a 12-well plate, rather than a 24-well plate, led to substantially fewer viable cells and an associated lower editing (<1%).

[0277] To assess viability, cells were imaged 48 h post-electroporation with a Leica DMil microscope (Thomas Scientific) using the 5x magnification objective.

[0278] Quantification of transgene expression via RT-ddPCR To assess transgene expression, mRNA was isolated from HEK293T cells or HuH7 cells 4 days post-transfection using the RNeasy Plus kit (Qiagen). 400-800 ng of isolated RNA was treated with RQ1 RNase-free DNase (Promega) for 1 h at 37 °C in a 10 pL reaction, and then combined with 1 pL RQ1 DNase Stop Solution (Promega) and incubated at 60 °C for 10 min. 9 pL of DNase-treated RNA was used as input for a 20-pL reverse transcription reaction containing the SuperScript IV Vilo Master Mix (ThermoFisher Scientific), performed according to manufacturer’s protocols. 1 pL of the reverse transcription reaction was used as input for a 25-pL (final volume) ddPCR reaction containing a final concentration of IxddPCR Supermix for Probes (no dUTP) (BioRad), 900 nM of each reference primer, 900 nM of each target primer, 250 nM reference probe (HEX-labeled), 250 nM target probe (F AM-labeled), and 0.2 U / pL Hindlll (New England BioLabs). For MECP2 quantification in HEK293T cells, the target primer pair / probe was designed to be specific to the exon 1-exon 2 junction, where exon 2 of the integrated transgene was recoded such that the target primer pair / probe did not detect endogenous MECP2 expression. For F9 quantification in HuH7 cells, the target primer pair / probe was designed to be specific to the ALB exon 1-F9 exon 2 junction. The reference primer pair and probe targeted TBP (BioRad, unique assay ID: dHsaCPE5O58363). ddPCR and analysis was performed as described above for quantification of integration efficiencies. Transgene expression was reported as concentration of target transcripts divided by the concentration of TBP transcripts.

[0279] Phage -assisted noncontinuous evolution (PANCE) PANCE was performed as described previously (S. M. Miller, et al., Nat Protoc 15, 4101-4127 (2020), incorporated herein by reference). In brief, S2060 host cells transformed with selection plasmids were made chemically competent and transformed with mutagenesis plasmid (MP6), then plated on 2xYT agar containing 100 mM glucose and appropriate antibiotics. 8-12 colonies were picked into individual wells of a 96-well deep well plate (Axygen) containing 1 mL of DRM and appropriate antibiotics. Colonies were resuspended and serially diluted 10-fold, seven times into DRM. The plate was grown at 37°C with shaking at 230 RPM overnight for 16-18 h. Wells containing dilutions with OD600 -0.3-0.4 were combined, then treated with 10 mM arabinose to induce mutagenesis. This mixture was distributed into 1-mL cultures in a 96-well deep well plate (Axygen). The cultures were then infected with SP at the indicated dilution (aiming for ~1E5 input phage). Infected cultures were grown overnight for 16-18 h at 37 °C and harvested the nextday by centrifugation for 10 min at 3400 g. 100 pL of the SP-containing supernatant was transferred to a 96-wcll PCR plate (ThermoFisher Scientific), sealed with foil and stored at 4°C. SP were then used to infect the next passage, and the process was repeated for the duration of the selection. Phage titers were determined by qPCR as described previously or by the plaque assay described above. If titers were low (<1E4 pfu / mL), a passage of drift was performed. For drift passages, S2208 encoding MP6 were used instead of selection strains. In drift passages, SP were only allowed to propagate for 6-8 h instead of overnight to minimize the likelihood of recombination of gill into the SP genome. Following completion of a PANCE campaign (upon a noticeable change in phage propagation on the selection strain), SP were plaqued using S2208 cells or the selection strain. The evolved genes of interest from individual plaques were then amplified by PCR, as described previously, and submitted for Sanger (Quintara Biosciences) or Nanopore (Plasmidsaurus) sequencing to generate inputs for Mutato analyses (hub . docker, com / r / araguram / mutato) .

[0280] Phage-assisted continuous evolution (PACE) PACE was performed as previously described (S. M. Miller, et al., Nat Protoc 15, 4101-4127 (2020), incorporated herein by reference). Briefly, host cells containing the mutagenesis plasmid were prepared as described for PANCE above. 12 colonies were picked into individual wells of a 96-well deep well plate (Axygen) containing 1 mL of DRM and appropriate antibiotics. Colonies were resuspended and serially diluted 10-fold, seven times into DRM. The plate was grown at 37°C with shaking at 230 RPM overnight for 16-18 h. Wells containing dilutions with OD600 ~0.3-0.4 were combined and used to inoculate a chemostat containing 100 mL of DRM. The chemostat was grown to OD600 -0.4-0.8 and then continuously diluted with fresh DRM at a rate of 1-1.5 chemostat volumes per hour to keep cell density constant. The chemostat was maintained at a volume of 80-100 mL.

[0281] Before SP infection, lagoons were filled with 15 mL of culture from the chemostat and pre-induced with 10 mM arabinose for at least 1 h. Lagoons were infected with SP at a high starting titer (typically -108 pfu / mL). To increase stringency, the lagoon dilution rates were increased over time as indicated. During the evolution, samples (-500 pL) of the lagoon were collected from the lagoon waste lines at the indicated times. Samples were centrifuged at 4,000 g for 5 minutes, and the SP-containing supernatant was stored at 4°C. Titers of SP samples were determined by plaque assays.

[0282] Following completion of a PACE campaign (after the lagoon dilution rate was 3 vol / h for >24 h), final SP samples were plaqucd on the S2060 strain to determine whether glll- recombinant SP had formed during selection (which ‘cheat’ the selection by enabling activityindependent propagation). If cheating was observed (plaques on S2060 strain), individual plaques were amplified overnight in 1 mL DRM at 37 °C with shaking at 230 RPM, and then the culture was centrifuged at 4,000 g for 5 min. The cell pellet, containing SP-infected E. coli, was miniprepped to isolate the double- stranded DNA replicative form of the SP. This isolated SP DNA was then sent for Nanopore sequencing (Plasmidsaurus) to determine the sequence of the gill-recombinant SP, allowing inspection of the mechanism of gill acquisition. If cheating was not detected (no plaques on S2060 strain), final SP samples were plaqued using S2208 cells or the selection strain. The evolved genes of interest from individual plaques were then amplified by PCR, as described previously, and submitted for Sanger (Quintara Biosciences) or Nanopore (Plasmidsaurus) sequencing to generate inputs for Mutato analyses (hub . docker, com / r / araguram / mutato) .SequencesTn7016-TnsA (SEQ ID NO: 1)MYIRNLRKPSPNKNVFKFASTKVSSVVMCESSLEFDACFHHEYNDLIESFGSQPEGFKYE FMGKSLPYTPDALISYTDKTQKYHEYKPYSKIASPLFRAEFAAKRAASLKLGIDLVLVTD RQIRVNPILNNLKLLHRYSGVYGISGIQKELLSFIHKSGVIKLNDISSQVGIPIGETRSFLFG LMHKGLVKADEGCDDETNNPTEWATPTn7016-TnsB (SEQ ID NO: 2) MTDFFNEFDESLVPLKPQTPTQYVKLDDANLIQRDLDTFSDTFKNQALQRYKLISTIDKK LSRGWTQRNLDPILDELFKGGDVVRPNWRTVARWRKKYIESNGDIASLADKNHKMGN RTNRIKGDDKFFDKALERFLDAKRPTIATAYQYYKDLIVIENESIVEGKIPIISYNAFNKRI KAIPPYAVAVARHGKFKADQWFAYCAAHVPPTRILERVEIDHTPLDLILLDDELLIPIGRP YLTLLIDVFSGCVLGFHLSYKSPSYVSAAKAITHAIKPKSLDALNIELQNDWPCFGKFEN LVVDNGAEFWSKNLEHACQSAGINIQYNPVRKPWLKPFIERFFGVMNEYFLPELPGKTF SNILEKEEYKPEKDAIMRFSTFVEEFHRWIADVYHQDSNSRETRIPIKRWQQGFDAYPPL TMNEEEETRFSMLMRISDSRTLTRNGFKYQELMYDSTALADYRKHYPQTKETVKKLIK VDPDDISKIYVYLEELESYLEVPCTDPTGYTDGLSIYEHKTIKKINREVIRESKDSLGLAK ARMAIHERVKQEQEVFIESKTKAKITAVKKQAQIADVSNTGTSTIKVSEESAAPVQKHIS NDNSDDWDDDEEAFETn7016-TnsC (SEQ ID NO: 3) MNALTEIQIEKLRNFSDCIVMHPQIKTIFNDFDELRLNRKFQSDQQCMLLIGDTGVGKSH TINHYKKRVLATQNYSRNTMPVLVSRISRGKGLDATLVQMLADLELFGSSQIKKRGYKT DLTKKLVESLIKAQVELLIINEFQELIEFKSVQERQQIANGLKFISEEAKVPIVLVGMPWAAKIAEEPQWASRLVRKRKLEYFSLKNDSKYFRQYLMGLAKKMPFDVPPKLESKNTTIALFAACRGENRALKHLLLEALKLALSCNEYLENKHFITAYDKFDFFNDKEKLKSKNPFKQDIKDIEIYEVIKNSSYNPNALDPEDMLTDRVFAIVKTn7016-TniQ (SEQ ID NO: 4)MAFLFSPKARAFSDESLESYLLRVVSENFFDSYEGLSLAIREELHELDFEAHGAFPVDLK RLNVYHAKHNSHFRMRALGLLETLLDLPRYELQKLALLKSDIKFNSSVALYNNGVDIPL RFIRHHAEEAVDSIPVCSQCLAEEAYIKQSWHIKWVNACTKHQCALLHNCPECYAPINYI ENESITHCSCGFELSCASTSPVNTLSIEHLNKLLDKGERNDSNPLFNNMTLTERFAALLW YQERYSQTDNFCLNDAVNYFSKWPAVFNTELDELSKNAEMKLIDLFNKTEFKFIFGDAI LACPSTQKQSESHFIYRALLDYLVTLVESNPKTKKPNAADLLVSVLEAATLLGTSVEQV YRLYQNGILQTAFRHKMNQRINPYKGAFFLRHVIEYKTSFGNDKARMYLSAWTn7016-Cas8-Cas5 fusion (SEQ ID NO: 5)MHLKELLEITDTTERDRSLRRAFSPYTAMIDITGSEAVALIILLNLTYRKNQVDDLLDKKL AKQALKSEDHINKCIKEIAWFHTHNLKYPDIRVSKQNLAVEPPTLHSYVLSSANYPKAY GWSHNSAKVNFAKLFVSYFKWQNQVSWLAQVLATNSDNWKSAFTSLGLSVKAFKSLC VTVKNSLPEEAIPDSVDRYSRQIRMPYHDGYLAVTPVISHVVQSKIQQAAIDKRARFSNV EFTRPAAVSMLAASLGGVINVLNYPPYIRSKYHGLSNSRAFKLNNGQTVFNVEALLKPE LIKALEGIIFSNNALALKQRRQQKVKNIKELRNTLLEWFSPVFEWRLDAIENGYDLEQLE SASERLEYKILSLPDNELPSLTIPLFRLLNEMLGGVSMTQRYAFHPKLMSPLKAALQWLL VNLTDQKHVLIEEDDEHYRYLHLSGIRVFDAQALSNPYCSGIPSLTAVWGMIHSYQRKL NEALGTNVRFTSFSWFIRNYSAVAGKKLPELSLQGAQQSRLKRPGIIDGKYCDLVFDLII HIDGYEDDLQAVDSKPDILKAHFPSNFAGGVMHQPELNSNINWCCLYSNENQLFEKLRR LPLSGCWVMPTEHKIQDLDELLLLLNSDSKLSPSMMGYMLLTEPMARVGSLERLHCYA EPAIGVVKYEAATSVRLKGIGNYFNSAFWMLDAQEKFMLMKKVTn7016-Cas7 (SEQ ID NO: 6)MELCNILKYDRSLYPGKAVFFYKTADSDFVPLEADINKIRGPKSGFTEAFTPQFSPKNISP QDLTHNNILTLEECYVPPNVEHIFCRFSLRVQANSLVPSGCSDPEVFSLLKELAETFKECG GYKELAVRYCRNILIGTWLWRNQNTGNTQIEIKTSKGSCYLIDNTRKLAWESKWASDDL KVLEELSNEIESALTDPNVFWSADITAKIEASFCQEIYPSQILNDKVKQGEASKQFVKAKC ADGRYAVSFNSVKIGAALQSIDDWWDEDASKRLRVHEFGADKEIGVARRPPDSEQNFY SIFKNTEWYLSALKNCITNKNEKIDPAIYYLFSVLIKGGMFQKKAEAKKATn7016-Cas6 (SEQ ID NO: 7)MQRYYFTVHFLPKQANLALLTGRCISIMHGFILKHNIEGMGVTFPAWSDSSIGNEIAFVY TDKEILNTLKDQAYFVDMQDCGFFKVSQVLAVPDSCEEVRFIRNQAVAKIFTGESRRRL KRLQKRALARGEDFNPKKIEAPREIDIFHRVAMTSKSSQEDYILHIQKQDVDCQAEPYFS NYGLASNEKFKGTVPDLSPSIDRNEngineered Cas7 (SEQ ID NO: 224)MELCNILKYDRSLYPGKAVFFYKTADSDFVPLEADINKIRGPKSGFTEAFTPQFSPKNISP QDLTHNNILTLEECYVPPNVEHIFCRFSLRVQANSLVPSGCSDPEVFSLLKELAETFKECG GYKELAVRYCRNILIGTWLWRNQNTGNTQIEIKTSKGSCYLIDNTRKLAWESKWASDDL KVLEELSNE1ESALTDPNVFWSAD1TAK1EASFCQE1YPSQ1LNDKVKQGEASKQFVKAKC ADGRYAVSFNSVKIGAALQSIDDWWDEDASKRLRVHEFGADKEIGVARRPPDSEQNFY SIFKNTEWYLSALKNCITNKNEKIDPAIYYLFSVLIKGGMFQKKAEKKKAEngineered Cas8 (SEQ ID NO: 225)MHLKELLEITDTTERDRSLRRAFSPYTAMIDITGSEAVALIILLNLTYRKNQVDDLLDKKLAKQALKSEDHINKCIKEIAWFHTHNLKYPDIRVSKQNLAVEPPTLHSYVLSSANYPKAYGWSHDSAKVNFAKLFVSYFKWQNQVSWLAQVLATNSDNWKSAFTSLGLSVKAFKSLCVTVKNSLPEEA1PDSVDRYSRQ1RMPYHDGYLAVTPV1SHVVQSK1QQAA1DKRARFSNVEFTRPANVSMLAASLGGVINVLNYPPYIRSKYHGLSNSRAFKLNNGQTVFNVEALLKPELIKALEGIIFSNNALALKQRRQQKVKNIKELRNTLLEWFSPVFEWRLDAIENGYDLEQLESASERLEYKILSLPDNELPSLTIPLFRLLNEMLGGVSMTQRYAFHPKLMSPLKRALQWLLVNLTDQKHVLIEEDDEHYRYLHLSGIRVFDAQALSNPYCSGIPSLTAVWGMIHSYQRKLNEALGTNVRFTSFSWFIRNYSAVAGKKLPELSLQGAQQSRLKRPGIIDGKYCDLVFDLIIHIDGYEDDLQAVDSKPDILKAHFPSNFAGGVMHQPELNSNINWCCLYSNENQLFEKLRRLPLSGCWVMPTEHKIQDLDELLLLLNSDSKLSPSMMGYMLLTEPMARVGSLERLHCYAEPAIGVVKYEAATSVRLKGIGNYFNSAFWMLDAQEKFMLMKKVEvolved TnsA (SEQ ID NO: 226)MYIRNLRKPSPNKNVFKFASTKVSSVVMCESSLEFDACFHHEYNDLIESFGSQPEGFKYEFMGKSLPYTPDALISYTDKTQKYHEYKTYSKIASPLFRAEFAAKRAASLKLGIDLVLVTDRQIRVNPILNNLKLLHRYSGVYGISGVQKELLSFIHKSGVIKLNDISSQLGIPIGETRSLLLGLMHKGLVKADLGCDDLTNNPTLWATPEvolved TnsB (SEQ ID NO: 227)MTDFFNEFDESLVPLKPQTPTQYVKLDDANLIQRDLDTFSDTSKNQALQRYKLISTIDKKLSRGWTQRNLDPILDELFKGGDVVRPNWRTVARWRKKYIESNGDIASLADKNHKMGNRTNRIKGDDKFFDKALERFLDAKRPTIATAYQYYKDLIVIENESIVEGKIPIISYNAFNKRIKAIPPYAVAVARHGKFKADQWFAYCAAHVPPTRILERVEIDHTPLDLILLDDELLIPIGRPYLTLLIDVFSGCVLGFHLSYKSPSYVSAAKAITHAIKPKSLDALNIELQNDWPCFGKFENLVVDNGAEFWSKNLEHACQSAGINIQYNPVRKPWLKPFIERFFGVMNENFLTELPGKTFSNILEKEEYKPEKDAIMRFSTFVEEFHRWIVDVYHQNSNSRETRIPIKRWKQGFDAYPPLTMNEEEETRFSMLMRISDSRTLTRNGFKYQELMYDSTALADYRKRYPQTKETVKKLIKVDPDDISKIYVYLEELESYLEVPCTDPTGYTDGLSIYEHKTIKKINREEIRESKDSLGLAKARMAIHERVKREQEVFIESKTKAKITAVKKQAQIADVSNTGTSTIKVSEESAAPVLKHISNDNSDDWDDDLEAFEEvolved TnsC (SEQ ID NO: 228)MNALTEIQIEKLRNFSDCIVMHPQIKTIFNDFDELRLNRKFQSDQQCMLLIGDTGVGKSHTINHYKKRVLATQNYSRNTMPVLVSRISRGKGLDATLVQMLADLELFGSSQIKKRGYKTDLTKKLVESLIKAQVELLIINEFQELIEFKSVQERQQIANGLKFISEEAKVPIVLVGMPWAAKIAEEPQWASRLVRKIKLEYFSLKNDSKYFRQYLMGLAKKMPFDVPPKLESKNTTIALFAACRGENRALKHLLLEALKLALSCNEYLENKHFITAYDKFDFFNDKEKLKSKNPFKQDIKDIEIYEVIKNSSYKPNALDPEDMLTDRVFAIVKTable 2 - Nucleic acid sequences encoding the engineered components - Nuclear localization sequence(s) are shown in bold, linker sequences are shown in italics, and engineered / evolved CAST protein components are underlinedTable 3 - E. coli expression plasmidsTable 4 - Selection phages (SPs)Table 5 - Mammalian cell expression plasmidsTable 6 - CAST crRNAs used in E. coli experimentsTable 7 - CAST crRNAs used in mammalian cell experimentsTable 8 - eePASSIGE pegRNAsTable 9 - Transposon cargoes used in mammalian cell experiments

[0283] The scope of the present invention is not limited by what has been specifically shown and described hereinabove. Those skilled in the art will recognize that there are suitable alternatives to the depicted examples of materials, configurations, constructions, and dimensions. Variations, modifications, and other implementations of what is described herein will occur to those of ordinary skill in the art without departing from the spirit and scope of the invention.

[0284] Numerous references, including patents and various publications, are cited and discussed in the description of this invention. The citation and discussion of such references is provided merely to clarify the description of the present invention and is not an admission that any reference is prior art to the invention described herein. All references cited and discussed in this specification are incorporated herein by reference in their entirety.

Claims

CL IMSWhat is claimed is:

1. A system for nucleic acid modification comprising: an engineered Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)- associated transposon (CAST) system or one or more nucleic acids encoding the engineered CAST system, wherein the engineered CAST system comprises: a) one or more Cas proteins comprising: a Cas8-Cas5 fusion protein comprising an amino acid sequence having amino acid substitutions at positions 125, 244, and 410 relative to SEQ ID NO: 5, a Cas7 protein comprising an amino acid sequence having an amino acid substitutions at position 347 relative to SEQ ID NO: 6; and / or b) one or more transposon-associated proteins comprising: a TnsA protein comprising an amino acid sequence having amino acid substitutions at positions 88, 147, 170, 180, and 182 relative to SEQ ID NO: 1, a TnsB protein comprising an amino acid sequence having amino acid substitutions at positions 43, 349, 352, 390, 396, 410, 464, 526, 549, and 594 relative to SEQ ID NO: 2, a TnsC protein comprising an amino acid sequence having amino acid substitutions at positions 197 and 314 relative to SEQ ID NO: 3.

2. The system of claim 1, wherein the Cas8-Cas5 fusion protein comprises an amino acid sequence having amino acid substitutions N125D, A244N, and A410R relative to SEQ ID NO: 5; the Cas7 protein comprises an amino acid sequence having an amino acid substitution A347K relative to SEQ ID NO: 6; the TnsA protein comprises an amino acid sequence having amino acid substitutions P88T, 1147V, V170L, F180L, and F182L relative to SEQ ID NO: 1, the TnsB protein comprises an amino acid sequence having amino acid substitutions F43S, Y349N, P352T, A390V, D396N, Q410K, H464R, V526E, Q549R, and Q594L relative to SEQ ID NO: 2, and / orthe TnsC protein comprises an amino acid sequence having amino acid substitutions R197I and N314K relative to SEQ ID NO: 3.

3. The system of claim 1 or 2, wherein the Cas8-Cas5 fusion protein comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 225; the Cas7 protein comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 224; the TnsA protein comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 226, the TnsB protein comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 227, and / or the TnsC protein comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 228.

4. The system of any of claims 1-3, wherein the Cas8-Cas5 fusion protein comprises an amino acid sequence of SEQ ID NO: 225; the Cas7 protein comprises an amino acid sequence of SEQ ID NO: 224; the TnsA protein comprises an amino acid sequence of SEQ ID NO: 226, the TnsB protein comprises an amino acid sequence of SEQ ID NO: 227, and / or the TnsC protein comprises an amino acid sequence of SEQ ID NO: 228.

5. The system of any of claims 1-4, wherein the one or more Cas proteins further comprise a Cas6 protein comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 7; and / or the one or more transposon-associated proteins further comprise a TniQ protein comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 4.

6. The system of any of claims 1-5, wherein the TnsA protein and the TnsB protein are provided as a TnsA-TnsB fusion protein.

7. The system of any of claims 1-6, wherein any or all of the one or more Cas proteins and the one or more transposon-associated proteins comprise at least one nuclear localization sequence (NLS).

8. The system of any of claims 1 -7, wherein the one or more nucleic acids comprises one or more messenger RNAs, one or more vectors, or a combination thereof.

9. The system of any of claims 1-8, wherein the engineered CAST system further comprises a gRNA complementary to at least a portion of the target nucleic acid sequence, or a nucleic acid encoding the at least one gRNA.

10. The system of any of claim 1-9, further comprising a target nucleic acid sequence.

11. The system of any of claims 1-10, further comprising a donor nucleic acid flanked by at least one transposon end sequence.

12. A method for DNA integration comprising contacting a target nucleic acid sequence with a system of any of claims 1-12 or a composition thereof.

13. The method of claim 12, wherein the target nucleic acid sequence is in a cell and the contacting a target nucleic acid sequence comprises introducing the system into the cell.

14. The method of claim 13, wherein introducing the system into the cell comprises administering the system to a subject.

15. The method of claim 14, wherein the administering comprises in vivo administration or transplantation of ex vivo treated cells comprising the system.

Citation Information

Patent Citations

  • Crispr-transposon systems for DNA modification

    WO2022261122A1

  • Adaptations for high efficiency i-f3-crispr-CAS systems for guide RNA-directed transposition in human cells

    WO2023154826A2

  • Crispr-transposon systems and components

    WO2024173573A1