Modulators of transposon systems for genome editing

The transposon system with a CAST modulator enhances genome editing efficiency in diverse prokaryotic cells, addressing low editing rates in existing CAST systems and improving industrial and therapeutic applications.

WO2026059860A2PCT designated stage Publication Date: 2026-03-19RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

CRISPR-associated transposases (CASTs) face limitations in achieving efficient genome editing across a broad range of organisms, with low editing efficiencies in species like Corynebacterium glutamicum, human cells, and complex gut microbiomes, limiting their utility for industrial and therapeutic applications.

Method used

A transposon system comprising a nucleotide sequence encoding a CRISPR-associated transposase (CAST) complex, a guide RNA, a transposon flanked by recognition sites, and a CAST modulator that enhances editing efficiency, including small molecules promoting homologous recombination.

Benefits of technology

The system significantly improves genome editing efficiency in various prokaryotic cells, achieving higher editing rates compared to traditional CAST systems, particularly when combined with λ-Red recombineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000025_0001
    Figure IMGF000025_0001
  • Figure IMGF000026_0001
    Figure IMGF000026_0001
  • Figure IMGF000027_0001
    Figure IMGF000027_0001
Patent Text Reader

Abstract

The present disclosure provides a transposon system comprising i) a nucleotide sequence encoding polypeptides that form a CRISPR-associated transposase (CAST) complex; ii) a nucleotide sequence encoding a guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome; iii) a transposon, or an insertion site for a transposon, wherein the transposon or the transposon insertion site is flanked by recognition sites that are recognized by the CAST complex; and iv) a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt: BERK-543WO MODULATORS OF TRANSPOSON SYSTEMS FOR GENOME EDITING CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Patent Application No.63 / 693,059 filed September 10, 2024, which application is incorporated herein by reference in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with government support under Grant Number DE-AC02-05CH11231 awarded by the US Department of Energy. The government has certain rights in the invention. INCORPORATION BYREFERENCE OFSEQUENCELISTINGPROVIDED AS A SEQUENCE LISTING XML FILE

[0003] A Sequence Listing is provided herewith as a Sequence Listing XML, “BERK- 543WO_SEQLIST.xml”, created on September 5, 2025 and having a size of 52,585,732 bytes. The contents of the Sequence listing XML are incorporated herein by reference in their entirety. INTRODUCTION

[0004] Site-specific insertion of large DNA segments into the genome of a single microorganism in isolation, or in a mixture of cells (a synthetic or natural microbial community, for example), remains challenging. CRISPR-associated transposases (CASTs) represent a powerful addition to the genome-editing toolbox. Unlike traditional CRISPR-Cas systems that introduce targeted double-stranded breaks and rely on endogenous repair machinery to introduce edits, CASTs integrate the RNA-guided targeting of CRISPR-Cas and Tn7-like transposition to make large programmable insertions. This unique mechanism circumvents the lethality often associated with double-strand breaks in bacteria due to inefficient non-homologous end joining. For these reasons, CASTs have been used for bacterial editing, manipulating microbial communities and even modifying human genomes.

[0005] Despite the established value of CASTs as a genome-editing tool, CASTs face limitations in achieving efficient editing across a broad range of organisms. For example, in the industrially important species Corynebacterium glutamicum, a CAST system derived from Vibrio cholerae (VchCAST) achieved an editing efficiency between 0-0.027%. In human cells, the challenges are even more pronounced. For instance, attempts to use VchCAST in HEK293T cells resulted in editing efficiencies below 0.01% for genomic targets, with slightly higher rates of 0.1-1% observed only for episomal targets. In complex gut microbiome samples, for example, VchCAST has shown editing efficiencies as low as 0.001%, significantly limiting its utility for in situ microbiome engineering. These low efficiencies restrict the potential of CAST systems for environmental, industrial, and therapeutic genome editing applications. Theres is a need in the art for improving the editing efficiencies of CAST systems – such is provided herein.Atty. Dkt: BERK-543WO SUMMARY

[0006] The present disclosure provides a transposon system comprising i) a nucleotide sequence encoding polypeptides that form a CRISPR-associated transposase (CAST) complex; ii) a nucleotide sequence encoding a guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome; iii) a transposon, or an insertion site for a transposon, wherein the transposon or the transposon insertion site is flanked by recognition sites that are recognized by the CAST complex; and iv) a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides. In some cases, the CAST modulator comprises a small molecule that promotes homologous recombination. The present disclosure provides a prokaryotic cell comprising a subject transposon system. The present disclosure additionally provides methods for editing the genome of a target prokaryotic cell(s) utilizing a subject transposon system. BRIEF DESCRIPTION OF THE DRAWINGS SUMMARY

[0007] FIG.1: Is a schematic depiction of a genome-wide screen for identifying modulators of CAST editing efficiency and testing of a screen-informed CAST modulator in diverse prokaryotes.

[0008] FIG.2A-2B: (A) Depicts a schematic a genome-wide screen to identify modulators of CAST editing efficiency. (B) Shows genes identified in the genome-wide screen that activate or inhibit CAST editing.

[0009] FIG.3A-3B: (A) Shows the effect that gene knockouts of CAST activators and inhibitors identified in the screen have on VchCAST editing efficiency compared to a negative control. (B) Shows the effect that gene knockouts of rec genes have on VchCAST editing efficiency compared to a Δyicl control.

[0010] FIG.4A-4C: (A) Depicts a schematic of the bacteriophage λ-Red genes (exo, beta, and gam) cloned into a VchCAST plasmid (R6K, PPmtl-catP cargo). (B) Shows that induced λ-Red VchCAST treatment increases editing efficiency compared to a VchCAST control. (C) Shows the effect that λ-gam VchCAST, λ-exo-beta VchCAST, and λ-Red VchCAST have on E. Coli compared to VchCAST and non-targeting λ-Red VchCAST-NT controls.

[0011] FIG.5: Shows λ-Red VchCAST editing efficiency in E. coli with varying levels of induction by Crystal Violet (CV).

[0012] FIG.6A-6C: (A) Depicts a schematic of a cointegrant insert, a T-RL simple insert, and a T-LR simple insert. (B) Shows the percentages of cointegrant, T-RL simple, and T-LR simple inserts in E. coli generated by λ-Red VchCAST and VchCAST. (C) shows a gel image of cointegrant, T-RL simple, and T-LR simple insert products.

[0013] FIG.7A-7B: (A) Shows λ-Red VchCAST editing efficiency in Pseudomonas putida (P. putida) compared to VchCAST and non-targeting λ-Red VchCAST-NT controls. (B) Shows λ-RedAtty. Dkt: BERK-543WO VchCAST editing efficiency in Klebsiella michiganensis (K. michiganensis) compared to VchCAST and non-targeting λ-Red VchCAST-NT controls.

[0014] FIG.8A-8E (A) Shows survival of K. michiganensis in the presence of different antibiotic dosages. (B) Shows survival of P. putida in the presence of different antibiotic dosages. (C) Shows the frequency survival of P. putida in the presence of 50 µg / ml kanamycin following VchCAST editing with two guides targeting different safe sites. (D) Shows λ-Red VchCAST editing efficiency in P. putida with varying levels of induction by Crystal Violet (CV). (E) Shows λ-Red VchCAST editing efficiency in K. michiganensis with varying levels of induction by Crystal Violet (CV).

[0015] FIG.9A-9D provide amino acid sequences of Scytonema hofmanni CAST polypeptides. (SEQ ID NOs: 377-380, respectively)

[0016] FIG.10A-10G provide amino acid sequences of Vibrio cholerae CAST polypeptides. (SEQ ID NOs: 381-387, respectively)

[0017] FIG.11A-11R provide amino acid sequences of CAST polypeptides suitable for use in an ShCAST-type complex. (SEQ ID NOs: 388-405, respectively)

[0018] FIG.12A-12U provide amino acid sequences of CAST polypeptides suitable for use in a VcCAST-type complex. (SEQ ID NOs: 406-426, respectively)

[0019] FIG.13 shows the phylogenetic distribution of 11 genes, including the genes identified in the genome-wide screen that activate or inhibit CAST editing, across 92 bacterial phyla.

[0020] FIG.14A-14B. Genome-wide screen identifies putative inhibitors and activators of VchCAST integration. (A) Schematic of the RB-TnSeq screen to identify E. coli genes affecting VchCAST integration efficiency. (B) Fitness scores of E. coli genes from two independent RB-TnSeq screens with fitness scores >1 or <-1 in both screens. Library screens 1 & 2 refer to repeated screens with gentamicin and chloramphenicol selection cargo, respectively. Asterisks denote candidate mutants that were experimentally validated. FIG.15A-15B. Validation of putative VchCAST activators and inhibitors in single deletion mutants. (A) Relative VchCAST editing efficiency in E. coli Keio collection mutants of candidate regulators identified from the RB-TnSeq screen. Editing efficiencies were normalized to the ΔyicI neutral fitness control strain. Mutants of screen-identified activators with significantly reduced editing efficiency relative to ΔyicI are shown in blue and inhibitors are shown in red (with ihfB previously described (10)); mutants of factors with non-significant (ns) difference in editing efficiency compared to ΔyicI are denoted by black. (B) Relative VchCAST editing efficiency in E. coli Keio collection mutants of RecBCD complex components and downstream homologous recombination effector RecA normalized to the ΔyicI negative control strain. Data values at zero represent samples with no viable colonies above the detection limit. Asterisks denote degree of significance (One sample t-test) for treatments compared to the ∆yicI control (* P ≤ 0.05, ** P ≤ 0.006, *** P ≤ 0.0008, **** P ≤ 0.0001). FIG.16A-16D. λ-Red recombineering system improves VchCAST editing efficiency in E. coli. (A) Schematic of the λ-Red VchCAST vector design. (B) Relative editing efficiency ofAtty. Dkt: BERK-543WO VchCAST vectors testing the different λ-Red genes (beta, exo, gam). The editing efficiency of λ- Red VchCAST as well as the non-targeting biological replicates were normalized to the paired VchCAST biological replicates. Asterisks denote statistical significance determined by one sample t-test compared to hypothetical mean (1) of VchCAST grown on crystal violet (one-tailed p values: * P ≤ 0.05, ** P ≤ 0.01) (C) On-target insertion frequency of E. coli transconjugants via whole genome sequencing. (D) WGS read distribution (%) of insertions downstream (bp) of PAM target site. FIG.17A-17F. λ-Red improves VchCAST editing efficiency in Pseudomonas putida and Klebsiella michiganensis. (A) Normalized editing efficiency of VchCAST compared to λ-Red VchCAST in P. putida. CV induction was performed at 1µM. Asterisks denote the degree of significance determined by one-sample t-test (one-tailed p-value: * P ≤ 0.05) compared to the normalized VchCAST control (1). (B) On-target insertion frequency across the P. putida genome via whole genome sequencing. (C) Distribution of cargo insertion loci downstream (bp) of the PAM in P. putida. (D) Normalized editing efficiency of VchCAST compared to λ-Red VchCAST in K. michiganensis with CV inductions performed at 0.5 µM. (E) On-target insertion frequency in K. michiganensis (F) Distribution of cargo insertion loci in K. michiganensis. FIG.18. Phylogenetic distribution of VchCAST activator and inhibitor genes. The phylogenetic distribution of 11 E. coli regulatory genes was mapped across 80,789 representative bacterial genomes from 92 phyla in Genome Taxonomy Database, release 214.0 (GTDB v214.0; (26, 27) with at least 10 members. Homologs were identified using AnnoTree (28) and confirmed with the eggNOG database (29).

[0021] FIG.19A-19B. VchCAST construct and conjugative delivery-insertion efficiency assay. (A) Plasmid schematic for the standard type I-F VchCAST editing vector. (B) Conceptual schematic for conjugation-based VchCAST insertion efficiency assay.

[0022] FIG.20. Comparison of relative fitness scores for single-gene insertion mutants generated using mariner and VchCAST systems. Scatterplots display the log₂ fold change in fitness from two independent RB-TnSeq library screens: screen 1 (top) and screen 2 (bottom). Each point represents the mean fitness score of technical replicates for a given gene. Points colored red indicate an absolute difference in fitness scores greater than 1 between the two systems (|mariner – VchCAST| > 1), while blue points fall below this threshold. The dashed diagonal line represents the line of identity (y = x). Genes of interest are labeled.

[0023] FIG.21A-21C. Comparison of fitness scores for mutants generated by VchCAST and mariner transposon systems. Side-by-side box plots display the log₂ fold change in fitness for selected genes across both VchCAST and mariner-edited mutant libraries, aggregated from two independent RB-TnSeq screens. Each panel represents an individual gene, with points indicating replicate measurements. Genes are categorized into (A) Activators, (B) Inhibitors, and (C) Neutral / Control, based on their putative roles in the phenotype of interest.Atty. Dkt: BERK-543WO

[0024] FIG.22. Ciprofloxacin (CIP) exposure increases VchCAST editing efficiency in E. coli. VchCAST’s editing efficiency in E. coli BW25113 incubated with full (160 ng / mL) and one-half (80 ng / mL) MIC CIP is normalized to that of the no CIP (0 ng / mL) treatment (* P ≤ 0.05).

[0025] FIG.23. CV induction optimization for λ-Red expression in E. coli. Percentage of insertion efficiency is reported in Log10. A Kruskal-Wallis test with multiple comparisons was used to determine if any significant differences were observed across the CV concentrations. Ns denotes not significant. Pink is used to denote treatments with crystal violet induction. Triangles represent the λ-Red VchCAST treatments, while circles represent the VchCAST treatment.

[0026] FIG.24A-24E. Comparison of insertion products for VchCAST and λ-Red VchCAST in BW25113 E. coli. (A) Schematic of primer binding and amplification for insertion product verification and orientation analysis by colony PCR (cPCR). (B) Representative gel image of cPCR analysis of VchCAST insertion products in E. coli BW25113. Oligo pairs used for amplification are detailed in part A of this figure. CO is cointegrate, SI is simple insert, and the orientations are denoted as RL (right-left) and LR (left-right). (C) Side-by-side comparison of product type and orientation (%) of VchCAST and λ-Red VchCAST in E. coli BW25113 as characterized by cPCR. (D) Percentage of cointegrate reads determined by whole genome sequencing across three sequencing runs. A welch’s T-test determined no significant difference (ns). (E) Side-by-side comparison of product type and orientation (%) of in E. coli BW25113 as characterized by whole genome sequencing.

[0027] FIG.25A-25F. P. Putida and K. michiganensis VchCAST editing. (A) Antibiotic susceptibility testing for P. putida. The dotted line depicts the limit of detection 2.56E-8 (B) Antibiotic susceptibility testing for K. michiganensis. The dotted line depicts the limit of detection 1.33E-8 (C) Guide design for P. putida Safe site 6 (SS6). (D) Safe site guide insertion efficiency (%) in P. putida. (E) CV induction optimization for λ-Red expression in P. putida. (F) CV induction optimization for λ-Red expression in K. michiganensis.

[0028] FIG.26. Percent distribution of 11 VchCAST activator and inhibitor genes across type I-F CAST-containing genomes (n = 1069) compiled from Rybarski et al. (2021), Peters et al. (2017), and Klompe et al. (2023). DEFINITIONS

[0029] The terms “polynucleotide” and “nucleic acid,” used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, this term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.

[0030] By "hybridizable" or “complementary” or “substantially complementary" it is meant that a nucleic acid (e.g. RNA, DNA) comprises a sequence of nucleotides that enables it to non- covalently bind, i.e. form Watson-Crick base pairs and / or G / U base pairs, “anneal”, or “hybridize,” to another nucleic acid in a sequence-specific, antiparallel, manner (i.e., a nucleic acid specificallyAtty. Dkt: BERK-543WO binds to a complementary nucleic acid) under the appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. Standard Watson-Crick base-pairing includes: adenine (A) pairing with thymidine (T), adenine (A) pairing with uracil (U), and guanine (G) pairing with cytosine (C) [DNA, RNA]. In addition, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization of a DNA molecule with an RNA molecule (e.g., when a DNA target nucleic acid base pairs with a guide RNA, etc.): guanine (G) can also base pair with uracil (U). For example, G / U base-pairing is at least partially responsible for the degeneracy (i.e., redundancy) of the genetic code in the context of tRNA anti-codon base-pairing with codons in mRNA. Thus, in the context of this disclosure, a guanine (G) (e.g., of dsRNA duplex of a guide RNA molecule; of a guide RNA base pairing with a target nucleic acid, etc.) is considered complementary to both a uracil (U) and to an adenine (A). For example, when a G / U base-pair can be made at a given nucleotide position of a dsRNA duplex of a guide RNA molecule, the position is not considered to be non-complementary, but is instead considered to be complementary.

[0031] Hybridization and washing conditions are well known and exemplified in Sambrook, J., Fritsch, E. F. and Maniatis, T. Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 therein; and Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001). The conditions of temperature and ionic strength determine the "stringency" of the hybridization.

[0032] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases are possible. The conditions appropriate for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, variables well known in the art. The greater the degree of complementarity between two nucleotide sequences, the greater the value of the melting temperature (Tm) for hybrids of nucleic acids having those sequences. For hybridizations between nucleic acids with short stretches of complementarity (e.g. complementarity over 35 or less, 30 or less, 25 or less, 22 or less, 20 or less, or 18 or less nucleotides) the position of mismatches can become important (see Sambrook et al., supra, 11.7-11.8). Typically, the length for a hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more). Temperature, wash solution salt concentration, and other conditions may be adjusted as necessary according to factors such as length of the region of complementation and the degree of complementation.

[0033] It is understood that the sequence of a polynucleotide need not be 100% complementary to that of its target nucleic acid to be specifically hybridizable or hybridizable. Moreover, a polynucleotide may hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a bulge, a loop structure or hairpin structure, etc.). A polynucleotide can comprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity to a target region within the target nucleic acid sequence to whichAtty. Dkt: BERK-543WO it will hybridize. For example, an antisense nucleic acid in which 18 of 20 nucleotides of the antisense compound are complementary to a target region, and would therefore specifically hybridize, would represent 90 percent complementarity. In this example, the remaining noncomplementary nucleotides may be clustered or interspersed with complementary nucleotides and need not be contiguous to each other or to complementary nucleotides. Percent complementarity between particular stretches of nucleic acid sequences within nucleic acids can be determined using any convenient method. Example methods include BLAST programs (basic local alignment search tools) and PowerBLAST programs (Altschul et al., J. Mol. Biol., 1990, 215, 403- 410; Zhang and Madden, Genome Res., 1997, 7, 649-656), the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), e.g., using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489), and the like.

[0034] The terms ''peptide," ''polypeptide," and "protein" are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.

[0035] "Binding" as used herein (e.g. with reference to an RNA-binding domain of a polypeptide, binding to a target nucleic acid, and the like) refers to a non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid; between a CAST polypeptide / guide RNA complex and a target nucleic acid; and the like). While in a state of non-covalent interaction, the macromolecules are said to be “associated” or “interacting” or “binding” (e.g., when a molecule X is said to interact with a molecule Y, it is meant the molecule X binds to molecule Y in a non-covalent manner). Not all components of a binding interaction need be sequence-specific (e.g., contacts with phosphate residues in a DNA backbone), but some portions of a binding interaction may be sequence-specific. Binding interactions are generally characterized by a dissociation constant (KD) of less than 10-6M, less than 10-7M, less than 10-8M, less than 10-9M, less than 10-10M, less than 10-11M, less than 10-12M, less than 10-13M, less than 10-14M, or less than 10-15M. "Affinity" refers to the strength of binding, increased binding affinity being correlated with a lower KD.

[0036] As used herein, a “promoter” or a "promoter sequence" is a DNA regulatory region capable of binding RNA polymerase and initiating transcription of a downstream (3' direction) coding or non- coding sequence. For purposes of the present disclosure, the promoter sequence is bounded at its 3' terminus by the transcription initiation site and extends upstream (5' direction) to include the minimum number of bases or elements necessary to initiate transcription at levels detectable above background. Within the promoter sequence will be found a transcription initiation site, as well as protein binding domains responsible for the binding of RNA polymerase. Eukaryotic promoters will often, but not always, contain "TATA" boxes and "CAT" boxes. Various promoters, including inducible promoters, may be used to drive expression by the various vectors of the present disclosure.Atty. Dkt: BERK-543WO

[0037] "Operably linked" refers to a juxtaposition wherein the components so described are in a relationship permitting them to function in their intended manner. For instance, a promoter is operably linked to a coding sequence (or the coding sequence can also be said to be operably linked to the promoter) if the promoter affects its transcription or expression.

[0038] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0039] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0040] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, the preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.

[0041] It must be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a conjugative nucleic acid construct” includes a plurality of such constructs and reference to “the CAST complex” or “the CAST modulator” includes reference to one or more CAST complexes or CAST modulators and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.

[0042] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various embodiments andAtty. Dkt: BERK-543WO elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.

[0043] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed. DETAILED DESCRIPTION

[0044] The present disclosure provides a transposon system comprising i) a nucleotide sequence encoding polypeptides that form a CRISPR-associated transposase (CAST) complex; ii) a nucleotide sequence encoding a guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome; iii) a transposon, or an insertion site for a transposon, wherein the transposon or the transposon insertion site is flanked by recognition sites that are recognized by the CAST complex; and iv) a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides. In some cases, the CAST modulator comprises a small molecule that promotes homologous recombination. The present disclosure provides a prokaryotic cell comprising a subject transposon system. The present disclosure additionally provides methods for editing the genome of a target prokaryotic cell(s) utilizing a subject transposon system. TRANSPOSONSYSTEM

[0045] The present disclosure provides a transposon system comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence(s) encoding one or more guide RNAs; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites, and iv) a Cast modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some embodiments, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides. In some embodiments, comprises a small molecule that promotes homologous recombination.

[0046] In some cases, i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence(s) encoding one or more guide RNAs; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and iv) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides are all present on the same nucleic acid construct; i.e., are present on a single nucleic acid construct. In some cases, the nucleic acid construct is a conjugative construct. A conjugative construct comprises an origin of transfer, e.g., a nucleotide sequence that provides for transfer of the construct from a first prokaryotic cell to a second prokaryotic cell. In some cases, a conjugative construct of the present disclosure is a non-Atty. Dkt: BERK-543WO replicative construct. Thus, the present disclosure provides a single conjugative construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence(s) encoding one or more guide RNAs; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and iv) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides. In some cases, a conjugative construct of the present disclosure is a replicative construct. In some cases, a conjugative construct of the present disclosure is replicative, but is lost from a host cell comprising the conjugative construct when the host cell is cultured at 37°C or at a temperature that is higher than 37°C.

[0047] In some cases, i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence(s) encoding one or more guide RNAs; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and iv) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides are all present on the same nucleic acid construct; i.e., are present on a single nucleic acid construct. In some cases, the nucleic acid construct is a conjugative construct. A conjugative construct comprises an origin of transfer, e.g., a nucleotide sequence that provides for transfer of the construct from a first bacterium to a second bacterium. In some cases, the conjugative construct is a non-replicative construct. Thus, the present disclosure provides a single conjugative construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and iv) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides.

[0048] In some cases, i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence(s) encoding one or more guide RNAs; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites are all present on the same nucleic acid construct; i.e., are present on a single nucleic acid construct. In some cases, the nucleic acid construct is a conjugative construct. In some cases, the conjugative construct is a non- replicative construct. Thus, the present disclosure provides a single conjugative construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites.

[0049] In some cases, i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence(s) encoding one or more guide RNAs; and iii) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides are present on a first nucleic acid construct; and iv) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites is present on a second nucleic acid construct. Thus, in some cases, a system of the present disclosure comprises: a) a first nucleic acid comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence(s) encoding one or more guide RNAs; and iii) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides; and b) a second nucleic acid comprising a transposon, or an insertion site for a transposon, flanked byAtty. Dkt: BERK-543WO CAST complex recognition sites. In some cases, the nucleic acid constructs are both conjugative constructs.

[0050] In some cases, i) a nucleotide sequence encoding polypeptides that form a CAST complex; and ii) a nucleotide sequence(s) encoding one or more guide RNAs are present on a first nucleic acid construct; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and iv) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides is present on a second nucleic acid construct. Thus, in some cases, a system of the present disclosure comprises: a) a first nucleic acid comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; and ii) a nucleotide sequence(s) encoding one or more guide RNAs; and b) a second nucleic acid comprising: i) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and ii) a nucleotide sequence(s) encoding one or more CAST modulator polypeptides. In some cases, the nucleic acid constructs are both conjugative constructs.

[0051] In some cases, i) a nucleotide sequence encoding polypeptides that form a CAST complex; and ii) a nucleotide sequence(s) encoding one or more guide RNAs are present on a first nucleic acid construct; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites is present on a second nucleic acid construct. Thus, in some cases, a system of the present disclosure comprises: a) a first nucleic acid comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; and ii) a nucleotide sequence(s) encoding one or more guide RNAs; and b) a second nucleic acid comprising a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites. In some cases, the nucleic acid constructs are both conjugative constructs.

[0052] In some cases, a nucleic acid construct of a transposon system of the present disclosure comprises a selectable marker. In some cases, a nucleic acid construct of a transposon system of the present disclosure does not comprise a selectable marker. Selectable markers include polypeptides that provide for antibiotic resistance. Antibiotic resistance includes, e.g., ampicillin resistance, kanamycin resistance, chloramphenicol resistance, streptomycin resistance, spectinomycin resistance, tetracycline resistance, erythromycin resistance, neomycin resistance, gentamycin resistance and the like. Polypeptides that provide for antibiotic resistance are known in the art and include, e.g., gentamycin acetyltransferase, beta-lactamase, neomycin phosphotransferase, and the like. Thus, a transposon system of the present disclosure can be used for negative selection (e.g., antimicrobial resistance).

[0053] In some cases, a nucleic acid construct of a transposon system of the present disclosure comprises a screenable marker (e.g., for positive selection), such as a fluorescent polypeptide. Suitable fluorescent proteins include, but are not limited to, green fluorescent protein (GFP) or variants thereof, blue fluorescent variant of GFP (BFP), cyan fluorescent variant of GFP (CFP), yellow fluorescent variant of GFP (YFP), enhanced GFP (EGFP), enhanced CFP (ECFP), enhanced YFP (EYFP), GFPS65T, Emerald, Topaz (TYFP), Venus, Citrine, mCitrine, GFPuv, destabilised EGFP (dEGFP), destabilised ECFP (dECFP), destabilised EYFP (dEYFP), mCFPm,Atty. Dkt: BERK-543WO Cerulean, T-Sapphire, CyPet, YPet, mKO, HcRed, t-HcRed, DsRed, DsRed2, DsRed-monomer, J- Red, dimer2, t-dimer2(12), mRFP1, pocilloporin, Renilla GFP, Monster GFP, paGFP, Kaede protein and kindling protein, Phycobiliproteins and Phycobiliprotein conjugates including B- Phycoerythrin, R-Phycoerythrin and Allophycocyanin. Other examples of fluorescent proteins include mHoneydew, mBanana, mOrange, dTomato, tdTomato, mTangerine, mStrawberry, mCherry, mGrape1, mRaspberry, mGrape2, mPlum (Shaner et al. (2005) Nat. Methods 2:905-909), and the like. Any of a variety of fluorescent and colored proteins, e.g., those from Anthozoan species or modified versions thereof, as described in, e.g., Matz et al. (1999) Nature Biotechnol. 17:969-973, is suitable for use.

[0054] As another example, in some cases, a nucleic acid construct of a transposon system of the present disclosure comprises a nucleotide sequence encoding a polypeptide that, when exhibited on the surface of a cell, can be targeted by an antibody specific for the polypeptide. Such polypeptides include, e.g., epitope tags.

[0055] In some cases, a nucleic acid construct of a transposon system of the present disclosure comprises a nucleic acid comprising nucleotide sequences encoding one or more polypeptides that can provide for metabolic selection (positive selection). For example, the ability to utilize a particular carbon source that is not normally a carbon source utilized by a particular bacterium can be selected. Such carbon sources include, e.g., lactose. CAST

[0056] CRISPR-associated transposases (CASTs) include a CRISPR-associated polypeptide and one or more additional polypeptides that, in complex with one another, mediate transposition of a target transposon. ShCAST

[0057] In some cases, a CAST comprises: i) a Cas 12k polypeptide; ii) a TnsC polypeptide; iii) a TnsB polypeptide; and iv) a TniQ polypeptide. An example of such a CAST is a Scytonema hofmanni CAST (ShCAST).

[0058] A Cas12k polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni Cas12k amino acid sequence depicted in FIG.9A. A Cas12k polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from 500 amino acids to 639 amino acids (e.g., from 500 amino acids (aa) to 550 aa, from 550 aa to 575 aa, from 575 aa to 600 aa, from 600 aa to 625 aa, or from 625 aa to 639 aa) of the S. hofmanni Cas12k amino acid sequence depicted in FIG.9A. In some cases, the Cas12k polypeptide has a length of from about 600 amino acids to 650 amino acids (e.g., from 600 amino acids (aa) to 625 aa, or from 625 aa to 650 aa). In some cases, the Cas12k polypeptide has a length of 639 aa.

[0059] Non-limiting examples of other suitable Cas12k polypeptides are provided in FIG.11F-11J. For example, a Cas12k polypeptide can comprise an amino acid sequence having at least 50%, at leastAtty. Dkt: BERK-543WO 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas12k polypeptide amino acid sequences depicted in FIG.11F-11J.

[0060] A TnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni TnsB amino acid sequence depicted in FIG.9B. A TnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 500 amino acids to 584 amino acids (e.g., from about 500 amino acids (aa) to 525 aa, from 525 aa to 550 aa, from 550 aa to 575 aa, or from 575 aa to 584 aa) of the S. hofmanni TnsB amino acid sequence depicted in FIG.9B. In some cases, the TnsB polypeptide has a length of from about 500 amino acids to about 600 amino acids (e.g., from about 500 amino acids (aa) to 525 aa, from 525 aa to 550 aa, from 550 aa to 575 aa, or from 575 aa to 600 aa). In some cases, the TnsB polypeptide has a length of 584 aa.

[0061] Non-limiting examples of other suitable TnsB polypeptides are provided in FIG.11A-11E. For example, a TnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the TnsB polypeptide amino acid sequences depicted in FIG.11A-11E.

[0062] A TnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni TnsC amino acid sequence depicted in FIG.9C. A TnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 200 amino acids to 276 amino acids (e.g., from about 200 amino acids (aa) to 225 aa, from 225 aa to 250 aa or from 250 aa to 276 aa) of the S. hofmanni TnsC amino acid sequence depicted in FIG.9C. In some cases, the TnsC polypeptide has a length of from about 200 amino acids to 276 amino acids (e.g., from about 200 amino acids (aa) to 225 aa, from 225 aa to 250 aa or from 250 aa to 276 aa). In some cases, the TnsC polypeptide has a length of 276 aa.

[0063] Non-limiting examples of other suitable TnsC polypeptides are provided in FIG.11K-11N. For example, a TnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the TnsC polypeptide amino acid sequences depicted in FIG.11K-11N.

[0064] A TniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni TniQ amino acid sequence depicted in FIG.9D. A TniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%,Atty. Dkt: BERK-543WO at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 100 amino acids to 167 amino acids (e.g., from 100 amino acids (aa) to 125 aa, from 125 aa to 150 aa, or from 150 aa to 167 aa) of the S. hofmanni TniQ amino acid sequence depicted in FIG.9D. In some cases, the TniQ polypeptide has a length of from about 100 amino acids to 167 amino acids (e.g., from 100 amino acids (aa) to 125 aa, from 125 aa to 150 aa, or from 150 aa to 167 aa). In some cases, the TniQ polypeptide has a length of 167 amino acids.

[0065] Non-limiting examples of other suitable TniQ polypeptides are provided in FIG.11O-11R. For example, a TniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the TniQ polypeptide amino acid sequences depicted in FIG.11O-11R. VcCAST

[0066] In some cases, a CAST comprises: i) a Cas6 polypeptide; ii) a Cas7 polypeptide; iii) a Cas8 polypeptide; iv) a TnsA polypeptide; v) a TnsB polypeptide; vi) a TnsC polypeptide; and vii) a TniQ polypeptide. An example of such a CAST is a Vibrio cholerae CAST (VcCAST or, interchangeably, VchCAST).

[0067] A Cas6 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera Cas6 polypeptide amino acid sequence depicted in FIG.10G. A Cas6 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 125 amino acids to 199 amino acids (e.g., from about 125 amino acids (aa) to 150 aa, from 150 aa to 175 aa, or from 175 aa to 199 aa) of the V. cholera Cas6 polypeptide amino acid sequence depicted in FIG.10G. A Cas6 polypeptide can have a length of from about 125 amino acids to 199 amino acids (e.g., from about 125 amino acids (aa) to 150 aa, from 150 aa to 175 aa, or from 175 aa to 199 aa). A Cas6 polypeptide can have a length of 199 aa.

[0068] Non-limiting examples of other suitable Cas6 polypeptides are provided in FIG.12M-12O. For example, a Cas6 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas6 polypeptide amino acid sequences depicted in FIG.12M-12O.

[0069] A Cas7 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera Cas7 polypeptide amino acid sequence depicted in FIG.10F. A Cas7 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 275 amino acids to 352 amino acids (e.g.,Atty. Dkt: BERK-543WO from about 275 amino acids (aa) to 300 aa, from 300 aa to 325 aa, or from 325 aa to 352 aa) of the V. cholerae Cas7 polypeptide amino acid sequence depicted in FIG.10F. A Cas7 polypeptide can have a length of from about 275 amino acids to 352 amino acids (e.g., from about 275 amino acids (aa) to 300 aa, from 300 aa to 325 aa, or from 325 aa to 352 aa). A Cas7 polypeptide can have a length of 352 aa.

[0070] Non-limiting examples of other suitable Cas7 polypeptides are provided in FIG.12P-12R. For example, a Cas7 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas7 polypeptide amino acid sequences depicted in FIG.12P-12R.

[0071] A Cas8 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera Cas8 polypeptide amino acid sequence depicted in FIG.10E. A Cas8 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 575 amino acids to 640 amino acids (e.g., from about 575 amino acids (aa) to 600 aa, from 600 aa to 625 aa, or from 625 aa to 640 aa) of the V. cholerae Cas8 polypeptide amino acid sequence depicted in FIG.10E. A Cas8 polypeptide can have a length of from about 575 amino acids to 640 amino acids (e.g., from about 575 amino acids (aa) to 600 aa, from 600 aa to 625 aa, or from 625 aa to 640 aa). A Cas8 polypeptide can have a length of 640 aa.

[0072] Non-limiting examples of other suitable Cas8 polypeptides are provided in FIG.12S-12U. For example, a Cas8 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas8 polypeptide amino acid sequences depicted in FIG.12S-12U.

[0073] A tnsA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tnsA polypeptide amino acid sequence depicted in FIG.10A. A tnsA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 150 amino acids to 222 amino acids (e.g., from about 150 amino acids (aa) to 175 aa, from 175 aa to 200 aa, or from 200 aa to 222 aa) of the tnsA amino acid sequence depicted in FIG.10A. A tnsA polypeptide can have a length of from about 150 amino acids to 222 amino acids (e.g., from about 150 amino acids (aa) to 175 aa, from 175 aa to 200 aa, or from 200 aa to 222 aa). A tnsA polypeptide can have a length of 222 amino acids.

[0074] Non-limiting examples of other suitable tnsA polypeptides are provided in FIG.12A-12C. For example, a tnsA polypeptide can comprise an amino acid sequence having at least 50%, at leastAtty. Dkt: BERK-543WO 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tnsA polypeptide amino acid sequences depicted in FIG.12A-12C.

[0075] A tnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tnsB polypeptide amino acid sequence depicted in FIG.10B. A tnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 525 amino acids to 603 amino acids (e.g., from about 525 amino acids (aa) to 550 aa, from 550 aa to 575 aa, or from 575 aa to 603 aa) of the tnsB amino acid sequence depicted in FIG.10B. A tnsB polypeptide can have a length of from about from about 525 amino acids to 603 amino acids (e.g., from about 525 amino acids (aa) to 550 aa, from 550 aa to 575 aa, or from 575 aa to 603 aa). A tnsB polypeptide can have a length of 603 amino acids.

[0076] Non-limiting examples of other suitable tnsB polypeptides are provided in FIG.12D-12F. For example, a tnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tnsB polypeptide amino acid sequences depicted in FIG.12D-12F.

[0077] A tnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tnsC polypeptide amino acid sequence depicted in FIG.10C. A tnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 225 amino acids to 330 amino acids (e.g., from about 225 amino acids (aa) to 250 aa, from 250 aa to 300 aa, or from 300 aa to 330 aa) of the tnsC amino acid sequence depicted in FIG.10C. A tnsC polypeptide can have a length of from about 225 amino acids to 330 amino acids (e.g., from about 225 amino acids (aa) to 250 aa, from 250 aa to 300 aa, or from 300 aa to 330 aa). A tnsC polypeptide can have a length of 330 amino acids.

[0078] Non-limiting examples of other suitable tnsC polypeptides are provided in FIG.12G-12I. For example, a tnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tnsC polypeptide amino acid sequences depicted in FIG.12G-12I.

[0079] A tniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tniQ polypeptide amino acid sequence depicted in FIG.10D. A tniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at leastAtty. Dkt: BERK-543WO 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 300 amino acids to 394 amino acids (e.g., from about 300 amino acids (aa) to 325 aa, from 325 aa to 350 aa, from 350 aa to 375 aa, or from 375 aa to 394 aa) of the tniQ amino acid sequence depicted in FIG.10D. A tniQ polypeptide can have a length of from about 300 amino acids to 394 amino acids (e.g., from about 300 amino acids (aa) to 325 aa, from 325 aa to 350 aa, from 350 aa to 375 aa, or from 375 aa to 394 aa). A tniQ polypeptide can have a length of 394 amino acids.

[0080] Non-limiting examples of other suitable tniQ polypeptides are provided in FIG.12J-12L. For example, a tniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tniQ polypeptide amino acid sequences depicted in FIG.12J-12L. CAST Modulators

[0081] CAST modulators of the present disclosure can enhance the editing efficiency of a subject transposon system compared to the editing efficiency of the subject transposon system without the CAST modulator. The “editing efficiency” of a subject transposon can be considered the probability that a successful edit occurs in a cell that has been introduced to the subject transposon system. Alternatively, the “editing efficiency” of a subject transposon can be considered the frequency of cells, in a population of cells, that have been introduced to the subject transposon system, in which a successful edit occurs. By way of example, the “editing efficiency” of a subject transposon can be considered the frequency of cells, in a population of cells, that have been conjugated by a donor cell harboring the subject transposon and in which the transposon has been successfully integrated. In some embodiments, the CAST modulator is a nucleotide sequence(s) encoding one or more CAST modulator polypeptides. A CAST modulator polypeptide is a polypeptide that, either alone or in combination with one or more other CAST modulator polypeptides, enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator polypeptide. In some embodiments, the CAST modulator is a small molecule that enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator small molecule. CAST Modulator Polypeptides

[0082] In some embodiments, the CAST modulator comprises a nucleotide sequence(s) encoding one or more CAST modulator polypeptides. In some embodiments, the CAST modulator polypeptide is a polypeptide that promotes / enhances editing efficiency (i.e., the CAST modulator is an activator). In some embodiments, the CAST modulator polypeptide is a polypeptide that, alone or in combination with another CAST modulator polypeptide, promotes homologous recombination (e.g., promotes homologous recombination DNA repair pathways). In some cases, the one or more CAST modulator polypeptides may comprise polypeptides of a recombineering system. Recombineering system polypeptides are known in the art and described, for example, in Li,Atty. Dkt: BERK-543WO Ruijuan, et al. "The emerging role of recombineering in microbiology." Engineering Microbiology 3.3 (2023): 100097, the entirety of which is incorporated by reference herein.

[0083] Suitable CAST modulator polypeptides include polypeptides comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the amino acid sequences set forth in Table 1 (e.g., SEQ ID NOs: 1-20). In some embodiments, the CAST modulator comprises a nucleotide sequence(s) encoding a combination of CAST modulator polypeptides, each polypeptide comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the amino acid sequences set forth in Table 1 (e.g., SEQ ID NOs: 1-20).

[0084] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides, wherein the CAST modulator polypeptides comprise: i) an Exo polypeptide; ii) a Beta polypeptide; and iii) a Gam polypeptide.

[0085] A Exo polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the λ phage Exo amino acid sequence set forth in SEQ ID NO:1 (or a sequence of wild type homolog thereof from another bacteriophage species). In some cases, the Exo polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the λ phage Exo amino acid sequence set forth in SEQ ID NO:1 (or a sequence of wild type homolog thereof from another bacteriophage species).A Beta polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the λ phage Beta amino acid sequence set forth in SEQ ID NO:2 (or a sequence of wild type homolog thereof from another bacteriophage species). In some cases, the Beta polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the λ phage Beta amino acid sequence set forth in SEQ ID NO:2 (or a sequence of wild type homolog thereof from another bacteriophage species).

[0086] A Gam polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the λ phage Gam amino acid sequence set forth in SEQ ID NO:3 (or a sequence of wild type homolog thereof from another bacteriophage species). In some cases, the Gam polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the λ phage Gam amino acid sequence set forth in SEQ ID NO:3 (or a sequence of wild type homolog thereof from another bacteriophage species).

[0087] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides, wherein the CAST modulator polypeptides comprise: i) a RecE polypeptide; and ii) a RecT polypeptide.Atty. Dkt: BERK-543WO

[0088] A RecE polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli prophage RecE amino acid sequence set forth in SEQ ID NO:4 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the RecE polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli prophage RecE amino acid sequence set forth in SEQ ID NO:4 (or a sequence of wild type homolog thereof from another bacterial species).

[0089] A RecT polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli prophage RecT amino acid sequence set forth in SEQ ID NO:5 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the RecT polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli prophage RecT amino acid sequence set forth in SEQ ID NO:5 (or a sequence of wild type homolog thereof from another bacterial species).

[0090] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides, wherein the CAST modulator polypeptides comprise: i) a addA polypeptide; and ii) an addB polypeptide.

[0091] An addA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the Bacillus addA amino acid sequence set forth in SEQ ID NO:18 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the addA polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the Bacillus addA amino acid sequence set forth in SEQ ID NO:18 (or a sequence of wild type homolog thereof from another bacterial species).

[0092] An addB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the Bacillus addB amino acid sequence set forth in SEQ ID NO:19 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the addB polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the Bacillus addB amino acid sequence set forth in SEQ ID NO:19 (or a sequence of wild type homolog thereof from another bacterial species).

[0093] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides (activator polypeptides), wherein the CAST modulator polypeptides comprise: a YnfA polypeptide, a YneK polypeptide, a YnbE polypeptide, a FimF polypeptide, a UmuD polypeptide, a NudG polypeptide, a Mdh polypeptide, a CspC polypeptide, a ChaB polypeptide, a Lpp polypeptide, a IhfA polypeptide, a IhfB polypeptide, a RecA polypeptide, or any combination thereof. See below for details regarding each of the listed proteins and their sequences.Atty. Dkt: BERK-543WO

[0094] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides (activator polypeptides), wherein the CAST modulator polypeptides comprise: a IhfB polypeptide, a Lpp polypeptide, a ChaB polypeptide, a CspC polypeptide, a Mdh polypeptide, a NudG polypeptide, a UmuD polypeptide, a FimF polypeptide, a YnbE polypeptide, a YneK polypeptide, a YnfA polypeptide, or any combination thereof. See below for details regarding each of the listed proteins and their sequences.

[0095] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides (activator polypeptides), wherein the CAST modulator polypeptides comprise: a IhfB polypeptide, a CspC polypeptide, a YnbE polypeptide, a UmuD polypeptide, a ChaB polypeptide, a NudG polypeptide, a YneK polypeptide, or any combination thereof. See below for details regarding each of the listed proteins and their sequences.

[0096] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides (activator polypeptides), wherein the CAST modulator polypeptides comprise: a IhfB polypeptide, a CspC polypeptide, a YnbE polypeptide, a UmuD polypeptide, a ChaB polypeptide, or any combination thereof. See below for details regarding each of the listed proteins and their sequences.

[0097] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides (activator polypeptides), wherein the CAST modulator polypeptides comprise: a IhfB polypeptide, a CspC polypeptide, a YnbE polypeptide, a UmuD polypeptide, a ChaB polypeptide, a YnfA polypeptide, a FimF polypeptide, a Mdh polypeptide, a Lpp polypeptide, or any combination thereof. See below for details regarding each of the listed proteins and their sequences.

[0098] In some cases, the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides, wherein the CAST modulator polypeptides comprise: i) a IhfA polypeptide; and ii) an IhfB polypeptide. See below for details regarding each of the listed proteins and their sequences.

[0099] A YnfA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli YnfA amino acid sequence set forth in SEQ ID NO:6 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the YnfA polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli YnfA amino acid sequence set forth in SEQ ID NO:6 (or a sequence of wild type homolog thereof from another bacterial species).A YneK polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli YneK amino acid sequence set forth in SEQ ID NO:7 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the YneK polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) aminoAtty. Dkt: BERK-543WO acid sequence identity to the E. Coli YneK amino acid sequence set forth in SEQ ID NO:7 (or a sequence of wild type homolog thereof from another bacterial species).

[0100] In some embodiments, a YneK polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the YneK amino acid sequences set forth in SEQ ID NOs: 7 and 227-230. In some cases, the YneK polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the YneK amino acid sequences set forth in SEQ ID NOs: 7 and 227-230.A YnbE polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli YnbE amino acid sequence set forth in SEQ ID NO:8 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the YnbE polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli YnbE amino acid sequence set forth in SEQ ID NO:8 (or a sequence of wild type homolog thereof from another bacterial species).

[0101] In some embodiments, a YnbE polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the YnbE amino acid sequences set forth in SEQ ID NOs: 8 and 367-376. In some cases, the YnbE polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the YnbE amino acid sequences set forth in SEQ ID NOs: 8 and 367-376.

[0102] A FimF polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli FimF amino acid sequence set forth in SEQ ID NO:9 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the FimF polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli FimF amino acid sequence set forth in SEQ ID NO:9 (or a sequence of wild type homolog thereof from another bacterial species).

[0103] A UmuD polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli UmuD amino acid sequence set forth in SEQ ID NO:10 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the UmuD polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli UmuD amino acid sequence set forth in SEQ ID NO:10 (or a sequence of wild type homolog thereof from another bacterial species).

[0104] In some embodiments, a UmuD polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the UmuD amino acid sequences set forth in SEQ ID NOs: 10 and 181-226. In some cases, the UmuD polypeptideAtty. Dkt: BERK-543WO comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the UmuD amino acid sequences set forth in SEQ ID NOs: 10 and 181-226.

[0105] A NudG polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli NudG amino acid sequence set forth in SEQ ID NO:11 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the NudG polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli NudG amino acid sequence set forth in SEQ ID NO:11 (or a sequence of wild type homolog thereof from another bacterial species).

[0106] In some embodiments, a NudG polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the NudG amino acid sequences set forth in SEQ ID NOs: 11 and 171-180. In some cases, the NudG polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the NudG amino acid sequences set forth in SEQ ID NOs: 11 and 171-180.

[0107] A Mdh polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli Mdh amino acid sequence set forth in SEQ ID NO:12 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the Mdh polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli Mdh amino acid sequence set forth in SEQ ID NO:12 (or a sequence of wild type homolog thereof from another bacterial species).

[0108] A CspC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli CspC amino acid sequence set forth in SEQ ID NO:13 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the CspC polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli CspC amino acid sequence set forth in SEQ ID NO:13 (or a sequence of wild type homolog thereof from another bacterial species).

[0109] In some embodiments, a CspC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the CspC amino acid sequences set forth in SEQ ID NOs: 13 and 58-142. In some cases, the CspC polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the CspC amino acid sequences set forth in SEQ ID NOs: 13 and 58-142.

[0110] A ChaB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli ChaB amino acid sequence set forth in SEQ ID NO:14Atty. Dkt: BERK-543WO (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the ChaB polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli ChaB amino acid sequence set forth in SEQ ID NO:14 (or a sequence of wild type homolog thereof from another bacterial species).

[0111] In some embodiments, a ChaB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the ChaB amino acid sequences set forth in SEQ ID NOs: 14 and 21-57. In some cases, the ChaB polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the ChaB amino acid sequences set forth in SEQ ID NOs: 14 and 21-57.

[0112] A Lpp polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli Lpp amino acid sequence set forth in SEQ ID NO:15 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the Lpp polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli Lpp amino acid sequence set forth in SEQ ID NO:15 (or a sequence of wild type homolog thereof from another bacterial species).

[0113] A IhfB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli IhfB amino acid sequence set forth in SEQ ID NO:16 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the IhfB polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli IhfB amino acid sequence set forth in SEQ ID NO:16 (or a sequence of wild type homolog thereof from another bacterial species).

[0114] In some embodiments, a IhfB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the IhfB amino acid sequences set forth in SEQ ID NOs: 16 and 157-170. In some cases, the IhfB polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the IhfB amino acid sequences set forth in SEQ ID NOs: 16 and 157-170.

[0115] A IhfA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli IhfA amino acid sequence set forth in SEQ ID NO:17 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the IhfA polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli IhfA amino acid sequence set forth in SEQ ID NO:17 (or a sequence of wild type homolog thereof from another bacterial species).

[0116] In some embodiments, a IhfA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, atAtty. Dkt: BERK-543WO least 99%, or 100% amino acid sequence identity to any one of the IhfA amino acid sequences set forth in SEQ ID NOs: 17 and 143-156. In some cases, the IhfA polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the IhfA amino acid sequences set forth in SEQ ID NOs: 17 and 143-156.

[0117] A RecA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the E. Coli RecA amino acid sequence set forth in SEQ ID NO:235 (or a sequence of wild type homolog thereof from another bacterial species). In some cases, the RecA polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the E. Coli RecA amino acid sequence set forth in SEQ ID NO: 235 (or a sequence of wild type homolog thereof from another bacterial species).

[0118] In some embodiments, a RecA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the RecA amino acid sequences set forth in SEQ ID NOs: 235-327. In some cases, the RecA polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the RecA amino acid sequences set forth in SEQ ID NOs: 235-327.

[0119] In some cases, the CAST modulator comprises a nucleotide sequence encoding a ICP8 polypeptide.

[0120] A ICP8 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the Herpes Simplex virus 1 ICP8 amino acid sequence set forth in SEQ ID NO:20 (or a sequence of wild type homolog thereof from another virus species). In some cases, the ICP8 polypeptide comprises at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to the Herpes Simplex virus 1 ICP8 amino acid sequence set forth in SEQ ID NO:20 (or a sequence of wild type homolog thereof from another virus species).

[0121] Table 1: Examples of CAST modulator polypeptides Name NCBI Amino Acid Sequence SEQ ID R f NO:Atty. Dkt: BERK-543WO KDEAERIVENTAYTAERQPERDITPVNDETMQEIN TLLIALDKTWDDDLLPLCSQIFRRDIRASSELTQAE AVKALGFLKQKAAEQKAAAAtty. Dkt: BERK-543WO FRGNDTLYNEKVQFKLTYPARNGRHKENIEFQVVI (also see NLSPIYLDNFRHDGEINIFCAPNPKPVTMGRVFQT 227-230) GVERVLFLFLNDFIEQFPMINPGVPIKRAHTPHIEPAtty. Dkt: BERK-543WO KSAIEELPFYQYVKEDIAMVLNGAKEKLLRALELT KAPGGPAPRADNFLDDLAQIDELIQHQDDFSELYK RVPAVSFKRAKAVKGDEFDPALLDEATDLRNGAKAtty. Dkt: BERK-543WO GLDLAEVYYGLALQMLTYLDLSITHSADWLGMR ATPAGVLYFHIHDPMIQSNLPLGLDEIEQEIFKKFK MKGLLLGDQEVVRLMDTTLQEGRSNIINAGLKKDCAST Modulator Inhibitors

[0122] In some embodiments, the CAST modulator comprises an inhibitor of gene that decreases the editing efficiency of a subject transposon system (i.e., a CAST modulator can inhibit an inhibitor of gene editing efficiency such as, e.g., hha or RecD). Thus, in some cases, the CAST modulator may inhibit function and / or expression of a target gene (e.g., RecD or hha) that decreases the editing efficiency of a subject transposon system, thereby resulting in an increase of editing efficiency. In particular embodiments, a CAST modulator may comprise a nucleotide sequence that is anti-sense to (i.e., the reverse complement of) at least a part of the transcript of a target gene (e.g., RecD or hha) that decreases the editing efficiency of a subject transposon system. In some cases (e.g., if using a CAST modulator in a eukaryotic cell) the CAST modulator inhibitsAtty. Dkt: BERK-543WO expression of the target gene through RNA interference (RNAi). In some cases, the CAST modulator may comprise a nucleotide sequence that encodes a nucleotide sequence that is anti- sense to (i.e., the reverse complement of) at least a part of the transcript of a target gene that decreases the editing efficiency of a subject transposon system.

[0123] In some cases, an inhibitor of gene editing efficiency is a protein encoded by recD (i.e., a RecD polypeptide). In some cases, an inhibitor of gene editing efficiency is a protein encoded by hha (i.e., a Hha polypeptide). In some cases, an inhibitor of gene editing efficiency is a protein encoded by cysQ (i.e., a CysQ polypeptide). In some cases, an inhibitor of gene editing efficiency is a protein encoded by asnV (i.e., an AsnV polypeptide). In some cases, an inhibitor of gene editing efficiency is encoded by recD (i.e., a RecD polypeptide) or encoded by hha (i.e., a Hha polypeptide). In some cases, an inhibitor of gene editing efficiency is encoded by recD (i.e., a RecD polypeptide), or encoded by hha (i.e., a Hha polypeptide), or encoded by asnV (i.e., an AsnV polypeptide), or encoded by hha (i.e., a Hha polypeptide).

[0124] In some cases, a CAST modulator may comprise a nucleotide sequence that is, or encodes a nucleotide sequence that is, anti-sense to (i.e., the reverse complement of) at least a part of a transcript encoding an amino sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the hha amino acid sequences set forth in SEQ ID NOs: 231-234.

[0125] In some cases, a CAST modulator may comprise a nucleotide sequence that is, or encodes a nucleotide sequence that is, anti-sense to (i.e., the reverse complement of) at least a part of a transcript encoding an amino sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the RecD amino acid sequences set forth in SEQ ID NOs: 328-365.

[0126] In some cases, a CAST modulator may comprise a nucleotide sequence that is, or encodes a nucleotide sequence that is, anti-sense to (i.e., the reverse complement of) at least a part of a transcript encoding an amino sequence having at least 80% (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or 100%) amino acid sequence identity to any one of the amino acid sequences set forth in SEQ ID NOs: 231-234 and 328-365. CAST Modulator Small Molecules

[0127] In some embodiments, the CAST modulator comprises a small molecule. In some embodiments, CAST modulator small molecule is a small molecule that promotes homologous recombination (e.g., promotes homologous recombination DNA repair pathways). Examples of small molecules that promote homologous recombination (e.g., promote homologous recombination in eukaryotes) include, without limitation, RS-1 (CAS No.312756-74-4), SCR7 (CAS No.14892-97-8), and L755507 (CAS No.159182-43-1).Atty. Dkt: BERK-543WO Promoters

[0128] The nucleotide sequence encoding the CAST complex polypeptides and / or the nucleotide sequence encoding the guide RNA can be operably linked to a promoter that is functional in a prokaryotic cell. In some embodiments, the nucleotide sequence encoding the CAST complex polypeptides and / or the nucleotide sequence encoding the guide RNA can be operably linked to a promoter that is functional in a eukaryotic cell. In some cases, the nucleotide sequence encoding the CAST complex polypeptides is operably linked to a first promoter; and the nucleotide sequence encoding the guide RNA is operably linked to a second promoter. In some cases, the nucleotide sequence encoding the CAST complex polypeptides and the nucleotide sequence encoding the guide RNA are operably linked to the same promoter.

[0129] In some embodiments, where the CAST modulator comprises a nucleotide sequence(s) encoding one or more CAST modulator polypeptides, the nucleotide sequence(s) encoding the one or more CAST modulator polypeptides can be operably linked to a promoter that is functional in a prokaryotic cell. In some embodiments, the nucleotide sequence(s) encoding the one or more CAST modulator polypeptides can be operably linked to a promoter that is functional in a eukaryotic cell.

[0130] In some cases, a nucleotide sequence encoding one or more CAST polypeptides is operably linked to a single promoter, that is, each of the nucleotide sequence(s) encoding a CAST polypeptide is operably linked to the same, single promoter. In some cases, each of the nucleotide sequence(s) encoding a CAST polypeptide is operably linked to a different promoter.

[0131] Suitable promoters include, constitutive promoters and inducible promoters. Inducible promoters include sugar-inducible promoters (e.g., lactose-inducible promoters; arabinose- inducible promoters); amino acid-inducible promoters; alcohol-inducible promoters; and the like. Suitable promoters include, e.g., lactose-regulated systems (e.g., lactose operon systems, sugar- regulated systems, isopropyl-beta-D-thiogalactopyranoside (IPTG) inducible systems, arabinose regulated systems (e.g., arabinose operon systems, e.g., an ARA operon promoter, pBAD, pARA, portions thereof, combinations thereof and the like), synthetic amino acid regulated systems, fructose repressors, a tac promoter / operator (pTac), tryptophan promoters, PhoA promoters, recA promoters, proU promoters, cst-1 promoters, tetA promoters, cadA promoters, nar promoters, PLpromoters, cspA promoters, and the like, or combinations thereof. In certain cases, a promoter comprises a Lac-Z, or portions thereof. In some cases, a promoter comprises a Lac operon, or portions thereof. In some cases, an inducible promoter comprises an ARA operon promoter, or portions thereof. In certain embodiments an inducible promoter comprises an arabinose promoter or portions thereof. An arabinose promoter can be obtained from any suitable bacteria. In some cases, an inducible promoter comprises an arabinose operon of E. coli or B. subtilis. In some cases, an inducible promoter is activated by the presence of a sugar or an analog thereof. Non-limiting examples of sugars and sugar analogs include lactose, arabinose (e.g., L-arabinose), glucose, sucrose, fructose, IPTG, and the like. Suitable promoters include a T7 promoter; a pBAD promoter; a lacIQ promoter; and the like. In some cases, the promoter is a J23119 promoter. Many bacterialAtty. Dkt: BERK-543WO promoters are known in the art; bacterial promoters can be found on the internet at parts(dot)igem(dot)org / promoters.

[0132] In some cases, suitable promoters include eukaryotic promoters. Examples of suitable eukaryotic promoters include, without limitation, cytomegalovirus (CMV) immediate early promoters, herpes simplex virus (HSV) thymidine kinase promoters, SV40 early and late promoters, Rous-Sarcoma virus (RSV) promoters, β-actin promoters, tubulin promoters, and EF1α promoters. Transposons

[0133] A transposon suitable for inclusion in a nucleic acid construct of a system of the present disclosure can have a length of up to about 100 kilobases (kb). For example, a transposon can have a length of from 0.1 kb to 0.5 kb, from 0.5 kb to 1 kb, from 1 kb to 5 kb, from 5 kb to 10 kb, from 10 kb to 15 kb, from 15 kb to 20 kb, from 20 kb to 25 kb, from 25 kb to 30 kb, from 30 kb to 35 kb, from 35 kb to 40 kb, from 40 kb to 45 kb, from 45 kb to 50 kb, from 50 kb to 55 kb, from 55 kb to 60 kb, from 60 kb to 65 kb, from 65 kb to 70 kb, from 70 kb to 75 kb, from 75 kb to 80 kb, from 80 kb to 85 kb, from 85 kb to 90 kb, from 90 kb to 95 kb, or from 95 kb to 100 kb.

[0134] A transposon suitable for inclusion in a nucleic acid construct of a system of the present disclosure can comprise one or more of: a) one or more nucleotide sequences encoding one or more polypeptides that confer on a prokaryotic cell resistance to one or more antibiotics; b) one or more nucleotide sequences encoding one or more enzymes in a biosynthetic pathway; c) one or more nucleotide sequences encoding one or more enzymes in a carbon utilization pathway (e.g., a polysaccharide utilization pathway); d) one or more nucleotide sequences encoding one or more polypeptides comprising a light-oxygen-voltage-sensing domain (LOV domain); e) a screenable marker (a detectable polypeptide; e.g., a polypeptide that provides a detectable signal such as a fluorescent signal); f) a polypeptide that provides for detection of an analyte in a bacterium; g) one or more nucleotide sequences encoding one or more therapeutic polypeptides; h) one or more nucleotide sequences encoding one or more nutritional polypeptides; i) one or more nucleotide sequences encoding one or more polypeptides that confer antibiotic sensitivity on a target prokaryotic cell; j) one or more nucleotide sequences encoding one or more polypeptides that facilitate isolation of a target prokaryotic cell; k) one or more nucleotide sequences encoding one or more enzymes in a nitrogen utilization pathway; l) one or more nucleotide sequences encoding one or more enzymes in a sulfur utilization pathway; m) one or more nucleotide sequences encoding one or more enzymes that degrade an allergen; n) one or more nucleotide sequences encoding one or more polypeptides that confers resistance to a phage (e.g., a bacteriophage); o) one or more nucleotide sequences encoding one or more polypeptides that provide for mobility of a gene edit; and the like. A transposon can include one or more nucleotide sequences encoding one or more polypeptides that allow establishment of a unique metabolic niche, e.g., ability to utilize a particular carbon source that is not normally a carbon source utilized by a particular bacterium (e.g., lactose, porphyrin, and the like).Atty. Dkt: BERK-543WO

[0135] A transposon can function to knock out an endogenous nucleic acid in a target bacterium, e.g., to delete all or a portion of an endogenous nucleic acid in a target prokaryotic cell or to introduce a loss-of-function mutation in an endogenous nucleic acid in a target prokaryotic cell. A “knockout” includes deletion of all or a portion of a nucleic acid; and includes introduction of a loss-of-function mutation in a nucleic acid. For example, a transposon can function to delete all or a portion of an endogenous nucleic acid in a target prokaryotic cell (e.g., target bacterium; target archaeon), or to introduce a loss-of-function mutation in an endogenous nucleic acid in a target prokaryotic cell, where the endogenous nucleic acid comprises one or more nucleotide sequences encoding one or more polypeptides that confer on a prokaryotic cell resistance to one or more antibiotics. A transposon can function to generate an auxotroph, e.g., an amino acid auxotroph (see, e.g., FIG.23 to FIG.25). A transposon can function to knock out an essential gene (e.g., a nucleic acid encoding one or more polypeptides that are essential to cell survival, cell proliferation, cell metabolism, etc.). A transposon can function to knock out a nucleic acid encoding a toxin. A transposon can function to knock out a counter-selectable gene, or a gene that confers a fitness advantage in a certain growth condition or medium composition (e.g., a galK knockout can grow in presence of 2-deoxygalactose; a pyrF knockout can grow in presence of 5- fluoroorotic acid; a thyA knockout can grow in presence of trimethoprim; etc.)

[0136] A transposon can comprise one or more nucleotide sequences encoding one or more polypeptides that confer resistance to one or more antibiotics in a target prokaryotic cell.

[0137] A transposon can comprise: a) one or more nucleotide sequences encoding magnetosome biosynthetic pathway polypeptides; b) one or more nucleotide sequences encoding gas vesicle biosynthetic polypeptides; c) one or more nucleotide sequences encoding one or more polypeptides in a porphyrin polysaccharide utilization pathway; d) one or more nucleotide sequences encoding one or more polypeptides in a glycosaminoglycan utilization pathway; e) one or more nucleotide sequences encoding one or more polypeptides in a glycosaminoglycan utilization pathway; f) one or more nucleotide sequences encoding one or more polypeptides in a non-caloric artificial sweetener utilization pathway; f) one or more nucleotide sequences encoding one or more polypeptides in a B-vitamin biosynthetic pathway; g) one or more nucleotide sequences encoding one or more polypeptides in an ethanolamine utilization pathway; h) one or more nucleotide sequences encoding one or more polypeptides in a sucrose utilization pathway; i) one or more nucleotide sequences encoding one or more polypeptides in a mevalonate biosynthetic pathway; j) one or more nucleotide sequences encoding one or more polypeptides in a polyketide biosynthetic pathway; and the like.

[0138] A transposon can comprise one or more nucleotide sequences encoding one or more polypeptides that provide for isolation of a target prokaryotic cell; e.g., a FLASH tag; FAST; iLOV; phiLOV; smURFP, IFP2.0; evoglow-Pp1; UnaG; a SNAP tag; a CLIP tag; a Halo tag; a spinach aptamer; mango aptamer; and the like. See, e.g., Thorn (2017) Mol. Biol. Cell 28:848; and Wang et al. (2017) Mol. Bhiochem. Parasitol.216:1. A transposon can comprise one or more nucleotide sequences encoding one or more polypeptides fluorescent proteins or tags that areAtty. Dkt: BERK-543WO detectable in anaerobic conditions, such as an anaerobic green fluorescent protein (GFP); see, e.g., Landete et al. ((2015) App. Microbiol. Biotechnol.99:6865) and Streett et al. (2019) Appl. Environmental Microbiol.85:e00622. Tagging surface exposed proteins with FLAG tag, His tag, Myc tag and the like, to be immunolabeled with fluorescence / magnetic-conjugated antibodies. Also suitable are tetracysteine tags to enable staining with biarsenical dyes (e.g., for staining with FlAsH and ReAsH dyes).

[0139] A transposon can comprise a nucleotide sequence encoding a fluorescent polypeptide. Suitable fluorescent proteins include, but are not limited to, green fluorescent protein (GFP) and variants thereof, blue fluorescent protein (BFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced GFP (EGFP), enhanced CFP (ECFP), enhanced YFP (EYFP), GFPS65T, Emerald, Topaz (TYFP), Venus, Citrine, mCitrine, GFPuv, destabilised EGFP (dEGFP), destabilized ECFP (dECFP), destabilized EYFP (dEYFP), mCFPm, Cerulean, T-Sapphire, CyPet, YPet, mKO, HcRed, t-HcRed, DsRed, DsRed2, DsRed-monomer, J-Red, dimer2, t-dimer2(12), mRFP1, pocilloporin, Renilla GFP, Monster GFP, paGFP, Kaede protein and kindling protein, Phycobiliproteins and Phycobiliprotein conjugates including B-Phycoerythrin, R-Phycoerythrin and Allophycocyanin. Other examples of fluorescent proteins include mHoneydew, mBanana, mOrange, dTomato, tdTomato, mTangerine, mStrawberry, mCherry, mGrape1, mRaspberry, mGrape2, mPlum (Shaner et al. (2005) Nat. Methods 2:905-909), and the like. See, e.g., Thorn (2017) Mol. Biol. Cell 28:848. CAST-recognition sites

[0140] As noted above, a transposon system of the present disclosure comprises a transposon or an insertion site for a transposon, where the transposon or the insertion site for a transposon is flanked by recognition sites (nucleotide sequences) that are bound by and cleaved by a CAST complex. The recognition sites are referred to as “left end” and “right end.” Recognition sites bound by and cleaved by a CAST complex are known in the art.

[0141] For example, “left end” and “right end” recognition sites bound by and cleaved by a VcCAST are: TGTTGATGCAACCATAAAGTGATATTTAATAATTATTTATAATCAGCAACTTAACCAC AAAACAACCATATATTGATATCTCACAAAACAACCATAAGTTGATATTTT (left end; SEQ ID NO:21); and GCAATATCAATTTATGGGTGTGATAATTATCAATTTATGGGTGTAATTATCATTTTATG GTTGTATCAACA (right end; SEQ ID NO:22).

[0142] As another example, “left end” and “right end” recognition sites bound by and cleaved by an ShCAST are: TGTACAGTGACAAATTATCTGTCGTCGGTGACAGATTAATGTCATTGTGACTATTTAAT TGTCGTCGTGACCCATCAGCGTTGCTTAATTAATTGATGACAAATTAAATGTCA (left end; SEQ ID NO:23); and CGACAGTCAATTTGTCATTATGAAAATACACAAAAGCTTTTTCCTATCTTGCAAAGCG ACAGCTAATTTGTCACAATCACGGACAACGACATCTATTTTGTCACTGCAAAGAGGTTAtty. Dkt: BERK-543WO ATGCTAAAACTGCCAAAGCGCTATAATCTATACTGTATAAGGATTTTACTGATGACAA TAATTTGTCACAACGACATATAATTAGTCACTGTACA (right end; SEQ ID NO:24). Guide RNA

[0143] As noted above, a transposon system of the present disclosure comprises a nucleotide sequence encoding one or more guide RNAs. The guide RNA comprises: i) a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic genome; and ii) a nucleotide sequence that binds to a polypeptide in the CAST complex. For example, the guide RNA comprises: i) a targeter RNA that comprises a nucleotide sequence (“guide sequence”) that hybridizes to a target nucleotide sequence in a prokaryotic genome; and ii) an activator RNA that comprises a nucleotide sequence that binds to a polypeptide in the CAST complex. A CAST forms a complex with a guide RNA. A CAST / guide RNA complex directs a transposon to a genomic site complementary to a guide RNA. See, e.g., Klompe et al. (2019) Nature 571:219; and Peters et al. (2019) Mol. Microbiol.112:1635.

[0144] In some cases, a transposon system of the present disclosure comprises a nucleotide sequence encoding a single guide RNA. In some cases, a transposon system of the present disclosure comprises nucleotide sequences encoding two or more guide RNAs, each guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome. For example, in some cases, a transposon system of the present disclosure comprises nucleotide sequences encoding 2, 3, 4, or 5 (or more than 5) different guide RNAs, each targeted to a different target nucleic acid.

[0145] A nucleic acid that binds to a polypeptide in a CAST complex, forming a CAST / guide nucleic acid complex, and targets the CAST / guide nucleic acid to a specific target sequence within a target DNA (e.g., prokaryotic genome) is referred to herein as a “guide RNA.” It is to be understood that in some cases, a hybrid DNA / RNA can be made such that a guide RNA includes DNA bases in addition to RNA bases - but the term “guide RNA” is still used herein to encompass such hybrid molecules. A subject guide RNA includes a guide sequence (also referred to as a “spacer”)(that hybridizes to target sequence of a target DNA) and a constant region (e.g., a region that is adjacent to the guide sequence and binds to a polypeptide in the CAST complex). A “constant region” can also be referred to herein as a “scaffold.”

[0146] The guide sequence has complementarity with (hybridizes to) a target sequence of the target DNA. In some cases, the guide sequence is 15-35 nucleotides (nt) in length (e.g., 15-26, 15- 24, 15-22, 15-20, 15-18, 16-28, 16-26, 16-24, 16-22, 16-20, 16-18, 17-26, 17-24, 17-22, 17-20, 17- 18, 18-26, 18-24, 30-32, 28-32, or 18-22 nt in length). In some cases, the guide sequence is 18-24 nucleotides (nt) in length. In some cases, the guide sequence is at least 15 nt long (e.g., at least 16, 18, 20, or 22 nt long). In some cases, the guide sequence is at least 17 nt long. In some cases, the guide sequence is at least 18 nt long. In some cases, the guide sequence is at least 20 nt long. In some cases, the guide sequence is 32 nt long. In some cases, VcCAST guides are included in a CRISPR array (repeat-spacer-repeat). In some cases, a ShCAST guides includes a 23-nt target complementarity.Atty. Dkt: BERK-543WO

[0147] In some cases, the guide sequence has 80% or more (e.g., 85% or more, 90% or more, 95% or more, or 100% complementarity) with the target sequence of the target DNA. In some cases, the guide sequence is 100% complementary to the target sequence of the target DNA. In some cases, the target DNA includes at least 15 nucleotides (nt) of complementarity with the guide sequence of the guide RNA.

[0148] In some cases, the constant region of a guide RNA is 15 or more nucleotides (nt) in length (e.g., 18 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, 29 or more, 30 or more, 31 or more nt, 32 or more, 33 or more, 34 or more, or 35 or more nt in length). In some cases, the constant region of a guide RNA is 18 or more nt in length.

[0149] In some cases, the guide RNA is a dual-molecule guide RNA. In some cases, the guide RNA is a single-molecule RNA (also referred to as a “single guide RNA” or “sgRNA”).

[0150] As an example, a crRNA for VcCAST system is GTGAACTGCCGAGTAGGTAGCTGATAACGAGACCTCGTTTACCTATCGGTCTCGTGAA CTGCCGAGTAGGTAGCTGATAAC (SEQ ID NO:25).

[0151] As an example, a sgRNA for the ShCas12k is ATATTAATAGCGCCGCAATTCATGCTGCTTGCAGCCTCTGAATTTTGTTAAATGAGGGT TAGTTTGACTGTATAAATACAGTCTTGCTTTCTGACCCTGGTAGCTGCTCACCCTGATG CTGCTGTCAATAGACAGGATAGGTGCGCTCCCAGCAATAAGGGCGCGGATGTACTGCT GTAGTGGCTACTGAATCACCCCCGATCAAGGGGGAACCCTAAATGGGTTGAAAGGGA GACCGAGATCTCGAGGTCTCC (SEQ ID NO:26). Target prokaryotic cells

[0152] Target prokaryotic cells include bacteria and archaea. In some cases, the target prokaryotic cells are bacteria. In some cases, the target prokaryotic cells are archaea.

[0153] In some cases, target prokaryotic cells include bacteria and / or archaea that have not yet been cultured or isolated in a laboratory in monoculture. This would include most phyla of the candidate phyla radiation, most archaeal phyla, and numerous phyla of bacteria. See, e.g., FIG.2 of Hug et al. (2016) Nature Microbiol.1:16048.

[0154] Target prokaryotic cells include prokaryotic cells found in a natural environment such as the gastrointestinal tract of a mammal (e.g., a human); the microbiome of a human; the microbiome of a non-human animal soil; hot springs; oceans; marshland; swamps; etc. Target prokaryotic cells include prokaryotic cells found in wastewater, agricultural runoff, and the like. Target prokaryotic cells include prokaryotic cells involved in food processing (e.g., fermentations to produce beverages or food that rely on a mixed community of cells such as with kimchi, soy sauce, or kombucha). Target prokaryotic cells include prokaryotic cells present in the rhizosphere. Target prokaryotic cells include prokaryotic cells present on the plant surface microbiome (the plant microbiome). Target prokaryotic cells include prokaryotic cells found in industrial processes relying on communities of microorgansisms such as industrial wastewater treatment or bioreactorsAtty. Dkt: BERK-543WO used for bioremediation of wastes (i.e. thiocyanate (SCN) degradation reactors used for gold mining runoff). Target prokaryotic cells include prokaryotic cells that find use in and / or are found in one or more of: the plant microbiome, food processing (e.g., wine, cheese, yogurt, etc.), bioremediation, and industrial processes. Target prokaryotic cells include pathogens (e.g., human pathogens) that harbor or help drive the spread of antibiotic resistance genes.

[0155] Target bacteria include gram-negative bacteria (e.g., gram-negative bacteria relevant in industrial biotechnology and bioremediation applications or relevant in the spread of antibiotic resistance genes). Target gram-negative bacteria include bacteria of the phyla Proteobacteria, Cyanobacteria, Spirochaetota, Chlorobiota, and Chloroflexota. Target gram-negative bacteria include bacteria of the genera Acenitobacter, Enterobacter, Escherichia, Helicobacter, Klebsiella, Legionella, Moraxella, Pseudomonas, Salmonella, Shigella, and Stenotrophomonas. In particular embodiments, target bacteria include Pseudomonas putida and / or Klebsiella michiganensis.

[0156] Target bacteria include bacteria present in the human gastrointestinal tract. Target bacteria include bacteria of the phyla Firmicutes, Bacteroidetes, Actinobacteria, and Proteobacteria. Target bacteria include bacteria of the genera Lactobacillus, Bacteroides, Clostridum, Faecalibacterium, Eubacterium, Ruminococcus, Peptococcus, Roseburia, Peptostreptococcus, Bifidobacterium, Alistipes, Parabacteroides, Porphyromonas, Prevotella, Collinsalla, Escherichia, and Desulfovibrio. See, e.g., Rinninella et al. (2019) Microoganisms 7:14. Examples of target bacteria include, e.g., Bacteroides fragilis ssp. vulgatus, Collinsella aerofaciens, Bacteroides fragilis ssp. thetaiotaomicron, Peptostreptococcus productus II, Parabacteroides distasonis, Faecalibacterium prausnitzii, Coprococcus eutactus, Peptostreptococcus productus I, Ruminococcus bromii, Bifidobacterium adolescentis, Gemmiger formicilis, Bifidobacterium longum, Eubacterium siraeum, Ruminococcus torques, Eubacterium rectale, Eubacterium eligens, Bacteroides eggerthii, Clostridium leptum, Bacteroides fragilis ssp. A, Eubacterium biforme, Bifidobacterium infantis, Eubacterium rectale, Coprococcus comes, Pseudoflavonifractor capillosus, Ruminococcus albus, Dorea formicigenerans, Eubacterium hallii, Eubacterium ventriosum, Fusobacterium russi, Ruminococcus obeum, Eubacterium rectale, Clostridium ramosum, Lactobacillus leichmannii, Ruminococcus callidus, Butyrivibrio crossotus, Acidaminococcus fermentans, Eubacterium ventriosum, Bacteroides fragilis ssp. fragilis, Coprococcus catus, Aerostipes hadrus, Eubacterium cylindroides, Eubacterium ruminantium. Staphylococcus epidermidis, Eubacterium limosum, Tissirella praeacuta, Fusobacterium mortiferum, Fusobacterium naviforme, Clostridium innocuum, Clostridium ramosum, Propionibacterium acnes, Ruminococcus flavefaciens, Bacteroides fragilis ssp. ovatus, Fusobacterium nucleatum, Fusobacterium mortiferum, Escherichia coli, Gemella morbillorum, Finegoldia magnus, Streptococcus intermedius, Ruminococcus lactaris, Eubacterium tenue, Eubacterium ramulus, Bacteroides clostridiiformis ssp. clostridliformis, Bacteroides coagulans, Prevotella oralis, Prevotella ruminicola, Odoribacter splanchnicus, and Desuifomonas pigra.

[0157] Target bacteria include bacteria present in the gastrointestinal tract of an ungulate (e.g., a bovine; an equine; an ovine; a caprine; etc.).Atty. Dkt: BERK-543WO

[0158] Other target bacteria include, e.g., bacteria associated with nosocomial infections in humans. Other target bacteria include soil bacteria.

[0159] In some cases, a target prokaryotic cell is one that is refractory to genetic modification by electroporation. In some cases, a target prokaryotic cell is one that is refractory to genetic modification by chemically-induced competence (e.g., competence induced by calcium chloride, rubidium chloride, and the like). In some cases, a target prokaryotic cell is one that is refractory to genetic modification by heat shock. In some cases, a target prokaryotic cell is one that is refractory to natural transformation. In some cases, a target prokaryotic cell is one that is refractory to isolation. In some cases, a target prokaryotic cell is one that is refractory growth in monoculture (e.g., in an industrial setting, a research laboratory setting, or the like).

[0160] Archaea that are suitable target prokaryotic cells include, e.g., archaea any species in any of the phyla Aenigmarchaeota, Diapherotrites, Nanoarchaeota, Nanohaloarchaeota, Micrarchaeota, Pacearchaeota, Parvarchaeota, Woesearchaeota, Aigarchaeota, Bathyarchaeota, Crenarchaeota, Geoarchaeota, Korarchaeota, Thaumarchaeota, Lokiarchaeota, Thorarchaeota, Odinarchaeota, Heimdallarchaeota, and the like. Target eukaryotic cells

[0161] Target eukaryotic cells include fungal cells, plant cells, animal cells, mammalian cells, human cells, and the like). Suitable eukaryotic cells can include, without limitation, a plant cell (e.g., cells from plant crops, fruits, vegetables, grains, soy bean, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, angiosperms, ferns, clubmosses, hornworts, liverworts, mosses, dicotyledons, monocotyledons, etc.), an algal cell, (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, and the like), seaweeds (e.g. kelp) a fungal cell (e.g., a yeast cell, a cell from a mushroom), an animal cell, a cell from an invertebrate animal (e.g., fruit fly, cnidarian, echinoderm, nematode, etc.), a cell from a vertebrate animal (e.g., fish, amphibian, reptile, bird, mammal), a cell from a mammal (e.g., an ungulate (e.g., a pig, a cow, a goat, a sheep); a rodent (e.g., a rat, a mouse); a non-human primate; a human; a feline (e.g., a cat); a canine (e.g., a dog); etc.), and the like. In some cases, the cell is a cell that does not originate from a natural organism (e.g., the cell can be a synthetically made cell; also referred to as an artificial cell).

[0162] A eukaryotic cell can be an in vitro cell (e.g., established cultured cell line). A cell can be an ex vivo cell (cultured cell from an individual). A eukaryotic cell can be an in vivo cell (e.g., a cell in an individual). A eukaryotic cell can be an isolated cell. A eukaryotic cell can be a cell inside of an organism. A eukaryotic cell can be an organism. A eukaryotic cell can be a cell in a cell culture (e.g., in vitro cell culture). A eukaryotic cell can be one of a collection of cells. A eukaryotic cell can be a plant cell or derived from a plant cell. A eukaryotic cell can be an animal cell or derived from an animal cell. A eukaryotic cell can be an invertebrate cell or derived from an invertebrate cell. A eukaryotic cell can be a vertebrate cell or derived from a vertebrate cell. A eukaryotic cell can be a mammalian cell or derived from a mammalian cell. A eukaryotic cell canAtty. Dkt: BERK-543WO be a rodent cell or derived from a rodent cell. A eukaryotic cell can be a human cell or derived from a human cell. A eukaryotic cell can be a fungi cell or derived from a fungi cell. A eukaryotic cell can be an insect cell. A eukaryotic cell can be an arthropod cell. A eukaryotic cell can be a protozoan cell. A eukaryotic cell can be a helminth cell. GENETICALLY MODIFIEDPROKARYOTICCELLS

[0163] The present disclosure provides a prokaryotic cell comprising a transposon system of the present disclosure. A prokaryotic cell of the present disclosure can be a “donor” bacterium, i.e., one that comprises a subject transposon system that is to be transferred to a target bacterium (a “recipient” bacterium). A prokaryotic cell of the present disclosure can be a “donor” archaeon, i.e., one that comprises a subject transposon system that is to be transferred to a target archaeon (a “recipient” archaeon).

[0164] The present disclosure also provides a genetically modified prokaryotic cell, where the genetically modified has been genetically modified by virtue of contact with a “donor” bacterium of the present disclosure; i.e., the genetically modified has been genetically modified with a transposon that is present in the transposon system present in the “donor” bacterium.

[0165] The present disclosure also provides a genetically modified prokaryotic cell, where the genetically modified has been genetically modified by virtue of contact with a “donor” archaeon of the present disclosure; i.e., the genetically modified has been genetically modified with a transposon that is present in the transposon system present in the “donor” archaeon.

[0166] The present disclosure provides a heterogeneous population of genetically modified prokaryotic cells, where the population comprises a plurality of genetically modified prokaryotic cells, which prokaryotic cells are the recipients of transposons present in a library of the present disclosure (e.g., are the recipients of a member of a library of the present disclosure). The heterogeneous population can comprise from 10 to 109different prokaryotic cells; e.g., from 10 to 102, from 102to 103, from 103to 104, from 104to 105, from 105to 106, from 106to 107, from 107to 108, or from 108to 109different prokaryotic cells, which comprise different transposons from a library of the present disclosure. In some cases, the population of prokaryotic cells are of the same genus. In some cases, the population of prokaryotic cells comprise bacteria of 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 (e.g., from 10 to 20, from 20 to 30, from 30 to 40, from 40 to 50,or more than 50), different genus and / or species. A heterogeneous population of genetically modified prokaryotic cells is also referred to as a “community” or a “prokaryotic cell community” or a “microbial community.”

[0167] The present disclosure provides a heterogeneous population of genetically modified bacteria, where the population comprises a plurality of genetically modified, which bacteria are the recipients of transposons present in a library of the present disclosure. The heterogeneous population can comprise from 10 to 109different bacteria; e.g., from 10 to 102, from 102to 103, from 103to 104, from 104to 105, from 105to 106, from 106to 107, from 107to 108, or from 108to 109different bacteria, which comprise different transposons from a library of the presentAtty. Dkt: BERK-543WO disclosure. In some cases, the population of bacteria are of the same genus. In some cases, the population of bacteria comprise bacteria of 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 (e.g., from 10 to 20, from 20 to 30, from 30 to 40, from 40 to 50,or more than 50), different genus and / or species. LIBRARIES

[0168] The present disclosure provides a library of nucleic acids comprising a plurality of member conjugative nucleic acid constructs of the present disclosure. Each member conjugative nucleic acid construct comprises: a) a nucleotide sequence encoding CAST complex polypeptides; b) a nucleotide sequence encoding one or more guide RNAs, each guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome; c) a transposon, wherein the transposon is flanked by recognition sites that are cleaved by the transposase, and d) a nucleotide sequence encoding one or more CAST modulator polypeptides.

[0169] In some cases, nucleotide sequence encoding the CAST complex polypeptides and / or the nucleotide sequence encoding the guide RNA can be operably linked to a promoter that is functional in a prokaryotic cell. In some cases, the nucleotide sequence encoding the CAST complex polypeptides is operably linked to a first promoter; and the nucleotide sequence encoding the guide RNA is operably linked to a second promoter. In some cases, the nucleotide sequence encoding the CAST complex polypeptides and the nucleotide sequence encoding the guide RNA are operably linked to the same promoter. In some cases, the nucleotide sequence encoding one or more CAST polypeptides is operably linked to a single promoter. That is, each of the nucleotide sequence(s) encoding a CAST polypeptide is operably linked to the same, single promoter. In some cases, each of the nucleotide sequence(s) encoding a CAST polypeptide is operably linked to a different promoter. Suitable promoters are described above.

[0170] In some cases, each member conjugative nucleic acid construct comprises a nucleotide sequence that provides a unique nucleotide sequence barcode that identifies the member (e.g., identifies the transposon present in each member and / or identifies the guide RNA(s) encoded by each member and / or identifies the promoter, etc.).

[0171] A library of the present disclosure can comprise from 10 to 109different members; e.g., from 10 to 102, from 102to 103, from 103to 104, from 104to 105, from 105to 106, from 106to 107, from 107to 108, or from 108to 109different member conjugative nucleic acid constructs of the present disclosure.

[0172] In some cases, a single member of the library can include a nucleotide sequence encoding two or more guide RNAs, each guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome. For example, in some cases, a single member of the library can a nucleotide sequence encoding 2, 3, 4, or 5 (or more than 5) different guide RNAs, each targeted to a different target nucleic acid.

[0173] A library of the present disclosure can be used to target more than one gene (nucleic acid) in a prokaryotic cell. A library of the present disclosure can be used to target a subset of genes (nucleic acids) in a prokaryotic cell. A library of the present disclosure can be used to target a single gene, or more than one gene (nucleic acid), in a specific species of prokaryotic cell. AAtty. Dkt: BERK-543WO library of the present disclosure can be used to target a single gene, or more than one gene (nucleic acid), in a subset of species (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 species) of prokaryotic cell present in a prokaryotic cell community. A library of the present disclosure can be used to target a single gene, or more than one gene (nucleic acid), in all members of a prokaryotic cell community. In some cases, the libraries of the present disclosure include genes encoding polypeptides involved in conjugation. In some cases, the libraries of the present disclosure lack genes encoding polypeptides involved in conjugation. METHODS OF EDITING THE GENOME OF A TARGET PROKARYOTIC CELL

[0174] The present disclosure provides a method of editing the genome of a target prokaryotic cell, the method comprising introducing into the target prokaryotic cell a transposon system of the present disclosure. The present disclosure provides a method of editing the genome of a target prokaryotic cell, the method comprising introducing into the target prokaryotic cell a single conjugative construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites, and a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some cases, provided is a method of editing the genome of a target prokaryotic cell, the method comprising introducing into the target prokaryotic cell a single conjugative construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites, and iv) a nucleotide sequence encoding one or more CAST modulator polypeptides. In some cases, provided is a method of editing the genome of a target prokaryotic cell, the method comprising introducing into the target prokaryotic cell a single construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some cases, provided is a method of editing the genome of a target prokaryotic cell, the method comprising introducing into the target prokaryotic cell a single construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and iv) a nucleotide sequence encoding one or more CAST modulator polypeptides.

[0175] The present disclosure provides a method of editing the genome of a target bacterium, the method comprising introducing into the target bacterium a transposon system of the present disclosure. The present disclosure provides a method of editing the genome of a target bacterium, the method comprising introducing into the target bacterium a single conjugative construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) aAtty. Dkt: BERK-543WO nucleotide sequence encoding a guide RNA; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some cases, provided is a method of editing the genome of a target bacterium, the method comprising introducing into the target bacterium a transposon system of the present disclosure. The present disclosure provides a method of editing the genome of a target bacterium, the method comprising introducing into the target bacterium a single conjugative construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites; and iv) a nucleotide sequence encoding one or more CAST modulator polypeptides.

[0176] In some cases, the transposon system is introduced via conditions that promote introduction of nucleic acid into prokaryotic cells including by electroporation, heat shock, use of chemically induced competence or other methods known in the art.

[0177] In some cases, a method of the present disclosure for editing the genome of a target prokaryotic cell comprises contacting one or more target bacteria with one or more “donor” prokaryotic cells of the present disclosure, where the one or more “donor” prokaryotic cells comprise a transposon system of the present disclosure or a single conjugative construct of the present disclosure. The transposon system of the present disclosure or the single conjugative construct of the present disclosure is transmitted conjugatively from the one or more “donor” prokaryotic cells to the one or more target (“recipient”) prokaryotic cell. Suitable target prokaryotic cells are described above.

[0178] In some cases, an editing method of the present disclosure further comprises identifying, within the contacted target prokaryotic cells, cells that have an edited genome. In other words, in some cases, the method further comprises identifying, within the contacted target prokaryotic cells, cells that are genetically modified by the method and that, as a result of the genetic modification, have a genetically modified genome. Identification can be carried out in a number of ways, depending on the transposon transmitted to the recipient target cells. For example, where the transposon comprises a nucleotide sequence encoding a fluorescent polypeptide, recipient cells that have an edited genome can be identified by detecting fluorescence in recipient target cells.

[0179] In some cases, an editing method of the present disclosure further comprises enriching the contacted target prokaryotic cells for target cells comprising an edited genome. Enriching can be carried out by selection. For example, where the transposon comprises one or more nucleotide sequences encoding one or more polypeptides that provide for resistance to one or more antibiotics, an editing method of the present disclosure can further comprise selecting target prokaryotic cells for antibiotic resistance. The enriching step can result in an enriched population in which from 50% to more than 99% of the cells (e.g., from 50% to 60%, from 60% to 70%, from 70% to 80%, from 80% to 90%, from 90% to 95%, from 95% to 99%, or more than 99%) of the cells have a genome that has been edited as a result of the contacting step.Atty. Dkt: BERK-543WO METHODS OF EDITING THE GENOME OF A TARGET EUKARYOTIC CELL

[0180] The present disclosure provides a method of editing the genome of a target eukaryotic cell, the method comprising introducing into the target eukaryotic cell a transposon system of the present disclosure. The present disclosure provides a method of editing the genome of a target eukaryotic cell, the method comprising introducing into the target eukaryotic cell a single construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; and iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites, and a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator. In some cases, provided is a method of editing the genome of a target eukaryotic cell, the method comprising introducing into the target eukaryotic cell a single construct comprising: i) a nucleotide sequence encoding polypeptides that form a CAST complex; ii) a nucleotide sequence encoding a guide RNA; iii) a transposon, or an insertion site for a transposon, flanked by CAST complex recognition sites, and iv) a nucleotide sequence encoding one or more CAST modulator polypeptides.

[0181] In some cases, the transposon system is introduced via known methods that promote introduction of nucleic acid into eukaryotic cells including e.g., viral infection, transfection, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct micro injection, nanoparticle-mediated nucleic acid delivery (see, e.g., Panyam et., al Adv Drug Deliv Rev.2012 Sep 13. pii: S0169-409X(12)00283-9. doi: 10.1016 / j.addr.2012.09.023), and the like.

[0182] In some cases, an editing method of the present disclosure further comprises identifying, within the target eukaryotic cells that have been introduced to a subject transposon system, cells that have an edited genome. In other words, in some cases, the method further comprises identifying, within the target eukaryotic cells, cells that are genetically modified by the method and that, as a result of the genetic modification, have a genetically modified genome. Identification can be carried out in a number of ways, depending on the transposon transmitted to the recipient target cells. For example, where the transposon comprises a nucleotide sequence encoding a fluorescent polypeptide, recipient cells that have an edited genome can be identified by detecting fluorescence in recipient target cells.

[0183] In some cases, an editing method of the present disclosure further comprises enriching the target prokaryotic cells that have been introduced to a subject transposon system for target cells comprising an edited genome. Enriching can be carried out by selection. For example, where the transposon comprises one or more nucleotide sequences encoding one or more polypeptides that provide for resistance to one or more antibiotics, an editing method of the present disclosure can further comprise selecting target prokaryotic cells for antibiotic resistance. The enriching step can result in an enriched population in which from 50% to more than 99% of the cells (e.g., from 50% to 60%, from 60% to 70%, from 70% to 80%, from 80% to 90%, from 90% to 95%, from 95% toAtty. Dkt: BERK-543WO 99%, or more than 99%) of the cells have a genome that has been edited as a result of the introducing step. EXAMPLES

[0184] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present invention, and are not intended to limit the scope of what the inventors regard as their invention nor are they intended to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g. amounts, temperature, etc.) but some experimental errors and deviations should be accounted for. Unless indicated otherwise, parts are parts by weight, molecular weight is weight average molecular weight, temperature is in degrees Celsius, and pressure is at or near atmospheric. Standard abbreviations may be used, e.g., bp, base pair(s); kb, kilobase(s); pl, picoliter(s); s or sec, second(s); min, minute(s); h or hr, hour(s); aa, amino acid(s); kb, kilobase(s); bp, base pair(s); nt, nucleotide(s); i.m., intramuscular(ly); i.p., intraperitoneal(ly); s.c., subcutaneous(ly); and the like.Atty. Dkt: BERK-543WO Example 1 Introduction

[0185] A significant barrier to broader and more efficient use of VchCAST is the limited understanding of the host factors that enhance and restrict its function. For instance, the addition of ClpX was found to facilitate VchCAST functionality in human HEK293T cells; however, ClpX has not been shown to be necessary for its function in bacteria. Within bacteria, integration host factor (IHF) is the only protein complex known to affect VchCAST function. A screen of variant VchCAST transposon ends in E. coli showed that the presence of IHF is required for efficient transposition. However, a significant opportunity remains for systematic identification of activators and inhibitors for VchCAST transposition.

[0186] In this study, we conducted a genome-wide screen to identify host factors that impact VchCAST activity, resulting in the validation of seven activators and two inhibitors. Informed by the identification of homologous recombination-involved effectors, RecD and RecA, we explored whether incorporating the efficient phage-derived λ-Red recombination system could enhance VchCAST editing performance. By integrating the λ-Red genes into our VchCAST vectors, we achieved a ~25-fold increase in editing efficiency. This strategy also facilitated more efficient integration in other industrially, environmentally, and medically relevant bacteria. Our research advances the understanding of CAST systems, revealing key regulatory factors and developing strategies to enhance editing efficiency. These improvements pave the way for broader application of CASTs across various host organisms, potentially enabling more efficient and versatile genome editing in both research and applied settings. Materials and Methods RB-TnSeq screen for E. coli regulators of CAST integration

[0187] To systematically identify E. coli genes that contribute to VchCAST-mediated integration, we conducted a genome-wide loss-of-function mutant screen and selected for VchCAST edits after conjugation in the KEIO_ML9 E. coli RB-TnSeq library. This approach allowed us to generate and select for VchCAST edits across a diverse array of insertion mutants. The library comprises 152,018 uniquely barcoded single-gene transposon insertion mutants, covering 3,728 nonessential protein-coding genes out of a total of 4,146 in the E. coli genome. To account for the library's inherent kanamycin resistance (kanR), two different antibiotic selection cargos (PPmtl-gmR, PPmtl-catP) were utilized in the RB-TnSeq and conjugation efficiency assays, with a mariner transposase system serving as a control.

[0188] The KEIO_ML9 RB-TnSeq library was inoculated into LB containing kanamycin (25 µg / mL) and grown at 37ºC with shaking. Donor strains harboring the VcDART and mariner vectors respectively in E. coli WM3064 (DAP auxotroph, pir+, RP4+) were grown at 37ºC with shaking in LB containing diaminopimelic acid (DAP; 0.3 mM) and gentamicin (screen 1; 50 µg / mL) or chloramphenicol (screen 2; 34 µg / mL). Three 10 OD*mL samples of the library overnight were washed and pelleted before being frozen at -80ºC as three technical replicate time zero (T=0) samples. After washing and resuspending in LB containing DAP, 1 OD*mL of donorAtty. Dkt: BERK-543WO was combined with 1 OD*mL of recipient in separate 1.5 mL Eppendorf tubes. Each combined donor-recipient sample was plated onto a plain LB agar petri plate lined with a MF-Millipore™ Membrane Filter for conjugation. After 6 hours of conjugation at 30ºC, cells from two conjugation plates were scraped into 20 mL of LB as a single technical replicate (six conjugation plates for three technical replicates for each donor-recipient combination).100 uL of the resuspended cells were set aside for 10-fold serial dilution and spot plating. The remaining cells from each technical replicate were plated onto selective BioAssay dishes of LB agar containing kanamycin and gentamicin or chloramphenicol. The three technical replicates of the library not introduced to any donor were plated on nonselective BioAssay dishes of LB agar with kanamycin. The BioAssay dishes were incubated at 30ºC for 12 hours. Cells on BioAssay dishes were then scraped and resuspended in 20 mL of LB. A two mL aliquot of the resuspended cells was taken for genomic DNA extraction via the QIAGEN DNeasy PowerSoil Pro Kit. Barcodes were amplified with the BarSeq_v3 primers following the BarSeq PCR protocol outlined in Wetmore et al.2015. Amplicons were pooled into a synthetic amplicon library and submitted for sequencing by Illumina NovaSeq PE150.

[0189] BarSeq sequencing reads were processed using the FEBA pipeline described previously. Samples that failed to meet the FEBA pipeline quality control criteria were excluded from further analysis. The pipeline-generated fitness scores (fit_logratios_good) were further analyzed using R (version 4.4.0) with the tidyverse package (version 2.0.0). Within an experimental trial, the difference between the VchCAST fitness scores and the mariner fitness scores was calculated. To identify genes consistently displaying strong differential effects, we focused on those with an absolute fitness difference greater than one in both experimental trials. Positive fitness values generated by the FEBA pipeline indicate a putative inhibitor of VchCAST, as knocking out the gene increases the fitness of the mutant. Conversely, negative fitness values indicate putative activators of VchCAST function. Validating activators and inhibitors with KEIO E. coli mutants

[0190] To validate the involvement of putative host factors identified in the RB-TnSeq screen, VchCAST editing was performed in KEIO collection deletion mutants of hits with the largest absolute fitness values. Primers were designed to confirm the kanamycin resistance gene insertion within each target locus of the expected KEIO mutants (data not shown). The ΔyicI mutant was selected as the negative control as it is a context-neutral genomic locus with no documented adverse fitness effects.

[0191] To assess how the presence or absence of individual E. coli genes affect VchCAST integration into the target genome, we implemented a conjugation-based editing efficiency assay (data not shown). Briefly, donor and recipient strains were grown for 16 hours, washed, resuspended in LB containing DAP, combined in a 1:1 ratio (0.1 OD*mL each), spotted on LB agar within a 24 well block, and allowed to conjugate for 6 hours at 30ºC. Afterward, spots were resuspended in one mL LB media. Ten-fold serial dilutions were performed with resuspended cells, spotted onto LB agar and LB agar with antibiotics, and incubated overnight at 30ºC. SerialAtty. Dkt: BERK-543WO dilution spot plates were left to grow until individual colonies formed. Finally, the plates were imaged and cell colonies were counted to compute editing efficiency, with transconjugants reflecting successful conjugation and insertion of selective cargo. Fold change in editing efficiency was computed by normalizing the editing efficiency of candidate regulator hits of interest to that of the ΔyicI control. Three biological replicates were used for each experimental condition. Construction and testing λ-Red VchCAST in E. coli

[0192] The DNA of λ-Red and Jungle Express (pJEx) were purchased as gblock gene fragments (IDT) and added to VchCAST vectors using Gibson and Golden Gate assemblies. For testing the effect of the phage-derived homologous recombination system on editing efficiency we cloned λ- Red onto a VchCAST backbone just after the origin of transfer, ensuring quick transcription in the recipient cell during conjugation. The tightly regulated and strong inducible Jungle Express promoter (pJEx) was used to minimize leaky λ-Red expression and toxicity of the vector. pJEx is inducible with crystal violet (CV) and transcriptional control is robust in Pseudomonadota. Plasmids, strains, synthesized DNA, and oligonucleotides (IDT, Coralville, USA) used in the study are listed in supplemental materials (data not shown). VchCAST vectors were assembled through multipart Golden Gate cloning. High-fidelity PCRs were performed with Q5 Hot Start High-Fidelity DNA polymerase (NEB). Golden Gate assembly enzymes (e.g. BsmBI-V2, BbsI, BsaI-HFV2, and T4 ligase) were ordered from NEB and used with the reported buffers following previously reported protocols. VchCAST assemblies were electroporated into electrocompetent E. coli EC100Dpir+ cells (LGC Biosearch). Clones were screened by colony PCR (cPCR) with 2x GoTaq Green Mastermix (Promega) and plasmids were isolated with a QIAprep Spin Miniprep Kit (Qiagen). Guide assemblies were electroporated into E. coli WM3064 pir+ and grown on the appropriate antibiotics plus DAP. All vectors were confirmed with whole-plasmid sequencing (Plasmidsaurus Labs).

[0193] Crystal violet induction was tested in E. coli to determine the concentration that produces the highest insertion efficiency (Figure 5). Conjugations were performed on agar plates containing the crystal violet at the optimal induction concentration (0.01 µM) for 6 hours before resuspension and selection. To account for variability in editing efficiency between biological replicates, λ-Red VchCAST treatments were paired and normalized to the negative VchCAST control. Each experiment was repeated with three biological replicates. Testing λ-Red VchCAST in P. putida and K. michiganensis

[0194] We first performed a quantitative assay in candidate strains to determine the frequency of mutants resistant to the antibiotics streptomycin / spectinomycin / carbenicillin (100 µg / mL and 400 µg / mL), chloramphenicol (34 µg / mL and 68 µg / mL), kanamycin (25 µg / mL, 50 µg / mL, 100 µg / mL and 200 µg / mL), and gentamicin (10 µg / mL, 20 µg / mL, and 40 µg / mL). Overnight cultures of each species were grown in LB at 30ºC, 10x serially diluted, and spotted onto LB agar without antibiotics (control) and onto each of the antibiotic concentrations. Colonies were counted after 16-40 hours of growth, and the antibiotic concentration exhibiting minimal or noAtty. Dkt: BERK-543WO detectable growth was chosen as a selection marker for genome editing experiments (Figure 8A- 8B). Based on these results, we constructed a VchCAST vector containing the Ppmtlpromoter driving the kanamycin resistance gene for selection in P. putida and K. michiganensis.

[0195] We identified safe sites and designed guides in P. putida (GCF_000007565.2) following previously reported methods and used a previously tested safe site guide in K.michiganensis. Briefly, intergenic regions between converging genes with a distance of 300-600 nucleotides were selected (data not shown). Candidate safe sites were excluded under any of the following circumstances: the region was located within or adjacent to predicted mobile genetic elements (MGEs), the region was flanked by essential genes (inspected using BioCyc), or the region contained non-coding RNA (ncRNA) features (inspected using Rfam). Within the selected safe site regions, protospacer adjacent motifs (PAMs) with the sequence 5'-CN-3' were identified. For each PAM, 32 nucleotides were added to generate a list of potential guides, which were chosen based on having a GC content of 40-60%, and ensuring that the insertion loci, approximately 49 bp downstream of the guide, remained within the intergenic region and that it would not accidentally disrupt a terminator sequence. Three guides per safe site were selected, and their off- target potential was assessed using a local BLASTn search (-dust no -word_size 4). Guides with off-target hits exhibiting the highest e-values and with minimal complementarity to off-targets in the seed region (first ~10 nt) were prioritized. Two guides targeting different safe sites were cloned and tested (Figure 8C).

[0196] Conjugation experiments in P. putida and K. michiganensis, were performed following the protocol described for E. coli with some modifications. Recipient and donor cells were grown 16 hours, at 30 ºC and 37 ºC, respectively. Conjugations were done for 20 hours at 30 ºC with crystal violet induction concentrations of 0.5 µM and 1 µM used for K. michiganensis and P. putida, respectively (Figure 8D-8E). Transconjugants were selected on LB plates containing kanamycin (50 µg / mL). Drip plates of 10 µL for each serial dilution were plated to increase sensitivity and reduce technical error. Three biological replicates were performed across different days. Insertion orientation and on-target analysis in E. coli, P. putida and K. michiganensis

[0197] Insertion analysis was preliminarily performed by cPCR on transconjugants following selection in all strains. Insertion orientation primers were designed to identify clones with right- left (T-RL) and left-right (T-LR) simple insertion, cointegration, and no integration outcomes (Figure S3A, S3C). Cointegration oligos were designed to amplify the junction between the VchCAST vector backbone and the resistance marker in the cargo. Simple insert oligos were designed to amplify off of the flanking genomic DNA to capture the full VchCAST cargo insert. Another set of oligos were designed to distinguish between T-RL versus T-LR simple insert products. Additional screening for insertion cointegration versus simple insertion was performed for E. coli by patch plating of transconjugants on LB agar plates containing 100 μg / mL carbenicillin, the VchCAST vector backbone resistance marker. Carbenicillin resistance in P.Atty. Dkt: BERK-543WO putida and K. michiganensis did not permit cointegrate screening by double selection and required insert orientation analysis after selection either by cPCR or WGS.

[0198] Insertion products from all recipient strains in this study (E. coli, P. putida and K. michiganensis) were further assayed for off-targets by whole-genome sequencing. To address colony heterogeneity, which has been observed previously, about a thousand colonies of transconjugant cells transformed with λ-Red VchCAST, were scraped from selection plates and resuspended. For each species investigated, resuspensions from three biological replicates were pooled together and replated on the appropriate dilution to yield thousands of single colonies. Colonies were scraped from selection plates and resuspended in a volume of LB media equivalent to OD600=3-4. An aliquot of this resuspension was used for high-molecular-weight genomic DNA (gDNA) extraction using the MasterPure™ Complete DNA and RNA Purification Kit (Biosearch Technologies). Genomic DNA samples were submitted for Oxford Nanopore long-read sequencing to Plasmidsaurus Labs.

[0199] The bioinformatics analysis was performed using custom scripts. The demultiplexed raw reads were filtered using nanofilt python package (v 2.8.0) to select reads with a Phred quality score of 20 (>Q20) and a minimum length of 150 base pairs (bp) (16). To identify the inserted cargo, the reads were aligned to the last 71-bp of the transposon right end, which consists of three 20-bp TnsB binding sites and a 8-bp terminal repeat. The alignment was performed using minimap2 software and reads mapping to less than 70% of the right end of cargo were filtered out.

[0200] For downstream analysis, a pipeline developed by Vo et al. (2021) was employed with some modifications. The reads containing the mapped region were extracted and aligned with the complete reference and plasmid genomes. The reads were classified as genomic, plasmid, or both based on whether more than 500 bp was mapped to the respective genomes.

[0201] The genomic coordinates of the mapped region were recorded to generate the genome- wide histograms of the integration location. The mapped region was categorized as on-target if it fell within the 100-bp window of the 3'end of the target site. Results

[0202] A whole genome mutant screen and validation to identify VchCAST regulators

[0203] To perform a genome-wide survey of genes involved with VchCAST function, we screened a library of loss-of-function mutants for the ability of VchCAST to integrate into a safe site (Figure 2A). The screen was conducted with a preexisting E. coli RB-TnSeq transposon mutant library. VchCAST was introduced to the library via conjugation, and the inserted cargo was selected for, ensuring all members of the library contained both the VchCAST and the background loss-of-function mutation (Figure 2A). The efficiency of VchCAST insertion into each mutant was then determined by measuring the abundance of each VchCAST-containing mutant. Based on their consistent phenotypes (absolute gene fitness > 1) in two screens conducted in parallel, we identified four genes as putative activators and eleven as putativeAtty. Dkt: BERK-543WO inhibitors (Figure 2B). Notably, the beta subunit of integration host factor (ihfB) emerged as the top activator from the screen (Figure 2B). As IHF is the only known activator of VchCAST insertion (8), this finding suggests the biological relevance of the screen results. The other subunit, ihfA, did not have a high enough abundance in the T=0 starting library to be considered in the analysis pipeline.

[0204] We validated hits with the largest absolute fitness values from the pooled RB-TnSeq screens by testing VchCAST editing in clonal gene deletion mutants. To prioritize candidates most likely to have direct interaction with VchCAST, we filtered for those possessing DNA- interacting, RNA-interacting, and protein-interacting functions as well as hypothetical proteins, as per the Gene Ontology (GO) database (data not shown). We then performed conjugation assays to quantify the relative editing efficiency of VchCAST in Keio deletion mutants of each of the nine filtered hits (Figure 3A). Editing efficiencies were compared to ΔyicI, a deletion mutant with no documented adverse fitness effects. Overall, mutants of putative activators (ihfB, cspC, ynbE, umuD, chaB, nudG, yneK) exhibited reduced VchCAST editing efficiency by 7.7 ± 21.7% to 99.9 ± 0.1% compared to the control. Knockouts of putative inhibitors (hha, recD) exhibited increased editing efficiency (4.0 ± 1.6-fold and 21.8 ± 7.2-fold) in comparison to the ΔyicI control (Figure 3A). The results from the whole-genome screen and the individual mutant validation were largely consistent, with recD and ihfB showing up as the strongest inhibitor and activator, respectively, in both experiments. Taken together, the validation results support the involvement of these factors in VchCAST’s genomic integration.

[0205] We were particularly compelled by the strong inhibitor classification of recD in VchCAST insertion revealed by the screen and validation, and sought to explore the rec genes further. RecD’s role in the RecBCD complex leads to the inhibition of homologous recombination, suggesting a beneficial role of factors involved in DNA repair during VchCAST integration. To test this hypothesized role, we delivered VchCAST via conjugation into Keio knock-out mutants of each rec gene (recABCD) and quantified editing efficiency normalized to ΔyicI (Figure 3B). Consistent with the screen results, the putative inhibitor ΔrecD showed a 9.8 ± 2.6-fold increase in editing efficiency compared to the neutral fitness mutant control. Knockouts of recB and recC slightly decreased insertion efficiency relative to the control (45.3 ± 25.5% and 15.0 ± 6.6%, respectively). ΔrecA, which did not have a high enough abundance in the T=0 starting library to be considered in the screen, had no viable transconjugant colonies above the limit of detection. The results of this experiment identify recA as an activator and support the role of homologous recombination in promoting VchCAST integration. Leveraging λ-Red to improve VchCAST editing efficiency in E. coli

[0206] Considering the role of the RecBCD complex and potentially homologous recombination more broadly in VchCAST integration, we hypothesized that introducing higher-efficiency recombination machinery may improve editing outcomes. To this end, the bacteriophage λ-Red genes (exo, beta, and gam) were cloned into a VchCAST plasmid (R6K, PPmtl-catP cargo) (Figure 4A). We optimized the induction level of λ-Red by crystal violet (Figure 5) for editing in E. coli.Atty. Dkt: BERK-543WO The induced λ-Red VchCAST treatment significantly increased editing efficiency by 25.7 ± 0.6 fold in BW25113 E. coli compared to the VchCAST control (Figure 4B). The NT control did not yield colonies within the level of detection for our assay. Preliminary insertion analysis of E. coli transconjugants by cPCR (n = 60 colonies) showed the same simple insert and co-integration rates between λ-Red VchCAST (86.7% simple insert, 13.3% co-integrate) and VchCAST (86.7% simple insert, 13.3% co-integrate) (Figure 6B), corresponding to previously reported rates. Further inspection of simple insertion orientation between λ-Red VchCAST (92.3% T-RL, 7.7% T-LR) and VchCAST (96.2% T-RL, 3.9% T-LR) by cPCR revealed similar rates to each other (Figure 6B) and the reported literature. WGS analysis confirmed that 100% of λ-Red VchCAST insertions were on-target in E. coli (data not shown). All together, these data support λ-Red increasing VchCAST editing efficiency without impacting the orientation or on-target specificity of edits. Co-integrate, simple insert, and simple insert orientations were also confirmed by gel analysis (Figure 6C).

[0207] We next tested which of the λ-Red genes on VchCAST were necessary for the observed increase in editing efficiency. Constructs containing inducible exo-beta, gam, and the full λ-Red operon on the backbone of the CAST editing vector were tested. Exo & Beta promote homologous recombination through their exonuclease and single-stranded binding activity respectively, while Gam inhibits RecBCD. We found that while the exo-beta construct does improve editing efficiency by 6.9 ± 1.8-fold, it doesn’t to the extent of all three lambda genes (28.8 ± 18.5-fold) compared to the control VchCAST (Figure 4C). Unsurprisingly, the presence of gam alone on the VchCAST vector did not significantly affect editing efficiency (Figure 4C). Taken together, these results suggest that the complete λ-Red system is necessary for maximizing the efficiency of VchCAST integration. λ-Red-assisted VchCAST editing in additional Gram-negative bacteria

[0208] Motivated by the successful implementation of λ-Red to increase VchCAST-directed editing efficiency in E. coli, we evaluated the impact of λ-Red on VchCAST-mediated DNA integration in other Gram-negative bacteria. We selected P. putida KT2440, a relevant species for applications in industrial biotechnology and bioremediation, and K.michiganensis M5a1, an important human pathogen driving the spread of antibiotic-resistance genes. The VchCAST and λ-Red VchCAST vectors were conjugated into both species and queried for their effect on insertion efficiency, insert orientation and on-target frequency as reported in E. coli. In both species, expression of λ-Red from the VchCAST vector resulted in increased insertion efficiency compared to VchCAST alone (Figure 7A-7B). In P. putida, we observed an increase without the inducer (4.7 ± 1.6 fold-increase) and in the presence of 1 uM CV (3.6 ± 1.1 fold-increase) (Figure 7A). This suggests that there may be leaky expression from the pJEx promoter in this strain, and basal levels of λ-Red are sufficient to enhance editing efficiency. In K. michiganensis, upon induction with 0.5 µM CV, there is a 7.7 ± 7.0-fold increase in percent editing efficiency (Figure 7B).Atty. Dkt: BERK-543WO We analyzed the VchCAST and λ-Red VchCAST insertion products by WGS analysis in both P. putida and K. michiganensis. In both bacteria, RNA-guided DNA insertions were 100% on-target, with the confidence corresponding to the depth of sequencing coverage (data not shown). Discussion

[0209] In this study, we conducted a genome-wide mutant screen to identify genes in E. coli that influence the efficiency of VchCAST, a promising tool for precise DNA insertion of large cargos in bacteria and eukaryotes. We screened for and then validated nine candidate genes, with eight being novel to this study, that either positively or negatively affect VchCAST insertion activity when disrupted (Figure 2B, 3A). Notably, our results highlighted the role of the RecBCD complex in CAST integration, with the disruption of recD increasing editing efficiency, while a recA deletion strongly decreased it (Figure 3B). Building on these insights, we leveraged the bacteriophage λ-Red genes (exo, beta, and gam) to enhance CAST insertion efficiency (Figure 4A). By optimizing the expression of these genes, we achieved improved editing efficiency not only in E. coli (Figure 4B), but also in P. putida (Figure 7A), an industrial model strain, and K. michiganensis (Figure 7B), an important human pathogen. This work provides a comprehensive survey of host factors influencing VchCAST integration and presents an approach to enhance its efficiency across various bacterial species.

[0210] Beyond the homologous recombination-associated RecD, we discovered seven other activators and inhibitors of the CAST complex that may further contribute to understanding the mechanism of VchCAST transposition (Figure 3A). Putative activator UmuD, through its role in inhibiting DNA polymerase III from binding to ssDNA, may protect the integration site from premature replication fork collision. Cold shock protein C (CspC), known to bind to single- stranded nucleotides, could stabilize the RNA guide or the single-stranded gap of the post-strand transfer intermediate. Interestingly, ClpX, which facilitated more efficient editing in human cells (6), potentially due to its speculated role in disassembling CASTs, was not identified in our screen. It is possible that in bacteria, ClpX’s functional role in human cells is redundant or provided by other proteins. The new activators and inhibitors found in this study provide valuable targets for further experiments to understand and control the function of VchCAST.

[0211] Among the strong activators and inhibitors identified and validated within this study, we were particularly interested in the role of the RecBCD complex. RecD was identified as the strongest inhibitor of VchCAST integration via the screen and Keio mutant validation (Figure 2B, 3A). RecD functions as an inhibitor to homologous recombination by blocking RecA loading, so it was expected that RecA - a single-stranded binding protein essential for Rec- mediated double-strand break repair - showed an opposing phenotype to RecD when single mutants were tested for VchCAST integration efficiency (Figure 3B). The weaker phenotypes of RecBC, but not RecA, may be explained by some amount of functional redundancy in Rec- mediated repair to other repair pathways, such as the RecF homologous recombination pathway. In the Mu transposon, the RecBCD complex facilitates the repair of a double-strand break that occurs during resolution of the insertion. Similarly, we hypothesize that the highly stableAtty. Dkt: BERK-543WO VchCAST post-transposition complex could stall replication forks and introduce a double-strand break at the target site, similar to Mu’s mechanism. RecBCD and RecA would then repair the double-strand break. This putative function is supported by our finding that the highly efficient homologous recombination system, λ-Red, improves the efficiency of VchCAST. Such a homologous recombination-dependent mechanism for VchCAST would contrast with the commonly stated assumption that transposition occurs independently of this DNA repair pathway.

[0212] In our experiments, we delivered VchCAST via suicide vector, which leads to lower insertion efficiencies but provides a more dynamic range to examine the effects of other activating and inhibiting proteins. Furthermore, using a suicide vector is advantageous for eventual microbiome editing applications since it is eliminated after cargo delivery, thus presenting fewer biocontainment concerns. The highest efficiencies reported in the literature are achieved in studies that deliver the VchCAST on a replicating plasmid which is then selected for, where insertions can be introduced to nearly 100% of cells. This is likely due to the presence of multiple copies of the VchCAST vector persisting in cells for an extended time. Given that λ-Red has often been used to increase homologous recombination efficiency on a replicating vector, we expect that our λ-Red VchCAST system would also work in that context.

[0213] Our findings on improving VchCAST efficiency in E. coli prompted us to investigate the broader applicability of this approach in other bacterial species. When the λ-Red VchCAST system was tested outside of E. coli, we found that it enhanced editing efficiencies, but to a lesser extent (Figure 7A-7B). This data suggests that insights from this screen will be most efficiently applied in a host-dependent manner. For example, to improve efficiency outside of E. coli, VchCAST could use recombination-promoting systems specific to the host of interest. In the case of strong activator IHF, it is common in Pseudomonadota and therefore may not be an imperative regulator for successful CAST insertion in that phylum; however, it could be integral for targeting strains without endogenous IHF. These considerations highlight the importance of tailoring optimization strategies to the specific genetic characteristics of the target organism when applying VchCAST systems more broadly.

[0214] We are excited about the potential of this work to bring the powerful CAST editing toolset to more biological systems. VchCAST has mostly been applied in Gammaproteobacteria. Understanding of host factors that enable and inhibit these systems is an important barrier to their use across the phylogenetic tree of bacteria. In the complex microbiomes that are most relevant for human health and the environment, editing efficiency is a major bottleneck for delivering functional cargo insertions via VchCAST. This work provides strategies for more efficient delivery so that the function of these communities can be better probed and controlled at a genetic level. Overall, our findings provide valuable insights into the factors influencing CAST efficiency and offer strategies for improving its performance across diverse organisms, paving the way for broader applications in genome editing.

[0215] ReferencesAtty. Dkt: BERK-543WO 1. Hsieh, S.-C. and Peters, J.E. (2024) Natural and Engineered Guide RNA-Directed Transposition with CRISPR-Associated Tn7-Like Transposons. Annu. Rev. Biochem., 93, 139–161. 2. Chang, C.-W., Truong, V.A., Pham, N.N. and Hu, Y.-C. (2024) RNA-guided genome engineering: paradigm shift towards transposons. Trends Biotechnol., 10.1016 / j.tibtech.2024.02.006. 3. Vento,J.M., Crook,N. and Beisel,C.L. (2019) Barriers to genome editing with CRISPR in bacteria. J. Ind. Microbiol. Biotechnol., 46, 1327–1341. 4. Gelsinger, D.R., Vo, P.L.H., Klompe, S.E., Ronda, C., Wang, H.H. and Sternberg, S.H. (2024) Bacterial genome engineering using CRISPR-associated transposases. Nat. Protoc., 10.1038 / s41596-023-00927-3. 5. Rubin,B.E., Diamond,S., Cress,B.F., Crits-Christoph,A., Lou,Y.C., Borges,A.L., Shivram,H., He,C., Xu,M., Zhou,Z., et al. (2022) Species- and site-specific genome editing in complex bacterial communities. Nat Microbiol, 7, 34–47. 6. Lampe, G.D., King, R.T., Halpin-Healy, T.S., Klompe, S.E., Hogan, M.I., Vo, P.L.H., Tang, S., Chavez, A. and Sternberg, S.H. (2024) Targeted DNA integration in human cells without double-strand breaks using CRISPR-associated transposases. Nat. Biotechnol.42, 87–98. 7. Yang,S., Zhu,J., Zhou,X., Zhang,J., Li,Q., Bian,F., Zhu,J., Yan,T., Wang,X., Zhang,Y., et al. (2023) RNA-Guided DNA Transposition in Corynebacterium glutamicum and Bacillus subtilis. ACS Synth. Biol., 12, 2198–2202. 8. Walker, M.W.G., Klompe, S.E., Zhang, D.J. and Sternberg, S.H. (2023) Novel molecular requirements for CRISPR RNA-guided transposition. Nucleic Acids Res., 51, 4519–4535. 9. Wetmore,K.M., Price,M.N., Waters,R.J., Lamson,J.S., He,J., Hoover,C.A., Blow,M.J., Bristow,J., Butland,G., Arkin,A.P., et al. (2015) Rapid quantification of mutant fitness in diverse bacteria by sequencing randomly bar-coded transposons. MBio, 6, e00306–15. 10. Baba,T., Ara,T., Hasegawa,M., Takai,Y., Okumura,Y., Baba,M., Datsenko,K.A., Tomita,M., Wanner,B.L. and Mori,H. (2006) Construction of Escherichia coli K-12 in-frame, single-gene knockout mutants: the Keio collection. Mol. Syst. Biol., 2, 2006.0008. 11. Egbert, R.G., Rishi, H.S., Adler, B.A., McCormick, D.M., Toro, E., Gill, R.T. and Arkin, A.P. (2019) A versatile platform strain for high-fidelity multiplex genome editing. Nucleic Acids Res., 47, 3244– 3256. 12. Ruegg,T.L., Pereira,J.H., Chen,J.C., DeGiovanni,A., Novichkov,P., Mutalik,V.K., Tomaleri,G.P., Singer,S.W., Hillson,N.J., Simmons,B.A., et al. (2018) Jungle Express is a versatile repressor system for tight transcriptional control. Nat. Commun., 9, 3617. 13. De Coster, W., D’Hert, S., Schultz, D.T., Cruts, M. and Van Broeckhoven, C. (2018) NanoPack: visualizing and processing long-read sequencing data. Bioinformatics, 34, 2666–2669. 14. Li, H. (2018) Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34, 3094– 3100. 15. Vo, P.L.H., Ronda, C., Klompe, S.E., Chen, E.E., Acree, C., Wang, H.H. and Sternberg, S.H. (2021) CRISPR RNA-guided integrases for high-efficiency, multiplexed bacterial genome engineering. Nat. Biotechnol., 39, 480–489. 16. Amundsen, S.K., Taylor, A.F. and Smith, G.R. (2000) The RecD subunit of the Escherichia coli RecBCD enzyme inhibits RecA loading, homologous recombination, and DNA repair. Proc. Natl. Acad. Sci. U. S. A., 97, 7399–7404. 17. Vo, P.L.H., Acree, C., Smith, M.L. and Sternberg, S.H. (2021) Unbiased profiling of CRISPR RNA- guided transposition products by long-read sequencing. Mob. DNA, 12, 13.Atty. Dkt: BERK-543WO urphy, K.C. (2007) The lambda Gam protein inhibits RecBCD binding to dsDNA ends. J. Mol. Biol., 371, 19–24. osberg, J.A., Lajoie, M.J. and Church, G.M. (2010) Lambda red recombineering in Escherichia coli occurs through a fully single-stranded intermediate. Genetics, 186, 791–799. eimer, A., Kohlstedt, M., Volke, D.C., Nikel, P.I. and Wittmann, C. (2020) Industrial biotechnology of Pseudomonas putida: advances and prospects. Appl. Microbiol. Biotechnol., 104, 7745–7766.u, Z., Li, S., Li, Y., Jiang, Z., Zhou, J. and An, Q. (2018) Complete genome sequence of N2-fixing model strain Klebsiella sp. nov. M5al, which produces plant cell wall-degrading enzymes and siderophores. Biotechnol Rep (Amst), 17, 6–9. rah, I., Nukui, Y., Yamaoka, S. and Saito, R. (2022) Emergence of a High-Risk Klebsiella michiganensis Clone Disseminating Carbapenemase Genes. Front. Microbiol., 13, 880248.haurasiya, K.R., Ruslie, C., Silva, M.C., Voortman, L., Nevin, P., Lone, S., Beuning, P.J. and Williams, M.C. (2013) Polymerase manager protein UmuD directly regulates Escherichia coli DNA polymerase III α binding to ssDNA. Nucleic Acids Res., 41, 8959–8968. hadtare, S. and Inouye, M. (1999) Sequence-selective interactions with RNA by CspB, CspC and CspE, members of the CspA family of Escherichia coli. Mol. Microbiol., 33, 1004–1014.illingham, M.S. and Kowalczykowski, S.C. (2008) RecBCD enzyme and the repair of double- stranded DNA breaks. Microbiol. Mol. Biol. Rev., 72, 642–71, Table of Contents. hoi, W., Jang, S. and Harshey, R.M. (2014) Mu transpososome and RecBCD nuclease collaborate in the repair of simple Mu insertions. Proc. Natl. Acad. Sci. U. S. A., 111, 14112–14117. ang, S., Sandler, S.J. and Harshey, R.M. (2012) Mu insertions are repaired by the double-strand break repair pathway of Escherichia coli. PLoS Genet., 8, e1002642. lompe, S.E., Vo, P.L.H., Halpin-Healy, T.S. and Sternberg, S.H. (2019) Transposon-encoded CRISPR–Cas systems direct RNA-guided DNA integration. Nature, 571, 219–225. ei, J. and Li, Y. (2023) CRISPR-based gene editing technology and its application in microbial engineering. Engineering Microbiology, 3, 100101. oberts, A., Nethery, M.A. and Barrangou, R. (2022) Functional characterization of diverse type I-F CRISPR-associated transposons. Nucleic Acids Res., 50, 11670–11681. els, U., Gevaert, K. and Van Damme, P. (2020) Bacterial Genetic Engineering by Means of Recombineering for Reverse Genetics. Front. Microbiol., 11, 548410. hang,Y., Sun,X., Wang,Q., Xu,J., Dong,F., Yang,S., Yang,J., Zhang,Z., Qian,Y., Chen,J., et al. (2020) Multicopy Chromosomal Integration Using CRISPR-Associated Transposases. ACS Synth. Biol., 9, 1998–2008. ang, S., Zhang, Y., Xu, J., Zhang, J., Zhang, J., Yang, J., Jiang, Y. and Yang, S. (2021) Orthogonal CRISPR-associated transposases for parallel and multiplexed chromosomal integration. Nucleic Acids Res., 49, 10192–10202. anta, A.B., Myers, K.S., Ward, R.D., Cuellar, R.A., Place, M., Freeh, C.C., Bacon, E.E. and Peters, J.M. (2024) A Targeted Genome-scale Overexpression Platform for Proteobacteria. bioRxiv, 10.1101 / 2024.03.01.582922. u, J., Sun, Y., Wu, J., Yang, S. and Yang, L. (2024) Chromosome recombination and modification by LoxP-mediated evolution in Vibrio natriegens using CRISPR-associated transposases. Biotechnol. Bioeng., 121, 1163–1172.Atty. Dkt: BERK-543WO 36. Garza Elizondo, A.M. and Chappell, J. (2024) Targeted Transcriptional Activation Using a CRISPR- Associated Transposon System. ACS Synth. Biol., 13, 328–336. Example 2 Methods

[0216] We examined the phylogenetic distribution of 11 genes across 92 bacterial phyla (see Figure 13), which included the nine identified and validated by the RB-TnSeq screen plus ihfA and recA. To simplify viewing, only phyla with more than 10 members in the AnnoTree database (v beta, based on GTDB R214.0), were included in this analysis. Briefly, the genes in the AnnoTree database were annotated using PFAM v27.0, TIGRFAM v15.0, and KEGG orthology identifiers assigned to the UniRef100 database. The KEGG or PFAM annotation assigned to individual genes was used to extract homologs from 80789 representative genomes from GTDB R214.0. Genes such as ihfB, recD, and umuD are found in close association with their respective enzymatic complex subunits, ihfA, recBC, and umuC. To maintain consistency, the homologs of these co-occurring genes were extracted using the same methods as for the genes of primary interest, and only genomes containing both sets of genes in the same contig were retained for further analysis. The gene yneK lacks a PFAM or KEGG annotation. To address this, protein sequences from the 80,789 representative genomes in the GTDB database were downloaded and subjected to BLASTp analysis against the E. coli yneK sequence. Proteins with greater than 50% sequence similarity and a significant FASTA hit (E ≤ 1e-5) were retained. To further validate our findings, a BLASTp search was conducted against a UniProt database subset containing sequences sharing at least 90% similarity to E. coli yneK. Both approaches yielded comparable results. While the yneK gene is prevalent among closely related E. coli and Shigella strains, only a limited number of these strains are present in the 80789 representative genomes in the GTDB database. The cspC gene is a member of the bacterial Cold Shock Protein (CSP) family, which can contain multiple paralogs with high sequence similarity (>70%;. To accurately differentiate these paralogs, cspC proteins identified by AnnoTree were re-analyzed using eggNOG-mapper v5.0. This tool leverages evolutionary relationships to distinguish between proteins within the CSP family excluding non-cspC family members. Finally, gene presence per phyla was visualized in iTOL v6 with the publicly available GTDB R214.0 bacterial tree inferred from the concatenation of 120 proteins. Results

[0217] To understand the broader relevance of our findings, we investigated the phylogenetic distribution of VchCAST activators, inhibitors and associated genes across the bacterial tree of life. We analyzed the presence of our seven validated RB-TnSeq activators, two inhibitors, as well as ihfA and recA (shown to be a strong activator in this study) across all bacterial phyla with more than 10 members in the AnnoTree database. Our analysis revealed diverse conservation patterns among these genes. RecA has near-universal presence across the microbial tree of life and was observed in over 90% of species in most phyla. In contrast, other genes identified fromAtty. Dkt: BERK-543WO the screen such as yneK, hha, ynbE, and chaB are absent across most phyla beyond E. coli containing Pseudomonadota and present only in small portions of the phyla where they exist. This landscape of VchCAST activators and inhibitors should provide insight into where VchCAST may be used effectively and what modifications may be necessary to make it function in organisms where it currently does not.

[0218] Beyond Pseudomonadota, where type I-F CASTs are largely confined, many of the activators identified in this study have limited presence (Figure 13). Notably, IHF, which is essential for efficient VchCAST transposition is absent from the majority of the bacterial tree of life. Given IHF’s importance in E. coli, it may be integral for targeting strains without endogenous IHF. For instance, members of Bascillota and Actinomycetota, where VchCAST has shown to have low efficiency, lack close sequence homologs of ihfA and ihfB. While Actinomycetota does contain sequence-divergent but functionally homologous IHF-like proteins, their ability to support transposition is unknown. The addition of ihfA and ihfB to VchCAST editing vectors in these phyla, as well as other phyla where their homologs are absent, may increase editing efficiency. These considerations highlight the relevance of tailoring optimization strategies to the specific genetic characteristics of the target organism when applying VchCAST systems more broadly. Example 3

[0219] This example includes some information from the above examples, but also includes additional information.

[0220] CRISPR-Associated Transposases (CASTs) hold tremendous potential for microbial genome editing due to their ability to integrate large DNA cargos in a programmable and site- specific manner. However, the widespread application of CASTs has been hindered by the unknown role of host-associated molecular requirements that facilitate editing with CASTs. In an effort to address this gap in knowledge, we conducted the first genome-wide screen for host factors impacting Vibrio cholerae CAST (VchCAST) activity and used the findings to increase VchCAST editing efficiency in E. coli and beyond. A genome-wide loss-of-function mutant library in E. coli was screened to identify 15 genes that impact VchCAST transposition. Of these, seven factors were validated to improve VchCAST activity, and two were found to be inhibitory. Informed by homologous recombination involved effectors, RecD and RecA, we tested the λ- Red recombineering system in our VchCAST editing vectors, which increased its insertion- mediated editing efficiency by 55.2-fold in E. coli, while maintaining high target specificity and similar insertion arrangements. Furthermore, λ-Red-enhanced VchCAST achieved increased editing efficiency in the industrially important bacteria Pseudomonas putida and the emerging pathogen Klebsiella michiganensis. This study improves understanding of factors impacting VchCAST activity and enhances its efficiency as a bacterial genome editor.Atty. Dkt: BERK-543WO INTRODUCTION

[0221] CRISPR-Associated Transposases (CASTs) represent a powerful new addition to the genome editing toolbox. Traditional CRISPR-Cas systems create double-stranded breaks to facilitate editing, which are often lethal in bacteria. In contrast, CASTs combine RNA-guided targeting with Tn7-like transposition to enable large programmable insertions without this limitation. Unlike homologous recombination, a standard alternative for microbial editing that requires lengthy homology arms, CASTs are programmed with simple ~30 bp guides that can be multiplexed for simultaneous edits. Among the currently discovered and tested CASTs— including those with a single Cas component (Type V) and multiple Cas components (Type I)— the Type I systems are significantly more accurate in E. coli (4–6). Within this group, the Vibrio cholerae type I-F CAST Tn6677 (VchCAST) (Figure 19A) has emerged as a valuable platform for bacterial editing (7), manipulating microbial communities (6), and even modifying human genomes (8).

[0222] Despite its established value as a genome editing tool, VchCAST currently faces limitations in achieving efficient editing across a broad range of organisms. For example, in the industrially important species Corynebacterium glutamicum, VchCAST achieved an editing efficiency between 0-0.027%. In complex gut microbiome samples, VchCAST has shown editing efficiencies as low as 0.001%, significantly limiting its utility for in situ microbiome engineering. These low efficiencies restrict VchCAST's potential for environmental, industrial, and therapeutic genome editing applications, highlighting the need for optimization.

[0223] Broader and more efficient use of VchCAST would be facilitated by an understanding of the host factors that enhance and restrict its function. For instance, the addition of ClpX was found to facilitate VchCAST functionality in human HEK293T cells (8); however, ClpX has not been found necessary for its function in bacteria. Within bacteria, integration host factor (IHF) and, to a lesser degree, factor for inversion stimulation (FIS) are the only proteins known to affect VchCAST function (10). Additionally, CAST editing efficiency rates differ between endogenous hosts and E. coli recipients, suggesting more, yet still unidentified, molecular requirements of CAST integration. A significant opportunity remains for systematic identification of activators and inhibitors for VchCAST transposition.

[0224] In this study, we conducted a genome-wide screen to identify host factors that impact VchCAST activity, resulting in the validation of seven activators and two inhibitors. Informed by the identification of homologous recombination-involved effectors, RecD and RecA, we explored whether incorporating the efficient phage-derived λ-Red recombination system could enhance VchCAST editing performance (Figure 19B). By integrating the λ-Red genes into our VchCAST vectors, we achieved a 55.2-fold increase in editing efficiency. This strategy also facilitated more efficient integration in other industrially, environmentally, and medically relevant bacteria. Our research advances the understanding of CAST systems, revealing key regulatory factors and developing strategies to enhance editing efficiency.Atty. Dkt: BERK-543WO RESULTS A whole genome mutant screen and validation to identify VchCAST regulators

[0225] We performed a genome-wide survey of genes involved with VchCAST function by screening a library of loss-of-function mutants for the ability of VchCAST to integrate into a safe site. The screen was conducted with a preexisting E. coli RB-TnSeq transposon mutant library. VchCAST was introduced to the library via conjugation, and the inserted cargo was selected for in liquid media overnight, ensuring all members of the library contained both the VchCAST and the background loss-of-function mutation (Figure 14A). The efficiency of VchCAST insertion into each mutant was determined by sequencing the unique barcode of each insertion-containing mutant after outgrowth and comparing its abundance to that of a mariner transposase control (See Materials and Methods). Mariner was chosen as an unrelated transposon system to control for screening hits that affect delivery efficiency and variables beyond VchCAST integration itself.

[0226] We identified 11 genes as putative activators and four as putative inhibitors, based on consistent absolute gene fitness > 1 across the two parallel screens (Figure 14B; Table 2). We report the raw fitness scores and compare the distribution of log2 fold change for VchCAST vs. the mariner control across both screens (Figure 20 & Figure 21) to inform gene selection for further validation. Notably, the beta subunit of integration host factor (ihfB) emerged as the top activator from the screen (Figure 14B). The other subunit, ihfA, did not have a high enough abundance in the T=0 starting library to be considered in the analysis pipeline. This finding supports previous research identifying IHF as a strong activator of VchCAST insertion (10), lending credibility to the screen’s results.

[0227] We validated hits with the largest absolute fitness values from the pooled RB-TnSeq screens by testing VchCAST editing in clonal gene deletion mutants. To prioritize candidates most likely to have direct interaction with VchCAST, we filtered for those possessing DNA- interacting, RNA-interacting, and protein-interacting functions as well as hypothetical proteins, as per the Gene Ontology (GO) database. We then performed conjugation assays to quantify the relative editing efficiency of VchCAST in Keio deletion mutants of each of the nine filtered hits (Figure 15A). Editing efficiencies were compared to ΔyicI, a deletion mutant with no documented adverse fitness effects (13). Mutants of putative activators (ihfB, cspC, ynbE, umuD, chaB) exhibited significantly reduced VchCAST editing efficiency (64.6 ± 19.4% to 99.9 ± 0.1%) compared to the ∆yicI neutral fitness control (Figure 15A). Knockouts of putative inhibitors (hha, recD) exhibited a significant increase in editing efficiency (4.0 ± 1.6-fold and 21.8 ± 7.2-fold) in comparison to the control (Figure 15A). The results from the whole genome screen and the individual mutant validation were largely consistent, with recD as the strongest inhibitor and ihfB as the strongest activator in both experiments. Taken together, the validation results support the potential involvement of these host factors in VchCAST’s genomic integration.Atty. Dkt: BERK-543WO

[0228] We were particularly compelled by the strong inhibitor classification of recD in VchCAST insertion revealed by the screen and validation, and sought to explore the rec genes further. RecD inhibits RecBCD-mediated homologous recombination, suggesting a beneficial role of factors involved in DNA repair during VchCAST integration. To test this hypothesized role, we delivered VchCAST via conjugation into Keio knock-out mutants of select, key rec genes (recABCD) and quantified editing efficiency normalized to ΔyicI (Figure 15B). Consistent with the screen results, the putative inhibitor ΔrecD showed a 9.8 ± 2.6-fold increase in editing efficiency compared to the neutral fitness mutant control. Knockouts of recB and recC slightly decreased insertion efficiency relative to the control (45.3 ± 25.5% and 15.0 ± 6.6%, respectively). ΔrecA, which did not have a high enough abundance in the T=0 starting library to be considered in the RB-TnSeq screen, had no viable transconjugant colonies above the limit of detection. The results of this experiment identify recA as an activator and, in combination with the identification of recD as an inhibitor, support the role of RecBCD-mediated homologous recombination in promoting VchCAST integration.

[0229] Table 2.11 genes were identified as putative activators and 4 genes were identified as putative inhibitors, based on consistent absolute gene fitness > 1 across the two parallel screens. Gene RB-TnSeq RB-TnSeq library libraryFurther evidence suggesting homologous recombination’s role in VchCAST integrationAtty. Dkt: BERK-543WO

[0230] Informed by RecA and RecD's Keio validation results, we sought to further support homologous recombination's involvement in VchCAST integration by assessing VchCAST’s editing efficiency in cells exposed to ciprofloxacin (CIP). This fluoroquinolone antibiotic has been shown to stimulate homologous recombination in E. coli at sublethal concentrations by inducing DNA damage and consequently the SOS response (15–17). We hypothesized that stimulating homologous recombination during VchCAST transposition would increase the efficiency of transposition. BW25113 E. coli recipients incubated with MIC CIP before conjugation showed a 58.0 ± 16.9-fold increase in VchCAST editing efficiency compared to the no CIP control (Figure 22). Similarly, BW25113 recipients incubated with one-half MIC CIP saw a 25.6 ± 11.9-fold increase in VchCAST editing efficiency compared to the control (Figure 22). These results provide further evidence in support of the positive correlation between homologous recombination frequency and VchCAST integration. Leveraging λ-Red to improve VchCAST editing efficiency in E. coli

[0231] Considering the role of the RecBCD complex and potentially homologous recombination more broadly in VchCAST integration, we hypothesized that introducing higher-efficiency recombination machinery may improve editing outcomes. To this end, the bacteriophage λ-Red genes (exo, beta, and gam) were cloned into a VchCAST plasmid (R6K, PPmtl-catP cargo) (Figure 16A). We optimized the induction level of λ-Red by crystal violet (Figure 23) for editing in E. coli. The induced λ-Red VchCAST treatment significantly increased editing efficiency by 55.2 ± 6.9 fold in BW25113 E. coli compared to the VchCAST control (Figure 16B).

[0232] To determine which components λ-Red are responsible for the increase in VchCAST efficiency, we also tested constructs containing inducible exo only, beta only, gam only, exo- beta, for effect on VchCAST insertion. Exo and Beta promote homologous recombination through their exonuclease and single-stranded binding activity respectively, while Gam inhibits RecBCD by binding to and blocking the DNA-interacting domain of RecB (18–20). We found that while the induced exo-beta construct does improve editing efficiency relative to the VchCAST control by 5.6 ± 2.0-fold in E. coli, it does not improve editing to the extent of all three λ-Red genes (55.2 ± 6.9-fold) (Figure 16B). The presence of exo, beta, or gam alone on the VchCAST vector did not affect editing efficiency (Figure 16B). These results suggest that the complete system is required to realize the full impact of λ-Red on editing efficiency in VchCAST.

[0233] Whole genome sequencing analysis confirmed that 100% of λ-Red VchCAST insertions were on-target in E. coli, matching the on-target efficiency of the VchCAST vector alone (Figure 16C). Furthermore, both vectors shared an equal proportion of reads that inserted at the expected ~49 bp downstream of the PAM in E. coli (Figure 16D). Analysis of transconjugants for cointegrates, where the duplicated transposon and the entire plasmid backbone are integrated, were performed by cPCR (Figure 24A-C) and whole genome sequencing (Figure 24D-E). We consistently observed equivalent cointegration rates between λ-Red VchCAST and VchCASTAtty. Dkt: BERK-543WO around 12% across both treatments and analysis methods (Figure 24C-E). The vast majority of the screened simple insert clones were observed in the T-RL orientation for both λ-Red VchCAST and VchCAST (Figure 24C & S6E). These data support λ-Red’s positive impact on VchCAST editing efficiency without affecting insertion outcomes. λ-Red-assisted VchCAST editing in additional gram-negative bacteria

[0234] Motivated by the successful implementation of λ-Red to increase VchCAST-directed editing efficiency in E. coli, we evaluated the impact of λ-Red on VchCAST editing in gram- negative bacteria with lower baseline editing efficiencies. We selected P. putida KT2440, a relevant strain for industrial biotechnology and bioremediation (21), and K. michiganensis M5a1, a plant-associated nitrogen-fixing strain and close relative to an important human pathogen (22, 23). VchCAST and λ-Red VchCAST vectors delivering a kanamycin resistance cargo were queried for their effect on insertion efficiency (Figur 25A-D).

[0235] In both species, the presence of λ-Red on the VchCAST vector yielded an increased insertion efficiency with comparable accuracy to VchCAST alone (Figure 17). We tested a range of crystal violet induction concentrations in P. putida and K. michiganensis and selected 1 µM, and 0.5 µM, respectively (Figure 25E-F). There was no statistical difference between the induced and uninduced VchCAST controls in both species (Figure 17A, 4D). In P. putida, we observed a significant increase in editing efficiency with and without the crystal violet inducer (3.1 ± 1.2 fold and 5.6 ± 2.3 fold, respectively) compared to the uninduced VchCAST control (1) (one- tailed, one-sample t-test). Whole genome sequencing confirmed that the edits in P. putida with both the VchCAST and λ-Red VchCAST were on target (Figure 17B) and followed similar insertion distribution patterns as expected (Figure 17C). In K. michiganensis, we observed a significant but variable increase in editing efficiency (10.8 ± 8.4 fold) in the presence of the crystal violet inducer (Figure 17D). WGS determined 100% on-target editing efficiency (Figure 17B, 4E) and comparable relative insertion distributions for both vectors tested (Figure 17C, 4F) for both target species. Altogether, these results suggest the potential of using λ-Red VchCAST to enhance on-target VchCAST editing efficiency for precise genome engineering in diverse gram-negative bacteria. Distribution of VchCAST inhibitors and activators across diverse bacteria

[0236] To understand the broader relevance of our findings, we investigated the phylogenetic distribution of VchCAST activator and inhibitor genes across the bacterial tree of life. We analyzed the presence of our seven screen-identified VchCAST activators, two inhibitors, as well as ihfA (10) and recA (shown to be a strong activator in this study) across all bacterial phyla with more than 10 members in the AnnoTree database. Our analysis revealed diverse conservation patterns among these genes (Figure 5). RecA has near-universal presence across the microbial tree of life and was observed in over 90% of species in most phyla (Figure 5). In contrast, other genes identified from the screen, such as yneK, hha, ynbE, and chaB are absent across most phylaAtty. Dkt: BERK-543WO beyond the E. coli-containing Pseudomonadota and occur only sporadically within the few phyla where they are found (Figure 5). Type I-F CAST systems are almost exclusively found in Pseudomonadota (24, 25), representing just a single branch on this phylum-level tree. Notably, key regulatory genes such as ihfA, ihfB, recA, and recD occur in over 90% of these type I-F CAST-containing genomes (Figure 26). Our analysis aims to serve as a guide for identifying where these regulators might enhance editing efficiency, particularly in taxonomic branches that naturally lack type I-F systems. DISCUSSION

[0237] In this study, we conducted a genome-wide mutant screen to identify genes in E. coli that influence the efficiency of VchCAST, a promising tool for precise DNA insertion of large cargos in bacteria and eukaryotes. We screened for and then validated nine candidate genes, with eight being novel to this study, that either positively or negatively affect VchCAST insertion activity when disrupted (Figure 14B, Figure 15A). Notably, our results highlighted a role for the RecBCD complex in CAST integration, with the disruption of recD increasing editing efficiency, while a recA deletion strongly decreased it (Figure 15B). Building on these insights, we leveraged the bacteriophage λ-Red genes (exo, beta, and gam) as an alternative recombination system to enhance CAST insertion efficiency (Figure 16A). By optimizing the expression of these genes, we achieved improved editing efficiency not only in E. coli (Figure 16B), but also in P. putida (Figure 17A), an industrial model strain (21), and K. michiganensis (Figure 17D), a plant-associated nitrogen-fixing strain and close pathogen relative (23). This work provides a comprehensive survey of host factors influencing VchCAST integration and presents an approach to enhance its efficiency across various bacterial species.

[0238] Beyond the homologous recombination-associated RecD, we discovered eight other activators and inhibitors of the CAST complex that may further contribute to understanding VchCAST transposition (Figure 15A). Putative activator UmuD, in complex with UmuC, performs translesion synthesis following activation by DNA damage. In the context of VchCAST integration, where DNA damage signals are absent, intact UmuD would instead be expected to inhibit translesion synthesis, potentially preventing premature replication fork collision and allowing RecBCD sufficient time to repair the insertion intermediate through homologous recombination. Putative activator cold shock protein C (CspC), known to bind to single-stranded nucleotides, may stabilize the RNA guide or the single-stranded gap of the post-strand transfer intermediate. Interestingly, ClpX, a sequence-specific AAA+ ATPase with protein unfolding capabilities that facilitated more efficient editing in human cells, was not identified in our screen. It is possible that in E. coli, one or more partially redundant proteolytic systems such as ClpAP, ClpYQ, and the Lon protease can perform ClpX's activity in CAST integration. Although the human proteasome provides analogous protein degradation functions, the absence of close homologs to these bacterial proteases may explain why ClpX addition specifically enhancesAtty. Dkt: BERK-543WO VchCAST function in human cells. The new activators and inhibitors found in this study provide valuable targets for further experiments to understand and control the function of VchCAST.

[0239] Among the strong activators and inhibitors identified and validated within this study, we were particularly interested in the role of the RecBCD complex. RecD was identified as the strongest inhibitor of VchCAST integration via the screen and Keio mutant validation (Figure 14B, Figure 15A). Notably, this result is consistent with observations made for Tn7 transposons, the native system from which all type I-F CASTs originally derived from, where activity is increased in RecD- E. coli. RecD functions as an inhibitor to homologous recombination by blocking RecA loading. Therefore, it was unsurprising that RecA, a single-stranded binding protein essential for Rec-mediated double-strand break repair, showed an activating phenotype when single mutants were tested for VchCAST integration efficiency (Figure 15B). Based on this low editing efficiency in the absence of recA, we would expect decreased efficacy of VchCAST in the ∆recA genotypes of many commercial E. coli strains. The weaker phenotypes of RecBC, may be explained by some amount of functional redundancy in RecBCD-mediated repair compared to other repair pathways, such as the RecF homologous recombination pathway. In the Mu transposon, the RecBCD complex facilitates the repair of a double-strand break that occurs during resolution of the insertion. Similarly, we hypothesize that the highly stable VchCAST post-transposition complex could stall replication forks and introduce a double-strand break at the target site, similar to Mu’s mechanism. RecBCD and RecA would then repair the double- strand break. This putative function is supported by our finding that CIP, the DNA damaging and SOS response / homologous recombination-inducing agent, and λ-Red, the highly efficient homologous recombination system, improve the efficiency of VchCAST. Such a homologous recombination-dependent mechanism for VchCAST would contrast with the commonly stated assumption that transposition occurs independently of this DNA repair pathway.

[0240] To examine the role of RecBCD and other effector proteins in VchCAST function, we delivered VchCAST via suicide vector. This approach leads to lower insertion efficiencies but provides an opportunity to examine the effects of activating and inhibiting proteins, while ensuring that all selected transconjugants are bonafide edited cells. Furthermore, a suicide vector is advantageous for eventual microbiome editing applications since it is eliminated after cargo delivery, thus presenting fewer biocontainment concerns. The efficiencies we report make up a small portion of the total cells but are consistent with those seen with homologous recombination and transposon insertion with transient presence of the editor. The highest efficiencies reported in the literature are achieved in studies that deliver the VchCAST on a replicating plasmid which is then selected for, where insertions can be introduced to nearly 100% of cells. This is likely due to the presence of multiple copies of the VchCAST vector persisting in cells for an extended time. Given that λ-Red has often been used to increase homologous recombination efficiency on a replicating vector, we expect that our λ-Red VchCAST system also works in that context.

[0241] Our findings on improving VchCAST efficiency in E. coli prompted us to investigate the broader applicability of this approach in other bacterial species. When the same λ-Red VchCASTAtty. Dkt: BERK-543WO system was tested in Pseudomonadota outside of E. coli, we found that it enhanced editing efficiencies, but to a lesser extent and with more variation (Figure 17A, 4D). In P. putida, we observed increased editing efficiency with uninduced λ-Red VchCAST (Figure 17A). These results suggest that leaky expression of λ-Red is sufficient to significantly increase editing efficiency, which could be due to an incompatible binding site for the repressor, EilR. Additionally, we chose the crystal violet concentration for P. putida based on a previous characterization of pJEx in this strain (42) and P. putida’s ability to grow on crystal violet as a sole carbon source (43). Incorporating analogous recombination machinery, such as the rac prophage-derived recombination system recET, could be a viable approach to further increase editing efficiency in genomes with higher GC content, such as P. putida (44). In K. michiganensis, we observed tight regulation of λ-Red via pJEx, but high day-to-day variability in the effect of λ-Red VchCAST (Figure 17D), suggesting the potential for compensatory mutagenesis of λ-Red when highly expressed, which was observed in early λ-Red constructs driven by TetR. These data suggest that insights from this screen will be most efficiently applied in a host-dependent manner. For example, commandeering the recombination machinery from a phage infecting the bacteria of interest and integrating it with the VchCAST vector would likely allow for more effective editing.

[0242] Beyond Pseudomonadota, where type I-F CASTs are largely confined, many of the activators identified in this study have limited presence (Figure 5). Notably, IHF, which is essential for efficient VchCAST transposition is absent from the majority of the bacterial tree of life. Given IHF’s importance in E. coli, it may be integral for targeting strains without endogenous IHF. For instance, members of Bacillota and Actinomycetota, where VchCAST has shown to have low efficiency, lack close sequence homologs of IhfA and IhfB. While Actinomycetota does contain sequence-divergent but functionally homologous IHF-like proteins (45, 46), their ability to support transposition is unknown. The addition of ihfA and ihfB to VchCAST editing vectors in these phyla, as well as other phyla where their homologs are absent, may increase editing efficiency. These considerations highlight the relevance of tailoring optimization strategies to the specific genetic characteristics of the target organism when applying VchCAST systems more broadly.

[0243] Beyond type I-F systems like VchCAST, type V CAST systems show promise due to their smaller size. However, they have demonstrated significantly lower on-target specificity. In their native Cyanobacterial hosts, crRNA matches with surrounding target sequences suggest higher accuracy. This indicates potential host-associated genes in these organisms that could enhance efficiency when the systems are applied to distant taxa.

[0244] We expect this work to bring the powerful CAST editing toolset to more biological systems. VchCAST has mostly been applied in Gammaproteobacteria. This work provides strategies for more efficient delivery so that the function of these communities can be better probed and controlled at a genetic level. Overall, our findings provide insights into the factorsAtty. Dkt: BERK-543WO influencing CAST efficiency and offer strategies for improving its performance across diverse organisms, paving the way for broader applications in genome editing. MATERIALS AND METHODS RB-TnSeq screen for E. coli regulators of CAST integration

[0245] To systematically identify E. coli genes that contribute to VchCAST-mediated integration, we conducted a genome-wide loss-of-function mutant screen and selected for VchCAST edits after conjugation in the KEIO_ML9 E. coli RB-TnSeq library (12, 53). This approach allowed us to generate and select for VchCAST edits across a diverse array of insertion mutants. The library comprises 152,018 uniquely barcoded single-gene transposon insertion mutants, covering 3,728 nonessential protein-coding genes out of a total of 4,146 in the E. coli genome (12). To account for biases introduced by the antibiotic resistance marker used, two different antibiotic selection cargos (PPmtl-gmR, PPmtl-catP) were utilized in parallel for the RB- TnSeq and conjugation efficiency assays, with a mariner transposase system serving as a control.

[0246] The KEIO_ML9 RB-TnSeq library was inoculated into LB containing kanamycin (25 µg / mL) and grown at 37ºC with shaking. Donor strains harboring the VchCAST and mariner vectors respectively in E. coli WM3064 (54) (DAP auxotroph, pir+, RP4+) were grown at 37ºC with shaking in LB containing diaminopimelic acid (DAP; 0.3 mM) and gentamicin (screen 1; 50 µg / mL) or chloramphenicol (screen 2; 34 µg / mL). Three 10 OD*mL samples of the library overnight were washed and pelleted before being frozen at -80ºC as three technical replicate time zero (T=0) samples. After washing and resuspending in LB containing DAP, 1 OD*mL of donor was combined with 1 OD*mL of recipient in separate 1.5 mL Eppendorf tubes. Each combined donor-recipient sample was plated onto a plain LB agar petri plate topped with a MF-Millipore™ Membrane Filter for conjugation. After 6 hours of conjugation at 30ºC, cells from two conjugation plates were scraped into 20 mL of LB as a single technical replicate (six conjugation plates for three technical replicates for each donor-recipient combination).100 uL of the resuspended cells were set aside for 10-fold serial dilution and spot plating. One mL of the resuspended cells from each technical replicate was used to inoculate 100 mL of LB containing kanamycin and gentamicin or chloramphenicol. The three technical replicates of the library not introduced to any donor were inoculated into nonselective liquid cultures of LB with kanamycin. The liquid cultures were grown at 30ºC for 12 hours with shaking. Cells from liquid cultures of each technical replicate were pelleted and resuspended in 40 mL of LB. A two mL aliquot of the resuspended cells was taken for genomic DNA extraction via the QIAGEN DNeasy PowerSoil Pro Kit. Barcodes were amplified with the BarSeq_v3 primers following the BarSeq PCR protocol (12). Amplicons were pooled into a synthetic amplicon library and submitted for sequencing by Illumina NovaSeq PE150.

[0247] BarSeq sequencing reads were processed using the FEBA pipeline described previously (12). In this pipeline, the final gene fitness value was calculated by averaging the fitness scores of all strains with independent transposon insertions located in the central region of the gene,Atty. Dkt: BERK-543WO excluding those near the beginning or end. The pipeline-generated fitness scores (fit_logratios_good) were further analyzed using R (version 4.4.0) with the tidyverse package (version 2.0.0). Within an experimental trial, treatment replicates were averaged and then the difference between the VchCAST fitness scores and the mariner fitness scores was calculated. To identify genes consistently displaying strong differential effects across the gmR and catP insertion screens, we focused on those with an absolute fitness difference greater than one in both experimental trials. Positive fitness values generated by the FEBA pipeline indicate a putative inhibitor of VchCAST, as mutating the gene increases the fitness of the strain. Conversely, negative fitness values indicate putative activators of VchCAST function. Validating activators and inhibitors with Keio E. coli mutants

[0248] We next validated the involvement of putative host factors identified in the RB-TnSeq screen by performing VchCAST editing in E. coli Keio collection deletion mutants corresponding to the genes with the largest fitness values (absolute gene fitness > 1) and possessing molecular functions of interest as detailed on the Gene Ontology (GO) database (55). Factors that do not possess molecular functions of interest, particularly membrane-associated factors, or those that are absent from the Keio collection were not validated or pursued in further experiments. To confirm the Keio strains, primers were designed to amplify the kanamycin resistance gene insertion within each target locus of the expected Keio mutants (56). The ΔyicI mutant was selected as the negative control as it is a context-neutral genomic locus with no documented adverse fitness effects (13).

[0249] To assess how the presence or absence of individual E. coli genes affect VchCAST integration into the target genome, we implemented a conjugation-based editing efficiency assay (7) (Figure 19B). Briefly, donor and recipient strains were grown for 16 hours, washed, resuspended in LB containing DAP, combined in a 1:1 ratio (0.1 OD*mL each), spotted on LB agar within a 24 well block, and allowed to conjugate for 6 hours at 30ºC. Afterward, spots were resuspended in one mL LB media.10-fold serial dilutions were performed with resuspended cells, spotted onto LB agar and LB agar with antibiotics, and incubated overnight at 30ºC. Serial dilution spot plates were left to grow until individual colonies formed. Finally, the plates were imaged and cell colonies were counted to compute editing efficiency (6, 7), with transconjugants reflecting successful conjugation and insertion of selective cargo. Fold change in editing efficiency was computed by normalizing the editing efficiency of candidate regulator hits of interest to that of the ΔyicI control. Due to day-to-day variation in editing efficiencies, normalizations were performed by pairing the treatment to the control of that day. Three biological replicates were used for each experimental condition. A one-sample t-test was used to determine statistical significance compared to the ΔyicI normalized control (hypothetical value = 1), thus asking whether the treatment condition (normalized) is significantly different than the normalized control. While multiple comparisons to the control were made, the statistical approach chosen works to compare normalized data with unmatched variance. A one-tailed p-Atty. Dkt: BERK-543WO value is reported for each treatment based on hypotheses (e.g. activator, inhibitor) generated from the genome-wide mutant screen. Ciprofloxacin-stimulated homologous recombination correlates with increased VchCAST activity

[0250] We further explored the involvement of homologous recombination in VchCAST integration by assessing the editing efficiency of VchCAST in BW25113 E. coli recipients incubated with ciprofloxacin (CIP). CIP is a fluoroquinolone antibiotic reported to increase homologous recombination in E. coli by inducing DNA damage and triggering the SOS response when exposed at sublethal concentrations (15), (16), (17). We tested the minimum inhibitory concentration (MIC) of CIP in liquid culture according to NCCLS recommendations (57)and observed an MIC of 160 ng / mL and one-half MIC of 80 ng / mL, corresponding to previously described concentrations (15).

[0251] To assess how ciprofloxacin-induced increase in homologous recombination influences VchCAST integration, we performed the conjugation-based editing efficiency assay with some minor changes (Figure 19B). After 16 hours of growth, MIC (160 ng / mL) and one-half MIC (80 ng / mL) CIP were added to overnight cultures of select BW25113 recipients and then allowed to incubate for two additional hours (15). Fold change in editing efficiency was computed by normalizing the editing efficiency of CIP-exposed recipients to that of the no CIP control. Three biological replicates were used for each experimental condition. A one-sample t-test was used to determine statistical significance compared to the normalized control (1). A one-tailed p-value is reported for the experimental treatments based on the hypothesis that increasing homologous recombination would increase editing efficiency. Construction and testing λ-Red VchCAST in E. coli

[0252] We constructed versions of the VchCAST vector that included different permutations of the λ-Red recombineering system (exo, beta, gam) onto the backbone. The DNA of λ-Red and Jungle Express (pJEx) were purchased as gblock gene fragments (IDT) and added to VchCAST vectors using Gibson and Golden Gate assemblies. The sequence for λ-Red was sourced from the BioDesignER E. coli strain to maximize recombination efficiency while reducing toxicity (13). We cloned λ-Red onto a VchCAST backbone just after the origin of transfer, ensuring quick transcription in the recipient cell during conjugation. Our experimental design deliberately employs non-replicative vectors (R6K origin of replication) to maintain a system with sufficient dynamic range to detect both activating and inhibiting effects. The tightly regulated and strong, inducible Jungle Express promoter (pJEx) was used to minimize leaky λ-Red expression and toxicity of the vector. pJEx is inducible with crystal violet (CV) and transcriptional control is robust in Pseudomonadota (42). VchCAST vectors were assembled through multipart Golden Gate cloning. High-fidelity PCRs were performed with Q5 Hot Start High-Fidelity DNA polymerase (NEB). Golden Gate assembly enzymes (e.g. BsmBI-V2, BbsI, BsaI-HFV2, and T4Atty. Dkt: BERK-543WO ligase) were ordered from NEB and used with the reported buffers following previously reported protocols (6). VchCAST assemblies were electroporated into electrocompetent E. coli EC100Dpir+ cells (LGC Biosearch). Clones were screened by colony PCR (cPCR) with 2x GoTaq Green Mastermix (Promega) and plasmids were isolated with a QIAprep Spin Miniprep Kit (Qiagen). Guide assemblies were electroporated into E. coli WM3064 pir+ and grown on the appropriate antibiotics plus DAP. All vectors were confirmed with whole plasmid sequencing (Plasmidsaurus Labs).

[0253] Crystal violet induction was tested in E. coli to determine the concentration that produces the highest insertion efficiency (Figure 23). Conjugations were performed on agar plates containing the crystal violet at the optimal induction concentration (0.01 µM) for six hours before resuspension and selection. To account for variability in editing efficiency between biological replicates, induced λ-Red VchCAST treatments were paired and normalized to induced VchCAST control. Each experiment was repeated with three biological replicates. A one-sample T-test was used to determine statistical significance compared to the normalized control (1). A one-tailed p-value is reported for the experimental treatments based on the hypothesis that introducing an alternative recombination machinery would increase VchCAST editing efficiency compared to the control. Testing λ-Red VchCAST in P. putida and K. michiganensis

[0254] We first performed a quantitative assay in candidate strains to determine the frequency of mutants resistant to the antibiotics streptomycin / spectinomycin / carbenicillin (100 µg / mL and 400 µg / mL), chloramphenicol (34 µg / mL and 68 µg / mL), kanamycin (25 µg / mL, 50 µg / mL, 100 µg / mL and 200 µg / mL), and gentamicin (10 µg / mL, 20 µg / mL, and 40 µg / mL). Overnight cultures of each species were grown in LB at 30ºC, 10x serially diluted, and spotted onto LB agar without antibiotics (control) and onto each of the antibiotic concentrations. Colonies were counted after 16-40 hours of growth, and the antibiotic concentration exhibiting minimal or no detectable growth was chosen as a selection marker for genome editing experiments (Figure 25A-B). Based on these results, we constructed a VchCAST vector containing the Ppmtl promoter driving the kanamycin resistance gene for selection in P. putida and K. michiganensis.

[0255] We identified safe sites and designed guides in P. putida (GCF_000007565.2) following previously reported methods (6, 7) and used a previously tested safe site guide in K. michiganensis (6). Briefly, intergenic regions between converging genes with a distance of 300- 600 nucleotides were selected (Figure 25D). Candidate safe sites were excluded under any of the following circumstances: the region was located within or adjacent to predicted mobile genetic elements (MGEs), the region was flanked by essential genes (inspected using BioCyc), or the region contained non-coding RNA (ncRNA) features (inspected using Rfam). Within the selected safe site regions, protospacer adjacent motifs (PAMs) with the sequence 5'-CN-3' were identified. For each PAM, 32 nucleotides were added to generate a list of potential guides, which were chosen based on having a GC content of 40-60%, and ensuring that the insertion loci,Atty. Dkt: BERK-543WO approximately 49 bp downstream of the guide, remained within the intergenic region and that it would not accidentally disrupt a terminator sequence. Three guides per safe site were selected, and their off-target potential was assessed using a local BLASTn search (-dust no -word_size 4). Guides with off-target hits exhibiting the highest e-values and with minimal complementarity to off-targets in the seed region (first ~10 nt) were prioritized. Two guides targeting different safe sites were cloned and tested (Figure 25C).

[0256] Conjugation experiments in P. putida and K. michiganensis, were performed following the protocol described for E. coli with some modifications. Recipient and donor cells were grown 16 hours, at 30 ºC and 37 ºC, respectively. Conjugations were done for 20 hours at 30 ºC with crystal violet induction concentrations of 0.5 µM and 1 µM used for K. michiganensis and P. putida, respectively (Figure 25E-F). Transconjugants were selected on LB plates containing kanamycin (50 µg / mL). Spot plates of 10 µL for each serial dilution were plated to increase sensitivity and reduce technical error. Three biological replicates were performed across different days. A one-sample T-test was used to determine statistical significance compared to the normalized uninduced VchCAST control (1). A one-tailed p-value is reported for the experimental treatments based on the hypothesis that increasing homologous recombination would increase editing efficiency. Insertion analysis in E. coli, P. putida and K. michiganensis

[0257] Insertion analysis was preliminarily screened by cPCR on transconjugants following selection in all strains. Insertion orientation primers were designed to identify clones with right- left (T-RL) and left-right (T-LR) simple insertion, cointegration, and no integration outcomes (7) (Figure 24A-B). Cointegration oligos were designed to amplify the junction between the VchCAST vector backbone and the resistance marker in the cargo. Simple insert oligos were designed to amplify from the genomic DNA on either side of the insertion site to capture the full VchCAST cargo insert. Another set of oligos were designed to distinguish between T-RL versus T-LR simple insert products. Additional screening for insertion cointegration versus simple insertion was performed for E. coli by patch plating of transconjugants on LB agar plates containing 100 μg / mL carbenicillin, the VchCAST vector backbone resistance marker. Carbenicillin resistance in P. putida and K. michiganensis did not permit cointegrate screening by double selection and analysis by cPCR.

[0258] Insertion products from all recipient strains in this study (E. coli, P. putida and K. michiganensis) were further assayed for off-targets by whole genome sequencing. To address colony heterogeneity, which has been observed previously (7), about a thousand colonies of transconjugant cells transformed with VchCAST and λ-Red VchCAST, were scraped from selection plates and resuspended. Three biological replicates were sequenced separately for E. coli. Colonies were scraped from selection plates and resuspended in a volume of LB media equivalent to OD600=3-4. An aliquot of this resuspension was used for high-molecular-weight genomic DNA (gDNA) extraction using the MasterPure™ Complete DNA and RNA PurificationAtty. Dkt: BERK-543WO Kit (Biosearch Technologies). Genomic DNA samples were submitted for Oxford Nanopore long-read sequencing to Plasmidsaurus Labs.

[0259] The bioinformatics analysis was performed using custom scripts. The demultiplexed raw reads were filtered using nanofilt python package (v 2.8.0) to select reads with a Phred quality score of 20 (>Q20) and a minimum length of 150 base pairs (bp) (58). We focused our analysis on reads exceeding 10,000 bp to guarantee the inclusion of complete cargo and insertion site information. Shorter reads, which were abundant in our dataset, were often fragmented and thus unsuitable for this purpose. To identify the inserted cargo, the reads were mapped to the complete sequence of the transposon using Blastn (59), and reads mapping to less than 80% of the right end of cargo were filtered out.

[0260] For downstream analysis, a pipeline developed by Vo et al. (2021) was employed with some modifications (36). The reads flanking the mapped region were extracted and aligned with the complete reference and plasmid genomes. The reads were classified as genomic, plasmid, or both based on whether more than 100 bp was mapped to the respective genomes. The read classification was further validated by the Bakta annotation software (60), to confirm the presence and correct annotation of genes associated with plasmids and genomic reference.

[0261] Finally, the position of the transposon’s right and left ends as well as the PAM and safe site, were determined by Blastn (59). Transposon orientation (RL or LR) was assigned based on the distances between the right and left ends and the PAM / safe site (Figure 24D-E). The genomic coordinates of the mapped region were recorded to generate the genome-wide histograms of the integration locations. The mapped region was categorized as on-target if it fell within the 100-bp window of the 3'end of the target site. Bioinformatic homology search for regulators

[0262] We examined the phylogenetic distribution of 11 genes across 92 bacterial phyla, which included the nine identified and validated by the RB-TnSeq screen plus ihfA and recA. To simplify viewing, only phyla with more than 10 members in the AnnoTree database (v beta, based on GTDB R214.0 (28)), were included in this analysis. Briefly, the genes in the AnnoTree database were annotated using PFAM v27.0 (61), TIGRFAM v15.0 (62), and KEGG orthology identifiers assigned to the UniRef100 database (63). The KEGG or PFAM annotation assigned to individual genes was used to extract homologs from 80789 representative genomes from GTDB R214.0 (64). Genes such as ihfB, recD, and umuD are found in close association with their respective enzymatic complex subunits, ihfA, recBC, and umuC (65–67). To maintain consistency, the homologs of these co-occurring genes were extracted using the same methods as for the genes of primary interest, and only genomes containing both sets of genes in the same contig were retained for further analysis. The gene yneK lacks a PFAM or KEGG annotation. To address this, protein sequences from the 80,789 representative genomes in the GTDB database (64) were downloaded and subjected to BLASTp analysis against the E. coli yneK sequence. Proteins with greater than 50% sequence similarity and a significant FASTA hit (E ≤ 1e-5) wereAtty. Dkt: BERK-543WO retained. To further validate our findings, a BLASTp search was conducted against a UniProt database subset containing sequences sharing at least 90% similarity to E. coli yneK. Both approaches yielded comparable results. While the yneK gene is prevalent among closely related E. coli and Shigella strains, only a limited number of these strains are present in the 80789 representative genomes in the GTDB database. The cspC gene is a member of the bacterial Cold Shock Protein (CSP) family (68), which can contain multiple paralogs with high sequence similarity (>70%; (69). To accurately differentiate these paralogs, cspC proteins identified by AnnoTree were re-analyzed using eggNOG-mapper v5.0 (29). This tool leverages evolutionary relationships to distinguish between proteins within the CSP family excluding non-cspC family members. Finally, gene presence per phyla was visualized in iTOL v6 (70) with the publicly available GTDB R214.0 bacterial tree inferred from the concatenation of 120 proteins (26, 27).

[0263] To identify genomes possessing type I-F CAST, a systematic review of published literature was conducted. A dataset of 1,064 genomes containing type I-F CAST, derived from three key studies (24, 25, 71), was compiled for subsequent bioinformatic analysis. Protein- coding genes were annotated using Prokka (72), and functional annotation was achieved through KEGG pathway mapping with kofamscan (73). Following the annotation method developed above, KEGG IDs were utilized to retrieve annotations for nine genes, eggNOG-mapper was used for cspC, and BLASTp was employed to identify yneK.

[0264] REFERENCES 1. S.-C. Hsieh, J. E. Peters, Natural and Engineered Guide RNA-Directed Transposition with CRISPR- Associated Tn7-Like Transposons. Annu. Rev. Biochem.93, 139–161 (2024). 2. J. M. Vento, N. Crook, C. L. Beisel, Barriers to genome editing with CRISPR in bacteria. J. Ind. Microbiol. Biotechnol.46, 1327–1341 (2019). 3. C.-W. Chang, V. A. Truong, N. N. Pham, Y.-C. Hu, RNA-guided genome engineering: paradigm shift towards transposons. Trends Biotechnol., doi: 10.1016 / j.tibtech.2024.02.006 (2024). 4. S. E. Klompe, P. L. H. Vo, T. S. Halpin-Healy, S. H. Sternberg, Transposon-encoded CRISPR–Cas systems direct RNA-guided DNA integration. Nature 571, 219–225 (2019). 5. J. Strecker, A. Ladha, Z. Gardner, J. L. Schmid-Burgk, K. S. Makarova, E. V. Koonin, F. Zhang, RNA-guided DNA insertion with CRISPR-associated transposases. Science 365, 48–53 (2019). 6. B. E. Rubin, S. Diamond, B. F. Cress, A. Crits-Christoph, Y. C. Lou, A. L. Borges, H. Shivram, C. He, M. Xu, Z. Zhou, S. J. Smith, R. Rovinsky, D. C. J. Smock, K. Tang, T. K. Owens, N. Krishnappa, R. Sachdeva, R. Barrangou, A. M. Deutschbauer, J. F. Banfield, J. A. Doudna, Species- and site-specific genome editing in complex bacterial communities. Nat. Microbiol.7, 34–47 (2022). 7. D. R. Gelsinger, P. L. H. Vo, S. E. Klompe, C. Ronda, H. H. Wang, S. H. Sternberg, Bacterial genome engineering using CRISPR-associated transposases. Nat. Protoc., doi: 10.1038 / s41596-023- 00927-3 (2024). 8. G. D. Lampe, R. T. King, T. S. Halpin-Healy, S. E. Klompe, M. I. Hogan, P. L. H. Vo, S. Tang, A. Chavez, S. H. Sternberg, Targeted DNA integration in human cells without double-strand breaks using CRISPR-associated transposases. Nat. Biotechnol.42, 87–98 (2024).Atty. Dkt: BERK-543WO S. Yang, J. Zhu, X. Zhou, J. Zhang, Q. Li, F. Bian, J. Zhu, T. Yan, X. Wang, Y. Zhang, J. Yang, Y. Jiang, S. Yang, RNA-Guided DNA Transposition in Corynebacterium glutamicum and Bacillus subtilis. ACS Synth. Biol.12, 2198–2202 (2023). M. W. G. Walker, S. E. Klompe, D. J. Zhang, S. H. Sternberg, Novel molecular requirements for CRISPR RNA-guided transposition. Nucleic Acids Res.51, 4519–4535 (2023). J. E. Alejandre-Sixtos, K. Aguirre-Martínez, J. Cruz-López, A. Mares-Rivera, S. M. Álvarez- Martínez, D. Zamorano-Sánchez, Insights on the regulation and function of the CRISPR / Cas transposition system located in the pathogenicity island VpaI-7 from Vibrio parahaemolyticus RIMD2210633. Infect. Immun., e0016925 (2025). K. M. Wetmore, M. N. Price, R. J. Waters, J. S. Lamson, J. He, C. A. Hoover, M. J. Blow, J. Bristow, G. Butland, A. P. Arkin, A. Deutschbauer, Rapid quantification of mutant fitness in diverse bacteria by sequencing randomly bar-coded transposons. MBio 6, e00306–15 (2015). R. G. Egbert, H. S. Rishi, B. A. Adler, D. M. McCormick, E. Toro, R. T. Gill, A. P. Arkin, A versatile platform strain for high-fidelity multiplex genome editing. Nucleic Acids Res.47, 3244–3256 (2019). S. K. Amundsen, A. F. Taylor, G. R. Smith, The RecD subunit of the Escherichia coli RecBCD enzyme inhibits RecA loading, homologous recombination, and DNA repair. Proc. Natl. Acad. Sci. U. S. A.97, 7399–7404 (2000). E. López, J. Blázquez, Effect of subinhibitory concentrations of antibiotics on intrachromosomal homologous recombination in Escherichia coli. Antimicrob. Agents Chemother.53, 3411–3415 (2009). T. C. Barrett, W. W. K. Mok, A. M. Murawski, M. P. Brynildsen, Enhanced antibiotic resistance development from fluoroquinolone persisters after a single exposure to antibiotic. Nat. Commun.10, 1177 (2019). A. Nath, D. Roizman, N. Pachaimuthu, J. Blázquez, J. Rolff, A. Rodríguez-Rojas, Antibiotic- induced recombination in bacteria requires the formation of double-strand breaks, bioRxiv (2022). https: / / doi.org / 10.1101 / 2022.03.08.483535. K. C. Murphy, The lambda Gam protein inhibits RecBCD binding to dsDNA ends. J. Mol. Biol.371, 19–24 (2007). J. A. Mosberg, M. J. Lajoie, G. M. Church, Lambda red recombineering in Escherichia coli occurs through a fully single-stranded intermediate. Genetics 186, 791–799 (2010). M. Wilkinson, L. Troman, W. A. Wan Nur Ismah, Y. Chaban, M. B. Avison, M. S. Dillingham, D. B. Wigley, Structural basis for the inhibition of RecBCD by Gam and its synergistic antibacterial effect with quinolones. Elife 5 (2016). A. Weimer, M. Kohlstedt, D. C. Volke, P. I. Nikel, C. Wittmann, Industrial biotechnology of Pseudomonas putida: advances and prospects. Appl. Microbiol. Biotechnol.104, 7745–7766 (2020). Z. Yu, S. Li, Y. Li, Z. Jiang, J. Zhou, Q. An, Complete genome sequence of N2-fixing model strain Klebsiella sp. nov. M5al, which produces plant cell wall-degrading enzymes and siderophores. Biotechnol Rep (Amst) 17, 6–9 (2018). I. Prah, Y. Nukui, S. Yamaoka, R. Saito, Emergence of a High-Risk Klebsiella michiganensis Clone Disseminating Carbapenemase Genes. Front. Microbiol.13, 880248 (2022). J. R. Rybarski, K. Hu, A. M. Hill, C. O. Wilke, I. J. Finkelstein, Metagenomic discovery of CRISPR-associated transposons. Proc. Natl. Acad. Sci. U. S. A.118, e2112279118 (2021). J. E. Peters, K. S. Makarova, S. Shmakov, E. V. Koonin, Recruitment of CRISPR-Cas systems byAtty. Dkt: BERK-543WO Tn7-like transposons. Proc. Natl. Acad. Sci. U. S. A.114, E7358–E7366 (2017). D. H. Parks, M. Chuvochina, D. W. Waite, C. Rinke, A. Skarshewski, P.-A. Chaumeil, P. Hugenholtz, A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat. Biotechnol.36, 996–1004 (2018). D. H. Parks, M. Chuvochina, P.-A. Chaumeil, C. Rinke, A. J. Mussig, P. Hugenholtz, A complete domain-to-species taxonomy for Bacteria and Archaea. Nat. Biotechnol.38, 1079–1086 (2020). K. Mendler, H. Chen, D. H. Parks, B. Lobb, L. A. Hug, A. C. Doxey, AnnoTree: visualization and exploration of a functionally annotated microbial tree of life. Nucleic Acids Res.47, 4442–4448 (2019). J. Huerta-Cepas, D. Szklarczyk, D. Heller, A. Hernández-Plaza, S. K. Forslund, H. Cook, D. R. Mende, I. Letunic, T. Rattei, L. J. Jensen, C. von Mering, P. Bork, eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res.47, D309–D314 (2019). Z. Wang, Translesion synthesis by the UmuC family of DNA polymerases. Mutat. Res.486, 59–70 (2001). S. Phadtare, M. Inouye, Sequence-selective interactions with RNA by CspB, CspC and CspE, members of the CspA family of Escherichia coli. Mol. Microbiol.33, 1004–1014 (1999). A. T. Hagemann, N. L. Craig, Tn7 transposition creates a hotspot for homologous recombination at the transposon donor site. Genetics 133, 9–16 (1993). M. S. Dillingham, S. C. Kowalczykowski, RecBCD enzyme and the repair of double-stranded DNA breaks. Microbiol. Mol. Biol. Rev.72, 642–71, Table of Contents (2008). W. Choi, S. Jang, R. M. Harshey, Mu transpososome and RecBCD nuclease collaborate in the repair of simple Mu insertions. Proc. Natl. Acad. Sci. U. S. A.111, 14112–14117 (2014). S. Jang, S. J. Sandler, R. M. Harshey, Mu insertions are repaired by the double-strand break repair pathway of Escherichia coli. PLoS Genet.8, e1002642 (2012). P. L. H. Vo, C. Ronda, S. E. Klompe, E. E. Chen, C. Acree, H. H. Wang, S. H. Sternberg, CRISPR RNA-guided integrases for high-efficiency, multiplexed bacterial genome engineering. Nat. Biotechnol.39, 480–489 (2021). J. Wei, Y. Li, CRISPR-based gene editing technology and its application in microbial engineering. Engineering Microbiology 3, 100101 (2023). J. A. Sawitzke, L. C. Thomason, N. Costantino, M. Bubunenko, S. Datta, D. L. Court, Recombineering: in vivo genetic engineering in E. coli, S. enterica, and beyond. Methods Enzymol. 421, 171–199 (2007). S. S. Naorem, J. Han, S. Y. Zhang, J. Zhang, L. B. Graham, A. Song, C. V. Smith, F. Rashid, H. Guo, Efficient transposon mutagenesis mediated by an IPTG-controlled conditional suicide plasmid. BMC Microbiol.18, 158 (2018). A. Roberts, M. A. Nethery, R. Barrangou, Functional characterization of diverse type I-F CRISPR- associated transposons. Nucleic Acids Res.50, 11670–11681 (2022). U. Fels, K. Gevaert, P. Van Damme, Bacterial Genetic Engineering by Means of Recombineering for Reverse Genetics. Front. Microbiol.11, 548410 (2020). T. L. Ruegg, J. H. Pereira, J. C. Chen, A. DeGiovanni, P. Novichkov, V. K. Mutalik, G. P. Tomaleri, S. W. Singer, N. J. Hillson, B. A. Simmons, P. D. Adams, M. P. Thelen, Jungle Express is a versatile repressor system for tight transcriptional control. Nat. Commun.9, 3617 (2018).Atty. Dkt: BERK-543WO C.-C. Chen, H.-J. Liao, C.-Y. Cheng, C.-Y. Yen, Y.-C. Chung, Biodegradation of crystal violet by Pseudomonas putida. Biotechnol. Lett.29, 391–396 (2007). K. R. Choi, S. Y. Lee, Protocols for RecET-based markerless gene knockout and integration to express heterologous biosynthetic gene clusters in Pseudomonas putida. Microb. Biotechnol.13, 199–209 (2020). N. Sharadamma, Y. Harshavardhana, A. Ravishankar, P. Anand, N. Chandra, K. Muniyappa, Molecular dissection of Mycobacterium tuberculosis integration host factor reveals novel insights into the mode of DNA binding and nucleoid compaction. J. Biol. Chem.289, 34325–34340 (2014). J. P. Swiercz, T. Nanji, M. Gloyd, A. Guarné, M. A. Elliot, A novel nucleoid-associated protein specific to the actinobacteria. Nucleic Acids Res.41, 4171–4184 (2013). S.-C. Hsieh, J. E. Peters, Tn7-CRISPR-Cas12K elements manage pathway choice using truncated repeat-spacer units to target tRNA attachment sites, Molecular Biology (2021). https: / / www.biorxiv.org / content / 10.1101 / 2021.02.06.429022v1.full. Y. Zhang, X. Sun, Q. Wang, J. Xu, F. Dong, S. Yang, J. Yang, Z. Zhang, Y. Qian, J. Chen, J. Zhang, Y. Liu, R. Tao, Y. Jiang, J. Yang, S. Yang, Multicopy Chromosomal Integration Using CRISPR- Associated Transposases. ACS Synth. Biol.9, 1998–2008 (2020). S. Yang, Y. Zhang, J. Xu, J. Zhang, J. Zhang, J. Yang, Y. Jiang, S. Yang, Orthogonal CRISPR- associated transposases for parallel and multiplexed chromosomal integration. Nucleic Acids Res. 49, 10192–10202 (2021). A. B. Banta, K. S. Myers, R. D. Ward, R. A. Cuellar, M. Place, C. C. Freeh, E. E. Bacon, J. M. Peters, A Targeted Genome-scale Overexpression Platform for Proteobacteria. bioRxiv, doi: 10.1101 / 2024.03.01.582922 (2024). J. Xu, Y. Sun, J. Wu, S. Yang, L. Yang, Chromosome recombination and modification by LoxP- mediated evolution in Vibrio natriegens using CRISPR-associated transposases. Biotechnol. Bioeng. 121, 1163–1172 (2024). A. M. Garza Elizondo, J. Chappell, Targeted Transcriptional Activation Using a CRISPR-Associated Transposon System. ACS Synth. Biol.13, 328–336 (2024). B. A. Adler, A. E. Kazakov, C. Zhong, H. Liu, E. Kutter, L. M. Lui, T. N. Nielsen, H. Carion, A. M. Deutschbauer, V. K. Mutalik, A. P. Arkin, The genetic basis of phage susceptibility, cross-resistance and host-range in Salmonella. Microbiology 167 (2021). P. Wang, Z. Yu, B. Li, X. Cai, Z. Zeng, X. Chen, X. Wang, Development of an efficient conjugation- based genetic manipulation system for Pseudoalteromonas. Microb. Cell Fact.14, 11 (2015). T. Baba, T. Ara, M. Hasegawa, Y. Takai, Y. Okumura, M. Baba, K. A. Datsenko, M. Tomita, B. L. Wanner, H. Mori, Construction of Escherichia coli K-12 in-frame, single-gene knockout mutants: the Keio collection. Mol. Syst. Biol.2, 2006.0008 (2006). K. A. Datsenko, B. L. Wanner, One-step inactivation of chromosomal genes in Escherichia coli K- 12 using PCR products. Proc. Natl. Acad. Sci. U. S. A.97, 6640–6645 (2000). Methods for Determining Bactericidal Activity of Antimicrobial Agents (1999). W. De Coster, S. D’Hert, D. T. Schultz, M. Cruts, C. Van Broeckhoven, NanoPack: visualizing and processing long-read sequencing data. Bioinformatics 34, 2666–2669 (2018). C. Camacho, G. Coulouris, V. Avagyan, N. Ma, J. Papadopoulos, K. Bealer, T. L. Madden, BLAST+: architecture and applications. BMC Bioinformatics 10, 421 (2009). O. Schwengers, L. Jelonek, M. A. Dieckmann, S. Beyvers, J. Blom, A. Goesmann, Bakta: rapid andAtty. Dkt: BERK-543WO standardized annotation of bacterial genomes via alignment-free sequence identification. Microb. Genom.7 (2021). 61. R. D. Finn, P. Coggill, R. Y. Eberhardt, S. R. Eddy, J. Mistry, A. L. Mitchell, S. C. Potter, M. Punta, M. Qureshi, A. Sangrador-Vegas, G. A. Salazar, J. Tate, A. Bateman, The Pfam protein families database: towards a more sustainable future. Nucleic Acids Res.44, D279–85 (2016). 62. D. H. Haft, J. D. Selengut, O. White, The TIGRFAMs database of protein families. Nucleic Acids Res.31, 371–373 (2003). 63. B. E. Suzek, Y. Wang, H. Huang, P. B. McGarvey, C. H. Wu, UniProt Consortium, UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics 31, 926–932 (2015). 64. D. H. Parks, M. Chuvochina, C. Rinke, A. J. Mussig, P.-A. Chaumeil, P. Hugenholtz, GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res.50, D785–D794 (2022). 65. H. Haluzi, D. Goitein, S. Koby, I. Mendelson, D. Teff, G. Mengeritsky, H. Giladi, A. B. Oppenheim, Genes coding for integration host factor are conserved in gram-negative bacteria. J. Bacteriol.173, 6297–6299 (1991). 66. M. D. Sutton, T. Opperman, G. C. Walker, The Escherichia coli SOS mutagenesis proteins UmuD and UmuD’ interact physically with the replicative DNA polymerase. Proc. Natl. Acad. Sci. U. S. A. 96, 12373–12378 (1999). 67. A. Bernheim, D. Bikard, M. Touchon, E. P. C. Rocha, A matter of background: DNA repair pathways as a possible cause for the sparse distribution of CRISPR-Cas systems in bacteria. Philos. Trans. R. Soc. Lond. B Biol. Sci.374, 20180088 (2019). 68. P. L. Graumann, M. A. Marahiel, A superfamily of proteins that contain the cold-shock domain. Trends Biochem. Sci.23, 286–290 (1998). 69. A. Catalan-Moreno, C. J. Caballero, N. Irurzun, S. Cuesta, J. López-Sagaseta, A. Toledo-Arana, One evolutionarily selected amino acid variation is sufficient to provide functional specificity in the cold shock protein paralogs of Staphylococcus aureus. Mol. Microbiol.113, 826–840 (2020). 70. I. Letunic, P. Bork, Interactive Tree of Life (iTOL) v6: recent updates to the phylogenetic tree display and annotation tool. Nucleic Acids Res.52, W78–W82 (2024). 71. S. E. Klompe, N. Jaber, L. Y. Beh, J. T. Mohabir, A. Bernheim, S. H. Sternberg, Evolutionary and mechanistic diversity of Type I-F CRISPR-associated transposons. Mol. Cell 82, 616–628.e5 (2022). 72. T. Seemann, Prokka: rapid prokaryotic genome annotation. Bioinformatics 30, 2068–2069 (2014). 73. T. Aramaki, R. Blanc-Mathieu, H. Endo, K. Ohkubo, M. Kanehisa, S. Goto, H. Ogata, KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold. Bioinformatics 36, 2251–2252 (2020).

[0265] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparent to those of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.Atty. Dkt: BERK-543WO

[0266] Accordingly, the preceding merely illustrates the principles of the invention. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

[0267] The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of present invention is embodied by the appended claims. In the claims, 35 U.S.C. §112(f) or 35 U.S.C. §112(6) is expressly defined as being invoked for a limitation in the claim only when the exact phrase "means for" or the exact phrase "step for" is recited at the beginning of such limitation in the claim; if such exact phrase is not used in a limitation in the claim, then 35 U.S.C. § 112 (f) or 35 U.S.C. §112(6) is not invoked.

Claims

Atty. Dkt: BERK-543WO CLAIMS What is claimed is:

1. A transposon system comprising: a) a nucleotide sequence encoding polypeptides that form a CRISPR-associated transposase (CAST) complex; b) a nucleotide sequence encoding a guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome; c) a transposon, or an insertion site for a transposon, wherein the transposon or the transposon insertion site is flanked by recognition sites that are recognized by the CAST complex; and d) a CAST modulator, wherein the CAST modulator enhances the editing efficiency of the transposon system compared to a transposon system without the CAST modulator.

2. The system of Claim 1, wherein the CAST modulator comprises a nucleotide sequence encoding one or more CAST modulator polypeptides.

3. The system of Claim 2, wherein the one or more CAST modulator polypeptides comprise: i) an Exo polypeptide, a Beta polypeptide, and a Gam polypeptide; ii) a RecE polypeptide and a RecT polypeptide; iii) an addA polypeptide and an addB polypeptide; and / or iv) a YnfA polypeptide, a YneK polypeptide, a YnbE polypeptide, a FimF polypeptide, a UmuD polypeptide, a NudG polypeptide, a Mdh polypeptide, a CspC polypeptide, a ChaB polypeptide, a Lpp polypeptide, a IhfA polypeptide, a IhfB polypeptide, a recA polypeptide, an AddA polypeptide, an addB polypeptide, or any combination thereof.

4. The system of Claim 3, wherein: i) the Exo polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 1; ii) the Beta polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 2; and iii) the Gam polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO:

3.

5. The system of Claim 3, wherein: i) the RecE polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 4; and ii) the RecT polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 5.Atty. Dkt: BERK-543WO 6. The system of Claim 3, wherein: i) the addA polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 18; and ii) the addB polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO:

19.

7. The system of Claim 3, wherein: i) the YnfA polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 6; ii) the YneK polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 7; iii) the YnbE polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 8; iv) the FimF polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 9; v) the UmuD polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 10; vi) the NudG polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 11; vii) the Mdh polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 12; viii) the CspC polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 13; ix) the ChaB polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 14; x) the Lpp polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 15; xi) the IhfB polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 16; xii) the IhfA polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO: 17; and xiii) the RecA polypeptide comprises an amino acid sequence that is at least 50% identical to the amino acid sequence set forth in SEQ ID NO:

235. 8 The system of Claim 1, wherein the CAST modulator comprises a small molecule that promotes homologous recombination.

9. The system of any one of claims 1-8, wherein:Atty. Dkt: BERK-543WO (a) and (b) are present on the same nucleic acid construct; or (a), (b), and (c) are all present on the same nucleic acid construct.

10. The system of any one of claims 2-7, wherein (a), (b), (c), and (d) are all present on the same nucleic acid construct.

11. The system of Claim 10, wherein the construct comprises a promoter operably linked to the nucleotide sequence encoding the one or more CAST modulator polypeptides, wherein the promoter is functional in a prokaryotic cell.

12. The system of any one or claims 1-11, wherein the construct comprises a promoter operably linked to the nucleotide sequence encoding the CAST complex polypeptides and to the nucleotide sequence encoding the guide RNA, wherein the promoter is functional in a prokaryotic cell.

13. The system of any one of claims 1-12 wherein the nucleic acid construct is a conjugative nucleic acid construct.

14. The system of any one of claims 1-13, wherein the CAST complex comprises: i) a Cas6 polypeptide, a Cas7 polypeptide, a Cas8 polypeptide, a tnsA polypeptide, a tnsB polypeptide, a tnsC polypeptide, and a tniQ polypeptide; or ii) a Cas12k polypeptide, a tnsC polypeptide, a tnsB polypeptide, and a tniQ polypeptide.

15. The system of Claim 14, wherein: i) the Cas6 polypeptide comprises an amino acid sequence having at least 50% amino acid sequence identity to the amino acid sequence depicted in any one of FIG.10G and FIG.12M-12O; ii) the Cas7 polypeptide comprises an amino acid sequence having at least 50% amino acid sequence identity to the amino acid sequence depicted in any one of FIG.10F and FIG.12P-12R; iii) the Cas8 polypeptide comprises an amino acid sequence having at least 50% amino acid sequence identity to the amino acid sequence depicted in any one of FIG.10E and FIG.12S-12U; iv) the tnsA polypeptide comprises an amino acid sequence having at least 50% amino acid sequence identity to the amino acid sequence depicted in any one of FIG.10A and FIG.12A-12C; v) the tnsB polypeptide comprises an amino acid sequence having at least 50% amino acid sequence identity to the amino acid sequence depicted in any one of FIG.10B and FIG.12D-12F; vi) the tnsC polypeptide comprises an amino acid sequence having at least 50% amino acid sequence identity to the amino acid sequence depicted in any one of FIG.10C and FIG.12G-12I; and vii) the tniQ polypeptide comprises an amino acid sequence having at least 50% amino acid sequence identity to the amino acid sequence depicted in any one of FIG.10D and FIG.12J and 12L.

16. The system of any one of claims 1-15, wherein the transposon has a size of up to 100 kb.Atty. Dkt: BERK-543WO 17. The system of any one of claims 1-16, wherein the construct comprises a selectable marker.

18. The system of any one of claims 1-16, wherein the construct does not comprise a selectable marker.

19. The system of any one of claims 1-18, wherein the transposon comprises: i) one or more nucleotide sequences encoding one or more polypeptides that confer antibiotic resistance on a bacterium; ii) one or more nucleotide sequences encoding one or more enzymes in a biosynthetic pathway; iii) one or more nucleotide sequences encoding a polypeptide that inhibits viability and / or growth of a prokaryotic cell; iv) one or more nucleotide sequences encoding one or more enzymes in a carbon utilization pathway; or v) one or more nucleotide sequences encoding one or more detectable markers.

20. The system of claim 19, wherein the carbon utilization pathway of (iv) is a polysaccharide utilization pathway.

21. The system of claim 19, wherein the detectable marker of (v) is a fluorescent polypeptide.

22. A prokaryotic cell comprising the system of any one of claims 1-21.

23. A library of nucleic acids comprising a plurality of member conjugative nucleic acid constructs, wherein each member conjugative nucleic acid construct comprises: a) a nucleotide sequence encoding CRISPR-associated transposase (CAST) complex polypeptides; b) a nucleotide sequence encoding a guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome; c) a transposon, wherein the transposon is flanked by recognition sites that are cleaved by the transposase; and d) a nucleotide sequence encoding one or more CAST modulator polypeptides.

24. The library of claim 23, wherein each member conjugative nucleic acid construct comprises a nucleotide sequence that provides a unique nucleotide sequence barcode that identifies the member.

25. A library of prokaryotic cells comprising the library of claim 23 or 24.Atty. Dkt: BERK-543WO 26. A method of editing the genome of a target prokaryotic cell, the method comprising introducing into the target bacterium the transposon system of any one of claims 1-21.

27. The method of Claim 26, wherein said introducing comprises contacting one or more target prokaryotic cells with one or more prokaryotic cells according to claim 21, and wherein the construct is transmitted conjugatively from said one or more prokaryotic cells to the one or more target prokaryotic cell.

28. The method of Claim 26 or 27, wherein the one or more target prokaryotic cells are: a) Pseudomonas putida; or b) Klebsiella michiganensis.

29. The method of Claim 26 or 27, wherein the one or more target prokaryotic cells are: a) one or more prokaryotic cells present in or enriched from a natural environment; or b) one or more prokaryotic cells present in a synthetic community of prokaryotic cells.

30. The method of Claim 29, wherein the one or more one target prokaryotic cells are one or more gut bacteria.

31. The method of Claim 29, wherein the natural environment comprises soil.

32. The method of any one of claims 29-31, wherein the one or more target prokaryotic cells are refractory to genetic modification by electroporation and / or heat shock.

33. The method of any one of claims 29-32, wherein the target prokaryotic cells are a heterogeneous population of prokaryotic cells.

34. The method of any one of claims 26-33, wherein said introducing comprises contacting a population of target prokaryotic cells with said one or more prokaryotic cells, and wherein the method comprises, after said introducing, identifying target cells, within the contacted population of target prokaryotic cells, that have an edited genome and / or enriching the contacted population of target prokaryotic cells for target cells having an edited genome.

35. The method of Claim 34, wherein said identifying comprises high throughput nucleic acid sequencing.Atty. Dkt: BERK-543WO 36. The method of Claim 35, wherein the transposon comprises a distinguishable marker and said enriching is based on a phenotype associated with the presence or absence of the distinguishable marker.

37. The method of Claim 36, wherein the distinguishable marker is a screenable marker.

38. The method of Claim 37, wherein the screenable marker is: a) a fluorescent protein encoded by the transposon; b) an epitope encoded by the transposon; or c) a fluorescent aptamer encoded by the transposon.

39. The system of Claim 3, wherein: i) the YneK polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NOs: 7 and 227-230; ii) the UmuD polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NO: 10 and 181-226; iii) the NudG polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NO: 11 and 171-180; iv) the CspC polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NO: 13 and 58-142; v) the ChaB polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NO: 14 and 21-57; vi) the IhfB polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NO: 16 and 157-170; vii) the IhfA polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NO: 17 and 143-156; and viii) the recA polypeptide comprises an amino acid sequence that is at least 50% identical to any one of the amino acid sequences set forth in SEQ ID NO: 235-327.