RNA-guided DNA integration using Tn7-like transposons
The RNA-guided CRISPR-Cas and Tn7-like transposon system addresses the limitations of CRISPR-Cas9 by integrating donor DNA proximal to a target site without DSBs, ensuring precise gene integration in diverse cell types, including non-dividing cells.
Patent Information
- Application Number
- JP2021552850
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-18
- Filing Date
- 2020-03-06
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2040-03-06
AI Technical Summary
Current CRISPR-Cas9 systems face challenges in precise gene integration, including risks of off-target mutations, heterogeneous repair, low HDR efficiency, and inability to integrate into non-dividing cells, necessitating the use of DSBs and laborious donor template preparation.
An RNA-guided DNA integration method using an engineered CRISPR-Cas system derived from type I CRISPR-Cas and a Tn7-like transposon system, which integrates donor DNA proximal to a target site without requiring DSBs, utilizing TnsA, TnsB, TnsC, and TnsD/TniQ components.
Enables precise and efficient gene integration in various cell types, including non-dividing cells, reducing off-target effects and simplifying the integration process.
Smart Images

Figure 0007733576000008 
Figure 0007733576000009 
Figure 0007733576000010
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application No. 62 / 815,187, filed March 7, 2019; U.S. Provisional Application No. 62 / 822,544, filed March 22, 2019; U.S. Provisional Application No. 62 / 845,218, filed May 8, 2019; U.S. Provisional Application No. 62 / 855,814, filed May 31, 2019; and U.S. Provisional Application No. 62 / 866,270, filed June 25, 2019. , which claims the benefit of U.S. Provisional Application No. 62 / 873,455, filed July 12, 2019, U.S. Provisional Application No. 62 / 875,772, filed July 18, 2019, U.S. Provisional Application No. 62 / 884,600, filed August 8, 2019, and U.S. Provisional Application No. 62 / 902,171, filed September 18, 2019, the contents of each of which are incorporated herein by reference.
[0002] FIELD OF THE INVENTION The present invention relates to methods and systems for modifying DNA and other nucleic acids and for gene targeting. In particular, the present invention relates to systems and methods for genetic engineering using engineered transposon-encoded CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas systems. [Background technology]
[0003] The CRISPR / Cas system is a prokaryotic immune system that confers resistance to foreign genetic elements such as plasmids and bacteriophages. The CRISPR / Cas9 system utilizes RNA-guided DNA binding and sequence-specific cleavage of target DNA. The guide RNA (gRNA) is complementary to the target DNA sequence upstream of the PAM (protospacer adjacent motif) site. The Cas (CRISPR-associated) 9 protein binds to the gRNA and the target DNA, introducing a double-strand break (DSB) at a specific location upstream of the PAM site. Geurts et al.,Science 325,433(2009);Mashimo et al.,PLoS ONE 5,e8870(2010);Carbery et al.,Genetics 186,451-459(2010);Tesson et al.,Nat.Biotech.29,695-696(2011);Wiedenheft et al. al.Nature 482,331-338(2012);Jinek et al.Science 337,816-821(2012);Mali et al.Science 339,823-826(2013);Cong et al.Science 339, 819-823 (2013) (all of which are incorporated herein by reference). The CRISPR-Cas9 system can be programmed to cut not only viral DNA but also other genes, opening up new avenues for genome engineering.
[0004] However, there are currently significant limitations and risks associated with using CRISPR-Cas9 and other programmable nucleases to insert large gene cargoes into eukaryotic genomes. CRISPR-Cas9-mediated gene integration requires the introduction of DSBs or the use of synthetic repair donor templates with appropriately designed homologous arms. DSBs are necessary precursors for the CRISPR-Cas9-mediated HDR pathway for gene integration, but they are known to pose a risk to cells. DSBs at off-target sites can introduce off-target mutations, induce the DNA damage response (Haapaniiemi et al., Nat. Med. 24, 927-930 (2018) (incorporated herein by reference)), select for p53-null cells, which increase the risk of tumorigenesis (Ihry et al., Nat. Med. 24, 939-946 (2018) (incorporated herein by reference)), and repair of DSBs at on-target sites can result in large gene deletions, inversions, or chromosomal translocations (Kosicki et al., Nat. Biotechnol. 36, 765-771 (2018) (incorporated herein by reference)). Homologous donors work most efficiently when supplied as recombinant AAV vectors or ssDNA, but are also extremely difficult to produce (see, e.g., Li et al., BioRxiv, 1-24 (2017) (incorporated herein by reference)). Furthermore, cloning a dsDNA donor template with homologous arms can be time-consuming and laborious.
[0005] In addition, gene integration using CRISPR-Cas9 and donor templates relies on homology-directed repair (HDR) for proper integration of the donor template. However, HDR efficiency is known to be very low in many different cell types, and DSBs preceding HDR are always repaired in a heterogeneous manner across the entire cell population. While some cells undergo HDR on one or both alleles, a much larger percentage of cells undergo non-homologous end joining (NHEJ) on one or both alleles, resulting in the introduction of small insertions or deletions at the target site (reviewed in Pawelczak et al., ACS Chem Biol. 13, 389-396 (2018) (incorporated herein by reference)). This means that while only a small percentage of cells across the entire cell population undergo the desired site-specific gene integration (e.g., for editing in therapeutic or experimental applications), a much higher percentage of cells undergo heterogeneous repair. The endogenous machinery for HDR is virtually absent in post-mitotic cells (i.e., non-dividing cells that do not undergo DNA replication), such as neurons and terminally differentiated cells, and therefore there are no options for precise, targeted gene integration in such cell types.
[0006] Many gene therapy products, both commercialized and in clinical trials, use randomly integrating viruses to deliver therapeutic agents into the genome of patient cells (Naldini et al., Science 353, 1101-1102 (2016) (incorporated herein by reference)). The methods of the present invention precisely integrate these therapeutic genes into known safe harbor loci within the genome, ensuring stable expression and completely avoiding the risk of insertional mutagenesis (Bokhoven et al., J Virol. 83, 283-294 (2009) (incorporated herein by reference)). Summary of the Invention
[0007] The systems and methods for RNA-guided DNA integration of the present invention eliminate the above-mentioned risks by eliminating the need to introduce DSBs. The systems and methods of the present invention have significant utility in genetic engineering, including mammalian cell genome engineering.
[0008] In some embodiments, the present disclosure provides a system for RNA-guided DNA integration, the system comprising: (i) an engineered clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system, the CRISPR-Cas system being derived from a type I CRISPR-Cas system and comprising a target site-specific guide RNA (gRNA); and (ii) an engineered transposon system derived from a Tn7-like transposon system, the engineered Tn7-like transposon system comprising TnsA, TnsB, TnsC, and TnsD / TniQ.
[0009] The present disclosure provides methods for RNA-guided DNA integration. In some embodiments, the methods may include introducing into a cell: (i) an engineered clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system, the CRISPR-Cas system being derived from a type I CRISPR-Cas system and including a target site-specific guide RNA (gRNA); (ii) an engineered transposon system derived from a Tn7-like transposon system, including TnsA, TnsB, TnsC, and TnsD / TniQ; and (iii) a donor DNA to be integrated, the donor DNA including a cargo nucleic acid flanked by transposon end sequences; wherein the engineered CRISPR-Cas system binds to the target site and the engineered transposon system integrates the cargo DNA proximal to the target site.
[0010] The method may involve introducing one or more or all of the components of the system of the invention into a cell.
[0011] The systems of the invention can include (i) one or more vectors encoding an engineered CRISPR-Cas system; and (ii) one or more vectors encoding an engineered transposon system; In this case, the CRISPR-Cas system and the transposon system are on the same vector or at least two different vectors.
[0012] The engineered CRISPR-Cas system can include Cas6, Cas7, Cas5, and Cas8. In one embodiment, the stoichiometry of Cas6, Cas7, Cas5, and Cas8 is 1:6:1:1. In some embodiments, Cas5 and Cas8 are combined as a functional fusion protein. In some embodiments, Cas5 and Cas8 are separate.
[0013] The CRISPR-Cas system can include an IF-type variant CRISPR-Cas system. In some embodiments, the engineered transposon system is derived from the Tn7-like transposon systems of Vibrio cholerae, Vibrio cholerae, Photobacterium iliopiscarium, Pseudoalteromonas sp. P1-25, Pseudoalteromonas ruthenica, Photobacterium ganghwense, Shewanella sp. UCD-KL21, Vibrio diazotrophicus, Vibrio sp. 16, Vibrio sp. F12, Vibrio splendidus, Aliivibrio wodanis, and Parashewanella spongiae. In some embodiments, the engineered transposon system is derived from a bacterium selected from the group consisting of Vibrio cholerae strain 4874, Photobacterium iliopiscarium strain NCIMB, Pseudoalteromonas sp. P1-25, Pseudoalteromonas ruthenica strain S3245, Photobacterium ganghwense strain JCM, Shewanella sp. UCD-KL21, Vibrio cholerae strain OYP7G04, Vibrio cholerae strain M1517, Vibrio diazotrophicus strain 60.6F, Vibrio sp. 16, Vibrio sp. F12, Vibrio splendidus strain UCD-SED10, Aliivibrio wodanis 06 / 09 / 160, and Parashewanella spongiae strain HJ039. In one exemplary embodiment, the engineered transposon system is derived from Vibrio cholerae Tn6677.
[0014] Engineered CRISPR-Cas systems can be nuclease-deficient.
[0015] The system of the present invention can further include donor DNA, which includes a cargo nucleic acid flanked by transposon end sequences.
[0016] Integration can be from about 40 base pairs (bp) to about 60 bp, about 48 bp to about 50 bp, about 48 bp, about 49 bp, or about 50 bp from the 3' end of the target site.
[0017] The cell may be a eukaryotic cell or a bacterial cell. The eukaryotic cell may be a mammalian cell, an avian cell, a plant cell, or a fish cell. The mammalian cell may be derived from a human, a primate, a cow, an sheep, a pig, a dog, a mouse, or a rat cell. In one embodiment, the mammalian cell is a human cell. The plant cell may be derived from rice, soybean, corn, tomato, banana, peanut, field pea, sunflower, canola, tobacco, wheat, barley, oat, potato, cotton, carnation, sorghum, or lupin. The avian cell may be derived from a chicken, a duck, or a goose.
[0018] In some embodiments, the systems and methods involve integration of donor DNA without homologous recombination.
[0019] The target site may be flanked by protospacer adjacent motifs (PAMs).
[0020] In some embodiments, provided herein are systems for RNA-guided DNA integration, the systems comprising one or more vectors encoding: a) an engineered clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system, the CRISPR-Cas system comprising Cas5, Cas6, Cas7, and Cas8; and b) an engineered Tn7-like transposon system, the Tn7-like transposon system comprising i) TnsA, ii) TnsB, iii) TnsC, and iv) TnsD and / or TniQ.
[0021] In some embodiments, the CRISPR-cas system is an IB-type CRISPR-cas system. In some embodiments, the CRISPR-cas system is an IF-type CRISPR-cas system. In some embodiments, the CRISPR-cas system is an IF-type variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, the TnsD or TniQ comprises TniQ. In some embodiments, the system further comprises a guide RNA (gRNA) specific to the target site. In some embodiments, the system further comprises donor DNA to be integrated, the donor DNA comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, the cargo nucleic acid sequence being flanked by the first transposon end sequence and the second transposon end sequence.
[0022] In some embodiments, the first transposon end sequence and the second transposon end sequence are Tn7 transposon end sequences. In some embodiments, the CRISPR-Cas system and the Tn7-like transposon system are on the same vector. In some embodiments, the engineered Tn7-like transposon system is derived from Vibrio cholerae Tn6677. In some embodiments, the engineered CRISPR-Cas system is nuclease-deficient. In some embodiments, one or more vectors are plasmids.
[0023] In certain embodiments, at least one cas protein of the CRISPR-cas system is derived from a type V CRISPR-cas system. In some embodiments, the at least one cas protein is C2c5. In some embodiments, at least one cas protein of the CRISPR-cas system is derived from a type II-A CRISPR-cas system, wherein the at least one Cas protein is Cas9. In some embodiments, the engineered CRISPR-cas system and engineered transposon system are derived from a type I CRISPR-cas system and transposon system, wherein the system further comprises a second engineered CRISPR-cas system and a second engineered transposon system, both of which are derived from a type V CRISPR-cas system and transposon system.
[0024] In some embodiments, provided herein is a method for RNA-guided DNA integration, comprising introducing into a cell: i) an engineered CRISPR-Cas system and / or one or more vectors encoding the engineered CRISPR-Cas system; ii) an engineered transposon system and / or one or more vectors encoding the engineered transposon system; and iii) a donor sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein, when one or more vectors are used, the CRISPR-Cas system and the transposon system are on the same or different vector(s); the cell comprises a nucleic acid sequence having a target site; the CRISPR-cas system comprises (a) at least one cas protein and (b) a guide RNA (gRNA); and wherein the CRISPR-cas system binds to the target site; and the transposon system integrates the donor sequence downstream of the target site.
[0025] In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the at least one Cas protein is derived from a type I CRISPR-Cas system. In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the type I CRISPR-Cas system is type IB or type IF. In some embodiments, the type I CRISPR-Cas system is a type IF variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC. In some embodiments, the transposon system is derived from a Tn7-like transposon system.
[0026] In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC. In some embodiments, the Tn7 transposon system is derived from Vibrio cholerae. In some embodiments, the transposon system comprises i) TnsA, TnsB, and TnsC, and ii) TnsD and / or TniQ. In some embodiments, at least one Cas protein of the CRISPR-Cas system is derived from a type V CRISPR-Cas system. In some embodiments, at least one Cas protein is C2c5. In some embodiments, at least one Cas protein of the CRISPR-Cas system is derived from a type II-A CRISPR-cas system. In some embodiments, at least one Cas protein is Cas9. In some embodiments, one or more vectors are plasmids (e.g., the only plasmid). In some embodiments, the engineered CRISPR-cas system and engineered transposon system are derived from a type I CRISPR-cas system and transposon system, and the system further comprises a second engineered CRISPR-cas system and a second engineered transposon system, both of which are derived from a type V CRISPR-cas system and transposon system.
[0027] In some embodiments, provided herein are systems for RNA-guided DNA integration, the systems comprising one or more vectors encoding: a) an engineered clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system, the CRISPR-Cas system comprising Cas5, Cas6, Cas7, and Cas8; and b) an engineered Tn7-like transposon system, the Tn7-like transposon system comprising i) TnsA, ii) TnsB, iii) TnsC, and iv) TnsD and / or TniQ.
[0028] In some embodiments, the CRISPR-Cas system is an IB or IF CRISPR-Cas system. In some embodiments, the CRISPR-Cas system is an IF variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, Cas5 and Cas8 are expressed as separate, non-fused proteins. In some embodiments, one or more vectors are plasmids.
[0029] In some embodiments, the system further comprises a guide RNA (gRNA) specific to the target site. In some embodiments, the system further comprises donor DNA to be integrated, the donor DNA comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, the cargo nucleic acid sequence being flanked by the first transposon end sequence and the second transposon end sequence. In some embodiments, the donor DNA is at least 2 kb in length (e.g., 2 kb, 5 kb, 10 kb, or more). In certain embodiments, the CRISPR-Cas system and the Tn7-like transposon system are on the same vector. In some embodiments, the engineered Tn7-like transposon system is derived from Vibrio cholerae Tn6677. In some embodiments, the engineered CRISPR-Cas system is nuclease-deficient.
[0030] In some embodiments, provided herein are methods for RNA-guided DNA integration, the methods comprising administering into a cell: a) one or more vectors encoding an engineered transposon-encoded CRISPR-Cas system, the CRISPR-Cas system comprising: i) an engineered clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system comprising: A) Cas5, Cas6, Cas7, and Cas8; and B) a target site-specific guide RNA (gRNA); and ii) an engineered Tn7-like transposon system comprising: A) TnsA, B) TnsB, C) TnsC, and D a) one or more vectors comprising a Tn7-like transposon system comprising TnsD and / or TniQ; and b) donor DNA to be integrated, the donor DNA comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, the cargo nucleic acid sequence being flanked by the first transposon end sequence and the second transposon end sequence, wherein the engineered transposon-encoded CRISPR-Cas system integrates the donor DNA proximal to the target site, and the transposon-encoded CRISPR-Cas system and the donor DNA are on the same vector or at least two different vectors.
[0031] In some embodiments, the CRISPR-cas system is an IB-type or IF-type CRISPR-cas system. In some embodiments, the CRISPR-cas system is an IF-type variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, one or more vectors encode an engineered CRISPR-Cas system, one or more vectors encode an engineered Tn7-like transposon system, and the CRISPR-Cas system and the Tn7-like transposon system are on at least two different vectors. In some embodiments, the donor DNA is integrated about 40 base pairs (bp) to about 60 bp 3' from the target site. In some embodiments, the donor DNA is integrated about 48 bp to about 50 bp 3' from the target site. In some embodiments, the donor DNA is integrated about 50 bp 3' from the target site.
[0032] In some embodiments, the cell is a eukaryotic cell or a bacterial cell. In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the engineered Tn7-like transposon system is derived from Vibrio cholerae Tn6677. In some embodiments, the engineered CRISPR-Cas system is nuclease-deficient. In some embodiments, the target site is flanked by a protospacer adjacent motif (PAM). In some embodiments, provided herein are cells comprising the systems described above and herein.
[0033] In some embodiments, provided herein are kits comprising: a) one or more vectors encoding i) an engineered clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system, the CRISPR-Cas system comprising Cas5, Cas6, Cas7, and Cas8; and ii) an engineered Tn7-like transposon system, the Tn7-like transposon system comprising A) TnsA, B) TnsB, C) TnsC, and D) TnsD and / or TniQ; and b) at least one component selected from the group consisting of i) an infusion device, ii) an intravenous solution bag, iii) a vial having a stopper pierceable by a hypodermic needle, iv) a buffer solution, v) a control plasmid, and vi) a sequencing primer.
[0034] In some embodiments, one or more vectors are plasmids. In some embodiments, Cas5 and Cas8 are expressed as separate, non-fused proteins. In some embodiments, the CRISPR-cas system is an IF-type variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, the kit further comprises a donor nucleic acid sequence, the donor nucleic acid sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence.
[0035] In some embodiments, provided herein are methods for inactivating a microbial gene, the methods comprising introducing into one or more cells a) an engineered transposon-encoded CRISPR-Cas system, and / or b) one or more vectors encoding an engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) a guide RNA (gRNA) specific for a target site proximal to the microbial gene; iii) the engineered transposon system; and iv) donor DNA, wherein the transposon-encoded CRISPR-Cas system inserts the donor DNA into the microbial gene.
[0036] In some embodiments, the microbial gene is a bacterial antibiotic resistance gene, a virulence gene, or a metabolic gene. In some embodiments, the donor DNA comprises a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence. In some embodiments, the cargo nucleic acid sequence encodes a CRISPR-Cas system encoded by the engineered transposon.
[0037] In some embodiments, the one or more cells are bacterial cells, and wherein the introducing comprises contacting an initial cell comprising the transposon-encoded CRISPR-Cas system with a recipient cell, such that the transposon-encoded CRISPR-Cas system is passed to the recipient cell via the bacterial conjugate.
[0038] In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the at least one Cas protein is from a Type I CRISPR-cas system. In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the Type I CRISPR-cas system is Type IB or Type IF. In some embodiments, the Type I CRISPR-cas system is a Type IF variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion.
[0039] In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC. In some embodiments, the transposon system is derived from the Tn7 transposon system. In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC. In some embodiments, the Tn7 transposon system is derived from Vibrio cholerae. In some embodiments, the transposon system comprises i) TnsA, TnsB, and TnsC, and ii) TnsD and / or TniQ. In some embodiments, at least one Cas protein of the CRISPR-cas system is derived from a type V CRISPR-cas system. In some embodiments, at least one Cas protein is C2c5. In some embodiments, at least one Cas protein of the CRISPR-cas system is derived from a type II-A CRISPR-Cas system. In some embodiments, at least one Cas protein is Cas9. In some embodiments, the engineered CRISPR-cas system and engineered transposon system are derived from a type I CRISPR-cas system and transposon system, and the system further comprises a second engineered CRISPR-cas system and a second engineered transposon system, both of which are derived from a type V CRISPR-cas system and transposon system.
[0040] In some embodiments, the present disclosure provides a method for producing a CRISPR-Cas system comprising: a) contacting a sample with i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more vectors encoding an engineered transposon-encoded CRISPR-Cas system, wherein the sample comprises an input nucleic acid sequence comprising: A) a double-stranded nucleic acid sequence of interest (NASI), B) a double-stranded first flanking region on one side of the NASI, and C) a double-stranded second flanking region on the other side of the NASI, and wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas The method includes contacting a sample comprising: ii) an engineered transposon system; iii) a first left transposon end sequence; iv) a first right transposon end sequence that is not covalently linked to the first left transposon end sequence; and v) a first guide RNA (gRNA-1) that targets the first left transposon end sequence and the first right transposon end sequence to a first flanking region; and b) incubating the sample under conditions such that the first left transposon end sequence and the first right transposon end sequence are integrated into the first flanking region.
[0041] In some embodiments, methods described herein include a) contacting a sample with i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more vectors encoding an engineered transposon-encoded CRISPR-Cas system, wherein the sample comprises an input nucleic acid sequence comprising: A) a double-stranded nucleic acid sequence of interest (NASI); B) a double-stranded first flanking region on one side of the NASI; and C) a double-stranded second flanking region on the other side of the NASI, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) the engineered transposon system; iii) a first left transposon end sequence; iv) a first right transposon end sequence that is not covalently linked to the first left transposon end sequence; and v) a first right transposon end sequence that is not covalently linked to the first left transposon end sequence. a) contacting a sample comprising: i) a second left transposon end sequence; vi) a second right transposon end sequence that is not covalently linked to the second left transposon end sequence; vii) a first guide RNA (gRNA-1) that targets the first left transposon end sequence and the first right transposon end sequence to a first flanking region; and viii) a second guide RNA (gRNA-2) that targets the second left transposon end sequence and the second right transposon end sequence to a second flanking region; and b) incubating the sample under conditions such that i) the first left transposon end sequence and the first right transposon end sequence are integrated into the first flanking region, and ii) the second left transposon end sequence and the second right transposon end sequence are integrated into the second flanking region.
[0042] In some embodiments, the method further includes c) contacting the sample with i) a first primer specific to the first left transposon end sequence or the first right transposon end sequence, ii) a second primer specific to the second left transposon end sequence or the second right transposon end sequence, and iii) a polymerase, and d) treating the sample under amplification conditions such that the NASI is amplified to produce an amplified NASI. In some embodiments, the method further includes e) sequencing the amplified NASI. In some embodiments, the sequencing is next-generation sequencing (NGS).
[0043] In some embodiments, the first left transposon end sequence or the first right transposon end sequence comprises a first adapter sequence, and the second left transposon end sequence or the second right transposon end sequence comprises a second adapter sequence. In some embodiments, the method further includes c) contacting the sample with i) a first primer specific to the first adapter sequence, ii) a second primer specific to the second adapter sequence, and iii) a polymerase; and d) treating the sample under amplification conditions such that the NASI is amplified to produce an amplified NASI. In some embodiments, the method further includes e) sequencing the amplified NASI. In some embodiments, the sequencing is next-generation sequencing (NGS). In some embodiments, the first adapter sequence and the second adapter sequence are next-generation sequencing adapters. In some embodiments, the left transposon end sequence comprises a first UMI sequence, and the right transposon end sequence comprises a second UMI sequence.
[0044] In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the at least one Cas protein is derived from a type I CRISPR-cas system. In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the type I CRISPR-cas system is type IB or type IF. In some embodiments, the type I CRISPR-cas system is a type IF variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC. In some embodiments, the transposon system is derived from a Tn7-like transposon system. In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC.
[0045] In some embodiments, the Tn7 transposon system is derived from Vibrio cholerae. In some embodiments, the transposon system comprises i) TnsA, TnsB, and TnsC, and ii) TnsD and / or TniQ. In some embodiments, at least one Cas protein of the CRISPR-Cas system is derived from a type V CRISPR-cas system. In some embodiments, at least one Cas protein is C2c5. In some embodiments, at least one Cas protein of the CRISPR-cas system is derived from a type II-A CRISPR-Cas system. In some embodiments, at least one Cas protein is Cas9. In some embodiments, the engineered CRISPR-cas system and engineered transposon system are derived from a type I CRISPR-cas system and transposon system, wherein the system further comprises a second engineered CRISPR-cas system and a second engineered transposon system, both of which are derived from a type V CRISPR-cas system and transposon system.
[0046] In some embodiments, provided herein are methods for RNA-guided DNA integration in plant cells, the methods comprising introducing into the plant cell a) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more vectors encoding the engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system comprises i) at least one Cas protein, ii) a target site-specific guide RNA (gRNA), iii) the engineered transposon system, and iv) donor DNA, wherein the transposon-encoded CRISPR-Cas system integrates the donor DNA proximal to a target nucleic acid site in the plant cell.
[0047] In some embodiments, the plant cell is a cell of rice, soybean, corn, tomato, banana, peanut, field pea, sunflower, canola, tobacco, wheat, barley, oat, potato, cotton, carnation, sorghum, lupine, Solanum lycopersicum, Glycine max, Arabidopsis thaliana, Medicago truncatula, Brachypodium distachion, Oryza sativa, Sorghum bicolor, Zea mays, or Solanum tuberosum. In some embodiments, the plant cell is a cell of petunia, Atropa, rutabaga, celery, switchgrass, apple, Nicotiana benthamiana, or Setaria viridis. In some embodiments, the plant cell is a cell of a monocotyledonous or dicotyledonous plant.
[0048] In some embodiments, integration of the donor DNA results in a change in one or more of the following traits in the plant cell: kernel number, kernel size, kernel weight, panicle size, tiller number, aroma, nutritional value, shelf life, lycopene content, starch content, and / or ii) reduced gluten content, reduced toxin levels, reduced steroid glycoalkaloid levels, replacement of meiosis with mitosis, asexual propagation, improved haploid propagation, and / or reduced growth time. In some embodiments, integration of the donor DNA results in a change in one or more of the following traits in the plant cell: herbicide tolerance, drought tolerance, male sterility, insect resistance, abiotic stress tolerance, modified fatty acid metabolism, modified carbohydrate metabolism, modified seed yield, modified oil percent, modified protein percent, resistance to bacterial disease, resistance to fungal disease, and resistance to viral disease.
[0049] In some embodiments, the transposon-encoded CRISPR-Cas system integrates the donor DNA into the genome of the plant cell. In some embodiments, one or more vectors encoding the transposon-encoded CRISPR-Cas system are introduced into the plant cell via Agrobacterium-mediated transformation of the plant cell.
[0050] In some embodiments, the donor DNA comprises a first transposon end sequence and a second transposon end sequence. In some embodiments, the transposon system is a bacterial Tn7-like transposon system. In some embodiments, the transposon-encoded CRISPR-Cas system comprises TnsD and / or TniQ. In some embodiments, the transposon-encoded CRISPR-Cas system comprises TnsA, TnsB, and TnsC. In some embodiments, the transposon-encoded CRISPR-Cas system is nuclease-deficient. In some embodiments, the transposon-encoded CRISPR-Cas system is derived from a type I CRISPR-Cas system. In some embodiments, the transposon-encoded CRISPR-Cas system comprises a Cascade complex.
[0051] In some embodiments, the transposon-encoded CRISPR-Cas system is derived from a type II CRISPR-Cas system. In some embodiments, the transposon-encoded CRISPR-Cas system is derived from a type V CRISPR-Cas system. In some embodiments, the transposon-encoded CRISPR-Cas system comprises C2c5. In some embodiments, the target site is flanked by protospacer adjacent motifs (PAMs). In some embodiments, the donor DNA integrates approximately 46 bp to 55 bp downstream of the target site. In some embodiments, the donor DNA integrates approximately 47 bp to 51 bp downstream of the target site.
[0052] In certain embodiments, provided herein are modified plant cells produced by the methods described above and herein. In certain embodiments, provided herein are plants or seeds comprising such plant cells. In some embodiments, provided herein are fruits, plant parts, or propagation material of such plants.
[0053] In some embodiments, provided herein are methods for RNA-guided DNA integration in animal cells, the methods comprising introducing into the animal cell a) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more vectors encoding the engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system comprises i) at least one Cas protein, ii) a guide RNA (gRNA) specific for a target site, iii) the engineered transposon system, and iv) donor DNA, wherein the transposon-encoded CRISPR-Cas system integrates the donor DNA proximal to the target site in the animal cell.
[0054] In some embodiments, the animal cell is a mouse, rat, rabbit, cow, sheep, pig, chicken, horse, buffalo, camel, turkey, or goose cell. In some embodiments, the animal cell is a mammalian cell. In some embodiments, the mammal is an orangutan, monkey, horse, cow, sheep, goat, pig, donkey, dog, rabbit, cat, rat, or mouse. In some embodiments, the animal cell is a livestock animal cell. In some embodiments, the transposon-encoded CRISPR-Cas system integrates the donor DNA into the genome of the animal cell.
[0055] In some embodiments, the donor DNA comprises transposon end sequences. In some embodiments, the transposon system is a bacterial Tn7-like transposon system. In some embodiments, the transposon-encoded CRISPR-Cas system comprises TnsD and / or TniQ. In some embodiments, the transposon-encoded CRISPR-Cas system comprises TnsA, TnsB, and TnsC. In some embodiments, the transposon-encoded CRISPR-Cas system is nuclease-deficient. In some embodiments, the transposon-encoded CRISPR-Cas system is derived from a Type I CRISPR-Cas system. In some embodiments, the transposon-encoded CRISPR-Cas system comprises a Cascade complex. In some embodiments, the transposon-encoded CRISPR-Cas system is derived from a Type II CRISPR-Cas system. In some embodiments, the transposon-encoded CRISPR-Cas system is derived from a Type V CRISPR-Cas system. In some embodiments, the transposon-encoded CRISPR-Cas system comprises C2c5. In some embodiments, the target site is adjacent to a protospacer adjacent motif (PAM). In some embodiments, the donor DNA integrates approximately 46 bp to 55 bp downstream of the target site. In some embodiments, the donor DNA integrates approximately 47 bp to 51 bp downstream of the target site. In some embodiments, the Tn7-like transposon system is derived from Vibrio cholerae.
[0056] In some embodiments, provided herein are modified non-human animal cells produced by the methods described above and herein. In some embodiments, provided herein are genetically modified non-human animals comprising such animal cells. In some embodiments, provided herein are populations of cells, tissues, or organs comprising such animal cells.
[0057] In some embodiments, provided herein are compositions comprising: a) an engineered transposon-encoded CRISPR-Cas system; and / or b) one or more nucleic acid sequence(s) encoding an engineered transposon-encoded CRISPR-Cas system, wherein the engineered transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) a guide RNA (gRNA) specific for a target site within human DNA; iii) the engineered transposon system; and iv) a donor nucleic acid comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence.
[0058] In some embodiments, provided herein is a kit comprising: a) the composition described above; and b) a device for holding the composition. In some embodiments, the device is selected from the group consisting of an infusion device, an intravenous solution bag, and a vial having a stopper pierceable by a hypodermic needle.
[0059] In some embodiments, provided herein are methods of treating a subject (e.g., a human) comprising: a) administering (e.g., intravenously) to the mammalian subject one or more compositions comprising subject cells and microbiome cells, wherein the one or more compositions comprise i) an engineered transposon-encoded CRISPR-Cas system and / or ii) one or more nucleic acid sequence(s) encoding an engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system encodes i) at least one Cas protein and ii) a gene encoding the genome of the subject cell. the donor nucleic acid comprising: iii) a guide RNA (gRNA) specific for a target site within the genome of the subject cells or microbiome cells; iii) an engineered transposon system; and iv) a donor nucleic acid comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence, wherein the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid proximal to the target site within the genome of at least one of the subject cells and / or at least one of the microbiome cells.
[0060] In certain embodiments, provided herein are methods of treating cells in vitro, the methods comprising: a) contacting at least one cell in vitro with a composition comprising i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more nucleic acid sequence(s) encoding an engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) a guide RNA (gRNA) specific for a target site within the genome of the cell; iii) the engineered transposon system; and iv) a donor nucleic acid sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence, and wherein the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid proximal to the target site within the genome of the at least one cell.
[0061] In some embodiments, provided herein are methods for RNA-guided nucleic acid integration in cells, the methods comprising: (a) introducing into a population of cells i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more nucleic acid sequence(s) encoding the engineered transposon-encoded CRISPR-Cas system, wherein the engineered transposon-encoded CRISPR-Cas system comprises: A) at least one Cas protein; B) a guide RNA (gRNA) specific for a target site within the genome of the cells; C) the engineered transposon system; and D) a donor nucleic acid at least 2 kb in length, the donor nucleic acid comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence; and (b) culturing the cells under conditions such that the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid proximal to the target site within the genome of the cells. In some embodiments, the donor nucleic acid sequence is at least 10 kb, at least 50 kb, at least 100 kb, or 20-60 kb in length. In some embodiments, the cells are bacterial cells and the conditions include culturing the bacterial cells at a temperature at least 5°C below the optimal growth temperature of the bacterial cells. In some embodiments, the bacterial cells are E. coli cells and the E. coli cells are cultured at a temperature of 30°C or less.
[0062] In some embodiments, the cell is a human cell, a plant cell, a bacterial cell, or an animal cell. In some embodiments, the one or more nucleic acid sequence(s) comprise one or more vectors. In some embodiments, the one or more nucleic acid sequence(s) comprise at least one mRNA sequence.
[0063] In some embodiments, the subject is a human. In some embodiments, the subject is a human with a disease selected from the group consisting of cancer, Duchenne muscular dystrophy (DMD), sickle cell disease (SCD), β-thalassemia, and hereditary tyrosinemia type 1 (HT1). In some embodiments, the cargo nucleic acid sequence comprises a therapeutic sequence.
[0064] In some embodiments, the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid sequence using a cut-and-paste transposition pathway. In some embodiments, at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8, and the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) TniQ. In some embodiments, at least one of the following applies: I) Cas5 and Cas8 form a Cas5-Cas8 fusion protein; II) TniQ and Cas6 form a TniQ-Cas6 fusion protein; and / or III) TnsA and TnsB form a TnsA-TnsB fusion protein. In some embodiments, TniQ is fused to at least one Cas protein to generate a TniQ-Cas fusion polypeptide. In some embodiments, at least one Cas protein is Cas6.
[0065] In some embodiments, at least one Cas protein is derived from a type I CRISPR-Cas system. In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the type I CRISPR-Cas system is type IB or type IF. In some embodiments, the type I CRISPR-Cas system is an IF variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC. In some embodiments, the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) TnsD and / or TniQ. In some embodiments, TnsA and TnsB are expressed as TnsA-TnsB fusion proteins. In some embodiments, the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) TniQ family proteins.
[0066] In some embodiments, the methods, compositions, and kits further comprise a second guide RNA (gRNA-2), where gRNA-2 directs the donor DNA to integrate proximal to a second, different target site. In some embodiments, the methods, compositions, and kits further comprise a third guide RNA (gRNA-3), where gRNA-3 directs the donor DNA to integrate proximal to a third, different target site.
[0067] In some embodiments, the transposon system is derived from a Tn7-like transposon system. In some embodiments, the Tn7 transposon system is derived from Vibrio cholerae. In some embodiments, at least one Cas protein of the CRISPR-cas system is derived from a type V CRISPR-cas system. In some embodiments, the at least one Cas protein comprises C2c5. In some embodiments, the engineered transposon-encoded CRISPR-Cas system is derived from Scytonema hofmannii PCC 7110. In some embodiments, at least one Cas protein of the CRISPR-cas system is derived from a type II-A CRISPR-cas system. In some embodiments, the at least one Cas protein is Cas9. In some embodiments, the engineered CRISPR-cas system and engineered transposon system are derived from a type I CRISPR-cas system and transposon system, wherein the system further comprises a second engineered CRISPR-cas system and a second engineered transposon system, both of which are derived from a type V CRISPR-cas system and transposon system.
[0068] In some embodiments, the donor nucleic acid is at least 2 kb in length. In some embodiments, the donor nucleic acid is at least 10 kb in length. In some embodiments, the one or more nucleic acid sequences are one or more viral vectors selected from the group consisting of retroviral vectors, lentiviral vectors, adenoviral vectors, adeno-associated viral vectors, and herpes simplex viral vectors. In some embodiments, the one or more nucleic acid sequence(s) further comprise one or more promoters. In some embodiments, the one or more nucleic acid sequences are the only vector. In some embodiments, the only vector comprises the only promoter.
[0069] In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8, wherein Cas5 and Cas8 form a fusion protein. In some embodiments, the first transposon end sequence is a left transposon end sequence, and the second transposon end sequence is a right transposon end sequence.
[0070] In some embodiments, the left transposon end sequence and / or the right transposon end sequence are variant sequences that increase the efficiency of integration of the donor nucleic acid sequence compared to the corresponding wild-type left transposon end sequence and / or right transposon end sequence. In some embodiments, the left transposon end sequence and / or the right transposon end sequence alters the directional bias of the donor nucleic acid sequence when integrated proximal to a target site in a genome compared to the corresponding wild-type left transposon end sequence and / or right transposon end sequence. In some embodiments, the directional bias is imparted to a tRL. In some embodiments, the directional bias is imparted to a tLR.
[0071] In some embodiments, the first transposon end sequence and / or the second transposon end sequence encode a functional protein linker sequence. In some embodiments, the genome of the subject cell or microbiome cell comprises a target protein-encoding gene, wherein the cargo nucleic acid sequence encodes an amino acid sequence of interest, and the donor nucleic acid sequence is inserted adjacent to or within the target protein-encoding gene to generate a fusion protein-encoding sequence, wherein the fusion protein comprises the amino acid sequence of interest appended to the target protein. In some embodiments, the amino acid sequence of interest is selected from the group consisting of a fluorescent protein, an epitope tag, and a degron tag.
[0072] In some embodiments, the genome of the cell or microbiome cell comprises a target protein-encoding gene, wherein the cargo nucleic acid sequence comprises i) an amino acid sequence of interest coding region (AASIER), ii) a splice acceptor and / or donor site flanking the AASIER, and the donor nucleic acid sequence is inserted adjacent to or within the target protein-encoding gene to generate a synthetic, engineered exon that allows in-frame tagging of the target protein with the amino acid sequence of interest.
[0073] In some embodiments, the engineered transposon-encoded CRISPR-Cas system is derived from a bacterium selected from the group consisting of Vibrio cholerae, Photobacterium iliopiscarium, Pseudoalteromonas sp. P1-25, Pseudoalteromonas ruthenica, Photobacterium ganghwense, Shewanella sp. UCD-KL21, Vibrio diazotrophicus, Vibrio sp. 16, Vibrio sp. F12, Vibrio splendidus, Aliivibrio wodanis, and Parashewanella spongiae. In some embodiments, the engineered transposon-encoded CRISPR-Cas system is derived from a bacterium selected from the group consisting of Vibrio cholerae strain 4874, Photobacterium iliopiscarium strain NCIMB, Pseudoalteromonas sp. P1-25, Pseudoalteromonas ruthenica strain S3245, Photobacterium ganghwense strain JCM, Shewanella sp. UCD-KL21, Vibrio cholerae strain OYP7G04, Vibrio cholerae strain M1517, Vibrio diazotrophicus strain 60.6F, Vibrio sp. 16, Vibrio sp. F12, Vibrio splendidus strain UCD-SED10, Aliivibrio wodanis 06 / 09 / 160, and Parashewanella spongiae strain HJ039.
[0074] In some embodiments, the cargo nucleic acid sequence comprises an element selected from the group consisting of a natural transcription promoter, a synthetic transcription promoter, an inducible transcription promoter, a constitutive transcription promoter, a natural transcription terminator, a synthetic transcription terminator, an origin of replication, a replication termination sequence, a centromere sequence, and a telomere sequence. In some embodiments, the cargo nucleic acid sequence encodes at least one of a therapeutic protein, a metabolic pathway, and / or a biosynthetic pathway.
[0075] In some embodiments, provided herein are methods of treating a cell, the method comprising: a) contacting at least one cell with a composition comprising i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more nucleic acid sequence(s) encoding an engineered transposon-encoded CRISPR-Cas system, wherein the engineered transposon-encoded CRISPR-Cas system comprises i) at least one Cas protein, and ii) a guide RNA (gRNA) specific for a target site within the genome of the cell. iii) an engineered transposon system; and iv) a donor nucleic acid sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence, and the cargo nucleic acid sequence is at least 2 kb in length (e.g., 2 kb, 5 kb, 50 kb, 100 kb, or more), wherein the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid proximal to a target site in the genome of at least one cell.
[0076] In some embodiments, provided herein are compositions comprising: i) an engineered transposon-encoded CRISPR-Cas system; and / or ii) one or more nucleic acid sequence(s) encoding an engineered transposon-encoded CRISPR-Cas system, wherein the engineered transposon-encoded CRISPR-Cas system comprises: a) at least one Cas protein; b) a guide RNA (gRNA) specific for a target site within the genome of a cell; c) the engineered transposon system; and d) a donor nucleic acid sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence, and the cargo nucleic acid sequence is at least 2 kb in length (e.g., 2 kb, 5 kb, 50 kb, 100 kb, or more).
[0077] In some embodiments, provided herein are compositions comprising a self-transposable nucleic acid sequence comprising: a) a mobilizable nucleic acid sequence encoding a transposon-encoded CRISPR-Cas system; and b) a first transposon end sequence and a second transposon end sequence flanking the mobilizable nucleic acid sequence, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) a guide RNA (gRNA) specific for a target site; and iii) an engineered transposon system.
[0078] In some embodiments, provided herein are methods for targeting cancer cells, the methods comprising introducing into the cancer cells i) an engineered transposon-encoded CRISPR-Cas system and / or ii) one or more nucleic acid sequence(s) encoding the engineered transposon-encoded CRISPR-Cas system, wherein the engineered transposon-encoded CRISPR-Cas system comprises: A) at least one Cas protein; B) a guide RNA (gRNA) specific for a target site within the genome of the cancer cell; C) the engineered transposon system; and D) a donor nucleic acid sequence comprising a first transposon end sequence and a second transposon end sequence. In certain embodiments, the introducing is performed under conditions such that the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid sequence proximal to the target site within the genome of the cancer cell. In some embodiments, the target site is within a genomic sequence associated with an oncogene. In some embodiments, the donor nucleic acid disrupts pathogenic expression of the oncogene.
[0079] In some embodiments, the composition further comprises a vector, wherein the self-transferring nucleic acid sequence is present within the vector. In some embodiments, the composition further comprises a cell having genomic DNA, wherein the self-transferring nucleic acid sequence is present within the genomic DNA.
[0080] In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the at least one Cas protein is derived from a type I CRISPR-cas system. In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the type I CRISPR-cas system is type IB or type IF. In some embodiments, the type I CRISPR-cas system is an IF variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. In some embodiments, the transposon system comprises TnsA, TnsB, and TnsC. In some embodiments, the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) TnsD and / or TniQ. In some embodiments, TnsA and TnsB are expressed as TnsA-TnsB fusion proteins. In some embodiments, TniQ is fused to at least one Cas protein to generate a TniQ-Cas fusion polypeptide. In some embodiments, at least one Cas protein is Cas6. In some embodiments, the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) a TniQ family protein.
[0081] In some embodiments, the transposon system is derived from a Tn7-like transposon system. In some embodiments, the Tn7 transposon system is derived from Vibrio cholerae. In some embodiments, at least one Cas protein of the CRISPR-cas system is derived from a type V CRISPR-cas system. In some embodiments, at least one Cas protein is C2c5. In some embodiments, at least one Cas protein of the CRISPR-Cas system is derived from a type II-A CRISPR-Cas system. In some embodiments, at least one Cas protein is Cas9. In some embodiments, at least one Cas protein comprises Cas2, Cas3, Cas5, Cas6, Cas7, and Cas8. In some embodiments, at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8, and the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) TniQ. In some embodiments, at least one of the following applies: I) Cas5 and Cas8 form a Cas5-Cas8 fusion protein; II) TniQ and Cas6 form a TniQ-Cas6 fusion protein; and / or III) TnsA and TnsB form a TnsA-TnsB fusion protein.
[0082] In some embodiments, the first transposon end sequence is a left transposon end sequence, and the second transposon end sequence is a right transposon end sequence. In some embodiments, the left transposon end sequence and / or the right transposon end sequence are variant sequences that increase the efficiency of integration of the donor nucleic acid sequence compared to the corresponding wild-type left transposon end sequence and / or right transposon end sequence. In some embodiments, the left transposon end sequence and / or the right transposon end sequence changes the directional bias of the donor nucleic acid sequence when integrated proximal to a target site in a genome compared to the corresponding wild-type left transposon end sequence and / or right transposon end sequence. In some embodiments, the directional bias is imparted to a tRL. In some embodiments, the directional bias is imparted to a tLR.
[0083] In some embodiments, the first transposon end sequence and / or the second transposon end sequence encode a functional protein linker sequence. In some embodiments, the engineered transposon-encoded CRISPR-Cas system is derived from a bacterium selected from the group consisting of Vibrio cholerae, Photobacterium iliopiscarium, Pseudoalteromonas sp. P1-25, Pseudoalteromonas ruthenica, Photobacterium ganghwense, Shewanella sp. UCD-KL21, Vibrio diazotrophicus, Vibrio sp. 16, Vibrio sp. F12, Vibrio splendidus, Aliivibrio wodanis, and Parashewanella spongiae. In some embodiments, the engineered transposon-encoded CRISPR-Cas system is derived from a bacterium selected from the group consisting of Vibrio cholerae strain 4874, Photobacterium iliopiscarium strain NCIMB, Pseudoalteromonas sp. P1-25, Pseudoalteromonas ruthenica strain S3245, Photobacterium ganghwense strain JCM, Shewanella sp. UCD-KL21, Vibrio cholerae strain OYP7G04, Vibrio cholerae strain M1517, Vibrio diazotrophicus strain 60.6F, Vibrio sp. 16, Vibrio sp. F12, Vibrio splendidus strain UCD-SED10, Aliivibrio wodanis 06 / 09 / 160, and Parashewanella spongiae strain HJ039. In some embodiments, the engineered transposon-encoded CRISPR-Cas system is derived from Scytonema hofmannii PCC 7110.
[0084] In some embodiments, provided herein are methods of administering the compositions described above and herein to a subject (e.g., a human). In some embodiments, provided herein are methods of contacting a cell (e.g., a human cell) in vitro with the compositions described above and herein. In some embodiments, the engineered CRISPR-cas system and engineered transposon system are derived from a type I CRISPR-cas system and transposon system, and the system further comprises a second engineered CRISPR-cas system and a second engineered transposon system, both of which are derived from a type V CRISPR-cas system and transposon system.
[0085] In some embodiments, provided herein are methods of treating cells, the methods comprising: a) contacting at least one cell with a composition comprising i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more nucleic acid sequence(s) encoding the engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) at least one guide RNA (gRNA) specific to a target site within the genome of the at least cell; iii) the engineered transposon system; and iv) a donor nucleic acid comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence, wherein the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid proximal to the target site within the genome of the at least one cell.
[0086]
[0010] In some embodiments, provided herein are methods of treating a cell, the method comprising: a) contacting at least one cell with a composition comprising: i) an engineered transposon-encoded CRISPR-Cas system; and / or ii) one or more nucleic acid sequence(s) encoding the engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) the engineered transposon system; and iii) a donor nucleic acid sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence, and at least a portion of the cargo nucleic acid sequence encodes at least one guide RNA (gRNA) specific for a target site within the genome of the cell; and wherein the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid proximal to the target site within the genome of the at least one cell.
[0087] In some embodiments, provided herein are methods of treating a cell, the method comprising: a) contacting at least one cell with i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) a composition comprising one or more nucleic acid sequence(s) encoding the engineered transposon-encoded CRISPR-Cas system, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) at least one guide RNA (gRNA) specific for a target site; and iii) a nucleic acid sequence comprising: A) TnsA, B) TnsB, C) TnsC, and D) a nucleic acid sequence encoding the engineered transposon-encoded CRISPR-Cas system. D) an engineered transposon system comprising a TniQ family protein, wherein TnsA comprises one or more inactivating point mutations; and iv) a donor nucleic acid sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence, wherein the transposon-encoded CRISPR-Cas system integrates a copy of the donor nucleic acid proximal to a target site in the genome of at least one cell using a copy-and-paste transposition pathway involving replicative transposition.
[0088] In some embodiments, provided herein are methods of treating cells, the methods comprising: a) contacting at least one cell with a composition comprising i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more nucleic acid sequence(s) encoding an engineered transposon-encoded CRISPR-Cas system, wherein the first transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) a first RNA (gRNA) specific for a first target site; iii) the engineered transposon system; and iv) a first donor nucleic acid sequence comprising a first cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the first cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence. and a first donor nucleic acid sequence encoding a CRISPR-Cas system that encodes a second transposon comprising: i) at least one Cas protein; ii) a second RNA (gRNA) specific for a second target site; iii) an engineered transposon system; and iv) a second donor nucleic acid sequence comprising a second cargo nucleic acid sequence and a third transposon end sequence and a fourth transposon end sequence, wherein the second cargo nucleic acid sequence is flanked by the third transposon end sequence and the fourth transposon end sequence; wherein the first transposon-encoded CRISPR-Cas system integrates the first donor nucleic acid proximal to the first target site in the at least one cell, and the second transposon-encoded CRISPR-Cas system integrates the second donor nucleic acid proximal to the second target site in the at least one cell.
[0089] In some embodiments, provided herein are methods of a) contacting a sample with i) an engineered transposon-encoded CRISPR-Cas system, and / or ii) one or more vectors encoding an engineered transposon-encoded CRISPR-Cas system, wherein the sample comprises an input nucleic acid sequence comprising: A) a double-stranded nucleic acid sequence of interest (NASI), B) a double-stranded first flanking region on one side of the NASI, and C) a double-stranded second flanking region on the other side of the NASI, wherein the transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) the engineered transposon system; iii) a first left transposon end sequence; iv) a first right transposon end sequence that is not covalently linked to the first left transposon end sequence; v) a second left transposon end sequence; vi) a second right transposon end sequence that is not covalently linked to the second left transposon end sequence; and vii) a first left transposon end sequence that is not covalently linked to the second left transposon end sequence. and (iii) a second guide RNA (gRNA-2) that targets a second left transposon end sequence and a second right transposon end sequence to a second flanking region; and (ix) a third guide RNA (gRNA-3); and b) incubating the sample under conditions such that i) the first left transposon end sequence and the first right transposon end sequence are integrated into the first flanking region, and ii) the second left transposon end sequence and the second right transposon end sequence are integrated into the second flanking region, thereby generating a transposable sequence comprising a NASI flanked by the first transposon end sequence and the second transposon end sequence; and iii) the transposable sequence is cleaved from its location in the genome by the engineered transposon system and pasted at a different location in the genome guided by gRNA-3.
[0090] In some embodiments, provided herein are methods of treating cells, the methods comprising: a) contacting at least one cell with a composition comprising: i) one or more nucleic acid sequence(s) encoding first and second engineered transposon-encoded CRISPR-Cas systems, and / or a first engineered transposon-encoded CRISPR-Cas system and a second engineered transposon-encoded CRISPR-Cas system, wherein the first transposon-encoded CRISPR-Cas system comprises: i) at least one Cas protein; ii) a first RNA (gRNA) specific for a first target site within the genome of the cell; iii) an engineered transposon system; and iv) a first donor nucleic acid sequence comprising a first cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence, wherein the first cargo nucleic acid sequence is flanked by the first transposon end sequence and the second transposon end sequence; a) contacting the cell with an ISPR-Cas system comprising: i) at least one Cas protein; ii) a second RNA (gRNA) specific for a second target site within the genome of the cell; iii) an engineered transposon system; and iv) a second donor nucleic acid sequence comprising a second cargo nucleic acid sequence and a third transposon end sequence and a fourth transposon end sequence, wherein the second cargo nucleic acid sequence is flanked by the third transposon end sequence and the fourth transposon end sequence; and b) contacting the cell with: i) at least one Cas protein; ii) a second RNA (gRNA) specific for a second target site within the genome of the cell; iii) an engineered transposon system; and iv) a second donor nucleic acid sequence comprising a second cargo nucleic acid sequence and a third transposon end sequence and a fourth transposon end sequence, wherein the second cargo nucleic acid sequence is flanked by the third transposon end sequence and the fourth transposon end sequence; a CRISPR-Cas system encoded by a first transposon integrates a first donor nucleic acid proximal to a first target site in the genome of the at least one cell; and ii) a CRISPR-Cas system encoded by a second transposon integrates a second donor nucleic acid proximal to a second target site in the genome of the at least one cell, thereby generating a transposable sequence comprising a first transposon end sequence, a fourth transposon end sequence, and a region of the genome between the first transposon end sequence and the fourth transposon end sequence;iii) incubating under conditions such that the engineered transposon system excises the transposable sequence from its location in the genome and pastes it into a different location in the genome;
[0091] In some embodiments, the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) a TniQ family protein. In some embodiments, the at least one guide RNA comprises at least two different gRNAs, each of which directs the donor nucleic acid to integrate proximal to a different target site. In certain embodiments, the at least one guide RNA comprises at least 10 different gRNAs, each of which directs the donor nucleic acid to integrate at a different target site.
[0092] In some embodiments, the first transposon end sequence is a left transposon end sequence, and the second transposon end sequence is a right transposon end sequence. In some embodiments, the left transposon end sequence and / or the right transposon end sequence are variant sequences that increase the efficiency of integration of the donor nucleic acid sequence compared to the corresponding wild-type left transposon end sequence and / or right transposon end sequence. In some embodiments, the left transposon end sequence and / or the right transposon end sequence changes the directional bias of the donor nucleic acid sequence when integrated proximal to a target site in a genome compared to the corresponding wild-type left transposon end sequence and / or right transposon end sequence. In some embodiments, the directional bias is imparted to a tRL. In some embodiments, the directional bias is imparted to a tLR.
[0093] In some embodiments, the first transposon end sequence and / or the second transposon end sequence encode a functional protein linker sequence. In some embodiments, the genome of the subject cell comprises a target protein-encoding gene, wherein the cargo nucleic acid sequence encodes an amino acid sequence of interest, and the donor nucleic acid sequence is inserted adjacent to or within the target protein-encoding gene to generate a fusion protein-encoding sequence, wherein the fusion protein comprises the amino acid sequence of interest appended to the target protein. In some embodiments, the amino acid sequence of interest is selected from the group consisting of a fluorescent protein, an epitope tag, and a degron tag. In some embodiments, the genome of the cell comprises a target protein-encoding gene, wherein the cargo nucleic acid sequence comprises i) an amino acid sequence-encoding region of interest (AASIER), and ii) a splice acceptor and / or donor site flanking the AASIER, and the donor nucleic acid sequence is inserted adjacent to or within the target protein-encoding gene to generate a synthetic, engineered exon that allows in-frame tagging of the target protein with the amino acid sequence of interest.
[0094] In some embodiments, the at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. In some embodiments, the type I CRISPR-cas system is an IF type variant. In some embodiments, the IF type variant is derived from a bacterium selected from the group consisting of Vibrio cholerae, Photobacterium iliopiscarium, Pseudoalteromonas sp. P1-25, Pseudoalteromonas ruthenica, Photobacterium ganghwense, Shewanella sp. UCD-KL21, Vibrio diazotrophicus, Vibrio sp. 16, Vibrio sp. F12, Vibrio splendidus, Aliivibrio wodanis, and Parashewanella spongiae. In certain embodiments, the IF-type variant is derived from a bacterium selected from the group consisting of Vibrio cholerae strain 4874, Photobacterium iliopiscarium strain NCIMB, Pseudoalteromonas species P1-25, Pseudoalteromonas ruthenica strain S3245, Photobacterium ganghwense strain JCM, Shewanella species UCD-KL21, Vibrio cholerae strain OYP7G04, Vibrio cholerae strain M1517, Vibrio diazotrophicus strain 60.6F, Vibrio species 16, Vibrio species F12, Vibrio splendidus strain UCD-SED10, Aliivibrio wodanis 06 / 09 / 160, and Parashewanella spongiae strain HJ039. In some embodiments, the IF-type variant is derived from Vibrio cholerae strain HE-45.
[0095] In some embodiments, at least one Cas protein of the CRISPR-cas system is derived from a Type V CRISPR-cas system. In some embodiments, the Type V CRISPR-Cas system is derived from Scytonema hofmannii PCC 7110.
[0096] In some embodiments, the transposon-encoded CRISPR-Cas system integrates the donor nucleic acid sequence using a cut-and-paste transposition pathway. In some embodiments, at least one gRNA comprises an extended length guide sequence that targets an extended length target site, where the extended length guide sequence is at least 25 nucleotides in length (e.g., 25, 30, 40, 50, or more). In some embodiments, at least one gRNA comprises an extended length guide sequence.
[0097] In some embodiments, the engineered transposon system comprises i) TnsA, ii) TnsB, iii) TnsC, and iv) a TniQ family protein. In some embodiments, TnsA and TnsB are fused into a single TnsA-TnsB fusion polypeptide. In some embodiments, TniQ is fused to at least one Cas protein to generate a TniQ-Cas fusion polypeptide.
[0098] In some embodiments, the cargo nucleic acid sequence comprises an element selected from the group consisting of a natural transcription promoter, a synthetic transcription promoter, an inducible transcription promoter, a constitutive transcription promoter, a natural transcription terminator, a synthetic transcription terminator, an origin of replication, a replication termination sequence, a centromere sequence, and a telomere sequence. In some embodiments, the cargo nucleic acid sequence encodes at least one of a therapeutic protein, a metabolic pathway, and / or a biosynthetic pathway.
[0099] In some embodiments, provided herein is a system for RNA-guided DNA integration, the system comprising a vector (or other nucleic acid sequence) comprising, in 5' to 3' order: a) a nucleic acid encoding one or more transposon system proteins; b) a nucleic acid encoding a guide RNA; and c) a nucleic acid encoding a donor nucleic acid comprising a first transposon end and a second transposon end and a cargo nucleic acid.
[0100] In some embodiments, the nucleic acid encoding the guide RNA is proximal to the end of the first transposon, thereby preventing proximal self-targeting of the guide RNA. In some embodiments, the nucleic acid encoding the guide RNA is proximal to the donor nucleic acid, thereby preventing proximal self-targeting of the guide RNA.
[0101] In some embodiments, the nucleic acid encoding the guide RNA is within 10,000 bases of the end of the first transposon (e.g., within 10,000, 5000, 2000, 1000, 500, 200, 100, 50, 20, 10 bases of the end of the first transposon). In some embodiments, the nucleic acid encoding the guide RNA is within 1000 or 500 bases of the end of the first transposon.
[0102] In some embodiments, the transposon system proteins include one or more of TnsA, TnsB, TnsC, and TnsD and / or TniQ. In some embodiments, the vector further includes a nucleic acid expressing one or more cas proteins located between the nucleic acid encoding the one or more transposon system proteins and the nucleic acid encoding the donor. In some embodiments, the one or more Cas proteins include Cas5, Cas6, Cas7, and Cas8, or c2C5.
[0103] In some embodiments, provided herein are methods for reducing self-targeting of an RNA-guided DNA integration system, comprising expressing the above-described vector (or other nucleic acid sequence) in a cell, in some embodiments, the cell is of a cell type whose fitness is affected by maintenance of the vector. [Brief explanation of the drawings]
[0104] [Figure 1A] Figure 1A shows RNA-guided DNA integration by V. cholerae transposons. Figure 1A is an exemplary scenario of Tn6677 transposition into a plasmid or genomic target site complementary to the gRNA. [Figure 1B] Figure 1B shows RNA-guided DNA integration by a V. cholerae transposon. Figure 1B is a schematic diagram of an exemplary plasmid in a transposition experiment where the transposon is transmobilized. The CRISPR array contains two repeats (gray diamonds) and one spacer (maroon rectangle). [Figure 1C] Figure 1C shows RNA-guided DNA integration by V. cholerae transposons. The genomic loci targeted by gRNA-1 and gRNA-2, the two potential transposition products, and PCR primer pairs for selectively amplifying them. [Figure 1D] Figure 1D shows RNA-guided DNA integration by V. cholerae transposons. PCR analysis of transposition by non-targeting (nt) gRNA and gRNA-1, separated by agarose gel electrophoresis. [Figure 1E] Figure 1E shows RNA-guided DNA integration by V. cholerae transposons. PCR analysis of transposition by gRNA-nt, gRNA-1, and gRNA-2 using four different primer pairs, separated by agarose gel electrophoresis. [Figure 1F]RNA-guided DNA integration by V. cholerae transposons is shown. Figure 1F shows Sanger sequencing chromatograms of the upstream and downstream junctions of genome-integrated transposons from experiments using gRNA-1 and gRNA-2. Overlapping peaks for gRNA-2 suggest the presence of multiple integration sites. The distance between the 3' end of the protospacer and the first base of the transposon sequence is designated "d." TSD: target site overlap. [Figure 1G] Figure 1G shows RNA-guided DNA integration by V. cholerae transposons. Next-generation sequencing (NGS) analysis of the distance between the cascade target site and the transposon integration site was quantified for gRNA-1 and gRNA-2 with four primer pairs. [Figure 1H] Figure 1H shows RNA-guided DNA integration by V. cholerae transposons. Genomic loci targeted by gRNA-3 and gRNA-4. [Figure 1I] Figure 1I shows RNA-guided DNA integration by V. cholerae transposons. PCR analysis of transposition by gRNA-nt, gRNA-3, and gRNA-4, separated by agarose gel electrophoresis. [Figure 2A] Figure 2A shows that TniQ forms a complex with Cascade and is used for RNA-guided DNA integration. PCR analysis of transposition by gRNA-4 and a panel of gene deletions or point mutations were separated by agarose gel electrophoresis. [Figure 2B] This indicates that TniQ forms a complex with Cascade and is used for RNA-guided DNA integration. Figure 2B shows SDS-PAGE analysis of purified TniQ, Cascade, and the TniQ-Cascade co-complex. * indicates HptG contamination. [Figure 2C] Figure 2C shows that TniQ forms a complex with Cascade and is used for RNA-guided DNA integration. [Figure 2D]This demonstrates that TniQ forms a complex with Cascade and is used for RNA-guided DNA integration. Figure 2D shows RNA sequencing analysis of RNA co-purification with Cascade (top). Reads mapping to the CRISPR array reveal the mature gRNA sequence (sequence number 1655, bottom). [Figure 2E] This demonstrates that TniQ forms a complex with the Cascade and is used for RNA-guided DNA integration. Figure 2E shows a PCR analysis (left) of a transposition experiment testing whether general R-loop formation or artificial TniQ tethering can direct targeted integration. V. cholerae transposons and TnsA-TnsB-TnsC were combined with DNA targeting components containing either the V. cholerae Cascade (Vch), the P. aeruginosa Cascade (Pae), or the S. pyogenes dCas9-RNA (dCas9). TniQ was expressed either alone from pTnsABCQ or as a fusion to the targeting complex (pCas-Q) at the Cas6 C-terminus (6), the Cas8 N-terminus (8), or the dCas9 N-terminus (N) or C-terminus (C). A schematic diagram (right) shows some of the embodiments being tested. [Figure 2F] This demonstrates that TniQ forms a complex with Cascade and is used for RNA-guided DNA integration. Figure 2F shows a schematic diagram of the R-loop formed upon binding of target DNA by Cascade, with the approximate locations of each protein subunit indicated. The distance to the putative TniQ binding site and primary integration site is indicated. [Figure 3A] Figure 3A shows the effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration. Figure 3A is a schematic diagram of alternative integration directions and primer pairs for their selective detection by qPCR. [Figure 3B] Figure 3B shows the effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration. Figure 3B shows qPCR-based quantification of the transposition efficiency in both directions by gRNA-nt, gRNA-3, and gRNA-4. [Figure 3C] The effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration are shown. Figure 3C shows the overall integration efficiency of gRNA-4 as a function of transposon size. The arrow indicates the "wild-type" pDonor used in most assays throughout this study. [Figure 3D] The effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration are shown. Figure 3D shows a schematic of gRNAs overlaid along the lacZ gene at 1-bp increments relative to gRNA-4 (4.0) (top) and quantification by qPCR of the resulting integration efficiencies (bottom). Data are normalized to gRNA-4.0, and the two-nucleotide PAM for each gRNA is shown. [Figure 3E] Figure 3E shows the effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration. Figure 3E is a heatmap showing the integration site distribution (x-axis) for each of the tiled gRNAs (y-axis) in Figure 3D, quantified by NGS. The 49-bp distance for each gRNA is indicated by a black box. [Figure 3F] The effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration are shown in Figure 3F. A schematic of gRNA mutations in 4-nt blocks to introduce gRNA-target DNA mismatches (top) and the resulting integration efficiencies quantified by qPCR (bottom) are shown. Data are normalized to gRNA-4. [Figure 3G] The effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration are shown. Figure 3G shows the gRNA-4 spacer length shortened or lengthened by 12 nt (top) and the resulting integration efficiencies quantified by qPCR (bottom). Data were normalized to gRNA-4. The inset shows a comparison of the integration site distributions of gRNA-4 and gRNA-4+12 as quantified by NGS. [Figure 3H]The effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration are shown. Figure 3H shows another example of the overall integration efficiency by gRNA-4 as a function of transposon cargo size. The listed size includes the cargo and transposon ends, and the arrow indicates the original pDonor. [Figure 3I] The effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration are shown. Figure 3I shows a third example of the overall integration efficiency by gRNA-4 as a function of transposon cargo size. The listed sizes do not include the left and right end sequences. [Figure 3J] Figure 3J shows the effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration. Figure 3J compares the distribution of integration sites in gRNA-4 and gRNA-4(mm29-32). [Figure 3K] The effects of cargo size, PAM sequence, and gRNA mismatch on RNA-guided DNA integration are shown. Figure 3K shows the results after shortening or lengthening the gRNA-4 spacer length in 6-nt increments, and the resulting integration efficiencies quantified by qPCR (left). Data are normalized to gRNA-4. A comparison of the integration site distributions of gRNA-4 and gRNA-4(+12nt) is shown on the right. Data in Figures 3B-3D, 3F, and 3G are shown as the mean ± SD of n = 3 biologically independent samples. [Figure 4A] Genome-wide analysis of programmable RNA-guided DNA integration. Figure 4A shows a schematic diagram of the genomic loci targeted by gRNAs 4 to 8 (top) and PCR analysis of transposition (bottom), separated by agarose gel electrophoresis. [Figure 4B] Genome-wide analysis of programmable RNA-guided DNA integration. Figure 4B is a schematic diagram of an exemplary Tn-seq workflow for deep sequencing of genome-wide transposition events. [Figure 4C]Genome-wide analysis of programmable RNA-guided DNA integration. Figure 4C shows mapped Tn-seq reads from transposition experiments using mariner transposons and V. cholerae transposons programmed with either gRNA-nt or gRNA-4. The target site of gRNA-4 is indicated by a maroon triangle. [Figure 4D] Genome-wide analysis of programmable RNA-guided DNA integration. Figure 4D shows the sequence logo of all Mariner Tn-seq reads, highlighting TA dinucleotide target site preference. [Figure 4E] Genome-wide analysis of programmable RNA-guided DNA integration. Figure 4E shows a comparison of integration site distribution in gRNA-4 determined by PCR amplicon sequencing and Tn-seq for T-RL products, with the distance between the cascade target site and the transposon integration site indicated. [Figure 4F] Genome-wide analysis of programmable RNA-guided DNA integration. Figure 4F shows a zoomed-in view of Tn-seq read coverage at the primary integration site for an experiment using gRNA-4, highlighting the 5-bp target site overlap (TSD), with the distance from the cascade target site indicated. [Figure 4G] Genome-wide analysis of programmable RNA-guided DNA integration. Figure 4G shows the genome-wide distribution of genome-mapping Tn-seq reads obtained from transposition experiments using gRNAs 9–16 against V. cholerae transposons. The location of each target site is indicated by a maroon triangle. [Figure 5A]This is a proposed model for RNA-guided DNA integration by Tn7-like transposons encoding CRISPR-Cas systems. The V. cholerae Tn6677 transposon encodes a programmable RNA-guided DNA-binding complex called Cascade, which forms a novel co-complex with TniQ. The TniQ-Cascade complex monitors cells for DNA target sites found on host chromosomes or mobile genetic elements. Upon target binding and R-loop formation, DNA-bound TniQ recruits the non-sequence-specific DNA-binding protein TnsC. Based on previous studies of E. coli Tn7, this can ultimately lead to the formation of a large megadalton-sized structure known as a transpososome, containing TniQ-Cascade-bound target DNA, TnsC, and TnsAB-bound transposon donor DNA. The transposon itself is bound to the left and right ends by TnsA and TnsB to form a so-called paired-end complex, which is then recruited to the target DNA by TnsC. Excision of the transposon from its donor site allows targeted integration at a fixed distance downstream of the DNA-binding TniQ cascade, resulting in a 5-bp target site duplication. [Figure 5B]This is a proposed model for RNA-guided DNA integration by Tn7-like transposons encoding CRISPR-Cas systems. The V. cholerae Tn6677 transposon encodes a programmable RNA-guided DNA-binding complex called Cascade, which forms a novel co-complex with TniQ. The TniQ-Cascade complex monitors the cell to adapt to DNA target sites that may be found on the host chromosome or mobile genetic elements. Upon target binding and R-loop formation, DNA-bound TniQ recruits the non-sequence-specific DNA-binding protein TnsC. Based on previous studies of E. coli Tn7, this can ultimately lead to the formation of a large megadalton-sized structure known as a transpososome, containing TniQ-Cascade-bound target DNA, TnsC, and TnsAB-bound transposon donor DNA. The transposon itself is bound to the left and right ends by TnsA and TnsB to form a so-called paired-end complex, which is then recruited to the target DNA by TnsC. Excision of the transposon from its donor site allows targeted integration at a fixed distance downstream of the DNA-binding TniQ cascade, resulting in a 5-bp target site duplication. [Figure 6A] Figure 6A shows the gene architecture of the E. coli Tn7 transposon transposition and the Tn6677 transposon from V. cholerae. Figure 6A shows the genomic organization of the native E. coli Tn7 transposon flanked by known attachment sites (attTn7) within the glmS gene. [Figure 6B] Figure 6B shows the gene architecture of the E. coli Tn7 transposon transposition and the Tn6677 transposon from V. cholerae. Figure 6B is a schematic diagram of exemplary expression and donor plasmids in a Tn7 transposition experiment. [Figure 6C]Figure 6C shows the gene architecture of the E. coli Tn7 transposon transposition and the Tn6677 transposon from V. cholerae. Figure 6C is a schematic diagram of the genomic locus containing the conserved TnsD binding site (attTn7), including predicted alternative orientations of Tn7 transposition products and PCR primer pairs for selectively amplifying them. [Figure 6D] Figure 6D shows the gene architecture of the E. coli Tn7 transposon and the Tn6677 transposon from V. cholerae. PCR analysis of Tn7 transposition, separated by agarose gel electrophoresis. Amplification of rssA serves as a loading control. [Figure 6E] Figure 6E shows the gene architecture of the E. coli Tn7 transposon and the Tn6677 transposon from V. cholerae. Figure 6E shows the Sanger sequencing chromatograms of both the upstream and downstream junctions of the genome-integrated Tn7. TSD: target site overlap. [Figure 6F]The gene architecture of the E. coli Tn7 transposon and the Tn6677 transposon from V. cholerae are shown. Figure 6F shows the genomic organization of the native V. cholerae strain HE-45 Tn6677 transposon. Highlighted genes are conserved between Tn6677 and the E. coli Tn7 transposon, as well as between Tn6677 and the canonical IF-type CRISPR-Cas system from Pseudomonas aeruginosa. The cas1 and cas2-3 genes, which mediate spacer acquisition and DNA degradation during the adaptation and interference phases of adaptive immunity, respectively, are missing from the CRISPR-Cas system encoded by Tn7-like transposons. Similarly, the tnsE gene, which facilitates sequence-nonspecific transposition, is also absent. The V. cholerae HE-45 genome contains another Tn7-like transposon (located within GenBank accession ALED01000025.1), which lacks an encoded CRISPR-Cas system and has low sequence similarity to the Tn6677 transposon examined in this study. [Figure 7A] Analysis of E. coli cultures and strain isolates harboring lacZ-integrated transposons. Figure 7A shows the genomic loci targeted by gRNA-3 and gRNA-4, including both potential transposition products and PCR primer pairs for their selective amplification (top). Next-generation sequencing (NGS) analysis quantified the distance between the Cascade target site and the transposon integration sites of gRNA-3 (left) and gRNA-4 (right) using two alternative primer pairs. [Figure 7B]Analysis of E. coli cultures and strain isolates harboring the lacZ integrated transposon. Figure 7B shows a schematic diagram of the lacZ locus with and without an integrated transposon after a transposition experiment using gRNA-4 (top). T-LR and T-RL indicate transposition products in which the left and right ends of the transposon are proximal to the target site, respectively. Primer pair g and h (external-internal) selectively amplify the integrated locus, while primer pair i (external-external) amplifies both the non-integrated and integrated loci. PCR analysis of 10 colonies after 24 h of growth on +IPTG plates (bottom left) shows that all colonies contain integration events in both orientations (primer pairs g and h), but the efficiency is low enough that non-integrated products predominate after amplification with primer pair i. After resuspending the cells, clonal growth continued for an additional 18 hours on −IPTG plates. The same PCR analysis was performed on 10 colonies (bottom right), revealing that 3 / 10 colonies showed clonal integration in the T-LR orientation (compare primer pairs h and i). The remaining colonies showed low levels of integration in both orientations, likely due to leaky expression, which occurred during the additional 18 hours of growth. These analyses demonstrate that colonies are genetically heterogeneous after growth on +IPTG plates and that RNA-guided DNA integration occurs in only a proportion of cells within the growing colony. I: integration product; U: non-integration product; *: mispriming product also present in the negative (non-integration) control. [Figure 7C] Analysis of E. coli cultures and strain isolates harboring lacZ integrative transposons. Figure 7C shows a photograph of an LB-agar plate used for blue-white colony screening. Cells from IPTG-containing plates were replated onto X-gal-containing plates, and white colonies predicted to harbor lacZ-inactivating transposon inserts were selected for further characterization. [Figure 7D]Analysis of E. coli cultures and strain isolates harboring lacZ integrative transposons. Figure 7D is a PCR analysis of E. coli strains harboring clonal integrative transposons identified by blue-white colony screening as shown in Figure 7B. [Figure 7E] Analysis of E. coli cultures and strain isolates harboring lacZ integrative transposons. Figure 7E is a schematic representation of Sanger sequencing coverage across the lacZ locus in the strains shown in Figure 7D. [Figure 7F] Analysis of E. coli cultures and strain isolates harboring the lacZ integration transposon. Figure 7F shows PCR analysis of transposition experiments using gRNA-4 after serial dilution of lysates from clonal integration strains with lysates from control strains to simulate variable integration efficiencies, as shown in Figure 7B. Transposition products can be reliably detected by PCR using the outside-inside primer pair with efficiencies greater than 0.5%, but PCR bias results in preferential amplification of non-integration products using the outside-outside primer pair with efficiencies substantially anywhere less than 100%. [Figure 7G] Analysis of E. coli cultures and strain isolates harboring lacZ-integrated transposons. Figure 7G shows a schematic diagram of the lacZ locus with and without integrated Tn7 (top) and colony PCR analysis of Tn7 transposition experiments with gRNA4 using primer pair a (middle) or primer pair b (bottom) (separated by agarose gel electrophoresis and shown in Figure 7B). [Figure 8A] Analysis of the V. cholerae cascade and TniQ-cascade complex. Figure 8A is a schematic diagram of an exemplary expression vector for recombinant protein or ribonucleoprotein complex purification. [Figure 8B]Analysis of V. cholerae cascade and TniQ-cascade complex. Figure 8B shows an SDS-PAGE analysis (left) of purified TniQ, cascade, and TniQ-cascade complex, highlighting the protein bands excised for in-gel trypsin digestion and mass spectrometry analysis. The table (right) shows the E. coli and recombinant proteins identified from these data, as well as the spectral counts of their associated peptides. Note that the cascade and TniQ-cascade samples used in this analysis are different from those shown in Figure 2. [Figure 8C] Analysis of the V. cholerae cascade and the TniQ-cascade complex. Figure 8C shows size-exclusion chromatography of the TniQ-cascade complex on a Superose 6 10 / 300 column (left) and a calibration curve generated using protein standards (right). The measured retention time of the TniQ-cascade (maroon) was consistent with a complex with a molecular weight of approximately 440 kDa. [Figure 8D] Analysis of the V. cholerae cascade and TniQ-cascade complex. Figure 8D shows the RNase A and DNase I sensitivity of nucleic acids co-purified with the cascade and TniQ-cascade and resolved by denaturing urea-PAGE. [Figure 8E] Analysis of V. cholerae Cascade and TniQ-Cascade complexes. Figure 8E shows results from TniQ, Cascade, and Cascade + TniQ combined reactions separated by size exclusion chromatography (left), and the indicated fractions were analyzed by SDS-PAGE (right). * indicates HptG contamination. [Figure 9]Control experiments demonstrating efficient DNA targeting by Cas9 and the P. aeruginosa cascade. (A) Schematic of exemplary plasmid expression systems for S. pyogenes Cas9-sgRNA (type II-A, left) and the P. aeruginosa cascade (PaeCascade) and Cas2-3 (type IF, right). The Cas2-3 expression plasmid was omitted from the experiment described in Figure 2E. (B) Graph of the results of cell killing experiments using S. pyogenes Cas9-sgRNA (left) or PaeCascade and Cas2-3 (right), monitored by determining colony-forming units (CFU) upon plasmid transformation. Complexes were programmed with gRNAs targeting the same genomic lacZ site as V. cholerae gRNA-3 and gRNA-4, resulting in lethality and reduced transformation efficiency through efficient DNA targeting and degradation. (C) Graph of qPCR-based quantification of transposition efficiency from experiments using V. cholerae transposon donors and TnsA-TnsB-TnsC with DNA targeting components, including either the V. cholerae cascade (Vch), P. aeruginosa cascade (Pae), or S. pyogenes dCas9 RNA. TniQ was expressed either alone or as a fusion to the targeting complex (pCas-Q) from pTnsABCQ with either the Cas6 C-terminus (6), the Cas8 N-terminus (8), or the dCas9 N-terminus (N) or C-terminus (C). The exact same sample lysates as in Figure 2E were used. Data in B and C are shown as the mean ± SD for n = 3 biologically independent samples. [Figure 10A] qPCR-based quantification of RNA-guided DNA integration efficiency. Figure 10A shows a schematic diagram of potential lacZ transposition products in either orientation for both gRNA-3 and gRNA-4, as well as qPCR primer pairs for selectively amplifying them. T-LR and T-RL indicate transposition products in which the left and right ends of the transposon are proximal to the target site, respectively. [Figure 10B]qPCR-based quantification of RNA-guided DNA integration efficiency. Figure 10B includes graphs comparing simulated integration efficiencies in the T-LR and T-RL directions, generated by mixing clonal integration and non-integration lysates in known ratios, with experimentally determined integration efficiencies measured by qPCR. [Figure 10C] qPCR-based quantification of RNA-guided DNA integration efficiency. Figure 10C is a graph comparing simulated mixtures of bidirectional integration efficiencies in gRNA-4, generated by mixing clonal integration and non-integration lysates in known ratios, with experimentally determined integration efficiencies measured by qPCR. [Figure 10D] qPCR-based quantification of RNA-guided DNA integration efficiency. Figure 10D is a graph of RNA-guided DNA integration efficiency as a function of IPTG concentration for gRNA-3 and gRNA-4, as measured by qPCR. [Figure 10E] qPCR-based quantification of RNA-guided DNA integration efficiency. Figure 10E shows a graph of bidirectional integration efficiency measured by qPCR for simulated mixtures of gRNA4 bidirectional integration efficiencies generated by mixing clonal integration and non-integration lysates in known ratios. Data in Figures 10B-10C are shown as the mean ± SD for n=3 biologically independent samples. [Figure 11A] The effect of transposon end sequences on RNA-guided DNA integration is shown. Figure 11A shows the sequence (top) and schematic diagram (bottom) of the left and right end sequences of V. cholerae Tn6677. Putative TnsB binding sites (blue) were determined based on sequence similarity to TnsB binding sites. 8-bp ends are shown in yellow, and the empirically determined minimal end sequences required for transposition are shown in a dashed red box. [Figure 11B]Figure 11B shows the effect of transposon end sequences on RNA-guided DNA integration. Figure 11B is a graph of integration efficiency by gRNA-4 as a function of transposon end length, as determined by qPCR. [Figure 11C] Figure 11C shows the effect of transposon end sequences on RNA-guided DNA integration. Figure 11C is a graph of the relative proportion of both integration orientations as a function of transposon end length, as determined by qPCR. ND: Not determined. [Figure 11D] Figure 11D shows the effect of transposon end sequences on RNA-guided DNA integration. Figure 11D shows a graph of gRNA-4 integration efficiency as a function of transposon end truncation (bottom), determined independently in both directions by qPCR. The empirically determined minimum required end sequences are shown as dashed boxes. Data in Figures 11B and 11C are shown as means ± SD for n = 3 biologically independent samples. [Figure 12A] Analysis of RNA-guided DNA integration in PAM-tiled gRNAs and gRNAs with extended spacer lengths. Figure 12A shows a graph of integration site distribution for all gRNAs listed in Figures 3D-3E with normalized transposition efficiencies greater than 20%, as determined by NGS. [Figure 12B] Analysis of RNA-guided DNA integration in PAM-tiled gRNAs and spacer-extended gRNAs. Figure 12B shows the distribution of integration sites in gRNAs containing mismatches at positions 29-32 compared to the distribution in gRNA-4, as determined by NGS. [Figure 12C] Analysis of RNA-guided DNA integration in PAM-tiled gRNAs and spacer-extended gRNAs. Figure 12C shows the resulting integration efficiency determined by qPCR after shortening or extending the gRNA-4 spacer length by 6 nt. Data are normalized to gRNA-4 and shown as the mean ± SD for n = 3 biologically independent samples. [Figure 12D]Analysis of RNA-guided DNA integration in PAM-tiled gRNAs and extended spacer length gRNAs. Figure 12D is a graph of integration site distribution in extended length gRNAs compared to the distribution with gRNA-4, as determined by NGS. [Figure 13A] Development and analysis of transposon insertion sequencing (Tn-seq) are shown. Figure 13A is a schematic diagram of the V. cholerae transposon end sequence. The 8-bp end sequence of the transposon is boxed and highlighted in light yellow. The mutations generated to introduce the MmeI recognition site are shown in red, and the resulting recognition site is highlighted in red. Cleavage by MmeI occurs 17–19 bp away from the transposon end, generating a 2-bp overhang. [Figure 13B] Development and analysis of transposon insertion sequencing (Tn-seq) are shown. Figure 13B is a graph comparing integration efficiencies in wild-type and MmeI-containing transposon Donors, as determined by qPCR. Labels on the x-axis indicate which plasmid was transformed last. Higher integration efficiencies were reproducibly observed when pQCascade was transformed last (gRNA-4) than when pDonor was transformed last. A transposon containing an MmeI site at the "right" transposon end (R*-L pDonor) was used for all Tn-seq experiments. Data are shown as the mean ± SD of n = 3 biologically independent samples. [Figure 13C] Figure 13C shows the development and analysis of transposon insertion sequencing (Tn-seq). Figure 13C is a schematic diagram of the plasmid expression system for Himar1C9 and mariner transposons. [Figure 13D] Development and analysis of transposon insertion sequencing (Tn-seq) are shown. Figure 13D is a scatter plot showing correlation between two biological replicates of a Tn-seq experiment and mariner transposons. Reads were binned by E. coli gene annotation, and linear regression fits and Pearson linear correlation coefficients (r) are shown. [Figure 13E] Development and analysis of transposon insertion sequencing (Tn-seq) are shown. Figure 13E shows a schematic of the 100-bp binning approach used for Tn-seq analysis of V. cholerae transposon-mediated transposition experiments, with bin-1 defined as the first 100 bp immediately downstream (PAM-distal) of the cascade target site. [Figure 13F] Development and analysis of transposon insertion sequencing (Tn-seq) are shown. Figure 13F is a scatter plot showing the correlation between biological replicates of a Tn-seq experiment and the V. cholerae transposon programmed with gRNA-4. Highly sampled reads all fell within bin-1, and low-level but reproducible widespread integration was also observed in 100-bp bins immediately upstream and downstream of the primary integration site (bins-1, 2, and 3). [Figure 13G] Figure 13G shows the development and analysis of transposon insertion sequencing (Tn-seq). Figure 13G is a scatter plot showing the correlation between biological replicates of Tn-seq experiments and V. cholerae transposons programmed with gRNA-nt. [Figure 13H] Development and analysis of transposon insertion sequencing (Tn-seq) are shown. Figure 13H is a scatter plot showing the correlation between biological replicates of a Tn-seq experiment and a V. cholerae transposon expressing TnsA-TnsB-TnsC-TniQ but not Cascade. In Figures 13F-13H, bins are plotted only if they contain at least one read in either dataset. [Figure 14A]Tn-seq data for additional gRNAs tested. Figure 14A shows the genome-wide distribution of genome-mapping Tn-seq reads from a transposition experiment using V. cholerae transposons programmed with gRNAs 1–8. The location of each target site is indicated by a maroon triangle. † The lacZ target site for gRNA-3 was found to overlap within the λDE3 prophage, as were the transposon integration sites. Tn-seq reads in this dataset were mapped to both genomic loci for visualization purposes only; their origins could not be determined. [Figure 14B] Figure 14B shows the genome-wide distribution of genome-mapping Tn-seq reads from a transposition experiment using V. cholerae transposons programmed with gRNAs 17-24. [Figure 14C] Figure 14C shows the Tn-seq data for additional gRNAs tested. Figure 14C shows a graph of the integration site distribution for gRNAs 1–8 determined from the Tn-seq data, showing the distance between the cascade target site and the transposon integration site. Data for both integration directions are overlaid, with the solid blue bar representing the T-RL direction and the dark outline representing the T-LR direction. The values in the upper right corner of each graph indicate the on-target specificity (%) (calculated as the percentage of reads resulting from integration within 100 bp of the primary integration site relative to the total number of reads aligning to the genome) and the directional bias (X:Y) (calculated as the ratio of T-RL reads to T-LR reads). The majority of gRNAs prefer integration in the T-RL direction, 49–50 bp downstream of the cascade target site. gRNA-21 is grayed out because its predicted primary integration site lies within a DNA repeat, preventing confident mapping of reads. * indicates samples where more than 1% of genome mapping reads could not be uniquely mapped are marked. [Figure 14D]Tn-seq data for additional gRNAs tested. Figure 14D shows a graph of the integration site distribution for gRNAs 9–16 determined from Tn-seq data, showing the distance between the cascade target site and the transposon integration site. Data for both integration directions are overlaid, with the solid blue bar representing the T-RL direction and the dark outline representing the T-LR direction. The values in the upper right corner of each graph indicate on-target specificity (%) (calculated as the percentage of reads resulting from integration within 100 bp of the primary integration site relative to the total number of reads aligning to the genome) and directional bias (X:Y) (calculated as the ratio of T-RL reads to T-LR reads). The majority of gRNAs prefer integration in the T-RL direction, 49–50 bp downstream of the cascade target site. gRNA-21 is grayed out because its predicted primary integration site lies within a DNA repeat, preventing confident mapping of reads. * indicates samples where more than 1% of genome mapping reads could not be uniquely mapped are marked. [Figure 14E]Tn-seq data for additional gRNAs tested. Figure 14E shows a graph of the integration site distribution for gRNAs 17–24 determined from Tn-seq data, showing the distance between the cascade target site and the transposon integration site. Data for both integration directions are overlaid, with the solid blue bar representing the T-RL direction and the dark outline representing the T-LR direction. The values in the upper right corner of each graph indicate on-target specificity (%) (calculated as the percentage of reads resulting from integration within 100 bp of the primary integration site relative to the total number of reads aligning to the genome) and directional bias (X:Y) (calculated as the ratio of T-RL reads to T-LR reads). The majority of gRNAs prefer integration in the T-RL direction, 49–50 bp downstream of the cascade target site. gRNA-21 is grayed out because the predicted primary integration site is within a repeat of the DNA, preventing confident mapping of reads. * indicates samples where more than 1% of genome mapping reads could not be uniquely mapped. [Figure 15]This figure shows that bacterial transposons also harbor a V-U5-type CRISPR-Cas system encoding C2c5. Representative genomic loci from various bacterial species include identifiable transposon ends (blue boxes, L and R), genes with homology to tnsB-tnsC-tniQ (shades of yellow), a CRISPR array (maroon), and the CRISPR-associated gene c2c5 (blue). The H. byssoidea example (top) highlights the target site duplication and terminal repeats, as well as genes found within the transposon cargo portion. Similar to Tn7-like transposons containing type I CRISPR-Cas systems, transposons containing type V CRISPR-Cas systems appear to preferentially harbor genes associated with innate immune system functions, such as restriction-modification systems. The C2c5 gene is often flanked by the predicted transcriptional regulator merR (light blue), and C2c5-containing transposons typically appear to be immediately upstream of tRNA genes (green). This phenomenon has also been observed for other prokaryotic integration elements: analysis of 50 spacers from eight CRISPR arrays represented by CRISPRTarget revealed six spacers with imperfectly matched targets (average of six mismatches), none of which mapped to bacteriophages, plasmids, or the same bacterial genome harboring the transposon itself. [Figure 16]Exemplary schematics of transposition via cut-and-paste and copy-and-paste mechanisms are shown. A is a schematic of cut-and-paste transposition. The E. coli Tn7 transposon mobilizes via the cut-and-paste mechanism. TnsA and TnsB cleave both strands of the transposon DNA at both ends, cleanly excising a linear dsDNA containing short 3-nucleotide 5'-overhangs at both ends (not shown). TnsB then attacks phosphodiester bonds on both strands of the target DNA using its free 3'-OH terminus as a nucleophile, resulting in a concerted transesterification reaction. After gap-filling, the transposition reaction is complete, and the integrated transposon is flanked on both ends by 5-bp target site overlaps (TSDs) as a result of the gap-filling reaction. B is a schematic of copy-and-paste (replicative) transposition. Some transposons alternatively mobilize via the copy-and-paste pathway, also known as replicative transposition. This occurs when the 5' end of the transposon donor DNA is not cleaved during the excision step, as occurs when the tnsA endonuclease gene is absent from the gene operon encoding the transposition protein. In this case, the 3'-OH end remains free and can participate in a staggered transesterification reaction with the target DNA catalyzed by TnsB (inset, center right), while the 5' end of the transposon remains covalently linked to the rest of the DNA within the donor DNA molecule, which can be a genome or a plasmid vector. This copy-and-paste reaction results in what is known as a Shapiro intermediate (center), in which the entire donor DNA (including the transposon sequence itself and adjacent sequences) is ligated to the cleaved target DNA. This intermediate only degrades during subsequent DNA replication (bottom left), resulting in the so-called cointegrate product. This cointegrate contains two copies of the transposon itself (orange rectangle), flanked on one side by the TSD. Importantly, the cointegrate also contains the entire donor DNA molecule and the entire target DNA molecule. Thus, if the transposon is encoded on a plasmid vector, the entire vector will ligate into the target DNA during replication transfer.At some frequency, the cointegration product can be resolved into the product shown on the right by the action of dedicated resolvase proteins (e.g., the TniR protein in Tn5090 / Tn5053) or by endogenous homologous recombination due to the high homology between the two copies of the transposon itself in the cointegration product. Cointegration resolution results in a target DNA with a single transposon flanked by TSDs and a regenerated version of the donor DNA molecule. [Figure 17]Figure 1 shows a comparison of the transposable genes in transposons with type IF and type V CRISPR-Cas systems. A is a schematic diagram of Tn7 and Tn7-like transposons described in the literature. (Panel reproduced and adapted from Figure 9.1b in Peters et al., Mol Microbiol 93, 1084-1092 (2014)). B is a schematic diagram of a representative Tn7-like transposon harboring genes encoding the Cascade complex. The Tn6677 transposon from Vibrio cholerae, which mediates RNA-guided DNA insertion, is a member of this family. Note the similarity of the transposable genes found in Tn6677 and related transposons as well as Tn7. While the tnsA-tnsB-tnsC operon is maintained, a tnsD homolog known as tniQ is encoded within an operon encoding the Cas8-Cas7-Cas6 proteins that collectively form the RNA-guided TniQ cascade complex. The protein products of TnsA and TnsB mediate transposon excision, while TnsB mediates transposon integration into target DNA. Figure C shows a schematic diagram of a representative Tn7-like transposon harboring a V-type CRISPR-Cas system with a gene encoding Cas12k (also known as C2c5). These transposons contain the tnsB, tnsC, and tniQ genes but lack the tnsA gene. This indicates that these transposons do not encode the machinery required to mediate cut-and-paste transposition. Instead, these transposons may proceed via copy-and-paste replicative transposition, resulting in cointegration products rather than pure integration products. [Figure 18]This expression strategy involves individual vectors for each component. Each component required for RNA-guided DNA integration with the Vibrio cholerae-derived CRISPR-Tn7 system is encoded on a separate mammalian expression plasmid. The protein-encoding gene is human codon-optimized (hCO), cloned downstream of a CMV promoter, and contains an N-terminal nuclear localization signal (NLS). In other embodiments, the NLS can be introduced in tandem or at the C-terminus of the protein. The CRISPR array encoding the gRNA is designed as a repeat-spacer-repeat array cloned downstream of a human U6 (hU6) promoter and processed by Cas6. Specific spacer sequences (maroon) are selected to correspond to the desired DNA target site. In this embodiment, all eight plasmids are cotransfected to reconstitute TniQCascade and TnsABC in cells, which, together with pDonor, can mediate RNA-guided integration. [Figure 19]Figure 1 shows an exemplary expression strategy involving polycistronic vectors. pTnsABC_hCO encodes human codon-optimized versions of TnsA, TnsB, and TnsC, with the NLS and T2A peptides indicated. pQCascade_hCO encodes human codon-optimized versions of TniQ, Cas6, Cas7, and Cas8, with the CRISPR array encoding the gRNA. The promoters for both vectors are indicated. In other embodiments, the gene order is altered to optimize expression, and the position and identity of the NLS and T2A peptide are modified. The CRISPR array encoding the gRNA is designed as a repeat-spacer-repeat array cloned downstream of the human U6 (hU6) promoter and processed by Cas6. Specific spacer sequences (maroon) are selected to correspond to the desired DNA target site. In this embodiment, both plasmids are co-transfected to reconstitute the TniQ-cascade and TnsABC in cells, which, together with pDonor, can mediate RNA-guided integration. The pQCascade_hCO variant (pSL1079) encodes a gRNA targeting a lacZ-specific sequence from E. coli, which in one embodiment is cloned into pTarget for RNA-guided DNA integration experiments in eukaryotic cells. [Figure 20] Possible delivery approaches are shown. A shows one embodiment, in which HEK293T cells are transfected with vectors encoding the respective proteins and RNA machinery, recapitulating RNA-guided DNA integration. B shows another embodiment, in which 5'-capped (red circle) and 3'-polyadenylated mRNA is synthesized along with precursor gRNA (shown) or fully processed mature gRNA (not shown), and HEK293T cells are then transfected with a mixture of mRNA and gRNA. C shows another embodiment, in which all necessary protein and RNA components are recombinantly purified, and HEK293T cells are then transfected with the purified protein and ribonucleoprotein components. The above strategy can be combined with delivery of donor DNA (e.g., as in the case of pDonor). [Figure 21]Figure 1 shows an exemplary experimental strategy for RNA-guided DNA integration in HEK293T cells. Figure 1A is a schematic diagram of one embodiment, in which HEK293T cells are co-transfected with a CRISPR-Tn7 expression vector along with both pDonor and pTarget. pDonor contains a mini-transposon construct with Tn7 transposon ends ("L" and "R") flanking the gene cargo of interest, and pTarget contains a target site (maroon) complementary to the gRNA spacer. Successful RNA-guided DNA integration involves excision of the transposon from pDonor (mediated by TnsA and TnsB) followed by RNA-guided integration of the transposon into pTarget at a fixed distance from the target site. pDonor and pTarget may contain fluorescent reporter genes and / or drug resistance markers to allow for selection of cells that undergo integration events. Figure B is a schematic diagram of another embodiment, where the transposon is again encoded on pDonor, but the gRNA is designed to target RNA-guided DNA integration to a site within the human genome (diagrammed by the chromosome in red). This results in genomic integration of the transposon at a fixed distance from the target site (maroon). The sequences of the plasmids represent only one possible design for each plasmid. pTarget_Int refers to the integration product after RNA-guided DNA integration into pTarget. The integrated transposon can be detected and further analyzed by PCR, qPCR, and / or next-generation sequencing. [Figure 22]Figure 1 shows an exemplary experimental strategy for selecting and / or detecting RNA-guided DNA integration in HEK293T cells. Figure 1A is a schematic diagram of one embodiment, referred to as the promoter capture approach. HEK293T cells are co-transfected with a CRISPR-Tn7 expression vector along with pDonor. This pDonor contains a mini-transposon construct with Tn7 transposon ends ("L" and "R") flanking a gene cargo containing a puromycin resistance gene (puroR) connected to the EGFP gene via a 2A peptide. Because the gene cargo does not contain a promoter element, it will not be expressed unless RNA-guided DNA integration places the cargo downstream of a eukaryotic promoter element. The targeted promoter can be within a plasmid (e.g., pTarget) or the genome. Once integrated, a reporter gene is activated, and integration can be detected via flow cytometry and / or drug selection. pA refers to the polyadenylation signal, and the promoter (black arrow) can be a CMV promoter or other constitutive or inducible promoter. (B) Schematic representation of a target site chosen such that integration also disrupts another fluorescent reporter gene encoding mCherry. In this experimental setup, RNA-guided DNA integration results in both an increase in GFP signal and a loss of mCherry signal. (C) Schematic representation of another embodiment, in which the reporter in pDonor also contains promoter elements within the gene cargo, such that the pDonor plasmid itself expresses EGFP and a puromycin resistance gene. In this scenario, integration of the gene cargo into the genome or pTarget plasmid results in expression, regardless of whether promoter elements are present adjacent to the integration site. [Figure 23]Illustrative expression construct designs for reducing promoter number. A is a schematic diagram of the previously described pQCascade plasmid (pSL0828 encoding gRNA-4), which contains two separate T7 promoters, one driving CRISPR RNA expression and the other driving the TniQ-Cas8-Cas7-Cas6 operon. B is a schematic diagram of engineered pQCascade-B and pQCascade-C, which contain only a single T7 promoter driving both CRISPR RNA and the TniQ-Cas8-Cas7-Cas6 operon expression. The CRISPR array is positioned at either the 5' or 3' end of the transcript. C is a schematic diagram of an RNA-guided DNA integration experiment utilizing pDonor (pSL0527), which contains a gene cargo flanked by both Tn7 transposon ends, and pTnsABC, which encodes the TnsA-TnsB-TnsC operon. D. Results of RNA-guided DNA integration experiments performed in E. coli BL21(DE3) cells and quantified by qPCR. The overall integration efficiency is plotted for experiments utilizing pDonor (pSL0527), pTnsABC (pSL0283), and either pQCascade-B (pSL1016) or pQCascade-C (pSL1018). [Figure 24A] Figure 24A is a schematic diagram of pTQC-A (pSL1020), an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery from a single plasmid. Figure 24A is a schematic diagram of pTQC-A (pSL1020), which encodes the CRISPR array and the TniQ-Cas8-Cas7-Cas6-TnsA-TnsA-TnsB operon from two T7 promoters. [Figure 24B] Figure 24B is a schematic diagram of pTQC-B (pSL1022), an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery from a single plasmid. Figure 24B is a schematic diagram of pTQC-B (pSL1022), encoding the CRISPR array and the TniQ-Cas8-Cas7-Cas6-TnsA-TnsA-TnsB operon from a single T7 promoter. [Figure 24C]Figure 24C is a schematic diagram of pTQC-C (pSL1024), an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery from a single plasmid, encoding the TnsA-TnsB-TnsC operon and the TniQ-Cas8-Cas7-Cas6-CRISPR operon from two T7 promoters. [Figure 24D] Figure 24D is a schematic diagram of pTQC-D (pSL1026), an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery from a single plasmid, encoding the TnsA-TnsB-TnsC-TniQ-Cas8 / Cas5 fusion protein-Cas7-Cas6-CRISPR operon from a single T7 promoter. [Figure 24E] Figure 24E shows an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery from a single plasmid. Figure 24E is a schematic diagram of the fusion mRNA and CRISPR RNA transcripts encoded by pTQC-B (left) and pTQC-D (right). Enzymatic CRISPR RNA processing by Cas6 liberates the mature gRNA without interfering with the remaining mRNA transcript, which encodes all protein components. [Figure 24F] Figure 24F shows an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery from a single plasmid. Figure 24F shows the results of RNA-guided DNA integration experiments performed in E. coli BL21(DE3) cells and quantified by qPCR. The overall integration efficiency in experiments utilizing pDonor(pSL0527) and either pTQC-A, pTQC-B, pTQC-C, or pTQC-D is plotted as shown. [Figure 25A]Figure 1 shows an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery and a mini-transposon donor from a single plasmid. (A) Schematic diagram of pAIO-A(pSL1120), encoding the CRISPR array and the TniQ-Cas8-Cas7-Cas6-TnsA-TnsA-TnsB operon from a single T7 promoter, with downstream mini-transposon donor DNA containing the Tn7 transposon ends ("L" and "R") flanking the cargo of interest. (B) Schematic diagram of pAIO-A(pSL1120), encoding the CRISPR array and the TniQ-Cas8-Cas7-Cas6-TnsA-TnsA-TnsB operon from a single T7 promoter. The entire expression cassette is cloned into the mini-transposon donor DNA containing the Tn7 transposon ends ("L" and "R"). RNA-guided DNA integration using this construct mobilizes genetic components encoding CRISPR and Tn7-related machinery within the donor DNA itself. [Figure 25B] Figure 1 shows an exemplary expression construct design for expressing all CRISPR and Tn7-related machinery and a mini-transposon donor from a single plasmid. (A) Schematic diagram of pAIO-A(pSL1120), encoding the CRISPR array and the TniQ-Cas8-Cas7-Cas6-TnsA-TnsA-TnsB operon from a single T7 promoter, with downstream mini-transposon donor DNA containing the Tn7 transposon ends ("L" and "R") flanking the cargo of interest. (B) Schematic diagram of pAIO-A(pSL1120), encoding the CRISPR array and the TniQ-Cas8-Cas7-Cas6-TnsA-TnsA-TnsB operon from a single T7 promoter. The entire expression cassette is cloned into the mini-transposon donor DNA containing the Tn7 transposon ends ("L" and "R"). RNA-guided DNA integration using this construct mobilizes genetic components encoding CRISPR and Tn7-related machinery within the donor DNA itself. [Figure 26]Exemplary expression construct designs for optimizing promoter strength, plasmid copy number, and cargo size for all-in-one RNA-guided DNA integration experiments. A shows pAIO-A(pSL1120) further modified to contain one of four constitutive E. coli promoters, and the introduction of the entire expression cassette into four different vector backbones (left). The resulting 4 x 4 matrix is tested for RNA-guided DNA integration activity in E. coli BL21(DE3) cells and analyzed by PCR, qPCR, and / or next-generation sequencing. These experiments reveal optimal expression levels for a given copy number of expression plasmid. B is a schematic diagram of pAIO-A(pSL1120) modified to contain gene cargos ranging in size from 0.17 kilobase pairs (kbp) to 10 kbp. The resulting plasmids are tested for RNA-guided DNA integration activity in E. coli BL21(DE3) cells and analyzed by PCR, qPCR, and / or next-generation sequencing. These experiments reveal the dependence of cargo size on different expression constructs and designs. [Figure 27] 1 shows an exemplary promoter strategy for expression and reconstitution of RNA-guided DNA integration in a heterologous host of choice. The all-in-one expression vector pAIO-A (pSL1120) is further modified to carry alternative promoters (red) that are recognized and expressed in a variety of other expression hosts (shown in italics). In one embodiment (bottom right), the selected promoter has broad host range activity and can be recognized by a variety of known human commensal and pathogenic bacteria. In further embodiments, additional promoters are selected to match additional target host bacterial species. [Figure 28]Bioinformatics analysis of C2c5 homologs. After performing multiple sequence alignment of C2c5 proteins, a phylogenetic tree was constructed and visualized using Interactive Tree of Life. Based on multiple criteria, including sequence diversity, gene architecture, and readily identifiable transposon end sequences, five homologs and their associated Tn7-like transposon components were selected for further experimental studies. They are labeled with bacterial species information and highlighted with red arrows. [Figure 29] Genetic architecture of a Tn7-like transposon harboring a V-U5-type CRISPR-Cas system encoding C2c5. Representative genomic loci from five selected bacterial species are shown. The Tn7-like transposon ends (dark blue rectangles), the Tn7-related genes tnsB-tnsC-tniQ (shades of yellow), the CRISPR array (maroon), and the CRISPR-related gene c2c5 (blue) are shown. Similar to Tn7 transposons containing type I CRISPR-Cas systems, T7-like transposons containing type V CRISPR-Cas systems overwhelmingly harbor genes related to innate immune system functions, such as restriction-modification systems. The C2c5 gene is often flanked by the predicted transcriptional regulator merR (gray), and Tn7-like transposons containing C2c5 appear to almost always be located immediately upstream of a tRNA gene (green). This phenomenon has also been observed with other prokaryotic integration factors. [Figure 30] Figure 1 shows an exemplary experimental setup for studying RNA-guided DNA integration by Tn7-like transposons, including C2c5. (A) Schematic of a general plasmid expression system for Tn7-C2c5 transposition experiments. The CRISPR array contains two repeat sequences (gray diamonds) and one spacer sequence (maroon rectangle). The mini-transposon on pDonor is mobilized by trans-expressed transposase. (B) Schematic of the lacZ genomic locus targeted by a synthetic gRNA, including two potential Tn7 transposition products and PCR primer pairs for selectively amplifying them. [Figure 31]Experimental data demonstrating transposition by a Tn7-like transposon from Cyanobacterium aponinum IPPAS B-1202 (Cap). (A) Schematic of the genomic sites within lacZ targeted by six different gRNAs. Different PAM sequences (yellow) are indicated, and the target site is in maroon. (B) PCR-based detection of integration events, separated by agarose gel electrophoresis. A single upstream primer specific for the 3' end of the lacZ gene was used in combination with a primer reading through the left transposon end (primer pair c2, diagrammed in Figure 30B). Reactions with both 1:10 and 1:100 diluted lysates are shown, along with a positive control (+C) run with a lysate targeting the same region as the Tn7 transposon from V. cholerae. Potential integration events were detected at the PAM sequence and are indicated for gRNAs 4, 5, and 6. [Figure 32A] Representative existing approaches for targeted DNA enrichment. Figure 32A is a schematic diagram outlining the PCR process for DNA enrichment. PCR amplicons are generated to enrich the desired DNA target in either a uniplex format, a multiplex format with multiple primer pairs, or a custom emulsion-based technology (e.g., Rainstorm). [Figure 32B] A representative existing approach for targeted DNA enrichment is shown in Figure 32B, which shows a schematic of molecular inversion probe (MIP) annealing to input DNA adjacent to the region of interest for enrichment, leading to probe circularization by gap filling and ligation. [Figure 32C]Representative existing approaches for targeted DNA enrichment. Figure 32C is a schematic diagram of the most widely used approach for targeted DNA enrichment, in which a pool of oligonucleotide-based probes is used to hybridize with the target sequence either in an array format (solid support) or in solution, followed by a washing and elution step. This figure is reproduced from Mamanova et al., Nat Meth 7, 111-118 (2010) (incorporated herein by reference). [Figure 33A]Figure 33A shows a schematic diagram of targeted DNA enrichment using RNA-guided DNA integration by CRISPR-Tn7. In Figure 33A, input DNA, which can be purified genomic DNA, contains the desired sequence of interest (blue). gRNAs are designed against target sites (target-1 and target-2) adjacent to the sequence of interest, which themselves are flanked by protospacer adjacent motifs (PAMs) (5'-CC-3' in one embodiment of the V. cholerae CRISPR-Tn7 sequence). Purified TniQ-Cascade complexes carrying gRNA-1 and gRNA-2 bind to both target sites, resulting in the recruitment of TnsC, followed by the recruitment of paired-end complexes (PECs) containing TnsA, TnsB, and the transposon ends (L and R). Successful recruitment results in the RNA-guided integration of transposon end sequences located a certain distance downstream of the target site that are complementary to both gRNAs. Integration fragments the input DNA at the integration site while also adding transposon end sequences and, in one embodiment, adapter sequences, which can be used for downstream PCR amplification and / or NGS library preparation and next-generation sequencing (NGS). The stoichiometry of TnsA and TnsB in the paired-end complex, as well as that of TnsC, is unknown. The L- and R-ends of the transposon are shown in light purple and light orange, respectively, with optional adapter sequences shown in dark purple and dark orange. Sequences of interest can be selectively amplified (e.g., enriched) in a subsequent PCR step by designing primers to either the transposon end sequences, the adapter sequences, or both. Sample-specific indicators can also be added in this subsequent PCR amplification step. [Figure 33B]Figure 33B is a schematic diagram showing possible derivatives of transposon end sequences. In one embodiment, paired-end complexes contain two unique transposon ends (purple and orange), which allow for the integration of unique sequences onto the Watson-Crick strand of input DNA for downstream PCR amplification. In another embodiment, the transposon ends are further engineered so that the modified left (L*) or modified right (R*) end is recognized and faithfully incorporated by TnsB during RNA-guided DNA integration, resulting in uniform integration of the same transposon end sequence, thereby enabling downstream PCR amplification using a single primer that recognizes both ends. In a further embodiment, the transposon ends are engineered or modified so that one end remains "dark" in the subsequent PCR amplification step, allowing for the direction-specific integration of the L and R ends to enable the targeted amplification of only certain DNA sequences of interest for targeted DNA enrichment. Alternatively, the "dark" ends may simply be R and L ends that are functionally eliminated during the PCR amplification step. The bottom row shows the transposon end sequences (dark purple, dark orange) without the addition of adapter sequences. [Figure 33C] Schematic diagram of targeted DNA condensation using RNA-guided DNA integration by CRISPRTn7. Figure 33C shows the geometries of possible target and integration sites, which differ in the relative location of the target site compared to the DNA sequence of interest, resulting in different outcomes for what is retained during subsequent steps (e.g., PCR amplification of the integrated transposon ends). In embodiment 1, target 2 is retained; in embodiment 2, both target 1 and target 2 are retained; in embodiment 3, target 1 is retained; and in embodiment 4, neither target is retained. In embodiment 5, the target is selected to be within the DNA sequence of interest, with the PAM-in configuration such that RNA-guided DNA integration of the transposon end occurs just outside the sequence of interest. Further embodiments combine this strategy at one end with a target located outside the sequence of interest at the other end. [Figure 33D] Figure 33D is a schematic diagram of targeted DNA enrichment using RNA-guided DNA integration by CRISPRTn7. Figure 33D is a schematic diagram of a library of gRNAs used to direct highly multiplexed RNA-guided DNA integration in input DNA, which then allows for targeted enrichment of many target DNA sequences. [Figure 34] Figure 1 shows schematic diagrams of existing methods for generating random fragment libraries from input DNA. (A) A is a schematic diagram of a conventional approach involving mechanical (e.g., sonication) or enzymatic (e.g., dsDNA fragmentase (NEB)) fragmentation of input DNA (which can be purified genomic DNA). After end-polishing and A-tailing, sequencing adapters are then added to all dsDNA ends, and PCR amplification using primers complementary to the universal adapters results in a DNA library spanning the entire input DNA. This can be sequenced in a subsequent step using massively parallel DNA sequencing (e.g., next-generation sequencing with an Illumina platform). (B) A schematic diagram of tagmentation using an engineered Tn5 transposase (e.g., as in the Nextera kit) combines DNA fragmentation and adapter insertion in a single, rapid step, potentially saving considerable time, cost, and labor. The transposon ends or transposase adapters directly prime the subsequent PCR amplification, followed by next-generation sequencing. Figure is taken from Adey et al., Genome Biol 11, R119 (2010). [Figure 35A] Figure 35A is a schematic diagram of the preparation of recombinant CRISPR-Tn7 components for in vitro RNA-guided DNA integration. Figure 35A is a schematic diagram of exemplary expression plasmids cloned to recombinantly express and purify each individual protein component of the V. cholerae CRISPR-Tn7 machinery. Each plasmid encodes an N-terminal decahistidine tag upstream of the protein of interest, an MBP solubilization tag, and a TEV protease recognition sequence. [Figure 35B] Figure 35B is a schematic diagram of the preparation of recombinant CRISPR-Tn7 components for in vitro RNA-guided DNA integration. Figure 35B is a schematic diagram of gRNA generation by either in vitro transcription from dsDNA (shown, top) or partial ssDNA / dsDNA (not shown) templates, transcription of longer transcripts containing self-cleaving ribozymes (middle), or chemical synthesis (bottom). Libraries of gRNAs are generated by designing DNA template libraries or chemically synthesizing gRNA libraries. [Figure 35C] Figure 35C shows a schematic diagram of the preparation of recombinant CRISPR-Tn7 components for in vitro RNA-guided DNA integration. Figure 35C shows another embodiment in which the TniQ-cascade is recombinantly purified as a complex containing TniQ, Cas8, Cas7, Cas6, and gRNA using the indicated expression plasmids. The pCRISPR plasmid shown (pSL0915) encodes gRNA-3, which targets lacZ, but this may be replaced with other plasmids encoding different gRNAs. In another embodiment, the TniQ cascade is purified from a heterogeneous pool of cells expressing a library of different gRNAs (right). [Figure 35D] Figure 35D shows a schematic diagram of the preparation of recombinant CRISPR-Tn7 components for in vitro RNA-guided DNA integration. Figure 35D shows another embodiment in which TnsA and TnsB are purified as heterodimers using the indicated expression plasmids (left), or TnsA, TnsB, and TnsC are purified as a co-complex using the indicated expression plasmids (right). [Figure 35E] Figure 35E is a schematic diagram of the preparation of recombinant CRISPR-Tn7 components for in vitro RNA-guided DNA integration. Figure 35F is a schematic diagram of the polycistronic expression plasmid. [Figure 36]PCR amplification of integrated DNA for next-generation sequencing. In one embodiment, the transposon end sequences (orange lines) serve as primer binding sites for PCR amplification after targeted RNA-guided DNA integration adjacent to the DNA sequence of interest (see Figure 33). The PCR primers may also contain additional sequences on the overhangs for indexing and / or addition of sequences required for downstream next-generation sequencing (e.g., p5 / p7 sequences required for bridge amplification in Illumina sequencing platforms). After PCR and standard cleanup steps, the sample can be used directly for next-generation DNA sequencing. [Figure 37] Integration of a unique molecular identifier (UMI) during RNA-guided DNA integration. The transposon end sequences used during RNA-guided DNA integration (upstream steps not shown) are designed so that a unique molecular identifier (labeled UMI and shown in various colors in the diagram) is incorporated into one of the transposon end donor sequences. This ensures that different molecules of the same intended target sequence (shades of blue) carry a unique tag, which is preserved and amplified in a subsequent PCR step that adds the adapters required for next-generation DNA sequencing. [Figure 38]This figure shows a method for generating a sequencing library by flanking a sequence of interest with target and integration sites. In this embodiment, the sequence of interest (blue), which may be known or unknown, is flanked on one side by a known sequence (maroon) that serves as a target site to which a complementary gRNA can be designed. RNA-guided DNA integration by the CRISPR-Tn7 system results in the integration of transposon ends (orange / purple in the illustrated embodiment) approximately 50 bp downstream of the target site. This configuration allows the sequence of interest to be selectively amplified in a downstream PCR step by designing primers specific to the target site (maroon) and one of the transposon end sequences (orange). Additionally, adapters for next-generation sequencing (gray) can be added as overhangs in the PCR step, enabling downstream next-generation sequencing. This method can be multiplexed across many different sequences of interest. [Figure 39] Figure 1 shows different exemplary plasmid designs for the expression of protein and RNA components required for RNA-guided DNA integration. A is a schematic of one embodiment in which a three-plasmid approach is used to express the RNA-guided DNA integration (INTEGRATE) components. B is a schematic of another embodiment in which an all-in-one single plasmid is used for rational expression and delivery of the RNA-guided DNA integration (INTEGRATE) components. A simplified diagram is also shown (top). [Figure 40] Schematic representation of the formation of fusion products by replication, copy-and-paste transposition, and final resolution to the final product by homologous recombination. [Figure 41] FIG. 1 is a schematic diagram of the design of a selectable expanded construct using erythromycin resistance (ErmR), which is expressed only after the construct has integrated into a transcribed genomic locus. [Figure 42] FIG. 1 is a schematic diagram of an exemplary method for modulating antibiotic resistance. [Figure 43A]Overall architecture of the V. cholerae TniQ-Cascade complex. Figure 43A shows the genetic architecture of the Tn6677 transposon (top) and the plasmid construct used to express and purify the TniQ-Cascade co-complex. Selected cryo-EM reference-free 2D classes in multiple orientations are shown on the right. [Figure 43B] Overall architecture of the V. cholerae TniQ-Cascade complex. Figure 43B shows an orthogonal view of the cryo-EM map for the TniQ-Cascade complex, showing Cas8 (pink), six Cas7 monomers (green), Cas6 (salmon), crRNA (gray), and TniQ monomers (blue, yellow). The complex adopts a helical architecture with protuberances at both ends. [Figure 43C] Overall architecture of the V. cholerae TniQ-Cascade complex. Figure 43C shows the flexible domain of Cas8, including residues 277-385 (gray), which could only be visualized in the low-pass filtered map. The unsharpened map is shown as a semi-transparent gray map overlaid on the post-processed map segmented and colored according to Figure 43A. [Figure 43D] Overall architecture of the V. cholerae TniQ-Cascade complex. Figure 43D is a refined model of the TniQ cascade complex derived from the cryo-EM map shown in Figure 43B. [Figure 44A] This shows that TniQ binds to Cascade in a dimeric head-to-tail configuration. Figure 44A (left) shows a general view of the TniQ-Cascade cryo-EM unsharpened map (gray) overlaid on the post-processing map segmented and colored as in Figure 43. Figure 44A (right) shows the cryo-EM map (top) and refined model (bottom) of the TniQ dimer. The two monomers interact with each other in a head-to-tail configuration and are anchored to Cascade via Cas6 and Cas7.1. [Figure 44B]This shows that TniQ binds to Cascade in a dimeric head-to-tail configuration. Figure 44B shows the secondary structure of the TniQ dimer, with 11 α-helices organized into an N-terminal helix-turn-helix (HTH) domain and a C-terminal TniQ domain. The dimeric interaction between H3 and H11 is shown, as are the interaction sites of Cas6 and Cas7.1. [Figure 44C] This shows that TniQ binds to Cascade in a dimeric head-to-tail configuration. Figure 44C shows the cryo-EM density of the H3-H11 interaction, showing distinct side chain features (top) and allowing accurate modeling of the interaction (bottom). [Figure 44D] This shows that TniQ binds to Cascade in a dimeric head-to-tail configuration. Figure 44D is a schematic of the dimeric interactions, showing the key dimerization interface between the HTH and TniQ domains. [Figure 45A] Figure 45A shows that Cas6 and Cas7.1 form a binding platform for TniQ. Figure 45A shows the top magnified region showing the interaction site between Cascade and the TniQ dimer. Cas6 and Cas7.1 are displayed as molecular van der Waals surfaces, the crRNA is shown as a gray sphere, and the TniQ monomer is shown as a ribbon. [Figure 45B] Figure 45B shows the loop connecting TniQ. α-helices H6 and H7 (blue) bind within the hydrophobic cavity of Cas6. [Figure 45C] Cas6 and Cas7.1 form a binding platform for TniQ. Figure 45C shows that Cas7.1 interacts with the HTH domain of the TniQ.2 monomer (yellow), primarily through H2 and the loop connecting H2 and H3. [Figure 45D] Figure 45D shows experimental cryo-EM densities observed during TniQ-Cas6 interaction, demonstrating that Cas6 and Cas7.1 form a binding platform for TniQ. [Figure 45E] Figure 45E shows experimental cryo-EM densities observed in the interaction of TniQ-Cas7.1 (Figure 45E), demonstrating that Cas6 and Cas7.1 form a binding platform for TniQ. [Figure 46A] The DNA-bound structure of the TniQ-Cascade complex. Figure 46A shows a schematic diagram of a portion of the crRNA and dsDNA substrate, as experimentally observed in the electron density map of the DNA-bound TniQ-Cascade. The target strand (TS), non-target strand (NTS), and PAM and seed regions are shown. [Figure 46B] Figure 46B shows the DNA-bound structure of the TniQ-Cascade complex. Figure 46B shows a selected cryo-EM reference-free 2D class of DNA-bound TniQ-Cascade, where the density corresponding to dsDNA could be directly observed protruding from the Cas8 component in the 2D average (white arrow). [Figure 46C] DNA-bound structure of the TniQ-Cascade complex. Figure 46C shows a cryo-EM map of the DNA-bound TniQ Cascade. The crRNA is in dark gray, and the DNA is in red. To the right and below are detailed views of the PAM and seed recognition regions of the map, and the refined model is represented as sticks in the electron density. Cas8 is in pink, Cas7 is in green, crRNA is in gray, and DNA is in red. [Figure 46D] DNA-binding structure of the TniQ-Cascade complex. Figure 46D shows a V. cholerae transposon encoding a TniQ-Cascade co-complex that utilizes the sequence content of the crRNA to bind to complementary DNA target sites (left). The incomplete R-loop observed in the structure (center) may represent an intermediate state that may precede the downstream "locking" step involving proofreading of RNA-DNA complementarity. TniQ is located at the PAM end of the DNA-binding Cascade complex and may interact with TnsC during the downstream step of RNA-guided DNA insertion. [Figure 47A]Cryo-EM sample optimization and image processing workflow. Figure 47A shows a representative negative staining micrograph of 500 nM TniQ-Cascade. [Figure 47B] Cryo-EM sample optimization and image processing workflow. Figure 47B (left) shows a representative cryo-EM image of the 2 μM TniQ cascade. A small dataset of 200 images was collected on a Tecnai F20 microscope equipped with a Gatan K2 camera. Figure 47B (right) shows the reference-free 2D class average of this initial cryo-EM dataset. [Figure 47C] Cryo-EM sample optimization and image processing workflow. Figure 47C (left) shows a representative image from a large dataset collected on a Tecnai Polara microscope equipped with a Gatan K3 detector. Figure 47C (center) shows the detailed 2D class averages used for initial model generation obtained using the SGD algorithm implemented in Relion3 (Figure 47C (right)). [Figure 47D] Cryo-EM sample optimization and image processing workflow. Figure 47D shows the image processing workflow used to identify two major classes of TniQ cascade complexes in open and closed conformations. Local refinement with soft masks was used to improve the quality of maps within the terminal projections of the complexes. These maps aided in de novo modeling and refinement of the initial model. [Figure 48A] Fourier shell correlation (FSC) curves, local resolution, and unsharpened filter maps for the TniQ-Cascade complex in the closed conformation. Figure 48A shows the gold standard FSC curve using the half map, with an estimated global resolution of 3.4 Å based on an FSC of 0.143. [Figure 48B]Fourier shell correlation (FSC) curves, local decomposition, and unsharpened filter maps for the TniQ-Cascade complex in closed conformation. Figure 48B shows cross-validated model vs. map FSC. Blue curve: FSC between the refined shacked model and half map 1; red curve: FSC for half map 2 (not included in refinement); black curve: FSC between the final model and final map. The observed overlap between the blue and red curves ensures a non-overfitted model. [Figure 48C] Figure 47C shows the Fourier shell correlation (FSC) curve, local resolution, and unsharpened filter map of the TniQ-Cascade complex in the closed conformation. Figure 47D shows the unsharpened map colored according to local resolution as reported by RESMAP. [Figure 48D] Figure 48D shows the Fourier shell correlation (FSC) curves, local decomposition, and unsharpened filter maps for the TniQ-Cascade complex in the closed conformation. Figure 48E shows the final model colored according to the B-factor calculated by REFMAC. [Figure 48E] Figure 48E shows the Fourier shell correlation (FSC) curve, local resolution, and unsharpened filter map of the TniQ-Cascade complex in the closed conformation. Figure 48E shows the flexible Cas8 domain encompassing residues 277-385, which contacts the TniQ dimer on the other side of the crescent. Applying a width-expanding Gaussian filter to the unsharpened map allows for better visualization of this flexible region. [Figure 49]The TniQ-Cascade is superimposed with structurally similar Cascade complexes. The V. cholerae IF variant TniQ-Cascade complex (left) is superimposed with Pseudomonas aeruginosa IF Cascade11 (also known as the Csy complex; center, PDB ID: 6B45) and Escherichia coli IE Cascade9 (right, PDB ID: 4TVX). Shown are superimpositions of the entire complex (top), the Cas8 and Cas5 subunits with their 5' crRNA handles (top center), the Cas7 subunit with a fragment of the crRNA (bottom center), and the Cas6 subunit with its 3' crRNA handle (bottom). [Figure 50A] Representative cryo-EM densities for all components of the TniQ-Cascade complex in closed conformation. Figure 50A shows the final refined model of the TniQ Cascade, with Cas8 in purple, Cas7 monomers in green, Cas6 in red, TniQ monomers in blue and yellow, and crRNA in gray. [Figure 50B] Figure 50B shows representative cryo-EM densities for all components of the TniQ-Cascade complex in the closed conformation. Figure 50B shows the final refined model inserted into the final cryo-EM densities for selected regions of all molecular components of the TniQ-Cascade complex. Residues are numbered. [Figure 50C] Figure 50C shows representative cryo-EM densities for all components of the TniQ-Cascade complex in the closed conformation. Figure 50D shows the final refined model inserted into the final cryo-EM densities for selected regions of all molecular components of the TniQ-Cascade complex. Residues are numbered. [Figure 50D] Figure 50D shows representative cryo-EM densities for all components of the TniQ-Cascade complex in the closed conformation. Figure 50E shows the final refined model inserted into the final cryo-EM densities for selected regions of all molecular components of the TniQ-Cascade complex. Residues are numbered. [Figure 50E]Figure 50E shows representative cryo-EM densities for all components of the TniQ-Cascade complex in the closed conformation. Figure 50F shows the final refined model inserted into the final cryo-EM densities for selected regions of all molecular components of the TniQ-Cascade complex. Residues are numbered. [Figure 50F] Figure 50F shows representative cryo-EM densities for all components of the TniQ-Cascade complex in the closed conformation. Figure 50F shows the final refined model inserted into the final cryo-EM densities for selected regions of all molecular components of the TniQ-Cascade complex. Residues are numbered. [Figure 50G] Figure 50G shows representative cryo-EM densities for all components of the TniQ-Cascade complex in the closed conformation. Figure 50G shows the final refined model inserted into the final cryo-EM densities for selected regions of all molecular components of the TniQ-Cascade complex. Residues are numbered. [Figure 50H] Figure 50H shows representative cryo-EM densities for all components of the TniQ-Cascade complex in the closed conformation. Figure 50H shows the final refined model inserted into the final cryo-EM densities for selected regions of all molecular components of the TniQ-Cascade complex. Residues are numbered. [Figure 51]The interaction of Cas8 and Cas6 with crRNA is shown. (i) A refined model of the TniQ cascade is shown as a ribbon inserted into a semi-transparent van der Waals surface and colored as in Figure 1. (ii) and (iii) are enlarged images of Cas8 interacting with the 5' end of crRNA. The inset shows electron density for the highlighted region, with the base of nucleotide C1 stabilized by stacking interactions with arginine residues R584 and R424. (iv) Cas6 interacts with the 3' end of the crRNA "handle" (nucleotides 45-60). (v) The arginine-rich α-helix is inserted deep into the major groove of the terminal stem-loop. This interaction is mediated by electrostatic interactions between basic residues of Cas6 and the negatively charged phosphate backbone of crRNA. (vi) Cas6 (red) also interacts with Cas7.1 (green), establishing a β-sheet formed by β-strands contributed by both proteins. [Figure 52] Schematic representation of crRNA and target DNA recognition by the TniQ-cascade. A shows the TniQ-cascade residues that interact with crRNA. The approximate locations of all protein components of the complex, as well as the location of the Cas7 "fingers," are also shown. B shows the TniQ-cascade residues that interact with crRNA and target DNA, as shown in A. [Figure 53A] Fourier shell correlation (FSC) curves, local resolution, and local refinement maps for the TniQ-Cascade complex in the open conformation. Figure 53A shows the gold standard FSC curve using the half map, with an estimated global resolution of 3.5 Å by the FSC 0.143 criterion. [Figure 53B]Fourier shell correlation (FSC) curves, local decomposition, and local refinement maps for the TniQ-Cascade complex in the open conformation. Figure 53B shows cross-validated model vs. map FSC. Blue curve: FSC between the refined shacked model and half map 1; red curve: FSC for half map 2 (not included in refinement); black curve: FSC between the final model and final map. The overlap between the blue and red curves ensures a non-overfitted model. [Figure 53C] Fourier shell correlation (FSC) curves, local resolution, and local refinement maps for the TniQ-Cascade complex in the open conformation. Figure 53C shows unsharpened maps colored according to local resolution as reported by RESMAP. The right shows a slice of the map shown on the left. [Figure 53D] Fourier shell correlation (FSC) curves, local resolution, and local refinement maps for the TniQ-Cascade complex in the open conformation. Figure 53D shows that local refinement using a soft mask improved the map within the flexible region. The region of the map corresponding to the TniQ dimer is shown. Unsharpened maps, colored according to the local resolution estimate, are shown before (left) and after (right) mask refinement. [Figure 53E] Figure 53E shows the Fourier shell correlation (FSC) curve, local decomposition, and local refinement map for the TniQ-Cascade complex in the open conformation. Figure 53F shows the final model of the TniQ dimer region colored according to the local B-factor calculated by REFMAC. [Figure 54]This shows that TniQ possesses an HTH domain that is involved in protein-protein interactions within the TniQ dimer. A DALI search using the refined TniQ model as a probe revealed significant similarity between the N-terminal domains of TniQ and PDB entries 4r24 (A) and 3ucs (B) (Z score 4.1 / 4.1, rmsd 3.8 / 5.1). Both proteins contain a helix-turn-helix (HTH) domain, which is often involved in nucleic acid recognition and mediates protein-protein interactions. Figure C shows that the TniQ dimer is stabilized in a head-to-tail configuration by interactions mediated by the HTH and TniQ domains from both monomers. [Figure 55A] Fourier shell correlation (FSC) curves, local resolution, and unsharpened filter maps for the DNA-bound TniQ-Cascade complex. Figure 55A shows the gold standard FSC curve using the half map, with an estimated global resolution of 2.9 Å by the FSC 0.143 criterion. [Figure 55B] Fourier shell correlation (FSC) curves, local decomposition, and local refinement maps for the open conformation of the TniQ-Cascade complex. Figure 53B shows cross-validated model vs. map FSC. Blue curve: FSC between the refined shacked model and half map 1; red curve: FSC for half map 2 (not included in refinement); black curve: FSC between the final model and final map. The observed overlap between the blue and red curves ensures a non-overfitted model. [Figure 55C] Fourier shell correlation (FSC) curves, local resolution, and unsharpened filter maps of the DNA-bound TniQ-Cascade complex. Figure 55C (left) shows the unsharpened map colored according to local resolution as reported by RESMAP. dsDNA can be seen protruding outside the complex in the upper right. Figure 54C (right) shows the final model colored according to the B-factor calculated by REFMAC. [Figure 56]The DNA-bound TniQ-Cascade is superimposed with the structurally similar Cascade complex. The DNA-bound structure of the V. cholerae IF variant TniQ-Cascade complex (left) was superimposed with the DNA-bound structures of Pseudomonas aeruginosa IF Cascade11 (also known as the Csy complex; center, PDB ID: 6B44) and Escherichia coli IE Cascade9 (right, PDB ID: 5H9F). Shown are superimpositions of the entire complex (top), the Cas8 and Cas5 subunits with their 5' crRNA handles and duplex PAM DNA (top center), the Cas7 subunit with its fragment of crRNA (bottom center), and the Cas6 subunit with its 3' crRNA handle (bottom). [Figure 57A] Pairwise sequence identity between C2c5 homologs. [Figure 57B] Pairwise sequence identity between C2c5 homologs. [Figure 57C] Pairwise sequence identity between C2c5 homologs. [Figure 57D] Pairwise sequence identity between C2c5 homologs. [Figure 57E] Pairwise sequence identity between C2c5 homologs. [Figure 57F] Pairwise sequence identity between C2c5 homologs. [Figure 58A] Analysis of the C2c5 genomic locus of C2c5 homologs from Figure 57. [Figure 58B] Analysis of the C2c5 genomic locus of C2c5 homologs from Figure 57. [Figure 58C] Analysis of the C2c5 genomic locus of C2c5 homologs from Figure 57. [Figure 59]Multiple sequence alignment of TnsAs from Vch (Vibrio cholerae) (SEQ ID NO: 141); Ecl (Enterobacter cloacae) (SEQ ID NO: 1715); Asa (Aeromonas salmonicida) (SEQ ID NO: 716); Pmi (Proteus mirabilis) (SEQ ID NO: 1717); Eco (Escherichia coli) (SEQ ID NO: 1714). Conserved catalytic residues are indicated by red triangles. [Figure 60] Multiple sequence alignment of TnsB from Vch (Vibrio cholerae) (SEQ ID NO: 143); Ecl (Enterobacter cloacae) (SEQ ID NO: 1719); Asa (Aeromonas salmonicida) (SEQ ID NO: 1720); Pmi (Proteus mirabilis) (SEQ ID NO: 1721); Eco (Escherichia coli) (SEQ ID NO: 1718). Conserved catalytic residues are indicated by red triangles. [Figure 61] Multiple sequence alignment of TnsCs from Vch (Vibrio cholerae) (SEQ ID NO: 145); Ecl (Enterobacter cloacae) (SEQ ID NO: 1723); Asa (Aeromonas salmonicida) (SEQ ID NO: 1724); Pmi (Proteus mirabilis) (SEQ ID NO: 1725); and Eco (Escherichia coli) (SEQ ID NO: 1722). The Walker A and Walker B motifs characteristic of AAA+ ATPases are indicated, and active site residues involved in ATPase activity are indicated with blue triangles. Several TnsC homologs are annotated as TniB. [Figure 62]Figure 1: Multiple sequence alignment of TniQ / TnsD from Vch (Vibrio cholerae) (SEQ ID NO: 147); Ecl (Enterobacter cloacae) (SEQ ID NO: 1727); Asa (Aeromonas salmonicida) (SEQ ID NO: 1728); Pmi (Proteus mirabilis) (SEQ ID NO: 1729); Eco (Escherichia coli) (SEQ ID NO: 1726). VchTniQ is aligned to members of the TniQ / TnsD family. Conserved zinc finger motif residues are indicated by blue arrows. [Figure 63] Multiple sequence alignment of Cas6 from Vch (Vibrio cholerae) (SEQ ID NO: 153); Rho (Rhodanobacter sp.) (SEQ ID NO: 1730); Bpl (Burkholderia plantarii) (SEQ ID NO: 1731); Idi (Idiomarina sp. H105) (SEQ ID NO: 1732); and Pae (Pseudomonas aeruginosa) (SEQ ID NO: 1733). VchCas6 is aligned to other IF Cas6 proteins (often annotated as Cas6f or Csy4). Conserved catalytic residues are indicated by red arrows. [Figure 64] Multiple sequence alignment of Cas7 from Vch (SEQ ID NO: 151) (Vibrio cholerae); Rho (Rhodanobacter sp.) (SEQ ID NO: 1734); Bpl (Burkholderia plantarii) (SEQ ID NO: 1735); Idi (Idiomarina sp. H105) (SEQ ID NO: 1736); Pae (Pseudomonas aeruginosa) (SEQ ID NO: 1737). VchCas7 is aligned to other IF Cas7 proteins (often annotated as Csy3). [Figure 65A]Multiple sequence alignment of Cas8 and Cas5 from Vch (Vibrio cholerae) (SEQ ID NO: 149); Rho (Rhodanobacter sp.) (SEQ ID NOs: 1738 and 1742); Bpl (Burkholderia plantarii) (SEQ ID NOs: 1739 and 1743); Idi (Idiomarina sp. H105) (SEQ ID NOs: 1740 and 1744); and Pae (Pseudomonas aeruginosa) (SEQ ID NOs: 1741 and 1745), respectively. The naturally occurring Cas8-Cas5 fusion protein, VchCas8, is aligned to another IF Cas8 protein (often annotated as Csy1). [Figure 65B] Figure 1: Multiple sequence alignment of Cas8 and Cas5 from Vch (Vibrio cholerae) (SEQ ID NO: 149); Rho (Rhodanobacter sp.) (SEQ ID NOs: 1738 and 1742); Bpl (Burkholderia plantarii) (SEQ ID NOs: 1739 and 1743); Idi (Idiomarina sp. H105) (SEQ ID NOs: 1740 and 1744); and Pae (Pseudomonas aeruginosa) (SEQ ID NOs: 1741 and 1745), respectively. The naturally occurring Cas8-Cas5 fusion protein, VchCas8, is aligned against another IF Cas5 protein (often annotated as Csy2). [Figure 66]Schematic diagram of the generation of tnsA-tnsB fusions in Tn7-like transposons encoding IF-type CRISPR-Cas systems. The genetic organization of the transposons and CRISPR-Cas machinery are from selected transposons, including E. coli Tn7 (top), V. cholerae Tn6677 (second from the top), Parashewanella spongiae (second from the bottom), and a new candidate Tn7-like transposon from Aliivibrio wodanis (bottom). In the bottom two examples, there is a natural fusion between tnsA and tnsB. Genes from the CRISPR-Cas operon are also shown (tniQ, cas8, cas7, cas6, and CRISPR array). Protein accession IDs for the bottom two systems are shown below the gene schematics. "R" and "L" indicate the right and left ends of the transposon, respectively. [Figure 67A] Design and testing of engineered TnsA-TnsB fusion proteins derived from the V. cholerae Tn6677 transposon. Starting with the pTnsABC vector encoding the native TnsA, TnsB, and TnsC operons from V. cholerae, a synthetic fusion of TnsA-TnsB was constructed based on alignment with other native TnsA-TnsB fusions to generate a new modified pTns(AB)fC vector, pSL1738 (SEQ ID NO: 935). E. coli BL21(DE3) competent cells already containing the mini-transposon plasmid Donor (pDonor; pSL0527, SEQ ID NO: 7) and a plasmid encoding the TniQ-Cascade(crRNA-4) complex (pSL0828, SEQ ID NO: 14) were transformed with either the empty vector (pSL0008, SEQ ID NO: 3) as a control, the original pTnsABC vector (encoding TnsA, TnsB, and TnsC), or the new engineered vector (pSL1738) containing the TnsA-TnsB fusion protein together with TnsC. [Figure 67B]Design and testing of engineered TnsA-TnsB fusion proteins derived from the V. cholerae Tn6677 transposon. Integration efficiency was quantified by qPCR for both of the two possible integration orientations downstream of target-4, tRL and tLR. The engineered fusion proteins exhibited wild-type activity similar to that of the pSL0283 / pTnsABC (SEQ ID NO: 13) construct, demonstrating that the engineered TnsA-TnsB fusion proteins are functional in vivo for RNA-guided DNA integration. [Figure 68] Figure 11C is a graph showing the effect of truncating the right transposon end sequence on the preferred RNA-guided DNA integration direction, confirming the results from Figure 11C at four additional target sites. The x-axis indicates the length of the right transposon end sequence. Blue shades indicate T-LR (where the R-end of the transposon is proximal to the target site) integration, and orange shades indicate T-RL (where the R-end of the transposon is proximal to the target site) integration. Truncation of the right transposon end to 97 bp or less shifted integration preference toward the TRL direction (approximately 95% of integration events) and was consistent across all target sites tested. [Figure 69] FIG. 1 is a schematic diagram of an exemplary approach for generating and testing engineered transposon end sequences in a pooled library experiment. [Figure 70]Figure 1 shows a schematic diagram of an exemplary cloning approach for generating separate transposon end libraries from oligo pools. The right transposon end library was generated by digesting the insert and vector with HindIII and BamHI. The left transposon end library was generated by digesting with KpnI and XbaI. For library a), all possible combinations of TnsB binding sites for three different positions were generated. For library b), all possible combinations of TnsB binding sites for two different positions were generated. Library c) contained 2-bp mutations along the entire right end. Library d) constructed all possible 1-bp mutations along the 8-bp right end. Library e) contained missense mutations affecting three different possible open reading frames at the right transposon end. Library f) varied the distance between the TnsB binding sites at positions 1 and 2. For left transposon end library g), the distance between the TnsB binding sites between positions 1 and 2 or positions 2 and 3 was varied. The same spacing sequences were also mutated separately to compare the effects of distance and sequence identity. [Figure 71A]This graph shows the relative integration efficiency of members of the "Right Three Binding Sites" library (library a). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL)) and red (T-LR (tLR)). Relative integration efficiency was calculated for the variant END.1.2.3, which is most similar to the natural transposon end (END.1.2.3 is a 90-bp truncated version of the standard pDonor, and its orientation bias is expected to be heavily biased toward tRL). In this library, the positions of the three TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the identities of the six different TnsB binding sites, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 (SEQ ID NOs: 936 to 942, respectively). The x-axis indicates which TnsB binding site identities (1-7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 71B] This graph shows the relative integration efficiency of members of the "Right Three Binding Sites" library (library a). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL)) and red (T-LR (tLR)). Relative integration efficiency was calculated for the variant END.1.2.3, which is most similar to the natural transposon end (END.1.2.3 is a 90-bp truncated version of the standard pDonor, and its orientation bias is expected to be heavily biased toward tRL). In this library, the positions of the three TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the identities of the six different TnsB binding sites, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 (SEQ ID NOs: 936 to 942, respectively). The x-axis indicates which TnsB binding site identities (1-7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 71C]This graph shows the relative integration efficiency of members of the "Right Three Binding Sites" library (library a). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL)) and red (T-LR (tLR)). Relative integration efficiency was calculated for the variant END.1.2.3, which is most similar to the natural transposon end (END.1.2.3 is a 90-bp truncated version of the standard pDonor, and its orientation bias is expected to be heavily biased toward tRL). In this library, the positions of the three TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the identities of the six different TnsB binding sites, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 (SEQ ID NOs: 936 to 942, respectively). The x-axis indicates which TnsB binding site identities (1-7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 71D] This graph shows the relative integration efficiency of members of the "Right Three Binding Sites" library (library a). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL)) and red (T-LR (tLR)). Relative integration efficiency was calculated for the variant END.1.2.3, which is most similar to the natural transposon end (END.1.2.3 is a 90-bp truncated version of the standard pDonor, and its orientation bias is expected to be heavily biased toward tRL). In this library, the positions of the three TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the identities of the six different TnsB binding sites, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 (SEQ ID NOs: 936 to 942, respectively). The x-axis indicates which TnsB binding site identities (1-7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 71E]This graph shows the relative integration efficiency of members of the "Right Three Binding Sites" library (library a). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL)) and red (T-LR (tLR)). Relative integration efficiency was calculated for the variant END.1.2.3, which is most similar to the natural transposon end (END.1.2.3 is a 90-bp truncated version of the standard pDonor, and its orientation bias is expected to be heavily biased toward tRL). In this library, the positions of the three TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the identities of the six different TnsB binding sites, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 (SEQ ID NOs: 936 to 942, respectively). The x-axis indicates which TnsB binding site identities (1-7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 71F] This graph shows the relative integration efficiency of members of the "Right Three Binding Sites" library (library a). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL)) and red (T-LR (tLR)). Relative integration efficiency was calculated for the variant END.1.2.3, which is most similar to the natural transposon end (END.1.2.3 is a 90-bp truncated version of the standard pDonor, and its orientation bias is expected to be heavily biased toward tRL). In this library, the positions of the three TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the identities of the six different TnsB binding sites, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 (SEQ ID NOs: 936 to 942, respectively). The x-axis indicates which TnsB binding site identities (1-7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 71G]This graph shows the relative integration efficiency of members of the "Right Three Binding Sites" library (library a). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL)) and red (T-LR (tLR)). Relative integration efficiency was calculated for the variant END.1.2.3, which is most similar to the natural transposon end (END.1.2.3 is a 90-bp truncated version of the standard pDonor, and its orientation bias is expected to be heavily biased toward tRL). In this library, the positions of the three TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the identities of the six different TnsB binding sites, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 (SEQ ID NOs: 936 to 942, respectively). The x-axis indicates which TnsB binding site identities (1-7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 72] Figure 7 shows a graph of relative integration efficiency in members of the "Right Two Binding Sites" library (library b). The two different orientations in which the transposon can integrate are shown in blue (T-RL (tRL), top) and red (T-LR (tLR), bottom). Relative integration efficiency was calculated for variant END.1.2.3. In this library, the positions of the two TnsB binding sites at the right end were maintained, but their identities were varied to generate all possible combinations of binding sites. In addition to the six different TnsB binding site identities, the location of a naturally occurring palindromic sequence immediately inside the right transposon end was also tested. These seven different sequences are numbered 1 to 7 as in Figure 71. The x-axis indicates which TnsB binding site identities (1 to 7) were present at positions 1 and 2, counting from the distal right transposon end (see Figure 68). [Figure 73]Figure 1 shows a graph of relative integration efficiency in members of the "right-side 2 bp mutant" library (library c). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated for variant END.1.2.3. The x-axis shows the position of the affected base, counting from the rightmost transposon end base. [Figure 74] Figure 1 shows a graph of relative integration efficiency in members of the "right-end mutant" library (library d). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated for variant END.1.2.3. The x-axis shows both the position of the changed base, counted from the most terminal base pair, and the identity of the new nucleotide. [Figure 75A] Figure 1 shows a graph of relative integration efficiency in members of the "right linker sequence" library (library e). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated for variant END.1.2. The x-axis shows the amino acid change caused by the mutation. [Figure 75B] Figure 1 shows a graph of relative integration efficiency in members of the "right linker sequence" library (library e). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated for variant END.1.2. The x-axis shows the amino acid change caused by the mutation. [Figure 75C] Figure 1 shows a graph of relative integration efficiency in members of the "right linker sequence" library (library e). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated for variant END.1.2. The x-axis shows the amino acid change caused by the mutation. [Figure 76]Figure 1 shows a graph of relative integration efficiency in members of the "right-spacing" library (library f). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). Relative integration efficiency was calculated for variant END.1.2.3. Library f) has variable spacing between the first and second TnsB binding sites from the terminal right transposon end. The x-axis shows the distance between binding sites. [Figure 77A] This graph shows the relative integration efficiency in members of the "left-spacing" library (library g). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated relative to the unmutated truncated (122 bp) version of standard pDonor (based on the truncation data published in Klompe et al., Nature 571, 219-225 (2019), incorporated herein by reference, which is expected to have a directional bias of 0.60 (T-RL):0.40 (T-LR)). In addition, the right side of all of these clones contains an MmeI recognition site, which reduced integration efficiency by approximately 40% compared to the wild-type. The x-axis of each graph indicates the type of mutation present in that particular variant. If the change affects the distance between binding sites, this is shown as the number of base pairs comprising the interval. If there is a change in sequence identity, the position of the affected base (counting from the most terminal base within the interval) is indicated. [Figure 77B]This graph shows the relative integration efficiency in members of the "left-spacing" library (library g). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated relative to the unmutated truncated (122 bp) version of standard pDonor (based on the truncation data published in Klompe et al., Nature 571, 219-225 (2019), incorporated herein by reference, which is expected to have a directional bias of 0.60 (T-RL):0.40 (T-LR)). In addition, the right side of all of these clones contains an MmeI recognition site, which reduced integration efficiency by approximately 40% compared to the wild-type. The x-axis of each graph indicates the type of mutation present in that particular variant. If the change affects the distance between binding sites, this is shown as the number of base pairs comprising the interval. If there is a change in sequence identity, the position of the affected base (counting from the most terminal base within the interval) is indicated. [Figure 77C] This graph shows the relative integration efficiency in members of the "left-spacing" library (library g). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated relative to the unmutated truncated (122 bp) version of standard pDonor (based on the truncation data published in Klompe et al., Nature 571, 219-225 (2019), incorporated herein by reference, which is expected to have a directional bias of 0.60 (T-RL):0.40 (T-LR)). In addition, the right side of all of these clones contains an MmeI recognition site, which reduced integration efficiency by approximately 40% compared to the wild-type. The x-axis of each graph indicates the type of mutation present in that particular variant. If the change affects the distance between binding sites, this is shown as the number of base pairs comprising the interval. If there is a change in sequence identity, the position of the affected base (counting from the most terminal base within the interval) is indicated. [Figure 77D]This graph shows the relative integration efficiency in members of the "left-spacing" library (library g). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated relative to the unmutated truncated (122 bp) version of standard pDonor (based on the truncation data published in Klompe et al., Nature 571, 219-225 (2019), incorporated herein by reference, which is expected to have a directional bias of 0.60 (T-RL):0.40 (T-LR)). In addition, the right side of all of these clones contains an MmeI recognition site, which reduced integration efficiency by approximately 40% compared to the wild-type. The x-axis of each graph indicates the type of mutation present in that particular variant. If the change affects the distance between binding sites, this is shown as the number of base pairs comprising the interval. If there is a change in sequence identity, the position of the affected base (counting from the most terminal base within the interval) is indicated. [Figure 77E] This graph shows the relative integration efficiency in members of the "left-spacing" library (library g). The two different orientations in which the transposon can integrate are shown in blue (T-RL) and red (T-LR). The relative integration efficiency was calculated relative to the unmutated truncated (122 bp) version of standard pDonor (based on the truncation data published in Klompe et al., Nature 571, 219-225 (2019), incorporated herein by reference, which is expected to have a directional bias of 0.60 (T-RL):0.40 (T-LR)). In addition, the right side of all of these clones contains an MmeI recognition site, which reduced integration efficiency by approximately 40% compared to the wild-type. The x-axis of each graph indicates the type of mutation present in that particular variant. If the change affects the distance between binding sites, this is shown as the number of base pairs comprising the interval. If there is a change in sequence identity, the position of the affected base (counting from the most terminal base within the interval) is indicated. [Figure 78]1 is an exemplary flowchart for bioinformatics identification and selection of candidate CRISPR-transposon systems. Each box, in the order indicated by the arrow, highlights the steps used to assemble a large set of candidate CRISPR-transposon systems for experimental study. Certain steps are indicated as optional, and the entire pipeline can be limited based on various seeding strategies. For example, in the exemplary flowchart shown, the entire search algorithm is seeded based on the tnsB gene. In other embodiments, the search is seeded based on other transposon-associated genes, CRISPR-associated genes, the CRISPR array itself, or transposon end sequences. [Figure 79] Bioinformatics identification of CRISPR-transposon systems using an IF-type variant CRISPR-Cas system in which tnsA and tnsB are fused is shown. The two species shown contain CRISPR-transposon systems, with the tnsA and tnsB genes naturally found in fusion genes. The locations of the remaining components required for RNA-guided DNA integration are shown, along with the NCBI protein accession IDs. In the tnsA-tnsB gene sequence from Parashewanella spongiae strain HJ039, HHpred analysis confirmed the presence of distinctive Pfams for both TnsA (PF05367.11) and TnsB (PF09039.11 and PF02914.15). [Figure 80]This figure shows a vector approach for RNA-guided DNA integration experiments involving CRISPR-transposon homologs. The gRNA and all protein components were expressed from pCQT (representing the three modules present: the CRISPR array, the tniQ-cas8-cas7-cas6 genes, and the tnsA-tnsB-tnsB genes). A single T7 promoter then drives expression of the precursor guide RNA and a longer mRNA encoding all seven protein components (A). pCQT (a single expression effector plasmid) was combined with pDonor (A), which contains the DNA cargo flanked by transposon end sequences (left (L) and right (R)). The two vectors encoded spectinomycin and carbenicillin resistance. B shows a list of organisms from which the engineered CRISPR-transposon systems were derived. The left column shows biological information, the second column contains identifier information for the plasmids used for pCQT in each system (SEQ ID NOs: 855, 1623, 1624, 1625, 1626, 1627, 1628, 1903, 1629, 1904, 1905, 1630, 1906, 1907, 1908, respectively), and the third column contains identifier information for the plasmids used for pDonor in each system (SEQ ID NOs: 1614, 1615, 1616, 1617, 1618, 1619, 1620, 1897, 1621, 1898, 1899, 1622, 1900, 1901, 1902, respectively). Each pair of pCQT and pDonor plasmids can mate because the transposon end sequences on pDonor are specifically recognized by protein components on the cognate pCQT vector. The CRISPR transposon systems from Aliivibrio wodanis and Parashewanella spongiae encode tnsA-tnsB fusion proteins. [Figure 81]Figure 1 shows a graph of RNA-guided DNA integration data for modified pDonor vector backbones. Integration efficiencies were determined by qPCR on pDonor derivatives using the CRISPR-transposon system from Vibrio cholerae strain HE-45. In comparison to pSL0527 (SEQ ID NO: 7), pSL0921 (SEQ ID NO: 1613) has a deletion in the irrelevant lac promoter, and pSL1235 (SEQ ID NO: 1614) has additional irrelevant sequences removed. pSL0001 (SEQ ID NO: 5) is an empty vector control equivalent to pUC19, and pSL1209 (SEQ ID NO: 1612) is an empty vector control with the same irrelevant sequences removed that are not present in pSL1235. Integration efficiencies in both the tRL and tLR orientations are plotted and shown in red and blue, respectively. Because the pSL0921 and pSL1235 Donor plasmids exhibit slightly higher integration efficiency than pSL0527, pSL1235 was designed to serve as a benchmark pDonor vector for other homologous CRISPR-transposon systems. [Figure 82A] Figure 82A shows PCR detection of RNA-guided DNA integration products from a transposition assay using a homologous CRISPR-transposon system. Figure 82A is a schematic diagram of an experiment in which target 4 within the E. coli lacZ gene is targeted for proximal DNA integration. The mini-transposon donor DNA can be inserted in one of two orientations: tRL (up, down) and tLR (down, down), and different primer pairs are used to detect each orientation by PCR. [Figure 82B]Figure 82B shows PCR detection of RNA-guided DNA integration products from transposition assays using the homologous CRISPR-transposon system. Figure 82B shows PCR analysis of E. coli BL21(DE3) cells transformed with the plasmids indicated in the legend. In each experiment, cells were transformed with both plasmids and grown on LB agar plates containing inducers. Cells were then scraped, lysates were prepared, and PCR analysis was performed to detect integration products. PCR reactions were separated by 1% agarose gel electrophoresis. The top left panel shows results for a primer pair designed to amplify tRL products, and the bottom left panel shows results for the exact same set of lysates, but using a primer pair designed to amplify tLR products. Reactions were tested with CRISPR-transposon homologs from the following organisms: 1) a negative control for the system derived from Vibrio cholerae strain HE-45 lacking pDonor; 2) Vibrio cholerae strain HE-45; 3) Vibrio cholerae strain 4874; 4) Photobacterium iliopiscarium strain NCIMB; 5) Pseudoalteromonas species P1-25; 6) Pseudoalteromonas ruthenica strain S3245; 7) Photobacterium ganghwense strain JCM; 8) Shewanella species UCD-KL21; 9) Vibrio cholerae strain OYP7G04; and 10) Vibrio cholerae strain M1517. [Figure 82C]Figure 82C shows PCR detection of RNA-guided DNA integration products from transposition assays using the homologous CRISPR-transposon system. Figure 82C shows PCR analysis of E. coli BL21(DE3) cells transformed with the plasmids indicated in the legend. In each experiment, cells were transformed with both plasmids and grown on LB agar plates containing inducers. Cells were then scraped, lysates were prepared, and PCR analysis was performed to detect integration products. PCR reactions were separated by 1% agarose gel electrophoresis. The top left panel shows results for a primer pair designed to amplify tRL products, and the bottom left panel shows results for the exact same set of lysates, but using a primer pair designed to amplify tLR products. Reactions tested CRISPR-transposon homologs from the following organisms: 1) Vibrio diazotrophicus strain 60.6F; 2) Vibrio species 16; 3) Vibrio species F12; 4) Vibrio splendidus strain UCD-SED10; 5) Aliivibrio wodanis 06 / 09 / 160; and 6) Parashewanella spongiae strain HJ039. Note that the CRISPR-transposon systems in reactions / lanes 5 and 6 encode TnsA-TnsB fusion proteins. * indicates a nonspecific PCR amplicon. [Figure 83]Figure 1 shows vector layouts for testing RNA-guided DNA integration by type V CRISPR-Cas system-associated transposons. A is a schematic representation of different exemplary vector layouts. Experiments are performed with an all-in-one vector (pAIO, top) or a vector expressing the mechanism (pCCT, middle) combined with a separate donor vector (pDonor, bottom). The left and right transposon end sequences are represented by "L" and "R," respectively. B is the plasmid ID of exemplary vectors used to test type V CRISPR-Cas-associated transposons from Scytonema hofmannii strain PCC 7110: pSL1117 (SEQ ID NO: 1767), pSL1114 (SEQ ID NO: 1632), and pSL0948 (SEQ ID NO: 1631). "NT / cloning" indicates that these plasmids encode full-length sgRNAs, but the guides do not have targets in E. coli and are therefore non-targeting (NT). Additionally, these vectors allow for the easy cloning of new guide sequences. [Figure 84A] Figure 84A shows RNA-guided DNA integration using a V-type system. Figure 84A shows an exemplary schematic diagram for separately targeting four different sites located upstream of lacZ and one within the cynX gene. Integration events were analyzed using a combination of a genome-specific primer and one of two transposon-specific primers to reveal the different orientations in which the minitransposon can integrate. [Figure 84B] Figure 84B shows RNA-guided DNA integration using the V-type system. PCR and subsequent agarose gel electrophoresis analysis reveals successful site-specific integration for all four guides tested, biased towards integration in the tLR orientation over the tRL orientation. [Figure 84C]Figure 84C shows RNA-guided DNA integration using the V-type system. Figure 84C is a graph of quantitative analysis completed using qPCR at different target sites. These data confirm the directional bias revealed in Figure 84B and show efficient integration in all targeting guides tested. [Figure 84D] Figure 84D shows RNA-guided DNA integration using the V-type system. Figure 84D is a schematic diagram and results from a proof-of-principle experiment demonstrating that the all-in-one version of the system also promotes RNA-guided DNA integration. [Figure 85A] Genome-wide specificity of three different CRISPR-transposon systems: two V-type and one I-type related system. Figure 85A shows the genome-wide specificity of one V-type CRISPR-transposon system. For each system (top and middle columns), two different guides were tested, with the tSL number indicated above each plot. The corresponding target sites are shown as maroon triangles on the x-axis. The percentage of reads mapping to on-target sites is shown in red next to the peak, where possible. For each system, the y-axis is expanded to 0.5% of reads (bottom row). On-target specificity is shown in bold red. [Figure 85B] Genome-wide specificity of three different CRISPR-transposon systems: two V-type and one I-type related system. Figure 85B shows the genome-wide specificity of one V-type CRISPR-transposon system. For each system (top and middle columns), two different guides were tested, with the tSL number indicated above each plot. The corresponding target sites are shown as maroon triangles on the x-axis. The percentage of reads mapping to the on-target site is shown in red next to the peak, where possible. For each system, the y-axis is expanded to 0.5% of reads (bottom row). On-target specificity is shown in bold red. [Figure 85C]Genome-wide specificity of three different CRISPR-transposon systems: two type V and one type I-related system. Figure 85C shows the genome-wide specificity of one type I CRISPR-transposon system. For each system (top and middle columns), two different guides were tested, with the tSL number indicated above each plot. The corresponding target sites are shown as maroon triangles on the x-axis. The percentage of reads mapping to on-target sites is shown in red next to the peak, where possible. For each system, the y-axis is expanded to 0.5% of reads (bottom row). On-target specificity is shown in bold red. [Figure 86A] Figure 86A shows an overview of engineered vector designs to streamline expression and reconstitution of RNA-guided DNA integration. Figure 86A is a schematic diagram of the process of RNA-guided DNA integration, involving DNA targeting by the CRISPR-Cas system and integration of donor DNA proximal to the target site by a transposon system. [Figure 86B] Figure 86B shows an overview of engineered vector designs to streamline expression and reconstitution of RNA-guided DNA integration. Figure 86B is a schematic diagram illustrating how targeting of a 32-bp genomic target site adjacent to a protospacer adjacent motif (PAM) by an IF-type variant CRISPR-Cas system results in integration of donor DNA approximately 47–51 bp downstream. The donor DNA can be inserted in one of two possible orientations, indicated by the order of the transposon end closest to the target site. Thus, tRL results from insertion of the right end of the transposon proximal to the target site, while tLR results from insertion of the left end of the transposon proximal to the target site. [Figure 86C]Figure 86C shows an overview of engineered vector designs to streamline expression and reconstitution of RNA-guided DNA integration. Figure 86C is a schematic diagram of a three-plasmid system for reconstituting RNA-guided integration. pQCascade is driven by a T7 promoter and encodes gRNA, and also encodes TniQ, Cas8, Cas7, and Cas6 from a single operon, also driven by a T7 promoter. pTnsABC is driven by a T7 promoter and encodes TnsA, TnsB, and TnsC in a single operon. pDonor contains donor DNA flanked by transposon end sequences. [Figure 86D] Figure 86D shows an overview of the engineered vector design to streamline expression and reconstitution of RNA-guided DNA integration. Figure 86D is a schematic diagram of a two-plasmid system for reconstituting RNA-guided DNA integration. pCQT encodes the gRNA and all seven protein components under the control of a single T7 promoter. There is a single transcription terminator at the 3' end of the operon. The donor DNA remains encoded on pDonor (pSL1119). [Figure 86E] Figure 86E shows an overview of engineered vector designs to streamline expression and reconstitution of RNA-guided DNA integration. Figure 86E is a schematic of a single engineered all-in-one (AIO) plasmid system for reconstituting RNA-guided DNA integration. pAIO encodes the gRNA and all seven protein components and also contains donor DNA. [Figure 86F] Figure 86F shows an overview of engineered vector designs to streamline expression and reconstitution of RNA-guided DNA integration. Figure 86F is a schematic diagram showing how the single long transcript from pCQT / pAIO (containing the precursor CRISPR RNA 5' of a single-operon mRNA) can be readily processed by Cas6 in a type I CRISPR-Cas system into a mature gRNA (also called CRISPR RNA, or crRNA), leaving the downstream mRNA intact for translation by ribosomes. [Figure 86G]Figure 86G shows an overview of engineered vector designs to streamline expression and reconstitution of RNA-guided DNA integration. Figure 86G is a schematic diagram illustrating how the single long transcript derived from pCQT / pAIO (containing the precursor CRISPR RNA 3' of a single-operon mRNA) can be readily processed by Cas6 of the type I CRISPR-Cas system into a mature gRNA (also called CRISPR RNA, or crRNA), leaving the upstream mRNA intact for translation by the ribosome. pCQT in panel D is exemplified by pSL1022 (SEQ ID NO: 855) (all plasmid sequences can be found in SEQ ID NOs: 9, 848-861, and 1746-1764), and pDonor in panels C and D is exemplified by pSL1119 (SEQ ID NO: 1755). [Figure 87A] Figure 87A (left panel) shows the optimization of engineered vectors containing fewer vector and promoter elements. Figure 87A (left panel) shows an overview of the iterative screening of engineered vectors in which expression of the gRNA and TniQ-Cas8-Cas7-Cas6 operon is driven by a single T7 promoter rather than two separate T7 promoters. Three derivative plasmids (pQCascade, pQCascade-B, and pQCascade-C) were cloned and tested for RNA-guided DNA integration with pTnsBC and pDonor in E. coli BL21(DE3) cells. All three plasmids showed similar activity (Figure 87A, right panel), indicating that a single T7 promoter can drive efficient production of all required molecular components. In Figure 87A, pQCascade = pSL0828 (sequence number 14), pQCascade-B = pSL1016 (sequence number 849), pQCascade-C = pSL1018 (sequence number 851), pTnsABC = pSL0283 (sequence number 6), pDonor = pSL1119 (sequence number 1755). [Figure 87B]Figure 87B (left panel) shows the optimization of engineered vectors containing fewer vector and promoter elements. Figure 87B (left panel) shows an overview of the iterative screening of engineered vectors in which expression of the gRNA and TniQ-Cas8-Cas7-Cas6-TnsA-TnsB-TnsC operon is driven by a single T7 promoter rather than two or three. pC7QT, pCQT, pT7QC, and pTQC vectors with variable component order and number of T7 promoters were cloned and tested for RNA-guided DNA integration in E. coli BL21(DE3) cells. The right panel of Figure 87B shows a graph of quantified integration efficiency (measured by qPCR). pCQT shows improved efficiency compared to the other vectors. In Figure 87B, pC7QT = pSL1020 (sequence number 853), pCQT = pSL1022 (sequence number 855), pT7QC = pSL1024 (sequence number 857), pTQC = pSL1026 (sequence number 859). [Figure 88A] This figure shows a graph of integration efficiency analysis with variable vector backbones and specific gRNAs. Derivatives of the all-in-one pAIO vector were cloned, and the exact same construct was transfected into several different vector backbones, including pCDF, pUC19, pSC101, and pBBR1. These vectors differ in antibiotic resistance and, importantly, steady-state copy number. BL21(DE3) cells were transformed with each vector, and RNA-guided DNA integration efficiency was quantified by qPCR. The data show that the pBBR1 and pSC101 vector backbones were the most efficient for RNA-guided DNA integration in this comparative study. In panel A, "pCDF" is exemplified by pSL1213 (SEQ ID NO: 1751), "pUC19" is exemplified by pSL1121 (SEQ ID NO: 861), "pSC101" is exemplified by pSL1220 (SEQ ID NO: 1752), and "pBBR1" is exemplified by pSL1222 (SEQ ID NO: 1753). [Figure 88B]This graph shows an analysis of integration efficiencies with variable vector backbones and specific gRNAs. The efficiency of RNA-guided DNA integration at five different target sites was systematically compared between an all-in-one plasmid design (pAIO) and a three-plasmid design with multiple T7 promoters and vectors driving gRNAs, a TniQ-Cas8-Cas7-Cas6 operon, and a TnsA-TnsB-TnsC operon. The efficiencies of the three-plasmid systems were normalized to 1, and the relative efficiencies of the pAIO plasmids were plotted. In all cases, the overall efficiency of the single all-in-one plasmid system was 2- to 5-fold higher than that of the three-plasmid system. [Figure 88C] Figure 88C shows a graph of integration efficiency analysis with variable vector backbones and specific gRNAs. Figure 88C shows genome-wide RNA-guided DNA insertion specificity assessment by Tn-seq in engineered all-in-one (pAIO) vectors. After performing Tn-seq-based experiments to assess genome-wide specificity, the percentage of on-target integration relative to the total number of on-target integration sites was calculated by considering the total number of reads mapping to the on-target integration site. All five gRNAs in the pAIO vector backbone directed integration with approximately 100% target specificity. [Figure 89] Tn-seq data for the engineered all-in-one pAIO vector. The genome-wide specificity of gRNA-1, gRNA-4, gRNA-12, gRNA-13, and gRNA-17 in the pAIO vector is shown by plotting all Tn-seq reads across the 5.6 Mbp E. coli genome. The inset on the right shows a zoomed-in on-target peak, and tabulates the on-target specificity (text line 2) and tRL:tLR orientation ratio (text line 3) for the same gRNA-1. [Figure 90A]Figure 90A shows engineered vectors with various promoters for RNA-guided DNA integration. Starting with an all-in-one pAIO plasmid containing an inducible T7 promoter, the promoter was replaced with either a variety of synthetic biology promoters (J series) of variable expression strength, as well as a broad-host-range promoter derived from previous work developing methods for in situ bacterial engineering using the lac promoter or conjugative plasmids (Ronda, C., Chen, SP, Cabral, V., Yaung, SJ & Wang, HH Nat Meth 16, 167-170 (2019) (incorporated herein by reference)). After cloning the desired plasmids, E. coli BL21(DE3) cells were transformed with pAIOs containing the promoters listed, and the efficiency of RNA-guided DNA integration was quantified by qPCR. The most powerful J23119 promoter showed optimal activity, and integration efficiency decreased with decreasing promoter strength. In panel A, "J23119" is exemplified by pSL1130 (SEQ ID NO: 864), "J23114" is exemplified by pSL1133 (SEQ ID NO: 867), and "MAGIC-1" is exemplified by pSL1279 (SEQ ID NO: 1750). [Figure 90B] Figure 90B shows engineered vectors with various promoters for RNA-guided DNA integration. Genome-wide specificity measurement using Tn-seq shows that variable expression levels of the machinery or variable absolute integration efficiency do not change genome-wide specificity. [Figure 90C]Figure 9 shows engineered vectors with various promoters for RNA-guided DNA integration. RNA-guided DNA integration assays were performed using the all-in-one pAIO vector containing variable promoter strengths, culturing transformed E. coli cells at either 37°C (red), 30°C (yellow), or 25°C (blue). Integration efficiency (Figure 90B) was then quantified after 24 hours of solid-medium culture by qPCR. Results demonstrate that low-efficiency constructs with low activity at 37°C (e.g., the weak J23114 promoter) achieve approximately 100% integration efficiency when cells are cultured at lower temperatures. These experiments provide a facile experimental strategy for increasing integration efficiency under otherwise suboptimal vector or promoter conditions at high temperatures. In panel C, T7-lacO is exemplified by pSL1213 (SEQ ID NO: 1751), "J23119" is exemplified by pSL1130 (SEQ ID NO: 864), and "J23114" is exemplified by pSL1133 (SEQ ID NO: 867). [Figure 91A] These results demonstrate that RNA-guided DNA integration proceeds independently of specific host and recombination factors. The all-in-one pAIO vector, containing the strong constitutive promoter J23119, was transformed into several different E. coli strains, including MG1655, BW25113, and BL21(DE3). Figure 91A shows the genome-wide specificity of RNA-guided DNA integration analyzed within each genetic background. The plotted data represent integration events at on-target sites. In addition, the text in the upper right corner of each plot reports on-target specificity (line 2), measured by dividing the reads at on-target sites by the total genome-mapped reads, as well as the tRL:tLR directional bias. These experiments demonstrate that the favorable specificity profile and near-exclusive directional preference for tRLs are reproducible across several different E. coli strains. [Figure 91B]This indicates that RNA-guided DNA integration proceeds independently of specific host and recombination factors. Multiple Keio knockout strains were transformed with the all-in-one pAIO vector containing the strong constitutive promoter J23119 (exemplified by pSL1130, SEQ ID NO: 864). Gene knockouts are indicated along the x-axis. For each strain, the integration efficiency relative to the wild-type BW25113 strain is plotted. These results indicate that the RecA recombinase, as well as recD, recF, and mutS factors, are completely dispensable for RNA-guided DNA integration. [Figure 92A] We demonstrate that RNA-guided DNA integration can be stimulated by low-temperature incubation, enabling highly efficient insertion of large gene payloads exceeding 10 kb. For RNA-guided DNA integration experiments, we used a two-plasmid system containing pDonor and pCQT, driven by a T7 promoter and targeting the E. coli genome with crRNA-4. A negative control experiment (non-targeting crRNA "nt"; no donor DNA) showed no integration as measured by qPCR. When transformed E. coli cells were cultured on solid medium at 37°C, integration efficiency decreased significantly as the gene payload size increased from 0.98 kb to 10 kb. However, when the exact same transformed cells were cultured on solid medium at 30°C, integration efficiency remained approximately 100%, regardless of the size of the gene payload inserted into pDonor between the transposon ends. [Figure 92B] These results demonstrate that RNA-guided DNA integration can be stimulated by low-temperature incubation, allowing for highly efficient insertion of large gene payloads exceeding 10 kb. For RNA-guided DNA integration experiments, we used a two-plasmid system containing pDonor and pCQT, driven by a T7 promoter and targeting the E. coli genome with crRNA-4. In Figure 92B, a similar experiment was performed, except that the expression vector used the J23119 promoter instead of the T7 promoter. Again, incubation at lower temperatures demonstrated a consistent and statistically significant increase in overall integration efficiency, regardless of payload size, compared with incubation at 37°C. [Figure 92C] These results demonstrate that RNA-guided DNA integration can be stimulated by low-temperature incubation, enabling highly efficient insertion of large gene payloads exceeding 10 kb. For RNA-guided DNA integration experiments, we used a two-plasmid system containing pDonor and pCQT, driven by a T7 promoter and targeting the E. coli genome with crRNA-4. In Figure 92C, a similar experiment was performed, except that the expression vector used the J23119 promoter instead of the T7 promoter and crRNA-13 instead of crRNA-4. Again, incubation at lower temperatures demonstrated a consistent and statistically significant increase in overall integration efficiency, regardless of payload size, compared with incubation at 37°C. pCQT is exemplified by pSL1022 (SEQ ID NO: 855). pDonor is exemplified by pSL1119 (SEQ ID NO: 1755) for the 0.98 kb version and pSL1619 (SEQ ID NO: 1756) for the 10 kb version. [Figure 93] We demonstrate that a fully autonomous, self-mobilizable mobile genetic element undergoes highly efficient RNA-guided DNA integration. We constructed an autonomous all-in-one plasmid (gRNA) in which a promoter-driven operon encoding the gRNA and all seven protein components (TniQ-Cas8-Cas7-Cas6-TnsA-TnsB-TnsC) was inserted directly between the left and right ends of the transposon (A). This converts the mini-transposon into a self-mobilizable element, encoding the machinery that directs RNA-guided DNA integration to insert the donor DNA into the target site and then continuously mobilize the same donor DNA at any target site programmed within the CRISPR array. Despite the large size of the gene payload (>10 kb), RNA-guided DNA integration of the donor DNA in pAAIO (B) proceeds with approximately 100% efficiency without any drug selection when transformed E. coli cells are cultured at 30°C versus 37°C. pAAIO is exemplified by pSL1184 (SEQ ID NO: 1747). [Figure 94A]Figure 94A shows multiplexed RNA-guided DNA integration using a multiplexed spacer CRISPR array. By encoding multiple different spacers within the expanded CRISPR array, as shown in Figure 94A, the engineered CRISPR-transposon system can be easily transformed into a multiplexed platform for inserting DNA proximal to multiple target sites within the same genomic DNA. In type I CRISPR-Cas systems, which use Cas6 for ribonucleolytic processing, long precursor CRISPR RNAs can be easily processed. [Figure 94B] This figure shows multiplexed RNA-guided DNA integration using a multiplexed spacer CRISPR array. As shown in Figure 94B (left), CRISPR arrays were constructed in which the maroon spacer sequence was absent (top), the only spacer present (second from the top), or one of several different spacers located within different positions on the CRISPR array relative to the transcription start site of the CRISPR array. For each different construct, RNA-guided DNA integration experiments were performed in E. coli BL21(DE3) cells, and the efficiency of RNA-guided DNA integration proximal to the genomic target site programmed by the maroon spacer was measured by qPCR. The total efficiency is plotted against the efficiency when the maroon spacer was the only spacer in the array (Figure 94B (right)). These results demonstrate that the maroon spacer, even when present as one of three different spacers, can still direct RNA-guided DNA integration with greater than 50% wild-type efficiency and has the highest activity when closest to the 5' transcription start site. [Figure 94C]This figure shows multiplexed RNA-guided DNA integration using a multiplexed spacer CRISPR array. Genome-wide specificity analysis from Tn-seq libraries was generated from cells that underwent multiplexed donor DNA integration using CRISPR arrays encoding three different spacer sequences. Tn-seq analysis revealed that 99.6% of reads were exclusively located at one of the three target sites, demonstrating the extremely high efficiency and on-target accuracy of multiplexed integration. Because ligation efficiency is known to be sequence-dependent and other confounding factors contribute to the overall height of the peaks from next-generation sequencing, no conclusions can be drawn from the Tn-seq profiles regarding the relative efficiency of DNA integration at these three sites. A two-spacer array construct is exemplified by pSL1202 (SEQ ID NO: 1757), and a three-spacer array construct is exemplified by pSL1341 (SEQ ID NO: 1758). [Figure 95]We demonstrate that multiplexed RNA-guided DNA integration results in predictable phenotypes. A multispacer CRISPR array was constructed containing one spacer target, thrC, for insertional inactivation and a second spacer target, lysA, for insertional inactivation (A, top). Cells undergoing multiplexed RNA-guided DNA integration should become auxotrophic for threonine and lysine because knockout insertions within these two genes prevent the synthesis of these amino acids from a carbon source. To test this hypothesis, E. coli cells were transformed and the resulting transformants were plated on either M9 minimal medium, M9 minimal medium + lysine, M9 minimal medium + threonine, or M9 intermediate medium + threonine and lysine. Because auxotrophic cells could only grow on plates containing the corresponding amino acids, the efficiency of multiplexed RNA-guided DNA integration was directly determined by relative colony counts on various LB agar plates. These experiments showed that after this single-step multiplex RNA-guided DNA integration activity, approximately 20% of the cells were immediately double-auxotrophic (A, bottom). To further confirm these results, clones isolated from the various plates were grown in liquid culture in the presence of various media sources, and their growth was then measured over time in a shaking microplate incubator and reader. The results (B) demonstrate that the predicted double-auxotrophic strains were indeed completely unable to grow in minimal medium alone, but instead required both threonine and lysine ("TL") in M9 minimal medium to survive. The construct in panel A is exemplified by pSL1642 (SEQ ID NO: 1759). [Figure 96A]Figure 96A shows an engineered CRISPR-transposon system for mobilizing donor DNA within cells. Tn7-like transposons exhibit target immunity, where the presence of one genomic integrated transposon suppresses the same target site from undergoing another round of integration. Figure 96A outlines an exemplary workflow for studying immunity. On the left, a temperature-sensitive all-in-one plasmid (pAIO-ts) is used to subject the genome to RNA-guided DNA integration, allowing cells to recover from the plasmid after a successful integration event. These cells are then made chemically competent and subjected to the next round of transformation, in which the protein-RNA machinery is delivered with a different, traceable pDonor (pCQT). If this system exhibits target immunity, the same target site should not be able to serve as an efficient receiver for another donor DNA molecule. pAIO-ts in panel A is exemplified by pSL1223 (SEQ ID NO: 1754). pCQT in panel A is exemplified by pSL1022 (SEQ ID NO: 855). [Figure 96B] Figure 96B shows an engineered CRISPR-transposon system for mobilizing donor DNA within cells. Tn7-like transposons exhibit target immunity, where the presence of one genomically integrated transposon suppresses the same target site from undergoing another round of integration. Figure 96B shows an exemplary experiment testing the distance range of target immunity. Starting with a cell line containing genomically integrated donor DNA (the "immunized" state), pCQT was transformed with gRNAs targeting variable target sites upstream of the existing donor DNA (target sites ranging from 0 to 5003 bp, up to more than 1 Mb from the first donor DNA site). Relative integration efficiencies were then calculated by measuring the local integration efficiency of naive wild-type strains and immunized strains by qPCR. Ratios were plotted, and the results demonstrate that target immunity can operate on a long-distance scale relative to the distance between target DNA binding and donor DNA integration. [Figure 96C]Figure 96C shows an engineered CRISPR-transposon system for mobilizing donor DNA in cells. Tn7-like transposons exhibit targeting immunity, and the presence of one genomic integrated transposon suppresses the same target site from undergoing another round of integration. In another embodiment, Figure 96C, the machinery encoded by pCQT is delivered to the immunized strain, but another copy of pDonor is not delivered. In this embodiment, the machinery can excise donor DNA from its existing site in the genome and mobilize it to a new target site based on the spacer content in pCQT. This embodiment provides a method for performing programmed translocation in cells with existing donor DNA that has transposon ends recognized by the CRISPR-transposon system. [Figure 97A]This demonstrates that the two engineered CRISPR-transposon systems do not cross-react and can therefore be used as an orthogonal RNA-guided DNA integration system. Figure 97A is a schematic diagram of the orthogonal RNA-guided integrase. The IF-variant CRISPR-transposon system from Vibrio cholerae strain HE-45 (left) is used to reconstitute RNA-guided DNA integration in E. coli using the pDonor plasmid and the pCQT expression plasmid. The V-type CRISPR-transposon system from Scytonema hofmannii strain PCC 7110 (right) is used to reconstitute RNA-guided DNA integration in E. coli using the pDonor plasmid (Sho-pDonor), a plasmid encoding an sgRNA under the control of a T7 promoter, and the Cas12k-TnsB-TnsC-TniQ operon (Sho-PCCT) under the control of a second T7 promoter. Experiments were performed to determine whether Vch-pCQT could mobilize the Shop-pDonor donor and whether Sho-pCCT could mobilize the Vch-pDonor donor DNA. Various combinations of the plasmids shown above the gel were used to transform E. coli BL21(DE3) cells, and primer pairs were used to detect RNA-guided DNA integration products. Different primer pairs were chosen to selectively amplify tRL or tLR products. In panel A, Vch-pCQT is exemplified by pSL1022 (SEQ ID NO: 855), Vch-pDonor is exemplified by pSL1119 (SEQ ID NO: 1755), Sho-pCCT is exemplified by pSL1115, and Sho-pDonor is exemplified by pSL0948 (SEQ ID NO: 1631). [Figure 97B]This indicates that the two engineered CRISPR-transposon systems do not cross-react and can therefore be used as orthogonal RNA-guided DNA integration systems. The results in Figure 97B clearly demonstrate that Vch-pCQT catalyzed RNA-guided DNA integration using its own Vch-Donor donor DNA but was unable to direct RNA-guided DNA integration using Sho-Donor donor DNA, and vice versa. However, when the expression plasmid was paired with the cognate donor DNA plasmid, both systems were able to catalyze efficient and robust RNA-guided DNA integration. [Figure 98A] We demonstrate that the engineered CRISPR-transposon system functions robustly in multiple other bacterial species. Using the CRISPR-transposon system from Vibrio cholerae strain HE-45, we generated a modified, engineered all-in-one plasmid in which the machinery and donor DNA were cloned into the broad-host-range pBBR1 backbone (pAIO-BBR1). Within this vector, we utilized the strong constitutive J23119 promoter, known to be recognized by a variety of Gram-negative bacteria. Using this plasmid, we cloned different spacer sequences to direct RNA-guided DNA integration in Klebsiella oxytoca and Pseudomonas putida. P. putida and K. oxytoca were electroporated with pAIO-BBR1 containing spacers targeting multiple different genes, and successful integrations were probed for either the tRL or tLR orientation using one of four different primer pairs a–d (Figure 98B), examining both the upstream and downstream genome-transposon junctions. [Figure 98B]We demonstrate that the engineered CRISPR-transposon system functions robustly in multiple other bacterial species. Using the CRISPR-transposon system from Vibrio cholerae strain HE-45, we generated a modified, engineered all-in-one plasmid in which the machinery and donor DNA were cloned into the broad-host-range pBBR1 backbone (pAIO-BBR1). Within this vector, we utilized the strong constitutive J23119 promoter, known to be recognized by a variety of Gram-negative bacteria. Using this plasmid, we cloned different spacer sequences to direct RNA-guided DNA integration in Klebsiella oxytoca and Pseudomonas putida. P. putida and K. oxytoca were electroporated with pAIO-BBR1 containing spacers targeting multiple different genes, and successful integrations were probed for either the tRL or tLR orientation using one of four different primer pairs a–d (Figure 98B), examining both the upstream and downstream genome-transposon junctions. [Figure 98C] We demonstrate that the engineered CRISPR-transposon system functions robustly in multiple other bacterial species. Figure 98C shows PCR analysis of RNA-guided DNA integration in the indicated bacterial species (top) analyzed by agarose gel electrophoresis. Data from a gRNA targeting one of two target genes is shown on the gel (see gene labels at the top of the figure), and cell lysates were probed with one of four primer pairs a, b, c, and d. The band at the top of the gel indicates robust RNA-guided DNA integration, which was confirmed by subsequent Zanger sequencing. The PCR above the gel amplifies a reference housekeeping gene and serves as a loading control for the lysate preparation. Genomic DNA was purified from transformed cells and subjected to Tn-seq analysis of the genome-wide specificity of RNA-guided DNA integration. [Figure 98D]We demonstrate that the engineered CRISPR-transposon system functions robustly in multiple other bacterial species. In Figure 98D, Tn-seq analysis demonstrated that approximately 95–100% of integration events occurred at the expected target site for both Klebsiella oxytoca and Pseudomonas putida, with the same distance rules previously observed in E. coli. The two P. putida guides showed much lower specificity, likely due to the presence of highly similar off-target sequences elsewhere in the genome. The pAIO-BBR1 construct used in K. oxytoca is exemplified by pSL1813 (SEQ ID NO: 1761). The pAIO-BBR1 construct used in P. putida is exemplified by pSL1802 (SEQ ID NO: 1760). [Figure 99A] This paper presents a method for avoiding self-inactivation of the CRISPR-transposon system. The CRISPR-transposon system from Vibrio cholerae strain HE-45 can target its own PAM sequence within the 3' end of the CRISPR array repeat sequence (5'-AC-3'), albeit with low efficiency, making the system susceptible to self-inactivation. That is, if the machinery indiscriminately targets its own target (encoding the gRNA) present within the CRISPR array itself, integration of the donor DNA downstream can inactivate the machinery (indicated by the red X in Figure 99A) and / or cause plasmid instability. This effect is mitigated under conditions where maintaining the plasmid incurs a fitness cost to the cell or when the desired RNA-guided DNA integration event incurs a fitness cost to the cell. [Figure 99B]This figure shows a method for avoiding the self-inactivation of the CRISPR-transposon system. In Figure 99B, an experiment using an engineered CRISPR-transposon system to target both bdhA and nirC for insertional inactivation via RNA-guided DNA integration showed clear evidence of the system's self-inactivation. By analyzing Tn-seq data, which provides an unbiased assessment of all integration sites genome-wide, we found that the self-targeting of CRISPR-encoded spacers results in a significant excess of reads compared to the low number of reads mapped to the genome. [Figure 99C] A method for circumventing self-inactivation of the CRISPR-transposon system is shown. To circumvent this issue, in Figure 99C, a reverse orientation all-in-one plasmid was cloned onto the pBBR1 backbone (designated pRAIO-BBR1). In this plasmid, the CRISPR array is at the 3' end of the polycistronic construct, followed by the mRNA protein encoding TnsA-TnsB-TnsC-TniQ-Cas8-Cas7-Cas6. This alternative orientation placed the self-target in close proximity to the donor DNA on the pRAIO-BBR1 vector, potentially preventing any escape self-targeting by immune mechanisms. [Figure 99D] This figure shows a method for avoiding the self-inactivation of the CRISPR-transposon system. When the experiment from Figure 99B was repeated, but using a new pRAIO-BBR1 vector, the self-inactivation problem was completely eliminated, and all reads were mapped to the target site in the genome, and no reads were observed that resulted from self-inactivation and RNA-guided DNA integration downstream of the CRISPR array. Therefore, this engineered system is desirable for use in experiments that have a fitness advantage when cells inactivate the CRISPR-transposon system. [Figure 99E]This figure shows a method for avoiding self-inactivation of the CRISPR-transposon system. To further confirm the utility of the engineered pRAIO-BBR1 vector, in Figure 99E, the percentage of all Tn-seq reads mapping to on-target sites was plotted, revealing that the new engineered pRAIO-BBR1 vector exhibited excellent on-target specificity for both difficult-to-knockout genes. pAIO-BBR1 is exemplified by pSL1802 (SEQ ID NO: 1760), and pRAIO-BBR1 is exemplified by pSL1780 (SEQ ID NO: 1763). [Figure 100A] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100B] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100C] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100D]Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100E] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100F] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100G] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100H]Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100I] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 100J] Table of guide RNAs and genomic target sites. *Coordinates are relative to the E. coli BL21(DE3) genome (GenBank accession CP001509). †PAM sequences indicate the two nucleotides immediately 5' to the target (V. cholerae and P. aeruginosa cascades) or the three nucleotides immediately 3' to the target (S. pyogenes Cas9) on the non-target strand. [Figure 101A] 1 is a table of oligonucleotides used in PCR. [Figure 101B] 1 is a table of oligonucleotides used for qPCR. [Figure 101C] 1 is a table of oligonucleotides used in NGS. [Figure 102A] Table of potential CRISPR-transposon systems. [Figure 102B] Table of potential CRISPR-transposon systems. [Figure 102C] Table of potential CRISPR-transposon systems. [Figure 103]Figure 1 shows the generation of a pooled gRNA library of RNA-guided DNA integration events across a population of cells. Figure 2 shows that a gRNA library is cloned by designing and synthesizing an oligo array library containing the desired spacer or guide sequences. Using standard molecular biology and molecular cloning methods, these oligos are converted into double-stranded DNA and cloned into an expression plasmid within a CRISPR array. Consequently, transcription of the CRISPR array produces gRNAs or gRNA precursors, which are processed by Cas6 into mature gRNAs. The expression plasmid can contain only a CRISPR array, or a CRISPR array and one or more protein-coding genes (e.g., genes involved in RNA-guided DNA integration). The CRISPR array can also be contained within the donor DNA itself. The pooled gRNA library plasmids are then used to transform target cells of interest, resulting in a corresponding library of distinct RNA-guided DNA insertion events across the cell population. In an optional next step, the cell population can be subjected to a selection step to enrich for the desired phenotypes procured by the insertion library. Finally, sequencing or next-generation sequencing (NGS) is used to identify the gRNA from the pooled library that caused the desired phenotype. In one embodiment of this process, the pooled gRNA library is first generated in plasmid DNA and then converted to a lentiviral gRNA library for experiments in eukaryotic cells. Cells from the pooled library experiment (B) contain a CRISPR array containing one of the members of the gRNA library and a donor DNA insertion proximal to the target site complementary to the gRNA. The gRNA locus, the insertion site, or both can be sequenced. (C) is a schematic diagram of one embodiment in which a CRISPR array encoding the gRNA is inserted directly into the donor DNA cargo. In another embodiment, the pooled gRNA library is cloned into the donor DNA cargo.In this embodiment, RNA-guided DNA integration results in the preservation of the gRNA within the donor DNA, so that information about the gRNA that drove DNA insertion into that particular genomic region is preserved within the donor element itself. Next-generation sequencing (NGS) analysis of the insertion site is then used to extract both the integration site and gRNA information, for example, by transposon insertion sequencing. [Figure 104A] This shows that the gRNA encoded by the donor DNA directs efficient RNA-guided DNA integration. Figure 104A is a schematic diagram of an engineered two-plasmid system for RNA-guided DNA integration. The effector plasmid (pCQT; pSL1022, exemplified by SEQ ID NO: 855) encodes all protein components (including TniQ-Cas8-Cas7-Cas6-TnsA-TnsB-TnsC in this embodiment) in addition to the gRNA (via the CRISPR array). The donor plasmid (pDonor; pSL0527, exemplified by SEQ ID NO: 7) contains donor DNA flanked by the left and right ends of the transposon. [Figure 104B] This shows that the gRNA encoded by the donor DNA directs efficient RNA-guided DNA integration. Figure 104B is a schematic diagram of a modified, engineered two-plasmid system for RNA-guided DNA integration. The effector plasmid (pQT; exemplified by pSL1466, SEQ ID NO: 2001) encodes all protein components (including TniQ-Cas8-Cas7-Cas6-TnsA-TnsB-TnsC in this embodiment). The Donor_CRISPR plasmid (pDonor_CRISPR-R; exemplified by pSL1805, SEQ ID NO: 2002) contains donor DNA flanked by the left and right ends of the transposon, and the CRISPR array encoding the gRNA is contained within the cargo donor DNA itself near the right end of the transposon. In another embodiment, the pDonor_CRISPR plasmid further removes the lac operator sequence downstream of the T7 promoter (exemplified by pSL1766, SEQ ID NO: 2005). [Figure 104C]This demonstrates that the gRNA encoded by the donor DNA directs efficient RNA-guided DNA integration. Figure 104C is a schematic diagram of a modified version of pDonor_CRISPR, in which a CRISPR array is contained either at the left transposon end (pSL1632, SEQ ID NO: 2003) or near the center of the cargo (pSL1631, SEQ ID NO: 2004). [Figure 104D] This demonstrates that the gRNA encoded by the donor DNA directs efficient RNA-guided DNA integration. Figure 104D is a graph of RNA-guided DNA integration activity in E. coli BL21(DE3) cells using a gRNA targeting lacZ. The identities of the two plasmids used in each experiment are indicated below the bar graph. Integration efficiency was quantified by qPCR using cell lysates after overnight culture on solid LB agar medium. The efficiency of the pDonor_CRISPR-R plasmid is much higher when the CRISPR array is contained near the right transposon end. DETAILED DESCRIPTION OF THE INVENTION
[0105] In certain embodiments, the systems and methods of the present invention use a Tn7-like transposon that encodes a CRISPR-Cas system for programmable RNA-guided DNA integration. Specifically, the CRISPR-Cas machinery directs Tn7 transposon-associated proteins to integrate DNA downstream of a target site (e.g., a genomic target site) recognized by a guide RNA (gRNA).
[0106] 1. RNA-guided DNA integration The RNA-guided transposase mechanism for gene integration does not proceed through a double-strand break (DSB) intermediate, and therefore does not result in non-homologous end joining (NHEJ)-mediated insertion or deletion. Rather, targeting DNA leads to direct integration via a concerted transesterification reaction, without any alternative off-pathway. Because targeting relies on gRNA, the method and system of the present invention eliminates the need to redesign homologous arms for each new target site.
[0107] For therapeutic purposes, gRNAs can be designed to target specific genes or chromosomal regions (e.g., genes or chromosomal regions associated with a disease, disorder, or condition).
[0108] The systems and methods of the present invention may produce any desired effect. In one embodiment, the systems and methods of the present invention may produce a decrease in transcription of a target gene.
[0109] The systems and methods of the present invention can target any target site or insert donor DNA into any site within DNA, for example, within a coding or non-coding region, within or adjacent to a gene (e.g., a leader sequence, trailer sequence, or intron), or within a non-transcribed region, either upstream or downstream of a coding region. The target site or target sequence can comprise any polynucleotide (e.g., a DNA or RNA polynucleotide).
[0110] The RNA-guided DNA integration systems and methods of the invention allow for DNA integration in a variety of cell types, including post-mitotic and non-dividing cells (e.g., neurons and terminally differentiated cells). Accordingly, cells comprising the RNA-guided DNA integration systems of the invention are also provided.
[0111] The systems and methods of the present invention can be derived from bacterial or archaeal transposons that possess CRISPR-Cas systems (e.g., Tn7-like transposons). In one embodiment, the Tn7-like transposon system is derived from Vibrio cholerae Tn6677. This system can include gain-of-function Tn7 mutants (Lu et al., EMBO 19(13):3446-3457 (2000); U.S. Patent Publication No. 20020188105) and replicative Tn7 transposition mutants (May et al., Science 272:401-404 (1996)). Tn7-like transposons include, but are not limited to, Tn6677, Tn5090 / Tn5053, Tn6230, and Tn6022 transposons from Vibrio cholerae. See Peters et al., Recruitment of CRISPR-Cas systems by Tn7-like transposons, Proc Natl Acad Sci USA 114, E7358-E7366 (2017); Peters, JETn7. Microbiol Spectr. 2 (2014).
[0112] Tn7-like transposons can encode various types of CRISPR-Cas systems, including type I CRISPR-Cas systems (e.g., subtype IB, IF (including IF variants)) and type V CRISPR-Cas systems (e.g., V-U5).
[0113] In certain embodiments, the systems and methods of the present invention may include a type I CRISPR-Cas system. Type I systems may include a multi-subunit effector complex, such as a Cascade or Csy complex. In one embodiment, the Cascade complex is derived from the Vibrio cholerae Tn7 transposon, which includes an IF-type cascade and the TniQ protein. TniQ can bridge the CRISPR-Cas machinery with the Tn7-associated machinery for DNA integration. The systems of the present invention may be nuclease-deficient. In one embodiment, the Tn7-associated IF-type system may lack the Cas3 nuclease.
[0114] The Cascade complex in the canonical IF CRISPR-Cas system is encoded by four genes designated cas8 (or csy1), cas5 (or csy2), cas7 (or csy3), and cas6 (or csy4), which may be further classified with subtype-specific modifiers such as cas8f, cas5f, cas7f, and cas6f.
[0115] In one embodiment, a Tn7-like transposon comprises an IF-type variant CRISPR-Cas system, the genes of which encode the Cascade complex. Tn7-like transposons contain the tnsA-tnsB-tnSc operon, while a tnsD homolog known as tniQ is encoded within an operon encoding the Cas8 / Cas5 fusion-Cas7-Cas6 proteins that collectively form the RNA-guided TniQ-Cascade complex. The protein products of TnsA and TnsB mediate transposon excision, and TnsB mediates transposon integration into target DNA.
[0116] Tn7-like transposons can contain the transposases TnsA and TnsB. TnsA and TnsB can form heteromeric transposases. TnsB is a DDE-type transposase that catalyzes a concerted cleavage and religation reaction, ligating the 3'-hydroxyl group at the donor end to the 5'-phosphate group at the insertion site of the target DNA. TnsA is structurally similar to a restriction enzyme and nicks the opposite strand of the donor DNA molecule. The accessory protein TnsC can regulate the activity of heteromeric TnsAB transposases. TnsC can activate transposition when complexed with target DNA and the target selection protein TnsD or TnsE. TnsC variants can promote transposition in the absence of TnsD or TnsE. In certain embodiments, wild-type or variants of TnsA, TnsB, and / or TnsC (including variants with deletions, insertions, or amino acid substitutions compared to the wild-type proteins) can be used in the systems and methods of the present invention. The systems of the invention may include one or more of the following variants: TnsA S69N, TnsA E73K, TnsA A65V, TnsA E185K, TnsA Q261Z, TnsA G239S, TnsA G239D, TnsA Q261Z, TnsB M366I, TnsB A325T, and TnsB A325V (see Lu et al., (EMBO J. 9(3):3446-57, 2000)).
[0117] In one embodiment, the engineered transposon-encoded CRISPR-Cas system of the present invention is derived from V. cholerae HE-45 (designated Tn6677 and registered with the Transposon Registry). See Roberts et al., Revised terminology for transposable genetic elements, Plasmid 60, 167-173 (2008). Tn6677 refers to the native V. cholerae transposon sequence, and the compacted transposon construct containing the transposon ends and artificial cargo is more commonly referred to as mini-Tn6677 or mini-transposon (mini-Tn). The CRISPR-Cas system found in Tn6677 is an IF variant system, and the cascade operon contains a cas8-cas5 fusion gene (also referred to herein as cas8), cas7, and cas6, along with the upstream tniQ gene. Trans expression of transposon- and CRISPR-associated machinery serves to transfer mini-Tn6677 from the vector containing the donor DNA to the DNA integration site.
[0118] In one embodiment, the systems and methods of the invention comprise an engineered V. cholerae Tn7 transposon comprising TnsA, TnsB, TnsC, TniQ, a Cas8 / Cas5 fusion, Cas7, Cas6, and at least one gRNA.
[0119] In certain embodiments, the systems and methods of the present invention may include a V-type CRISPR-Cas system. V-type systems belong to class 2 CRISPR-Cas systems and are characterized by a single-protein effector complex programmed with a gRNA. In one embodiment, the Tn7-like transposon of the present invention includes a V-U5-type system encoding an enzyme such as C2c5 (S. Shmakov et al., Nat Rev Microbiol. 15, 169-182 (2017)). The systems of the present invention may be nuclease-deficient. In one embodiment, the systems of the present invention lack TnsA (the tnsA gene is missing).
[0120] C2c5 can be derived from Geminocystis species NIES-3709 (NCBI accession ID: WP_066116114.1). Transposon-associated type V CRISPR-Cas systems can be derived from Anabaena variabilis ATCC 29413 (or Trichormus variabilis ATCC 29413 (see GenBank CP000117.1)), Cyanobacterium aponinum IPPAS B-1202, filamentous cyanobacterium CCP2, Nostoc punctiforme PCC73102, and Scytonema hofmannii PCC7110.
[0121] In one embodiment, the systems and methods of the invention comprise an engineered Tn7-like transposon encoding a V-U5 type CRISPR-Cas system comprising TnsB, TnsC, TniQ, C2c5, and at least one gRNA.
[0122] The term "transposon" encompasses a DNA segment having cis-acting sites (which may contain heterologous DNA sequences) and a gene encoding a trans-acting protein that acts at those cis-acting sites to mobilize the DNA segment defined by those sites, regardless of how they are organized within the DNA. Transposons of the present invention (e.g., Tn7-like transposons) also encode a CRISPR-Cas system. The entire transposon is not required to practice the methods of the present invention. Therefore, as used herein, the terms "transposon derivative," "transposable element," or "insertion element" can also refer to DNA that minimally contains cis-acting sites at which trans-acting proteins act to mobilize the segment defined by those sites. It should also be understood that those sites may contain heterologous DNA. Proteins can be provided in the form of nucleic acids (protein-encoding DNA or RNA) or proteins (e.g., purified proteins).
[0123] As used herein, the term "Tn7 transposon" refers to the prokaryotic transposable element Tn7, and modified forms thereof or transposons that share homology with the Tn7 transposon ("Tn7-like transposons"). Tn7 has been most commonly studied in Escherichia coli. "Tn7 transposon" can encompass DNA forms that do not explicitly contain a Tn7 gene but can be made to undergo transposition through the use of the Tn7 gene products TnsA and TnsB (which together form the Tn7 transposase), or modifications thereof. Such DNA is bounded by 5' and 3' DNA sequences recognizable by the transposase, which can function as transposon end sequences. Examples of Tn7 transposon end sequences can be found in Arciszewska et al. (1991) J Biol Chem 266:21736-44 (PMID: 1657979); Tang et al. (1995) Gene 162:41-6 (PMID: 7557414); Tang et al. (1991) Nucleic Acids Res 19:3395-402 (PMID: 1648205); Biery et al. (2000) Nucleic Acids Res 28:1067-77 (PMID: 10666445); Craig (1995) Cur Top Microbiol Immunol 204:27-48 (PMID: 8556868), and other published sources, which should allow transposition given the appropriate Tns proteins. Without wishing to be bound by any theory, it is believed that the ends of the transposon are opposed to the donor DNA by TnsA and TnsB, and that these two Tns proteins cooperate to carry out the cleavage and ligation reactions that underlie transposition.
[0124] The Tn7 transposon contains characteristic left and right transposon end sequences and encodes five tns genes, tnsA–E, which collectively encode heteromeric transposases, TnsA and TnsB, catalytic enzymes that excise the transposon donor through coordinated double-strand breaks. TnsB, a member of the retroviral integrase superfamily, catalyzes DNA integration, while TnsD and TnsE constitute mutually exclusive targeting factors that identify the DNA integration site. TnSc is an ATPase that communicates between TnsAB and TnsD or TnsE. TnsD mediates site-specific Tn7 transposition to a conserved Tn7 binding site (attTn7) downstream of the E. coli glmS gene, while TnsE mediates random transposition to lagging-strand templates during replication. In E. coli, site-specific transposition involves binding of attTn7 by TnsD and subsequent interaction with the TnSc regulatory protein, directly mobilizing the TnsA-TnsB-donor DNA. TnSc, TnsD, and TnsE interact with target DNA to regulate transposase activity through two distinct pathways. TnsABC+TnsD induces transposition at high frequencies to attTn7, a discontinuous site on the E. coli chromosome, and at low frequencies to other loosely related "pseudo-att" sites. The alternative combination, TnsABC+E, induces transposition at low frequencies to many unrelated, non-attTn7 sites in the chromosome, preferentially targeting conjugative plasmids. Thus, attTn7 and conjugative plasmids contain positive signals that mobilize the transposon to these target DNAs. Alternative target site selection mechanisms allow Tn7 to explore various potential target sites within the cell and select the one most likely to ensure its survival.
[0125] As used herein, the term "transposase" refers to an enzyme that catalyzes transposition.
[0126] As used herein, the term "transposition" refers to a complex gene rearrangement process that involves moving a DNA sequence from one location and inserting it into another location, for example, between a genome and a DNA construct (e.g., a plasmid, bacmid, cosmid, and viral vector).
[0127] The present disclosure provides an engineered transposon-encoded CRISPR-Cas system for RNA-guided DNA integration in cells, the system comprising: (i) at least one Cas protein; (ii) a guide RNA (gRNA); and (iii) a Tn7-like transposon system.
[0128] Also encompassed by the present disclosure are systems and methods for RNA-guided DNA integration in cells, the systems and methods including: (i) one or more vectors encoding an engineered CRISPR-Cas system, the one or more vectors comprising (a) at least one Cas protein and (b) a guide RNA (gRNA); and (ii) one or more vectors encoding a Tn7-like transposon system, the CRISPR-Cas system and the transposon system being on the same or different vector(s).
[0129] The present disclosure provides engineered transposon-encoded CRISPR-Cas systems and methods for RNA-guided DNA integration in cells, the systems and methods comprising: (i) at least one Cas protein; (ii) a guide RNA (gRNA); and (iii) an engineered transposon system.
[0130] The present disclosure also provides systems and methods for RNA-guided DNA integration in cells, the systems and methods including: (i) one or more vectors encoding an engineered CRISPR-Cas system, the one or more vectors comprising (a) at least one Cas protein and (b) a guide RNA (gRNA); and (ii) one or more vectors encoding an engineered transposon system, the CRISPR-Cas system and the transposon system being on the same or different vector(s).
[0131] The present disclosure provides a method for RNA-guided DNA integration in a cell, the method comprising introducing into an animal cell an engineered transposon-encoded CRISPR-Cas system comprising (i) at least one Cas protein, (ii) a guide RNA (gRNA) specific for a target site, (iii) an engineered transposon system, and (iv) donor DNA, wherein the transposon-encoded CRISPR-Cas system integrates the donor DNA proximal to the target site.
[0132] The systems and methods of the invention may include TnsD or TniQ. The systems of the invention may include TnsA, TnsB, and TnsC. The systems of the invention may include TnsB and TnsC.
[0133] The systems and methods of the invention may be derived from Class 1 CRISPR-Cas systems. The systems and methods of the invention may be derived from Class 2 CRISPR-Cas systems. The systems and methods of the invention may be derived from Type I CRISPR-Cas systems (e.g., subtype IB, IF (including IF variants)). The systems and methods of the invention may be derived from Type V CRISPR-Cas systems (e.g., V-U5). The systems and methods of the invention may be derived from Type II CRISPR-Cas systems (e.g., II-A).
[0134] The systems and methods of the invention can be nuclease-deficient. The systems and methods of the invention can include Cas6, Cas7, and Cas5, and Cas8, either separately or as fusion proteins. The systems and methods of the invention can include Cas9.
[0135] The systems and methods of the present invention may include a cascade complex. The systems of the present invention may include C2c5.
[0136] The transposon-encoded CRISPR-Cas system can integrate donor DNA into the cell's genome.
[0137] The systems and methods of the present invention may further include a donor DNA comprising a cargo nucleic acid flanked by transposon end sequences. The transposon end sequences at both ends may be the same or different. The transposon end sequences may be endogenous Tn7 transposon end sequences or may comprise deletions, substitutions, or insertions. The endogenous Tn7 transposon end sequences can be truncated. In some embodiments, the transposon end sequences comprise a deletion of about 40 base pairs (bp) compared to the endogenous Tn7 transposon end sequence. In some embodiments, the transposon end sequences comprise a deletion of about 100 base pairs compared to the endogenous Tn7 transposon end sequence. This deletion can take the form of a truncation at the distal end (relative to the cargo) of the transposon end sequence.
[0138] Integration can be about 40 bp to about 60 bp, about 46 bp to about 55 bp, about 47 bp to about 51 bp, about 48 bp to about 50 bp, about 43 bp to about 57 bp, about 45 bp to about 50 bp, about 48 bp, about 49 bp, or about 50 bp downstream (3') from the target site.
[0139] The target site may be flanked by protospacer adjacent motifs (PAMs).
[0140] The present disclosure provides systems and methods for the transient expression or stable integration of DNA or polynucleotide(s) encoding one or more components of the systems of the invention.
[0141] The systems and methods of the present invention may be specific for one target site, or may be specific for two, three, four, five, six, seven, eight, nine, ten or more target sites.
[0142] In certain embodiments, the systems and methods of the invention may operate via a cut-and-paste mechanism (e.g., an IF-type CRISPR-Cas system, such as a system derived from E. coli Tn7 or V. cholerae Tn6677). In certain embodiments, the systems and methods of the invention may operate via a copy-and-paste mechanism (or replicative transposition) (e.g., a V-type CRISPR-Cas system including C2c5 (Cas12k)).
[0143] The systems and methods of the present invention can operate via a cut-and-paste mechanism, in which donor DNA is completely excised from the donor site and inserted into the target location (Bainton et al., Cell, 1991; 65(5), pp. 805-816). TnsA and TnsB cleave both strands of the transposon DNA at both ends, leaving a clean, linear dsDNA fragment containing short, three-nucleotide 5'-overhangs at both ends (not shown). TnsB then attacks phosphodiester bonds on both strands of the target DNA using its free 3'-OH terminus as a nucleophile, resulting in a concerted transesterification reaction. After gap-filling, the transposition reaction is complete, and the integrated transposon is flanked on both ends by 5-bp target site overlaps (TSDs) as a result of the gap-filling reaction.
[0144] The systems and methods of the present invention can operate via a copy-and-paste mechanism, also known as replicative transfer. This occurs when the 5' end of the transposon donor DNA is not cleaved during the excision step, as occurs when the tnsA endonuclease gene is absent from the gene operon encoding the transposition protein. In this case, the 3'-OH end remains free and can participate in a staggered transesterification reaction with the target DNA catalyzed by TnsB, while the 5' end of the transposon remains covalently linked to the rest of the DNA within the donor DNA molecule, which can be a genome or a plasmid vector. This copy-and-paste reaction results in what is known as a Shapiro intermediate, in which the entire donor DNA (including the transposon sequence itself and adjacent sequences) is ligated to the cleaved target DNA. This intermediate only degrades during subsequent DNA replication, resulting in the so-called cointegrate product. This cointegrate contains two copies of the transposon itself, flanked on one side by the TSD. Importantly, the cointegrate also contains the entire donor DNA molecule and the entire target DNA molecule. Therefore, when a transposon is encoded on a plasmid vector, the entire vector is ligated to the target DNA during replication transfer. At some frequency, the cointegration product can be resolved into the product shown on the right by the action of a dedicated resolvase protein (e.g., the TniR protein in Tn5090 / Tn5053) or by endogenous homologous recombination due to the high homology between the two copies of the transposon itself in the cointegration product. Cointegration resolution results in a target DNA with a single transposon flanked by TSDs and a regenerated version of the donor DNA molecule.
[0145] In one embodiment, the systems and methods of the present invention comprise a Tn7 or Tn7-like transposon with a single point mutation in the TnsA active site (TnsA D114A). DNA cleavage can occur at the 3' end of each strand of the donor (May and Craig, Science, 1996;272(5260):401-4). Failure to completely excise the donor DNA can result in the system switching to a replicative copy-and-paste mechanism, resulting in a cointegration product that is eventually resolved by recombination, producing two identical copies of the cargo. In another embodiment, the systems of the present invention comprise a Tn7 or Tn7-like transposon with a single point mutation (D90A) in the V. cholerae TnsA protein (TnsA D90A). In yet another embodiment, the cargo contains a site-specific recombinase (e.g., Cre or CinH) along with its recognition sequence to increase the efficiency of recombination and resolution of the cointegration product. This recombinase-assisted strategy has been shown to be utilized for integrant resolution in naturally occurring replicating transposons such as Tn3 and Mu (Nicolas et al. Microbiology Spectrum. 2015;3(4)).
[0146] In some embodiments, the nucleic acids encoding the Cas proteins, Tns proteins, and gRNA are provided on the same nucleic acid (e.g., vector). In some embodiments, the nucleic acids encoding the Cas proteins, Tns proteins, and gRNA are provided on different nucleic acids (e.g., different vectors), e.g., on two, three, four, five, six, or more vectors. Alternatively, or additionally, the Cas proteins and / or Tns proteins may be provided or introduced into the cell in protein form.
[0147] In some embodiments, the nucleotide sequences encoding the Cas and / or Tns proteins may be codon-optimized for expression in a host cell, hi some embodiments, one or more of the Cas and / or Tns proteins are homologs or orthologs of the wild-type proteins.
[0148] In some embodiments, the nucleotide sequences encoding the Cas and / or Tns proteins are modified to alter the activity of the proteins. Alternatively or additionally, the Cas and / or Tns proteins may be fused to another protein or portion thereof. In some embodiments, the Cas and / or Tns proteins are fused to a fluorescent protein (e.g., GFP, RFP, mCherry, etc.). In some embodiments, the Cas and / or Tns proteins fused to a fluorescent protein are used to label and / or visualize a genomic locus or to identify cells expressing the protein.
[0149] In certain embodiments, the systems of the present invention comprise one or more vector DNAs or polynucleotides comprising one or more nucleotide sequences selected from SEQ ID NOs: 1-139 and equivalents thereof. In certain embodiments, the systems of the present invention comprise one or more vectors comprising one or more nucleotide sequences that are about 80% to about 100% identical to a nucleotide sequence selected from SEQ ID NOs: 1-139. The vectors may comprise a nucleotide sequence that is at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to any of the nucleotide sequences set forth in SEQ ID NOs: 1-139.
[0150] In certain embodiments, the systems and methods of the invention comprise one or more vectors, DNA, or polynucleotides comprising one or more nucleotide sequences selected from SEQ ID NO:140 (TnsA), SEQ ID NO:142 (TnsB), SEQ ID NO:144 (TnsC), SEQ ID NO:146 (TniQ), SEQ ID NO:148 (Cas8 / Cas5 fusion), SEQ ID NO:150 (Cas7), SEQ ID NO:152 (Cas6), and equivalents thereof. In certain embodiments, the systems of the invention comprise one or more vectors, DNA, or polynucleotides comprising one or more nucleotide sequences that are about 80% to about 100% identical to a nucleotide sequence selected from SEQ ID NO:140, SEQ ID NO:142, SEQ ID NO:144, SEQ ID NO:146, SEQ ID NO:148, SEQ ID NO:150, and SEQ ID NO:152. The vector may comprise a nucleotide sequence that is at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to any of the nucleotide sequences set forth in SEQ ID NO:140, SEQ ID NO:142, SEQ ID NO:144, SEQ ID NO:146, SEQ ID NO:148, SEQ ID NO:150, and SEQ ID NO:152.
[0151] In certain embodiments, the systems and methods of the invention comprise one or more proteins having one or more amino acid sequences selected from SEQ ID NO: 141 (TnsA), SEQ ID NO: 143 (TnsB), SEQ ID NO: 145 (TnsC), SEQ ID NO: 147 (TniQ), SEQ ID NO: 149 (Cas8 / Cas5 fusion), SEQ ID NO: 151 (Cas7), SEQ ID NO: 153 (Cas6), and equivalents thereof. In certain embodiments, the systems of the invention comprise one or more proteins comprising one or more amino acid sequences that are about 80% to about 100% identical to an amino acid sequence selected from SEQ ID NO: 141 (TnsA), SEQ ID NO: 143 (TnsB), SEQ ID NO: 145 (TnsC), SEQ ID NO: 147 (TniQ), SEQ ID NO: 149 (Cas8), SEQ ID NO: 151 (Cas7), and SEQ ID NO: 153 (Cas6). The protein may comprise an amino acid sequence that is at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to any of the amino acid sequences set forth in SEQ ID NO: 141 (TnsA), SEQ ID NO: 143 (TnsB), SEQ ID NO: 145 (TnsC), SEQ ID NO: 147 (TniQ), SEQ ID NO: 149 (Cas8), SEQ ID NO: 151 (Cas7), and SEQ ID NO: 153 (Cas6).
[0152] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding TnsA, wherein the nucleotide sequence is SEQ ID NO: 140 or an equivalent thereof. The nucleotide sequence encoding TnsA can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 140.
[0153] The amino acid sequence of TnsA can comprise the amino acid sequence set forth in SEQ ID NO: 141, or an equivalent thereof. The amino acid sequence of TnsA can comprise an amino acid sequence that is at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 141.
[0154] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding TnsB, wherein the nucleotide sequence is SEQ ID NO: 142 or an equivalent thereof. The nucleotide sequence encoding TnsB can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 142.
[0155] The amino acid sequence of TnsB can comprise SEQ ID NO: 143 or an equivalent thereof. The amino acid sequence of TnsB can comprise an amino acid sequence that is at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 143.
[0156] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding TnsC, wherein the nucleotide sequence is SEQ ID NO: 144 or an equivalent thereof. The nucleotide sequence encoding TnsC can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 144.
[0157] The amino acid sequence of TnsC can comprise SEQ ID NO: 145 or an equivalent thereof. The amino acid sequence of TnsC can comprise an amino acid sequence that is about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 145.
[0158] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding TniQ, wherein the nucleotide sequence is SEQ ID NO: 146 or an equivalent thereof. The nucleotide sequence encoding TniQ can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 146.
[0159] The amino acid sequence of TniQ can comprise SEQ ID NO: 147 or an equivalent thereof. The amino acid sequence of TniQ can comprise an amino acid sequence that is about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 147.
[0160] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding Cas8 (Cas5 / Cas8), wherein the nucleotide sequence is SEQ ID NO: 148 or an equivalent thereof. The nucleotide sequence encoding Cas8 (Cas5 / Cas8) can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 148.
[0161] The amino acid sequence of Cas8 (Cas5 / Cas8) can include SEQ ID NO: 149 or an equivalent thereof. The amino acid sequence of Cas8 (Cas5 / Cas8) can include an amino acid sequence that is about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 149.
[0162] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding Cas7, wherein the nucleotide sequence is SEQ ID NO: 150 or an equivalent thereof. The nucleotide sequence encoding Cas7 can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 150.
[0163] The amino acid sequence of Cas7 can comprise SEQ ID NO: 151 or an equivalent thereof. The amino acid sequence of Cas7 can comprise an amino acid sequence that is about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 151.
[0164] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding Cas6, wherein the nucleotide sequence is SEQ ID NO: 152 or an equivalent thereof. The nucleotide sequence encoding Cas6 can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 152.
[0165] The amino acid sequence of Cas6 can comprise SEQ ID NO: 153 or an equivalent thereof. The amino acid sequence of Cas6 can comprise an amino acid sequence that is about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in SEQ ID NO: 153.
[0166] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding TnsA, wherein the nucleotide sequence is selected from SEQ ID NOs: 768, 1777, 1786, 1795, 1804, 1813, 1822, 1831, 1909, 1925, 1941, 1957, or equivalents thereof. The nucleotide sequence encoding TnsA can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in any of SEQ ID NOs: 1768, 1777, 1786, 1795, 1804, 1813, 1822, 1831, 1909, 1925, 1941, and 1957.
[0167] The amino acid sequence of TnsA may comprise the amino acid sequence set forth in any one of SEQ ID NOs: 1714 to 1717, 1840, 1847, 1854, 1861, 1868, 1875, 1882, 1889, 1896, 1918, 1934, 1950, or 1966, or an equivalent thereof. The amino acid sequence of TnsA may be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, or at least about 100% identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1714 to 1717, 1840, 1847, 1854, 1861, 1868, 1875, 1882, 1889, 1896, 1918, 1934, 1950, or 1966. The amino acid sequence may be at least about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence of the target gene.
[0168] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding TnsB, wherein the nucleotide sequence is selected from SEQ ID NOs: 1769, 1778, 1787, 1796, 1805, 1814, 1823, 1832, 1910, 1926, 1942, 1958, or an equivalent thereof. The nucleotide sequence encoding TnsB may be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in any of SEQ ID NOs: 1769, 1778, 1787, 1796, 1805, 1814, 1823, 1832, 1910, 1926, 1942, and 1958.
[0169] The amino acid sequence of TnsB may comprise the amino acid sequence set forth in any of SEQ ID NOs: 1841, 1848, 1855, 1862, 1869, 1876, 1883, 1890, 1919, 1935, 1951, 1967, or an equivalent thereof. The amino acid sequence of TnsB may comprise an amino acid sequence that is about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in any of SEQ ID NOs: 1841, 1848, 1855, 1862, 1869, 1876, 1883, 1890, 1919, 1935, 1951, or 1967.
[0170] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding a TnsA / TnsB fusion, wherein the nucleotide sequence is selected from SEQ ID NO: 1973, 1987, or an equivalent thereof. The nucleotide sequence encoding the TnsA / TnsB fusion can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in either SEQ ID NO: 1973 or 1987.
[0171] The amino acid sequence of the TnsA / TnsB fusion comprises the amino acid sequence set forth in either of SEQ ID NOs: 1981, 1995, or an equivalent thereof. The amino acid sequence of the TnsA / TnsB fusion may comprise an amino acid sequence that is at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in either of SEQ ID NOs: 1981 and 1995.
[0172] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding TnsC, wherein the nucleotide sequence is selected from SEQ ID NOs: 1770, 1779, 1788, 1797, 1806, 1815, 1824, 1833, 1911, 1927, 1943, 1959, 1974, 1988, or equivalents thereof. The nucleotide sequence encoding TnsC can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in any of SEQ ID NOs: 1770, 1779, 1788, 1797, 1806, 1815, 1824, 1833, 1911, 1927, 1943, 1959, 1974, and 1988.
[0173] The amino acid sequence of TnsC may comprise the amino acid sequence set forth in any one of SEQ ID NOs: 1842, 1849, 1856, 1863, 1870, 1877, 1884, 1891, 1920, 1936, 1952, 1968, 1982, and 1996, or an equivalent thereof. The amino acid sequence of TnsC may be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about It may comprise an amino acid sequence that is 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical.
[0174] In one embodiment, the systems and methods of the present invention comprise a nucleotide sequence encoding TniQ, wherein the nucleotide sequence is selected from SEQ ID NOs: 1771, 1780, 1789, 1798, 1807, 1816, 1825, 1834, 1912, 1928, 1944, 1960, 1975, 1989, or equivalents thereof. The nucleotide sequence encoding TniQ can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in any of SEQ ID NOs: 1771, 1780, 1789, 1798, 1807, 1816, 1825, 1834, 1912, 1928, 1944, 1960, 1975, and 1989.
[0175] The amino acid sequence of TniQ may comprise the amino acid sequence set forth in any one of SEQ ID NOs: 1843, 1850, 1857, 1864, 1871, 1878, 1885, 1892, 1921, 1937, 1953, 1969, 1983, and 1997, or an equivalent thereof. The amino acid sequence of TniQ may be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about It may comprise an amino acid sequence that is 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical.
[0176] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding Cas7, wherein the nucleotide sequence is selected from SEQ ID NOs: 1773, 1782, 1791, 1800, 1809, 1818, 1827, 1836, 1914, 1930, 1946, 1962, 1977, 1998, or an equivalent thereof. The nucleotide sequence encoding Cas7 can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in any of SEQ ID NOs: 1773, 1782, 1791, 1800, 1809, 1818, 1827, 1836, 1914, 1930, 1946, 1962, 1977, and 1998.
[0177] The amino acid sequence of Cas7 may comprise any of the amino acid sequences set forth in SEQ ID NOs: 1845, 1852, 1854, 1866, 1873, 1880, 1887, 1899, 1923, 1939, 1955, 1971, 1958, and 1999, or equivalents thereof. The amino acid sequence of Cas7 may be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about It may comprise an amino acid sequence that is 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical.
[0178] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding Cas6, wherein the nucleotide sequence is selected from SEQ ID NOs: 1774, 1783, 1792, 1801, 1810, 1819, 1828, 1837, 1915, 1931, 1947, 1963, 1978, 1992, or an equivalent thereof. The nucleotide sequence encoding Cas6 can be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence set forth in any of SEQ ID NOs: 1774, 1783, 1792, 1801, 1810, 1819, 1828, 1837, 1915, 1931, 1947, 1963, 1978, and 1992.
[0179] The amino acid sequence of Cas6 may comprise any of the amino acid sequences set forth in SEQ ID NOs: 1846, 1853, 1860, 1867, 1874, 1881, 1888, 1895, 1924, 1940, 1956, 1972, 1986, and 2000, or equivalents thereof. The amino acid sequence of Cas6 may be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about It may comprise an amino acid sequence that is 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical.
[0180] In one embodiment, the systems and methods of the invention comprise a nucleotide sequence encoding a Cas8 / Cas5 fusion, wherein the nucleotide sequence is selected from SEQ ID NOs: 1772, 1781, 1790, 1799, 1808, 1817, 1826, 1835, 1913, 1929, 1945, 1961, 1976, 1990, or an equivalent thereof. The nucleotide sequence encoding Cas8 / Cas5 is about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, or at least or about 83% of the amino acid sequence set forth in any of SEQ ID NOs: 1772, 1781, 1790, 1799, 1808, 1817, 1826, 1835, 1913, 1929, 1945, 1961, 1976, and 1990. , at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical.
[0181] The amino acid sequence of Cas8 / Cas5 may comprise any of the amino acid sequences set forth in SEQ ID NOs: 1844, 1851, 1858, 1865, 1872, 1879, 1886, 1893, 1922, 1938, 1954, 1970, 1984, and 1998, or equivalents thereof. The amino acid sequence of Cas8 / Cas5 may be about 80% to about 100%, at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 100%, or at least about 100%, at least or ... The amino acid sequence may be at least about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, or about 100% identical to the amino acid sequence of the target gene.
[0182] The systems and methods of the invention can include (i) one or more vectors encoding an engineered CRISPR-Cas system, and (ii) one or more vectors encoding an engineered transposon system, where the CRISPR-Cas system and the transposon system are on the same vector or on at least two different vectors. In one embodiment, a first vector encodes TnsB, TnsC, and TniQ (e.g., pTnsBCQ), a second vector encodes C2c5 (e.g., pC2c5), and a third vector encodes donor DNA (e.g., pDonor).
[0183] Proteins of the present systems and methods include wild-type proteins as well as any substantially homologous proteins and variants of the wild-type protein. The term "variant" of a protein is intended to mean a protein derived from the native protein by deletion (truncating), addition, and / or substitution of one or more amino acids in the native protein. Such variants may result, for example, from genetic polymorphism or human manipulation. A variant of a native protein may be "substantially homologous" to the native protein when at least about 80%, at least about 90%, or at least about 95% of its amino acid sequence is identical to the amino acid sequence of the native protein.
[0184] The systems and methods of the present invention provide for the insertion of nucleic acids into any DNA segment in any organism. Additionally, the systems and methods of the present invention provide for the insertion into any synthetic DNA segment.
[0185] Also provided is a self-transferring nucleic acid comprising a mobile nucleic acid sequence encoding a transposon-encoded CRISPR-cas system, as described above, and a first transposon end sequence and a second transposon end sequence flanking the mobile nucleic acid sequence. The cargo nucleic acid of the transposon-encoded CRISPR-cas system may be flanked by the transposon end sequences. The self-transferring nucleic acid may be present in a vector. A "vector" or "expression vector" is a replicon (e.g., a plasmid, phage, virus, or cosmid) into which another DNA segment (e.g., an "insert") can be attached or incorporated to replicate the attached segment within a cell. The self-transferring nucleic acid may be present in the genomic DNA of a cell.
[0186] a. donor DNA The donor DNA can be part of a bacterial plasmid, bacteriophage, plant virus, retrovirus, DNA virus, self-replicating extrachromosomal DNA element, linear plasmid, mitochondrial or other organelle DNA, chromosomal DNA, etc. The donor DNA comprises a cargo nucleic acid sequence flanked by transposon end sequences.
[0187] The donor DNA can also be of any suitable length depending on the extension of the cargo nucleic acid, including, for example, about 50-100 bp (base pairs), about 100-1000 bp, at least or about 10 bp, at least or about 20 bp, at least or about 25 bp, at least or about 30 bp, at least or about 35 bp, at least or about 40 bp, at least or about 45 bp, at least or about 50 bp, at least or about 55 bp, at least or about 60 bp, at least or about 65 bp, at least or about 70 bp, at least or about 75 bp, at least or about 80 bp, at least or about 85 bp, at least or about 90 bp, at least or about 9 Lengths of 5 bp, at least or about 100 bp, at least or about 200 bp, at least or about 300 bp, at least or about 400 bp, at least or about 500 bp, at least or about 600 bp, at least or about 700 bp, at least or about 800 bp, at least or about 900 bp, at least or about 1 kb (kilobase pairs), at least or about 2 kb, at least or about 3 kb, at least or about 4 kb, at least or about 5 kb, at least or about 6 kb, at least or about 7 kb, at least or about 8 kb, at least or about 9 kb, at least or about 10 kb, or less than 10 kb, or longer. The donor DNA and cargo nucleic acid can be at least or about 10 kb, at least or about 50 kb, at least or about 100 kb, 20 kb to 60 kb, or 20 kb to 100 kb.
[0188] b.CRISPR The CRISPR-Cas system has been successfully used to edit the genomes of a variety of organisms, including but not limited to bacteria, humans, fruit flies, zebrafish, and plants. For example, Jiang et al.,Nature Biotechnology(2013)31(3):233;Qi et al,Cell(2013)5:1173;DiCarlo et al.,Nucleic Acids Res.(2013)7:4336;Hwang et al.,Nat.Biotechnol(2013),3:227);Gratz et al. al.,Genetics(2013)194:1029;Cong et al.,Science(2013)6121:819;Mali et al.,Science(2013)6121:823;Cho et al.Nat.Biotechnol(2013)3:230;and Jiang et al.,Nucleic Acids See Research(2013)41(20):el88.
[0189] The systems of the invention may include Cas6, Cas7, Cas5, and Cas8. In some embodiments, Cas5 and Cas8 are combined as a functional fusion protein. The systems of the invention may include Cas9.
[0190] The systems of the invention may be derived from a Class 1 CRISPR-Cas system. The systems of the invention may be derived from a Class 2 CRISPR-Cas system. The systems of the invention may be derived from a Type I CRISPR-Cas system. The systems of the invention may be derived from a Type II CRISPR-Cas system. The systems of the invention may be derived from a Type V CRISPR-Cas system.
[0191] The system of the present invention may comprise a cascade complex. The system of the present invention may comprise C2c5.
[0192] c.gRNA The gRNA can be a crRNA / tracrRNA (or single guide RNA, sgRNA).
[0193] The terms "gRNA," "guide RNA," and "CRISPR guide sequence" are used interchangeably throughout and refer to a nucleic acid containing a sequence that determines the binding specificity of a CRISPR-Cas system. The gRNA hybridizes to a target nucleic acid sequence (e.g., a genome) in a host cell (either partially or fully complementary). The gRNA, or the portion thereof that hybridizes to the target nucleic acid (target site), can be 15-25 nucleotides, 18-22 nucleotides, or 19-21 nucleotides in length. In some embodiments, the gRNA sequence that hybridizes to the target nucleic acid is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. In some embodiments, the gRNA sequence that hybridizes to the target nucleic acid can be 10-30 nucleotides or 15-25 nucleotides in length. The gRNA or sgRNA used in this disclosure can be about 5-100 nucleotides in length or more (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 1 , 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides or more). In one embodiment, the gRNA or sgRNA(s) can be about 15 to about 30 nucleotides in length (e.g., about 15-29, 15-26, 15-25, 16-30, 16-29, 16-26, 16-25; or about 18-30, 18-29, 18-26, or 18-25 nucleotides in length).
[0194] Many computational tools have been developed to facilitate gRNA design (see Prykhozhij et al. (PLoS ONE, 10(3):(2015)); Zhu et al. (PLoS ONE, 9(9)(2014)); Xiao et al. (Bioinformatics. Jan 21(2014)); Heigwer et al. (Nat Methods, 11(2):122-123(2014)). Methods and tools for guide RNA design are discussed by Zhu (Frontiers in Biology, 10(4)pp289-296(2015)), which is incorporated herein by reference. In addition, there are many publicly available software tools that can be used to facilitate the design of sgRNA(s), including, but not limited to, Genscript Interactive CRISPR gRNA Design Tool, WU-CRISPR, and Broad Institute GPP sgRNA Design Tool. Such tools include the Gene Designer. Pre-designed gRNA sequences targeting many genes and locations within the genomes of many species (human, mouse, rat, zebrafish, C. elegans) are also publicly available, including but not limited to IDT DNA pre-designed Alt-R CRISPR-Cas9 guide RNAs, Addgene validated gRNA target sequences, and the GenScript genome-wide gRNA database.
[0195] In addition to a sequence that binds to a target nucleic acid, in some embodiments, the gRNA may also include a scaffold sequence (e.g., a tracrRNA). In some embodiments, such a chimeric gRNA may be referred to as a single guide RNA (sgRNA). Exemplary scaffold sequences will be apparent to those skilled in the art and can be found, for example, in Jinek, et al. Science (2012) 337(6096):816-821 and Ran, et al. Nature Protocols (2013) 8:2281-2308.
[0196] In some embodiments, the gRNA sequence does not include the scaffold sequence, and the scaffold sequence is expressed as a separate transcript. In such embodiments, the gRNA sequence further includes an additional sequence that is complementary to a portion of the scaffold sequence and functions to bind (hybridize) to the scaffold sequence.
[0197] In some embodiments, the gRNA sequence is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to the target nucleic acid. In some embodiments, the gRNA sequence is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to the 3' end of the target nucleic acid (e.g., the last 5, 6, 7, 8, 9, or 10 nucleotides of the 3' end of the target nucleic acid).
[0198] The gRNA may be a non-naturally occurring gRNA.
[0199] The target nucleic acid may be adjacent to a protospacer adjacent motif (PAM). The PAM site is a nucleotide sequence located proximal to the target sequence. For example, the PAM may be the DNA sequence immediately following the DNA sequence targeted by the CRISPR / Cas system.
[0200] The target sequence may or may not be adjacent to a protospacer adjacent motif (PAM) sequence. In certain embodiments, the nucleic acid-guided nuclease can cleave the target sequence only if the appropriate PAM is present (see, e.g., Doudna et al., Science, 2014, 346(6213):1258096, incorporated herein by reference). The PAM may be 5' or 3' to the target sequence. The PAM may be upstream or downstream from the target sequence. In one embodiment, the target sequence is immediately adjacent to the PAM sequence at its 3' end. The PAM may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides in length. In some embodiments, the PAM is 2-6 nucleotides in length. The target sequence may or may not be adjacent to the PAM sequence (e.g., a PAM sequence located immediately 3' to the target sequence) (e.g., in the case of type I CRISPR / Cas systems and type II CRISPR / Cas systems). In some embodiments, for example, in Type I systems, the PAM is on the alternative side (5' end) of the protospacer. Makarova et al. describe the names of all classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology 13:722-736 (2015)). Guide structures and PAMs are described by R. Barrangou (Genome Biol. 16:247 (2015)).
[0201] Non-limiting examples of PAM sequences include CC, CA, AG, GT, TA, AC, CA, GC, CG, GG, CT, TG, GA, AGG, TGG, T-rich PAMs (such as TTT, TTG, TTC, TTTT (SEQ ID NO: 385)), NGG, NGA, NAG, NGGNG, and NNAGAAW (W=A or T, SEQ ID NO: 912), NNNNGATT (SEQ ID NO: 913), NAAR (R=A or G), NNGRR (R=A or G), NNAGAA (SEQ ID NO: 914), and NAAAAC (SEQ ID NO: 915) (where "N" is any nucleotide).
[0202] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds (multiple bonds) with another nucleic acid sequence, either through traditional Watson-Crick or other non-traditional methods. The percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence. Perfect complementarity is not necessary, as long as there is sufficient complementarity to cause hybridization. Mismatches may exist distal to the PAM.
[0203] D transposon Any Tn7 transposon that encodes a CRISPR-Cas system can be used in the methods and systems of the present invention.
[0204] For example, type I Cascade complexes can be used in the methods and systems of the present invention. Type I CRISPR-Cas systems encode a multi-subunit protein-RNA complex called Cascade. Cascade utilizes crRNA (or guide RNA) to target double-stranded DNA during immune responses. Cascade itself does not have nuclease activity; instead, degradation of targeted DNA is mediated by a trans-acting nuclease known as Cas3. Interestingly, the IF and IB systems found within the Tn7 transposon consistently lack the Cas3 gene, suggesting that these systems no longer possess any DNA degradation capabilities and have been reduced to RNA-guided DNA-binding complexes. In additi...
Claims
1. 1. A system for RNA-guided DNA integration, comprising: containing one or more vectors heterologous to Vibrio cholerae, The vector a) an engineered Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system, comprising Cas5, Cas6, Cas7, and Cas8; and b) an engineered transposon 7 (Tn7)-like transposon system comprising i) TnsA, ii) TnsB, iii) TnsC, and iv) TnsQ; It is coded as The engineered Tn7-like transposon system is derived from Vibrio cholerae Tn6677, and a donor DNA to be integrated, the donor DNA comprising a cargo nucleic acid sequence, a first transposon end sequence, and a second transposon end sequence. system.
2. The CRISPR-Cas system is an IF type CRISPR-Cas system or an IF type variant CRISPR-Cas system in which Cas8 and Cas5 form a Cas8-Cas5 fusion product; The system of claim 1 .
3. Further comprising a guide RNA (gRNA) specific for the target site; The system of claim 1 .
4. The cargo nucleic acid sequence is sandwiched between the first transposon end sequence and the second transposon end sequence; each of the first transposon end sequence and the second transposon end sequence comprises at least one TnsB binding site; The system of claim 1 .
5. the first transposon end sequence and the second transposon end sequence are Tn7 transposon end sequences; The system of claim 4.
6. The CRISPR-Cas system and the Tn7-like transposon system are on the same vector. The system of claim 1 .
7. The engineered CRISPR-Cas system is nuclease-deficient. The system of claim 1 .
8. 1. A method for RNA-guided DNA integration, comprising: introducing into a cell in vitro or ex vivo: i) an engineered CRISPR-Cas system, and / or one or more vectors encoding said engineered CRISPR-Cas system; ii) an engineered transposon system, and / or one or more vectors encoding said engineered transposon system; and iii) a donor sequence comprising a cargo nucleic acid sequence and a first transposon end sequence and a second transposon end sequence; When one or more vectors are used, the CRISPR-Cas system and the transposon system are on the same or different vector(s); the cell comprises a nucleic acid sequence having a target site; The CRISPR-Cas system includes (a) at least one Cas protein selected from Cas5, Cas6, Cas7, and Cas8, and (b) a guide RNA (gRNA); The CRISPR-Cas system binds to a target site, The transposon system is derived from a transposon 7 (Tn7)-like transposon system and integrates the donor sequence downstream of the target site. method.
9. The Tn7-like transposon system is derived from Vibrio cholerae. The method of claim 8.
10. The at least one Cas protein is derived from a type I CRISPR-Cas system. The method of claim 8.
11. The at least one Cas protein comprises Cas5, Cas6, Cas7, and Cas8. The method of claim 8.
12. The type I CRISPR-Cas system is type I-B or type IF, The method of claim 10.
13. The type I CRISPR-Cas system is an IF type variant in which Cas8 and Cas5 form a Cas8-Cas5 fusion. The method of claim 12.
14. The at least one Cas protein of the CRISPR-Cas system is derived from a type V CRISPR-Cas system. The method of claim 8.
15. wherein the at least one Cas protein is C2c5.
15. The method of claim 14.
16. the engineered CRISPR-Cas system is derived from a Type I CRISPR-Cas system; The system further comprises a second engineered CRISPR-Cas system and a second engineered transposon system derived from a V-type CRISPR-Cas system. The method of claim 8.
17. The transposon system includes TnsA, TnsB, and TnsC. The method of claim 8.
18. The transposon system comprises i) TnsA, TnsB, and TnsC, and ii) TnsD and / or TniQ. The method of claim 8.
Citation Information
Patent Citations
delivery vehicle
JP2018522935A
Altering microbial populations & modifying microbiota
US20160333348A1
Polynucleotide enrichment using crispr-CAS systems
WO2016014409A1
CAS 9 retroviral integrase and CAS 9 recombinase systems for targeted incorporation of a DNA sequence into a genome of a cell or organism
WO2016161207A1
Novel crispr-associated transposases and uses thereof
WO2017117395A1